Method for determining a property of a metallurgical product and associated electronic device
By training individual models with shared lengthscales and applying them to a global model, the method addresses data scarcity and line disparities, improving predictive accuracy and efficiency in metallurgical production models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2026-03-26
AI Technical Summary
Existing metallurgical production models struggle with training data scarcity and uneven distribution, leading to overfitting and sensitivity to disparities between production lines, which affects predictive accuracy and computational efficiency.
Training individual Gaussian Process or kernel-based regression models with shared lengthscales for different production lines, followed by a global model with fixed shared lengthscales, to create a robust multi-line model that supplements data gaps and reduces sensitivity to line-specific disparities.
The method enhances predictive accuracy and computational efficiency by minimizing overfitting and leveraging diverse data from multiple lines, allowing accurate predictions even when data for specific chemical compositions or process settings are lacking.
Smart Images

Figure IB2024059231_26032026_PF_FP_ABST
Abstract
Description
Method for determining a property of a metallurgical product and associated electronic device
[0001] The technical field is that of metal making, for instance steelmaking. It concerns more particularly determining a product property resulting from processing the product in a metallurgical production line, using a product processing model, for instance for predicting this property or for monitoring such a process.
[0002] Being able to determine a property of a metallurgical product, that results from applying processing operations to the product in a metallurgical production line, using a numerical model, is very useful and has become a key challenge for the metal making industry.
[0003] Indeed, having an accurate model, modelling the effect on the product of such processing operations, allows for implementing efficient predictive control methods, for controlling on-line the processing operations (based on predicted properties for the product, expected at the end of the processing). Such a predictive control enables to reach high levels of quality and conformity and enables to reduce waste and scrap production or product downgrading.
[0004] Having such a model also allows for non-direct characterization of final properties of the product. In particular, it allows determining the product final properties (eg: mechanical properties), after its processing, based on its initial characteristics and based on the process parameters that were actually employed during the processing. This enables for instance to specify mechanical properties of a final product (for instance to specify these properties to a client), using such a model, based on the process parameters employed, instead of making direct tensile strength tests using physical samples taken from the product. Such a postcharacterization is more and more employed. Corresponding requirements in terms of accuracy and reproducibility are specified for instance in the Euronorm EN 10373.
[0005] Having such a modelling capacity also allows to plan adaptations of operation conditions, or of chemical compositions of the metal, to be made as a consequence of a change of the supply chain upstream the processing operations, for instance as a consequence of switching from an ore-based primary production of the metal to a recycled- material based primary production for which impurity or residual contents in the metal will be higher. This is the case in particular when switching from products coming from a Blast Furnace source to products coming from an Electric Arc Furnace source where steel scraps is recycled, which results in products whose residual contents are higher and for which subsequent processing operations have to be adapted to cope with these higher residualcontents (for instance a higher copper content). Being able to determine the adaptations to be made to the operation conditions, or to the chemical compositions of the metal, as a consequence of a change of the primary source of liquid metal, thanks to such a model, is thus very useful in the context of the decarbonization of mass production of metals, in part the mass production of steel.
[0006] Such a model is also useful for developing new metal grades having desired properties.
[0007] A metallurgical product processing model, as above mentioned, could be a physicsbased model, based on metallurgical transformations within the material. It could also be a data-based model, obtained using a machine-learning technique, more particularly a regression. Anyhow, in practice, in an industrial context in which many variables influence the resulting property, training well a data-based model is often difficult, due to a possible lack of training data (or to an uneven distribution of the training data) and to risks of overfitting.
[0008] In this context, amethod for training a global model configured for determining a property of a metallurgical product output by a metallurgical production line is provided, the method comprising:- training individual models respectively associated to different production lines, the individual models being Gaussian Process or kernel-based regression models, based on a same kernel function and being trained with the constraint that the values of the hyperparameters (e.g.: lengthscales) parametrizing the kernel function are identical for the different individual models, the training comprising determining optimized values for the hyperparameters;- training the global model using production data coming from the different production lines, the global model being a Gaussian Process or a kernel-based regression model based on the same kernel function as the individual models, the values for the hyperparameters remaining equal to the optimized values previously determined, when training the global model.
[0009] In particular, a method according to claim 1 is provided.
[0010] Using sets of production data coming from different production lines, usually installed in different plants, enables to gather data that are more diverse, varied than when using data sets coming from a single line and thus allow to obtain a potentially more robust global model.
[0011] Yet, using data coming from different production lines requires particular care, in the verification of calibrations, and standardization of data coming from different lines, and in thechoice of the type of the variables (having a significant influence on the resulting property, and available on the different lines).
[0012] Even with careful calibration checks and measurements normalization, with a global model, trained with the global production dataset (which gathers data from all the production lines), there is a significant risk for the global model to be sensitive to disparities between lines, constituting artifacts (measurement disparities, for example), these disparities being captured by the global model during training (as if they were significant variations; causing a kind of overfitting).
[0013] The instant technology makes it possible, among other things, to avoid this undesirable effect. Indeed, the lengthscales (more precisely: at least the “shared lengthscales”, and possibly all the lengthscales), to be used for the global model, are determined when the individual models are trained using their own, individual set of production data, and are therefore less sensitive to the spurious variations or disparities between different production lines. All the more so as the lengthscales (at least the “shared lengthscales”) are required to have the same, or essentially the same values (i.e. the same values to within + / - 10%), for the various individual models. And so, with this technique, the lengthscales values, and therefore the global model are more robust than if training the global model directly, based on the global production dataset, with no other precaution.
[0014] This constraint, interesting in terms of model training, is also well-grounded on a technical, scientific level. Indeed, it is expected that for similar transformations, the chemical elements would have the same effect. And so, when it comes to the chemical composition for example, one expects the lengthscale value associated with a chemical element (i.e.: the typical scale, for a variation in that element content, that leads to a significant variation in the product property) not to depend on production line specificities (but, on the contrary, to depend on aspects common to the different lines, as related to the relationship between the chemistry and the material transformations). A same lengthscale value for the different production lines is thus likely, from a physical point of view. The same applies to the main process parameters (such as an annealing or a hot-coiling temperature).
[0015] The instant technology is also beneficial in terms of computational demand, in particular in terms of RAM amount required to execute the method. Indeed, for at least some of the lengthscales (namely for the shared lengthscales), possibly for all of them, the optimization of the lengthscales values is achieved for the individual models. And, depending on the detailed implementation, it may be less demanding (in terms of RAM, and regarding the number of operations required) to achieve this optimization for J individual models than for a global model based on the J individual production data sets groups all together.
[0016] The individual models may in particular be Gaussian Process regression models. The optimization of the values of the shared lengthscales is then achieved by maximizing alikelihood of obtaining the values acquired for said property of the products, given the chemical composition CC of the products and the process parameters employed for processing the product.
[0017] Gaussian Process regression, in spite of its formal and computation complexity, has the advantage of providing uncertainties (variances), in addition to mean expected values, which is useful as it allows for instance evaluating a success probability for a given manufacturing route or other risk assessments. And it allows for taken advantage of possible tolerances regarding final product properties. In addition, Gaussian Process regression results can be considered as an optimal estimate (from a fully Bayesian point of view).
[0018] Yet, a kernel-based regression model different from Gaussian Process regression model, could be used instead (for instance a kernel ridge regression, or the training of a kernel-based support vector machine). Indeed, the technique of training individual models, but with shared parameters (for instance shared lengsthscales), and then training a global model with these fixed, predetermined parameters can be applied fruitfully in the same way to kernel-based regression models others than Gaussian Process ones.
[0019] The production data employed for training the models can be such that:- a first set among the sets of production data comprises production data for products whose chemical composition CC belongs to a given category of chemical compositions,- while a second set among of the sets of production data comprises no or almost no production data for products whose chemical composition CC belongs to said category of chemical compositions.
[0020] By almost no, it is meant for instance that the number of samples in the second set, with a chemical composition belonging to this specific category of chemical compositions, is less than 5% or even less than 1% of the total number of samples in the second data set.
[0021] The first set mentioned above (corresponding to production data for a first line) usefully supplements the second set (corresponding to production data for a second line), for which data corresponding to said category of chemical compositions are missing. Using the multi-line, global model, trained as above described, enables then to predict what would be obtained when processing, on the second line, a product having a chemical composition belonging to said category, while there is no (or almost no) past production data for the second line, for such chemical compositions.
[0022] In other words, an abstraction is made from each line specific details, to predict what would be obtained in the second line. This achieves better predictive accuracy than working only with the information available from the second line.
[0023] Test results presented further below in the description show that this technique actually enables to improve the prediction accuracy for the second line, for which data for that category of chemical compositions are missing, for instance.
[0024] Said category of chemical compositions may be defined by intervals for the contents of the different alloying elements considered. In practice, it the category in question may be a category grouping chemical compositions:- corresponding to one given type of steel (eg.: LC, DP, HSLA, IPS, TRIP), or to a given grade (normalized grade) of steel,- and corresponding to products made of a steel coming specifically from an Electric Arc Furnace (EAF) and produced by recycling steel scraps, or, on the contrary, corresponding to products made of a steel coming specifically from a Blast Furnace.
[0025] Indeed, a steel of the HSLA type for instance (Hight Strength Low Alloy), produced from liquid steel coming from an EAF where steel scraps are melted (together with other materials), has a different detailed chemical composition than a similar HSLA steel produced from liquid steel coming from a Blast Furnace, the later having lower residual contents (e.g.: copper content, chromium, nickel or cobalt contents) than the EAF-originating one.
[0026] In this context, using the data coming from a line where EAF-originating steel has been processed allows, thanks to the instant multi-line modelling, to predict what results are to be expected on another line, when this line will be provided with EAF-originating steel instead of Blast-Furnace-originating steel.
[0027] This helps a lot anticipating adaptations to be made to the production process and / or to the chemical compositions, for this second line, in view of a future change of primary source of steel, and thus facilitates deploying production solutions with reduced environmental impacts.
[0028] The same kind of technique of “data-bridging” can be employed for process parameters instead of chemical compositions. For instance, to predict results that could be expected on a given line after a revamping or an upgrading (allowing for instance higher cooling rates), based on production data coming from an already upgraded line.
[0029] In particular, the production data employed for training the models can be such that:- one set among of the sets of production data comprises production data for products processed using process parameters values belonging to a given category of process settings,- while another set among of the sets of production data comprises no or almost no production data for products processed using process parameters values belonging to said category of process settings
[0030] The method according to the instant technology may comprise one or several additional features, defined in claims 2 to 18, considered alone or in combination.
[0031] The instant technology concerns also a manufacturing method according to claim 19, and a method for training a method according to claim 20. This training method may also comprise one or several additional features defined in claims 2 to 14.
[0032] The instant technology concerns also a programmable electronic device, such as a computer, or a programmable controller for a production line, according to claim 21.
[0033] The instant technology concerns also a computer program, whose execution on a computer makes the computer to execute the method according to anyone of claims 1 to 18 (the computer being possibly connected to sensors and / or actuators of a metallurgical production line, or to a controller of the line, depending on the details of implementation of the method). It concerns also a non-volatile computer-readable storage medium comprising such a computer program.
[0034] The instant technology will now be described in more detail and illustrated by examples without introducing limitations, with reference to the appended figures.
[0035] Figure 1 is a schematic representation of a set of individual models together with a global model.
[0036] Figure 2 is a schematic representation of data availability in two individual data sets coming respectively from two production lines, these two data sets being used to train a global, multi-line model.
[0037] One objective of the instant method is to determine a final property of a metallurgical product at an output of a metallurgical production line in which processing operations are applied to the product, the final property being determined, using a metallurgical processing model, based on at least on:- characteristics of the product (prior to its processing in the line), in particular its chemical composition,- and on process parameters relative said processing operations.
[0038] The characteristics of the product, prior to its processing on the line, may be directly measured. Regarding the chemical composition, for instance, it may be a composition measured by analysing a sample of liquid steel taken at the tundish level, before the casting (the chemical composition remaining then constant, during subsequent processing steps, at least for the bulk of the product). The characteristics of the product prior to its processing on the line may also be determined non-directly, using a model.
[0039] The method is particularly well adapted to the field of steelmaking for which final properties of a product result from a strong interaction between initial product characteristics and conditions (temperatures, reductions, displacement speed...) employed for the product processing. The below description is made in the case of steel products and steelmakinglines. Yet the instant method may also be applicable fruitfully to other metallurgical products, such as aluminium-based products for instance. In the following, the metallurgical product and the metallurgical production line(s) are also designated, indifferently as the product and the production line(s) (without necessarily specifying “metallurgical”), or as the steel product and the steelmaking line(s).
[0040] The metallurgical production line in question is for instance:- a continuous casting line,- a hot-rolling line (including or not including a reheating furnace),- a pickling line,- a warm rolling line,- a cold rolling line,- a continuous annealing line,- a coating line (such as a hot dip galvanizing line, galvannealing line, electro-platting line, organic coating line),- a section of one of the above-listed lines (for instance the finishing rolling section or the run-out table of a hot rolling line), or a combination of one or more of those lines.
[0041] The metallurgical product is a semi-finished product, that is an intermediary sourceproduct, destinated, after processing, to become a part, a good, or another, more finished (semi-finished product). It is for instance a slab, a billet, a bloom, a broom, an ingot, a bar, a beam, a rail, a tube, a wire or a steel sheet (coiled, or to be coiled).
[0042] The final property of the product may be:-a mechanical property, such as its Yield Strength YS, its Ultimate Tensile Strength UTS, uniform elongation UE or its elongation at break e%,-a microstructural property (grain size, or phase fraction),-a surface property such as a roughness, a flatness, a surface-defects abundance, a near-surface chemical characteristics (such as an in-taken H2 content or an intergranular oxidation feature).
[0043] The characteristics of the product (prior to its processing in the line) comprise data representative of its chemical composition CC, for instance in the form of weight% contents of alloying elements (C, Mn, Si, Cu, Cr, Ni, ....) added to the base metal (Fe). The data representative of the chemical composition CC may comprise alloying element contents for one or more of the following element contents: aluminium, arsenic, boron, carbon, calcium, chromium, copper, manganese, molybdenum, nitrogen, niobium, nickel, phosphorus, lead, sulphur, antimony, silicon, tin, titanium, and vanadium. Said data may in particular comprise alloying element contents for: aluminium, chromium, copper and nickel. It may also comprise alloying element contents for: carbon, manganese, nitrogen, phosphorus, sulphur and silicon.
[0044] The characteristics of the product may also comprise other, initial characteristics of the product such as one or more initial dimensions (in particular an initial thickness), one or more initial microstructure properties, one or more initial mechanical property, and / or one or more initial surface property.
[0045] Here, the characteristics of the product are completed by an indication of the type of steel for this product (in other words, an indication of the steel family it belongs to), for instance one of: Dual-Phase steel DP, Hight Strength Low Alloy steel HSLA, Low-Carbon LC or Medium Carbon steel MC, Interstitial Free steel IFS, Transformation Induced Plasticity steel TRIP. They may be completed also by an indication of the initial source of liquid steel, among an Electric Arc Furnace EAF and a Basic Oxygen Furnace BOF (i.e.: a converter) which processes liquid iron coming from a Blast Furnace BF. Yet, these additional indications are not employed for the global model training (nor for individual models training); they are used here when testing the predictions of the trained model, to test the prediction capability of the model for different steel families in different scenarios (these tests are presented in more detail further below).
[0046] The process parameters relative to the processing operations comprise typically parameters representative of a thermal route followed the product during said processing, or representative of a part of a thermal route, such as a cooling rate or a temperature at a key step in the process like a reheating furnace exit temperature, a rolling temperature, a hot- coiling temperature, a soaking temperature or a tempering temperature. The process parameters may also comprise: a reduction ratio (during rolling) or, equivalently, a thickness after rolling, a displacement speed of the product in the processing line, an atmosphere composition, process parameters of the run-out table (ROT), a transit or residence time in such or such section of the line, a mechanical effort (rolling force or couple) applied on the product.Data type
[0047] The metallurgical processing model, for predicting the product final property, is a regression model, here a Gaussian Process regression model, trained using past production data.
[0048] Remarkably, the production data are acquired for a number J > 2 of distinct metallurgical production lines. These production lines are of a same type. For instance, they are all hot rolling lines comprising a reheating furnace and extending until a coiler. Or they are all continuous annealing and coating lines. As explained in more details below, using data coming from different lines, in a joint manner, allows for supplementing a lack of past data for one line by data coming from another line while adapting them to the data-missing line.
[0049] It is noted that two “distinct lines” may designate a same production line in a plant, but respectively before and after a revamping of a modification of the installation (replacement or upgrade of part of the line, for instance). Of course, two distinct lines may also correspond to two different lines installed in different locations, for instance two lines belonging to two different plants.
[0050] For a given production line among those lines, and for one product that has been processed on this line, the production data acquired comprise:- the above-mentioned characteristics of the product (prior to its processing in the line) and the process parameters employed during the processing operations; they play the role of predictor data, or in other words explanatory data (they could also be called input variables, or independent variables, or features, in regression and machine learning terms); for the needs of the description, these data are presented in the form of a feature vector x=[xi,...Xd,...XD] in the following; D is the number of (scalar) variables taken into account for predicting the value of the final property y of the product ; D is typically higher than 5 and often higher than 10 ; it is for instance from 15 to 150.- the final property of the product (e.g.: UTS, YTS, ...), as obtained for the product at the end of the processing operations; this property is noted y; here, y is measured (for instance by tensile testing using a sample piece taken from the product). The property y plays the role of resulting variable (it could also be called output variable, or independent variable, or label, in regression and machine learning terms).
[0051] The values of the process parameters are either measured (directly or non-directly - that is derived from measurements using an intermediary model) or given by control signals or setpoints employed for controlling the production line when processing the product considered. These values thus represent (describe) the process conditions employed for processing the product considered. Regarding the chemical composition, it is for instant read in an industrial database where the compositions of the products, produced upstream of the line considered, are recorded. The chemical composition, recorded in this industrial database, may be determined by a chemical analysis of a sample of liquid steel taken at a tundish or tapping level. It may also be derived from the known amounts of the different compounds melted to produce the steel.
[0052] The data (x,y), acquired for one product processed on a given line (among the above- mentioned lines), forms one sample (one example), for the training of the metallurgical processing model.
[0053] The notation (x,y) is a simplified one. Indeed, for one of the above-mentioned lines, for instance for the production line number j (with j from 1 to J), several samples,corresponding to a number n, of products successively processed on that line are acquired, n, is typically above 500, for instance from 1000 to 50000, or even from 5000 to 50000.
[0054] For the line number j, the n, samples thus acquired form a set of data ©;- withwith (x / ,y / ) the sample corresponding to the production data for the product number i processed on the production line number j. The feature vector x / can be written as x / ix; , . .„,x;T.... ,x(?J. For the sake of clarity, some or all of the subscripts and superscripts i, j, d may be omitted in the following, when not indispensable.
[0055] The global set of production datagroups all the production data of the individual data sets ®f, j=1 , ... , J. rocess regression from a single set of data
[0056] Gaussian process regression is presented below in the case of a single dataset, for instance the production data from one of the production line (the index j of which being omitted).
[0057] Probabilistic regression can be formulated as: given a training set © == I,..., !?} of n samples each gathering an input vector Xi and a noisy (real, scalar) output y, compute a predictive distribution for the valuesat different test “locations” x* (that is for other values of the input variables, here other values of the product initial characteristics and process parameters). The relationship between the function / '(x) (sometimes called latent function) and the noisy observations is y£= / (x, ) + s; witha random variable. In the context of Gaussian Process Regression, E£is typically assumed to follow a normal (i.e. gaussian) distribution with a zero mean and a variance noted(gaussian homoscedastic noise). The following developments use this assumption in terms of random noise. Besides, it is reminded that for a Gaussian Process Regression, f has a joint gaussian distribution.
[0058] The covariance cov(y ,y ) between two observations ypand yqat two different “locations” xpand xq(i.e.: for two different inputs) is assumed to take the following form:with $?<? the Kronecker delta, and k(xp,xq) the covariance function (referred to in many contexts as the kernel function).
[0059] The following notations are introduced:- X = [xn... X;. ... ,x?, p which is a matrix with n lines and D columns;- X, = x. p...,x.,m]Twhich is a matrix with m lines and D columns, m being the number of input vectors for which one wants to determine the mean value of f (in addition to the inputs vectors Xi , xnfor which this value has been measured); ch is an n-components vector;which is an m-cornponents vector, with j the value at xif;- K(X,X) which is an n x n covariance matrix whose component K(X,X)PrQ(line p column q) is equal to fc(xp,xt?);- K(X;.,X) which is ancovariance matrix whose component K(X,..X);! .;(line p column q) is equal toand similarly for K(X..X. ) and K(X,X,).
[0060] The mean values of the predictive distributionat the “test points” is then given by eqn-1 below (where I is the identity matrix):(eqn-1)
[0061] The covariance of the predictive distribution(whose diagonal values are the variances for the predictions at the test points) is given by eqn-2:(eqn-2)
[0062] The log-likelihood logp(y[X), of obtaining the observed values y, given the inputs X (likelihood which represents how well the regression model reflects the acquired samples forming the training data ®), is given by eqn-3:(eqn-3)is the determinant of matrixI.
[0063] The likelihood p(y[X) is a marginal likelihood (as the possible values of function f are marginalized out), and may also be referred to as the marginal likelihood, in the following.The log-likelihood logp(y[X) may also be referred to as the log marginal-likelihood in the following.
[0064] The regression model enables to compute, in other words to predict, the values of the final property of the product,s(e.g.: UTS, elongation at break ... ), for the inputs X*, that is for a set of values of the product characteristics prior the processing and the process parameters, for which no observation or measurement was available.
[0065] The results of the training of the regression model (results given by eqn-1 and eqn-2) depend directly on the kernel function k(x.,.xi). and also on the noise variance aj?ioL,e.
[0066] There are different possibilities for the kernel function. For instance, the kernel function can be a Radial Basis Function (RBF), that is a squared exponential; it could also be of the so-called Matern class (based on modified Bessel functions), or it could be a polynomial function with a compact support (truncated to that support).
[0067] Anyhow, here, the kernel function is stationary, meaning that
[0068] Anyhow, the kernel function is parametrized by the following hyperparameters (at least): an amplitudeof the kernel function (that is a multiplicative factor by which a ‘bare’, base kernel function is multiplied to obtain k), and- lengthscales 1 ,d =respectively associated to the different components d=1 ,... ,D of the input vectors xp; each lengthscaleis the typical amount by which the quantity xp,d- xq,dhas to vary to make the kernel function fc vary; In other words, it is the typical distance, between xp,dand xq,d, for which a significant variation of the covariance k between these two point is obtained (for instance: beyond which xp,dand xq,dare not correlated to each other anymore) ; put another way, it is the scaling to be applied to the variable xp,d- xq,dbefore inputting it in a generic, base kernel function.
[0069] Here, a number D of distinct lengthscales are employed, respectively associated to the D components of each input vector. Yet, in other embodiments, a same lengthscale could be used for two or more of these components, or even using just one lengthscale for all the inputs.
[0070] In the examples describe below the kernel function if a radial base function whose expression is:(eqn-4)
[0071] In practice, the hyperparameters ? , ld fd = 1, — D and a?iOi5eneed to be set (values have to be attributed to them), to be able to predict the mean and variance of f*. This setting can be done manually, based on expert knowledge. It can also be done, in a systematic and automated manner, by looking for values of the hyperparameters that maximize the marginal likelihood p(y[X) or the log marginal likelihood logp(y[X) defined by eqn-3 (it is the case for the examples presented later below). This maximization can be achieved for instance using a gradient descent (by minimizing the opposite of the log marginal-likehood). Alternatively, some of the hyperparameters could be set manually while others are set by maximizing the marginal likelihood or its logarithm.
[0072] Determining optimal values for the hyperparameters, and evaluating numerically the mean and variance of f , based on eqn-1 and eqn-3 is very computer-demanding when the number n of samples, or the number n of test “locations” is substantial (due to inversion oftypically when it is higher than 1000 (). And in practice, for the applications considered here, n is typically around 10000 or more.
[0073] Instead of directly invertingaCholesky decomposition thereof is usually employed, as it is faster and more stable from a numerical point of view. But even though, the computation demand, in particular the amount of RAM required, remains very high and can even be impractical. To overcome this difficulty, different approximate resolution techniques can be employed.
[0074] In the instant embodiments, the approximate resolution method employed is a method sparse approximation method, based on so-called inducing variables u. A set of muinducing inputs, u, is chosen in a way that computational costs are reduced. To achieve this, a joint prior (i.e.: prior distribution for the values for f* and f, before taking the observations y into account, distribution for which the covariance isis approximated through an assumption of conditional independence between f* and f given u. Different additional assumptions lead to different approximations, most of which offer a computational cost reduction (asymptotically) of a factor n / mu(a reduction from a computation cost of O(n3) to a computation cost of O(mu2)). In other words, the inducing variables u are variables introduced to obtain a kind of independence between training and test function values (the inducing variables being then a kind of bridge between the training and the test set), moreprecisely a conditional independence between f and fxgiven u . The number muof inducing variables is chosen smaller than n.
[0075] This inducing-variables approximation can be implemented according to one of the methods described respectively in sections 4 to 7 of the following article: “A Unifying View of Sparse Approximate Gaussian Process Regression”, by J. Quinonero-Candela and C; E. Rasmussen, Journal of Machine Learning Research 6 (2005), 1939-1959.
[0076] Among these inducing-variables techniques, for the instant multi-line modeling, the technique called Deterministic Training Conditional (DTC) approximation is well suited. It is described in: “Fast forward selection to speed up sparse Gaussian process regression" by M. Seeger at aL, in Ninth International Workshop on Artificial Intelligence and Statistics. Society for Artificial Intelligence and Statistics, 2003 (wherein the DTC approximation is called Project Latent Variables). This approximation uses the further assumption that the relationship between u and / is deterministic. modelling
[0077] According to the instant technology, the metallurgical processing model, also designated as the global regression model or as the global model or as the multi-line model, is trained using the following steps:- s1 ) for the J distinct metallurgical production lines, acquiring the corresponding sets of production data ©,, j=1 ,J;- s2) training J individual models respectively associated to the J production lines, the individual models being Gaussian Process regression models, here, each individual model being trained using the set of production datafor the production line number ](j=1..J) it is associated to, without considering the other sets of production data except through the commonly shared lengthscales values, the different individual models being based on a same covariance, kernel function fc parametrized by a set of lengthscales ld,d =at least some of the lengthscales, being designated as the shared lengthscales, the individual models being trained under the constraint that the values of the shared lengthscales are the same or substantially the same for the different individual models, the training of the individual models comprising determining optimized values of the shared lengthscales, by maximizing a likelihood of obtaining the values acquired for said property of the products, given the chemical composition CC of the products and the process parameters employed,- s3) grouping the J sets of production data to form the global set of production data- s4) training the global model using the global set of production data the global model being based on the same kernel function k as the individual models, for which the shared lengthscales have fixed values, that are the optimized values determined in step s2.
[0078] In the examples considered in more detail here, all the lengthscales parametrizing the kernel function k are “shared lengthscales”, that is lengthscales whose values are constrained to be the same for all the individual models.
[0079] In step s2, in practice, the likelihood maximization is achieved by maximizing the log marginal-likelihood using a sparse approximation, here a Deterministic Training Conditional approximation. During this maximization, the value of the noise variancethe value of the amplitude of the kernel function of and the positions of the inducing points are optimized, as well as the values of the lengthscales f.rf ,d = 1,.... D.
[0080] In step s4, even if the values of the lengthscalesremain fixed (equal to the optimized values obtained in step s2), the value of the noise variance a£olS(?and the value of the amplitude of the kernel function vj (as well as the positions of the inducing points) are adjusted again, optimized for the global data setThis optimization is achieved by maximizing the log-likelihood logp(y|X) of obtaining the values acquired for said property of the products (that is the values of y), given the chemical composition CC of the products and the process parameters employed for processing the products (that is given the values of x). This optimization, as well as the rest of the training of the global model, is achieved here:- as explained above in the section relative to Gaussian process regression from a single set of data, the single set of data being, in this case, the global set- except that the lengthscales values remain fixed, instead of being adjusted.
[0081] Regarding step s2, for maximizing the likelihood of obtaining the observed values for y, (given the values of x), for the ensemble of individual models, it’s the sum 5' = LyZ{logp(y[X); of individual log-likelihoods logpfrjX) / , for j from 1 to J, that is maximized (by adjusting the commonly shared values of the lengthscales). Each individual log- marginal-likelihood logp(y|X)j is a log-likelihood computed for one of the sets of production data,This individual log marginal likelihood k>gp(y[X);- (and its gradient, inview of maximizing the log-likelihood) is computed as above explained in the section relative to Gaussian process regression from a single set of data (see eqn-3, in particular), the single set of data being, in this case, the individual data set £>f.
[0082] In step s2, the different individual models may also be trained, like here, under the constraint that they have the same amplitude for the kernel function.
[0083] In practice, in step s2, the training of the J individual models, which comprises determining the optimized values for the lengthscales, may be achieved by training a joint model, wherein the joint model:- is a Gaussian Process regression model,- is trained using the global set of production data- is based on the same kernel function fc as the individual models, but has a - covariance matrixwhich is a block-diagonal matrix with zero covariance between samples belonging respectively to two different sets of production dataj1, each block on the diagonal being an individual covariance matrix K associated to one of the J sets of production data
[0084] Using this joint model enforces that the kernel function and its hyperparameters values are the same for the different production lines, while the individual predictions for each line are independent from the predictions for the other line(s); in other words, the predictions for each production line depend only on individual production datafrom that line (once the lengthscales and amplitude(s) set). In other words, the lengthscales and amplitude(s) being given, the predictions for one line are independent from data of the other lines.
[0085] This way to implement step s2 has the advantage that a single model (the joint model) has to be handled, making the implementation more convenient.
[0086] It is noted that the log-likelihood for this joint model, computed according to eqn-3 (with K = KoiJit), is indeed equal to themaximizing the loglikelihood for the joint model gives thus the same lengthscales values as maximizing S.
[0087] It is noted also that the joint model has the form of a coregionalised Gaussian process, with a coregionalization matrix which is the identity matrix. Step s2 can thus be implemented by a training a coregionalised Gaussian process, which is convenient as efficient and robust algorithms are available for training coregionalised Gaussian process models, in Gaussian Process routine libraries like the Python language GPy library maintained by the Machine Learning group of the University of Sheffield, for instance.Implementing step s2 based on this joint model is thus beneficial, in that it allows for a convenient and efficient numerical implementation.
[0088] Remarkably, the production data employed for training the models can be such that:- a first set among of the individual sets of production data ©, = 1.... J, comprises production data for products whose chemical composition CC belongs to a given category of chemical compositions,- while a second set among of the sets of production data comprises no or almost no production data for products whose chemical composition CC belongs to said category of chemical compositions.
[0089] As explained in more details in the section “summary” above, in this case, the mutli- line model achieves a kind beneficial data-generalization (with a kind of abstraction from the line-specificities, to retain the general, related to physical transformation, tendencies) allowing to predict results for the second line (for which data are missing for the category of chemical compositions in question), using (inter alia) the ones from the first line.
[0090] The category in question may, like here, be a category grouping chemical compositions with high residual contents, in particular with a copper content and / or a chromium content and / or a nickel above a given threshold (e.g.: with a copper content above 0.2%, or even above 0.3%).
[0091] In practice, products made of a steel coming from an Electric Arc Furnace (where steel scraps are melted, to recycle them) are more likely to belong to this category of chemical composition. In the example of figure 2, for instance, the products of the HSLA type, made of steel originating from an EAF, belong to this category of chemical compositions (compositions with a high metallic residuals content).
[0092] Figure 1 schematically represents (as gearwheels) the individual models and the global, multi-line model. Figure 2 schematically represents data availability in the individual data sets for the production lines 1 and 2, for an exemplary use case. In the tables of figure 2, a cell with a grey background means that production data are available for this case, while a white background means that no production data are available. For line 1 , production data are available for LC, MC, HSLA and DP steel products made of steel coming from a BOF, itself supplied with liquid iron coming from a Blast Furnace, while no production data are available for steel products made of steel coming from an EAF source. For line 2, production data are available for LC, MC, HSLA and DP steel products made of steel coming from the BOF source, and also for LC and HSLA steel products made of steel coming from the EAF source (but not for MC and DP product made of steel coming from an EAF source). Having production data for HSLA steel products made of EAF-originating steel (with a high residuals content), thanks to line 2 data set, enables, thanks to the multi-line model, to anticipate, topredict what would result from producing on line 1 HSLA products made of EAF-originating steel.
[0093] Test results are presented below in case for which J=2. The two lines are two hot- rolling lines, located in two different plants. The final product considered is a hot-coiled coil. For both lines, production data are available for LC, MC, HSLA and DP steel products. For line 1 , production data are available both for products made of EAF-originating steel and for products made of BOF-originating steel. For plant 2, production data are available only for products made of BOF originating steel. Different global, multi-line models have been trained, using the above-described technique, for different properties of the product respectively, for instance one global model for the Ultimate Tensile Strength, one other for the Yield strength, and yet another global model for the namely for elongation at break e%.
[0094] For each line, about 10000 samples are available. The training of the models, using sparse approximation, the number of inducing “points” is above 50, for instance from 50 to 500. Here, it is equal to 150, for instance.
[0095] For the example considered here, in addition to a value of the property y, each sample comprises 21 values defining the chemical composition CC of the product. These values correspond to contents (expressed for instance as weight%) in: soluble aluminium, aluminium, arsenic, boron, carbon, calcium, chromium, copper, manganese, molybdenum, nitrogen, niobium, nickel, phosphorus, lead, sulphur, antimony, silicon, tin, titanium, and vanadium. Besides, each sample comprises process parameters values for the following process parameters: a temperature (of the product) at the exit of the reheating furnace, a hot-coiling temperature, a reducing ratio for one or more finishing stand(s). Each sample comprises also values for a final thickness and a strip speed (in m / min) at the exit of the rolling section, here.
[0096] Each sample also comprises an indication of a grade code or grade family for the product (and indication specifying the line, either 1 or 2, it comes from). Yet, as above mentioned, these indications are not used for the training.
[0097] Each variable (except the grade code and origin of the steel) is standardized before being used for the models training.
[0098] Tests have been achieved to confirm that data missing on one line can be compensated if corresponding data are available on another line, thanks to the multi-line model above presented. For the different types of steel product (i.e.: for the different steel families), predictions have been made for the final properties of the product while removing voluntarily from the production data the data corresponding to one type of steel product (e.g.: removing the data for the HSLA products, in one of the individual production data sets ,©2, or in both of them. A corresponding prediction error has then been computed (namely aRoot Mean Square Error RMSE) using a test set which gathers production data for the type of steel product in question.
[0099] For each pair (production line, steel family), the following tests are achieved:- test called “family out”: in the set of production data for the line considered (e.g.: for line 1 ), the data for the steel family considered are removed (e.g.: for HSLA) when computing the expected, predicted values for(prediction according to eqn-1 , for instance), and a corresponding RMSE is computed;- test called “family completely removed”: the production data for the family considered are removed both from all the production data, for all the production lines;- test called “other line only”: the complete set of production data for the line considered is removed; the predictions are thus based only on production data from the other production line.
[0100] The corresponding results are groups in tables 1 , 2 and 3 below. The prediction error for the “family out” test is most of the time smaller than for the “family completely removed” test. This illustrates that having production data for the family considered, for one of the lines, enables to improve the prediction for that family for all the lines (including the ones not having such data).
[0101] Besides, the prediction error for the “family out” test is most of the time smaller than for the “other line only” test. It shows that the multi-line modelling actually enables to adapt the information retrieved from the other line, to the line considered for which data were missing, the information retrieved from the other line being used an intermediary to reconstruct information for the line considered (rather than using this information directly, as such, which would correspond to the “other line only” model).
[0102]
[0103]
[0104] ApPiications of the. trained., mujti-h
[0105] Once the training phase above described is achieved, the multi-line model can be employed, during a use phase, for different purposes some of which being presented below.
[0106] The multi-line model, trained as above explained, can be used for determining the property of a product depending on its chemical composition and on values of the process parameters some of which at least being values of the process parameters employed when processing the product on the production line considered.
[0107] In other words, in this case, in the use phase of the multi-line model, for at least some of the process parameters (possibly for all of them), the process parameters values input in the model are either measured on the production line (directly or non-directly), or given by control signals or setpoints employed for controlling the production line whenprocessing the product. It is thus the property of the actual, physical product, at the end of the processing, that is predicted (or post-determined) by the multi-line model.
[0108] For instance, the values of all the process parameters may be measured or given (i.e.: specified) by the control signals or setpoints. In this case, the model is employed a posteriori, once the product processed on the production line. This is useful for determining the property of product finally obtained, for instance its UTS, without having to achieve direct measurements on the final product (e.g..: direct tensile strength tests using physical samples taken from the product).
[0109] Alternatively: for a first part of the processing operations, the values of the process parameters are measured, or specified by control signals or setpoints, while for a second, remaining part of the processing operations, the values of the process parameters are values planned for this remaining part of the processing operations, intended to be used to process the product.
[0110] In this last case, the property of the product is a property that is predicted for the product, expected at the end of the processing operations given the procession operations already achieved (first part of the processing operations), and given the remaining, planned ones (second part of the processing operations). In this case, the multi-line model is employed for achieving line monitoring and possibly for predictive control. The value of the property of the product, output by the multi-line model, may be transmitted to a humanmachine interface and displayed on this interface (to allow operators for monitoring the processing).
[0111] The multi-line model may in particular be integrated in a line controller configured for achieving an on-line predictive control of the processing operations, based on a comparison between: the value of said property, predicted by the multi-line model for the final product (predicted as explained just above), and a desired, target value for this property of the product.
[0112] The result of this comparison is used for instance to adjust the process parameters for the second, remaining part of the processing operations, so as to obtain a property value closer to the target one.
[0113] The multi-line model can also be used for designing a new product (namely a chemical composition, and possibly some initial characteristics of the product), and associated processing conditions suitable to obtain a given, target value for the final property for the product, or for multiple final properties of the product (using multiple multi-lines models.
[0114] To this end, the multi-line model(s) may be used in conjunction with an optimizing algorithm configured for finding an optimum chemical composition, and optionally optimum initial characteristics of the product, which minimize(s) a difference between a value of the property, as predicted by the multi-line model, and the target value for this property. Values of the processing parameters, adequate for obtaining the desired value for this property, may be determined as well, during this optimization.
[0115] The product thus designed may then be produced on the production line, preferably using the values of the process parameters thus obtained.
[0116] The method for determining the property of a metallurgical product using the multi- line model, that has been described above, can be implemented using a programmable electronic device, such as a computer. The applications presented above, in which a producing line is monitored or controlled on-line using the multi-line model, can also be implemented using a programmable electronic device which would then take the form of a programmable line controller (such as a component of a module of a Distributed Control System). The method for manufacturing a product according to a chemical composition and process parameters designed as above explained may also be automatized using such a line controller.
Claims
1. A method for determining a property of a metallurgical product output by a metallurgical production line in which processing operations are applied to the product, said property being output by a global model whose inputs comprise at least a chemical composition CC of the product and process parameters relative to said processing operations, wherein the global model is a Gaussian Process regression model or a kernelbased regression model and has been trained during a training phase using the following steps: s1 ) for a number J > 2 of distinct metallurgical production lines, acquiring, for each production line, a respective set of production data that gathers at least, for several metallurgical products that have been processed on said production line: values of the chemical composition CC of the product and of the process parameters employed when processing the product and a value of said property of the product, s2) training J individual models respectively associated to the J production lines, the individual models being Gaussian Process regression models or kernel-based regression models, each individual model being trained using the set of production data for the production line it is associated to, the different individual models being based on a same kernel function parametrized by a set of lengthscales, at least some of the lengthscales being designated as the shared lengthscales, the individual models being trained under the constraint that the values of the shared lengthscales are the same or substantially the same for the different individual models, the training of the individual models comprising determining optimized values of the shared lengthscales, s3) grouping the J sets of production data to form a global set of production data, s4) training the global model using the global set of production data, the global model being based on the same kernel function as the individual models, for which the shared lengthscales have fixed values, that are the optimized values determined in step s2.
2. A method according to the preceding claim wherein the shared lenghtscales comprise all the lengthscales of said set.
3. A method according to anyone of the preceding claims wherein: a first set among of the sets of production data comprises production data for products whose chemical composition CC belongs to a given category of chemical compositions, a second set among of the sets of production data comprising no or almost no production data for products whose chemical composition CC belongs to said category of chemical compositions.
4. A method according to the preceding claim wherein said category of chemical compositions is a category grouping chemical compositions: corresponding to one given type of steel or grade of steel, and corresponding to products made of a steel coming specifically from an Electric Arc Furnace and produced by recycling steel scraps, or, on the contrary, corresponding to products made of a steel coming specifically from a Blast Furnace.
5. A method according to anyone of the preceding claims wherein: one set among of the sets of production data comprises production data for products processed using process parameters values belonging to a given category of process settings, another set among of the sets of production data comprising no or almost no production data for products processed using process parameters values belonging to said category of process settings.
6. A method according to anyone of the preceding claims wherein the global model and the individual models are Gaussian Process regression models and the optimization of the values of the shared lengthscales is achieved by maximizing a likelihood of obtaining the values acquired for said property of the products, given the chemical composition CC of the products and the process parameters employed for processing the product.
7. A method according to the preceding claim wherein, in step s2, maximizing said likelihood is achieved by maximizing a sum of individual log-likelihoods, each individual loglikelihoods being a log-likelihood, computed for one of the sets of production data, of obtaining the values acquired for said property, given the corresponding chemical composition CC and the process parameters.
8. A method according to claim 6 or 7 wherein the training of the J individual models in step s2, which comprises determining the optimized values for the shared lengthscales, is achieved by training a joint model, wherein the joint model:- is a Gaussian Process regression model,- is trained using the global set of production data,- is based on the same kernel function as the individual models but has a covariance matrix which is a block-diagonal matrix with zero covariance between samples belonging respectively to two different sets of production data, each block on the diagonal being an individual covariance matrix associated to one of the J sets of production data.
9. A method according to anyone of the claims 6 to 8 wherein, in step s2, the maximization of the likelihood is achieved using a sparse approximation techniques..
10. A method according to the preceding claim wherein the sparse approximation method employed is the Deterministic Training Conditional approximation.
11. A method according to anyone of claims 6 to 10, wherein the kernel function is parametrized also by an amplitude, and wherein step s4 comprises determining an optimized value of said amplitude and / or determining an optimized value of a measurement noise standard deviation cn0lse, by maximizing the likelihood of obtaining the values acquired for said property of the products, given the chemical composition CC of the products and the process parameters employed.
12. A method according to anyone of the preceding claims, wherein all the production lines are of a same type, said type being one of: a continuous casting line; a hot-rolling line including or not including a reheating furnace; a pickling line; a warm rolling line; a cold rolling line; a continuous annealing line; a coating line; a section of one of the above-listed lines; a combination of one or more of the above-listed lines.
13. A method according to anyone of the preceding claims, wherein said property of the product is one of- a mechanical property, among: a Yield Strength, an Ultimate Tensile Strength, an elongation at break;- a microstructural property among: a grain size, a phase fraction;- a surface property among: a roughness, a flatness, a surface-defects abundance, a near-surface chemical characteristics.
14. A method according to anyone of the preceding claims, wherein the process parameters comprise one or more of: a cooling rate; a reheating furnace exit temperature; a rolling temperature; a hot-coiling temperature; a soaking temperature; a tempering temperature; a set of parameters defining a thermal route followed the product during the processing operations.
15. A method according to anyone of the preceding claims wherein, in use phase of the global model, when determining the property of the metallurgical product using the global model: for at least some of the process parameters, the values of the process parameters values input in the global model are either measured on the production line or specified bycontrol signals or setpoints employed for controlling the production line when processing the product.
16. A method according to the preceding claim wherein, the property of the product is determined after the product had been output by the production line, based on the values of the process parameters employed during the processing operations applied to the product.
17. A method according to claim 15 wherein:- for a first part of the processing operations already applied to the product, the values of the process parameters are measured on the production line or specified by control signals or setpoints employed for controlling the production line when processing the product,- while for a second part of the processing operations, to be applied to the product, the values of the process parameters are values planned for this second part of the processing operations, wherein the value of the property of the product, determined by the global model based on said values of the process parameters, is compared to a target value for said property, and wherein the second part of the processing operations is controlled based on the result of said comparison.
18. A method according to anyone of claims 1 to 15 wherein, in use phase of the global model, adjusted values of the chemical composition of the product, and optionally of the process parameters are determined by minimizing a difference between a target value for the property of the product, and the value of said property output by the global model, given the chemical composition of the product, and the process parameters employed.
19. A manufacturing method, wherein a metallurgical product whose chemical composition is specified by the adjusted values determined according to claim 18, is processed in the production line.
20. A method for training a global model whose inputs comprise at least a chemical composition CC of a metallurgical product and process parameters relative to processing operations applied to the product, and whose output comprises a property of a metallurgical product output by a metallurgical production line in which the processing operations are applied to the product, the global model being a Gaussian Process regression model or a kernel-based regression model and being trained using the following steps:1 s1 ) for a number J > 2 of distinct metallurgical production lines, acquiring, for each production line, a respective set of production data that gathers at least, for several metallurgical products that have been processed on said production line: values of the chemical composition CC of the product and of the process parameters employed when processing the product and a value of said property of the product, s2) training J individual models respectively associated to the J production lines, the individual models being Gaussian Process regression models or kernel-based regression models, each individual model being trained using the set of production data for the production line it is associated to, the different individual models being based on a same kernel function parametrized by a set of lengthscales, at least some of the lengthscales being designated as the shared lengthscales, the individual models being trained under the constraint that the values of the shared lengthscales are the same or substantially the same for the different individual models, the training of the individual models comprising determining optimized values of the shared lengthscales, s3) grouping the J sets of production data to form a global set of production data, s4) training the global model using the global set of production data, the global model being based on the same kernel function as the individual models, for which the shared lengthscales have fixed values, that are the optimized values determined in step s2.21 . Programmable electronic device comprising at least a processor and a non-transitory memory, configured for executing the method according to anyone of the preceding claims.
22. Computer program comprising instructions whose execution on a computer make the computer to execute the method according to anyone of the preceding claims.
Citation Information
Patent Citations
LIBS quantitative analysis method based on Gaussian process regression
CN117929356A
System and predictive modeling method for smelting process control based on multi-source information with heterogeneous relatedness
US20180081339A1