A method for generating and optimizing converter charge data

The FD-VEGAN model generates complementary data that is highly consistent with the real data, which solves the problem of sparse label samples in converter steelmaking, and improves the training effect of soft measurement model and the efficiency of converter batching optimization.

CN120340716BActive Publication Date: 2025-08-22HEBEI UNIV OF TECH +3
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510820144.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

During the converter steelmaking process, due to the sparse label samples and incoherent distribution, the training effect of the generative model is limited, and it is difficult to generate complete data that is highly consistent with the real value, which affects the training and update effect of the soft measurement model.

Method used

Using the FD-VEGAN model, the potential space is constructed by introducing encoder and generator, combined with forced discriminators and improved TD3 algorithms, a complete batching data set that meets process constraints is generated and optimized to generate high-quality converter batching data.

Benefits of technology

The missing tag samples are effectively filled, and the completion data is generated that is highly consistent with the real data, which improves the training effect and online update performance of the soft measurement model, and optimizes the multi-objective optimization of the converter batching process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340716B_ABST
    Figure CN120340716B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating and optimizing converter batching data, comprising: obtaining an original industrial data set of a converter batching process, preprocessing the original industrial data set, inputting the preprocessed training data into an encoder, generating a mean vector and a variance vector of a latent variable distribution, and constraining the similarity between the latent variable distribution and a standard Gaussian distribution through a KL divergence loss; inputting the latent variables and a random noise vector into a generator to generate simulated samples, and distinguishing between real samples and generated samples through a discriminator; introducing a forced discriminator module to convert the discrimination result into an additional loss term; generating data for missing batching parameters, outputting a complete batching data set that meets process constraints, constructing a cost objective function, a quality objective function, and a resource consumption objective function, and constructing constraint conditions for the chemical element composition content of a target finished product, optimizing the converter batching data, and generating optimized converter batching data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet of Things industry, and in particular to a method for generating and optimizing converter batching data. Background Art

[0002] With the rapid development of the Internet of Things (IoT), industrial sites are able to collect massive amounts of data in real time. In the converter steelmaking process, various ingredient data significantly impact the final product, necessitating data collection to facilitate timely adjustments based on the data. However, this data often lacks completeness requirements. Specifically, auxiliary variables (such as temperature, pressure, and flow) used as inputs to soft sensor models are typically sampled frequently, while key quality variables (label variables) serving as model outputs are often sampled less frequently or even severely missing due to the difficulty or high cost of measuring them. This data imbalance results in a scarcity of labeled samples in industrial processes, severely limiting the training and online update performance of soft sensor models. Therefore, effectively filling in the gaps in labeled samples has become a research hotspot in the field of industrial process soft sensing. Generative models, with their powerful data generation capabilities, offer a new approach to addressing the problem of missing labeled samples. By learning the inherent distribution structure of the data, these models can generate complementary data that is highly consistent with the real samples, providing sufficient data support for the training and update of downstream soft sensor models. However, in actual industrial scenarios, the training effect of generative models is often limited due to the scarcity and incoherent distribution of labeled samples.

[0003] There are few comparative documents in this field. The patent with application number 202011352099.3 in a similar field discloses a method for expanding unbalanced time series data for industrial fault diagnosis. Step 1: prepare a training data set; Step 2: construct the network structure of GRUBEGAN; Step 3: train the constructed GRUBEGAN network model; Step 4: generate small sample type artificial data based on the trained GRUBEGAN generative adversarial network model. After training, the model inputs a simple random variable z|t to generate time series data that conforms to time t, and expands the generated data set to the small sample type of the original data. The industrial field environment is complex, and the data often contains noise and outliers, resulting in unstable sample quality, making it difficult for the generator to accurately capture key data features, and there may be significant deviations between the generated samples and the true values. Summary of the Invention

[0004] To achieve the above-mentioned and other related purposes, the present invention discloses a method for generating and optimizing converter charge data, comprising:

[0005] Step S1: obtaining an original industrial data set of a converter batching process and preprocessing the original industrial data set to obtain preprocessed training data, wherein the original industrial data set includes multidimensional sensor data and corresponding missing batching quality parameters;

[0006] Step S2: Input the preprocessed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity between the latent variable distribution and the standard Gaussian distribution through the KL divergence loss;

[0007] Step S3: Input the latent variables and random noise vectors into the generator to generate simulated samples, and use the discriminator to distinguish between real samples and generated samples; introduce a forced discriminator module, which performs compliance judgment on the generated samples based on the prior constraints defined in the converter metallurgical process knowledge base, and converts the judgment results into additional loss terms;

[0008] Step S4: Construct a loss function, including generative adversarial loss, KL divergence loss, additional loss of the forced discriminator, and reconstruction loss, and jointly optimize the generator and discriminator parameters through backpropagation;

[0009] Step S5: Use the trained FD-VEGAN model to generate data for the missing ingredient parameters and output a complete ingredient dataset that meets the process constraints;

[0010] Step S6: Construct a cost objective function, a quality objective function, and a resource consumption objective function, and construct constraints on the chemical element content of the target finished product. Use the improved TD3 algorithm to solve and optimize the converter batching data to generate optimized converter batching data.

[0011] Furthermore, step S2 includes:

[0012] The encoder encodes the input true sample into a mean vector and a variance vector;

[0013] On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss;

[0014] The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss;

[0015] The encoder optimizes its performance based on the regularization loss and reconstruction loss.

[0016] Furthermore, step S3 includes:

[0017] The generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples, generating the discriminator loss;

[0018] Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss;

[0019] The generator is adversarially trained based on the discriminator loss and the additional loss to generate high-quality samples.

[0020] Furthermore, the regularization loss includes:

[0021] ;

[0022] ;

[0023] in, is the probability distribution characteristic of the latent variable z obtained based on a specific input x, is the posterior distribution given by the encoder, is the prior distribution of the latent space, is the definition of KL divergence, where the subscripts prior and post refer to the prior distribution and posterior distribution generated from the encoder, respectively. and denote the mean and standard deviation of the prior distribution of the standard Gaussian distribution, and denote the mean and standard deviation of the posterior distribution of the encoder output, respectively.

[0024] Furthermore, the reconstruction loss includes:

[0025] ;

[0026] is a multivariate Gaussian distribution.

[0027] Furthermore, in the basic VAE model, the total loss of regularization loss and reconstruction loss is:

[0028] ;

[0029] in is the reconstruction loss, is the regularization loss.

[0030] Furthermore, in the dynamic β-VAE model, the total loss of regularization loss and reconstruction loss is:

[0031] ;

[0032] in is the reconstruction loss of the nth round, is the regularization loss of the nth round;

[0033] is a parameter that changes dynamically according to the round.

[0034] Furthermore, the loss of the generator in step S3 is:

[0035] ;

[0036] in, and are weighting coefficients, which are used to adjust the latent variables and noise The loss of generated samples, is the weight coefficient of reconstruction loss, It is the reconstruction loss, which ensures that the generator can reconstruct the input data according to the latent variables. The generator not only needs to generate noise samples To create new data, we also generate latent variables from the encoder To reconstruct the real sample.

[0037] Furthermore, the loss function of the discriminator is:

[0038] ;

[0039] in, is the score of the discriminator on the real sample, is the discriminator's response to the latent variable generated from the encoder The score of the generated samples, is the discriminator's response to the noise The scores of the generated samples.

[0040] Furthermore, suppose that in a batch of M samples, a generated virtual sample is , for the i-th element , which generates additional penalties The size is set as follows:

[0041] ;

[0042] Where, is the element involved in forced discrimination in the i-th sample, To set the elements in the sample according to prior knowledge in the forced discriminator The upper limit of the value, is the lower limit, is an adjustable parameter, and the sum of the additional penalties for this batch of samples is:

[0043] .

[0044] Furthermore, step S6 includes:

[0045] Monitor the progress of each objective function in real time and dynamically adjust the weight according to the relative rate of change of the objective function to avoid over-optimization of the objective function;

[0046] By adding historical experience samples and calculating the cosine similarity between the current state and historical experience, relevant experience is prioritized and combined with random experience selection to maintain training diversity, thereby accelerating strategy convergence and avoiding overfitting;

[0047] Noise is added during the solution process, and the noise scale is adaptively adjusted according to the reward variance and the number of training steps, so that the solution converges to the optimal strategy.

[0048] Furthermore, adaptively adjusting the noise scale based on reward variance and training steps includes:

[0049] After each training round, the average reward variance of the most recent reward is calculated. When the variance is lower than the preset value, the noise is increased; when the variance is higher than the preset value, the noise is reduced.

[0050] The noise also decays exponentially with the number of training steps.

[0051] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the method described is performed when the program is executed by a processor.

[0052] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the described method.

[0053] In a fourth aspect, the present invention provides a converter charge data generation and optimization system, comprising:

[0054] A data acquisition module is used to obtain an original industrial data set of the converter batching process and preprocess the original industrial data set to obtain preprocessed training data. The original industrial data set includes multi-dimensional sensor data and corresponding missing batching quality parameters;

[0055] The pre-training module is used to input the pre-processed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity of the latent variable distribution to the standard Gaussian distribution through the KL divergence loss;

[0056] A sample generation module is used to input the latent variables and random noise vectors into a generator to generate simulated samples, and a discriminator is used to distinguish between real samples and generated samples. A forced discriminator module is introduced. The forced discriminator performs compliance judgment on the generated samples based on the prior constraints defined in the converter metallurgical process knowledge base, converts the judgment results into additional loss terms, and jointly optimizes the parameters of the generator and discriminator.

[0057] The data generation module is used to generate data for missing ingredient parameters using the trained FD-VEGAN model and output a complete ingredient dataset that meets the process constraints;

[0058] The batching data optimization module is used to construct the cost objective function, quality objective function and resource consumption objective function, and to construct the constraint conditions of the chemical element composition content of the target finished product. It uses the improved TD3 algorithm to solve and optimize the converter batching data to generate the optimized converter batching data.

[0059] Furthermore, in the pre-training module:

[0060] The encoder encodes the input true sample into a mean vector and a variance vector;

[0061] On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss;

[0062] The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss;

[0063] The encoder optimizes its performance based on the regularization loss and reconstruction loss.

[0064] Furthermore, in the sample generation module:

[0065] The generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples, generating the discriminator loss;

[0066] Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss;

[0067] The generator is adversarially trained based on the discriminator loss and the additional loss to generate high-quality samples.

[0068] Through the above technical solution, first, a VAE model is introduced to construct a latent space. VAE is used to encode existing data, capturing the global characteristics of the data and constructing a structured latent space. Second, after obtaining the latent vectors generated by the VAE encoder, the GAN generator assigns different generation weights to different vectors in the latent space, thereby generating more representative samples. In addition, the encoder maps real samples to the latent space, forming a closed-loop bidirectional mapping mechanism with the samples generated by the generator, further enhancing the similarity and distribution consistency between the generated samples and the real data. Combined with the forced discriminator module, the method can effectively guide the generator to more accurately simulate key data features and generate high-quality samples. Then, an improved TD3 algorithm is used to solve and optimize the converter batching data, generating optimized converter batching data. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0070] Figure 1 is a flow chart of the present invention;

[0071] Figure 2 is a system block diagram of the present invention;

[0072] Figure 3 This is the framework diagram of the FD-VEGAN method;

[0073] Figure 4 It is a data enhancement module based on VAE;

[0074] Figure 5 A sample generation module based on FD-GAN;

[0075] Figure 6 Framework diagram of converter alloy batch optimization method based on improved TD3 algorithm

[0076] Figure 7 The comparison chart of the experimental results of each comparison model;

[0077] Figure 8 Comparison chart of experimental results of various ablation models;

[0078] Figure 9 This is a graph showing changes in total raw material amount G, raw material cost F, and product quality Z. DETAILED DESCRIPTION

[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0080] Reference Figure 1 The embodiment of the present invention provides a method for generating and optimizing converter charge data, comprising:

[0081] Step S1: obtaining an original industrial data set of a converter batching process and preprocessing the original industrial data set to obtain preprocessed training data, wherein the original industrial data set includes multidimensional sensor data and corresponding missing batching quality parameters;

[0082] Step S2: Input the preprocessed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity between the latent variable distribution and the standard Gaussian distribution through the KL divergence loss;

[0083] Step S3: Input the latent variables and random noise vectors into the generator to generate simulated samples, and use the discriminator to distinguish between real samples and generated samples; introduce a forced discriminator module, which performs compliance judgment on the generated samples based on the prior constraints defined in the converter metallurgical process knowledge base, and converts the judgment results into additional loss terms;

[0084] Step S4: Construct a loss function, including generative adversarial loss, KL divergence loss, additional loss of the forced discriminator, and reconstruction loss, and jointly optimize the generator and discriminator parameters through backpropagation;

[0085] Step S5: Use the trained FD-VEGAN model to generate data for the missing ingredient parameters and output a complete ingredient dataset that meets the process constraints.

[0086] Step S6: Construct a cost objective function, a quality objective function, and a resource consumption objective function, and construct constraints on the chemical element content of the target finished product. Use the improved TD3 algorithm to solve and optimize the converter batching data to generate optimized converter batching data.

[0087] Furthermore, step S2 includes:

[0088] The encoder encodes the input true sample into a mean vector and a variance vector;

[0089] On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss;

[0090] The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss;

[0091] The encoder optimizes its performance based on the regularization loss and reconstruction loss.

[0092] Furthermore, step S3 includes:

[0093] The generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples, generating the discriminator loss;

[0094] Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss;

[0095] The generator is adversarially trained based on the discriminator loss and the additional loss to generate high-quality samples.

[0096] Furthermore, the regularization loss includes:

[0097] ;

[0098] ;

[0099] in, is the posterior distribution given by the encoder, is the prior distribution of the latent space (usually a standard normal distribution), For the definition of KL divergence, the subscripts prior and post refer to the prior distribution and posterior distribution generated from the encoder, respectively.

[0100] Furthermore, the reconstruction loss includes:

[0101] ;

[0102] is a multivariate Gaussian distribution.

[0103] Furthermore, the total loss of regularization loss and reconstruction loss is:

[0104] ;

[0105] in is the reconstruction loss, is the regularization loss.

[0106] Furthermore, the total loss of regularization loss and reconstruction loss is:

[0107] ;

[0108] in is the reconstruction loss, is the regularization loss;

[0109] Dynamically changing Expressed as:

[0110] );

[0111] Among them, the sign function Returns any real number The symbol, ;

[0112] ;

[0113] ;

[0114] ;

[0115] Among them, the hyperparameters It was before All rounds of yes The last change in the round, where To control the weight decay rate of the regularization term loss, To control the weight enhancement rate of the reconstruction loss, is the momentum factor, used to stabilize the mutation of β value, are the smoothing coefficients of the historical loss reference window respectively;

[0116] Furthermore, the loss of the generator in step S3 is:

[0117] ;

[0118] in, and are weighting coefficients, which are used to adjust the latent variables and noise The loss of generated samples, It is the reconstruction loss, which ensures that the generator can reconstruct the input data according to the latent variables. The generator not only needs to generate noise samples To create new data, we also generate latent variables from the encoder To reconstruct the real sample.

[0119] Furthermore, the loss function of the discriminator is:

[0120] ;

[0121] in, is the score of the discriminator on the real sample, is the discriminator's response to the latent variable generated from the encoder The score of the generated samples, is the discriminator's response to the noise The scores of the generated samples.

[0122] Furthermore, suppose that in a batch of M samples, a generated virtual sample is , for the i-th element , which generates additional penalties The size is set as follows:

[0123] ;

[0124] Where, is the element involved in forced discrimination in the i-th sample, To set the elements in the sample according to prior knowledge in the forced discriminator The upper limit of the value, is the lower limit, is an adjustable parameter, and the sum of the additional penalties for this batch of samples is:

[0125] .

[0126] Step S6 specifically includes:

[0127] Monitor the progress of each objective function in real time and dynamically adjust the weight according to the relative rate of change of the objective function to avoid over-optimization of the objective function;

[0128] By adding historical experience samples and calculating the cosine similarity between the current state and historical experience, relevant experience is prioritized and combined with random experience selection to maintain training diversity, thereby accelerating strategy convergence and avoiding overfitting;

[0129] Noise is added during the solution process, and the noise scale is adaptively adjusted according to the reward variance and the number of training steps, so that the solution converges to the optimal strategy.

[0130] The following is an embodiment of the present invention, which explains the above steps.

[0131] To address the problem of scarce labeled samples affecting the training and updating of downstream soft sensor models in actual industrial scenarios, this embodiment proposes a sample filling method based on FD-VEGAN. First, a pre-training module based on VAE is proposed to construct a latent space, capture the global features and hidden structure of the input data, and generate a latent vector. Second, a sample generation model based on FD-GAN is proposed. After obtaining the latent vector generated by VAE, GAN further generates completed samples based on the sampling of the latent space. Combined with the discriminator and forced discriminator modules, it guides the generator to more accurately simulate key data features, solves the problems of unstable model training and mode collapse, and ensures the high quality and stability of generated samples.

[0132] Generating samples using generative models usually relies on a large number of high-quality training samples. That is, when labeled samples are scarce, high-quality available samples are scarce, which affects the training effect of the generative model. Figure 3 and Figure 4 This embodiment proposes a pre-training module based on VAE. First, the encoder encodes the input real sample into a mean vector and a variance vector. These variables are combined into a latent variable input generator on the one hand, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to regularize the encoder. Figure 4 and Figure 5 ,Secondly, the generator takes the latent variables as input and generates ,reconstructed samples similar to the original samples through the decoding process.,The difference between the reconstructed samples and the real samples is quantified ,through the reconstruction loss to optimize the performance of the encoder, ensure the ,reasonable distribution of the latent variables, and encourage the generator to generate high-quality ,samples.

[0133] In the input layer, VAE encodes the input into a multivariate latent distribution to improve the structure of the latent space. The encoder maps the input to a multivariate Gaussian distribution Approximately true, the decoder then reconstructs the input data from the latent variables, given by the density function It turns out that the encoder and decoder neural networks are and Parameterization, the loss function includes a reconstruction term and a regularization term. The reconstruction term measures the difference between the input sample and the generated sample. The regularization term is used to constrain the structure of the latent space and ensure that the posterior distribution of the latent variable is close to the prior distribution. The Kullback-Leibler (KL) divergence is usually used to calculate the difference between the posterior distribution and the prior distribution. The loss function is shown in Formula 1.1:

[0134] (1.1)

[0135] (1.2)

[0136] (1.3)

[0137] in, is the probability distribution characteristic of the latent variable z obtained based on a specific input x, is the posterior distribution given by the encoder, is the prior distribution of the latent space (usually a standard normal distribution), The definition of KL divergence, the subscripts prior and post refer to the prior distribution and posterior distribution generated from the encoder respectively. The total loss function of VAE is shown in formula 1.4:

[0138] (1.4)

[0139] (1.5)

[0140] The quality of the autoencoder reconstruction is determined by control, is the regularization loss. At a high level, the regularization term controls the smoothness and regularity of the latent space. The trade-off between these two losses affects the performance of VAE. This example proposes to apply a scaling trick to the regularization loss, resulting in a modified β-VAE objective function:

[0141] (1.6)

[0142] The role of is to balance the reconstruction and regularization losses, the lower Values ​​of will produce better reconstructions, while higher The value will cause the posterior to collapse. To solve this problem, we use - Annealing, i.e. Starting from a lower value and gradually increasing to a fixed point, in any given epoch n, the dynamic The objective function of VAE is:

[0143] (1.7)

[0144] Dynamically changing Expressed as:

[0145] (1.8)

[0146] Among them, the sign function Returns any real number The symbol, Formula 1.8 aims to optimize the reconstruction term and regularization term, respectively for increase and decrease.

[0147] (1.9)

[0148] (1.10)

[0149] (1.11)

[0150] Among them, the hyperparameters It was before All rounds of yes The last round changed.

[0151] when When is positive, because in formula 1.8 item, Decreases, which means that the reconstruction loss increases compared to the historical minimum reconstruction loss according to formula 1.9. When is positive, similarly, according to formula 1.10, it means that the regularization loss is increasing. The last term of formula 1.11 is Added a loss constraint for the entire model to slow down The changing trend of The value changes slowly.

[0152] Dynamic β-VAE changes dynamically through loss difference value, The change of value will affect the loss, and they will affect each other until a balance is reached. The value will change according to the loss value in formula 1.7-1.11. During the training process, Starting from 0, when and The model converges when the balance is maintained without increasing the total loss.

[0153] There are complex correlations between the various variables in industrial process data, and the data quality is unstable. Traditional generative models have difficulty simulating the true distribution of data, resulting in distorted generated samples. To address this problem, this embodiment proposes a sample generation module based on FD-GAN. First, the generator receives the latent variables and random noise output by the encoder as input, generates corresponding samples, and uses the discriminator to distinguish the generated samples from real samples. Secondly, a forced discriminator based on prior knowledge is introduced, and its discrimination result is added as an additional loss term to the optimization objective function, guiding the generator to avoid generating samples that do not conform to the prior knowledge. Finally, the generator and discriminator optimize the generator through adversarial training to generate high-quality samples, making the samples more suitable for training downstream models.

[0154] The loss functions of the generator and discriminator are derived from the classic GAN framework. The goal of the generator is to generate samples that are as realistic as possible, so that the discriminator cannot distinguish between the generated samples and the real samples. The loss function of the generator is defined by the following formula:

[0155] (1.12)

[0156] in, and are weighting coefficients, which are used to adjust the latent variables and noise The loss of the generated samples. is the reconstruction loss, which ensures that the generator can reconstruct the input data according to the latent variables. The generator not only needs to generate noise samples To create new data, we also generate latent variables from the encoder To reconstruct real samples, unlike traditional GAN, the generator here is not only used as a generator, but also acts as a decoder in VAE.

[0157] The task of the discriminator is to distinguish between real data and generated data (samples from the generator). The loss function of the discriminator is:

[0158] (1.13)

[0159] in, is the score of the discriminator on the real sample, is the discriminator's response to the latent variable generated from the encoder The score of the generated samples, is the discriminator's response to the noise The scores of the generated samples.

[0160] Next, we introduce prior knowledge to enhance the discriminator’s ability. By adding corresponding terms to the loss function, we give appropriate additional losses when the constraint conditions are triggered to guide the generator to generate useful samples. Suppose that in a batch of M samples, a generated virtual sample is , for the i-th element , which generates additional penalties The size is set as follows:

[0161] (1.14)

[0162] Where, is the element involved in forced discrimination in the i-th sample. To set the elements in the sample according to prior knowledge in the forced discriminator The upper limit of the value, is the lower limit. is an adjustable parameter. The sum of the additional penalties for this batch of samples is:

[0163] (1.15)

[0164] Add an additional penalty term to the generator loss:

[0165] (1.16)

[0166] The loss of the discriminator is still , and finally the objective function of GAN with forced discriminator is:

[0167] (1.17)

[0168] When using a generative model to fill in label samples of industrial production data, the problem of low quality of generated samples occurs due to the scarcity and incoherent distribution of label samples. In response to the above problems, this embodiment proposes a sample filling method based on FD-VEGAN. First, a VAE model is introduced to construct a latent space, and the existing data is encoded by VAE to capture the global characteristics of the data and construct a structured latent space; secondly, after obtaining the latent vector generated by the VAE encoder, the GAN generator assigns different generation weights to different vectors in the latent space, thereby generating more representative samples; in addition, the encoder maps the real samples to the latent space, forming a closed-loop bidirectional mapping mechanism with the samples generated by the generator, further enhancing the similarity and distribution consistency between the generated samples and the real data. Combined with the forced discriminator module, the method can effectively guide the generator to more accurately simulate key data features and generate high-quality samples.

[0169] Converter batching refers to the selection of appropriate raw materials and additives during the converter steelmaking process, based on the specific requirements of steelmaking, to ensure the smooth progress of the steelmaking process and the quality of the final product. The main raw materials for converter steelmaking include:

[0170] Molten iron and scrap iron: This is one of the main raw materials for converter steelmaking, providing the iron element required for steelmaking1.

[0171] Steelmaking pig iron: used to increase the carbon content and impurity elements in molten iron.

[0172] Alloy Return Steel: Used to adjust the chemical composition of steel.

[0173] Slag-forming materials: such as limestone, dolomite, etc., are used to control the composition and properties of slag.

[0174] Oxidants: such as oxygen, used to oxidize impurity elements.

[0175] Deoxidizers and alloying agents: used to adjust the chemical composition of steel and remove oxides.

[0176] Recarburizer: used to increase the carbon content in steel.

[0177] The selection and use of these raw materials requires precise calculation and adjustment based on the specific steelmaking requirements and the chemical composition of the target product. By rationally selecting and using these raw materials, the smooth operation of the converter steelmaking process can be ensured, and high-quality steel that meets the requirements can be obtained.

[0178] Afterwards, refer to Figure 6 The process of optimizing the alloying ingredients usually requires considering multiple objectives simultaneously, including cost minimization, quality optimization, and resource consumption minimization. These objectives often conflict with each other. First, an objective function weight adjustment module is proposed. By monitoring the progress of each objective function in real time and dynamically adjusting the weight according to the relative rate of change of the objective function, the over-optimization of certain objective functions is avoided while ensuring that the progress of other objectives is given sufficient attention. Secondly, a similarity experience replay module is proposed. By calculating the cosine similarity between the current state and historical experience, relevant experience is prioritized and combined with the selection of random experience to maintain training diversity, thereby accelerating the convergence of the strategy and avoiding overfitting. Finally, a noise adjustment module is proposed. The noise scale is adaptively adjusted according to the reward variance and the number of training steps, ensuring sufficient exploration in the early stage of training and gradual convergence to the optimal strategy in the later stage. Through the above methods, the multi-objective optimization process in alloying ingredients is optimized and a better alloying ingredient solution is provided.

[0179] Specifically include:

[0180] First, we construct the objective function. To reduce the cost of the factory's actual production process, increase resource utilization, optimize environmental benefits, and ensure product quality, we establish the objective function formulas (1.18), (1.19), and (1.20).

[0181] (1.18)

[0182] Among them, F represents the cost of raw materials, nc represents the number of nc kinds of raw materials, represents the unit price of the i-th raw material (yuan / kg), Indicates the amount of the i-th added raw material.

[0183] (1.19)

[0184] (1.19) Where G represents the total amount of raw materials used, nc represents the number of nc types of raw materials, Indicates the amount of the i-th added raw material.

[0185] (1.20)

[0186] in, It indicates the weight of the molten steel in the ladle when tapping, Z indicates the quality of the product, and mc indicates that there are mc elements. represents the amount of the i-th added raw material, represents the element content of the jth element contained in the i-th raw material, represents the element yield of the i-th raw material, represents the element content of the jth element in molten steel, It represents the optimal control point of the content of the jth element in the product. The smaller the value of Z, the better the product quality.

[0187] The following constraints are established based on the chemical element content requirements and quality requirements of the ingredients in the actual production process.

[0188] The chemical element content of the target finished product can be kept within a range, that is, the content of each element should be greater than the lower limit and less than the upper limit:

[0189] (1.21)

[0190] in, is the lower limit of the composition requirement of the jth element of the finished product, is the upper limit of the composition requirement of the jth element of the finished product, is the content of the jth element in the i-th raw material, Represents the element yield of the i-th raw material.

[0191] The amount of each raw material added The following non-negativity constraints should be satisfied: .

[0192] To address conflicts and weight imbalances between different objectives in converter batch optimization, the weights of each objective function are dynamically adjusted to ensure a balance between them. This module monitors the optimization progress of each objective function in real time and dynamically adjusts the weights of each objective function based on the actual progress. When an objective function makes significant progress, its weight is reduced accordingly to avoid over-optimization of that objective function. Conversely, when an objective function makes slow progress, its weight is increased to ensure that it receives adequate attention during subsequent optimization. This dynamic weight adjustment method effectively balances the importance of each objective function in the multi-objective optimization process, avoiding an imbalance in importance between objective functions, thereby improving the quality of the optimization results and ensuring the rationality of the final solution.

[0193] The rate of change of the objective function reflects the progress of the objective function in the current iteration. In the current iteration and the previous iteration The values ​​of and , then its relative rate of change can be expressed as:

[0194] (1.22)

[0195] This formula measures the objective function The magnitude of change in the current iteration. If the relative rate of change is large, it indicates that the objective function has made significant progress in the current iteration; otherwise, it indicates that the objective function has made little progress. For each objective function If the rate of change of the objective function in the current iteration is large, its weight should be reduced to avoid over-optimization of the objective function that has made great progress in the subsequent optimization process. On the contrary, if the rate of change of the objective function is small, it means that it has made little progress and needs to be given a larger weight. Based on this, a dynamic weight adjustment formula is proposed:

[0196] (1.23)

[0197] in, is a small constant used to prevent division by zero errors. This formula shows that when the objective function When the relative rate of change is large, the weight will decrease; conversely, when the rate of change is small, the weight will increase. In order to ensure that the sum of the weights is 1, the weights of all objective functions need to be standardized:

[0198] (1.24)

[0199] This weight normalization ensures that the sum of the weights of all objective functions is 1 after each optimization iteration, while dynamically adjusting their attention in the optimization process according to the progress of the objective functions.

[0200] Finally, the adjusted weights are used to construct the reward function. The reward function is in the form of:

[0201] (1.25)

[0202] in, is the dynamically adjusted target weight, is the standardized objective function value, is the final comprehensive reward function.

[0203] (1.26)

[0204] in, is the objective function In the current solution The value below, and The objective functions are Mean and standard deviation over all iterations of history:

[0205] (1.27)

[0206] (1.28)

[0207] To address the issues of low sample efficiency and the potential for local optima during training in converter batch optimization, we replay representative and similar experience samples to accelerate policy convergence, enhance exploration, and avoid local optima. Learning efficiency is improved by prioritizing experiences that are highly relevant to the current state. Random experience selection ensures training diversity, thus avoiding model overfitting and accelerating policy convergence.

[0208] The experience replay buffer is used to store the experience gained by the agent in interacting with the environment. Each experience is usually represented as a four-tuple, including the current state, the action taken by the agent, the reward obtained, and the new state transferred to. Specifically, each experience is represented as:

[0209] (1.29)

[0210] in, Represents the current state, indicating the environmental information of the agent network at the current moment, Represents the state The action selected by the agent network is It is the agent network that performs the action The immediate reward obtained from the environment, Is to perform an action The new state reached.

[0211] In the similarity experience playback mechanism, by calculating the cosine similarity between the current state and historical experience, the experience is divided into two categories: similar experience and random experience. Similar experience refers to the experience that is similar to the current state. Similar experiences help the agent network learn quickly in the context of the current task and optimize its strategy; random experiences refer to experiences randomly selected from the replay buffer, with the aim of maintaining training diversity and avoiding overfitting of the model that only focuses on the current situation.

[0212] In order to select the experience most relevant to the current state from the experience replay buffer, cosine similarity is used to measure the similarity between the current state and historical experience. The angle between the two vectors is calculated to determine their similarity. The calculation formula is:

[0213] (1.30)

[0214] in, and Respectively indicate the current status and historical experience The vector representation of and Represent their L2 norms respectively. The cosine similarity ranges from -1 to 1. The closer the value is to 1, the more similar the two are.

[0215] The number of similar experiences and random experiences is controlled by the similarity of the experiences and a certain ratio. In each training step, a hyperparameter is set To express the proportion of similar experiences. Specifically, assuming that the current playback buffer contains experiences, and the agent network wants to select from the buffer Experience ( can be greater or less than the total amount of experience in the buffer), according to the ratio , select a certain proportion of similar experiences and random experiences.

[0216] Number of similar experience options for:

[0217] (1.31)

[0218] in, is the selection ratio of similar experiences, It is the minimum number of similar experiences to prevent too few similar experiences from being selected.

[0219] Random experience selection amount for:

[0220] (1.32)

[0221] in, represents the number of randomly selected experiences, This is the minimum amount of random experience to ensure that there is enough variety in each training session.

[0222] After each experience is selected from the replay buffer, the agent network uses this experience to update its policy. The update process is driven by the temporal difference (TD) error. The TD error is used to evaluate the difference between the agent's predicted value and the actual reward received, and adjust the agent's policy based on this difference. The TD error is calculated as:

[0223] (1.33)

[0224] in, is the reward function, Is the Q value corresponding to the current experience, indicating that in state Take action The value of is a discount factor that controls the impact of future rewards, Is the next state The best action in , the experience with larger TD error is preferentially selected for updating.

[0225] Traditional converter batching optimization methods often have the problem of being easily trapped in local optimal solutions and lack sufficient exploration to avoid this situation. When selecting actions, TD-3 will directly add fixed-scale Gaussian noise to the output of the Actor network.

[0226] (1.34)

[0227] To ensure a certain degree of exploration. However, a fixed noise scale is difficult to adapt to the entire training process, which may lead to insufficient exploration in the early stages of training or interference with convergence in the later stages. To address this problem, this section proposes a noise adjustment module (NA), which enables the noise scale to be adaptively adjusted based on the strategy stability and the number of training steps. This mechanism mainly includes the following two steps:

[0228] Adaptive adjustment based on reward variance:

[0229] After each training round (set to 5 rounds), the agent calculates the average reward variance of the most recent rounds and uses it to measure the stability of the strategy. When the variance is low, noise is added to avoid falling into local optimality, and when the variance is high, noise is reduced to enhance the utilization effect:

[0230] (1.35)

[0231] in, Represents the average reward over several rounds.

[0232] Dynamic attenuation of noise with the number of training steps:

[0233] To ensure that the agent network converges to a stable optimal strategy in the later stages of training, the noise needs to increase with the number of training steps. Exponential decay. Finally, the strategy expression after noise adjustment is:

[0234] (1.36)

[0235] (1.37)

[0236] in, is the noise attenuation rate, is the number of training steps. In the early stages of training, due to the small number of steps and large noise, it is conducive to full exploration; while in the later stages of training, the noise gradually decays, making the strategy stable.

[0237] Next, the results of the batching generated by the method of the present invention are compared with the batching results generated by other methods.

[0238] The experimental data is derived from real factory data. To verify the effectiveness of the methods in this chapter, the production of HRB400E steel is used as an example. A total of eight raw materials are required, and their specific parameters are shown in Tables 1 and 2. Table 1 shows the upper and lower limits and control points of each element content in the finished product process. Table 2 shows the raw material price and the content of each element.

[0239] Table 1 Upper and lower limits of various element contents in HRB400E

[0240] Element symbols Element Name Lower limit (%) Upper limit (%) Optimal value (control point) C carbon 0.21 0.25 0.23 Si silicon 0.15 0.40 0.17 Mn manganese 1.35 1.45 1.37 P phosphorus 0 0.045 - S sulfur 0 0.045 -

[0241] Table 2 Content of various elements in raw materials

[0242]

[0243] To verify the effectiveness of the I-TD3 method, the evaluation indicators include raw material cost, raw material usage, total raw material usage and product quality. The average value of the experimental results of 100 heats is taken as the final result of each indicator. Product quality reflects the sum of the absolute values ​​of the differences between the content of each element in the finished product and the optimal control point. The smaller the product quality, the better the product quality.

[0244] The experimental parameters are set as follows: the number of training rounds is 1×10 5 The maximum number of iterations per round is 500, the batch size is 512, the soft update smoothing factor is 0.005, the discount factor is 0.7, and the learning rate of the evaluation network and the policy network is 3×10 -4 ,The activation function is ReLU function, the number of neurons in the input layer is 256, and three hidden layers are used, and the number of neurons in each layer is 128.

[0245] To verify the effectiveness of the I-TD3 model, this paper conducts experimental comparisons with a variety of comparative models. A brief description of each model is as follows:

[0246] MOEDO model: This model is a multi-objective optimization algorithm based on exponential distribution optimization. It combines non-dominated sorting and crowding distance to ensure the diversity and uniform distribution of solutions, and introduces an integrated information feedback mechanism to dynamically balance the model's global exploration and local development capabilities.

[0247] DDQN model: This model introduces two Q networks, one for action selection and the other for Q value estimation, to reduce the Q value overestimation problem in DQN.

[0248] A-TD3 model: This model introduces an asynchronous parallel mechanism based on TD3, optimizes the learning efficiency of the continuous action space through an adaptive update strategy, uses a parallel mechanism to improve the convergence speed, and designs two adaptive weight functions based on off-policy learning to dynamically adjust the weights of the local intelligent agent network.

[0249] From Table 3 and Figure 7 As can be seen, the I-TD3 model outperforms other models in terms of total raw material usage, raw material cost control, and product quality. Compared to MOEDO, DDQN, and A-TD3, the batching plan generated by I-TD3 significantly reduces raw material consumption while optimizing overall costs. This optimization demonstrates the model's more precise batching strategy, adaptively adjusting the proportions of various raw materials based on historical data and current heat demand, ensuring optimal raw material allocation and avoiding unnecessary waste. This is due to the introduction of the similarity experience replay module and the objective function weight adjustment module in I-TD3. By selecting experiences relevant to the current state, the agent can better learn the optimal batching plan, thereby reducing raw material usage. Simultaneously, the rate of change of each objective function is monitored in real time to avoid over-optimization of certain objectives while ensuring adequate attention to other objectives, thus balancing cost and quality. Through this approach, I-TD3 is able to significantly reduce batching costs while ensuring product quality. In terms of product quality, I-TD3 achieved the best performance, reaching 0.0883, significantly outperforming other models. This is due to I-TD3's noise adjustment module, which adaptively adjusts the noise scale, allowing the agent to fully explore in the early stages of training and gradually converge to the optimal strategy in the later stages. This gradual attenuation of noise helps the model avoid local optimal solutions, ultimately improving ingredient quality.

[0250] In summary, the I-TD3 model significantly improves the efficiency and quality of alloy batch optimization by introducing objective function weight adjustment, similarity experience replay, and noise adjustment modules. Compared with other traditional optimization algorithms, I-TD3 demonstrates superior performance in optimizing raw material usage, reducing costs, and improving product quality.

[0251] Table 3 Comparison of experimental results of various methods

[0252] index MOEDO DDQN A-TD3 I-TD3(Ours) Silicon manganese alloy (kg) 3185.4 2960.5 2801.5 2743.8 Medium carbon ferromanganese (kg) 142.3 115.7 58.2 34.7 Vanadium nitrogen alloy (kg) 24.8 18.2 15.4 14.5 Aluminum wire (kg) 32.7 27.5 14.1 9.2 High carbon ferrochrome (kg) 0 48.3 20.1 18.9 High nitrogen ferrovanadium (kg) 0 0 0 0 Recarburizer (kg) 18.6 105.9 93.4 98.6 Silicon manganese alloy ball (kg) 0 62.4 70.2 78.3 Total amount of raw materials G (kg) 3403.8 3338.5 3063.6 2998 Raw material cost F (yuan) 27289.44 25506.87 22479.34 21783.33 Product quality Z 0.1792 0.1522 0.0912 0.0883

[0253] To further explore the contributions of each module in the I-TD3 model, this section designed and conducted ablation experiments, gradually removing key modules from the complete model and observing their impact on model performance. Five model versions were designed for the experiment: the TD3 model, the I-TD3 w / o OFWA model, the I-TD3 w / oSER model, the I-TD3 w / oNA model, and the complete I-TD3 model. The ablation models are described below:

[0254] TD3 model: This model is a commonly used deep reinforcement learning method. All modules are removed. The weight of each objective function in the reward function is 1 / 3. The experience pool sampling uses priority experience replay without similar experience selection. The noise is fixed-scale Gaussian noise.

[0255] I-TD3w / oOFWA model: This model removes the objective function weight adjustment module and uses static weights. The weight of each objective function in the reward function is set to 1 / 3. Only the noise mechanism and experience replay are improved based on the TD3 algorithm.

[0256] I-TD3w / oSER model: This model removes the similarity experience replay module. The experience pool sampling adopts priority experience revisiting without similar experience selection. It only improves the noise mechanism and reward function weight based on the TD3 algorithm.

[0257] I-TD3w / oNA model: This model removes the noise adjustment module and selects fixed-scale Gaussian noise for exploration and utilization. It only performs experience revisit and reward function weight improvement based on the TD3 algorithm.

[0258] Table 4 Comparison of experimental results of various methods

[0259] index TD3 I-TD3w / oOFWA I-TD3w / oSER I-TD3w / oNA I-TD3(Ours) Silicon manganese alloy (kg) 2855.2 2790.1 2805.3 2747.6 2743.8 Medium carbon ferromanganese (kg) 87.6 56.3 65.4 49.2 34.7 Vanadium nitrogen alloy (kg) 15.8 15.6 17.2 15.1 14.5 Aluminum wire (kg) 21.3 15.4 19.1 13.6 9.2 High carbon ferrochrome (kg) 29.7 20.3 25.1 22.4 18.9 High nitrogen ferrovanadium (kg) 0 0 0 0 0 Recarburizer (kg) 101.2 95.4 102.3 98.1 98.6 Silicon manganese alloy ball (kg) 45.1 65.2 72.1 69.3 78.3 Total amount of raw materials G (kg) 3155.9 3034.5 3146.7 3009.8 2998 Raw material cost F (yuan) 23677.16 22472.56 22951.88 22017.15 21783.33 Product quality Z 0.1049 0.0965 0.1027 0.097 0.0883

[0260] From Table 4 and Figure 8As can be seen, I-TD3 outperforms other ablation models in terms of raw material consumption, cost control, and product quality. Compared to the ingredient solutions generated by other ablation models, the resulting solution effectively reduces total raw material usage and significantly lowers costs. This optimization demonstrates that the complete I-TD3 model leverages the synergy of its modules, enabling the agent network to more precisely adjust the ratio of different raw materials, ensuring a balance between cost efficiency and quality. The TD3 model, as the baseline, employs conventional deep reinforcement learning methods. Each objective function in the reward function is weighted as 1, using prioritized experience replay without any improvement modules. In contrast, the I-TD3w / oOFWA model removes the objective function weight adjustment module and uses static weights (all weighted as 1), making it incapable of flexible adjustment. Although this model reduces the usage of some raw materials (such as silicon-manganese alloy and aluminum wire), its overall performance does not significantly outperform the TD3 model. The use of silicon-manganese alloy decreased from 2855.2 kg to 2790.1 kg, but the overall raw material cost remained high (22,472.56 yuan), and the product quality was 0.0965, failing to achieve optimal performance. The I-TD3w / oSER model removed the similarity experience replay module and adopted traditional priority experience replay. This model performed relatively poorly in terms of total raw material usage and cost. Although the use of silicon-manganese alloy (2805.3 kg) was close to the I-TD3w / oOFWA model, the improvement in the use of other raw materials, such as medium-carbon ferromanganese and aluminum wire, was not significant. In particular, the total raw material usage was slightly reduced compared to the TD3 model, but still not optimal compared to other models. The I-TD3 w / oNA model removes the noise adjustment module and uses fixed-scale Gaussian noise to achieve a balance between exploration and exploitation. This model achieves significant improvements in both raw material usage and cost, reducing total raw material usage to 3009.8 kg and raw material costs to 22,017.15 yuan. This demonstrates that the introduction of the noise adjustment module significantly contributes to model performance, particularly in terms of resource utilization and cost optimization. The complete I-TD3 model combines the objective function weight adjustment module, the similarity experience replay module, and the noise adjustment module. This model achieves optimal performance across all key metrics, with a total raw material usage of 2998 kg, a raw material cost of 21,783.33 yuan, and a product quality of 0.0883. These results demonstrate that by incorporating these three modules, the I-TD3 model significantly optimizes raw material usage, cost control, and product quality, outperforming other ablation models.

[0261] In summary, the objective function weight adjustment module, similarity experience replay module, and noise adjustment module play a significant role in improving the performance of the I-TD3 model. The objective function weight adjustment module flexibly adjusts the weights of various objectives to optimize resource allocation; the similarity experience replay module improves learning efficiency by prioritizing experiences similar to the current task; and the noise adjustment module effectively balances exploration and exploitation, making the agent more stable and efficient when exploring new strategies. Therefore, the various modules of the I-TD3 model work together to improve the overall performance of the model.

[0262] In order to explore the impact of the discount factor on the model, a parameter sensitivity analysis experiment was conducted.

[0263] In the TD error formula, the discount factor acts on the Q value of the future state to adjust the current decision's dependence on future rewards. The larger the discount factor, the more the agent network focuses on future rewards; while the smaller the discount factor, the more the agent network tends to focus on immediate rewards. As shown in Table 5 and Figure 9 As shown in the figure, the value of the discount factor affects the model training results. The total raw material usage, raw material cost, and product quality first increase and then decrease with increasing discount factors. For example, the total raw material usage reaches the optimal value when the discount factor is 0.7, while it is suboptimal when it is 0.6. This is because when the discount factor is small, the agent network focuses more on immediate feedback, which may lead to local optimal solutions or overfitting in policy learning. While larger discount factors can increase the focus on long-term returns, in complex dynamic environments, long-term reliance on future rewards may lead to slow policy updates and even unstable learning. Therefore, a discount factor of 0.7 was selected as the final experimental parameter.

[0264] Table 5 The impact of discount factors on the model

[0265]

[0266] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the above method when executed by a processor.

[0267] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of any one of the methods described above.

[0268] Fourthly, refer to Figure 2 The present invention discloses a converter charge data generation and optimization system, which is characterized by comprising:

[0269] A data acquisition module is used to obtain an original industrial data set of the converter batching process and preprocess the original industrial data set to obtain preprocessed training data. The original industrial data set includes multi-dimensional sensor data and corresponding missing batching quality parameters;

[0270] The pre-training module is used to input the pre-processed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity of the latent variable distribution to the standard Gaussian distribution through the KL divergence loss;

[0271] A sample generation module is used to input the latent variables and random noise vectors into a generator to generate simulated samples, and a discriminator is used to distinguish between real samples and generated samples. A forced discriminator module is introduced. The forced discriminator performs compliance judgment on the generated samples based on the prior constraints defined in the converter metallurgical process knowledge base, converts the judgment results into additional loss terms, and jointly optimizes the parameters of the generator and discriminator.

[0272] The data generation module is used to generate data for missing ingredient parameters using the trained FD-VEGAN model and output a complete ingredient dataset that meets the process constraints;

[0273] The batching data optimization module is used to construct the cost objective function, quality objective function and resource consumption objective function, and to construct the constraint conditions of the chemical element composition content of the target finished product. It uses the improved TD3 algorithm to solve and optimize the converter batching data to generate the optimized converter batching data.

[0274] Furthermore, in the pre-training module:

[0275] The encoder encodes the input true sample into a mean vector and a variance vector;

[0276] On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss;

[0277] The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss;

[0278] The encoder optimizes its performance based on the regularization loss and reconstruction loss.

[0279] Furthermore, in the sample generation module:

[0280] The generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples, generating the discriminator loss;

[0281] Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss;

[0282] The generator is adversarially trained based on the discriminator loss and the additional loss to generate high-quality samples.

[0283] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.

[0284] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0285] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0286] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating and optimizing converter charge data, characterized in that: include: Step S1: obtaining an original industrial data set of the converter batching process and preprocessing the original industrial data set to obtain preprocessed training data. The original industrial data set includes multidimensional sensor data and corresponding missing batching quality parameters. Step S2: Input the preprocessed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity between the latent variable distribution and the standard Gaussian distribution through the KL divergence loss; Step S3: Input the latent variables and random noise vectors into the generator to generate simulated samples. The discriminator distinguishes between real samples and generated samples. A forced discriminator module is introduced. The forced discriminator performs compliance judgment on the generated samples based on the prior constraints defined in the converter metallurgical process knowledge base, and converts the judgment results into additional loss terms. Step S4: Construct a loss function, including generative adversarial loss, KL divergence loss, additional loss of the forced discriminator, and reconstruction loss. Jointly optimize the generator and discriminator parameters through backpropagation to train the generative model. Step S5: Use the model trained in step S4 to generate data for the missing ingredient parameters and output a complete ingredient data set that meets the process constraints; Step S6: Construct a cost objective function, a quality objective function, and a resource consumption objective function, and construct constraints on the chemical element composition content of the target finished product. Use the improved TD3 algorithm to solve it. The improved TD3 algorithm optimizes the converter batching data by introducing objective function weight adjustment, similarity experience playback, and noise adjustment modules to generate optimized converter batching data.

2. The method according to claim 1, characterized in that Step S2 includes: The encoder encodes the input true sample into a mean vector and a variance vector; On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss; The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss; The encoder optimizes its performance based on the regularization loss and reconstruction loss.

3. The method according to claim 1, characterized in that Step S3 includes: The generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples, generating the discriminator loss; Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss; The generator is adversarially trained based on the discriminator loss and an additional loss.

4. The method according to claim 2, characterized in that Regularization losses include: ; ; in, is the probability distribution characteristic of the latent variable z obtained based on a specific input x, is the posterior distribution given by the encoder, is the prior distribution of the latent space, is the definition of KL divergence, where the subscripts prior and post refer to the prior distribution and posterior distribution generated from the encoder, respectively. and denote the mean and standard deviation of the prior distribution of the standard Gaussian distribution, and denote the mean and standard deviation of the posterior distribution of the encoder output, respectively.

5. The method according to claim 4, characterized in that The reconstruction loss includes: ; is a multivariate Gaussian distribution.

6. The method according to claim 4, characterized in that In the basic VAE model, the total loss of regularization loss and reconstruction loss is: ; in is the reconstruction loss, is the regularization loss.

7. The method according to claim 4, characterized in that In the dynamic β-VAE model, the total loss of regularization loss and reconstruction loss is: ; in is the reconstruction loss of the nth round, is the regularization loss of the nth round; is a parameter that changes dynamically according to the round.

8. The method according to claim 1, characterized in that The loss of the generator in step S3 is: ; in, and are weighting coefficients, which are used to adjust the latent variables and noise The loss of generated samples, is the weight coefficient of reconstruction loss, is the reconstruction loss, is the discriminator's response to the latent variable generated from the encoder The score of the generated samples, is the discriminator's response to the noise The scores of the generated samples.

9. The method according to claim 8, characterized in that The loss function of the discriminator is: ; in, is the score of the discriminator on the real sample, is the discriminator's response to the latent variable generated from the encoder The score of the generated samples, is the discriminator's response to the noise The scores of the generated samples.

10. The method according to claim 1, characterized in that Step S6 includes: Monitor the progress of each objective function in real time and dynamically adjust the weight according to the relative rate of change of the objective function; Add historical experience samples and calculate the cosine similarity between the current state and historical experience; Noise is added during the solution process, and the noise scale is adaptively adjusted according to the reward variance and the number of training steps.

11. The method according to claim 10, characterized in that Adaptive adjustment of noise scale based on reward variance and number of training steps includes: After each training round, the average reward variance of the most recent round is calculated. When the variance is lower than the preset value, the noise is increased; when the variance is higher than the preset value, the noise is reduced. The noise also decays exponentially with the number of training steps.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

13. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 11.

14. A converter charge data generation and optimization system, characterized in that: include: The data acquisition module is used to obtain the original industrial data set of the converter batching process and preprocess the original industrial data set to obtain preprocessed training data. The original industrial data set contains multi-dimensional sensor data and the corresponding missing batching quality parameters; The pre-training module is used to input the pre-processed training data into the encoder to generate the mean vector and variance vector of the latent variable distribution, and constrain the similarity of the latent variable distribution to the standard Gaussian distribution through the KL divergence loss; The sample generation module is used to input latent variables and random noise vectors into the generator to generate simulated samples. The discriminator distinguishes between real samples and generated samples. The forced discriminator module is introduced. The forced discriminator verifies the compliance of generated samples based on the prior constraints defined in the converter metallurgical process knowledge base. The judgment results are converted into additional loss terms. The generator and discriminator parameters are jointly optimized. The generator performs adversarial training based on the discriminator loss and the additional loss to generate high-quality samples. The data generation module is used to generate data for missing ingredient parameters using the trained FD-VEGAN model and output a complete ingredient dataset that meets the process constraints; The batching data optimization module is used to construct the cost objective function, quality objective function and resource consumption objective function, and to construct the constraint conditions of the chemical element composition content of the target finished product. The improved TD3 algorithm is used to solve the problem. The improved TD3 algorithm optimizes the converter batching data by introducing objective function weight adjustment, similarity experience replay and noise adjustment modules to generate optimized converter batching data.

15. The system according to claim 14, wherein: In the pre-training module: The encoder encodes the input true sample into a mean vector and a variance vector; On the one hand, the above vectors are combined into a latent variable input generator, and on the other hand, the KL divergence is compared with the mean and variance vector of the standard Gaussian distribution to generate the regularization loss; The reconstruction loss is used to quantify the difference between the reconstructed sample and the real sample to generate the reconstruction loss; The encoder optimizes its performance based on the regularization loss and reconstruction loss.

16. The system according to claim 14, wherein: In the sample generation module, the generator receives random noise and the latent variable output by the encoder as input, generates corresponding noise samples and reconstructed samples, and uses the discriminator to distinguish the generated samples from the real samples to generate the discriminator loss; Introducing a forced discriminator based on prior knowledge and using its discrimination result as additional loss; The generator is adversarially trained based on the discriminator loss and the additional loss to generate high-quality samples.

Citation Information

Patent Citations

  • A method for augmenting unbalanced time-series data for industrial fault diagnosis

    CN112328588B

  • Cloth defect image generation system and method based on improved variational auto-encoder network

    CN114067168A

  • Training method and application method of distributed medical image processing model

    CN117058091A