Vacuum degassing furnace control method, system and device
By combining a dual-path architecture of mechanistic model and deep learning model, the problems of poor parameter adaptability and insufficient multi-objective optimization in vacuum degassing furnace control are solved, thereby improving production efficiency and product quality and providing precise time, temperature and material optimization.
Patent Information
- Application Number
- CN202511368430.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-12-19
AI Technical Summary
Existing vacuum degassing furnace control methods rely on operator experience and fixed parameters, which cannot adapt to fluctuations in raw materials and changes in equipment status. This leads to unstable production rhythm and product quality, lacks multi-process collaboration and multi-objective comprehensive optimization capabilities, and makes it difficult to achieve accurate prediction and control.
A dual-path architecture combining mechanistic models and deep learning models is adopted. The mechanistic model provides reliable predictions based on physical characteristics, while the deep learning model provides accurate predictions. Furthermore, the physical information neural network coordinates the various sub-networks to achieve multi-objective collaborative optimization.
It improves the production efficiency and product quality of vacuum degassing furnaces, achieves precision in time prediction, temperature control and material optimization, adapts to changes in production conditions, and ensures the flexibility and reliability of the system.
Smart Images

Figure CN121165481A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vacuum degassing furnace control in the field of steel metallurgy, and in particular to a vacuum degassing furnace control method, system and device. BACKGROUND
[0002] As a key link between the converter and the ladle refining (LF) furnace, the vacuum degassing (VD) furnace undertakes important metallurgical tasks such as decarburization, deoxidation, alloying, slagging desulfurization, and degassing (nitrogen, hydrogen), etc. The control effect is directly related to the quality of molten steel and production efficiency. With the increasing requirements of the steel industry on product quality, energy saving and emission reduction, and intelligent level, precise prediction and optimal control of the VD furnace process have become a key requirement for the development of the industry.
[0003] The existing VD furnace control method mainly relies on the experience of operators and simple empirical formulas. However, due to the use of fixed parameters and simplified assumptions in the above method, it cannot fully adapt to actual production conditions such as raw material fluctuations and equipment state changes, and therefore has the following shortcomings:
[0004] On the one hand, the traditional mechanism model has fixed parameters and poor adaptability, and it is difficult to dynamically adjust according to real-time production data, resulting in insufficient prediction accuracy of key parameters such as processing time and temperature changes, affecting production rhythm and product quality stability.
[0005] On the other hand, the existing method lacks modeling capability for complex nonlinear relationships and relies mainly on linear or locally linearized models, which cannot fully capture the deep coupling relationship between variables in the VD furnace process, resulting in inaccurate material addition amount and timing, causing resource waste and cost increase.
[0006] On the other hand, the existing control strategy often focuses on the optimization of a single process or a single target, lacks multi-process coordination and multi-target comprehensive optimization capability, and is difficult to balance time, energy consumption and economy while ensuring that the composition of molten steel meets the standard, and the overall control effect is limited.
[0007] Therefore, under this background, how to provide a vacuum degassing furnace control method with multi-target collaborative optimization capability, to realize precise prediction and control in time prediction, temperature control, material optimization, etc., and significantly improve production efficiency and product quality, is a technical problem to be solved. SUMMARY
[0008] In view of the above problems of the prior art, the present application provides a vacuum degassing furnace control method, system and device to provide a vacuum degassing furnace control method with multi-target collaborative optimization capability, to realize precise prediction and control in time prediction, temperature control, material optimization, etc., and significantly improve production efficiency and product quality.
[0009] To achieve the above object, the first aspect of the present application provides a vacuum degassing furnace control method, comprising the following steps:
[0010] Collecting initial parameters;
[0011] Selecting a model route; wherein the model route comprises a mechanism model route and a deep learning model route;
[0012] According to the initial parameters, the key parameters of each process of vacuum degassing are predicted by the selected model route;
[0013] According to the prediction results, the key operation of the vacuum degassing furnace is adjusted.
[0014] From the above, by containing two routes of mechanism model and deep learning model, and by selecting different routes to make the prediction method adapt to different conditions and working conditions, more accurate prediction results are obtained; the mechanism model route and the deep learning model route of the present application predict the key parameters in various processes in the metallurgical process, and adjust according to the prediction results, so that the control of the vacuum degassing furnace is more flexible and controllable, thereby coordinating the control of multi-objective optimization; thereby improving production efficiency and product quality.
[0015] As a possible implementation manner of the first aspect, using the mechanism model route for prediction comprises:
[0016] The initial parameters are input into the mechanism model, and the mechanism model comprises a decarburization mechanism model, a deoxidation mechanism model, an alloying mechanism model, a desulfurization and slagging model, a degassing mechanism model and a temperature mechanism model;
[0017] The mechanism model is a series of constraint equations established according to corresponding physical properties, and by solving the constraint equations, the key parameters of the vacuum degassing furnace are obtained;
[0018] Using the deep learning model route for prediction comprises:
[0019] The initial parameters are preprocessed and feature engineered;
[0020] The processed data are input into the deep learning model; the deep learning network is controlled by a physical information neural network master model to constrain each sub-network by physical laws, and the sub-networks comprise a decarburization prediction network, a temperature prediction network, a deoxidation prediction network, an alloying prediction network, a desulfurization and slagging prediction network, and a degassing prediction network;
[0021] The sub-networks predict the corresponding key parameters through the deep learning network;
[0022] The physical information neural network also coordinates and optimizes the prediction results of the sub-networks.
[0023] From the above, the dual route architecture design makes the system have good flexibility and reliability. When the data is insufficient or needs to be responded quickly, by deploying a mechanism model based on physical characteristics, it is ensured that the mechanism model can provide reliable prediction based on physical principles. In the case of large amount of data, when the data is sufficient, the deep learning model can provide more accurate prediction. The application of physical information neural network ensures that the model prediction result conforms to the basic principles of physics, and avoids unreasonable results that may be produced by pure data-driven methods.
[0024] As a possible implementation manner of the first aspect, using the mechanism model route for prediction further includes: performing global optimization on the prediction result of the mechanism model, the global optimization being constrained by component requirements, material supply inventory, and process processing time; and monitoring whether the prediction result of each transfer parameter and intermediate / final output violates a constraint condition in real time and adjusting model parameters, handling parameter out-of-range, calculation divergence, and logic conflict.
[0025] From the above, by introducing linear programming of components, inventory, time, etc. through global optimization, and monitoring transfer parameters and prediction results and handling exceptions, the problem of multi-objective collaborative optimization is solved.
[0026] As a possible implementation manner of the first aspect, using the deep learning model route for prediction further includes: updating model parameters of the deep learning model according to the new data.
[0027] From the above, the model parameters are continuously updated using new data generated in the production process, so that the model has self-adaptive ability and can continuously optimize the model parameters as the production conditions change, thereby maintaining the prediction accuracy. This continuous learning ability is of great significance for dealing with actual production problems such as raw material fluctuations and equipment aging.
[0028] As a possible implementation manner of the first aspect, the selection model route includes: selecting the deep learning model route when the amount of historical data accumulated by the system exceeds a threshold; and selecting the mechanism model route when the amount of historical data is insufficient.
[0029] From the above, an automatic model selection method is provided according to the amount of historical data, to ensure that the most suitable model is always used for prediction.
[0030] As a possible implementation manner of the first aspect, the selection model route can select to use the mechanism model route and the deep learning model route simultaneously; when the mechanism model route and the deep learning model route are used simultaneously, a consistency index of the prediction result is calculated, and if the deviation exceeds a threshold, an abnormality analysis process is triggered, including: identifying the reason for the deviation, and outputting the result to a vacuum degassing furnace control system.
[0031] From the above, the two models are allowed to run simultaneously and the results are compared, and when the deviation is large, abnormal analysis is triggered, which can provide self-fault diagnosis capability and ensure the safety robustness of the system. The two models can verify each other to improve the overall reliability of the system.
[0032] As a possible implementation of the first aspect, the initial parameters include at least one of the following: molten steel weight, initial temperature, initial composition, ladle age, bottom blowing flow, target composition;
[0033] The key parameters include at least one of the following: process processing time, temperature curve, material addition scheme, and final composition;
[0034] The key operations include at least one of the following: adjusting the material addition scheme, adjusting the machine working mode, speed, and power.
[0035] From the above, by limiting the initial parameters, key parameters and key operations described above, the metallurgical key links and processes of the vacuum degassing furnace are covered, which provides a direct processing method for the actual implementation of the scheme.
[0036] The second aspect of the application provides a vacuum degassing furnace control system, comprising: an input parameter layer, a model selection controller, a mechanism model, a deep learning model and an output layer;
[0037] The input parameter layer is configured to preprocess the input initial parameters;
[0038] The model selection controller is configured to select the mechanism model and / or the deep learning model;
[0039] The mechanism model includes at least one of the following mechanism models: decarburization mechanism model, deoxidation mechanism model, alloying mechanism model, desulfurization and slagging model, degassing mechanism model, temperature mechanism model; It also includes a material selection linear programmer, and a model coordination and timing controller;
[0040] The deep learning model includes a physical information neural network master model and at least one of the following sub-networks: decarburization prediction network, deoxidation prediction network, alloying prediction network, desulfurization prediction network, degassing prediction network, temperature prediction network; It also includes an online learning and parameter updating module;
[0041] The output layer is configured to integrate the prediction results of the two model routes.
[0042] The third aspect of the application provides a vacuum degassing furnace control device, comprising:
[0043] The data acquisition module is configured to acquire initial parameters;
[0044] a model selection module configured to select a model route; wherein the model route comprises a mechanism model route and a deep learning model route;
[0045] a parameter prediction module configured to predict key parameters of the vacuum degassing furnace according to the initial parameters through the selected model route;
[0046] a control adjustment module configured to adjust key operations of the vacuum degassing furnace according to the prediction result.
[0047] The fourth aspect of the present application provides a computing device, comprising a processor and a memory having program instructions stored thereon, the program instructions causing the processor to execute the vacuum degassing furnace control method of any one of the first aspect when executed by the processor.
[0048] The fifth aspect of the present application provides a computer-readable storage medium having program instructions stored thereon, the program instructions causing the computer to execute the vacuum degassing furnace control method of any one of the first aspect when executed by the computer.
[0049] The sixth aspect of the present application provides a computer program product comprising program instructions, the program instructions causing the computer to execute the vacuum degassing furnace control method of any one of the first aspect when executed by the computer. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 is a flowchart of the vacuum degassing furnace control method provided by the first embodiment of the present application;
[0051] Figure 2a is a flowchart of the vacuum degassing furnace control method provided by the second embodiment of the present application;
[0052] Figure 2b is a dual-route architecture diagram provided by the second embodiment of the present application;
[0053] Figure 2c is a mechanism model running framework diagram provided by the second embodiment of the present application;
[0054] Figure 2d is a deep learning neural network architecture diagram provided by the second embodiment of the present application;
[0055] Figure 2e is a ResNet residual block structure diagram provided by the second embodiment of the present application;
[0056] Figure 3 is a schematic diagram of the vacuum degassing furnace control system provided by the second embodiment of the present application;
[0057] Figure 4is a schematic diagram of a vacuum degassing furnace control device provided by an embodiment of the present application.
[0058] Figure 5 is a structural schematic diagram of a computing device provided by an embodiment of the present application.
[0059] It should be understood that in the above structural schematic diagram, the size and shape of each block diagram are only for reference and should not constitute an exclusive interpretation of the embodiments of the present application. The relative position and inclusion relationship between the block diagrams presented by the structural schematic diagram are only used to represent the structural association between the block diagrams, and are not intended to limit the physical connection mode of the embodiments of the present application. DETAILED DESCRIPTION
[0060] The technical solutions provided by the present application will be further described below in conjunction with the drawings and embodiments. It should be understood that the system structure and business scenarios provided in the embodiments of the present application are mainly used to illustrate possible implementation modes of the technical solutions of the present application, and should not be interpreted as the only limitation of the technical solutions of the present application. Those skilled in the art can know that the technical solutions provided by the present application are also applicable to similar technical problems as the system structure evolves and new business scenarios appear.
[0061] It should be understood that the vacuum degassing furnace control scheme provided by the embodiments of the present application includes a vacuum degassing furnace control method, system, device, computing device, computer readable storage medium and computer program product. Since the principles of solving problems of these technical solutions are the same or similar, in the introduction of the following specific embodiments, some repeated parts may not be described again, but should be regarded as mutual reference between these specific embodiments, which can be combined with each other.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. If there is any inconsistency, the meaning explained in the specification or the meaning derived from the content described in the specification shall prevail. In addition, the terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application. In order to accurately describe the technical content in the present application, and in order to accurately understand the present application, before the specific embodiments are described, the terms used in the specification are first explained as follows:
[0063] 1) Physics-Informed Neural Networks (PINN): A method that embeds physical laws (usually in the form of partial differential equations or constraints) into neural networks as part of the loss function. It penalizes the network output for violating physical laws, making the learning results satisfy both data and physical principles, enhancing the model's interpretability and extrapolation ability.
[0064] 2) Long Short-Term Memory (LSTM): A special type of Recurrent Neural Network (RNN) that introduces "gate" mechanisms (input gate, forget gate, output gate) to selectively remember and forget information, addressing the vanishing or exploding gradient problem that traditional RNNs face when dealing with long sequences, suitable for modeling long-term dependencies in time series data.
[0065] 3) Residual Network (ResNet): A deep convolutional neural network architecture that introduces "skip connections" or "shortcuts" that allow data to bypass one or more layers and go directly to the next. This makes it possible to train very deep networks, as the structure alleviates the gradient vanishing problem, allowing the network to learn the identity mapping.
[0066] 4) Graph Neural Network (GNN): A class of neural networks that directly operate on graph-structured data. It aggregates and updates node features through a message passing mechanism between neighboring nodes, capturing complex relationships between nodes. Suitable for non-Euclidean data such as social networks, molecular structures, and knowledge graphs.
[0067] 5) Q-Network, Deep Q-Learning: Q-Network is a method that uses a neural network to approximate the Q function (state-action value function) in reinforcement learning. Deep Q-Learning (DQN) is an algorithm that combines Q-learning with deep neural networks, using techniques such as experience replay and target networks to stabilize training, allowing the agent to learn the optimal policy directly from high-dimensional inputs.
[0068] 6) Deep Belief Network (DBN): A generative model composed of multiple layers of stochastic latent variables, usually consisting of stacked Restricted Boltzmann Machines (RBM). It uses layer-by-layer unsupervised greedy pre-training to initialize weights, and then fine-tunes with supervised learning, which was once an effective deep learning method.
[0069] 7) Linear Programming (LP): A mathematical optimization method used to find the maximum or minimum value of a linear objective function under the constraints defined by linear equations or inequalities. It is applied to problems that require optimal decision-making, such as resource allocation, production planning, and transportation scheduling.
[0070] 8) Transformer: A deep learning model architecture based entirely on self-attention mechanisms. It abandons the recurrent and convolutional structures, allowing efficient parallel processing of sequence data and capturing global dependencies between elements within the sequence.
[0071] 9) Convolutional Neural Network (CNN): A type of neural network designed specifically for handling grid-like data such as images. It automatically extracts local features and achieves translation invariance through convolutional layers, pooling layers, etc.
[0072] 10) Bayesian Deep Learning: A method that combines Bayesian probability principles with deep learning. Instead of taking a single value for model parameters, it learns their probability distribution (posterior distribution), quantifying the uncertainty of model predictions (epistemic uncertainty).
[0073] 11) Bayesian Optimization: A sequential strategy for optimizing black-box functions (high evaluation cost, no analytical expression). By constructing a surrogate model (such as Gaussian process) to model the objective function, and using acquisition functions (such as EI) to guide the next evaluation point, it finds the global optimal solution with few evaluations, often used for hyperparameter tuning.
[0074] 12) Neural Architecture Search (NAS): A subfield of AutoML that aims to automatically discover and design high-performance neural network architectures instead of relying on manual design. It finds the optimal network structure through three elements: search space, search strategy, and performance evaluation strategy.
[0075] 13) Elastic Weight Consolidation (EWC): A continuous learning algorithm that addresses the problem of catastrophic forgetting. EWC estimates the importance of old task parameters through the Fisher information matrix and penalizes changes to important parameters when learning new tasks, thereby remembering old knowledge.
[0076] 14) Model-Agnostic Meta-Learning (MAML): A model-agnostic meta-learning framework that aims to find a set of model initial parameters. This allows the model to quickly adapt to new tasks with only a small number of samples and a few gradient updates.
[0077] 15) Mixed-Integer Linear Programming (MILP): An extension of linear programming where some decision variables are restricted to integers (or binary). This increases the complexity of the problem (NP-hard), but can model a wide range of discrete decision-making problems such as equipment selection, path planning, etc.
[0078] 16) Branch and Bound (B&B): An algorithmic framework for solving optimization problems, especially discrete combinatorial optimization like MILP. It efficiently finds the optimal solution by systematically "branching" to enumerate the solution space and "bounding" to prune subspaces that cannot be optimal.
[0079] 17) Interior-Point Method: Interior-Point Method is a powerful algorithm for solving linear or nonlinear convex optimization problems, which approximates the optimal solution from the interior of the feasible region. Its "differentiable implementation" refers to the mathematical technique that makes the optimization solving process computable gradient, so that it can be embedded in neural networks and trained end-to-end.
[0080] 18) Recursive Least Squares (RLS): An online parameter estimation algorithm (parameter identification algorithm). RLS does not need to save all historical data, and dynamically updates the model parameter estimation through recursive formula, so that it can adapt to system changes in real time.
[0081] 19) Smooth L1 Loss: A loss function commonly used in regression tasks, combining the advantages of L1 loss (absolute error) and L2 loss (mean square error). It is smooth like L2 loss when the error is small, and is insensitive to outliers like L1 loss when the error is large, and is more robust.
[0082] 20) Policy Gradient Methods: A class of reinforcement learning algorithms that directly optimize the policy function. It calculates the gradient of the performance indicator (expected return) with respect to the policy parameters, and updates the parameters along the gradient direction to improve the policy, which is suitable for continuous action space and high-dimensional problems.
[0083] 21) Dropout Sampling: A regularization technique used in neural network training. It randomly and temporarily "drops out" (zeros) neurons in the network with a certain probability, preventing overfitting between neurons. In Bayesian deep learning, Dropout is enabled during testing, which can be used to approximate Bayesian inference and estimate model uncertainty.
[0084] The vacuum degassing furnace control scheme provided by the embodiments of the present application can collect initial parameters, select a model route, predict key parameters of the vacuum degassing furnace through the selected model route according to the initial parameters, and adjust the key operation of the vacuum degassing furnace according to the prediction result. This method can provide a vacuum degassing furnace control method with the ability of multi-target collaborative optimization, realize accurate prediction and control in time prediction, temperature control, material optimization and the like, and significantly improve production efficiency and product quality. The embodiments of the present application can be applied to steel metallurgy engineering or research in various industries and research fields, such as intelligent control, monitoring and automation of the vacuum degassing furnace. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0085] The first embodiment of the present application provides a vacuum degassing furnace control method, which will be described below in combination with Figure 1 , and the implementation of each step of the method will be specifically described, including steps S10-S40.
[0086] S10: Collect initial parameters.
[0087] In some embodiments, the initial parameter data is collected by a standardized interface of an upstream device, a real-time sensor or the like, and is transmitted to a computer (or various hosts, distributed systems) system for real-time processing and analysis. The collected data can also be preprocessed by, for example, taking an average value of multi-point measurement to improve the accuracy of the collected data.
[0088] In some embodiments, the method of the present application is installed and run on a computer (or various hosts, distributed systems) system in the form of a computer program.
[0089] In some embodiments, the initial parameters include at least one of the following: molten steel weight, initial temperature, initial composition, ladle age, bottom blowing flow, target composition. The molten steel weight determines the reference of the amount of all subsequent materials added, which can be received from the upstream converter through a standardized interface; the initial temperature can be preprocessed by taking an average value of multi-point measurement to improve the accuracy of the collected data; the initial composition includes the content of multiple elements such as carbon, silicon, manganese, phosphorus, sulfur, hydrogen and nitrogen, which can come from rapid analysis when the converter is tapped; the ladle age reflects the number of times the ladle is used and the heat preservation performance, which has an important influence on temperature prediction; the bottom blowing flow includes the set value of the bottom blowing argon flow at different stages of the entire process, which is a time series data; the target composition defines the quality requirements of the final product, which is the optimization target of the entire control process.
[0090] In some embodiments, the initial parameters also need to be preprocessed by outlier detection, missing value filling and standardization. For data beyond the normal range, the system will trigger an alarm and require manual confirmation.
[0091] S20: Selecting a model route; wherein the model route comprises a mechanism model route and a deep learning model route.
[0092] In some embodiments, the selection can be made by a model selection controller according to pre-set control instructions and current production status. The selection model route comprises: selecting a deep learning model route when the amount of historical data accumulated by the system exceeds a threshold value; and selecting a mechanism model route when the amount of historical data is insufficient.
[0093] In some embodiments, the model selection can be performed automatically under certain conditions, or manually according to actual needs.
[0094] In some embodiments, the selection model route can select to use both the mechanism model route and the deep learning model route simultaneously; when both the mechanism model route and the deep learning model route are used simultaneously, a consistency index of the predicted results is calculated, and if the deviation exceeds a threshold value, an abnormality analysis process is triggered, including: identifying the cause of the deviation, and outputting the results to the vacuum degassing furnace control system.
[0095] In some embodiments, the model selection controller evaluates the reliability of the current input parameters, and if a sensor failure or data anomaly is detected, it automatically switches to a more robust mechanism model route.
[0096] S30: According to the initial parameters, the key parameters of each process of vacuum degassing are predicted by the selected model route.
[0097] In some embodiments, the key parameters include at least one of the following: process processing time, temperature curve, material addition scheme, and final composition.
[0098] In some embodiments, in addition to the key parameters, pre-set model parameters are also included. The system first loads the corresponding model parameters from the parameter configuration file manager. The parameter configuration file is managed according to steel type, equipment number, and optimization version. The loading process includes parameter integrity check, value range verification, and version compatibility confirmation. For important parameters such as reaction rate constant and activation energy, the system compares with historical data, and if the deviation exceeds 30%, a parameter review process is triggered. After the parameter loading is completed, the system generates a parameter snapshot for subsequent tracing and analysis.
[0099] In some embodiments, the prediction using the mechanism model route comprises: inputting the initial parameters into a mechanism model, the mechanism model comprising a decarburization mechanism model, a deoxidation mechanism model, an alloying mechanism model, a desulfurization and slagging model, a degassing mechanism model, and a temperature mechanism model; the mechanism model is a series of constraint equations established according to corresponding physical properties, and by solving the constraint equations, the key parameters of the vacuum degassing furnace are obtained.
[0100] In some embodiments, the decarburization mechanism model can establish a reaction kinetics model based on the carbon-oxygen reaction mechanism for the decarburization process, and by solving the equation by the Runge-Kutta method, the decarburization time and the carbon content change curve are obtained. The deoxidation mechanism model can establish an equilibrium equation of Gibbs free energy with temperature change based on the thermodynamic equilibrium principle, calculate the deoxidation reaction equilibrium constant, and output the type of deoxidizer, the actual amount of deoxidizer added, and the deoxidation time. The alloying mechanism model can establish basic constraints based on the element balance principle, find the least cost alloy combination that meets the alloy melting kinetics model and satisfies the composition constraints, and output the optimal alloy ratio, addition sequence, and alloying time. The desulfurization and slagging model can control the basicity by a three-dimensional response surface of basicity-sulfur partition coefficient-temperature, determine the optimal basicity, calculate and obtain the slagging agent formula, predict the desulfurization rate and desulfurization time. The degassing mechanism model can establish a model based on the dissolution-diffusion-escape principle of gas in the steel liquid, determine the optimal vacuum degree change curve by multi-objective optimization of degassing rate-vacuum degree-energy consumption, and obtain the degassing time and the final hydrogen and nitrogen content prediction. The temperature mechanism model can constrain throughout the entire vacuum degassing process, establish a complete heat balance equation system, including modeling of chemical reaction heat, radiation heat dissipation, ladle heat dissipation, etc., and when the chemical reaction calculation of all processes is coupled, the final thermal state and temperature curve are output.
[0101] In some embodiments, the prediction using the mechanism model route further comprises: globally optimizing the prediction results of the mechanism model, the global optimization being constrained by composition requirements, material supply inventory, and process processing time; real-time monitoring whether the prediction results of the transfer parameters, intermediate / final output violate the constraint conditions and adjusting the model parameters, processing parameter out-of-bounds, calculation divergence, and logic conflicts.
[0102] In some embodiments, the global optimization can be performed by a material selection linear programer on the total cost objective function. A model coordinator can be responsible for the timing arrangement and parameter passing of each mechanism model, which employs a standardized interface to ensure that the output of an upstream process can be correctly passed as the input of a downstream process. A constraint checking module can be included to monitor the predicted results of each passing parameter, intermediate / final output in real time, and to adjust the constraint conditions, optimization objectives or model parameters immediately when a constraint violation is found. An exception handling mechanism can be included to deal with parameter out-of-bound, calculation divergence and logical conflict. For example, the end-point carbon content and decarburization time output by the decarburization model are passed as key parameters to the temperature model for heat balance calculation after the amount of CO gas generated by the model coordinator. When the decarburization model fails to converge, it can affect the downstream model to lack key inputs, and thus an alarm needs to be processed and the constraint conditions or parameter settings of the decarburization model need to be checked or prompted.
[0103] In some embodiments, the mechanism model also has a feedback mechanism to continuously collect actual production data, and to update the model parameters through a parameter identification algorithm, such as online estimation of model parameters by using the recursive least squares method.
[0104] In some embodiments, the use of the deep learning model route for prediction includes: pre-processing and feature engineering of the initial parameters; inputting the processed data into a deep learning model; the deep learning network is controlled by a physical information neural network master model to constrain each sub-network by physical laws, and the sub-networks include a decarburization prediction network, a temperature prediction network, a deoxidation prediction network, an alloying prediction network, a desulfurization and slagging prediction network, and a degassing prediction network; the sub-networks predict corresponding key parameters through the deep learning network; and the physical information neural network also coordinates and optimizes the prediction results of the sub-networks.
[0105] In some embodiments, the data used for deep learning model training includes historical data corresponding to each process stage, i.e., the weight of molten steel, initial temperature, initial composition, ladle age, bottom blowing flow, target composition of each process stage of decarburization, deoxidation, alloying, desulfurization and slagging, and desulfurization; and historical data of key parameters and intermediate process parameters measured during operation, such as steel type, deoxidizer type, temperature curve, slagging agent formula, sulfur partition coefficient, pressure of vacuum degassing furnace, vacuum pressure curve, etc.
[0106] In some embodiments, the deep learning model route first preprocesses the data, including standardization, mapping parameters of different dimensions to a unified scale, and dimensionless of physical quantities; feature encoding, such as using a learnable embedding vector to process discrete variables such as steel grade, and using one-dimensional convolution to extract local patterns of time series features such as bottom blowing flow patterns; data augmentation, expanding the training set by adding controlled noise and physically constrained perturbations; missing value filling, estimating and filling missing values according to the context and internal relations of the data.
[0107] In some embodiments, the physical information neural network master model can coordinate multiple sub-networks, responsible for feature fusion, physical constraint encoding, task allocation, and coordinated loss function. Among them, feature fusion can use a multi-layer perceptron structure, introduce residual connection and layer normalization; physical constraint encoding can embed physical laws as constraints into the network through a physical constraint encoder, such as using the conservation law constraint matrix as an item of the loss function to modify the loss function and guide the output results of the overall model; task allocation can dynamically allocate the weights of each sub-network using an attention mechanism. The coordinated loss function considers prediction accuracy, sub-network consistency, and physical constraints to coordinate the prediction results.
[0108] In some embodiments, the physical constraint encoding layer can be based on the principles of physical information neural network (PINN), such as by adding the residual of physical conservation laws (such as mass conservation, energy conservation, chemical equilibrium, phase equilibrium) as a penalty term to the loss function to constrain the output results of the sub-networks.
[0109] In some embodiments, the decarburization prediction network can use an LSTM architecture, introduce an attention mechanism to enhance attention to key moments, and output decarburization time, carbon content change curve, and decarburization rate. The temperature prediction network can use a ResNet architecture to output temperature curve and temperature change rate. The deoxidation prediction network can use GNN to model the complex reaction relationship between elements, represent elements and compounds as nodes of a graph, and reflect relationships as edges, and through a reinforcement learning module (Q-value network) to optimize the deoxidizer selection strategy, obtain the optimal type and amount of deoxidizer, and corresponding parameters such as deoxidization time. The alloying prediction network can use DBN to pre-train the probability distribution of the composition space (alloy composition) layer by layer, consider optimizing components and costs in the loss function, and find the optimal alloy ratio (amount of addition) through linear programming. The desulfurization and slagging prediction network can use the Transformer architecture to process the sequence evolution of slag components, and predict the time series of the slagging agent formula and the sulfur distribution coefficient. The degassing prediction network can use a CNN-LSTM hybrid architecture, extract local patterns of the vacuum curve through CNN, model time series evolution through LSTM, and optimize the vacuum curve through a policy gradient method to obtain the optimal degassing time and final hydrogen and nitrogen content under the vacuum curve.
[0110] In some embodiments, the deep learning model route further includes an integration layer at the end, which combines the outputs of each sub-network through learned weights, quantifies uncertainty through Bayesian deep learning methods, and optimizes hyperparameters through Bayesian optimization for architecture search.
[0111] In some embodiments, using the deep learning model route for prediction further includes updating the model parameters of the deep learning model according to the new data.
[0112] In some embodiments, updating the model parameters of the deep learning model according to the new data is performed by an online learning module, which uses an elastic weight consolidation algorithm to prevent catastrophic forgetting. Key historical samples are saved in an experience replay buffer. When a new steel grade needs to be produced, only a few furnace data are available, and traditional deep learning requires a large amount of data, which can be quickly adapted through meta-learning algorithms.
[0113] In some embodiments, model version management uses a Git-like mechanism to support branching, merging, and rollback, ensuring the stability of the production environment. An A / B testing framework allows new and old models to run in parallel, and determines the optimal version through statistical tests.
[0114] In some embodiments, the prediction results of the two model routes are integrated at the output layer, outputting the entire process schedule (including the processing time prediction of each process), the entire process temperature curve, the optimization scheme of material addition, the final composition prediction, and the quality prediction. The prediction result is also accompanied by confidence evaluation, which is calculated based on the uncertainty of the model parameters and the historical prediction accuracy. The output format follows the industry standard, facilitating integration with downstream systems.
[0115] S40: Adjusting the key operations of the vacuum degassing furnace according to the prediction results.
[0116] In some embodiments, the key operations include at least one of the following: adjusting the material addition scheme, adjusting the machine working mode, speed, and power. For example: based on the predicted optimal alloy ratio, specific control instructions can be generated to control the gate opening degree of the alloy feeding hopper, the frequency of the vibrating feeder, or the speed of the belt scale; based on the predicted slagging agent formula, the delivery gas flow and valve opening and closing timing of the slagging agent injection system can be controlled to inject powdered slagging agent into the ladle.
[0117] In some embodiments, the prediction results are delivered to the programmable logic controller (PLC) of the vacuum degassing furnace through standardized industrial communication protocols (such as Modbus TCP / IP, Profinet, etc.), and specific physical device control instructions are executed by the PLC.
[0118] In some embodiments, a detailed process report can be generated, including model selection basis, calculation process key parameters, prediction result confidence, etc., to provide support for production optimization and fault diagnosis. All data are automatically archived for subsequent model optimization and production analysis.
[0119] The second embodiment of the present application provides a vacuum degassing furnace control method. The following will be described with reference to the flowchart shown in the figure, the method provided by the second embodiment includes the following steps S200-S240. Figure 2a
[0120] S200: Collect initial parameters and pre-process.
[0121] The control method described in the present application can be installed and run on a computer (or various hosts, distributed systems) system in the form of a computer program. When started by the system, first, a comprehensive data collection is performed through the input parameter layer.
[0122] The system receives the molten steel weight W from the upstream converter through the standardized interface, which determines the reference for the subsequent addition amount of all materials. At the same time, the initial temperature T0 is obtained, and the temperature data needs to be measured at multiple points to improve accuracy. The initial composition C0 includes the content of various elements such as carbon, silicon, manganese, phosphorus, sulfur, hydrogen, and nitrogen. The ladle age A reflects the number of times the ladle is used and the heat preservation performance, which has an important influence on temperature prediction. The bottom blowing flow F includes the bottom blowing argon flow setting value at different stages of the entire process, which is a time series data. The target composition C target The quality requirements of the final product are defined, which is the optimization target of the entire control process.
[0123] According to the needs of each mechanism model or sub-network, the pressure of the vacuum degassing furnace, the vacuum pressure curve and other parameters provided by the equipment also need to be collected.
[0124] The above input parameters are pre-processed, which can include outlier detection, missing value filling and standardization processing. For example, data standardization can ensure that parameters of different dimensions are processed on the same scale, which can use:
[0125] x norm =(x-μ) / σ
[0126] Where μ is the mean of historical data, σ is the standard deviation of historical data, x is the currently collected input parameter data, and x norm is the standardized current input parameter data.
[0127] S210: Select a suitable model route, wherein the model route includes a mechanism model route and a deep learning model route.
[0128] For example, Figure 2b As shown, the scheme of the present application predicts the key parameters of each process in the vacuum degassing furnace through two model routes. Its characteristic is to organically combine the metallurgical mechanism model with the deep learning technology to build an intelligent control system with both physical interpretability and strong learning ability.
[0129] Among them, when selecting the model route, the model selection controller can make intelligent decisions according to the preset control instructions and the current production state. The decision logic includes: when the amount of historical data accumulated by the system exceeds the threshold (N min 1000 heats), the deep learning model route is preferred; when rapid response is needed or data is insufficient, the mechanism model route is selected; when the system is in the debugging stage or needs to be verified, both models can be run simultaneously for comparative analysis.
[0130] The model selection controller is also responsible for evaluating the reliability of the current input parameters. If a sensor failure for data collection or abnormal data collection is monitored, it will automatically switch to the more robust mechanism model route to ensure stable operation of the system under various working conditions.
[0131] S220: Select the mechanism model route to predict the key parameters.
[0132] When the mechanism model route is selected, the prediction is made through the mechanism model running framework as shown in Figure 2c The whole mechanism model adopts modular design, and each mechanism model can be independently run and cooperatively work, and the whole process optimization control is realized through model coordination and time sequence controller. The calculation method (S221-S222) of each mechanism model and the global optimization method (S223) of the mechanism model are described below. And how to continuously optimize the mechanism model by feeding back the output results (S224) is described.
[0133] S221: Load and verify the parameter configuration file.
[0134] The system first loads the model parameters in the parameter configuration file through a parameter configuration file manager. The parameter configuration file is managed in three levels according to the steel type, equipment number and optimization version. The loading process includes parameter integrity check, value range verification and version compatibility confirmation.
[0135] For example, for important parameters such as reaction rate constant k, activation energy E, etc., the system will compare with historical data, and if the deviation exceeds 30%, the parameter audit process will be triggered. After the parameter loading is completed, the system will generate a parameter snapshot for subsequent tracing and analysis.
[0136] S222: Use each mechanism model to calculate the key parameters of different processes.
[0137] As shown in Figure 2cAs shown, the mechanism model route includes a decarburization mechanism model, a deoxidation mechanism model, an alloying mechanism model, a desulfurization and slagging model, a degassing mechanism model, and a temperature mechanism model. The following will be explained in turn.
[0138] i. Decarburization mechanism model
[0139] Decarburization is the primary process of the vacuum degassing furnace, and the vacuum system of the vacuum degassing furnace is started to promote the decarburization reaction. The reaction kinetics model is based on the carbon-oxygen reaction mechanism:
[0140]
[0141] wherein, is the conversion rate of carbon content with time, i.e., the slope of the carbon content change curve; k c is the decarburization reaction rate constant; A is the reaction interface area; P CO is the carbon monoxide partial pressure, which is approximately equal to the current vacuum degree under deep vacuum, and the current vacuum degree can be obtained through a sensor; P eq is the equilibrium carbon monoxide partial pressure, which is obtained by calculation; f(T) is a temperature correction function; g(F) is a bottom-blown stirring correction function.
[0142] The above differential equation is solved by the Runge-Kutta method, and the time step is adaptively adjusted, with the step being automatically refined when the carbon content approaches the target value. Double standards are adopted for the decarburization endpoint judgment: the carbon content reaches the target value ± 0.002% or the decarburization rate is lower than the threshold. The decarburization time t1 and the carbon content change curve C(t) are output.
[0143] ii. Deoxidation mechanism model
[0144] After deep decarburization, the carbon content reaches the target, and a strong deoxidizer needs to be added to the molten steel for deoxidation. The deoxidation mechanism model calculates the stability of various deoxidation products through thermodynamic equilibrium. For commonly used deoxidizers, the system establishes a complete thermodynamic database, including the relationship between the standard Gibbs free energy of formation and temperature:
[0145] ΔG° i = A i +B i ×T+C i ×T×ln(T)
[0146] wherein, ΔG° i is the standard Gibbs free energy change of the i-th deoxidation reaction; A i is the constant term in the Gibbs free energy temperature relationship; B i is the first temperature coefficient; C i is the logarithmic temperature coefficient; T is the absolute temperature of the molten steel, which is obtained by pre-acquisition.
[0147] Based on the calculated Gibbs free energy, the deoxidizer selection algorithm calculates the equilibrium constant of the deoxidation reaction:
[0148]
[0149] Among them, K i R is the equilibrium constant for the i-th deoxygenation reaction; R is the ideal gas constant with a value of 8.314 J / (mol·K).
[0150] The equilibrium constant K is obtained. i Then, based on the oxygen activity coefficient and the activity of the deoxidation products, the type of deoxidizer and the theoretical deoxidizer requirement are determined. For example, taking the most typical aluminum deoxidation reaction as an example, the equilibrium aluminum content required for the deoxidation reaction to be completed is calculated, thereby calculating the theoretical deoxidizer requirement for various deoxidizers.
[0151] The actual amount of deoxidizer added should be adjusted based on the yield.
[0152] m 实际 =m 理论 / η(T,t,stirring intensity)
[0153] Where, m 实际 This is the actual amount of deoxidizer added; m 理论 This represents the theoretical demand for deoxidizer; η(T,t, stirring intensity) is the deoxidizer yield function. The yield model is based on a large amount of production data and considers the effects of temperature, reaction time, and stirring conditions.
[0154] The final deoxygenation mechanism model outputs the type of deoxygenator, the actual amount of deoxygenator added (m2), and the deoxygenation time (t2).
[0155] iii. Alloying Mechanism Model
[0156] After the deoxidation process, various alloying elements need to be added to the molten steel to improve its properties. The core of the alloying process is to achieve optimal cost while meeting compositional requirements. Therefore, the basic constraints are first established using a set of element balance equations:
[0157] C f,j =(C 0,j ×m0+∑(C i,j ×m i ×η i,j )) / m f
[0158] Among them, C f,j It is the final composition content of element j in molten steel, which is a pre-set target; C 0,j is the initial composition content of element j in the molten steel, where the initial composition is obtained by laboratory testing; m0 is the total mass of the molten steel before treatment, obtained from the secondary system before treatment; Ci,j is the percentage of element j in alloy i, which is entered into the secondary system when the material enters the warehouse, and is obtained from the secondary system when calculating; m i is the addition amount of the i-th alloy; m f is the total mass of molten steel after adding all alloys; η i,j is the yield of element j in alloy i.
[0159] Since it takes a long time to completely melt when adding large and dense alloys, it is also necessary to establish an alloy melting kinetics model to check whether the melting rate (melting time, addition sequence, etc. time sequence parameters) can meet the demand. The alloy melting kinetics model considers the effects of particle size, temperature and stirring:
[0160] dm / dt = k 熔 × A 颗粒 × (T - T 熔 ) × h(stirring)
[0161] Where dm / dt is the alloy melting rate; k 熔 is the melting rate constant; A 颗粒 is the total surface area of alloy particles, which is an estimated value; T is the current molten steel temperature, which is obtained from the previous process calculation; T 熔 is the melting point temperature of the alloy, which is obtained from the secondary system entered in advance; h(stirring) is the stirring enhancement function, which is calculated according to the bottom blowing flow.
[0162] When searching for an alloy combination, an alloy combination that satisfies both the element balance equation and the alloy melting kinetics model should be selected. By establishing a mixed integer linear programming problem, the alloy combination with the minimum cost is found under the condition of satisfying all component constraints.
[0163] Since the alloy type involves a discrete selection problem, the branch and bound algorithm can be used for solving. The alloying mechanism model finally outputs the optimal alloy ratio, addition sequence and alloying time t3.
[0164] iv. Desulfurization and slagging model
[0165] The next important process is desulfurization, which transfers sulfur in molten steel to slag through interfacial reaction to maximize the reduction of sulfur content in molten steel. The efficiency of the desulfurization process depends on the sulfur distribution coefficient of the slag-steel interface, which can be represented by the following distribution coefficient model that comprehensively considers the effects of slag composition, temperature and oxygen potential:
[0166] L S = K S × (f S / γ S ) × (a CaO / a CaS )^n × exp(Q / RT)
[0167] where L S is the distribution coefficient of sulfur between slag-steel; K S is the reference constant of the distribution coefficient of sulfur; f S is the activity coefficient of sulfur in the liquid steel, calculated by taking into account the mutual influence of sulfur with other elements of the liquid steel; y S is the activity coefficient of sulfides in the slag, calculated; a CaO is the activity of calcium oxide in the slag, calculated; a CaS is the activity of calcium sulfide in the slag, calculated; n is the reaction order; Q is the activation energy of the desulfurization reaction, which is a constant that needs to be empirically adjusted; R is the ideal gas constant, which is 8.314 J / (mol K); T is the reaction temperature.
[0168] The desulfurization slag-making model of the present application recommends the optimal slag-making agent formula through a slag system optimization algorithm. First, a three-dimensional response surface of basicity-sulfur distribution coefficient-temperature is established based on historical data and experiments, and through basicity control, the optimal basicity that maximizes the sulfur distribution coefficient is found at the current temperature through interpolation calculation, and the slag-making agent ratio is calculated according to the optimal basicity and the current basicity, thereby ensuring the desulfurization efficiency while controlling the slag amount. The relationship between basicity R and slag-making agent ratio is as follows:
[0169] R = (CaO + 1.4 x MgO) / (SiO2 + 0.6 x Al2O3)
[0170] where CaO, MgO, SiO2, and Al2O3 are the mass percentages of these oxides in the final slag after mixing.
[0171] After obtaining the slag-making agent formula, the desulfurization rate and desulfurization time t4 can be calculated and output together.
[0172] v. Degassing mechanism model
[0173] Finally, the gas dissolved in the molten steel, mainly hydrogen and nitrogen, is removed through the degassing process. The degassing process is based on the dissolution-diffusion-outgassing mechanism of gas in the molten steel. The vacuum degassing furnace breaks the dissolution equilibrium constantly through a vacuum pump, making the solubility of gas in the molten steel far exceed its new equilibrium solubility, forcing the gas to escape from the molten steel. The equilibrium solubility of hydrogen and nitrogen follows Sievert's law:
[0174]
[0175] where [H] eq is the equilibrium solubility of hydrogen in the molten steel; K H is the solubility constant of hydrogen; is the hydrogen partial pressure; [N] eq is the equilibrium solubility of nitrogen in the molten steel; KN is the solubility constant of nitrogen; is the partial pressure of nitrogen.
[0176] The diffusion modeling of the gas breaking the equilibrium solubility adopts Fick's second law, considering the enhancement of the stirring on the mass transfer coefficient:
[0177]
[0178] where, is the rate of change of the gas concentration with time; is the Laplace operator of the concentration; D eff is the effective diffusion coefficient, which is a constant, calculated according to the following formula:
[0179] D eff = D 分子 ×(1+β×Re 0.5 )
[0180] where, D 分子 is the molecular diffusion coefficient of the gas in the liquid steel, an empirical constant; β is the stirring enhancement coefficient, calculated by the bottom blowing flow; Re is the Reynolds number.
[0181] Finally, through multi-objective optimization of the degassing rate-vacuum degree-energy consumption, the optimal vacuum degree change curve P(t) is determined. Specifically, first, generate an initial vacuum degree change curve P(t) or a group of initial vacuum degree change curves P(t). According to the Sievert's law, the equilibrium solubility of the gas at different times is calculated, and according to the gas diffusion model of Fick's second law, the actual trajectory of the gas content in the molten steel with time and the time (degassing time t s ) used to reach the standard are further simulated, and the total energy consumption of the vacuum pump is calculated. The energy consumption and time are optimized to find the vacuum degree change curve that consumes less energy and takes less time. The degassing mechanism model finally outputs the degassing time t5 and the final hydrogen and nitrogen content under the vacuum degree change curve P(t).
[0182] vi. Temperature mechanism model
[0183] The temperature mechanism model is aimed at the whole treatment process of the vacuum degassing furnace. The heat input can include chemical reaction heat, electric arc heating, ladle roasting preheating, etc.; the heat loss can include radiation heat loss, convection heat loss, slag surface heat loss, etc. Taking the most important chemical reaction heat, radiation heat loss, and ladle heat loss as examples, the chemical reaction heat is calculated in real time according to the reaction progress of each process (here, real time refers to the prediction process of the mechanism model):
[0184]
[0185] where, Q 反应 is the heat released or absorbed by chemical reaction; ΔHi is the molar reaction enthalpy of the ith reaction, which is a standard constant; is the reaction rate of the ith reaction, which is obtained by calculation of the above model; M i is the molecular weight of the ith reaction;
[0186] The radiation heat dissipation adopts the Stefan-Boltzmann law, and the comprehensive emissivity of the molten steel surface and the ladle inner wall is considered:
[0187]
[0188] wherein Q 辐射 is the radiation heat dissipation power, ε 综合 is the comprehensive emissivity, which is obtained by calculation; σ is the Stefan-Boltzmann constant, which is 5.67x10 -8 W / (m 2 ·K 4 ); A 辐射 is the effective radiation area, which is obtained by empirical estimation; T 4 is the molten steel temperature; environment is the ambient temperature.
[0189] The ladle heat dissipation model considers the influence of the ladle age on the heat preservation performance:
[0190] Q 包壁 = U (ladle age) x A 包壁 x (T-T 环境 )
[0191] wherein Q 包壁 is the conduction heat dissipation power through the ladle wall; U (ladle age) is the heat transfer coefficient which changes with the ladle age; the ladle age is the number of times of using the ladle; A 包壁 is the inner wall area of the ladle, which is obtained by empirical estimation; T-T 环境 is the temperature difference, which is the potential difference driving the heat conduction.
[0192] The temperature change curve with time in the vacuum degassing furnace treatment process is calculated and obtained by the above temperature model, and the final heat state is output. Among them, for other mechanism models which need temperature as input, the heat state and the temperature change curve with time calculated by the last process are used as input.
[0193] S223: coordinate and optimize the prediction results of each mechanism model.
[0194] After each mechanism model is calculated, a material selection linear programmer is used for global optimization, specifically, by minimizing the total cost objective function:
[0195] min Z =∑(c i x i )+λ 时间×T 总 +λ 能耗 ×E 总
[0196] Where Z is the total cost objective function, i.e., the comprehensive cost index that needs to be minimized; c i x is the unit price of the i-th material; i It is the amount of the i-th material added; ∑(c i ×x i ) represents the total cost of materials; λ 时间 This is the time cost weighting coefficient, obtained through empirical estimation; T 总 This is the total processing time, obtained by calculating the material handling time; λ 能耗 It is the energy cost weighting coefficient, obtained through empirical estimation; E 总 It is the total energy consumption, which is estimated through experience.
[0197] Constraints include composition constraints, material supply constraints, and equipment capacity constraints.
[0198] A 成分 ×x≥b 成分 (Ingredient Requirements)
[0199] x≤x max (Inventory restrictions)
[0200] T 总 ≤T 限制 (Time Constraint)
[0201] Among them, A 成分 This is the component constraint coefficient matrix; x is the material addition vector; b 成分 It is the target component vector; x max This is the upper limit vector of inventory for each material; T 限制 This is the maximum allowed processing time.
[0202] After considering the constraints, the interior-point method is used to solve the total cost objective function. For ill-conditioned problems, preprocessing techniques are employed to improve numerical stability. For example, if the carbon constraint is set too stringently, scaling the carbon constraint equation during preprocessing can stabilize the numerical values and find a reliable optimal solution. After finding the optimal solution, the robustness of the solution can be evaluated using a sensitivity analysis module, and adjustment suggestions can be provided for sensitive parameters. For instance, the impact of the price of high-carbon ferromanganese on the final cost can determine whether the requirement for Mn content needs to be relaxed.
[0203] Furthermore, in order to handle the various transfer parameters between different mechanism models, such as calculating the total processing time, it is necessary to collect the time of each process step for calculation as follows:
[0204] T total =∑ti +∑Δt 过渡 +t 裕量
[0205] where T total is the total processing time, t i is the processing time of the ith process (i.e. t1-t5), t 裕量 is the safety margin, which is self-adaptively adjusted according to the uncertainty of model prediction; Δt 过渡 is the transition time between processes, which can be obtained according to the following empirical model:
[0206] Δt 过渡 = t 基准 +f(equipment state, process difference)
[0207] where t 基准 is the baseline transition time; f(equipment state, process difference) is the transition time correction function.
[0208] Then the total processing time can be calculated as:
[0209] The parameter transfer adopts standardized interfaces to ensure that the output of the upstream process can be correctly used as the input of the downstream process
[0210] The constraint checking module can be used to monitor the predicted results of the transfer parameters, intermediate / final outputs in real time, and adjust the constraint conditions, optimization objectives or model parameters immediately when the constraints are violated.
[0211] The abnormal handling mechanism can also be responsible for parameter out-of-range handling, calculation divergence handling and logic conflict handling. For example, the end-point carbon content and decarburization time output by the decarburization model are used as key parameters to transfer to the temperature model for heat balance calculation after the CO gas volume is calculated by the model coordinator. When the decarburization model cannot converge, it may affect the downstream model to lack key inputs, so an alarm needs to be processed, and the constraint conditions or parameter settings of the decarburization model need to be checked or prompted.
[0212] S224: The predicted results of the mechanism model are continuously optimized through a feedback mechanism.
[0213] The final output results of the mechanism model include the detailed prediction results of each process, the time arrangement of the whole process, the optimization scheme of material addition and the quality prediction of each mechanism model. The output format follows the industrial standard, which is convenient for integration with downstream systems.
[0214] The output results can be accompanied by a confidence assessment, which is calculated based on the uncertainty of the model parameters and the historical prediction accuracy. For example: according to the fitting results of historical data, the error statistics are calculated, the uncertainty is calculated, and the prediction interval is obtained according to the confidence level.
[0215] To continuously improve the mechanistic model during practical use, a feedback mechanism can be used to continuously collect actual production data, and the model parameters can be updated using parameter identification algorithms. For example, online parameter estimation can be achieved using the recursive least squares method.
[0216]
[0217] Where K is the gain matrix, which is adaptively adjusted according to data quality and model performance. θ(k+1) is the model parameter vector at time k+1; θ(k) is the model parameter vector at time k; K(k+1) is the gain matrix at time k+1; y(k+1) is the actual observation value at time k+1. It is the regression vector at time k+1; It is a predicted value based on the current parameters; This refers to prediction error. The updated parameters are validated and saved as a new version, enabling continuous model improvement.
[0218] S230: Select key parameters for deep learning model route prediction.
[0219] Deep learning model approach adopts, for example Figure 2d The architecture shown is based on a physical information neural network as the main control model. This architecture deeply integrates the powerful representation capabilities of deep learning with the laws of metallurgical physics, and coordinates multiple specialized sub-networks through the main control model to achieve end-to-end intelligent prediction and optimization.
[0220] The deep learning model approach is similar to the mechanistic model approach in its use, requiring only the initial parameters for prediction. However, while the mechanistic model predicts by progressively superimposing the results of each process based on physical principles, the deep learning model relies on trained sub-models and the master control model for prediction. Therefore, when training the deep learning model, it is not sufficient to use only historical data of the initial parameters of the initial process of the vacuum degassing furnace. Instead, it should use historical data of steel weight, initial temperature, initial composition, ladle age, bottom blowing flow rate, and target composition for each process stage; as well as historical data of key parameters and intermediate process parameters measured during operation, such as steel type, deoxidizer type, temperature curve, slagging agent formula, sulfur distribution coefficient, vacuum degassing furnace pressure, and vacuum pressure curve. The data acquisition method is similar to step S200 and will not be repeated here.
[0221] The following steps, S231-236, provide a detailed explanation of the data processing, master model coordination method, addition of physical constraints, prediction methods for each sub-model, optimization, and adaptation of this deep learning model roadmap.
[0222] S231: Input data preprocessing and feature engineering.
[0223] Data preprocessing is crucial for the success of deep learning models. The system first standardizes the raw input data, mapping parameters x with different dimensions to a uniform scale. norm :
[0224] x norm =(x-μ) train ) / (σ train +ε)
[0225] Where, μ train and σ train The mean and standard deviation of the training set are given by ε = 1e-8 to prevent division by zero errors.
[0226] Feature encoding employs a combination of techniques. For discrete variables such as steel grade, a learnable embedding vector e is used. 钢种 :
[0227] e 钢种 =Embedding(steel type ID, d embed )
[0228] Where, d embed For the embedding dimension, the steel type ID is the identifier of the steel type, and Embedding is the embedding operation.
[0229] For time-series features such as bottom-blowing flow patterns, one-dimensional convolution is used to extract local patterns F. 时序 :
[0230] F 时序 =Conv1D(F 原始 kernel size =5, filters=64)
[0231] The kernel size is 5, the number of kernels is 64, and Conv1D is a one-dimensional convolution operation.
[0232] Furthermore, physical quantities need to be dimensionless to ensure that the network learns the physical essence rather than numerical magnitude. For example, multiple dimensional original variables can be fused together using dimensionless numbers such as Reynolds number (Re) and Prandtl number (Pr) as additional feature inputs.
[0233] For samples with limited data, the training set can be augmented by adding controlled noise and physical constraint perturbations:
[0234] x aug =x + α × N(0, σ physics )
[0235] Where, x augThis is the enhanced data sample; x is the original data sample; α is the noise intensity coefficient; N(0,σ) physics () is a function with a mean of 0 and a standard deviation of σ. physics Gaussian noise; σ physics Determined based on the measurement accuracy of physical quantities, ensuring the physical rationality of enhanced data.
[0236] For missing values, intelligent filling can be performed using statistical features such as average values and time-slot relationships such as interpolation.
[0237] S232: Coordinates the various sub-networks through the physical information neural network master control model.
[0238] The physical information neural network master control model is the core of the entire system, responsible for feature fusion, physical constraint encoding, task allocation, and coordination of the loss function.
[0239] The feature fusion network uses a multilayer perceptron structure to extract features from input parameters of different modalities (such as local patterns of preprocessed temporal features in S231 and embedding vectors of discrete variables) and then concatenates and fuses them. However, residual connections and layer normalization are introduced.
[0240] H 1 =LayerNorm(X+MLP) 1 (X))
[0241] H 2 =LayerNorm(H 1 +MLP 2 (H 1 ))
[0242] Among them, H 1 It is the hidden representation after feature fusion in the first layer; X is the input feature vector; MLP 1 (X) is the output of the first multilayer perceptron; LayerNorm is the layer normalization operation; H 2 It is the hidden representation after the fusion of features from the second layer; MLP 2 (H 1 The output of the second multilayer perceptron is denoted as ; the residual connection is X+MLP. 1 (X) Ensure gradient flow and avoid gradient vanishing.
[0243] A physical constraint encoder ensures that the output satisfies physical constraints by translating abstract physical laws into a form that the network can understand. For example, for conservation laws, a constraint matrix is constructed, yielding the operator:
[0244] A 守恒 ×x=b 守恒
[0245] The constraints are embedded into the network by Lagrange multiplier method:
[0246] L physics =λ×||A 守恒 ×x-b 守恒 || 2
[0247] where A 守恒 is the conservation law constraint matrix, each row represents a conservation equation; x is the intermediate representation or output of the network; b 守恒 is the right constant vector of the conservation law; λ is the Lagrange multiplier, which controls the strength of the physical constraint; ||A 守恒 ×x-b 守恒 || 2 is the square norm of the constraint violation degree, which is used to convert the constraint condition into a penalty term; L physics is the physical constraint loss term, which is used in the total coordination loss function.
[0248] The role of the task allocator is to adjust the weights of the sub-networks to adaptively weight the contributions of each sub-network according to the input conditions, for example, the decarburization time is the primary prediction target, then the result of the decarburization prediction network is allocated a higher weight, and its predicted result is more accurate; here, the attention mechanism is used to dynamically allocate the weights of each sub-network:
[0249]
[0250] where α i is the attention weight of the i-th sub-network; softmax is a normalized exponential function, which ensures that the weights sum to 1; W query is the query projection matrix, which maps the features to the query space; H is the feature representation of the master model, which comes from the output of the feature fusion network; W key is the key projection matrix, which maps the features to the key space; d k is the dimension of the attention mechanism, which is used to scale the dot product result; is the scaling factor, which prevents the dot product result from being too large to cause gradient vanishing.
[0251] The physical information neural network master model uses a coordination loss function to comprehensively consider the prediction accuracy, sub-network consistency and physical constraints:
[0252] L coord =MSE(y pred ,y true )+λ1×∑||y i -y j || 2 +λ2×L physics
[0253] where Lcoord is the total coordination loss function; MSE(y pred , y true ) is the mean square error loss, measuring the prediction accuracy; y pred is the model predicted value; y true is the true label value; λ1 is the weight coefficient of the sub-network consistency loss, ∑||y i -y j || 2 is the accumulation of the difference between the outputs of different sub-networks, which is used to ensure that the prediction results of each professional model are coordinated, i.e. the prediction results of the different sub-networks for the cross part are coordinated with each other, avoiding contradictions, for example, the heat released by the decarburization amount (y i ) predicted by the decarburization prediction network and the heat released by the temperature change (y j ) predicted by the temperature prediction network should be similar; λ2 is the weight coefficient of the physical constraint loss term;
[0254] S233: providing physical constraints through the physical constraint encoding layer.
[0255] The core principle of the physical neural network is to embed physical constraints into the neural network (deep learning network) as part of the loss function to make the output results comply with physical laws. For the constraint matrix described in the physical encoder in S232, the operator is obtained from multiple conservation laws, for example, for the density value of element i, it should comply with the mass conservation constraint, which is realized by constructing the element balance equation:
[0256]
[0257] where ρ i is the density of element i, v i is the transport velocity, and R i is the reaction source term.
[0258] In the neural network, the partial derivative with respect to time, for example, element i, can be calculated by automatic differentiation:
[0259]
[0260] where ρ pred is the predicted density value (of element i) of the network; t is the time variable.
[0261] For the predicted temperature, it needs to comply with the energy conservation constraint, considering heat conduction, convection and radiation:
[0262]
[0263] where ρ is the density of the liquid steel, and C p is the constant-pressure specific heat capacity, which is a constant; It is the partial derivative of temperature with respect to time; k is the thermal conductivity, which is a constant. It is the heat conduction term; Q 源 It is the heat source term, including the heat of chemical reaction, obtained through calculation; Q 损 This is the heat loss term, which includes radiation and convection losses, and is obtained through calculation.
[0264] In addition, to satisfy the chemical equilibrium constraint, a chemical potential network μ can be introduced. net Predict the chemical potentials of each component to ensure the system tends towards equilibrium. Chemical equilibrium constraints are based on the principle of minimum Gibbs free energy.
[0265] G=∑n i ×μ i →min
[0266] Where G is the total Gibbs free energy of the system; n i It is the number of moles of component i; μ i It is the chemical potential of component i; ∑n i ×μ i It is the weighted sum of the chemical potentials of all components.
[0267] Chemical potential networks can also be used to constrain phase equilibrium and ensure the thermodynamic consistency of multiphase systems.
[0268]
[0269] Chemical potential of component i in the gas phase.
[0270] The aforementioned conservation law can be used as an operator described by the physical encoder in S232 to calculate the physical loss term on the output of the sub-network, which is used by the physical encoder to adjust the neural network parameters during backpropagation.
[0271] S234: Use each sub-network to predict key parameters for different processes.
[0272] like Figure 2d As shown, the sub-networks in the deep learning model include a decarburization prediction network, a temperature prediction network, a deoxidation prediction network, an alloying prediction network, a desulfurization and slag formation prediction network, and a degassing prediction network. These will be explained in turn below.
[0273] i. Decarbonization Prediction Network
[0274] The decarbonization network can employ an LSTM architecture to handle time-series dependencies. Network inputs include time series data for initial carbon content, temperature, vacuum degassing furnace pressure (built into the equipment), and bottom-blowing flow rate.
[0275] x t =[C t ,T t ,Pt ,F t ,t]
[0276] where x t is the input vector at time t; C t is the carbon content at time t; T t is the temperature at time t; P t is the pressure at time t; F t is the bottom blowing flow rate at time t; and t is the time step.
[0277] The LSTM architecture captures long-term dependencies through a gating mechanism:
[0278] f t = σ(W f × [h {t-1} , x t ] + b f ) # forget gate
[0279] i t = σ(W i × [h {t-1} , x t ] + b i ) # input gate
[0280]
[0281] o t = σ(W o × [h {t-1} , x t ] + b o ) # output gate
[0282] h t = o t × tanh(C t ) # hidden state
[0283] where f t is the forget gate output, taking values in the range [0, 1]; σ is the sigmoid activation function; W f is the forget gate weight matrix; h {t-1} is the hidden state at the previous time step; [h {t-1} , x t ] is the concatenation of the hidden state and the input; and b f is the forget gate bias vector.
[0284] i t is the input gate output; W i is the input gate weight matrix; and b i is the input gate bias.
[0285] is the candidate cell state; tanh is the hyperbolic tangent activation function; W C is the candidate value weight matrix; b C is the candidate value bias;
[0286] C t is the current cell state, storing long-term memory; C {t-1} is the previous cell state; f t ×C {t-1} is the history information after forgetting; is the new information after screening;
[0287] o t is the output gate, controlling the amount of output information; W o is the output gate weight matrix; b o is the output gate bias;
[0288] h t is the current hidden state, as the output and the input of the next time step.
[0289] The attention mechanism is introduced to the output of the LSTM to enhance the focus on key moments:
[0290] α t =softmax(W a ×h t )
[0291] c t =∑(α i ×h i )
[0292] where α t is the attention weight at the current time step; W a is the attention weight matrix to be learned; c t is the context vector after weighting by the attention weight at the i-th time step; α i is the attention weight at the i-th time step; h i is the hidden state at the i-th time step;
[0293] The decarburization prediction network obtains and outputs the predicted carbon content change curve, as well as the corresponding decarburization time and decarburization rate, by predicting the carbon content at the next time step:
[0294] [t 脱碳 ,C(t),v C(t) ]=Dense(concat([h T ,c T ]))
[0295] where t 脱碳is the predicted decarburization completion time; C(t) is the carbon content time curve; v C(t) is the decarburization rate curve. Dense is a fully connected layer; concat is a vector concatenation operation; h T is the hidden state at the last moment; c T is the final context vector.
[0296] When the network is trained, the decarburization kinetics equation in step 222 can be added to constrain the decarburization rate to comply with the kinetics law. The physical constraints of the subnetwork are similar to the physical constraint method in step 233, but are used for the subnetwork itself, and will not be described in detail.
[0297] ii. Temperature prediction network
[0298] The temperature prediction part obtains the temperature prediction network by learning the relationship between the historical temperature curve of the initial parameters (mainly the weight of the molten steel, the initial temperature, the ladle age, and can also include the initial composition, the bottom blowing flow, and the target composition). ResNet architecture is adopted to effectively solve the gradient disappearance problem of deep network. The basic residual block is designed as:
[0299] y = F(x, {W i}) + x
[0300] Where y is the output of the residual block; F(x, {W i}) is a residual mapping function that ensures gradient flow through a jump connection; x is the input of the residual block; {W i} is the weight parameter set of each layer in the residual block; the jump connection x ensures that at least the identity mapping is preserved.
[0301] The temperature prediction network has a depth of 20 (i.e., the network contains 20 residual blocks), and the structure of each residual block is as shown in Figure 2e . Among them, BN is a batch normalization layer that normalizes the features; ReLU is a rectified linear unit activation function used to introduce nonlinearity; Conv is a convolutional layer used to extract features; the first BN-ReLU-Conv sequence performs feature transformation; the second BN-ReLU-Conv sequence further processes the features; + indicates element-wise addition of the jump connection and the residual path.
[0302] Finally, the theoretical temperature drop model is introduced to enhance the prediction by physical embedding (different from soft constraints, which directly modify the prediction results):
[0303] T pred = T nn + ΔT 理论
[0304] Where T pred is the final temperature prediction value; T nnIt is the raw output of the neural network; ΔT 理论 It is a temperature correction based on the physical model, calculated based on the heat balance equation:
[0305] ΔT 理论 =-(Q 辐射 +Q 对流 )×Δt / (m×C p )
[0306] Among them, Q 辐射 It is the radiative heat dissipation power; Q 对流 Δt is the convective heat dissipation power; m is the time step; C is the mass of molten steel; p It is the specific heat capacity of molten steel.
[0307] The output of the temperature prediction network includes the temperature curve T(t) and the rate of temperature change dT / dt. A smooth L1 loss function is used during training, which is robust to outliers.
[0308] iii. Deoxygenation prediction network
[0309] Deoxidation prediction networks can use Generative Neural Networks (GNNs) to model the complex reaction relationships between various components after the addition of deoxidizers. Elements and compounds are represented as nodes in a graph, and the reaction relationships are represented as edges.
[0310] G = (V, E)
[0311] V = {element nodes ∪ compound nodes}
[0312] E = {reaction edge ∪ affinity edge}
[0313] Where G represents the graph structure of the chemical reaction network; V is the set of nodes in the graph; E is the set of edges in the graph; element nodes represent basic elements such as Fe, O, Al, and Si; compound nodes represent deoxygenation products such as Al2O3 and SiO2; reaction edges connect elements and compounds that can react; and affinity edges represent the strength of chemical affinity between elements.
[0314] Node characteristics are updated via a message passing mechanism:
[0315]
[0316] in, It is the feature vector of node v after the (k+1)th iteration; σ is the feature vector of node v in the k-th iteration; σ is the activation function, usually ReLU; W self It is a self-connection weight matrix, used to update the node's own features; W edge It is an edge weight matrix that aggregates neighbor information; is the eigenvector of the neighbor node u in the kth iteration; N(v) is the set of all neighbor nodes of node v.
[0317] The deoxidation prediction network will connect the above GNN model with a reinforcement learning module, and optimize the deoxidant selection strategy by controlling the learning reward (the reaction product and time of the GNN model output using various deoxidant selection strategies). The state space includes the current oxygen content, temperature and available deoxidant; the action space is the combination of deoxidant type and addition amount. The Q-value network is updated by deep Q learning:
[0318]
[0319] where s t is the state at time t, including oxygen content, temperature, etc.; a t is the action at time t, i.e. deoxidant selection and addition amount; Q(s t ,a t ) is the value function of state-action pair; α is the learning rate, controlling the update step size; r t is the immediate reward obtained at time t; γ is the discount factor, balancing immediate reward and future reward; max a Q(s {t+1} ,a) is the maximum Q value of the next state.
[0320] The reward function r (or r(t)) considers deoxidation effect, cost and side effects comprehensively:
[0321] r = -w 1 × |O 目标 -O 实际 | - w 2 × cost - w 3 × inclusion index
[0322] where r is the total reward value; w 1 is the deoxidation effect weight; O 目标 is the target oxygen content; O 实际 is the actual achieved oxygen content; |O 目标 -O 实际 | is the deoxidation deviation; w 2 is the cost weight; cost includes deoxidant price and amount; w 3 is the inclusion weight; the inclusion index reflects the impact of deoxidation product on steel quality.
[0323] The final deoxidation prediction network outputs the optimal deoxidant type and addition amount, as well as the corresponding parameters such as deoxidation time, etc.
[0324] iv. Alloying prediction network
[0325] The alloying prediction network adopts a coupling manner of DBN and linear programming to process the relationship between the alloying stage, the alloy adding amount, the molten steel weight, the initial composition, the initial temperature and the target composition. First, the DBN learns the probability distribution of the composition space (alloy composition) layer by layer for pre-training prediction:
[0326] P(v, h) = exp(-E(v, h)) / Z
[0327] wherein v is the visible layer (input composition), h is the hidden layer (feature representation), E is the energy function, and Z is the partition function.
[0328] The prediction of the composition space optimizes the composition prediction and the cost simultaneously by using a multi-task learning loss function:
[0329] L total = L 成分 + β × L 成本
[0330] wherein L total is the total loss function. L 成分 is the composition prediction loss. L 成本 is the cost prediction loss. β is a balance coefficient.
[0331] The composition space prediction adopts a vector output, and each element corresponds to the final content of one alloy element. The cost prediction is realized by an additional regression head.
[0332] In order to find the optimal alloy ratio in the composition space, linear programming is added to the DBN, and a differentiable optimization layer is realized:
[0333] x* = argmin {x} (c T × x) s.t. Ax ≤ b
[0334] wherein x* is the optimal solution, representing the best adding amount of various alloys; argmin represents the x value that minimizes the objective function; c T × x is the linear objective function, c is the cost vector (unit price of each alloy), and x is the decision variable (adding amount of each alloy); in the constraint condition Ax ≤ b, A is the constraint matrix, each row represents a constraint condition, such as the content limit of a certain element; x is the decision variable vector; b is the constraint right end vector, representing the upper limit value of each constraint.
[0335] For the above linear programming problem, the differentiable implementation of the interior point method is used to solve it, and the back propagation gradient can be obtained to obtain the network parameters.
[0336] The final alloying prediction model outputs the optimal alloy ratio or adding amount.
[0337] v. Desulfurization and slagging prediction network
[0338] The desulfurization slagging prediction network adopts a Transformer architecture to process the sequence evolution of slag components, and reverses the slagging agent formula and sulfur distribution coefficient through the expected sequence evolution. The self-attention mechanism of the Transformer is used to capture the mutual influence between different components:
[0339]
[0340] where Q is the query matrix, representing the information that the slag component at the current time wants to query; K is the key matrix, containing the feature identifiers of the slag components at each time; V is the value matrix, storing the actual slag component information at each time; QK T is the dot product of the query and the key, measuring the relevance; d k is the dimension of the key vector; is the scaling factor to prevent the dot product from being too large; softmax converts the relevance into attention weights.
[0341] Multi-head attention learns interactions from different subspaces:
[0342] MultiHead=Concat(head1,...,head h )×W O
[0343] where MultiHead is the output of multi-head attention; head i is the output of the i-th attention head; h is the number of attention heads, usually 8 or 16; Concat is the concatenation operation; W O is the output projection matrix that maps the concatenated result back to the model dimension.
[0344] The Transformer model uses relative position encoding for position encoding of the feature representation of the input sequence, which reflects the time series features:
[0345] PE(pos,2i)=sin(pos / 10000^(2i / d_model))
[0346] PE(pos,2i+1)=cos(pos / 10000^(2i / d_model))
[0347] where PE is the position encoding; pos is the position index in the sequence; i is the dimension index; d_model is the dimension of the model; 10000 is the base number, which controls the period of different dimensions; 2i corresponds to the use of sine for even dimensions; 2i+1 corresponds to the use of cosine for odd dimensions; relative position encoding can capture time interval information.
[0348] The desulfurization and slagging prediction network finally outputs the slagging agent formulation and the sulfur distribution coefficient L s The complete desulfurization process is generated by sequence-to-sequence manner through the time series of the vacuum degree curve.
[0349] vi. The degassing prediction network
[0350] The degassing prediction network adopts a CNN-LSTM hybrid architecture. The CNN extracts the local patterns of the vacuum degree curve historical data, and the LSTM models the time evolution. Among them, the feature extraction is as follows:
[0351] F CNN = MaxPool(ReLU(Conv1D(P(t))))
[0352] Where F CNN is the feature vector extracted by CNN, which captures the local pattern of the change of vacuum degree; P(t) is the change curve of vacuum degree with time; Conv1D is a one-dimensional convolution operation; ReLU is an activation function; MaxPool is a maximum pooling operation.
[0353] The LSTM time series modeling is as follows:
[0354] H LSTM = LSTM(F CNN ,h {t-1} )
[0355] Where H LSTM is the output state vector of LSTM; F CNN is the preprocessed feature as the input of LSTM; h {t-1} is the hidden state of the previous moment, carrying historical information; LSTM selectively remembers and forgets information through the gating mechanism, which will not be described here.
[0356] The above CNN-LSTM architecture can provide the degassing time and hydrogen and nitrogen content under any vacuum degree curve.
[0357] And for the vacuum degree curve, it can be optimized by the policy gradient method, and can be automatically used in the automatic control of the vacuum degassing furnace:
[0358]
[0359] Where, is the gradient of the objective function J with respect to the policy parameter θ; E represents the expected value; is the gradient of the logarithmic policy; π θ(a|s)is the policy network, which selects the probability of action a given state s, specifically, the current state [s = the output of CNN-LSTM under the current vacuum degree] selects the action [a = the next vacuum degree change rate]; A(s, a) is the advantage function, which measures the goodness of action a relative to the average level.
[0360] The final degassing prediction network outputs the degassing time t under the optimal vacuum degree curve 脱气 and the final hydrogen and nitrogen content [H] f , [N] f , and provides a confidence interval estimate.
[0361] S235: Coordinate and optimize the outputs of each sub-network.
[0362] The deep learning model route finally includes an integration layer, as shown in Figure 2d which includes integrated strategy, uncertainty quantification, and Bayesian optimization.
[0363] The integrated strategy refers to the combination of the outputs of each sub-network through learned weights. Compared to the preference settings of the physical information neural network master model task distributor for the prediction focus of each sub-network, the weight combination here is used to fine-tune the outputs of each sub-network according to their confidence. Higher weight means that the sub-network is more reliable and contributes more to the prediction result. Specifically:
[0364] Y final =∑(w i ×y i )+b ensemble
[0365] where Y final is the final output of the system, which integrates the prediction results of all sub-networks. w i is the weight of the i-th sub-network. y i is the output of the i-th sub-network, where y i refers to the overall time arrangement of the final output of the physical information neural network master model, the full-process temperature curve, the optimization scheme of material addition, and the final component prediction, etc. ∑(w i ×y i ) is the weighted sum of all sub-network outputs. b ensemble is the integrated bias term.
[0366] In order to quantify the uncertainty of the model prediction results. Adopt the Bayesian deep learning method:
[0367]
[0368] where, is the total uncertainty, is the data inherent uncertainty, For model uncertainty, estimate by dropout sampling.
[0369] To further optimize the structure and hyperparameters of the deep learning model, Bayesian optimization is adopted:
[0370] α next = argmax α (μ(α) + κ × σ(α))
[0371] where α next is the next combination of hyperparameters to be tried. argmax α indicates finding the value of α that maximizes the objective function. μ(α) is the expected performance at α predicted by the Gaussian process. σ(α) is the uncertainty at α. k is the exploration-exploitation trade-off coefficient: k is larger, it tends to explore uncertain areas, k is smaller, it tends to exploit known good areas.
[0372] S236: Continuously adapt to new data through online learning and incremental update.
[0373] The deep learning model can continuously adapt to new data through the online learning module. Specifically, it includes four parts: EWC algorithm, experience replay, meta-learning, and model version management (as shown in Figure 2d ).
[0374] Using the Elastic Weight Consolidation (EWC) algorithm can prevent catastrophic forgetting:
[0375]
[0376] where L is the total loss function. L new is the loss on new data. λ is the memory strength coefficient. F i is the i-th diagonal element of the Fisher information matrix, which quantifies the importance of parameter θ i for old tasks. θ i is the current parameter value, is the parameter value trained on old tasks. Penalize the deviation of parameters from the original value.
[0377] Use the experience replay buffer to save key historical samples:
[0378] B = {(x i , y i ) | importance(x i , y i ) > threshold}
[0379] where B stores those historical cases that are particularly important or rare. (x i , yi ) is an input-output pair.importance(x i ,y i ) is an importance score function.threshold is an importance threshold.
[0380] When a new steel grade needs to be produced, only a few furnace data are available. Traditional deep learning requires a large amount of data and cannot quickly adapt. Meta-learning MAML algorithm can be used to achieve fast adaptation:
[0381]
[0382] Where θ is the initial model parameter. θ' is the parameter updated for one step on the training set. α is the inner loop learning rate. is the gradient on the training data. θ final is the final parameter. β is the outer loop learning rate. f θ ′(x val ),y val indicates the prediction on the validation set using the updated parameter θ'.
[0383] The version management of the deep learning model adopts a Git-like mechanism, supporting branching, merging and rollback, ensuring the stability of the production environment. The A / B testing framework allows new and old models to run in parallel, and determines the optimal version through statistical tests.
[0384] S240: Process the prediction results of the outputs of the two model routes, and adjust the key operations of the vacuum degassing furnace according to the prediction results.
[0385] The calculation results of the two model routes can be integrated at the output layer, and the output includes: the time schedule of the whole process (including the processing time prediction of each process, the accumulation of t1 to t5), the precision reaches ± 30 seconds; the temperature curve T(t) of the whole process, one sampling point per minute; the addition scheme of various materials, including type, quantity and addition time; final composition prediction, including the content of all control elements and confidence interval; and quality prediction.
[0386] When the two models run simultaneously, the system calculates the consistency index of the prediction results. If the deviation exceeds the threshold, an exception analysis process is triggered to help identify the cause of the model deviation. The output results are transmitted to the VD furnace control system through a standardized interface, and are displayed on the human-machine interface for operators to monitor and intervene.
[0387] The method described in the present application can also generate detailed process reports, including model selection basis, calculation process key parameters, prediction result confidence, etc. to provide support for production optimization and fault diagnosis. All data are automatically archived for subsequent model optimization and production analysis.
[0388] The main innovation of the present application is to organically combine deep learning technology with metallurgical mechanism model, and construct an intelligent control system with both physical interpretability and strong learning ability. The application of physical information neural network ensures that the model prediction result conforms to the basic principles of metallurgy, avoiding unreasonable results that may be produced by pure data-driven methods.
[0389] The dual-line architecture design makes the system have good flexibility and reliability. When the data is sufficient, the deep learning model can provide more accurate prediction; when the data is insufficient or fast response is needed, the mechanism model can provide reliable prediction based on physical principles. The two models can verify each other, improving the overall reliability of the system.
[0390] The online learning mechanism makes the system have adaptive ability, which can continuously optimize model parameters as the production conditions change, maintaining the prediction accuracy. This continuous learning ability is of great significance for dealing with raw material fluctuations, equipment aging and other actual production problems.
[0391] The third embodiment of the present application provides a vacuum degassing furnace control system, as shown in Figure 3 The vacuum degassing furnace control system includes an input parameter layer, a model selection controller, a mechanism model, a deep learning model, and an output layer.
[0392] The input parameter layer is used for pre-processing the input initial parameters.
[0393] The model selection controller is used to select the mechanism model route and / or the deep learning model route.
[0394] The mechanism model includes at least one of the following mechanism models: decarburization mechanism model, deoxidation mechanism model, alloying mechanism model, desulfurization and slagging model, degassing mechanism model, temperature mechanism model; It also includes a material selection linear programmer and a model coordination and timing controller.
[0395] The deep learning model includes a physical information neural network master model and at least one of the following sub-networks: decarburization prediction network, deoxidation prediction network, alloying prediction network, desulfurization prediction network, degassing prediction network, temperature prediction network; It also includes an online learning and parameter updating module.
[0396] The output layer is used to output the prediction results of the two model routes.
[0397] The fourth embodiment of the present application provides a vacuum degassing furnace control device, which can be used to realize the vacuum degassing furnace control method in the above embodiments, as shown in Figure 4 The vacuum degassing furnace control device includes:
[0398] The data acquisition module is configured to acquire initial parameters. Specifically, the data acquisition module can be configured to implement step S10 in the first embodiment and optional embodiments thereof.
[0399] The model selection module is configured to select a model route. The model route includes a mechanism model route and a deep learning model route. Specifically, the model selection module can be configured to implement step S20 in the first embodiment and optional embodiments thereof.
[0400] The parameter prediction module is configured to predict key parameters of the vacuum degassing furnace according to the initial parameters and the selected model route. Specifically, the parameter prediction module can be configured to implement step S30 in the first embodiment and optional embodiments thereof.
[0401] The control adjustment module is configured to adjust the key operation of the vacuum degassing furnace according to the prediction result. Specifically, the control adjustment module can be configured to implement step S40 in the first embodiment and optional embodiments thereof.
[0402] Figure 5 FIG. 9 is a structural schematic diagram of a computing device 900 according to an embodiment of the present application. The computing device can execute the optional embodiments of the above method. The computing device can be a terminal or a chip or chip system inside the terminal. As shown in FIG. 9, the computing device 900 includes a processor 910, a memory 920, and a communication interface 930. Figure 5
[0403] It should be understood that Figure 5 The communication interface 930 in the computing device 900 shown in FIG. 9 can be configured to communicate with other devices, and can specifically include one or more transceiver circuits or interface circuits.
[0404] The processor 910 can be connected with the memory 920. The memory 920 can be configured to store program codes and data. Therefore, the memory 920 can be a storage unit inside the processor 910, or an external storage unit independent of the processor 910, or a component including the storage unit inside the processor 910 and the external storage unit independent of the processor 910.
[0405] Optionally, the computing device 900 can further include a bus. The memory 920 and the communication interface 930 can be connected with the processor 910 through the bus. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, and a control bus, etc. For ease of representation,Figure 5 A single bus or a single type of bus can be used. However, the bus 912 is the bus that is connected with the processor 910, and can be used to transmit the data between other components in the computing device 900 and the processor 910.
[0406] It should be appreciated that in the embodiments of the present application, the processor 910 can be a central processing unit (CPU). The processor can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. Alternatively, the processor 910 can be one or more integrated circuits for performing related programs to implement the technical solutions provided by the embodiments of the present application.
[0407] The memory 920 can include a read-only memory and a random access memory, and provide instructions and data for the processor 910. A part of the processor 910 can also include a non-volatile random access memory. For example, the processor 910 can also store device type information.
[0408] When the computing device 900 is running, the processor 910 executes computer-executed instructions in the memory 920 to perform any operation steps of the above method and any optional embodiments thereof.
[0409] It should be appreciated that the computing device 900 according to the embodiments of the present application can correspond to a subject performing the corresponding method according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 900 are respectively for implementing the corresponding flow of each method of the embodiments, and for brevity, will not be repeated here.
[0410] Those of ordinary skill in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0411] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0412] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic, and the division of the units is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0413] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.
[0414] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0415] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program codes that can be stored in the medium.
[0416] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The program is executed by a processor to execute the above method. The method includes at least one of the schemes described in the various embodiments.
[0417] The computer storage medium of the embodiments of the present application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0418] The computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave, in which computer-readable program code is embodied. Such propagated data signals can take a wide variety of forms, including but not limited to electro-magnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium that is not a storage medium, that is, that is not a tangible medium, and that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus or device.
[0419] The program code embodied on the computer-readable media can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0420] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, application specific circuitry, or field programmable gate array (FPGA) circuitry can execute the computer program code.
[0421] In addition, the words "first", "second", "third", etc., or "module A", "module B", "module C" and the like in the description and claims are used only to distinguish similar objects, and do not represent a specific order or sequence of the objects, and it is understood that the specific order or sequence can be interchanged, if permitted, so that the embodiments described herein can be implemented in other than the order or sequence described herein.
[0422] In the above description, the reference signs representing steps, such as S110, S120, etc., do not necessarily mean that the steps are executed in the order, and the order of the steps can be interchanged, or the steps can be executed simultaneously, if permitted.
[0423] The term "comprising" as used in the specification and claims should not be interpreted as limiting to the listed elements; it does not exclude other elements or steps. It means that the listed elements are present, but it does not exclude the presence or addition of one or more other features, integers, steps, components or groups thereof. Therefore, the expression "a device comprising A and B" should not be limited to a device only consisting of A and B.
[0424] The phrase "one embodiment" or "an embodiment" appearing in the specification is meant to convey a specific feature, structure, or characteristic described in connection with that embodiment includes in at least one embodiment of the application. Thus, use of the phrase "in one embodiment" or "in an embodiment" appearing in various places in the specification are not necessarily all referring to the same embodiment, although they can. Furthermore, in one or more embodiments, features, structures, or characteristics can be combined in any suitable manner in one or more embodiments, as would be apparent to one of ordinary skill in the art upon inspection of this disclosure.
[0425] Note that the above only describes the preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, reconfigurations and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and all fall within the scope of the present application.
Claims
1. A method for controlling a vacuum degassing furnace, characterized in that, Includes the following steps: Collect initial parameters; Select a model approach; wherein, the model approach includes a mechanistic model approach and a deep learning model approach; Based on the initial parameters, the key parameters of each step of the vacuum degassing process are predicted using the selected model route; The key operations of the vacuum degassing furnace are adjusted based on the predicted results.
2. The method according to claim 1, characterized in that, Using the aforementioned mechanistic model for route prediction includes: The initial parameters are input into the mechanism model, which includes a decarburization mechanism model, a deoxidation mechanism model, an alloying mechanism model, a desulfurization and slag formation model, a degassing mechanism model, and a temperature mechanism model. The mechanism model is a series of constraint equations established based on the corresponding physical characteristics. By solving the constraint equations, the key parameters of the vacuum degassing furnace are obtained. Using the deep learning model approach for prediction includes: The initial parameters are preprocessed and feature-engineered. The processed data is input into a deep learning model; the deep learning network uses a physical information neural network master model to constrain each sub-network with physical laws, and the sub-networks include a decarbonization prediction network, a temperature prediction network, a deoxidation prediction network, an alloying prediction network, a desulfurization slag formation prediction network, and a degassing prediction network. The sub-network predicts the corresponding key parameters through a deep learning network; The physical information neural network also coordinates and optimizes the prediction results of the sub-networks.
3. The method according to claim 2, characterized in that, Using the aforementioned mechanism model for prediction also includes: performing global optimization on the prediction results of the mechanism model, wherein the global optimization is constrained by component requirements, material supply inventory, and process processing time; It monitors in real time whether the transmitted parameters and intermediate / final output prediction results violate the constraints and adjusts the model parameters, handling parameter out-of-bounds errors, calculation divergence and logical conflicts.
4. The method according to claim 2, characterized in that, Using the deep learning model for prediction also includes updating the model parameters of the deep learning model based on the new data.
5. The method according to claim 1, characterized in that, The selection of the model route includes: selecting the deep learning model route when the amount of historical data accumulated by the system exceeds a threshold; and selecting the mechanism model route when the amount of historical data is insufficient.
6. The method according to claim 1, characterized in that, The selected model route can simultaneously use both the mechanistic model route and the deep learning model route. When using both the mechanistic model route and the deep learning model route, a consistency index of the prediction results is calculated. If the deviation exceeds the threshold, an anomaly analysis process is triggered, including: identifying the cause of the deviation and outputting the results to the vacuum degassing furnace control system.
7. The method according to claim 1, characterized in that, The initial parameters include at least one of the following: molten steel weight, initial temperature, initial composition, ladle age, bottom blowing flow rate, and target composition; The key parameters include at least one of the following: process time, temperature profile, material addition scheme, and final composition; The key operations include at least one of the following: adjusting the material addition scheme, adjusting the machine working mode, speed, and power.
8. A vacuum degassing furnace control system, characterized in that, include: Input parameter layer, model selection controller, mechanistic model, deep learning model, and output layer; The input parameter layer is used to preprocess the initial input parameters; The model selection controller is used to select the mechanistic model route and / or the deep learning model route; The mechanism model includes at least one of the following: decarburization mechanism model, deoxidation mechanism model, alloying mechanism model, desulfurization and slag formation model, degassing mechanism model, and temperature mechanism model; it also includes a material selection linear planner, as well as a model coordination and timing controller; The deep learning model includes a physical information neural network master model, and at least one of the following sub-networks: decarbonization prediction network, deoxidation prediction network, alloying prediction network, desulfurization prediction network, degassing prediction network, and temperature prediction network; it also includes an online learning and parameter update module. The output layer is used to output the prediction results of the two model routes.
9. A vacuum degassing furnace control device, characterized in that, include: The data acquisition module is used to collect initial parameters; The model selection module is used to select a model route; wherein, the model route includes a mechanistic model route and a deep learning model route; The parameter prediction module is used to predict the key parameters of the vacuum degassing furnace based on the initial parameters and through the selected model route. A control adjustment module is used to adjust the key operations of the vacuum degassing furnace based on the prediction results.
10. A computing device, characterized in that, include: processor, and A memory having stored program instructions that, when executed by the processor, cause the processor to perform the vacuum degassing furnace control method according to any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that, It stores program instructions that, when executed by a computer, cause the computer to perform the vacuum degassing furnace control method according to any one of claims 1 to 7.
12. A computer program product, characterized in that, It includes program instructions that, when executed by a computer, cause the computer to perform the vacuum degassing furnace control method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for intelligently controlling dosage of ammonia water serving as industrial flue gas absorbent
CN116203836A
Method, device and equipment for controlling temperature in closed high-temperature cavity
CN116225100A
Quality evaluation and analysis method for excavation support of underground powerhouse
CN119761147A
Agent-driven intelligent AI depth prediction method and system for power system
CN120257826A
Offshore wind plant generation power prediction method and system based on big data
CN120638331A