Method and system for optimizing metallurgical process of hypoxic hypoxic aluminum vanadium alloy

By constructing parameter-coupled state vectors and deep reinforcement learning networks, the complexity of oxygen and nitrogen content control in metallurgical processes was solved, achieving high-precision process parameter optimization and stability improvement, and reducing energy consumption and material waste.

CN120525141BActive Publication Date: 2025-11-04BAOJI JIACHENG RARE METAL MATERIALS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511029232.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-04
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In existing metallurgical processes, the control of oxygen and nitrogen content is difficult to accurately describe the complex nonlinear coupling relationship between process parameters. The lack of dynamic correlation modeling between parameters leads to low prediction accuracy and optimization strategies that cannot adapt to the dynamic changes in the process, resulting in large fluctuations in production stability and product quality.

Method used

A three-layer gated graph attention network and a temporal attention module are used to extract the correlation features between parameters. Thermodynamic equilibrium equation and long short-term memory network are combined to predict oxygen and nitrogen content. A deep reinforcement learning network is constructed to optimize parameters. Monte Carlo tree search is used to plan future process parameters, so as to achieve high-precision prediction and optimization of oxygen and nitrogen content.

Benefits of technology

It improves the accuracy and completeness of process state characterization, has strong adaptability, and can find the global optimal parameter combination under the constraint of oxygen and nitrogen content, which significantly improves the stability of metallurgical process and product quality, and reduces energy consumption and material waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525141B_ABST
    Figure CN120525141B_ABST
Patent Text Reader

Abstract

The application provides a hypoxic and hypoxic aluminum vanadium alloy metallurgical process optimization method and system, relates to the technical field of process optimization, and comprises the following steps: extracting the correlation characteristics among process parameters through a three-layer gated graph attention network, combining a time sequence attention module to construct a parameter coupling state vector; constructing a process prediction model based on the state vector, fusing a thermodynamic equilibrium equation and a long short-term memory network to predict oxygen and nitrogen contents; adopting a deep reinforcement learning network of a Soft Actor-Critic algorithm, taking the parameter coupling state vector and the oxygen and nitrogen content prediction values as state inputs, and outputting process parameter adjustment amounts; planning and iteratively training future adjustment strategies through Monte Carlo tree search, and finally obtaining optimal process parameter combinations that satisfy oxygen and nitrogen content constraints, so that intelligent optimization control of the metallurgical process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to process optimization technology, and more particularly to a method and system for optimizing the metallurgical process of low-oxygen and low-nitrogen aluminum-vanadium alloys. Background Technology

[0002] Controlling oxygen and nitrogen content in metallurgical processes has always been a challenge in the industry. Existing control methods mainly rely on empirical models and PID controllers, which are insufficient to handle the complex nonlinear coupling relationships between process parameters. The mechanisms by which process parameters such as temperature, pressure, and vacuum affect oxygen and nitrogen content are complex, and existing models cannot accurately describe their interactions, resulting in low model prediction accuracy. Furthermore, parameter optimization mostly adopts single-parameter adjustment strategies, ignoring the mutual influence between parameters and making it difficult to find the globally optimal parameter combination.

[0003] Existing technologies lack modeling methods for the dynamic correlation of process parameters, making it difficult to capture the time-varying coupling characteristics between parameters. Existing prediction models fail to effectively integrate theoretical knowledge with actual process data, resulting in large prediction biases when operating conditions change. Parameter optimization methods are mainly based on empirical rules or simple optimization algorithms, which are difficult to handle complex optimization problems with multiple parameters, multiple objectives, and multiple constraints, and lack the ability to predict long-term process trends.

[0004] Existing process parameter optimization strategies are typically static and cannot adapt to dynamic changes in the process. Metallurgical processes are complex and highly variable, and existing models struggle to accurately capture transient response characteristics, leading to significant discrepancies between predictions and actual operating conditions. Furthermore, the lack of effective decision-making and planning tools in current technologies hinders the optimization of process parameters across multiple time steps, resulting in blind operation and large fluctuations in production stability and product quality. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method and system for optimizing the metallurgical process of low-oxygen and low-nitrogen aluminum-vanadium alloys, which can solve the problems in existing technologies.

[0006] A first aspect of this invention provides a method for optimizing the metallurgical process of low-oxygen, low-nitrogen aluminum-vanadium alloys, comprising:

[0007] Temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally using a temporal attention module and residual connections to obtain the parameter coupling state vector.

[0008] A process prediction model is constructed based on the parameter-coupled state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, and uses a long short-time memory network to predict the actual oxygen and nitrogen content. The theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused to obtain the predicted oxygen and nitrogen content value.

[0009] A deep reinforcement learning network is constructed, with the parameter coupled state vector and the predicted oxygen and nitrogen content as the state space input, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment as the action space output. The Soft Actor-Critic algorithm structure is adopted, and the oxygen and nitrogen content is set to be less than the preset value in the reward function.

[0010] The parameter adjustment values ​​output from the action space are input into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

[0011] Optionally,

[0012] The steps of constructing parameter nodes from temperature, pressure, time, raw material ratio, and vacuum parameters, extracting the correlation features between these parameter nodes using a three-layer gated graph attention network, and performing temporal modeling of these correlation features through a temporal attention module and residual connections to obtain the parameter coupling state vector include:

[0013] Temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. The sensitivity values ​​of these parameter nodes to oxygen and nitrogen content are calculated. The parameter nodes are weighted based on the sensitivity values. The weighted parameter nodes are then input into a multi-head graph attention network. The attention coefficients of the multi-head graph attention network are adjusted using the sensitivity values. The correlation features between the parameter nodes are then extracted.

[0014] A dedicated subgraph attention network is constructed for the temperature, pressure, and vacuum parameters. In the message passing mechanism of the dedicated subgraph attention network, constraints on the inverse relationship between vacuum and pressure, the influence of temperature on vacuum, and the influence of pressure on oxygen and nitrogen solubility are introduced to extract the coupling features between the temperature, pressure, and vacuum parameters.

[0015] The associated features are divided into melting stage nodes, deoxidation stage nodes, and denitrification stage nodes. Temporal feature extractors are constructed for each of the melting stage nodes, deoxidation stage nodes, and denitrification stage nodes. The temporal feature extractor includes a temporal attention module and a residual connection module. The temporal attention module captures the temporal dependencies of parameters, and the residual connection module retains the original feature information to extract the temporal features of parameters at different stages.

[0016] The associated features, the coupling features, and the parameter temporal features are fused to obtain the parameter coupling state vector.

[0017] Optionally,

[0018] The steps of using a long short-term memory network to predict actual oxygen and nitrogen content, and then weighting and fusing the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content include:

[0019] A process prediction model is constructed based on the parameter coupling state vector, and the process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content.

[0020] The parameter-coupled state vector is decomposed using short-term, medium-term, and long-term time windows to obtain corresponding multi-scale state vectors. The multi-scale state vectors are then input into the corresponding long short-term memory networks to obtain the hidden layer states at different time scales. Inter-layer gating units are constructed, and feature interaction is performed on the hidden layer states at adjacent time scales through the inter-layer gating units to obtain multi-scale feature representations. The actual oxygen and nitrogen content is predicted based on the multi-scale feature representations.

[0021] Calculate the historical forecast mean absolute error, and construct the forecast confidence level based on the historical forecast mean absolute error; calculate the changes in temperature parameter, pressure parameter, and vacuum parameter, and construct the process parameter fluctuation index based on the changes in temperature parameter, pressure parameter, and vacuum parameter.

[0022] The multi-scale feature representation, the prediction confidence, and the process parameter fluctuation index are input into a multilayer perceptron network to obtain theoretical value weights; the theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused based on the theoretical value weights to obtain the predicted oxygen and nitrogen content.

[0023] A dynamic correction term is constructed based on the deviation between the measured oxygen and nitrogen content at the previous moment and the predicted oxygen and nitrogen content at the previous moment. The dynamic correction term is then superimposed on the predicted oxygen and nitrogen content to obtain the corrected predicted oxygen and nitrogen content.

[0024] Optionally,

[0025] Constructing a deep reinforcement learning network, using the parameter-coupled state vector and the predicted oxygen and nitrogen content as input to the state space, and using temperature, pressure, vacuum, and time parameter adjustments as outputs to the action space, employing a Soft Actor-Critic algorithm structure, the steps of setting constraints in the reward function that the oxygen and nitrogen content is less than a preset value include:

[0026] The parameter-coupled state vector and the predicted oxygen and nitrogen content are used as the state space input, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment are used as the action space output.

[0027] A deep reinforcement learning network is constructed using the Soft Actor-Critic algorithm, comprising a critic network and an actor network. The critic network employs a dual Q-value network structure, while the actor network employs a stochastic policy network structure. The dual Q-value network calculates the target Q-value based on the state space input and the action space output. The stochastic policy network predicts action distribution parameters based on the state space input and samples the optimized action space output from the action distribution.

[0028] A basic reward term is constructed based on the mean square error between the predicted oxygen content and the target oxygen content, and the mean square error between the predicted nitrogen content and the target nitrogen content. An action smoothing term is constructed based on the Euclidean distance of the action space output at adjacent time points. A penalty term is introduced when the oxygen and nitrogen contents are greater than the preset values. The reward function is obtained by weighting the basic reward term, the action smoothing term, and the penalty term.

[0029] A commenter network loss function is constructed, comprising an immediate reward term, a target Q-value term adjusted by a discount factor, and a policy entropy term adjusted by a temperature parameter. An actor network loss function is constructed based on the difference between the policy entropy adjusted by the temperature parameter and the target Q-value, wherein the policy entropy is used to achieve maximum entropy reinforcement learning. A temperature parameter loss function is constructed based on the difference between the policy entropy and the target entropy.

[0030] The state space input, action space output, immediate reward, and next-moment state are stored as transition samples in the experience replay buffer. Based on the transition samples in the experience replay buffer, the loss function of the critic network, the loss function of the actor network, and the temperature parameter loss function are optimized to update the parameters of the deep reinforcement learning network.

[0031] Optionally,

[0032] The steps for constructing an adaptive weight adjustment mechanism based on the reward function include:

[0033] Based on historical samples, oxygen content error sequences and nitrogen content error sequences are calculated. The mean of oxygen content error and the mean of nitrogen content error are calculated using the exponential weighted moving average method. The weight coefficients of oxygen content and nitrogen content items in the basic reward items are dynamically adjusted according to the mean error.

[0034] The penalty term includes a multi-level penalty function, which is divided into a first over-limit interval and a second over-limit interval according to the degree to which the oxygen and nitrogen content exceeds the preset value. A quadratic penalty function is used in the first over-limit interval, and an exponential penalty function is used in the second over-limit interval.

[0035] The rate of change of process parameters is calculated using a sliding time window. Based on the rate of change, velocity constraints and acceleration constraints are constructed. The velocity constraints and acceleration constraints are added to the reward function to achieve smooth adjustment of process parameters.

[0036] Optionally,

[0037] The steps of inputting the parameter adjustment amount output from the action space into the process prediction model for simulation verification, using Monte Carlo tree search to plan the adjustment strategy for future time steps, and iteratively training based on the planning results to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints include:

[0038] Input the parameter adjustment amount output from the action space into the process prediction model to obtain the oxygen and nitrogen content prediction values ​​under multiple simulation conditions;

[0039] A decision tree is constructed using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node. The decision tree is then expanded, and the node evaluation score is calculated based on the predicted oxygen and nitrogen content. The UCB value is obtained by combining the node visit count. The node with the highest UCB value in the future time step is selected for exploration. A single exploration is completed in four stages: selection, expansion, simulation, and backpropagation.

[0040] The path with the most visits is extracted from the decision tree as the optimal parameter adjustment sequence. The first step of the optimal parameter adjustment sequence is used as the parameter adjustment scheme at the current moment. The process parameters are updated to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints.

[0041] Optionally,

[0042] The steps of constructing a decision tree using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node, include:

[0043] Temperature, pressure, vacuum, time, and predicted oxygen and nitrogen content are used to construct a process state vector. Temperature, pressure, vacuum, and time adjustment values ​​are used to construct a parameter adjustment vector. Decision tree nodes are constructed based on the process state vector and the parameter adjustment vector.

[0044] An exponential evaluation function is constructed for the predicted oxygen content and the predicted nitrogen content respectively. A first evaluation value is calculated based on the deviation between the predicted oxygen content and the target oxygen content value, and a second evaluation value is calculated based on the deviation between the predicted nitrogen content and the target nitrogen content value. The Euclidean norm of the parameter adjustment vector is calculated to obtain a stability evaluation value. The first evaluation value, the second evaluation value, and the stability evaluation value are weighted to obtain a node evaluation score.

[0045] The exploration coefficient is calculated based on the number of search rounds, and the exploration coefficient decreases as the number of search rounds increases; the node evaluation score is used as the node value item, and the exploration coefficient is combined with the logarithm of the number of node visits to construct a confidence upper limit calculation formula, and the node UCB value is calculated based on the confidence upper limit calculation formula;

[0046] During the tree expansion phase, the node with the largest UCB value is selected for expansion, and child nodes are generated based on the physical constraints of parameter adjustment. During the tree search simulation phase, a combined evaluation function of immediate reward and future potential is constructed, and leaf nodes are quickly evaluated based on the combined evaluation function. During the backpropagation phase, the value and number of visits of the visited nodes are updated based on the results of the quick evaluation.

[0047] Secondly, it provides a metallurgical process optimization system for low-oxygen and low-nitrogen aluminum-vanadium alloys, including:

[0048] The first unit is used to construct parameter nodes from temperature parameters, pressure parameters, time parameters, raw material ratio parameters, and vacuum degree parameters. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally through a temporal attention module and residual connections to obtain a parameter coupling state vector.

[0049] The second unit is used to construct a process prediction model based on the parameter coupling state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, uses a long short-time memory network to predict the actual oxygen and nitrogen content, and weights and fuses the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content value.

[0050] The third unit is used to construct a deep reinforcement learning network. The parameter coupled state vector and the predicted oxygen and nitrogen content are used as input to the state space, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment are used as output to the action space. The Soft Actor-Critic algorithm structure is adopted, and the oxygen and nitrogen content is set to be less than the preset value in the reward function.

[0051] The fourth unit is used to input the parameter adjustment amount output from the action space into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

[0052] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0053] This invention employs a gated graph attention network and a temporal attention module to extract dynamic correlation features between parameters, accurately capturing the complex coupling relationships between parameters such as temperature, pressure, and vacuum, and establishing a dynamic influence network between parameters. By fusing parameter correlation features and temporal features, the resulting parameter coupling state vector comprehensively reflects the process state, providing high-quality feature representations for subsequent prediction and optimization, and significantly improving the accuracy and completeness of process state characterization.

[0054] This invention combines thermodynamic equilibrium equations and long short-term memory networks to achieve a deep integration of theoretical knowledge and data-driven approaches, overcoming the limitations of single models. The process prediction model can simultaneously consider theoretical equilibrium states and actual dynamic processes, exhibiting strong adaptability and high prediction accuracy. It provides a reliable basis for process parameter optimization and effectively solves the problem of large prediction deviations in traditional prediction methods under complex operating conditions.

[0055] This invention combines deep reinforcement learning networks with Monte Carlo tree search to achieve optimized planning of process parameters across multiple time steps, overcoming the limitations of traditional methods that only focus on immediate optimization. This method can find the globally optimal parameter combination while satisfying oxygen and nitrogen content constraints, and continuously optimizes the decision-making strategy through simulation verification and iterative training. This significantly improves the stability of metallurgical processes and product quality, while reducing energy consumption and material waste. Attached Figure Description

[0056] Figure 1 This is a schematic flowchart of the metallurgical process optimization method for low-oxygen and low-nitrogen aluminum-vanadium alloys according to an embodiment of the present invention.

[0057] Figure 2 This is a comparison chart showing the trends of oxygen and nitrogen content changes during the optimization process using different methods. Detailed Implementation

[0058] The technical solutions of the present invention will be described below with reference to the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0059] Figure 1 This is a schematic diagram of the optimized metallurgical process for low-oxygen, low-nitrogen aluminum-vanadium alloys according to the method of the present invention, as shown below. Figure 1 As shown, the method includes:

[0060] Temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally using a temporal attention module and residual connections to obtain the parameter coupling state vector.

[0061] A process prediction model is constructed based on the parameter-coupled state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, and uses a long short-time memory network to predict the actual oxygen and nitrogen content. The theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused to obtain the predicted oxygen and nitrogen content value.

[0062] A deep reinforcement learning network is constructed, with the parameter coupled state vector and the predicted oxygen and nitrogen content as the state space input, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment as the action space output. The SoftActor-Critic algorithm structure is adopted, and the oxygen and nitrogen content is set to be less than the preset value in the reward function.

[0063] The parameter adjustment values ​​output from the action space are input into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

[0064] Optionally,

[0065] The steps of constructing parameter nodes from temperature, pressure, time, raw material ratio, and vacuum parameters, extracting the correlation features between these parameter nodes using a three-layer gated graph attention network, and performing temporal modeling of these correlation features through a temporal attention module and residual connections to obtain the parameter coupling state vector include:

[0066] Temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. The sensitivity values ​​of these parameter nodes to oxygen and nitrogen content are calculated. The parameter nodes are weighted based on the sensitivity values. The weighted parameter nodes are then input into a multi-head graph attention network. The attention coefficients of the multi-head graph attention network are adjusted using the sensitivity values. The correlation features between the parameter nodes are then extracted.

[0067] A dedicated subgraph attention network is constructed for the temperature, pressure, and vacuum parameters. In the message passing mechanism of the dedicated subgraph attention network, constraints on the inverse relationship between vacuum and pressure, the influence of temperature on vacuum, and the influence of pressure on oxygen and nitrogen solubility are introduced to extract the coupling features between the temperature, pressure, and vacuum parameters.

[0068] The associated features are divided into melting stage nodes, deoxidation stage nodes, and denitrification stage nodes. Temporal feature extractors are constructed for each of the melting stage nodes, deoxidation stage nodes, and denitrification stage nodes. The temporal feature extractor includes a temporal attention module and a residual connection module. The temporal attention module captures the temporal dependencies of parameters, and the residual connection module retains the original feature information to extract the temporal features of parameters at different stages.

[0069] The associated features, the coupling features, and the parameter temporal features are fused to obtain the parameter coupling state vector.

[0070] In this embodiment, temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. The time parameter includes the duration of each process stage, and the raw material ratio parameter includes the percentage content of each element. Each parameter is normalized using a maximum-minimum standardization method, mapping the parameter value to the [0, 1] interval. Each parameter node is represented as a 64-dimensional vector. Initial values ​​are generated through a parameter embedding layer, which is a fully connected layer. The input is the normalized parameter value, and the output is a 64-dimensional feature vector.

[0071] The sensitivity values ​​of parameter nodes to oxygen and nitrogen content are calculated. The sensitivity calculation employs local perturbation analysis, applying a ±5% perturbation to each parameter and recording the variation in oxygen and nitrogen content. Analysis of historical data reveals the following sensitivity values: temperature 0.45, pressure 0.35, vacuum 0.38, time 0.25, and raw material ratio 0.28. These sensitivity values ​​are then normalized using Softmax and used as weighting coefficients for the parameter nodes. The eigenvectors of each parameter node are multiplied by their corresponding sensitivity values ​​to obtain a weighted eigenvector. For example, the temperature eigenvector is multiplied by 0.45, the pressure eigenvector by 0.35, and so on. This weighted eigenvector serves as the input to a multi-head graph attention network.

[0072] The weighted parameter nodes are input into a multi-head graph attention network. The network structure contains eight attention heads, each with a dimension of 32 and an output dimension of 256. During attention calculation, the input features are first converted into query vectors, key vectors, and value vectors using three weight matrices. Then, the dot product of the query vector and the key vector is calculated to obtain the original attention score. A sensitivity adjustment mechanism is introduced here, multiplying the original attention score by the parameter sensitivity value, and then normalizing it using the Softmax function to obtain the final attention coefficient. The outputs of each attention head are concatenated and linearly transformed to obtain the final output. The network consists of three stacked layers, with skip connections and layer normalization added between each layer to effectively prevent the gradient vanishing problem.

[0073] A dedicated subgraph attention network is constructed for temperature, pressure, and vacuum parameters. This network employs a graph neural network structure, where nodes represent three parameters and edges represent the physical relationships between them. Three physical constraints are introduced into the message passing mechanism: First, the inverse relationship between vacuum and pressure is implemented as follows: when calculating messages from pressure nodes to vacuum nodes, the message value is inversely proportional to the pressure value; if the pressure value doubles, the transmitted message value decreases by half. Second, the influence of temperature on vacuum is implemented as follows: as temperature increases, the message value transmitted to vacuum nodes increases, and the rate of increase accelerates with increasing temperature. Third, the influence of pressure on oxygen and nitrogen solubility is implemented as follows: according to Henry's Law, gas solubility is directly proportional to pressure; therefore, the message value transmitted from pressure nodes is directly proportional to the pressure value. Through these physical constraints, the subgraph network can learn parameter coupling relationships consistent with metallurgical theory.

[0074] The associated features are divided into melting stage nodes, deoxidation stage nodes, and denitrification stage nodes. This division is based on the time limits of the process stages: the melting stage is typically 0-30 minutes after the process starts, the deoxidation stage is 30-60 minutes, and the denitrification stage is 60-90 minutes. For each stage, a corresponding subset of parameters is selected as input based on its characteristics. The melting stage mainly focuses on temperature and raw material ratio parameters; the deoxidation stage mainly focuses on temperature, pressure, and vacuum parameters; and the denitrification stage mainly focuses on pressure, vacuum, and time parameters. Temporal feature extractors are constructed for different stage nodes. Each feature extractor includes a temporal attention module and a residual connection module. The temporal attention module uses a self-attention mechanism, taking a parameter sequence of 20 historical time steps as input, and obtaining attention weights by calculating the correlation between different time steps. In the specific implementation, a multi-head attention structure is used, with 4 heads, each with a dimension of 32. The residual connection module adds the original input features to the output features of the attention module, and then performs layer normalization to ensure stable feature distribution. In practical applications, the time-series feature extractor in the melting stage learns the relationship between the rate of temperature rise and the rate of raw material melting; the time-series feature extractor in the deoxidation stage captures the time lag relationship between vacuum fluctuations and changes in oxygen content; and the time-series feature extractor in the denitrification stage identifies the correlation between pressure maintenance time and the rate of nitrogen content decrease.

[0075] Feature fusion is performed on correlation features, coupling features, and time-series parameter features using an attention-weighted mechanism. First, a fully connected layer maps the three features to the same dimension (128-dimensional), then feature importance scores are calculated. The importance score for correlation features is 0.4, for coupling features it is 0.35, and for time-series features it is 0.25. The three features are weighted and summed according to their importance scores, then nonlinearly transformed through two fully connected layers (256 and 128 dimensions respectively), finally yielding a 128-dimensional parameter coupling state vector. Dimensions 1-32 of this state vector primarily represent temperature-related features, dimensions 33-64 primarily represent pressure and vacuum-related features, dimensions 65-96 primarily represent time-related features, and dimensions 97-128 primarily represent raw material ratio-related features.

[0076] This invention achieves accurate modeling of complex relationships between metallurgical process parameters. By integrating physical constraints and data-driven methods, it overcomes the problem of traditional models' inadequate capture of nonlinear relationships. The constructed parameter-coupled state vector comprehensively reflects the dynamic interaction relationships of process parameters, providing high-quality feature representations for subsequent process prediction and optimization. This effectively improves the accuracy of oxygen and nitrogen content control and process stability, reduces energy consumption and material waste, and is of great significance for improving the quality of metallurgical products.

[0077] Optionally,

[0078] The steps of using a long short-term memory network to predict actual oxygen and nitrogen content, and then weighting and fusing the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content include:

[0079] A process prediction model is constructed based on the parameter coupling state vector, and the process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content.

[0080] The parameter-coupled state vector is decomposed using short-term, medium-term, and long-term time windows to obtain corresponding multi-scale state vectors. The multi-scale state vectors are then input into the corresponding long short-term memory networks to obtain the hidden layer states at different time scales. Inter-layer gating units are constructed, and feature interaction is performed on the hidden layer states at adjacent time scales through the inter-layer gating units to obtain multi-scale feature representations. The actual oxygen and nitrogen content is predicted based on the multi-scale feature representations.

[0081] Calculate the historical forecast mean absolute error, and construct the forecast confidence level based on the historical forecast mean absolute error; calculate the changes in temperature parameter, pressure parameter, and vacuum parameter, and construct the process parameter fluctuation index based on the changes in temperature parameter, pressure parameter, and vacuum parameter.

[0082] The multi-scale feature representation, the prediction confidence, and the process parameter fluctuation index are input into a multilayer perceptron network to obtain theoretical value weights; the theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused based on the theoretical value weights to obtain the predicted oxygen and nitrogen content.

[0083] A dynamic correction term is constructed based on the deviation between the measured oxygen and nitrogen content at the previous moment and the predicted oxygen and nitrogen content at the previous moment. The dynamic correction term is then superimposed on the predicted oxygen and nitrogen content to obtain the corrected predicted oxygen and nitrogen content.

[0084] In this embodiment, a process prediction model is constructed based on a parameter-coupled state vector. The prediction model receives the parameter-coupled state vector as input and generates theoretical and actual oxygen and nitrogen content predictions through parallel theoretical calculation and data-driven branches, respectively. Key parameter information such as temperature, pressure, and vacuum degree in the parameter-coupled state vector is used for theoretical calculation of the thermodynamic equilibrium equation, while the complete 128-dimensional state vector is input to the data-driven branch for actual content prediction. Finally, the output results of the two branches are fused through an adaptive weighting mechanism to form the final prediction value.

[0085] The parameter coupling state vector is a 128-dimensional vector containing coupling information for parameters such as temperature, pressure, and vacuum. The thermodynamic equilibrium equation is based on Sieverts' law and the second law of thermodynamics, considering the relationship between the solubility of oxygen and nitrogen in the molten metal and temperature and pressure. In practice, for oxygen content calculation, the equilibrium constant K(O) under standard conditions is first obtained by looking up the temperature parameter in a table. Then, the oxygen partial pressure is calculated using the pressure and vacuum parameters to obtain the theoretical oxygen content. For nitrogen content calculation, the equilibrium constant K(N) is similarly obtained by looking up the temperature parameter in a table, and the nitrogen partial pressure is calculated using the pressure parameter to obtain the theoretical nitrogen content. For example, when the temperature is 1550℃, the pressure is 0.8MPa, and the vacuum is 50Pa, the calculated theoretical oxygen content is 12ppm and the theoretical nitrogen content is 35ppm.

[0086] The parameter-coupled state vector is decomposed into time-scaled sequences. The short-term time window uses data from the past 10 minutes with a sampling interval of 10 seconds; the medium-term time window uses data from the past 60 minutes with a sampling interval of 1 minute; and the long-term time window uses data from the past 6 hours with a sampling interval of 10 minutes. By using sliding windows with different sampling intervals, three different scale state vector sequences are obtained. These multi-scale state vectors are then input into corresponding Long Short-Term Memory (LSTM) networks. Each LSM network contains two layers: a first layer with 64 units and a second layer with 32 units. The network includes input gates, forget gates, output gates, and unit states. The input gates control the proportion of new information entering, the forget gates control the proportion of historical information retained, the output gates control the proportion of information output, and the unit states record long-term dependencies. Through forward propagation, each network outputs the hidden state corresponding to its time scale.

[0087] Inter-layer gating units are constructed to enable the interaction of features at different time scales. These units contain three gating mechanisms: an uplink gate, a downlink gate, and a lateral gate. The uplink gate controls the information flow from short-term features to medium-term features, the downlink gate controls the information flow from medium-term features to short-term features, and the lateral gate controls the interaction of features at the same time level. In practice, the similarity between the source and target features is first calculated. Then, the similarity is mapped to a value between 0 and 1 using the sigmoid function, which serves as the gating value. Finally, the gating value is multiplied by the source feature to obtain the transmitted information. For example, when short-term features indicate a rapid temperature increase while medium-term features indicate a stable temperature, the uplink gate increases, transmitting the abrupt change information to the medium-term network.

[0088] After feature interaction, the hidden state at different scales is concatenated and transformed through a fully connected layer to obtain a multi-scale feature representation. This feature representation has a dimension of 96 and includes information at short-term, medium-term, and long-term time scales. Based on the multi-scale feature representation, the actual oxygen and nitrogen content is predicted through a two-layer fully connected network. The first layer has 64 neurons, and the second layer has 2 neurons, corresponding to the predicted oxygen and nitrogen content values, respectively.

[0089] Calculate the mean absolute error (MAE) of historical predictions to construct prediction confidence levels. Select prediction results from the past 50 time points and calculate the MAE between predicted and measured values, separately for oxygen and nitrogen content. Convert the MAE into a confidence level value using an inverse proportional function; the smaller the error, the higher the confidence level. In practical applications, under stable operating conditions, the MAE for oxygen content prediction is approximately 1.5 ppm, corresponding to a confidence level of 0.85; the MAE for nitrogen content prediction is approximately 2.8 ppm, corresponding to a confidence level of 0.78.

[0090] First, the changes in temperature, pressure, and vacuum parameters over the past 10 time points are calculated. The absolute values ​​are then summed using weighted averages to obtain a comprehensive fluctuation index. The weights are assigned as follows: temperature 0.4, pressure 0.3, and vacuum 0.3. A larger fluctuation index value indicates more drastic changes in process conditions and a poorer applicability of the theoretical model. For example, when the temperature fluctuation is ±5℃, the pressure fluctuation is ±0.1MPa, and the vacuum fluctuation is ±10Pa, the calculated process parameter fluctuation index is 0.35.

[0091] Multi-scale feature representations, prediction confidence, and process parameter fluctuation indices are input into a multilayer perceptron network to obtain theoretical value weights. This network consists of two layers: the first layer has 32 neurons with the ReLU activation function; the second layer has 1 neuron with the Sigmoid activation function, and the output range is 0-1. When the prediction confidence is high and the process parameter fluctuation is small, the network tends to output smaller theoretical value weights; when the prediction confidence is low or the process parameter fluctuation is large, the network tends to output larger theoretical value weights. In a test case, the prediction confidence was 0.82, the process parameter fluctuation index was 0.25, and the calculated theoretical value weight was 0.35.

[0092] The theoretical and actual oxygen and nitrogen contents are weighted and fused based on theoretical values. The fusion formula is: Predicted value = Theoretical value weight × Theoretical value + (1 - Theoretical value weight) × Actual predicted value. For example, when the theoretical oxygen content is 12 ppm, the actual predicted oxygen content is 15 ppm, and the theoretical value weight is 0.35, the fused predicted oxygen content is 14.05 ppm.

[0093] Finally, a dynamic correction term is constructed to correct the prediction results in real time. A linear correction term is constructed based on the deviation between the measured and predicted oxygen and nitrogen content at the previous time step. The correction method is: corrected predicted value = current predicted value + attenuation coefficient × deviation at the previous time step. The attenuation coefficient is set to 0.7 to avoid overcorrection. For example, if the predicted oxygen content at the previous time step was 14 ppm, the measured value was 16 ppm, and the deviation was 2 ppm, then the correction term at the current time step would be 1.4 ppm. The current predicted value is 14.05 ppm, and the corrected value is 15.45 ppm.

[0094] This invention achieves high-precision prediction of oxygen and nitrogen content in metallurgical processes by integrating thermodynamic theoretical models with multi-scale deep learning methods. Multi-scale temporal feature extraction captures the patterns of process fluctuations at different time scales, an inter-layer gating mechanism enables cross-scale information exchange, an adaptive weighted fusion mechanism dynamically adjusts the weight ratio of the theoretical model and the data-driven model according to operating conditions, and a dynamic correction mechanism further improves prediction accuracy. By fully utilizing theoretical knowledge and historical data, it maintains high-precision prediction even under conditions of significant process fluctuations, providing a reliable basis for the precise control of metallurgical processes.

[0095] Optionally,

[0096] Constructing a deep reinforcement learning network, using the parameter-coupled state vector and the predicted oxygen and nitrogen content as input to the state space, and the temperature, pressure, vacuum, and time parameter adjustments as outputs to the action space, employing the SoftActor-Critic algorithm structure, and setting constraints in the reward function that the oxygen and nitrogen content is less than a preset value includes the following steps:

[0097] The parameter-coupled state vector and the predicted oxygen and nitrogen content are used as the state space input, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment are used as the action space output.

[0098] A deep reinforcement learning network is constructed using the SoftActor-Critic algorithm, comprising a critic network and an actor network. The critic network employs a dual Q-value network structure, while the actor network employs a stochastic policy network structure. The dual Q-value network calculates the target Q-value based on the state space input and the action space output. The stochastic policy network predicts action distribution parameters based on the state space input and samples the optimized action space output from the action distribution.

[0099] A basic reward term is constructed based on the mean square error between the predicted oxygen content and the target oxygen content, and the mean square error between the predicted nitrogen content and the target nitrogen content. An action smoothing term is constructed based on the Euclidean distance of the action space output at adjacent time points. A penalty term is introduced when the oxygen and nitrogen contents are greater than the preset values. The reward function is obtained by weighting the basic reward term, the action smoothing term, and the penalty term.

[0100] A commenter network loss function is constructed, comprising an immediate reward term, a target Q-value term adjusted by a discount factor, and a policy entropy term adjusted by a temperature parameter. An actor network loss function is constructed based on the difference between the policy entropy adjusted by the temperature parameter and the target Q-value, wherein the policy entropy is used to achieve maximum entropy reinforcement learning. A temperature parameter loss function is constructed based on the difference between the policy entropy and the target entropy.

[0101] The state space input, action space output, immediate reward, and next-moment state are stored as transition samples in the experience replay buffer. Based on the transition samples in the experience replay buffer, the loss function of the critic network, the loss function of the actor network, and the temperature parameter loss function are optimized to update the parameters of the deep reinforcement learning network.

[0102] In this embodiment, the state space input consists of two parts: a 128-dimensional parameter-coupled state vector and a 2-dimensional oxygen and nitrogen content prediction value, which are combined into a 130-dimensional state vector. The parameter-coupled state vector reflects the correlation between parameters such as temperature, pressure, and vacuum degree. The oxygen and nitrogen content prediction value includes the currently predicted oxygen and nitrogen content values, in ppm. The action space output is a 4-dimensional vector, corresponding to the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment, respectively. The temperature parameter adjustment range is [-10℃, +10℃], the pressure parameter adjustment range is [-0.2MPa, +0.2MPa], the vacuum degree parameter adjustment range is [-20Pa, +20Pa], and the time parameter adjustment range is [-5min, +5min]. All action values ​​are mapped to the [-1, 1] interval using the tanh function, and then multiplied by the corresponding adjustment range to obtain the actual adjustment amount.

[0103] A deep reinforcement learning network is constructed using the SoftActor-Critic algorithm. This network comprises two main components: a critic network and an actor network. The critic network employs a dual Q-value network structure, consisting of two structurally identical but parameter-independent Q-value networks. Each Q-value network comprises four fully connected layers with 256, 256, 128, and 1 neurons, respectively. The input is a concatenation of the state vector and action vector, and the output is the Q-value evaluation of that state-action pair. The actor network employs a stochastic policy network structure, consisting of four fully connected layers with 256, 256, 128, and 8 neurons, respectively. The input is the state vector, and the output is the action distribution parameters, including the mean and standard deviation of the four action dimensions. The mean is constrained to the range [-1, 1] using the tanh function, and the standard deviation is ensured to be positive using the Softplus function.

[0104] The state vector and action vector are concatenated and input into two Q-value networks to obtain two values, Q1 and Q2. The minimum of the two values ​​is then taken as the final Q-value. For example, in a certain state, the temperature adjustment is +5℃, the pressure adjustment is -0.1MPa, the vacuum adjustment is +10Pa, and the time adjustment is +2min. The calculated Q1 value is 15.3, and the Q2 value is 14.8. Therefore, the final Q-value is 14.8.

[0105] The stochastic policy network predicts action distribution parameters based on the state space input and samples the optimized action space output from the action distribution. Specifically, the state vector is first input into the actor network to obtain the mean and standard deviation of each action dimension; then, raw sampled values ​​are obtained from a normal distribution; next, a reparameterization technique is used to combine the raw sampled values, mean, and standard deviation to generate the final action value; finally, the action value is constrained to the range [-1, 1] using the tanh function. For example, if the mean of the temperature parameter adjustment is 0.45 and the standard deviation is 0.15, and the raw value obtained from the normal distribution is 0.2, then the final sampled temperature parameter adjustment is (0.45 + 0.15 × 0.2) mapped by the tanh function, which is approximately 0.47, corresponding to an actual temperature adjustment of +4.7℃.

[0106] The root mean square error (RMSE) between the predicted oxygen content and the target oxygen content, and the root mean square error between the predicted nitrogen content and the target nitrogen content, are calculated and multiplied by weighting coefficients of 0.6 and 0.4 respectively, then summed to obtain the basic reward term. For example, when the predicted oxygen content is 15 ppm and the target value is 10 ppm, and the predicted nitrogen content is 40 ppm and the target value is 30 ppm, the calculated basic reward term is -13.0. Euclidean distance is calculated based on the action space output of adjacent time steps to construct an action smoothing term. For example, if the temperature adjustment at the current time step is +5℃ and the previous time step was +3℃, the calculated temperature adjustment change is 2℃. Similarly, the changes in other parameters are calculated, and the resulting Euclidean distance is 3.5, which is multiplied by a weighting coefficient of -0.5 to obtain the action smoothing term -1.75. A penalty term is introduced when the oxygen and nitrogen contents exceed preset values. The preset oxygen content value is 12 ppm, and the preset nitrogen content value is 35 ppm. If the predicted oxygen content value exceeds the preset value, the penalty is the square of the excess multiplied by -5; if the predicted nitrogen content value exceeds the preset value, the penalty is the square of the excess multiplied by -4. For example, if the predicted oxygen content is 15 ppm and exceeds it by 3 ppm, the penalty is -45; if the predicted nitrogen content is 40 ppm and exceeds it by 5 ppm, the penalty is -100. Adding the base reward, smoothing effect, and penalty together yields a final reward of -159.75.

[0107] The commentator network loss function consists of three parts: an immediate reward term, a target Q-value term adjusted by a discount factor, and a policy entropy term adjusted by a temperature parameter. First, the immediate reward is calculated, as mentioned earlier, to be -159.75. Then, the target Q-value for the next state is calculated, assumed to be 10.5, multiplied by a discount factor of 0.99 to obtain 10.395. Next, the entropy value of the current policy is calculated, assumed to be 1.2, multiplied by the temperature parameter of 0.2 to obtain 0.24. Finally, the mean squared error between the target value and the current Q-value is calculated as the loss value. For example, if the target value is -159.75 + 10.395 + 0.24 = -149.115, and the current Q-value is -152.0, then the loss value is 8.32.

[0108] The actor network loss function is constructed based on the difference between the policy entropy adjusted by the temperature parameter and the target Q value. Actions are sampled from the current policy and input into the Q-value network to obtain the Q value. Then, the entropy value of the current policy is calculated and multiplied by the temperature parameter. Finally, the negative value of the weighted entropy value is subtracted from the Q value as the loss function. For example, if the Q value of the sampled action is 14.8, the policy entropy is 1.2, and the temperature parameter is 0.2, then the loss value is -(14.8 - 0.2 × 1.2) = -14.56.

[0109] Construct a temperature parameter loss function, and set the target entropy to the negative value of the action space dimension, i.e., -4. Calculate the difference between the current policy entropy and the target entropy, and multiply it by the temperature parameter to obtain the loss value. For example, if the current policy entropy is 1.2, the target entropy is -4, and the temperature parameter is 0.2, then the loss value is 0.2 × (1.2 - (-4)) = 1.04.

[0110] The state-space input, action-space output, immediate reward, and next-time state are stored as transition samples in the experience replay buffer. The buffer capacity is set to 100,000 and maintained using a first-in, first-out (FIFO) strategy. During each training iteration, 256 transition samples are randomly sampled from the buffer, and the loss for the commentator network, actor network, and temperature parameter are calculated. The parameters are then updated using the Adam optimizer. The learning rates for the commentator and actor networks are set to 0.0003, and the learning rate for the temperature parameter is set to 0.0001.

[0111] Figure 2 A comparative graph showing the trends of oxygen and nitrogen content changes during the optimization process using different methods is presented. The horizontal axis represents optimization time (minutes), and the vertical axis represents elemental content (ppm). The method of this invention demonstrates a significant advantage in reducing oxygen and nitrogen content. This method not only shows a significant advantage in the final oxygen and nitrogen content but also excels in the rate of reduction. The curves show a steeper downward trend for the method of this invention, indicating higher optimization efficiency. Furthermore, the method of this invention exhibits better stability after reaching the target content compared to the comparative methods, with a smaller fluctuation range, demonstrating better control precision. Comprehensive comparison shows that the deep reinforcement learning-based optimization method proposed in this invention, compared to traditional PID control and conventional DQN methods, has a faster optimization speed, lower final content value, and more stable control effect in optimizing the metallurgical process parameters of low-oxygen and low-nitrogen aluminum-vanadium alloys.

[0112] This invention achieves intelligent optimization of metallurgical process parameters through the SoftActor-Critic algorithm. It combines parameter-coupled state vectors with oxygen and nitrogen content prediction to construct a state representation capable of capturing complex relationships between process parameters. The structural design employing a dual Q-value network and a stochastic policy network effectively avoids overestimation and policy degradation. Through a multi-objective reward function design, it ensures smooth parameter adjustments while guaranteeing that oxygen and nitrogen content requirements are met. The maximum entropy reinforcement learning strategy promotes an exploration-exploitation balance, improving the algorithm's robustness and generalization ability, and providing reliable intelligent decision support for the precise control of metallurgical process parameters.

[0113] Optionally,

[0114] The steps for constructing an adaptive weight adjustment mechanism based on the reward function include:

[0115] Based on historical samples, oxygen content error sequences and nitrogen content error sequences are calculated. The mean of oxygen content error and the mean of nitrogen content error are calculated using the exponential weighted moving average method. The weight coefficients of oxygen content and nitrogen content items in the basic reward items are dynamically adjusted according to the mean error.

[0116] The penalty term includes a multi-level penalty function, which is divided into a first over-limit interval and a second over-limit interval according to the degree to which the oxygen and nitrogen content exceeds the preset value. A quadratic penalty function is used in the first over-limit interval, and an exponential penalty function is used in the second over-limit interval.

[0117] The rate of change of process parameters is calculated using a sliding time window. Based on the rate of change, velocity constraints and acceleration constraints are constructed. The velocity constraints and acceleration constraints are added to the reward function to achieve smooth adjustment of process parameters.

[0118] In this embodiment, an adaptive weight adjustment mechanism for the above-mentioned reward function is provided to optimize the oxygen and nitrogen content control process in metallurgical processes.

[0119] Data from the most recent 50 process cycles are selected as historical samples. Each cycle includes the target value, actual value, target value, and actual value of oxygen and nitrogen content. The oxygen content error (actual value minus target value) and nitrogen content error (actual value minus target value) for each cycle are calculated. For example, if the target oxygen content for a certain cycle is 10 ppm and the actual value is 13 ppm, the oxygen content error is 3 ppm; if the target nitrogen content is 30 ppm and the actual value is 34 ppm, the nitrogen content error is 4 ppm. This method yields oxygen and nitrogen content error sequences for 50 cycles. The exponentially weighted moving average method assigns higher weight to recent data, making the calculation results more reflective of the current process status. In specific implementation, a smoothing factor of 0.9 is chosen, and the calculation process is as follows: First, the first value in the error sequence is taken as the initial mean; then, for each new value in the sequence, the new mean is equal to the current value multiplied by (1 minus the smoothing factor) plus the old mean multiplied by the smoothing factor. In the above example, it is assumed that the calculated mean oxygen content error is 2.5 ppm and the mean nitrogen content error is 3.8 ppm.

[0120] The weight coefficients of the oxygen and nitrogen content items in the basic reward items are dynamically adjusted based on the mean error. The adjustment principle is that content items with larger mean errors should be assigned higher weights to strengthen the optimization of those items. The weight adjustment uses a normalization method, dividing the mean error of oxygen content and the mean error of nitrogen content by their sum to obtain the normalized weights. In the example above, the weight of the oxygen content item is 2.5 / (2.5+3.8)=0.4, and the weight of the nitrogen content item is 3.8 / (2.5+3.8)=0.6. To avoid excessive weight bias towards any one item, a limit is set on the range of weight changes, with a single adjustment not exceeding ±0.1, and the weight range limited to [0.3, 0.7]. Through this dynamic adjustment mechanism, the reinforcement learning algorithm can adaptively focus on content indicators with poor current control performance.

[0121] A multi-level penalty function is constructed to refine the penalty strategy for exceeding limits. Based on the degree to which oxygen and nitrogen content exceeds preset values, the system is divided into a first exceedance interval and a second exceedance interval. For oxygen content, the preset value is 12 ppm, the first exceedance interval is 12-15 ppm, and the second exceedance interval is greater than 15 ppm; for nitrogen content, the preset value is 35 ppm, the first exceedance interval is 35-40 ppm, and the second exceedance interval is greater than 40 ppm. A quadratic penalty function is used in the first exceedance interval. Specifically, the penalty value equals the square of the excess amount multiplied by a penalty coefficient. The penalty coefficient is set to oxygen content - 5 and nitrogen content - 4. For example, when the oxygen content is 14 ppm, exceeding the preset value by 2 ppm, the penalty value is 2. 2 ×(-5)=-20; When the nitrogen content is 38ppm, exceeding the preset value by 3ppm, the penalty value is 3. 2 ×(-4)=-36. The quadratic penalty function applies a lighter penalty for smaller exceedances and a heavier penalty for larger exceedances, prompting the algorithm to prioritize larger exceedances. An exponential penalty function is used in the second exceedance interval. Specifically, the penalty value equals the base penalty value multiplied by 1.5 raised to the power of the exceedance amount. The base penalty value is -20 for oxygen content and -16 for nitrogen content. For example, when the oxygen content is 17 ppm, exceeding the upper limit of the first exceedance interval by 2 ppm, the penalty value is -20 × 1.5. 2 ≈-45; When the nitrogen content is 43 ppm, exceeding the upper limit of the first over-limit range by 3 ppm, the penalty value is -16 × 1.5 3 ≈-54. The exponential penalty function imposes extremely heavy penalties on severe over-limit cases, ensuring that the algorithm prioritizes avoiding severe over-limits.

[0122] A sliding time window is used to calculate the rate of change of process parameters, constructing velocity and acceleration constraints. The sliding window size is set to 10 time steps, with a time step size of 10 seconds. For each parameter (temperature, pressure, vacuum, and time), its rate of change within the window is calculated. In practice, the parameter values ​​at two adjacent time points within the window are selected, the difference is calculated, and then divided by the time interval to obtain the rate of change at each time point. For example, if the temperature parameter is 1550℃ at time t and 1552℃ at time t+10, then the temperature change rate at time t+10 is 0.2℃ / second. The velocity constraint is used to limit the parameter adjustment speed to prevent excessively rapid changes from causing process instability. A maximum allowable rate of change is set for each parameter: 0.3℃ / second for temperature, 0.02MPa / second for pressure, 1Pa / second for vacuum, and 0.1 minutes / second for time. When the actual rate of change exceeds the maximum allowable value, the square of the excess is multiplied by a penalty coefficient as the velocity constraint penalty value. The penalty coefficient is set to -10. For example, if the temperature change rate is 0.4℃ / second, exceeding the maximum allowable value of 0.1℃ / second, the penalty value is 0.1. 2 ×(-10) = -0.1. The acceleration constraint term is used to limit the acceleration during parameter adjustment, preventing excessively drastic parameter changes. The difference in the rate of change between adjacent time points is calculated to obtain the changing acceleration. The maximum allowable acceleration is set to 0.05℃ / second for the temperature parameter. 2 The pressure parameter is 0.005 MPa / s. 2 The vacuum level parameter is 0.2 Pa / s. 2 The time parameter is 0.02 minutes / second. 2 When the actual acceleration exceeds the maximum allowable value, the square of the excess portion multiplied by a penalty coefficient is used as the acceleration constraint penalty value. The penalty coefficient is set to -15. For example, if the temperature change rate at time t is 0.2℃ / second, the rate of change at time t+10 is 0.3℃ / second, and the acceleration is 0.01℃ / second. 2 The penalty value is 0, as the maximum allowable value is not exceeded. Velocity and acceleration constraints are added to the reward function to achieve smooth adjustment of process parameters. The optimized reward function includes: a basic reward term, a motion smoothing term, a penalty term, a velocity constraint term, and an acceleration constraint term. The weights of each term are allocated as follows: basic reward term 1.0, motion smoothing term 0.5, penalty term 1.0, velocity constraint term 0.8, and acceleration constraint term 0.6. Through this combination of weight settings, the reinforcement learning algorithm can achieve smooth adjustment of process parameters while ensuring that oxygen and nitrogen content meet the standards.

[0123] In practical applications, when the oxygen content fluctuates significantly during the steelmaking process, the system will automatically increase the weight of the oxygen content term to strengthen the control of oxygen content; when the nitrogen content is detected to be seriously excessive, the multi-level penalty function will apply extremely heavy penalties, forcing the system to quickly adjust the process parameters; when the parameter adjustment is too drastic, the speed constraint term and acceleration constraint term will generate negative rewards, guiding the system to adopt a more gradual adjustment strategy.

[0124] This invention achieves intelligent optimization of the reinforcement learning reward function by constructing an adaptive weight adjustment mechanism. Dynamic weight adjustment enables the system to flexibly adjust the optimization objective according to the current process state, while multi-level penalty functions finely manage exceedance situations, and velocity and acceleration constraints ensure the stability of parameter adjustments. This multi-level, adaptive reward design significantly improves the applicability and effectiveness of reinforcement learning algorithms in metallurgical process optimization, achieving precise control of oxygen and nitrogen content and stable adjustment of process parameters, providing strong support for improving metallurgical product quality and optimizing production efficiency.

[0125] Optionally,

[0126] The steps of inputting the parameter adjustment amount output from the action space into the process prediction model for simulation verification, using Monte Carlo tree search to plan the adjustment strategy for future time steps, and iteratively training based on the planning results to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints include:

[0127] Input the parameter adjustment amount output from the action space into the process prediction model to obtain the oxygen and nitrogen content prediction values ​​under multiple simulation conditions;

[0128] A decision tree is constructed using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node. The decision tree is then expanded, and the node evaluation score is calculated based on the predicted oxygen and nitrogen content. The UCB value is obtained by combining the node visit count. The node with the highest UCB value in the future time step is selected for exploration. A single exploration is completed in four stages: selection, expansion, simulation, and backpropagation.

[0129] The path with the most visits is extracted from the decision tree as the optimal parameter adjustment sequence. The first step of the optimal parameter adjustment sequence is used as the parameter adjustment scheme at the current moment. The process parameters are updated to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints.

[0130] In this embodiment, the parameter adjustment amount obtained from reinforcement learning is input into the process prediction model for simulation verification. The adjustment strategy for multiple future time steps is planned through Monte Carlo tree search, thereby obtaining the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

[0131] The action space output includes adjustments for temperature, pressure, vacuum, and time parameters. For example, the action output at a certain moment might be: temperature adjustment +5℃, pressure adjustment -0.1MPa, vacuum adjustment +10Pa, and time adjustment +2min. Applying these adjustments to the current process parameters (assuming a current temperature of 1550℃, pressure of 1.0MPa, vacuum of 50Pa, and process time of 60min), the adjusted parameter values ​​are: temperature 1555℃, pressure 0.9MPa, vacuum 60Pa, and process time 62min. The adjusted parameters are then input into the process prediction model to predict the adjusted oxygen and nitrogen content. The predicted results are assumed to be: adjusted oxygen content of 11ppm and nitrogen content of 32ppm. To increase the reliability of the prediction results, multiple simulations are required for verification. By adding small random perturbations to the parameter adjustments (temperature ±2℃, pressure ±0.05MPa, vacuum ±5Pa, time ±1min), 10 sets of simulation conditions were generated, and 10 sets of predicted oxygen and nitrogen content values ​​were obtained. For example, the simulation results show that the predicted oxygen content range is 10-12ppm, and the predicted nitrogen content range is 30-34ppm.

[0132] A Monte Carlo tree search algorithm is used to construct a decision tree. The current process state includes the current process parameters (temperature, pressure, vacuum, and time) and the predicted oxygen and nitrogen content. For example, the root node state is: temperature 1550℃, pressure 1.0MPa, vacuum 50Pa, process time 60min, predicted oxygen content 13ppm, and predicted nitrogen content 36ppm. Parameter adjustments are used as optional branch nodes. Based on the output of the reinforcement learning network, five possible adjustment values ​​are generated for each parameter. For example, the adjustment options for the temperature parameter are {-5℃, -2.5℃, 0℃, +2.5℃, +5℃}, the adjustment options for the pressure parameter are {-0.2MPa, -0.1MPa, 0MPa, +0.1MPa, +0.2MPa}, the adjustment options for the vacuum parameter are {-20Pa, -10Pa, 0Pa, +10Pa, +20Pa}, and the adjustment options for the time parameter are {-5min, -2.5min, 0min, +2.5min, +5min}. These adjustment options form the branch nodes of the decision tree, with each node representing a possible combination of parameter adjustments.

[0133] The decision tree is expanded, and node evaluation scores are calculated based on the predicted oxygen and nitrogen content values. The node evaluation score reflects the quality of the parameter adjustment scheme corresponding to that node. The scoring calculation considers three factors: the degree to which the oxygen content approaches the target value, the degree to which the nitrogen content approaches the target value, and the smoothness of the parameter adjustment. Specifically, the following steps are taken: first, the square of the difference between the predicted oxygen content value and the target value (10 ppm) is calculated and multiplied by a weight of 0.6; then, the square of the difference between the predicted nitrogen content value and the target value (30 ppm) is calculated and multiplied by a weight of 0.4; finally, the sum of the squares of the parameter adjustments is calculated and multiplied by a smoothing factor of -0.2; these three items are then added together to obtain the node evaluation score. The higher the score, the better the adjustment scheme. For example, if a node's parameters are adjusted to predict an oxygen content of 11 ppm, a nitrogen content of 32 ppm, a temperature adjustment of +2.5℃, a pressure adjustment of -0.1 MPa, a vacuum adjustment of +10 Pa, and a time adjustment of 0 min, the calculated evaluation score is -3.8. The UCB value balances the node evaluation score (utilization) and the number of node visits (exploration). The specific calculation method is as follows: take the node evaluation score as the first term; divide the logarithm of the number of visits to the parent node by the number of visits to the node and then take the square root, multiply by the exploration coefficient 1.4, and take it as the second term; add the two terms together to get the UCB value. For example, if a node has an evaluation score of -3.8, a number of visits to the node is 5, and a number of visits to the parent node is 20, the calculated UCB value is -3.8 + 1.4 × square root (logarithm (20) / 5) ≈ -2.1. In each iteration, Monte Carlo tree search starts from the current node and selects the child node with the highest UCB value until a leaf node or an unexpanded node is reached. For example, starting from the root node, select the node with a UCB value of -2.1 as the next node to be explored. The parameters corresponding to this node are adjusted as follows: temperature +2.5℃, pressure -0.1MPa, vacuum degree +10Pa, time 0min.

[0134] The selection phase chooses an exploration path based on the UCB value; the expansion phase adds a new node at the end of the exploration path; the simulation phase starts from the new node and uses a fast simulation strategy (e.g., randomly selecting parameters for adjustment) to simulate the results for multiple future time steps; the backpropagation phase updates the simulation results to all nodes on the exploration path. In practice, the simulation depth is set to 5 time steps. After each simulation, the visit count of a node on the path is incremented by 1, and the node's evaluation score is updated to a weighted average of the original score and the new simulated score. For example, if a node's original score is -3.8 and its simulated score is -3.2, the updated score is 0.8 × (-3.8) + 0.2 × (-3.2) = -3.68.

[0135] Repeat the Monte Carlo tree search until the preset number of iterations or computation time limit is reached. The number of iterations is set to 1000, or the computation time limit is 2 seconds, whichever comes first. Through numerous iterations, the decision tree continuously expands and optimizes, and the UCB value gradually converges to a more accurate estimate.

[0136] The number of visits reflects how frequently the path is explored, and indirectly reflects its quality. For example, after 1000 iterations, the path with the most visits is: root node → (temperature +2.5℃, pressure -0.1MPa, vacuum +10Pa, time 0min) → (temperature +2.5℃, pressure -0.1MPa, vacuum +5Pa, time 0min) → (temperature 0℃, pressure 0MPa, vacuum 0Pa, time +2.5min) → (temperature -2.5℃, pressure +0.1MPa, vacuum -10Pa, time 0min) → (temperature 0℃, pressure 0MPa, vacuum 0Pa, time 0min). This path represents the optimal parameter adjustment sequence for the next 5 time steps. Taking the first step of the optimal parameter adjustment sequence as the parameter adjustment scheme for the current moment, according to the example above, the optimal parameter adjustment for the current moment is: temperature +2.5℃, pressure -0.1MPa, vacuum +10Pa, time 0min. Applying these adjustments to the current process parameters yields the updated parameters: temperature 1552.5℃, pressure 0.9MPa, vacuum 60Pa, and process time 60min. Verification using the process prediction model shows that the predicted oxygen and nitrogen contents for the updated parameters are: oxygen 11ppm and nitrogen 32ppm, satisfying the constraints of oxygen <12ppm and nitrogen <35ppm. The final optimal combination of process parameters is determined to be: temperature 1552.5℃, pressure 0.9MPa, vacuum 60Pa, and process time 60min.

[0137] This invention utilizes Monte Carlo Tree Search (MCS) to achieve strategy planning for future multi-timestep metallurgical processes, effectively overcoming the limitations of traditional optimization methods that only consider immediate optimization. This method combines reinforcement learning with predictive programming, constructing a decision tree and dynamically evaluating node value to achieve global optimization of process parameter adjustment paths. The UCB strategy balances exploration and exploitation, improving search efficiency and optimization quality. An evaluation mechanism based on extensive simulation validation enhances the reliability and stability of the optimization results. The final combination of process parameters simultaneously satisfies oxygen and nitrogen content constraints and parameter adjustment smoothness requirements, providing reliable decision support for precise control of metallurgical processes.

[0138] Optionally,

[0139] The steps of constructing a decision tree using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node, include:

[0140] Temperature, pressure, vacuum, time, and predicted oxygen and nitrogen content are used to construct a process state vector. Temperature, pressure, vacuum, and time adjustment values ​​are used to construct a parameter adjustment vector. Decision tree nodes are constructed based on the process state vector and the parameter adjustment vector.

[0141] An exponential evaluation function is constructed for the predicted oxygen content and the predicted nitrogen content respectively. A first evaluation value is calculated based on the deviation between the predicted oxygen content and the target oxygen content value, and a second evaluation value is calculated based on the deviation between the predicted nitrogen content and the target nitrogen content value. The Euclidean norm of the parameter adjustment vector is calculated to obtain a stability evaluation value. The first evaluation value, the second evaluation value, and the stability evaluation value are weighted to obtain a node evaluation score.

[0142] The exploration coefficient is calculated based on the number of search rounds, and the exploration coefficient decreases as the number of search rounds increases; the node evaluation score is used as the node value item, and the exploration coefficient is combined with the logarithm of the number of node visits to construct a confidence upper limit calculation formula, and the node UCB value is calculated based on the confidence upper limit calculation formula;

[0143] During the tree expansion phase, the node with the largest UCB value is selected for expansion, and child nodes are generated based on the physical constraints of parameter adjustment. During the tree search simulation phase, a combined evaluation function of immediate reward and future potential is constructed, and leaf nodes are quickly evaluated based on the combined evaluation function. During the backpropagation phase, the value and number of visits of the visited nodes are updated based on the results of the quick evaluation.

[0144] In this embodiment, the process state vector contains six elements: temperature parameter (°C), pressure parameter (MPa), vacuum parameter (Pa), time parameter (min), predicted oxygen content (ppm), and predicted nitrogen content (ppm). For example, the process state vector at a certain moment is [1550, 0.8, 50, 60, 13, 36], indicating that the current temperature is 1550°C, the pressure is 0.8 MPa, the vacuum is 50 Pa, the process time is 60 minutes, the predicted oxygen content is 13 ppm, and the predicted nitrogen content is 36 ppm. The parameter adjustment vector contains four elements, corresponding to the adjustment range of the four process parameters. According to the metallurgical process characteristics, the adjustment range of each parameter is set as follows: the temperature parameter adjustment range is [-10°C, +10°C], the pressure parameter adjustment range is [-0.2 MPa, +0.2 MPa], the vacuum parameter adjustment range is [-20 Pa, +20 Pa], and the time parameter adjustment range is [-5 min, +5 min]. For example, a set of parameter adjustment vectors is [+5, -0.1, +10, +2], which means that the temperature parameter is adjusted by +5℃, the pressure parameter by -0.1MPa, the vacuum parameter by +10Pa, and the time parameter by +2min.

[0145] Decision tree nodes are constructed based on process state vectors and parameter adjustment vectors. Each decision tree node contains the following attributes: state attribute, action attribute, visit count attribute, value attribute, parent node attribute, and child node attribute. The state attribute stores the current process state vector; the action attribute stores the parameter adjustment vector that caused the current state; the visit count attribute records the number of times the node has been visited, with an initial value of 0; the value attribute stores the node's evaluation score, with an initial value of 0; the parent node attribute points to the parent node; and the child node attribute stores a list of all possible child nodes, initially empty. The root node's action attribute is empty, representing the initial state.

[0146] Exponential evaluation functions were constructed for the predicted oxygen and nitrogen contents, respectively. These functions impose a stronger penalty on values ​​exceeding the target, encouraging the algorithm to prioritize addressing overshooting. Specifically, when the predicted value is less than the target value, the evaluation value is equal to the square of the difference multiplied by a coefficient; when the predicted value is greater than the target value, the evaluation value is equal to the square of the difference multiplied by the coefficient and then multiplied by 1.5 raised to the power of the difference. The coefficients were set to -0.6 for oxygen and -0.4 for nitrogen.

[0147] The first evaluation value is calculated based on the deviation between the predicted oxygen content and the target oxygen content. The target oxygen content is set at 10 ppm. The first evaluation value is calculated using the exponential evaluation function described above. In practical applications, when the predicted oxygen content is in the range of 7-13 ppm, the first evaluation value is between -12.2 and -5.4; when the predicted oxygen content exceeds 15 ppm, the first evaluation value drops sharply to below -30, reflecting a strong penalty for severe exceedances. The second evaluation value is calculated based on the deviation between the predicted nitrogen content and the target nitrogen content. The target nitrogen content is set at 30 ppm. The second evaluation value is calculated using the same exponential evaluation function. In practical applications, when the predicted nitrogen content is in the range of 25-35 ppm, the second evaluation value is between -8.0 and -4.0; when the predicted nitrogen content exceeds 40 ppm, the second evaluation value drops to below -20. The Euclidean norm of the parameter adjustment vector is calculated to obtain the stability evaluation value. The Euclidean norm is calculated as follows: divide each parameter adjustment by its maximum adjustment range to obtain a normalized value, and then calculate the square root of the sum of the squares of the normalized values. To ensure comparability of adjustments for different parameters, divide the temperature parameter by 10, the pressure parameter by 0.2, the vacuum parameter by 20, and the time parameter by 5. For example, the normalized value of the parameter adjustment vector [+5, -0.1, +10, +2] is [0.5, 0.5, 0.5, 0.4], and the Euclidean norm is 0.95. The stability assessment value equals the Euclidean norm multiplied by the stability coefficient -5, i.e., -4.75, representing the penalty for large parameter adjustments. The first assessment value, the second assessment value, and the stability assessment value are weighted to obtain the node assessment score. The weighting method is simple summation, as the three assessment values ​​already include their respective weight coefficients. For example, if the first assessment value is -8.5, the second assessment value is -6.0, and the stability assessment value is -4.75, then the node assessment score is -19.25. A higher evaluation score indicates a better parameter adjustment scheme for that node.

[0148] The exploration coefficient is calculated based on the number of search rounds. This coefficient balances exploration and utilization, with an initial value of 1.5 that decreases as the number of search rounds increases, thus guiding the algorithm from extensive initial exploration to deeper utilization in later stages. Specifically, the exploration coefficient equals the initial value multiplied by an exponential decay factor, where the decay factor is equal to 0.999 raised to the power of the number of search rounds.

[0149] The node evaluation score is used as the node value term, and the confidence upper limit formula is constructed by combining the exploration coefficient and the logarithm of the node visit count. The confidence upper limit formula is the core of the Monte Carlo Tree Search algorithm, used to balance node value (utilization) and exploration level (exploration). The UCB value calculation formula: the first term is the node evaluation score, representing the known node value; the second term is the exploration reward, used to encourage nodes with fewer visits. For example, for a process parameter adjustment scheme, assuming: the node evaluation score is -8.5 (oxygen content evaluation -4.2, nitrogen content evaluation -3.0, stability evaluation -1.3), the node visit count is 8, the parent node visit count is 40, and the exploration coefficient is 1.4, then the UCB value is calculated as follows:

[0150] ;

[0151] When the node visit count is 0, to avoid division by zero, the UCB value is set to positive infinity, ensuring that each new node is visited at least once. As the node visit count increases, the exploration reward gradually decreases, and the algorithm tends to select nodes with higher evaluation scores.

[0152] During the search tree expansion phase, the node with the largest UCB value is selected for expansion. Starting from the root node, the child node with the largest UCB value is selected at each level until a leaf node or an incompletely expanded node is reached. For example, if the root node has 5 child nodes with UCB values ​​of -17.6, -18.2, -18.7, -19.0, and -19.5, the node with a UCB value of -17.6 is selected to continue the search downwards. Physical constraints include parameter adjustment range limits and inter-parameter correlation constraints. Parameter adjustment range limits ensure that each parameter does not exceed the set range; inter-parameter correlation constraints reflect the physical relationship between parameters, such as the vacuum level not increasing significantly when the temperature rises. In practice, 5 possible adjustment amounts are set for each parameter, such as the adjustment amount options for the temperature parameter being {-5℃, -2.5℃, 0℃, +2.5℃, +5℃}. Considering all parameter combinations, there are theoretically 625 possible adjustment schemes, but by applying physical constraints, approximately 50 effective adjustment schemes are selected as child nodes. In the tree search simulation phase, a combined evaluation function of immediate reward and future potential is constructed. The immediate reward is based on the current node's evaluation score, and the future potential is estimated by quickly simulating the results of the next few steps. The quick simulation employs lightweight strategies, such as randomly selecting adjustment schemes or using simplified heuristics. The simulation depth is set to 3-5 steps, balancing evaluation accuracy and computational efficiency. The combined evaluation function is: immediate reward multiplied by 0.7 plus future potential multiplied by 0.3. For example, if a node's immediate reward is -19.25 and the simulated future potential is -15.0, then the combined evaluation value is -19.25 × 0.7 + (-15.0) × 0.3 = -17.98. In the backpropagation phase, the value and number of visits to the visited nodes are updated based on the quick evaluation results. Starting from the leaf nodes, the information of each node is updated upwards along the search path. The update rule is: the number of visits is incremented by 1; the node value is updated as a weighted average of the original value and the new evaluation value, with the weight related to the number of visits. Specifically, the new value equals the original value multiplied by the old number of visits plus the new evaluation value, then divided by the new number of visits. For example, a node originally had a value of -20.0 and an original number of visits of 10. Its new value is -17.98. After the update, its value is (-20.0×10+(-17.98)) / (10+1)≈-19.82, and its number of visits is updated to 11.

[0153] This invention achieves intelligent decision-making for optimizing metallurgical process parameters by constructing a decision tree based on Monte Carlo tree search. An exponential evaluation function imposes strong penalties for exceeding limits, ensuring precise control of oxygen and nitrogen content. Stability assessment of parameter adjustments promotes smooth adjustment of process parameters. Dynamic exploration coefficients and the UCB formula balance exploration and utilization, improving search efficiency. The introduction of physical constraints ensures that the generated adjustment schemes conform to process characteristics. Combined evaluation functions and a rapid simulation mechanism enable effective evaluation of future states. This multi-level, adaptive decision-making mechanism provides strong support for the intelligent optimization of metallurgical process parameters, achieving precise control of oxygen and nitrogen content and stable and efficient process operation.

[0154] Secondly, it provides a metallurgical process optimization system for low-oxygen and low-nitrogen aluminum-vanadium alloys, including:

[0155] The first unit is used to construct parameter nodes from temperature parameters, pressure parameters, time parameters, raw material ratio parameters, and vacuum degree parameters. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally through a temporal attention module and residual connections to obtain a parameter coupling state vector.

[0156] The second unit is used to construct a process prediction model based on the parameter coupling state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, uses a long short-time memory network to predict the actual oxygen and nitrogen content, and weights and fuses the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content value.

[0157] The third unit is used to construct a deep reinforcement learning network. The parameter coupled state vector and the predicted oxygen and nitrogen content are used as input to the state space, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment are used as output to the action space. The Soft Actor-Critic algorithm structure is adopted, and the oxygen and nitrogen content is set to be less than the preset value in the reward function.

[0158] The fourth unit is used to input the parameter adjustment amount output from the action space into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

[0159] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

Claims

1. A method for optimizing the metallurgical process of low-oxygen, low-nitrogen aluminum-vanadium alloys, characterized in that, include: Temperature, pressure, time, raw material ratio, and vacuum parameters are constructed as parameter nodes. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally using a temporal attention module and residual connections to obtain the parameter coupling state vector. A process prediction model is constructed based on the parameter-coupled state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, and uses a long short-time memory network to predict the actual oxygen and nitrogen content. The theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused to obtain the predicted oxygen and nitrogen content value. A deep reinforcement learning network is constructed, using the parameter-coupled state vector and the predicted oxygen and nitrogen content as the state space input, and the temperature, pressure, vacuum, and time parameter adjustments as the action space output. A Soft Actor-Critic algorithm structure is adopted, with a constraint condition that the oxygen and nitrogen content is less than a preset value set in the reward function. This includes: constructing a deep reinforcement learning network using the Soft Actor-Critic algorithm structure, comprising a critic network and an actor network. The critic network uses a double Q-value network structure, and the actor network uses a random policy network structure. The double Q-value network calculates the target Q-value based on the state space input and action space output. The random policy network predicts action distribution parameters based on the state space input and samples the optimized action space output from the action distribution. A basic reward term is constructed based on the mean squared error between the oxygen and nitrogen content and the target value, and an action smoothing term is constructed based on the Euclidean distance between action space outputs at adjacent time points. A penalty term is introduced when the oxygen and nitrogen content is greater than the preset value. The basic reward term and action smoothing term are then combined. A reward function is obtained by weighting the slip and penalty terms; a critic network loss function is constructed, which includes an immediate reward term, a target Q-value term adjusted by a discount factor, and a policy entropy term adjusted by a temperature parameter; an actor network loss function is constructed based on the difference between the policy entropy adjusted by the temperature parameter and the target Q-value; a temperature parameter loss function is constructed based on the difference between the policy entropy and the target entropy; the state space input, action space output, immediate reward, and next-time state are stored as transition samples in an experience replay buffer; the critic network loss function, the actor network loss function, and the temperature parameter loss function are optimized based on the transition samples in the experience replay buffer, and the parameters of the deep reinforcement learning network are updated. The parameter adjustment values ​​output from the action space are input into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

2. The method according to claim 1, characterized in that, The steps of constructing parameter nodes from temperature, pressure, time, raw material ratio, and vacuum parameters, extracting the correlation features between these parameter nodes using a three-layer gated graph attention network, and performing temporal modeling of these correlation features through a temporal attention module and residual connections to obtain the parameter coupling state vector include: The sensitivity values ​​of the parameter nodes to oxygen and nitrogen content are calculated. The parameter nodes are weighted based on the sensitivity values. The weighted parameter nodes are then input into a multi-head graph attention network. The attention coefficients of the multi-head graph attention network are adjusted using the sensitivity values. The correlation features between the parameter nodes are then extracted. A subgraph attention network is constructed for the temperature, pressure, and vacuum parameters. In the message passing mechanism of the subgraph attention network, constraints on the inverse relationship between vacuum and pressure, the influence of temperature on vacuum, and the influence of pressure on oxygen and nitrogen solubility are introduced to extract the coupling features between the temperature, pressure, and vacuum parameters. The associated features are divided into melting stage nodes, deoxidation stage nodes, and denitrification stage nodes, and temporal feature extractors are constructed for each stage. The temporal feature extractor includes a temporal attention module and a residual connection module. The temporal attention module captures the temporal dependencies of parameters, and the residual connection module retains the original feature information to extract the temporal features of parameters at different stages. The associated features, the coupling features, and the parameter temporal features are fused to obtain the parameter coupling state vector.

3. The method according to claim 1, characterized in that, The steps of using a long short-term memory network to predict actual oxygen and nitrogen content, and then weighting and fusing the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content include: The parameter-coupled state vector is decomposed using short-term, medium-term, and long-term time windows to obtain corresponding multi-scale state vectors. The multi-scale state vectors are then input into the corresponding long short-term memory networks to obtain the hidden layer states at different time scales. Inter-layer gating units are constructed, and feature interaction is performed on the hidden layer states at adjacent time scales through the inter-layer gating units to obtain multi-scale feature representations. The actual oxygen and nitrogen content is predicted based on the multi-scale feature representations. Calculate the historical forecast mean absolute error, and construct the forecast confidence level based on the historical forecast mean absolute error; calculate the changes in temperature parameter, pressure parameter, and vacuum parameter, and construct the process parameter fluctuation index. The multi-scale feature representation, the prediction confidence, and the process parameter fluctuation index are input into a multilayer perceptron network to obtain theoretical value weights; the theoretical oxygen and nitrogen content and the actual oxygen and nitrogen content are weighted and fused based on the theoretical value weights to obtain the predicted oxygen and nitrogen content. A dynamic correction term is constructed based on the deviation between the measured oxygen and nitrogen content at the previous moment and the predicted oxygen and nitrogen content at the previous moment. The dynamic correction term is then superimposed on the predicted oxygen and nitrogen content to obtain the corrected predicted oxygen and nitrogen content.

4. The method according to claim 1, characterized in that, The steps for constructing an adaptive weight adjustment mechanism based on the reward function include: Based on historical samples, oxygen content error sequences and nitrogen content error sequences are calculated. The mean of oxygen content error and the mean of nitrogen content error are calculated using the exponential weighted moving average method, and the weight coefficients of oxygen content and nitrogen content items in the basic reward items are dynamically adjusted. The penalty term includes a multi-level penalty function, which is divided into a first over-limit interval and a second over-limit interval according to the degree to which the oxygen and nitrogen content exceeds the preset value. A quadratic penalty function is used in the first over-limit interval, and an exponential penalty function is used in the second over-limit interval. The rate of change of process parameters is calculated using a sliding time window. Based on the rate of change, velocity constraints and acceleration constraints are constructed and added to the reward function.

5. The method according to claim 1, characterized in that, The steps of inputting the parameter adjustment amount output from the action space into the process prediction model for simulation verification, using Monte Carlo tree search to plan the adjustment strategy for future time steps, and iteratively training based on the planning results to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints include: Input the parameter adjustment amount output from the action space into the process prediction model to obtain the oxygen and nitrogen content prediction values ​​under multiple simulation conditions; A decision tree is constructed using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node. The decision tree is then expanded, and the node evaluation score is calculated based on the predicted oxygen and nitrogen content. The UCB value is obtained by combining the node visit count. The node with the highest UCB value in the future time step is selected for exploration. A single exploration is completed in four stages: selection, expansion, simulation, and backpropagation. The path with the most visits is extracted from the decision tree as the optimal parameter adjustment sequence. The first step of the optimal parameter adjustment sequence is used as the parameter adjustment scheme at the current moment. The process parameters are updated to obtain the optimal combination of process parameters that satisfies the oxygen and nitrogen content constraints.

6. The method according to claim 5, characterized in that, The steps of constructing a decision tree using the Monte Carlo tree search algorithm, with the current process state as the root node and the parameter adjustment amount as the optional branch node, include: Temperature, pressure, vacuum, time, and predicted oxygen and nitrogen content are used to construct a process state vector. Temperature, pressure, vacuum, and time adjustment values ​​are used to construct a parameter adjustment vector. Decision tree nodes are constructed based on the process state vector and the parameter adjustment vector. An exponential evaluation function is constructed for the predicted oxygen content and the predicted nitrogen content respectively. A first evaluation value is calculated based on the deviation between the predicted oxygen content and the target oxygen content value, and a second evaluation value is calculated based on the deviation between the predicted nitrogen content and the target nitrogen content value. The Euclidean norm of the parameter adjustment vector is calculated to obtain a stability evaluation value. The first evaluation value, the second evaluation value, and the stability evaluation value are weighted to obtain a node evaluation score. The exploration coefficient is calculated based on the number of search rounds, and the exploration coefficient decreases as the number of search rounds increases; the node evaluation score is used as the node value item, and the exploration coefficient is combined with the logarithm of the number of node visits to construct a confidence upper limit calculation formula, and the node UCB value is calculated based on the confidence upper limit calculation formula; During the tree expansion phase, the node with the largest UCB value is selected for expansion, and child nodes are generated based on the physical constraints of parameter adjustment. During the tree search simulation phase, a combined evaluation function of immediate reward and future potential is constructed, and leaf nodes are quickly evaluated based on the combined evaluation function. During the backpropagation phase, the value and number of visits of the visited nodes are updated based on the results of the quick evaluation.

7. A metallurgical process optimization system for low-oxygen, low-nitrogen aluminum-vanadium alloys, used to implement the method described in any one of claims 1-6, characterized in that, include: The first unit is used to construct parameter nodes from temperature parameters, pressure parameters, time parameters, raw material ratio parameters, and vacuum degree parameters. A three-layer gated graph attention network is used to extract the correlation features between the parameter nodes. The correlation features are then modeled temporally through a temporal attention module and residual connections to obtain a parameter coupling state vector. The second unit is used to construct a process prediction model based on the parameter coupling state vector. The process prediction model uses the thermodynamic balance equation to calculate the theoretical oxygen and nitrogen content, uses a long short-time memory network to predict the actual oxygen and nitrogen content, and weights and fuses the theoretical oxygen and nitrogen content with the actual oxygen and nitrogen content to obtain the predicted oxygen and nitrogen content value. The third unit is used to construct a deep reinforcement learning network. The parameter coupled state vector and the predicted oxygen and nitrogen content are used as input to the state space, and the temperature parameter adjustment, pressure parameter adjustment, vacuum degree parameter adjustment, and time parameter adjustment are used as output to the action space. The Soft Actor-Critic algorithm structure is adopted, and the oxygen and nitrogen content is set to be less than the preset value in the reward function. The fourth unit is used to input the parameter adjustment amount output from the action space into the process prediction model for simulation verification. Monte Carlo tree search is used to plan the adjustment strategy for future time steps. Based on the planning results, iterative training is performed to obtain the optimal combination of process parameters that meets the oxygen and nitrogen content constraints.

Citation Information

Patent Citations

  • New material production process parameter optimization method and system based on artificial intelligence

    CN119397914A