Converter steelmaking end point forecasting method and device and storage medium

By decoupling high-dimensional process features through multi-scale importance modeling and graph attention mechanism, and combining multi-granularity Transformer encoder to capture cross-furnace time-series features, the problem of high-dimensional variable coupling and cross-furnace evolution information forgetting in the prediction of the endpoint state of converter steelmaking is solved, and the accurate prediction of the endpoint state is achieved.

CN120930101AActive Publication Date: 2025-11-11TIANJIN DEV ZONE JINGNUOHANHAI DATA TECH CO LTD +1

Patent Information

Application Number
CN202511461487.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing methods for predicting the final state of converter steelmaking are ineffective in addressing issues such as high-dimensional variable coupling, heterogeneous structures depending on different prediction targets, and the forgetting of cross-furnace evolution information.

Method used

Multi-scale importance modeling and graph attention mechanism are used to decouple high-dimensional process features. Multi-granularity Transformer encoder is combined to capture cross-furnace time-series features. The endpoint state is accurately predicted through self-attention mechanism and regression prediction network.

Benefits of technology

It achieves dynamic decoupling of high-dimensional variables and in-depth mining of cross-furnace evolution information, improves the ability to predict the endpoint temperature and multi-element content in a coordinated and accurate manner, and solves the problems of insufficient model stability and generalization ability in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930101A_ABST
    Figure CN120930101A_ABST
Patent Text Reader

Abstract

The invention discloses a converter steelmaking endpoint forecasting method and device and a storage medium, and the method comprises the steps: carrying out the multi-scale importance modeling and graph attention mechanism modeling of the multi-dimensional process characteristics of the current heat, and generating the importance weight and the structure perception weight of each dimension of process characteristics; fusing the importance weight and the structure sensing weight, and multiplying the fused weight with the corresponding dimension process characteristics to generate a characteristic regulation and control vector of the current heat; performing multi-scale convolution operation on the feature regulation and control vector, extracting local mode features under different receptive fields, and outputting a spatial feature vector after the local mode features are enhanced by a self-attention mechanism; a plurality of groups of Transformer encoders are adopted to respectively process the furnace feature sequences with different historical window lengths, and cross-furnace time sequence feature vectors are extracted; and carrying out dimension alignment on the spatial feature vector and the time sequence feature vector, carrying out dimension-by-dimension weighted fusion through a learnable fusion weight, and outputting a steelmaking end point state prediction value of the current heat through a regression prediction network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus and storage medium for predicting the end point of converter steelmaking. Background Technology

[0002] Converter steelmaking is a typical integrated smelting and refining steelmaking process, such as... Figure 1 As shown, the process can generally be divided into two main stages: First, in the blowing stage, scrap steel and molten iron are added to the converter. Oxygen blowing triggers a vigorous oxidation reaction to remove impurities and raise the temperature. Simultaneously, auxiliary materials are added to absorb some heat, forming a slag system conducive to dephosphorization and desulfurization reactions. As the reaction progresses, the steelmaking process gradually approaches the blowing endpoint. At this point, the temperature of the molten steel must meet the blowing requirements to ensure smooth transitions to subsequent operations. Second, in the alloying and tapping stage, alloys are added based on the endpoint state to ensure the steel composition meets standard requirements. In this process, the endpoint state is both a direct result of the upstream oxygen blowing and auxiliary material control and a key basis for formulating the downstream alloy addition strategy, forming a pivotal link throughout the entire process. If the endpoint state deviates from expectations, it will not only lead to waste of upstream raw materials but may also cause drastic adjustments to the alloy addition amount, resulting in a chain reaction of problems such as composition fluctuations and increased costs. Therefore, constructing a high-precision endpoint state prediction model not only helps to anticipate the smelting trend and improve endpoint accuracy but is also a fundamental prerequisite and key bridge for achieving synergistic optimization of the "oxygen blowing—auxiliary materials—alloy" three-stage process.

[0003] From a process control perspective, the core task of predicting the endpoint state in converter steelmaking is to accurately forecast the temperature and carbon concentration of molten steel, while also considering the composition levels of key elements such as manganese, phosphorus, and sulfur. Temperature and carbon concentration together determine the heat balance and reactivity of the molten steel, which are core indicators for ensuring endpoint control accuracy and smelting stability. Phosphorus and sulfur are typical harmful impurities that need to be effectively removed before the endpoint to reduce the burden on subsequent refining and improve the quality of the finished product. Manganese, as a typical alloying element, has its endpoint content affecting alloy feeding schemes and costs.

[0004] To achieve these goals, numerous studies have focused on constructing endpoint prediction methods. For temperature and carbon concentration prediction, early methods often employed neural network structures for modeling. For example, some researchers proposed a BP neural network framework combining an improved particle swarm optimization algorithm to enhance the fitting ability to complex variable coupling relationships; others further introduced topology optimization mechanisms to improve the depth and generalization ability of feature representation. With the development of ensemble learning techniques, some researchers used feature correlation analysis to screen key variables and constructed prediction models based on extreme gradient boosting algorithms, improving interpretability while maintaining nonlinear expressive power. For endpoint prediction problems involving elements such as manganese, phosphorus, and sulfur, various methods have also been proposed. For instance, some researchers used particle swarm optimization to improve the training efficiency and generalization ability of BP networks; others used the firefly algorithm and extreme learning machines to achieve feature sensitivity modeling; still others constructed fusion structures by integrating random forests, adaptive reinforcement algorithms, and stacked ensembles, effectively improving the stability and accuracy of phosphorus content prediction. In summary, despite continuous progress in algorithm selection and feature processing in existing research, current endpoint prediction methods still face two key bottlenecks: First, the variables in converter steelmaking are highly coupled, and different objectives have significantly different dependencies on variables. Most existing methods adopt a static approach, assigning fixed weights to variables, making it difficult to dynamically model the importance distribution of features, resulting in difficulty in suppressing redundant interference and insufficient feature representation quality. Second, there are temporal correlations between furnace runs, such as hot residuals and the continuation of process rhythm, but most models still treat furnace runs as independent samples, modeling only based on the current furnace run, failing to effectively utilize historical evolution information, thus limiting the stability and generalization ability of the models.

[0005] To address the technical problems mentioned above, traditional converter steelmaking endpoint state prediction methods are unable to effectively handle high-dimensional variable coupling relationships, structural heterogeneity depending on different prediction targets, and the forgetting of cross-furnace evolution information. Currently, no effective solution has been proposed. Summary of the Invention

[0006] The embodiments of this disclosure provide a method, apparatus and storage medium for predicting the endpoint of converter steelmaking, so as to at least solve the technical problems existing in the prior art where traditional converter steelmaking endpoint state prediction methods are difficult to effectively deal with high-dimensional variable coupling relationships, heterogeneous structural dependence of different prediction targets, and the forgetting of cross-furnace evolution information.

[0007] According to one aspect of the present disclosure, a method for predicting the endpoint of converter steelmaking is provided, comprising: performing multi-scale importance modeling on the multi-dimensional process features of the current heat, generating importance weights for each process feature; modeling the structural coupling relationship between the multi-dimensional process features through a graph attention mechanism, generating structure-aware weights for each process feature; fusing the importance weights and structure-aware weights to obtain feature weights for each process feature; multiplying each process feature by its corresponding feature weight to generate a feature control vector for the current heat; performing multi-scale convolution operations on the feature control vector of the current heat to extract local pattern features under different receptive fields, and outputting the spatial feature vector of the current heat after enhancement by a self-attention mechanism; using multiple Transformer encoders to process heat feature sequences with different historical window lengths respectively, and extracting temporal feature vectors across heats; wherein the heat feature sequence consists of the feature control vectors corresponding to the current heat and its k preceding heats; and dimensionally aligning the spatial feature vector and the temporal feature vector, performing dimensional weighted fusion through learnable fusion weights, and outputting the predicted value of the steelmaking endpoint state of the current heat through a regression prediction network.

[0008] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0009] According to another aspect of the embodiments of this disclosure, a converter steelmaking endpoint prediction device is also provided, comprising: a feature soft selection module oriented to target differences, used to perform multi-scale importance modeling on the multi-dimensional process features of the current heat, generating importance weights for each process feature; modeling the structural coupling relationship between the multi-dimensional process features through a graph attention mechanism, generating structure-aware weights for each process feature; fusing the importance weights and structure-aware weights to obtain the feature weights for each process feature; multiplying each process feature by its corresponding feature weight to generate a feature control vector for the current heat; and a multi-scale spatial structure perception module, used to perform multi-scale convolution on the feature control vector of the current heat. The system employs a product operation to extract local pattern features under different receptive fields, which are then enhanced by a self-attention mechanism to output the spatial feature vector of the current furnace. A multi-granularity temporal dependency modeling module is used to process furnace feature sequences with different historical window lengths using multiple Transformer encoders, extracting temporal feature vectors across furnaces. The furnace feature sequence consists of feature adjustment vectors corresponding to the current furnace and its k preceding furnaces. A spatiotemporal feature adaptive fusion prediction module is used to align the dimensions of the spatial and temporal feature vectors, perform dimension-wise weighted fusion using learnable fusion weights, and output the predicted steelmaking endpoint state value of the current furnace via a regression prediction network.

[0010] According to another aspect of the present disclosure, a converter steelmaking endpoint prediction device is also provided, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions to perform the following processing steps: performing multi-scale importance modeling on the multi-dimensional process features of the current heat, generating importance weights for each process feature; modeling the structural coupling relationship between the multi-dimensional process features through a graph attention mechanism, generating structure-aware weights for each process feature; fusing the importance weights and structure-aware weights to obtain feature weights for each process feature; and multiplying each process feature by its corresponding feature weight to generate a feature control vector for the current heat. Multi-scale convolution is performed on the feature control vector of the current furnace to extract local pattern features under different receptive fields. After enhancement by a self-attention mechanism, the spatial feature vector of the current furnace is output. Multiple Transformer encoders are used to process the feature sequences of furnaces with different historical window lengths to extract the temporal feature vectors across furnaces. The furnace feature sequence consists of the feature control vectors corresponding to the current furnace and its k preceding furnaces. The spatial feature vector and the temporal feature vector are dimensionally aligned and fused dimensionally by learnable fusion weights. The predicted value of the steelmaking endpoint state of the current furnace is output by a regression prediction network.

[0011] This application first performs multi-scale gating weight generation and graph structure attention coupling analysis on the multi-dimensional steelmaking process characteristics of the current furnace, generating feature regulation vectors to achieve high-dimensional variable decoupling and target difference perception. Then, it extracts local process combination patterns through multi-core parallel convolution, and constructs spatial feature vectors after channel-independent self-attention enhancement. Simultaneously, it uses a multi-granularity Transformer encoder to process the historical furnace sequence composed of feature regulation vectors from historical furnaces, extracting temporal feature vectors across furnaces to characterize process evolution. Finally, it fuses spatiotemporal features through dynamic gating weighting, and outputs the predicted values ​​of endpoint temperature and multi-element content synchronously through a regression network. This paper proposes several key technologies for steelmaking. First, a multi-scale gating network is used to distinguish the dependencies of key variables for different prediction objectives. A graph attention mechanism is then used to explicitly model the physical relationships between process variables, enabling the discovery of key feature response structures for different output tasks. This dynamically adjusts the importance distribution of each variable, improving the ability to identify key variables under multiple objectives and solving the problems of high-dimensional variable coupling and heterogeneous target dependencies. Second, a convolutional neural network is used to extract the structural patterns of process variables in the current furnace, representing the static dependencies of the "spatial dimension." Simultaneously, a Transformer is introduced to model the historical evolution process across furnaces, capturing the dynamic relationships of the "temporal dimension," thus solving the problem of information forgetting across furnaces. Third, an adaptive fusion mechanism adjusts the contribution weights of spatiotemporal features according to dynamic operating conditions, achieving deep fusion of multi-source information and unified modeling of the endpoint state, supporting multi-objective collaborative optimization. This enables the dynamic decoupling of strongly coupled process variables, in-depth mining of cross-furnace evolution information, and accurate collaborative prediction of endpoint temperature and multi-element content. Furthermore, this solves the technical problems of traditional converter steelmaking endpoint state prediction methods, which struggle to effectively address high-dimensional variable coupling relationships, heterogeneous dependencies between different prediction objectives, and the forgetting of cross-furnace evolution information. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 It is a complete flow chart of the existing converter steelmaking process; Figure 2 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 3 This is a flowchart of the converter steelmaking endpoint prediction method according to Embodiment 1 of this application; Figure 4 This is a schematic diagram of the framework of the converter steelmaking endpoint prediction method according to Embodiment 1 of this application; Figure 5 This is a schematic diagram of the framework of the feature soft selection module for target differences as described in Embodiment 1 of this application; Figure 6 This is a schematic diagram of the framework of the multi-scale spatial structure sensing module according to Embodiment 1 of this application; Figure 7 This is a schematic diagram of the framework of the multi-granularity temporal dependency modeling module according to Embodiment 1 of this application; Figure 8 This is an error distribution curve of each method for the endpoint temperature prediction task according to Embodiment 1 of this application; Figure 9 This is an error distribution curve of each method for the endpoint carbon element content prediction task according to Embodiment 1 of this application; Figure 10 This is a comparison chart of the hit type (carbon temperature) percentages of different methods according to Embodiment 1 of this application; Figure 11 This is an error distribution diagram of each method for the endpoint phosphorus content prediction task according to Embodiment 1 of this application; Figure 12 This is an error distribution diagram of each method for the endpoint sulfur content prediction task according to Embodiment 1 of this application; Figure 13 This is a comparison chart of the hit types (phosphorus and sulfur) percentages of different methods according to Embodiment 1 of this application; Figure 14 This is an error distribution curve of each method for the endpoint manganese content prediction task according to Embodiment 1 of this application; Figure 15 This is a schematic diagram of the converter steelmaking endpoint prediction device according to Embodiment 2 of this application; Figure 16 This is a schematic diagram of the converter steelmaking endpoint prediction device according to Embodiment 3 of this application. Detailed Implementation

[0013] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0015] Example 1 According to this embodiment, a method embodiment for predicting the endpoint of converter steelmaking is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0016] The method embodiments provided in this example can be executed on a server or similar computing device. Figure 2 A hardware block diagram of a computational device for implementing a converter steelmaking endpoint prediction method is shown. Figure 2 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 2 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 2 The more or fewer components shown, or having the same Figure 2 The different configurations shown.

[0017] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this application. The data processing circuit can be embodied, in whole or in part, as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuit can be a single, independent processing module, or it can be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0018] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the converter steelmaking endpoint prediction method in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the converter steelmaking endpoint prediction method of the above-mentioned application. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0019] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0020] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0021] It should be noted here that, in some optional embodiments, the above... Figure 2 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 2 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0022] Under the above operating environment, according to the first aspect of this embodiment, a method for predicting the endpoint of converter steelmaking is provided. Figure 3 A flowchart illustrating the method is shown. Figure 4 The architecture diagram of this method is shown, for reference. Figure 3 and Figure 4 As shown, the method includes: S302: Perform multi-scale importance modeling on the multi-dimensional process features of the current furnace batch to generate importance weights for each process feature; model the structural coupling relationship between multi-dimensional process features through graph attention mechanism to generate structure-aware weights for each process feature; fuse importance weights and structure-aware weights to obtain feature weights for each process feature; multiply each process feature by its corresponding feature weight to generate the feature control vector for the current furnace batch. Optionally, the operation of performing multi-scale importance modeling on the multi-dimensional process features of the current furnace and generating importance weights for each process feature includes: dividing the multi-dimensional process features of the current furnace into local scale features, global scale features, and statistical scale features; wherein, local scale features are variables dynamically controlled during the blowing process, global scale features are static variables reflecting the furnace charge structure, and statistical scale features are statistical variables characterizing the process evolution trend; and generating importance weights for local scale features through a first type of gated subnetwork; generating importance weights for global scale features through a second type of gated subnetwork; and generating importance weights for statistical scale features through a third type of gated subnetwork; wherein, the first type of gated subnetwork adopts a single-layer nonlinear transformation structure, the second type of gated subnetwork adopts a multi-layer perceptron structure, and the third type of gated subnetwork adopts the Softplus activation function.

[0023] Optionally, the operation of modeling the structural coupling relationship between multi-dimensional process features through a graph attention mechanism and generating the structure-aware weight of each process feature includes: taking each dimension feature in the multi-dimensional process features of the current furnace as a graph node, establishing a connection edge between any two nodes to form a fully connected graph structure; and inputting the fully connected graph structure into the graph attention network to generate the structure-aware weight of each process feature.

[0024] Optionally, the operation of fusing importance weights and structure-aware weights to obtain feature weights for each process feature includes: for each process feature, dynamically fusing importance weights and structure-aware weights using dedicated learnable fusion weights to obtain the corresponding feature weights.

[0025] Specifically, this application, through research and analysis, found that the data characteristics in the industrial steelmaking process have multiple scales of expression, including local process quantities that vary with each heat, global variables reflecting the process settings, and statistical features extracted from historical processes. Directly applying isomorphic processing to all variables easily overlooks the differences in their granularity of influence and modeling requirements, leading to unstable modeling results and poor interpretability. Based on this, this application divides the multidimensional process characteristics of the current heat into three subspaces: local scale features, global scale features, and statistical scale features.

[0026] Local-scale features include process variables such as oxygen consumption and charge weight for the current heat, which mainly originate from dynamic control behaviors during steelmaking. These variables are typically adjusted frequently during process execution, reflecting short-term operational response characteristics. Such variables place higher demands on the model's local perception and dynamic adaptability.

[0027] Global scale features include molten iron weight, scrap steel ratio, and initial element content in molten iron. These features remain stable over a long period of time, reflecting background information such as furnace charge structure and original operating conditions, and constitute a static characterization of the global state.

[0028] Statistical scaling characteristics include statistical variables such as the mean temperature and elemental content (carbon, manganese, phosphorus, etc.) in the current furnace batch. These are typically calculated from data collected over multiple time periods within the furnace batch, reflecting the potential trends and control rhythms during process evolution. Compared to variables directly involved in operational control, these characteristics are more moderate and global.

[0029] Then, as Figure 5 As shown, an independent gating subnetwork is designed for each type of subspace. Specifically, to adapt to the semantic differences of the three types of feature subspaces, this application designs gating subnetworks with different structures to capture fast response, global stability influence, and temporal trend features respectively. The calculation method for generating the three gating weights is shown in formulas (1), (2), and (3).

[0030] (1) (2) (3) in, , , These represent local scale, global scale, and statistical scale, respectively. , , Let represent the gating weights of the input features in the local, global, and statistical subspaces, respectively. , , Let represent the input features of the local, global, and statistical subspaces, respectively. This represents the Sigmoid function, with an output value range of (0,1). It is used to generate the final gating weights, and each sub-network structure is controlled by independent parameters. , These represent the weights of two linear transformations in a local-scale network. , These represent the bias terms of the two linear transformations in the local-scale network. , , These represent the weights of the three linear transformation layers in a global-scale network. , , These represent the bias terms of the three linear transformations in a global-scale network. , These represent the weights of two linear transformations in the statistical scaling network. , These represent the bias terms and weight matrices of two linear transformations in a statistical scaling network, respectively. With bias term Not shared, and This represents the activation function, used to introduce nonlinearity.

[0031] The local-scale gating network advances dynamically changing variables. This sub-network adopts a shallow structure design, containing only one layer of nonlinear transformation and activation function, avoiding short-term information loss caused by overfitting or smoothing delays. This structure enhances the model's sensitivity to capturing "instantaneous process signals," enabling it to rapidly and accurately weight features of adjustments made during the steelmaking process. The global-scale channel focuses on pre-process conditions. These variables have a long-term, stable, fundamental impact on the final state. This sub-network adopts a multi-layered sensing structure and introduces multi-level nonlinear activation, which helps the model establish stable global expressive capabilities and provides contextual background characterization for different heats. The statistical-scale channel processes variables obtained from statistical data of the current heat, reflecting the potential trends and rhythms throughout the process. To match its mild characteristics of "non-abrupt, gradual change," this sub-network adopts... Activation function replacement In conjunction with parameter initialization suppression strategies, the gated output is guided to be more continuous and smooth, avoiding excessive response to process noise and enhancing its steady-state perception capability for slowly changing signals.

[0032] Furthermore, to avoid training bias caused by uneven gating strength, independent normalization and channel-level amplitude constraint mechanisms are introduced for the three types of sub-networks to ensure that they have relatively balanced and stable control capabilities in the overall feature space. During training, the weights of all gating sub-networks are automatically optimized by backpropagation of the loss function without the need for external labels or prior rules, exhibiting strong adaptability and end-to-end training characteristics.

[0033] Furthermore, this application also finds that structural coupling relationships often exist between various process variables during steelmaking. For example, the relationship between molten iron temperature and carbon content may exhibit a non-linear linkage, and minor changes in the raw material structure may lead to periodic fluctuations in the concentrations of elements such as phosphorus and sulfur. These relationships between variables are not always explicitly encoded through position or order, but are implicit in the industrial control logic and physical mechanisms. To model the implicit coupling relationships between these features, this application introduces a graph attention mechanism, treating input variables as nodes in a graph, and explicitly constructing information pathways between variables through a learnable attention mechanism.

[0034] Specifically, construct a containing A fully connected graph with nodes, where the current input vector (i.e., multidimensional process features) is... Each dimension of process features regarded as a diagram One of the nodes , Since the correlation between variables is often not fully understood before modeling, this application constructs a complete graph (fully connected graph), meaning that there is a connecting edge between any two variable nodes. This graph construction method ensures maximum coverage of structural modeling and provides an expressive basis for automatically learning the importance relationships between variables through an attention mechanism. To evaluate the strength of semantic dependencies between variables, this application employs a learnable attention mechanism. For any two nodes in the graph and The weights of the edges are assigned as shown in formulas (4) and (5).

[0035] (4) (5) in, and Representative node and The original characteristics, Given a linear transformation matrix as input, For attention weight vectors, This demonstrates the splicing operation. This mechanism essentially calculates the node... Its adjacent nodes The correlation, It is a non-linear activation function. This refers to the edge weight (attention score) between the two. Represents an exponential function. The number of nodes is represented by a softmax operation to normalize the weights of adjacent edges of each node to ensure a sum of 1, facilitating interpretable weighted aggregation. This normalization is applied to the edge weight coefficients. After the calculation is completed, the node The structural awareness weights can be obtained by weighted aggregation of its neighbor node information, as shown in formula (6).

[0036] (6) in, It is the structure-aware weight after incorporating structural information. The Sigmoid activation function implicitly encodes the variable. The importance and dependencies within the global structure are highlighted. Compared to methods that generate gating weights based on univariate response values, this approach allows the model to automatically focus on other variables with high semantic coupling to the current variable, generating gating vectors with greater global awareness. It also allows the model to learn higher-order interaction patterns between variables, overcoming the limitations of local awareness and linear assumptions, thus enabling the variable selection process to possess structural cognitive capabilities.

[0037] Furthermore, to further improve the expressive efficiency and robustness of the multi-channel gating mechanism, this application introduces a dynamic weighted fusion strategy. This strategy adaptively integrates one of the three gating methods—local scale gating, global scale gating, and statistical scale gating—with the graph structure gating weight, thereby forming a final gating result that is more discriminative and context-aware. Specifically, for any feature dimension... Its feature weights The importance weights come from two sources: the output of the corresponding scale channel. Structure-aware weights output by the graph structure channel Introducing a learnable fusion weight vector specific to each feature. After Softmax normalization, the weighted contribution for controlling the two gating channels is obtained, as shown in formula (7).

[0038] (7) in, This represents the feature weight obtained after fusing importance weight and structure-aware weight, i.e., the feature weight of the i-th dimension of process feature; Softmax regularization ensures the stability and controllability of the fusion process. This dynamic fusion mechanism not only preserves the perceptual advantages of each gating network for specific semantic features, but also automatically allocates attention focus according to the feature dependency structure of different working conditions and target tasks, achieving deep integration of four types of semantics: global, local, trend + structure. (Feature weights after fusion) This will be compared with the corresponding dimension of process characteristics. Multiplying these together yields the final eigenvector after adjustment, which is the eigenvector of the current furnace cycle.

[0039] Meanwhile, in the high-dimensional input space, although some features are retained in the model, their actual contribution is extremely low, and they may even be redundant or have spurious correlations. If left uncontrolled, this can easily lead to model overfitting, training non-convergence, or poor interpretability of the results. Therefore, this application introduces an L1 sparse regularization term into the loss function to gate the final fused weights. Apply constraints, as in formula (8).

[0040] (8) in, For the final total loss, The regression loss for the main task, This is a hyperparameter that controls the proportion of the influence of the regularization term on the total loss. L1 regularization, through which the model is subjected to "sparse pressure" during training, tends to compress the gating weights of some invalid variables to near zero, thereby achieving "soft selection" of features. Unlike hard selection, this mechanism allows weak features to be retained in the early stages and gradually suppressed in the later stages of training, which helps with stable convergence and information exploration. In addition, this mechanism can improve the interpretability of features and the engineering readability of variable selection, providing important decision-making basis and variable reduction support for subsequent multi-objective optimization.

[0041] S304: Perform multi-scale convolution operation on the feature control vector of the current furnace batch, extract local pattern features under different receptive fields, and output the spatial feature vector of the current furnace batch after being enhanced by a self-attention mechanism. Optionally, the operation of performing multi-scale convolution on the feature control vector of the current furnace to extract local pattern features under different receptive fields, and outputting the spatial feature vector of the current furnace after enhancement by a self-attention mechanism includes: expanding the feature control vector of the current furnace into a one-dimensional feature map; performing n sets of convolution operations with different receptive fields in parallel on the one-dimensional feature map to obtain n sets of feature maps; concatenating the n sets of feature maps along the channel dimension to form a fused feature map; and inputting the fused feature map into a channel-independent self-attention network, flattening the attention-enhanced fused feature map into a one-dimensional vector as the spatial feature vector of the current furnace.

[0042] In this embodiment of the invention, research has revealed that, although the static process variables of each heat in the converter steelmaking process are represented in one-dimensional form, they often contain local combination relationships and characteristic substructures, such as the ratio of total charge to oxygen supply, and the coupling effect of scrap steel content and initial temperature. These local patterns have multi-scale and multi-receptive field feature representation requirements. However, traditional single-layer convolutional neural networks, due to their fixed receptive fields, often struggle to capture key patterns at different scales simultaneously, limiting the model's ability to fully express process information. To improve the spatial perception capability of the input features of the current heat, this application designs a multi-scale convolutional structure, as shown in the following figure. Figure 6 As shown, by introducing parallel convolutional paths with different kernel widths, a feature pyramid is formed, and finally a feature map is obtained, which enhances the model's ability to capture the spatial relationships between features.

[0043] Specifically, let the feature control vector of the current furnace be... First, it is expanded into a one-dimensional feature map, and then n sets of convolution operations with different receptive fields are performed in parallel, as shown in Equation (9).

[0044] (9) in, Representing the Feature maps obtained by path convolution Representing the The convolution operation is performed on each path, followed by m layers of convolutions with the same kernel width for depth feature extraction, using the same number of output channels, ultimately resulting in n sets of feature maps. , respectively representing the local pattern expression under different receptive fields. Then, the n-way feature maps are spliced ​​in the channel dimension. In order to further enhance the global perception ability of the spliced ​​features at different positions, a lightweight self-attention mechanism is introduced to perform global modeling of the spliced ​​feature maps. This module takes the furnace feature map after channel fusion as input, constructs the correlation matrix between positions, generates dynamic attention weights through Query-Key dot product, and then performs weighted aggregation on the Value feature map to realize information interaction and fusion between different positions. This attention module maintains the independence between channels and does not introduce additional channel coupling. Finally, the attention-enhanced multi-scale feature map is flattened in the channel and position dimensions to construct a fixed-length furnace secondary feature vector, which serves as the final representation output of the spatial main branch, as shown in formula (10).

[0045] (10) in, The concatenated feature maps are the result of convolution. For flattening operation, Attention calculations ultimately yield multi-scale spatial feature vectors. This vector serves as the spatial representation of the current furnace's features and will be used for subsequent feature fusion and prediction tasks. This branch focuses on extracting key static structures from the current furnace's process information and is a crucial channel for spatial perception in the entire endpoint prediction model.

[0046] S306: Multiple Transformer encoders are used to process the furnace feature sequences with different historical window lengths to extract cross-furnace time-series feature vectors; wherein, the furnace feature sequence is composed of the feature control vectors corresponding to the current furnace and its k preceding furnaces; Optionally, the operation of using multiple Transformer encoders to process furnace feature sequences with different historical window lengths and extracting cross-furnace temporal feature vectors includes: setting multiple different window lengths and constructing a corresponding furnace feature sequence for each window length; wherein, the furnace feature sequence includes the feature control vectors corresponding to the current furnace and the k furnaces before it, and k is different in different furnace feature sequences; inputting each furnace feature sequence into an independent Transformer encoder for processing and outputting the feature representation corresponding to the window length; and concatenating the feature representations corresponding to multiple different window lengths along the feature dimension to obtain the cross-furnace temporal feature vector.

[0047] In this embodiment of the invention, research has revealed significant temporal correlations between different heats in the converter steelmaking process. For example, factors such as the residual heat state of the furnace lining, the slag residue from the previous heat, and the equipment heat load often affect the oxygen utilization efficiency, smelting reaction rate, and temperature loss of the current heat. If heats are treated as independent samples, it is difficult to capture this cross-heater information implicit in the time evolution process, affecting the model's ability to perceive changes in operating conditions and its predictive stability. In actual production, the temporal impact of the smelting process cannot be completely covered by a fixed window length. Short-term information has a rapid impact on the direct heat balance and reaction rate of the current heat, while medium- and long-term factors (such as equipment wear and tear, system operating cycles, etc.) may gradually accumulate their impact over a longer heat range. Therefore, this application introduces historical sequence lengths of various granularities to model short-term, medium-term, and long-term time series patterns respectively, thereby enhancing the model's sensitivity and expressive ability to multi-scale temporal changes.

[0048] To model multi-granularity historical furnace sequence, a Transformer encoder is introduced into the time-series branches. Leveraging its powerful context-awareness and long-term dependency modeling capabilities, it captures temporal evolution patterns at different scales. For example... Figure 7As shown, it mainly consists of feature embedding, position encoding, multi-head attention mechanism and feedforward network. The embedding layer is used to represent the input features of each furnace in high dimension; position encoding introduces sequence order information, enabling the model to have time-series awareness; the multi-head attention mechanism can model the dependencies between different positions in parallel, effectively improving the model's ability to model time-series features across step lengths.

[0049] Specifically, let the current number be... The characteristic control vector of each furnace is Multiple different window lengths are set, and for each window length, a corresponding furnace feature sequence is constructed. For example, but not limited to, the window lengths are 5, 7, and 9. When the window length is 5, the constructed furnace feature sequence is generated by the first... Feature control vector of each furnace The feature control vectors corresponding to its first 5 furnace cycles as well as Composition, that is Then, each sequence is first embedded into a unified dimensional space through linear mapping, and then fed into multiple independent Transformer encoders for modeling, as shown in Equation (11).

[0050] (11) in, Representing the The first furnace Granularity history sequence input, Representative for Granular Transformer encoder To output the features at the corresponding granularity, multiple output vectors of different granularities are finally concatenated to construct a multi-granularity temporal feature representation, as shown in formula (12).

[0051] (12) in, Represents the temporal feature vectors used in multi-granularity temporal dependency modeling. The multi-granularity design, representing splicing operations, further enhances the model's ability to perceive short-term, medium-term, and long-term furnace evolution trends, avoiding the limitations of fixed windows on modeling capabilities.

[0052] S308: Align the spatial feature vector and temporal feature vector in terms of dimensions, perform dimensional weighted fusion through learnable fusion weights, and output the predicted value of the steelmaking endpoint state of the current furnace through a regression prediction network.

[0053] Optionally, the operation of dimensionally aligning the spatial feature vector and temporal feature vector, performing dimensional weighted fusion through learnable fusion weights, and outputting the predicted steelmaking endpoint state value of the current heat through a regression prediction network includes: compressing or expanding the spatial feature vector and temporal feature vector to the same dimension through linear projection; concatenating the aligned spatial feature vector and temporal feature vector, and then inputting them into a gating network to generate a learnable fusion weight vector; performing dimensional weighted fusion of the spatial feature vector and temporal feature vector according to the generated learnable fusion weight vector to obtain a fused feature vector; and inputting the fused feature vector into a regression prediction network to output the predicted steelmaking endpoint state value of the current heat.

[0054] In this embodiment of the invention, research has revealed that the current furnace process state is not only determined by the structural relationships between its static variables, but also profoundly influenced by the evolution of previous furnaces. Spatial structure and temporal dependence are naturally complementary in prediction tasks. To fully integrate the static process characteristics of the current furnace with the dynamic evolution information in the historical sequence, this application designs a spatiotemporal feature adaptive fusion prediction module based on spatial and temporal branch modeling to achieve deep collaborative modeling of multi-source features. Unifying and integrating the features extracted from the first two modules is a crucial step in achieving the final prediction. Among them, the multi-scale spatial structure perception module focuses on modeling the combination relationships between process variables within the current furnace, possessing strong expressive power and local pattern sensitivity; while the multi-granularity temporal dependence modeling module focuses on the cross-temporal dependence structure between furnaces, possessing good contextual modeling capabilities and historical information memory. The two are significantly complementary in terms of feature sources and modeling perspectives. The former emphasizes local structure perception, while the latter strengthens global temporal modeling. After fusion, it is expected to construct a furnace-level feature representation with greater discriminative and generalizable capabilities.

[0055] Unlike traditional direct splicing methods, this application introduces an adaptive weighted fusion mechanism. After unifying the feature dimensions, it integrates spatial and temporal features dimension-wise using learnable fusion weights, thereby improving the discriminative and adaptive nature of feature fusion. Specifically, let the output spatial feature vector of the multi-scale spatial structure perception module be... The multi-granularity temporal dependency modeling module outputs a temporal feature vector as follows: After the two are mapped to the same dimension through linear projection or compression, the feature vectors are fused. The definition is as follows: (13) in, Represents element-wise multiplication, merging weight vectors This weight controls the contribution ratio of the two types of features in each dimension. It is generated by a set of fusion-gated networks, whose input is a concatenated vector of the two types of features, with the following structure: (14) in, Representative features spliced ​​together and Representing the weight matrix and bias terms, this gating mechanism allows the model to dynamically adjust the fusion strength of spatial and temporal features based on different samples, achieving fine-grained feature control. The fused representation The data is input into the regression prediction module, which further obtains the predicted endpoint state through nonlinear mapping. The regression prediction module consists of n fully connected layers, and its structure is shown below: (15) in, For the final prediction result, Represents weight, This represents the bias term.

[0056] To verify the effectiveness and generalization of the method proposed in this application, the experiment used real production data from a domestic steel plant's converter steelmaking process, with the steel grade being HRB400E. The production data included production variables such as molten iron information, scrap steel information, smelting process information, oxygen supply, and smelting endpoint information, spanning from January 1, 2023 to May 1, 2024. A total of 10,991 valid data points were selected as the original dataset, some of which are shown in Table 1. Twenty variables that are meaningful for the endpoint state were selected as the feature control vector for the current furnace cycle, as shown in Table 2.

[0057] Table 1 Production Variables Table

[0058] Table 2 Input Variables Table

[0059] To evaluate the predictive performance of the model, the mean absolute error was used. Root mean square error Accuracy within the error range As an evaluation criterion, it is defined as follows: (16) (17) (18) in, and These represent the true value and predicted value of the target variable, respectively. This represents the number of samples in the test set and is an indicator function. If the expression within the parentheses is true, the indicator function evaluates to 1; otherwise, it evaluates to 0. It is the absolute value of the given error range. It describes the average absolute difference between the predicted and actual values; It describes the average error between the predicted value and the actual value, and is more sensitive to outliers; A visual indicator that reflects predictive performance; the higher the indicator, the better the predictive performance.

[0060] This experiment first preprocesses the original dataset to construct the required input structure for the model and removes outliers. For multi-granularity time series, four historical window lengths (3 / 5 / 7 / 9 furnace batches) are set to capture short-, medium-, and long-range cross-furnace batch dependencies, respectively. In the multi-scale spatial structure perception module, a parallel four-path convolutional structure is used with kernel sizes of 1 / 3 / 5 / 7, and each path contains four convolutional layers, with the number of channels increasing layer by layer. In the multi-granularity temporal dependency modeling module, an 8-layer encoder layer and 8 multi-head attention heads are used. During training, the optimizer is Adam (with a fixed learning rate of 0.001), the batch size is 16, the maximum training epochs are 200, an early stopping strategy is adopted, and the basic loss function is L1 loss.

[0061] To comprehensively evaluate the performance of the proposed endpoint state prediction model under different prediction tasks, this application divides the modeling task into three sub-task groups, focusing on endpoint prediction problems for temperature and carbon content, phosphorus and sulfur content, and manganese content, respectively. For each group of experiments, representative classical models or mainstream deep learning methods in the field are selected as baselines for comparison. The method in this application is denoted as SGCT-Net, and the mean absolute error is uniformly used. Root mean square error And accuracy within the error range As an evaluation metric, the accuracy calculation sets tolerance ranges based on the process control requirements of each variable: temperature is set to ±15℃, carbon and manganese to ±0.05%, phosphorus to ±0.01%, and sulfur to ±0.005%. Samples whose absolute error between the predicted and actual values ​​does not exceed the threshold within this tolerance range are considered to be accurately predicted.

[0062] To further evaluate the applicability and stability of the model in multivariate collaborative control scenarios, this application introduces a joint hit rate metric in each experimental group: separately calculating the temperature and carbon... Phosphorus and sulfur The proportion of samples where the variable is simultaneously matched reflects the overall performance of the model in local multi-objective consistency control.

[0063] To verify the effectiveness of the proposed model in predicting traditional key control variables—endpoint temperature and carbon content—this application designed a set of comparative experiments. This task represents a key research area in steelmaking, and prediction accuracy directly impacts the hit rate of endpoint control and the stability of smelting quality. The experiment selected three representative comparative models in this field for performance comparison: FOA-GRNN: This method uses the fruit fly optimization algorithm to optimize the kernel width parameter of the generalized regression neural network, so as to improve the model's prediction accuracy of steelmaking endpoint temperature and carbon mass fraction. LWOA-TSVR: This method integrates the Logistic whale optimization algorithm and ε-insensitive dual support vector regression. It optimizes the model hyperparameters through LWOA to improve generalization performance. ε-TSVR has good nonlinear modeling ability and noise resistance, and is suitable for prediction scenarios with fluctuations and anomalies in steelmaking data. JITL: This method improves the model's adaptability to nonlinear and time-varying conditions by dynamically selecting similar data from historical samples based on the input characteristics of the current furnace batch during each prediction and building a personalized regression model based on these local samples.

[0064] The experimental results are shown in Table 3.

[0065] Table 3. Comparison of Predicted Endpoint Temperature and Carbon Content Experimental Results

[0066] As shown in Table 3, the experimental results demonstrate that the proposed SGCT-Net model comprehensively outperforms the comparative models in both the prediction of endpoint temperature and carbon content. Taking endpoint temperature as an example, SGCT-Net... and Compared to the second-best performing FOA-GRNN, SGCT-Net reduced accuracy by 13.3% and 14.3% respectively; within an error tolerance range of ±15℃, its accuracy reached 85.30%, an improvement of 9.58 percentage points over FOA-GRNN. Regarding the endpoint carbon content, SGCT-Net also... and It performed best, reducing accuracy by 13.2% and 7.5% compared to FOA-GRNN, respectively, with an accuracy rate as high as 94.21%. Furthermore, in the joint hit rate metric... In the above tests, SGCT-Net achieved a success rate of 80.85%, significantly higher than FOA-GRNN, LWOA-TSVR, and JITL. This indicates that the model can not only predict key variables separately but also demonstrates superior capabilities in multi-objective consistency control. This performance advantage is mainly attributed to SGCT-Net's use of a multi-scale spatial structure perception module to extract local combination patterns among process variables in the current furnace, and its combination with a multi-granularity temporal dependency modeling module to mine cross-furnace evolution patterns. Finally, through a spatiotemporal feature adaptive fusion module, it achieves deep fusion and differentiated weighted modeling of spatial and temporal information, thereby improving the model's predictive stability and generalization ability under complex and volatile operating conditions.

[0067] Prediction error distribution curve as shown Figure 8 and Figure 9 As shown, the error density distribution of SGCT-Net and three comparative models in the task of predicting the endpoint temperature and endpoint carbon content is illustrated. The horizontal axis represents the prediction error (predicted value minus the true value), and the vertical axis represents the sample density within the error interval. Each curve represents the error distribution characteristics of a model in this task. The black dashed line indicates the position where the error is 0. Ideally, more samples should be densely distributed near this line. The more concentrated the curve, the more accurate the model prediction and the smaller the fluctuation.

[0068] from Figure 8 It can be seen that the prediction error distribution of SGCT-Net is most concentrated near the error of 0, and the curve contracts sharply in the central region. Approximately 90% of the sample errors fall within the range of [–10℃, +10℃], with very few samples far from the center. This indicates that SGCT-Net possesses extremely strong accuracy and error control capabilities in predicting the final temperature. In contrast, FOA-GRNN and LWOA-TSVR have wider error distributions, with significantly higher density at the tails on both sides than SGCT-Net, indicating a greater number of large error samples. JITL, on the other hand, has a lower density in the central region, suggesting greater volatility in its prediction results. Figure 9 The image shows the prediction error distribution for the final carbon element. SGCT-Net also performs best near the error threshold of 0, with the error density rapidly concentrated in the range of [–0.025%, +0.025%], showing virtually no outliers significantly deviating from the center. This distribution pattern indicates that SGCT-Net possesses very strong robustness and stability in modeling low-amplitude, highly sensitive variables like carbon content. In contrast, the distributions of FOA-GRNN and LWOA-TSVR are biased towards the positive error region, exhibiting a systematic tendency to overestimate, while JITL's distribution is relatively dispersed near the center, lacking consistency. The advantage of SGCT-Net mainly stems from its introduction of a multi-scale spatial structure sensing module to accurately capture local combination patterns between variables, effectively reducing the sensitivity to process fluctuations in the prediction.

[0069] from Figure 10 As can be seen, the SGCT-Net proposed in this application significantly outperforms other models in the "double hit" metric, reaching approximately 80%, which is about 10-20 percentage points higher than other models, indicating stronger predictive ability for the two core variables of temperature and carbon. Furthermore, SGCT-Net accounts for the smallest proportion of samples with "only one hit" or "no hits," demonstrating that its predictions are not only accurate but also more consistent. This performance is mainly attributed to the target difference modeling and spatiotemporal feature fusion mechanism introduced by SGCT-Net. By specifically constructing differential feature channels for temperature and carbon, the model achieves more coordinated feature matching and error control across different targets, thereby effectively improving the single-hit rate for multiple targets.

[0070] To further verify the adaptability and robustness of the proposed method in modeling trace impurity elements, this application sets up a second set of comparative experiments, focusing on the joint prediction task of the endpoint phosphorus and sulfur content. Phosphorus and sulfur, as typical harmful impurities, have low and highly volatile contents, influenced by multiple factors such as slag composition and temperature fluctuations, making modeling them difficult and intolerant of prediction errors, thus posing a challenging task in intelligent modeling. Therefore, this experiment selects the following three models as comparative methods: PCA-BP: This method first uses principal component analysis (PCA) to reduce the dimensionality of high-dimensional input variables in the smelting process, extracts the main information components to reduce noise interference, and then uses backpropagation neural network (BP) to construct a nonlinear prediction model. Stacking: This method is based on thermodynamic and kinetic analysis of converter dephosphorization. It selects key process parameters from production data as input features and combines three sub-models, namely random forest, adaptive enhancement algorithm and multiple enhancement algorithm, to construct a Stacking ensemble regression model to predict the final phosphorus content of the converter. MLI-GA-BPNN: This method integrates metallurgical mechanisms and industrial data. First, it uses local linear embedding for feature selection and dimensionality reduction based on the dephosphorization mechanism. Then, it employs a backpropagation (BP) neural network to construct a prediction model and optimizes the model structure and parameters using a genetic algorithm. Mechanistic information is incorporated into the fitness function and loss function, improving the model's prediction accuracy and physical consistency.

[0071] The experimental results are shown in Table 4.

[0072] Table 4. Comparison of Experimental Results for Predicted Phosphorus and Sulfur Content at the Endpoint

[0073] As shown in Table 4, SGCT-Net achieved significant performance advantages in both the endpoint phosphorus and sulfur content prediction tasks. Specifically, in the endpoint phosphorus content prediction, SGCT-Net... and Compared to MLI-GA-BPNN, the accuracy was reduced by approximately 22.2% and 19.7%, respectively; within an error range of ±0.01%, the accuracy reached 89.98%, significantly higher than Stacking and PCA-BP. For the endpoint sulfur content, it improved by more than 4 percentage points compared to the second-best performing Stacking. Notably, the combined hit rate for phosphorus and sulfur was significantly higher. In the above tests, SGCT-Net achieved a score of 81.07%, significantly outperforming MLI-GA-BPNN and Stacking, fully demonstrating the model's stability and accuracy consistency in control scenarios involving trace impurity elements. This advantage mainly stems from SGCT-Net's ability to accurately capture the cross-features of weak signals and strongly coupled structures: its graph structure perception module constructs a fully connected variable graph to explicitly model the implicit coupling relationships between low-concentration elements such as phosphorus and sulfur and upstream and downstream process variables, improving the structural integrity of feature representation; simultaneously, the feature soft selection mechanism combines sparse regularization and multi-scale gating to highlight phosphorus and sulfur-sensitive variables in high-dimensional inputs, avoiding redundant interference.

[0074] Error distribution curve as shown Figure 11 and Figure 12 As shown, the error density distribution of each model in the final task of predicting phosphorus and sulfur content is illustrated.

[0075] from Figure 11 It can be seen that the prediction error of SGCT-Net is most concentrated near 0, with the highest concentration of the curve. The vast majority of sample errors fall within the range of [-0.01%, +0.01%], and the tails decay rapidly on both sides. This indicates that the model has a very strong error control capability in predicting the final phosphorus content. In contrast, the error distributions of PCA-BP and MLI-GA-BPNN are wider, and the density peaks deviate from 0, showing a systematic overestimation trend. Although Stacking has a relatively good concentration, there are still many samples with errors above ±0.01%, indicating that its boundary prediction is unstable. Figure 12Among the models, SGCT-Net exhibits the most compact error distribution in predicting the final sulfur element, with the vast majority of samples falling within the range of [-0.005%, +0.005%], and a significantly higher proportion of samples with errors close to 0 compared to other models. This distribution characteristic indicates that SGCT-Net possesses higher robustness and consistency when dealing with low-concentration, high-sensitivity variables. In contrast, the curves of PCA-BP and MLI-GA-BPNN are shifted to the right, indicating significant overestimation issues; while Stacking performs slightly better in error contraction, its density in the central region is still lower than that of SGCT-Net. SGCT-Net's performance in this task largely relies on its multi-granularity temporal modeling structure, which effectively enhances the model's ability to perceive complex dynamic processes by capturing subtle trend changes in furnace evolution.

[0076] from Figure 13 As can be seen, SGCT-Net significantly outperforms other models in the "double hit" rate for phosphorus and sulfur, exceeding 80%, which is 10-25 percentage points higher than other models, indicating its strongest predictive ability for trace impurity elements. Meanwhile, SGCT-Net has the lowest percentage of "no hits," further validating its stability and generalization ability in complex variable modeling. This performance advantage mainly stems from the synergistic effect of the graph structure attention mechanism and the multi-granularity temporal modeling module in SGCT-Net: the former can identify potential coupling variable combinations between phosphorus and sulfur, while the latter captures the implicit evolutionary trends in the furnace sequence, enabling the model to have stronger linkage perception and error control capabilities when simultaneously predicting two low-concentration, highly sensitive variables.

[0077] This application includes a third set of experiments to evaluate the prediction of the final manganese content. Manganese, as a typical alloying control element, not only affects the subsequent alloy feed amount but also directly impacts steel composition and cost control. Compared to phosphorus and sulfur, manganese content is generally higher but fluctuates more dramatically and is influenced by multiple factors such as temperature, slag composition, and reaction time, making modeling a challenging process. Therefore, this experiment selects the following three models as comparative methods: IWOA-LSSVM: This method utilizes the Improved Whale Optimization (IWOA) algorithm to optimize the LSSVM model parameters, thereby improving the prediction accuracy of the final manganese content. By introducing cosine perturbation and a nonlinear convergence mechanism, the model's fitting and generalization capabilities under nonlinear conditions are enhanced.

[0078] NN-Static: This method is based on a feedforward neural network to build a static prediction model. The input is a variety of process parameters and the output is the final manganese content. It is trained by the BP algorithm to achieve nonlinear modeling and prediction of manganese content.

[0079] SSA-LSTM: This method uses the Spiral Search Algorithm (SSA) to globally optimize the hyperparameters of the LSTM model, thereby improving its fitting ability and prediction performance in converter smelting data.

[0080] The experimental results are shown in Table 5.

[0081] Table 5. Comparison of Predicted Manganese Content at the Endpoint

[0082] As shown in Table 5, SGCT-Net also achieved optimal performance across the board in the endpoint manganese content prediction task. and In terms of two error metrics, SGCT-Net reduced errors by 18.9% and 17.9% respectively compared to SSA-LSTM, and was significantly superior to traditional static neural network models and IWOA-LSSVM based on optimized support vector machines. Within an error range of ±0.05%, SGCT-Net achieved an accuracy of 84.41%, an improvement of 8.46 percentage points compared to SSA-LSTM, fully demonstrating its significant advantage in modeling highly volatile alloy elements. Thanks to the multi-granularity time-series modeling module's modeling of reaction rhythm and residual effects in previous furnace cycles, the model exhibits stronger adaptability and anti-interference capabilities when dealing with alloy variables such as manganese, which are characterized by "large fluctuations, strong coupling, and high sensitivity," providing reliable support for precise control in actual steelmaking processes.

[0083] Error distribution curve as shown Figure 14 As shown, the error density distribution of each model in the final manganese content prediction task is illustrated.

[0084] from Figure 14 It can be seen that the prediction error of SGCT-Net is highly concentrated around 0, mainly distributed in the range of [-0.05%, +0.05%], with the highest density at the center, indicating that the model has stronger error compression and anomaly control capabilities in the final manganese content prediction. In contrast, the error distributions of SSA-LSTM and IWOA-LSSVM are wider, especially SSA-LSTM, which still has a large density at both ends, showing higher sensitivity to complex operating conditions; NN-Static has greater error volatility and its density curve is not smooth enough, indicating weaker prediction stability. The reason why SGCT-Net performs better in modeling manganese, a highly volatile and strongly coupled variable, is mainly due to its multi-granularity time-series modeling module, which can capture residual information between furnace cycles.

[0085] To systematically evaluate the specific contributions of each core module in the proposed model to the endpoint prediction performance, three ablation experiments were designed. Feature selection mechanism, spatial structure modeling capability, and time series modeling capability were systematically eliminated and replaced, respectively. The effectiveness and necessity of each module were verified through changes in model performance. The specific comparison model configurations are as follows: SGCT-Net (Complete Model): The endpoint state prediction model proposed in this application includes core modules: feature soft selection based on target differences, multi-scale spatial structure perception, multi-granularity temporal dependency modeling, and spatiotemporal feature adaptive fusion. The model possesses comprehensive capabilities such as dynamic feature modeling, complex variable structure representation, and furnace evolution trend perception.

[0086] SGCT-Net-noSelector (removal of soft feature selection module): The soft feature selection module is replaced with a static feature selection method based on Pearson correlation coefficient. The input feature with the highest correlation is selected for each prediction target to verify the modeling advantages of the learnable feature selection mechanism in high-dimensional coupled scenarios.

[0087] SGCT-Net-noSpatial (removal of spatial structure perception module): Replaces the multi-scale spatial structure perception branch with a single-scale shallow perceptron, weakening the model's ability to model complex variable combinations within the current furnace, and evaluating the contribution of spatial multi-scale structure extraction to local pattern recognition and prediction performance.

[0088] SGCT-Net-noTemporal (removal of temporal dependency modeling module): The multi-granularity temporal modeling branch is replaced with a normal fully connected network, and the historical furnace sequence is no longer used for context modeling. The impact of cross-furnace dynamic trend perception capability on prediction stability and robustness is tested.

[0089] The results of the ablation experiment are shown in Table 6.

[0090] Table 6 Ablation Experiment Results

[0091] As shown in Table 6, the ablation experiment results demonstrate that the complete SGCT-Net model exhibits optimal performance across all prediction tasks, fully validating the effectiveness and synergy of each core module in improving prediction accuracy and robustness. In contrast, the simplified model, by removing any module, shows significant performance degradation across various metrics, indicating that each sub-module plays an irreplaceable role in the overall model architecture.

[0092] Specifically, the SGCT-Net-noSelector model generally underperformed the full model in multiple prediction tasks, indicating that the feature soft selection module, which focuses on target differences, can effectively distinguish between task-related and irrelevant features in highly coupled input scenarios, improving the expressive power of key variables under different prediction targets. In contrast, the traditional Pearson correlation screening method lacks contextual modeling capabilities and struggles to address nonlinear coupling and feature redundancy issues between variables. The SGCT-Net-noSpatial model showed a significant increase in error for targets sensitive to local process changes, such as endpoint temperature and manganese content, reflecting the importance of the multi-scale spatial structure perception module in extracting complex variable interactions and short-term process disturbance patterns within a furnace, especially in modeling scenarios dominated by local information. The SGCT-Net-noTemporal model exhibited the most significant performance degradation in predicting elements with strong volatility and significant historical dependence, such as phosphorus, sulfur, and manganese, indicating that the multi-granular temporal dependency modeling module plays a crucial role in characterizing cross-furnace evolution trends and gradual changes, effectively improving the model's stability and trend response capabilities.

[0093] In summary, the ablation experiments systematically verified the necessity and synergistic benefits of the aforementioned modules from three dimensions: feature selection, structure awareness, and temporal modeling. These modules complement each other in modeling high-dimensional complex variables, capturing process disturbance features, and identifying temporal evolution patterns, forming the key modeling support for SGCT-Net. This significantly improves the model's robustness and generalization ability under complex operating conditions, providing a solid predictive foundation for subsequent collaborative optimization of the "oxygen blowing—auxiliary materials—alloy" three-stage process.

[0094] Therefore, to address the challenge of traditional endpoint state prediction methods failing to effectively handle high-dimensional variable coupling relationships and inter-furnace evolution trends, this application proposes an endpoint state prediction model, SGCT-Net, which integrates feature selection and spatiotemporal structure modeling. This model consists of four key modules: a feature soft selection module oriented towards target differences enhances the identification capability of key variables under multi-objective conditions; a multi-scale spatial structure perception module extracts complex local variable combination patterns in the current furnace; a multi-granularity temporal dependency modeling module characterizes historical process evolution features; and a spatiotemporal feature adaptive fusion module achieves unified expression and dynamic weighted fusion of cross-dimensional information. Multiple experiments conducted on real production data of typical steel grades demonstrate that SGCT-Net achieves optimal performance in predicting five endpoint variables: temperature, carbon, phosphorus, sulfur, and manganese, with the average error reduced by up to 18.9%, and a significantly improved multi-objective hit rate, exhibiting excellent prediction accuracy, stability, and engineering adaptability. This model lays a solid predictive foundation and data support for subsequent multi-objective collaborative optimization of the "oxygen blowing—auxiliary materials—alloy" three-stage process.

[0095] In addition, refer to Figure 2 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0096] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0098] Example 2 Figure 15 A converter steelmaking endpoint prediction device according to this embodiment is shown, which corresponds to the method described according to Embodiment 1. (Reference) Figure 15As shown, the device includes: a feature soft selection module 1510 for target difference, used to perform multi-scale importance modeling on the multi-dimensional process features of the current furnace, generating importance weights for each process feature; modeling the structural coupling relationship between multi-dimensional process features through a graph attention mechanism, generating structure-aware weights for each process feature; fusing importance weights and structure-aware weights to obtain feature weights for each process feature; multiplying each process feature with its corresponding feature weight to generate the feature control vector for the current furnace; and a multi-scale spatial structure perception module 1520, used to perform multi-scale convolution operations on the feature control vector of the current furnace to extract features under different receptive fields. The system consists of several modules: a local pattern feature, which is enhanced by a self-attention mechanism to output the spatial feature vector of the current furnace; a multi-granularity temporal dependency modeling module 1530, which uses multiple Transformer encoders to process furnace feature sequences with different historical window lengths to extract temporal feature vectors across furnaces; wherein the furnace feature sequence consists of feature adjustment vectors corresponding to the current furnace and its k preceding furnaces; and a spatiotemporal feature adaptive fusion prediction module 1540, which aligns the spatial feature vector and the temporal feature vector in dimensions, performs dimensional weighted fusion through learnable fusion weights, and outputs the predicted value of the steelmaking endpoint state of the current furnace through a regression prediction network.

[0099] Therefore, according to this embodiment, firstly, multi-scale gating weight generation and graph structure attention coupling analysis are performed on the multi-dimensional steelmaking process characteristics of the current furnace to generate feature control vectors, achieving high-dimensional variable decoupling and target difference perception. Then, local process combination patterns are extracted through multi-core parallel convolution, and spatial feature vectors are constructed after channel-independent self-attention enhancement. Simultaneously, a multi-granularity Transformer encoder is used to process the historical furnace sequence composed of feature control vectors from historical furnaces, extracting temporal feature vectors across furnaces to characterize process evolution. Finally, spatiotemporal features are fused through dynamic gating weighting, and the predicted values ​​of endpoint temperature and multi-element content are synchronously output through a regression network. This paper proposes several key technologies for steelmaking. First, a multi-scale gating network is used to distinguish the dependencies of key variables for different prediction objectives. A graph attention mechanism is then used to explicitly model the physical relationships between process variables, enabling the discovery of key feature response structures for different output tasks. This dynamically adjusts the importance distribution of each variable, improving the ability to identify key variables under multiple objectives and solving the problems of high-dimensional variable coupling and heterogeneous target dependencies. Second, a convolutional neural network is used to extract the structural patterns of process variables in the current furnace, representing the static dependencies of the "spatial dimension." Simultaneously, a Transformer is introduced to model the historical evolution process across furnaces, capturing the dynamic relationships of the "temporal dimension," thus solving the problem of information forgetting across furnaces. Third, an adaptive fusion mechanism adjusts the contribution weights of spatiotemporal features according to dynamic operating conditions, achieving deep fusion of multi-source information and unified modeling of the endpoint state, supporting multi-objective collaborative optimization. This enables the dynamic decoupling of strongly coupled process variables, in-depth mining of cross-furnace evolution information, and accurate collaborative prediction of endpoint temperature and multi-element content. Furthermore, this solves the technical problems of traditional converter steelmaking endpoint state prediction methods, which struggle to effectively address high-dimensional variable coupling relationships, heterogeneous dependencies between different prediction objectives, and the forgetting of cross-furnace evolution information.

[0100] Example 3 Figure 16 A converter steelmaking endpoint prediction device according to this embodiment is shown, which corresponds to the method described according to Embodiment 1. (Reference) Figure 16As shown, the device includes: a processor 1610; and a memory 1620, connected to the processor 1610, for providing the processor 1610 with instructions to perform the following processing steps: performing multi-scale importance modeling on the multi-dimensional process features of the current furnace batch, generating importance weights for each process feature; modeling the structural coupling relationship between the multi-dimensional process features through a graph attention mechanism, generating structure-aware weights for each process feature; fusing the importance weights and structure-aware weights to obtain the feature weights for each process feature; multiplying each process feature by its corresponding feature weight to generate the feature control vector for the current furnace batch; and performing multi-scale importance modeling on the multi-dimensional process features of the current furnace batch. The feature control vectors are subjected to multi-scale convolution operations to extract local pattern features under different receptive fields. After being enhanced by a self-attention mechanism, the spatial feature vector of the current furnace is output. Multiple Transformer encoders are used to process the furnace feature sequences with different historical window lengths to extract the temporal feature vectors across furnaces. The furnace feature sequence consists of the feature control vectors corresponding to the current furnace and its k preceding furnaces. The spatial feature vectors and temporal feature vectors are dimensionally aligned and fused one-way by a learnable fusion weight. The predicted value of the steelmaking endpoint state of the current furnace is output by a regression prediction network.

[0101] Therefore, according to this embodiment, firstly, multi-scale gating weight generation and graph structure attention coupling analysis are performed on the multi-dimensional steelmaking process characteristics of the current furnace to generate feature control vectors, achieving high-dimensional variable decoupling and target difference perception. Then, local process combination patterns are extracted through multi-core parallel convolution, and spatial feature vectors are constructed after channel-independent self-attention enhancement. Simultaneously, a multi-granularity Transformer encoder is used to process the historical furnace sequence composed of feature control vectors from historical furnaces, extracting temporal feature vectors across furnaces to characterize process evolution. Finally, spatiotemporal features are fused through dynamic gating weighting, and the predicted values ​​of endpoint temperature and multi-element content are synchronously output through a regression network. This paper proposes several key technologies for steelmaking. First, a multi-scale gating network is used to distinguish the dependencies of key variables for different prediction objectives. A graph attention mechanism is then used to explicitly model the physical relationships between process variables, enabling the discovery of key feature response structures for different output tasks. This dynamically adjusts the importance distribution of each variable, improving the ability to identify key variables under multiple objectives and solving the problems of high-dimensional variable coupling and heterogeneous target dependencies. Second, a convolutional neural network is used to extract the structural patterns of process variables in the current furnace, representing the static dependencies of the "spatial dimension." Simultaneously, a Transformer is introduced to model the historical evolution process across furnaces, capturing the dynamic relationships of the "temporal dimension," thus solving the problem of information forgetting across furnaces. Third, an adaptive fusion mechanism adjusts the contribution weights of spatiotemporal features according to dynamic operating conditions, achieving deep fusion of multi-source information and unified modeling of the endpoint state, supporting multi-objective collaborative optimization. This enables the dynamic decoupling of strongly coupled process variables, in-depth mining of cross-furnace evolution information, and accurate collaborative prediction of endpoint temperature and multi-element content. Furthermore, this solves the technical problems of traditional converter steelmaking endpoint state prediction methods, which struggle to effectively address high-dimensional variable coupling relationships, heterogeneous dependencies between different prediction objectives, and the forgetting of cross-furnace evolution information.

[0102] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0103] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0106] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0108] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for predicting the endpoint of converter steelmaking, characterized in that, include: Multi-scale importance modeling is performed on the multi-dimensional process characteristics of the current furnace, and importance weights for each process characteristic are generated. The structural coupling relationship between multi-dimensional process features is modeled by graph attention mechanism to generate the structure-aware weight of each process feature. By integrating importance weights and structure-aware weights, feature weights for each dimension of process features are obtained; Each process feature is multiplied by its corresponding feature weight to generate the feature control vector for the current furnace. Perform multi-scale convolution operation on the feature control vector of the current furnace batch to extract local pattern features under different receptive fields, and output the spatial feature vector of the current furnace batch after enhancement by a self-attention mechanism. Multiple Transformer encoders are used to process furnace feature sequences with different historical window lengths to extract cross-furnace temporal feature vectors; wherein, the furnace feature sequence is composed of the feature control vectors corresponding to the current furnace and its k preceding furnaces; The spatial feature vector and temporal feature vector are dimensionally aligned, and then weighted fusion is performed dimension-wise using learnable fusion weights. The predicted value of the steelmaking endpoint state for the current furnace is output by the regression prediction network.

2. The method according to claim 1, characterized in that, The operation of performing multi-scale importance modeling on the multi-dimensional process features of the current furnace and generating importance weights for each process feature includes: The multidimensional process characteristics of the current furnace are divided into local scale characteristics, global scale characteristics, and statistical scale characteristics. Among them, local scale characteristics are variables that are dynamically controlled during the blowing process, global scale characteristics are static variables that reflect the furnace charge structure, and statistical scale characteristics are statistical variables that characterize the process evolution trend. The first type of gated subnetwork generates importance weights for local scale features; the second type of gated subnetwork generates importance weights for global scale features; and the third type of gated subnetwork generates importance weights for statistical scale features. Among them, the first type of gated subnetwork adopts a single-layer nonlinear transformation structure, the second type of gated subnetwork adopts a multi-layer sensing structure, and the third type of gated subnetwork adopts the Softplus activation function.

3. The method according to claim 2, characterized in that, The operation of modeling the structural coupling relationship between multi-dimensional process features through graph attention mechanism and generating structure-aware weights for each dimension of process feature includes: Each dimension of the multi-dimensional process features in the current furnace is used as a graph node, and a connection edge is established between any two nodes to form a fully connected graph structure. The fully connected graph structure is input into the graph attention network to generate structure-aware weights for each dimension of process features.

4. The method according to claim 1, characterized in that, The operation of integrating importance weights and structure-aware weights to obtain the feature weights for each dimension of process features includes: For each dimension of process feature, the importance weight and structure awareness weight are dynamically fused using dedicated learnable fusion weights to obtain the corresponding feature weights.

5. The method according to claim 1, characterized in that, The process of performing multi-scale convolution operations on the feature control vector of the current furnace, extracting local pattern features under different receptive fields, and outputting the spatial feature vector of the current furnace after enhancement by a self-attention mechanism includes: Expand the feature control vector of the current furnace batch into a one-dimensional feature map; Perform n sets of convolution operations with different receptive fields in parallel on the one-dimensional feature map to obtain n sets of feature maps; The n sets of feature maps are concatenated along the channel dimension to form a fused feature map; For a self-attention network with independent input channels for the fused feature map, the attention-enhanced fused feature map is flattened into a one-dimensional vector, which is used as the spatial feature vector of the current furnace.

6. The method according to claim 1, characterized in that, The operation of using multiple Transformer encoders to process furnace feature sequences with different historical window lengths and extracting time-series feature vectors across furnaces includes: Multiple different window lengths are set, and for each window length, a corresponding furnace feature sequence is constructed; wherein, the furnace feature sequence includes the feature control vectors corresponding to the current furnace and the k furnaces before it, and k is different in different furnace feature sequences; Each furnace feature sequence is input into an independent Transformer encoder for processing, and the output is a feature representation of the corresponding window length. The feature representations corresponding to multiple different window lengths are concatenated along the feature dimension to obtain a time-series feature vector spanning multiple furnace cycles.

7. The method according to claim 1, characterized in that, The process of dimensional alignment of spatial and temporal feature vectors, dimensional weighted fusion through learnable fusion weights, and outputting the predicted end-point state of the current heat through a regression prediction network includes: Spatial feature vectors and temporal feature vectors are compressed or expanded to the same dimension through linear projection; The aligned spatial feature vector is concatenated with the temporal feature vector and then input into the gating network to generate a learnable fusion weight vector. The spatial feature vector and temporal feature vector are fused dimension-wise according to the generated learnable fusion weight vector to obtain the fused feature vector. The fused feature vector is input into the regression prediction network, which outputs the predicted value of the steelmaking endpoint state for the current heat.

8. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method described in any one of claims 1 to 7 is performed by a processor.

9. A converter steelmaking endpoint prediction device, characterized in that, include: The feature soft selection module for target differences is used to perform multi-scale importance modeling of the multi-dimensional process features of the current furnace and generate importance weights for each process feature. The structural coupling relationship between multi-dimensional process features is modeled by graph attention mechanism to generate the structure-aware weight of each process feature. By integrating importance weights and structure-aware weights, feature weights for each dimension of process features are obtained; Each process feature is multiplied by its corresponding feature weight to generate the feature control vector for the current furnace. The multi-scale spatial structure perception module is used to perform multi-scale convolution operations on the feature control vector of the current furnace, extract local pattern features under different receptive fields, and output the spatial feature vector of the current furnace after being enhanced by a self-attention mechanism. The multi-granularity temporal dependency modeling module is used to process furnace feature sequences with different historical window lengths using multiple sets of Transformer encoders to extract temporal feature vectors across furnaces; wherein, the furnace feature sequence consists of the feature adjustment vectors corresponding to the current furnace and its k preceding furnaces; The spatiotemporal feature adaptive fusion prediction module is used to align the spatial feature vector and the temporal feature vector in dimensions. It performs dimensional weighted fusion through learnable fusion weights and outputs the predicted value of the steelmaking endpoint state of the current furnace through a regression prediction network.

10. A converter steelmaking endpoint prediction device, characterized in that, include: processor; A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Multi-scale importance modeling is performed on the multi-dimensional process characteristics of the current furnace, and importance weights for each process characteristic are generated. The structural coupling relationship between multi-dimensional process features is modeled by graph attention mechanism to generate the structure-aware weight of each process feature. By integrating importance weights and structure-aware weights, feature weights for each dimension of process features are obtained; Each process feature is multiplied by its corresponding feature weight to generate the feature control vector for the current furnace. Perform multi-scale convolution operation on the feature control vector of the current furnace batch to extract local pattern features under different receptive fields, and output the spatial feature vector of the current furnace batch after enhancement by a self-attention mechanism. Multiple Transformer encoders are used to process furnace feature sequences with different historical window lengths to extract cross-furnace temporal feature vectors; wherein, the furnace feature sequence is composed of the feature control vectors corresponding to the current furnace and its k preceding furnaces; The spatial feature vector and temporal feature vector are dimensionally aligned, and then weighted fusion is performed dimension-wise using learnable fusion weights. The predicted value of the steelmaking endpoint state for the current furnace is output by the regression prediction network.

Citation Information

Patent Citations

  • Method and device for predicting phosphorus content at end point of converter

    CN117935968A

  • Converter end point carbon content prediction method based on graph structure working condition division

    CN119272026A

  • Converter steelmaking oxygen supply prediction method based on knowledge and data fusion driving

    CN119360997A

  • Method for controlling end point of steelmaking in converter

    JP1994264129A

  • Gas analysis-based dynamic control method for end-point carbon in whole converter smelting process

    WO2022198594A1

Cited By

  • Bridge technical condition prediction method based on fusion of graph neural network and time sequence modeling

    CN121959067A

  • Bridge technology condition prediction method based on fusion of graph neural network and time series modeling

    CN121959067B