Anode furnace refining sampling point key index prediction method, medium and computer equipment

By extracting flame video features using an improved R(2+1)D network and a 3D convolutional network, and combining a hybrid neural network of Bi-LSTM and MLP, a regression prediction model was constructed. This solved the problem of lag in key indicators during the anode furnace refining process, enabling real-time prediction and efficiency improvement.

CN121963014APending Publication Date: 2026-05-01NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEASTERN UNIV CHINA
Filing Date
2025-11-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the content of key indicators such as copper, oxygen, sulfur and nickel in the anode furnace refining process can only be obtained through sampling and testing, which results in a lag, leading to high production costs and low efficiency.

Method used

An improved R(2+1)D network and a 3D convolutional network are used to extract the spatiotemporal features of flame videos. A regression prediction model is constructed by combining a bidirectional long short-term memory network and a multilayer perceptron to predict key indicators in real time.

Benefits of technology

It enables real-time prediction of key indicators at the anode furnace refining sampling points, reducing the number of manual samplings, lowering production costs, and improving production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963014A_ABST
    Figure CN121963014A_ABST
Patent Text Reader

Abstract

The invention discloses an anode furnace refining sampling point key index prediction method, a medium and computer equipment, and the method comprises the steps: extracting the spatial-temporal characteristics of a flame video image through employing an R (2 + 1) D convolution network, introducing a 3D convolution-based spatial-temporal attention mechanism at a later sampling moment for weighting, dividing a training set, a verification set and a test set, and obtaining a training set, a verification set and a test set; a regression prediction model based on R (2 + 1) D-attention mechanism-BiLSTM is constructed by combining a hybrid neural network architecture of a bidirectional long short-term memory network and a multi-layer perceptron, key indexes of anode furnace refining sampling points are predicted, and finally evaluation is performed by adopting mean square errors and decision coefficient evaluation indexes. The key index prediction of the anode furnace refining sampling point can be realized, the problem that the index content test result has hysteresis quality is solved, the manual sampling frequency is reduced, the production cost is reduced, and the product quality and the production efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automated anode furnace refining technology, and in particular to a method, medium, and computer equipment for predicting key indicators at sampling points in anode furnace refining. Background Technology

[0002] Copper is primarily derived from pyrometallurgical processes, which mainly involve matte smelting, matte blowing, pyrometallurgical refining of blister copper, and electrolytic refining of anode copper. The purpose of pyrometallurgical refining of blister copper is to remove impurities and provide qualified anode copper for electrolytic refining; this process involves the anode furnace refining process. The anode furnace refining process is mainly divided into four stages: holding, oxidation, slag removal, and reduction. The end of the reduction stage represents the end of the entire refining process. Throughout the refining process, the main production cost comes from the consumption of fuel materials. Currently, the content of key indicators such as copper, oxygen, sulfur, and nickel before tapping can only be obtained through sampling and testing, and the results are lagging. Summary of the Invention

[0003] In view of this, this application provides a method, medium, and computer equipment for predicting key indicators at anode furnace refining sampling points, which can predict key indicators at anode furnace refining sampling points, avoid the problem of delayed sampling and testing results, thereby reducing production costs, improving production efficiency, and meeting capacity requirements.

[0004] According to one aspect of this application, a method for predicting key indicators at sampling points in an anode furnace refining process is provided, the method comprising: When the anode furnace reaction reaches the reduction period, the flame video at the furnace mouth is collected in real time, and the real values ​​of key indicators in the anode furnace at each sampling point during the reduction period are recorded synchronously. Multiple sampling points are set during the reduction period, and the sampling points during the reduction period represent the refining sampling points. The key indicators include at least one of copper content, oxygen content, sulfur content and nickel content. The acquired flame video is preprocessed to obtain multiple flame video images arranged in chronological order. The preprocessing includes video frame extraction at the sampling point, frame processing, and frame normalization. After improving the convolutional blocks of the R(2+1)D network, the improved R(2+1)D network is used to extract the original spatiotemporal feature matrix from the flame video image. The original spatiotemporal feature matrix is ​​obtained by splicing spatial features and temporal dynamics. The elements in the original spatiotemporal feature matrix represent spatiotemporal locations. The spatial features characterize the spatial information of the flame, including shape and / or texture. The temporal dynamics characterize the change law of the flame features over time. The improved R(2+1)D network is used again to extract the enhanced spatiotemporal feature matrix from the flame video images in the last preset time period of the flame video. The extracted enhanced spatiotemporal feature matrix is ​​convolved by a 3D convolution kernel to output an attention weight map with the same dimension as the original spatiotemporal feature matrix. After normalization, the attention weight map is multiplied element-wise with the original spatiotemporal feature matrix to obtain the weighted spatiotemporal features. The elements in the attention weight map represent the importance weight of the spatiotemporal position. After flattening the weighted spatiotemporal features into time series features, they are input together with the recorded true values ​​of key indicators into the regression prediction model. The regression prediction model is then trained based on the input weighted spatiotemporal features flattened into time series features and the true values ​​of key indicators until training is complete. New flame videos are acquired and new weighted spatiotemporal features are extracted. These features are then input into a trained regression prediction model, enabling the model to generate new predictions for key indicators.

[0005] According to another aspect of this application, a medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method for predicting key indicators at anode furnace refining sampling points.

[0006] According to another aspect of this application, a computer device is provided, including a medium, a processor, and a computer program stored on the medium and executable on the processor, wherein the processor executes the program to implement the above-described method for predicting key indicators at anode furnace refining sampling points.

[0007] Using the above technical solution, this application provides a method, medium, and computer equipment for predicting key indicators at anode furnace refining sampling points. First, it acquires and preprocesses video images of the flame at the anode furnace opening. Second, it uses an R(2+1)D convolutional network to extract the spatiotemporal features of the flame images, and introduces a 3D convolution-based spatiotemporal attention mechanism for weighting at later sampling times. Subsequently, it divides the data into training, validation, and test sets, and constructs a regression prediction model based on an R(2+1)D-attention mechanism-BiLSTM using a hybrid neural network architecture combining a bidirectional long short-term memory network (Bi-LSTM) and a multilayer perceptron (MLP) to predict key indicators at anode furnace refining sampling points. Finally, it uses mean squared error (MSE) and coefficient of determination (R²) as evaluation metrics. This method enables the prediction of key indicators at anode furnace refining sampling points, solving the problem of lag in indicator content testing results, thereby reducing the number of manual samplings, lowering production costs, improving product quality, and increasing production efficiency.

[0008] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a method for predicting key indicators at sampling points in an anode furnace refining process, as provided in an embodiment of this application, is shown. Figure 2 This illustration shows an architecture diagram for extracting the original spatiotemporal feature matrix according to an embodiment of this application; Figure 3 A schematic diagram of an attention mechanism provided in an embodiment of this application is shown. Detailed Implementation

[0010] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0011] This embodiment provides a method for predicting key indicators at anode furnace refining sampling points, such as... Figure 1 As shown, the method includes: Step 101: When the anode furnace reaction reaches the reduction period, the flame video at the furnace mouth is collected in real time, and the real values ​​of key indicators in the anode furnace at each sampling point during the reduction period are recorded synchronously. Multiple sampling points are set during the reduction period, and the sampling points during the reduction period represent refining sampling points. The key indicators include at least one of copper content, oxygen content, sulfur content and nickel content. Step 102: Preprocess the acquired flame video to obtain multiple flame video images arranged in chronological order. The preprocessing includes video frame extraction at the sampling point, frame processing, and frame normalization. Step 103: After improving the convolutional blocks of the R(2+1)D network, the improved R(2+1)D network is used to extract the original spatiotemporal feature matrix from the flame video image. The original spatiotemporal feature matrix is ​​obtained by splicing spatial features and temporal dynamics. The elements in the original spatiotemporal feature matrix represent spatiotemporal locations. Spatial features characterize the spatial information of the flame, including shape and / or texture. Temporal dynamics characterize the change law of flame features over time. Step 104: The improved R(2+1)D network is used again to extract the enhanced spatiotemporal feature matrix from the flame video images in the last preset time period of the flame video. The extracted enhanced spatiotemporal feature matrix is ​​convolved by a 3D convolution kernel to output an attention weight map with the same dimension as the original spatiotemporal feature matrix. After normalizing the attention weight map, it is multiplied element-wise with the original spatiotemporal feature matrix to obtain the weighted spatiotemporal features. The elements in the attention weight map represent the importance weight of the spatiotemporal position. Step 105: After flattening the weighted spatiotemporal features into time series features, input them together with the recorded true values ​​of key indicators into the regression prediction model, so that the regression prediction model can be trained based on the input weighted spatiotemporal features flattened into time series features and the true values ​​of key indicators until training is completed. Step 106: Acquire new flame videos and extract new weighted spatiotemporal features, then input them into the trained regression prediction model so that the trained regression prediction model can predict new values ​​for key indicators.

[0012] In the above embodiments of this application, an industrial camera can be used to capture flame videos of the furnace mouth during the reduction period of the anode furnace. The captured flame videos are preprocessed to obtain flame video images, including video frame extraction, frame processing, and frame normalization. An improved R(2+1)D network is used to extract the spatiotemporal features of the flame video images. That is, spatial features are first extracted using 2D convolution, and temporal dynamics are extracted using 1D convolution. The spatial features and temporal dynamics are then spliced ​​together to obtain spatiotemporal features. Next, the last 2 seconds (within the last preset time period) of the flame video sampling points are weighted using a spatiotemporal attention mechanism based on 3D convolution to train a regression prediction model using BiLSTM, which is used to predict the key indicators at each sampling point during the refining process.

[0013] Optionally, in step 103, the convolutional blocks of the R(2+1)D network are improved, including: Step 1031: Improve the residual connections in the R(2+1)D network into sequentially connected convolutional blocks, wherein the convolutional blocks include spatial convolutional blocks and temporal convolutional blocks, the spatial convolutional blocks are used to perform 2D convolution, and the temporal convolutional blocks are used to perform 1D convolution.

[0014] In the above embodiments of this application, in order to better adapt to the industrial environment, the R(2+1)D network can be improved, that is, instead of using residual connections, a sequential structure is adopted; the sequential structure is as follows: in the first convolutional block, the (1,5,5) kernel first performs convolution in the pure spatial dimension (H×W), and the subsequent (3,1,1) kernel focuses on feature extraction in the temporal dimension (T). The second convolutional block repeats the first convolutional block, and the sequential connection simplifies the structure.

[0015] Optionally, in step 103, an improved R(2+1)D network is used to extract the original spatiotemporal feature matrix from the flame video image, including: Step 1032: First, use the 2D convolution in the improved R(2+1)D network to extract the spatial features of the flame video image, and then use the 1D convolution in the improved R(2+1)D network to extract the temporal dynamics of the flame video image. Then, stitch together the spatial features and the temporal dynamics to obtain the original spatiotemporal feature matrix.

[0016] In the above embodiments of this application, the improved R(2+1)D network decomposes the three-dimensional convolutional kernel into two parts: a two-dimensional spatial convolutional kernel (2D convolution) and a one-dimensional temporal convolutional kernel (1D convolution). First, spatial features are extracted using 2D convolution, and then temporal dynamics are extracted using 1D convolution. The two-dimensional spatial convolutional kernel performs primary extraction of spatiotemporal features. The first layer can use a large (1,5,5) kernel to enhance local feature capture capabilities, and a stride of (1,2,2) is used to achieve spatial downsampling. The temporal convolutional kernel size can be 3 frames (3,1,1) to cover short-term motion patterns. Max pooling (1,3,3) further compresses the feature map size, preserving salient features, while the stride (1,2,2) reduces computational cost.

[0017] Specifically, such as Figure 2 As shown, the improved R(2+1)D network decomposes the traditional 3D convolution into 2D spatial convolution + 1D temporal convolution. It extracts spatial features and temporal dynamics through two independent convolutional blocks, and then concatenates them to obtain the original spatiotemporal feature matrix. Figure 2 This demonstrates two repeating convolutional block structures (first convolutional block and second convolutional block), each following a "spatial convolution - temporal convolution" sequence. Furthermore, hierarchical extraction and downsampling of spatiotemporal features are achieved through parameter design. Specifically: The first convolutional block performs primary extraction of spatiotemporal features. The first layer uses a large kernel of (1,5,5) to enhance the local feature capture capability, and a stride of (1,2,2) is used to achieve spatial downsampling. The temporal convolutional kernel size is 3 frames (3,1,1) to cover short-term motion modes. The feature map size is further compressed by max pooling of (1,3,3) to retain significant features, while the stride of (1,2,2) reduces the amount of computation. The second convolutional block performs advanced feature abstraction and regularization. The spatial kernel is shrunk to (1,3,3), and the spatial stride of (1,2,2) further compresses the feature map size to 1 / 4 of the original input. Dropout3d is used to mask random channels in the entire 64 channels with a 30% probability, and adaptive spatial pooling forces the output to have a spatial size of 4×4.

[0018] Furthermore, for the spatial feature extraction part of 2D convolution: the convolution kernel is designed as follows: the first layer uses a large kernel of (1,5,5) (the kernel size in the time dimension is 1 and the spatial dimension is 5×5), combined with a stride of (1,2,2). The large kernel can enhance the ability to capture local features (such as the shape and texture of the flame). The stride of (1,2,2) realizes spatial downsampling (halving the size of the H / W dimension) and reduces the amount of subsequent computation.

[0019] The operation process for the first convolutional block is as follows: 1. Spatial Convolution: Input a flame video image (dimensions B×T×H×W, where B is the batch size and T is the time step), and perform convolution in the spatial dimension (H×W) using a (1,5,5) convolution kernel to output intermediate features (the number of channels increases, T remains unchanged in the spatiotemporal dimension, and H / W may change).

[0020] 2. 3D Batch Normalization: Normalizes the convolution output to accelerate training convergence.

[0021] 3. GELU Activation: Introducing non-linearity enhances feature representation. Specifically, GELU is an activation function based on a Gaussian error function. Compared to activation functions like ReLU, GELU is smoother, which helps improve the convergence speed and performance of the training process. By replacing ReLU with GELU, the introduction of non-linearity enhances feature representation.

[0022] 4. Spatial Pooling: Using (1,3,3) max pooling (1 kernel for the temporal dimension, 3×3 spatial dimension), salient features are preserved, while a stride of (1,2,2) reduces spatial resolution. The type and position of the pooling layers are adjusted, such as adaptive pooling. The spatial size is dynamically adjusted to accommodate different input resolutions and compress redundant information. For extracting the temporal dynamics using 1D convolution: the convolution kernel is designed as a (3,1,1) kernel (3 kernels in the temporal dimension and 1×1 in the spatial dimension), covering short-term motion patterns (such as the dynamic changes of flames). This enables the capture of temporal correlations of flame features in the temporal dimension (T), such as short-term motion patterns (changes within a 3-frame window).

[0023] The operation flow for outputting the spatial features of the first convolutional block is as follows: 1. Temporal Convolution: The output of spatial convolution (which already contains spatial features) is convolved in the temporal dimension using a (3,1,1) convolution kernel to output temporal dynamic features (the T dimension may change, while H / W remains constant).

[0024] 2. Adaptive pooling / normalization Figure 2Example of the second convolutional block: The second convolutional block repeats the "spatial convolution-temporal convolution" structure, adding Dropout regularization or adaptive pooling to further compress the feature size and enhance robustness.

[0025] For the splicing of spatial features and temporal dynamics, the splicing method is to splice the spatial features (including shape and texture information) extracted by 2D convolution and the temporal dynamics (including temporal change information) extracted by 1D convolution in the channel dimension to form the original spatiotemporal feature matrix.

[0026] Assuming the spatial feature dimension is T×H′×W′×C1 and the temporal dynamic dimension is T′×H′×W′×C2, the concatenated original spatiotemporal feature matrix has a dimension of T′′×H′×W′×(C1+C2) (T / H / W dimensions aligned). This allows for the fusion of static spatial information and dynamic temporal information of flames, providing a foundation for subsequent attention mechanisms and prediction models.

[0027] Therefore, by Figure 2 As can be seen, sequentially connected convolutional blocks replace residual connections, and each convolutional block strictly follows the order of "spatial convolution - temporal convolution," which simplifies the computation process. The parameter adaptation uses a first-layer design with a (1,5,5) large kernel and a stride of (1,2,2), directly addressing the need for "enhanced local feature capture capability + spatial downsampling." The (3,1,1) temporal convolutional kernel covers short-term motion, matching the dynamic analysis scenario of flames. The hierarchical extraction uses a repeating structure of two convolutional blocks to achieve layer-by-layer abstraction of spatiotemporal features, adapting to the representation needs of complex flame videos.

[0028] In summary, the improved R(2+1)D network, through decomposing the convolutional kernel, sequentially connecting the convolutional blocks, and optimizing the parameters, can efficiently extract the original spatiotemporal feature matrix of flame videos, laying the foundation for subsequent enhanced feature extraction and key indicator prediction.

[0029] Optionally, the regression prediction model is a hybrid neural network architecture combining a bidirectional long short-term memory network and a multilayer perceptron. The hybrid neural network architecture includes a bidirectional LSTM, a vector result concatenation layer, and a regression prediction head. The bidirectional LSTM includes a forward LSTM and a backward LSTM. The forward LSTM processes data from the start position to the end position of the sequence until it captures the forward context information, and the backward LSTM processes data from the end position to the start position of the sequence until it captures the backward context information. The vector result concatenation layer is used to fuse the outputs of the bidirectional LSTM, connecting the final hidden states of the forward LSTM and the backward LSTM into a feature vector. The regression prediction head adopts a multilayer perceptron structure and outputs the predicted value.

[0030] In the embodiments described above, the regression prediction module uses a hybrid neural network architecture combining a bidirectional long short-term memory network (Bi-LSTM) and a multilayer perceptron (MLP). This architecture consists of three parts: a bidirectional LSTM, a vector result concatenation layer, and a regression prediction head. The bidirectional LSTM contains two LSTM layers, one forward and one backward, which process temporal features from the forward and backward directions respectively. The output features are fused into a unified feature vector by the vector result concatenation layer. The regression prediction head adopts a multilayer perceptron structure, which successively reduces the feature dimension from 256 dimensions to 128 dimensions and then to 64 dimensions, finally mapping it to a 4-dimensional output space to generate the prediction result (predicted value).

[0031] Furthermore, bidirectional LSTM specifically involves: forward LSTM processing data from the start position to the end position of the sequence, capturing forward context information, and backward LSTM processing data from the end position to the start position of the sequence, capturing backward context information; The vector concatenation layer specifically merges the outputs of the bidirectional LSTM and connects the final hidden states of the forward and backward LSTMs into a unified feature vector. The regression prediction head specifically employs a multilayer perceptron structure. First, the concatenated vectors are normalized to stabilize the training process, compressing high-dimensional features into a 128-dimensional space. Overfitting is prevented by random deactivation through a Dropout layer. Feature normalization is performed again to enhance training stability. Nonlinear transformation capability is introduced to further reduce the dimensionality to a 64-dimensional feature space. Dropout layer random deactivation regularization is applied again, and the normalization layer maintains the stability of the feature distribution. ReLU increases the nonlinear expressive capability. Finally, the result is mapped to a 4-dimensional output space to generate the prediction result (predicted value).

[0032] Optionally, in step 101, real-time acquisition of flame video at the furnace opening includes: Step 1011: Use an industrial camera to collect real-time video of the flame at the furnace opening. The collected flame video is pre-processed after being confirmed by the workers at the refining site.

[0033] In the above embodiments of this application, an industrial camera can be used to capture flame videos of the furnace mouth during the reduction period of the anode furnace. These flame videos during the refining and reduction period of the anode furnace are confirmed by on-site workers to ensure that the captured flame videos can be preprocessed. For each video segment, the same number of frames are uniformly sampled at specific time points and then standardized to eliminate size differences. During preprocessing, frame extraction and standardization are performed on the flame videos, and the frames are scaled to a uniform resolution to eliminate size differences.

[0034] Optionally, the acquired flame video is in BGR format. In step 102, the acquired flame video is preprocessed, including: Step 1021: Convert the BGR format flame video to RGB format and preprocess the RGB format flame video.

[0035] In the above embodiments of this application, the BGR format is converted to RGB format to adapt to the input requirements of deep learning models.

[0036] Optionally, in step 104, the extracted enhanced spatiotemporal feature matrix is ​​convolved using a 3D convolution kernel to output an attention weight map of the same dimension as the original spatiotemporal feature matrix, including: Step 1041: The extracted enhanced spatiotemporal feature matrix is ​​activated by the first layer of 3D convolution, BatchNorm and GELU, then regularized by Dropout, and finally activated by a 3D convolution that reduces the number of channels to 1 and Sigmoid to generate an attention weight map between 0 and 1.

[0037] In the above embodiments of this application, based on prior knowledge, the last 2 seconds of each flame video segment are considered the sampling point. Therefore, the video feature sequence extracted by the improved R(2+1)D network is divided into two parts: the first part retains the original features, and the second part is weighted through a spatiotemporal attention mechanism based on 3D convolution. Drawing inspiration from the idea of ​​generating attention weights through convolution in CBAM spatial attention, a single-channel attention map is generated through a series of convolution operations, and finally mapped to the [0,1] interval using Sigmoid. That is, after activation by the first layer of 3D convolution, BatchNorm, and GELU, dropout is used to increase regularity, and finally, an attention weight map between 0 and 1 is generated through a 3D convolution that reduces the number of channels to 1 and Sigmoid activation.

[0038] Specifically, such as Figure 3 As shown, for example, if the captured flame video is 5 seconds long, based on the prior knowledge that "the last 2 seconds (i.e., the last 12 frames) are the sampling point moments," the video is divided into two parts: The first 18 frames (first 3 seconds): serve as background timing information, retaining the original features unchanged; The last 12 frames (last 2 seconds): As the key period of the sampling point, its key spatiotemporal features need to be highlighted through the 3D attention mechanism (i.e. the processing object of step 1041).

[0039] The core of step 1041 is to generate an attention weight map of the same dimension from the enhanced spatiotemporal feature matrix of the last 12 frames through multi-step 3D convolution and nonlinear transformation. Figure 3 The dashed box on the right side illustrates this process in detail, with the specific correspondence as follows: The improved R(2+1)D network first extracts the enhanced spatiotemporal feature matrix from the last 12 frames (critical time periods), denoted as Fenhance=T×H×W×C, where: 1. Input to the enhanced spatiotemporal feature matrix: T=12: Time step (corresponding to the last 12 frames). H / W: Spatial dimensions of the flame image (e.g., 64×64); C: Number of feature channels (e.g., 64 or 128, determined by the depth of the R(2+1)D network).

[0040] 2. First layer 3D convolution ( Figure 3 (“3×3×3 convolution”): Apply a 3D convolution kernel size (3,3,3) (3 for the time dimension and 3×3 for the spatial dimension) to the Fenhance enhanced spatiotemporal feature matrix, and set the stride (1,1,1) and padding (1,1,1).

[0041] Dimension Preservation Principle: Based on the formula for the output dimension of 3D convolution: , After substituting the parameters: Time dimension (consistent with the input); , Spatial dimension (consistent with input): , , Channel Dimension: Mapped from C to an intermediate channel number C1 (e.g., 128, used to capture deep spatiotemporal correlations). Short-term spatiotemporal features are extracted within the last 12 frames (temporal dimension 3 covers 3 frames of dynamics, spatial dimension 3×3 covers local texture / shape), corresponding to... Figure 3 The "3×3×3 Convolution" module.

[0042] 3. BatchNorm and GELU activation ( Figure 3 (Batch normalization and GELU activation function in the text) The output of the first 3D convolution layer is processed sequentially as follows: Batch normalization: Normalizes the channel dimension C1C_1C1 to reduce feature distribution offset and accelerate model training convergence; GELU activation: Introduces nonlinear transformation to enhance the expressive power of features (such as capturing complex spatiotemporal relationships such as abrupt changes in flame shape and brightness).

[0043] 4. Dropout regularization: By using a Dropout layer (dropoutrate=0.5 can be set) to randomly discard some neurons, the model can be prevented from overfitting, the attention mechanism can be prevented from relying too much on features of a few frames or a few spatial locations, and the generalization ability can be enhanced.

[0044] 5. 3D convolution with channels reduced to 1 ( Figure 3 (“3×3×3 convolutional layer”): Using a 3D convolution kernel with the same size (3,3,3) (maintaining consistent spatiotemporal dimensions), but setting the number of kernels to 1, the number of intermediate channels C1 is compressed to 1 dimension. Simultaneously, the stride (1,1,1) and padding (1,1,1) are maintained to ensure: The spatiotemporal dimension remains 12×H×W (consistent with the input enhancement features); The channel dimension is reduced from C1 to 1 (generating a single-channel feature map, denoted as Fsingle=12×H×W×1).

[0045] 6. Sigmoid activation generates an attention weight map: Apply the Sigmoid activation function to the single-channel feature map Fsingle to map the feature values ​​to the [0,1] interval, generating the final attention weight map (denoted as Watt=12×H×W×1).

[0046] Key dimension matching: The spatiotemporal dimension (12×H×W) of the weight map is completely consistent with the original enhanced spatiotemporal feature matrix (12×H×W×C), satisfying the requirement of "element-by-element multiplication". Each element in the weight map corresponds to the importance of the same spatiotemporal position (a certain frame, a certain spatial pixel) in the original feature matrix (0 indicates irrelevant, 1 indicates key).

[0047] Application of attention weights and feature concatenation ( Figure 3 (Left half), after generating the attention weight map, the process returns to... Figure 3 The left half, specifically: 1. Weighted Enhanced Features: Multiply the attention weight map Watt element-wise with the enhanced spatiotemporal feature matrix Fenhance of the last 12 frames to obtain the weighted features of the last 12 frames (the dimensions are still 12×H×W×C). 2. Feature concatenation: The weighted features of the last 12 frames are concatenated with the original features of the first 18 frames (dimension 18×H×W×C) to obtain a complete spatiotemporal feature sequence (dimension 30×H×W×C). 3. Input into long-term time series model: The spliced ​​feature sequence is finally input into an LSTM long-term time series model for the prediction of key indicators (copper / oxygen / sulfur / nickel content).

[0048] Throughout the process, prior knowledge ("the last 2 seconds are the sampling point") is reflected through "focusing on the last 12 frames + attention weighting." Enhanced features are extracted and attention is applied only to the key time periods of the last 12 frames, avoiding interference from background information in the first 18 frames. The attention weight map accurately depicts the key spatiotemporal locations of the flame within the last 12 frames (such as sudden changes in flame color in certain frames, or abnormal flame shape in a certain spatial region), providing high signal-to-noise ratio feature input for subsequent predictions. Therefore, step 1041... Figure 3 The right half of the module, "3×3×3 convolution—batch normalization—GELU—Dropout—3×3×3 convolution—Sigmoid," implements the transformation from "enhanced spatiotemporal features to same-dimensional attention weight map"; simultaneously, Figure 3 The “feature segmentation-weighting-assembly” process in the left half combines prior knowledge with attention mechanisms to ultimately output a high-value spatiotemporal feature sequence for prediction.

[0049] Optionally, in step 104, after obtaining the weighted spatiotemporal features, the method further includes: Step 107: Based on the weighted spatiotemporal features and the true values ​​of key indicators, a set of datasets is formed. Based on the datasets obtained from multiple restoration periods, training set, test set and validation set are divided. Accordingly, in step 106, after the weighted spatiotemporal features are flattened into time series features, they are input together with the recorded true values ​​of key indicators into the regression prediction model. This allows the regression prediction model to be trained based on the input weighted spatiotemporal features flattened into time series features and the true values ​​of key indicators, until training is complete. This includes: Step 1061: After flattening the weighted spatiotemporal features in the training set into time series features, input them together with the true values ​​of the key indicators in the training set into the regression prediction model so that the regression prediction model can be trained. Step 1062: During model training, the weighted spatiotemporal features of the validation set are flattened into time series features and then input into the regression prediction model. The mean square error and the coefficient of determination are calculated using the predicted values ​​of the key indicators output by the regression prediction model and the true values ​​of the key indicators in the validation set. The hyperparameters of the regression prediction model are adjusted based on the calculation results until the optimized regression prediction model is obtained. Step 1063: Flatten the weighted spatiotemporal features of the test set into time series features and input them into the optimized regression prediction model. Use the predicted values ​​of the key indicators output by the optimized regression prediction model and the true values ​​of the key indicators in the test set to recalculate the mean square error and the coefficient of determination. Based on the calculation results, determine whether the optimized regression prediction model has reached the training completion standard. Step 1064: When the training completion standard is met, the optimized regression prediction model has completed model training. The completed regression prediction model outputs new predicted values ​​of key indicators through the input of new weighted spatiotemporal features.

[0050] In the above embodiments of this application, the ratio of training set, validation set, and test set can be 5:1:1. The training and validation sets are used to train and validate the model using pre-set training parameters, resulting in a regression prediction model based on R(2+1)D-attention mechanism-BiLSTM. The test set is then input into the obtained R(2+1)D-attention mechanism-BiLSTM regression prediction model for prediction, achieving prediction of key indicators at the anode furnace refining sampling point, and using mean squared error (MSE) and coefficient of determination (R²). 2 The model's prediction results are evaluated.

[0051] The training parameters can be set as follows: batch_size is 2, random seed is 42, training epochs are 100, learning rate is le-3, learning rate scheduling is OneCycleLR, optimizer is AdamW, weight decay is 0.03, precision mode is mixed precision (FP16), and loss function is mean squared error (MSE). Specifically, during training, the model parameters are updated using the training set, and the model convergence is monitored using the validation set to prevent overfitting.

[0052] By applying the technical solution of this embodiment, the flame video images of the anode furnace opening are first acquired and preprocessed. Then, the spatiotemporal features of the flame images are extracted using an R(2+1)D convolutional network, and a 3D convolution-based spatiotemporal attention mechanism is introduced for weighting at later sampling times. Subsequently, training, validation, and test sets are divided, and a regression prediction model based on R(2+1)D-attention mechanism-BiLSTM is constructed using a hybrid neural network architecture combining a bidirectional long short-term memory network (Bi-LSTM) and a multilayer perceptron (MLP) to predict key indicators at the anode furnace refining sampling points. Finally, the mean squared error (MSE) and coefficient of determination (R²) are used for evaluation. This method enables the prediction of key indicators at the anode furnace refining sampling points, solving the problem of lag in indicator content testing results, thereby reducing the number of manual samplings, lowering production costs, improving product quality, and increasing production efficiency.

[0053] Based on the above, Figures 1 to 3 Accordingly, this application also provides a medium on which a computer program is stored, which, when executed by a processor, implements the above-described method. Figures 1 to 3 The method for predicting key indicators at the anode furnace refining sampling point is shown.

[0054] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods described in various implementation scenarios of this application.

[0055] Based on the above, Figures 1 to 3 To achieve the above objectives, this application also provides a computer device, specifically a personal computer, server, network device, etc., as shown in the method. The computer device includes a medium and a processor; the medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figures 1 to 3 The method for predicting key indicators at the anode furnace refining sampling point is shown.

[0056] Optionally, the computer device may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB ports, card reader ports, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Bluetooth interfaces, Wi-Fi interfaces), etc.

[0057] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0058] The medium may also include an operating method and a network communication module. The operating method is a program that manages and stores the hardware and software resources of the computer device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the medium, as well as communication with other hardware and software within the physical device.

[0059] Through the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented using hardware. First, the flame video images of the anode furnace opening are acquired and preprocessed. Second, the spatiotemporal features of the flame images are extracted using an R(2+1)D convolutional network, and a spatiotemporal attention mechanism based on 3D convolution is introduced for weighting at later sampling times. Subsequently, training, validation, and test sets are divided, and a regression prediction model based on R(2+1)D-attention mechanism-BiLSTM is constructed using a hybrid neural network architecture combining a bidirectional long short-term memory network (Bi-LSTM) and a multilayer perceptron (MLP) to predict key indicators at the anode furnace refining sampling points. Finally, the mean squared error (MSE) and coefficient of determination (R²) are used for evaluation. This method can predict key indicators at the anode furnace refining sampling points, solving the problem of lag in indicator content test results, thereby reducing the number of manual samplings, lowering production costs, improving product quality, and increasing production efficiency.

[0060] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the methods of the embodiment can be distributed within the methods of the embodiment as described, or they can be modified and located in one or more methods different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.

[0061] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any modifications that can be made by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for predicting key indicators at anode furnace refining sampling points, characterized in that, The method includes: When the anode furnace reaction reaches the reduction period, the flame video at the furnace mouth is collected in real time, and the real values ​​of key indicators in the anode furnace at each sampling point during the reduction period are recorded synchronously. Multiple sampling points are set during the reduction period, and the sampling points during the reduction period represent the refining sampling points. The key indicators include at least one of copper content, oxygen content, sulfur content and nickel content. The acquired flame video is preprocessed to obtain multiple flame video images arranged in chronological order. The preprocessing includes video frame extraction at the sampling point, frame processing, and frame normalization. After improving the convolutional blocks of the R(2+1)D network, the improved R(2+1)D network is used to extract the original spatiotemporal feature matrix from the flame video image. The original spatiotemporal feature matrix is ​​obtained by splicing spatial features and temporal dynamics. The elements in the original spatiotemporal feature matrix represent spatiotemporal locations. The spatial features characterize the spatial information of the flame, including shape and / or texture. The temporal dynamics characterize the change law of the flame features over time. The improved R(2+1)D network is used again to extract the enhanced spatiotemporal feature matrix from the flame video images in the last preset time period of the flame video. The extracted enhanced spatiotemporal feature matrix is ​​convolved by a 3D convolution kernel to output an attention weight map with the same dimension as the original spatiotemporal feature matrix. After normalization, the attention weight map is multiplied element-wise with the original spatiotemporal feature matrix to obtain the weighted spatiotemporal features. The elements in the attention weight map represent the importance weight of the spatiotemporal position. After flattening the weighted spatiotemporal features into time series features, they are input together with the recorded true values ​​of key indicators into the regression prediction model. The regression prediction model is then trained based on the input weighted spatiotemporal features flattened into time series features and the true values ​​of key indicators until training is complete. New flame videos are acquired and new weighted spatiotemporal features are extracted. These features are then input into a trained regression prediction model, enabling the model to generate new predictions for key indicators.

2. The method according to claim 1, characterized in that, The extraction of the original spatiotemporal feature matrix from the flame video image using the improved R(2+1)D network includes: First, the spatial features of the flame video image are extracted using 2D convolution in the improved R(2+1)D network. Then, the temporal dynamics of the flame video image are extracted using 1D convolution in the improved R(2+1)D network. The spatial features and temporal dynamics are then spliced ​​together to obtain the original spatiotemporal feature matrix.

3. The method according to claim 2, characterized in that, The improvement to the convolutional blocks of the R(2+1)D network includes: The residual connections in the R(2+1)D network are improved into sequentially connected convolutional blocks, which include spatial convolutional blocks and temporal convolutional blocks. The spatial convolutional blocks are used to perform 2D convolutions, and the temporal convolutional blocks are used to perform 1D convolutions.

4. The method according to claim 1, characterized in that, The regression prediction model is a hybrid neural network architecture combining a bidirectional long short-term memory network (LSTM) and a multilayer perceptron. This architecture includes a bidirectional LSTM, a vector concatenation layer, and a regression prediction head. The bidirectional LSTM comprises a forward LSTM and a backward LSTM. The forward LSTM processes data from the start to the end of the sequence until forward context information is captured, while the backward LSTM processes data from the end to the start until backward context information is captured. The vector concatenation layer fuses the outputs of the bidirectional LSTMs, connecting the final hidden states of the forward and backward LSTMs into a feature vector. The regression prediction head uses a multilayer perceptron structure to output predicted values.

5. The method according to claim 1, characterized in that, Real-time acquisition of flame video from the furnace opening, including: Industrial cameras are used to capture real-time video of the flames at the furnace opening. The captured flame videos are then pre-processed after being confirmed by workers at the refining site.

6. The method according to claim 1, characterized in that, The acquired flame video is in BGR format. After improving the convolutional blocks of the R(2+1)D network, the acquired flame video undergoes preprocessing, including: Convert the BGR format flame video to RGB format, and then preprocess the RGB format flame video.

7. The method according to claim 1, characterized in that, After obtaining the weighted spatiotemporal features, the method further includes: Based on the weighted spatiotemporal features and the true values ​​of key indicators, a set of datasets is formed. Based on the datasets obtained from multiple restoration periods, training set, test set and validation set are divided. Accordingly, after flattening the weighted spatiotemporal features into time series features, they are input together with the recorded true values ​​of key indicators into the regression prediction model. This allows the regression prediction model to be trained based on the input flattened weighted spatiotemporal features (time series features) and the true values ​​of key indicators, until training is complete. This includes: After flattening the weighted spatiotemporal features in the training set into time series features, they are input together with the true values ​​of the key indicators in the training set into the regression prediction model so that the regression prediction model can be trained. During model training, the weighted spatiotemporal features of the validation set are flattened into time series features and then input into the regression prediction model. The mean square error and coefficient of determination are calculated using the predicted values ​​of the key indicators output by the regression prediction model and the true values ​​of the key indicators in the validation set. The hyperparameters of the regression prediction model are adjusted based on the calculation results until the optimized regression prediction model is obtained. The weighted spatiotemporal features of the test set are flattened into time series features and input into the optimized regression prediction model. The predicted values ​​of key indicators output by the optimized regression prediction model and the true values ​​of key indicators in the test set are used to recalculate the mean square error and the coefficient of determination. Based on the calculation results, it is determined whether the optimized regression prediction model has reached the training completion standard. When the training completion criteria are met, the optimized regression prediction model has completed model training. The completed regression prediction model outputs new predicted values ​​of key indicators by taking new weighted spatiotemporal features as input.

8. The method according to any one of claims 1 to 7, characterized in that, The extracted enhanced spatiotemporal feature matrix is ​​convolved using a 3D convolution kernel, outputting an attention weight map of the same dimension as the original spatiotemporal feature matrix, including: The extracted enhanced spatiotemporal feature matrix is ​​activated by a first-layer 3D convolution, BatchNorm, and GELU, then regularized by Dropout, and finally activated by a 3D convolution that reduces the number of channels to 1 and Sigmoid, generating an attention weight map between 0 and 1.

9. A medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for predicting key indicators at the anode furnace refining sampling point as described in any one of claims 1 to 8.

10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for predicting key indicators at the anode furnace refining sampling point as described in any one of claims 1 to 8.