Semi-supervised learning and free threshold-based magnesium electric smelting furnace working condition identification method and system

Through the semi-supervised learning and free threshold method, combined with the global channel space attention module and the CNN model of the dual-time fusion mechanism, the problem of high manual inspection dependence and labeling cost in the electromelting magnesium furnace operating condition recognition is solved, and efficient and accurate operating condition recognition is achieved.

CN120107646APending Publication Date: 2025-06-06HEFEI UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510004036.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art relies on manual inspection and a large amount of labeled data in the identification of the operating conditions of the electromelting magnesium furnace, which leads to high costs and prone to misjudgment or misjudgment, affecting the robustness of the model and the practical application effect.

Method used

Using semi-supervised learning and free threshold method, by collecting historical working conditions and video picture sample data of the electric melting magnesium furnace, using the global channel space attention module and the CNN model based on the timing dual-time fusion mechanism for feature extraction and model training, a qualified pseudo-label data set is generated, and the global and local thresholds are iteratively updated to improve the model performance.

Benefits of technology

It reduces the cost of manual marking, improves the accuracy and efficiency of the working condition recognition of the electromelting magnesium furnace, and enhances the robustness and practical application effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107646A_ABST
    Figure CN120107646A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial magnesium smelting intelligent control, and provides a semi-supervised learning and free threshold value-based electric smelting furnace working condition identification method, which comprises the following steps of: acquiring historical working conditions and video picture sample data of an electric smelting furnace, and dividing the historical working conditions and the video picture sample data into labeled data xa and unlabeled data xb; performing feature enhancement extraction on the preprocessed picture data by adopting a ViT feature extraction network based on a global channel space attention module; performing CNN model training of a time sequence-based dual-time fusion mechanism on the labeled feature extraction picture Xaout to obtain an initial model; sending the unlabeled image data Xbout into the initial model to predict the category probability, and generating a qualified pseudo-label data set by adopting a free threshold mechanism; and updating the training data set, carrying out model training again, calculating training loss, and carrying out iterative training continuously until the model performance meets a preset requirement. According to the method, the manual marking cost can be greatly reduced, and the working condition of the electric smelting magnesia furnace can be efficiently and accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control of industrial magnesium smelting, and in particular to a method and system for identifying working conditions of a molten magnesium furnace based on semi-supervised learning and free threshold. Background Art

[0002] Fused magnesia is made by melting natural magnesia or high-purity light-burned magnesia particles in an electric arc furnace. This material is known for its high melting point, dense structure, excellent chemical resistance, high compressive strength, strong corrosion resistance and stable chemical properties. It is a high-quality high-temperature electrical insulation material. It is not only a key raw material for the production of high-end magnesia bricks, magnesia carbon bricks and monolithic refractory materials, but also plays an important role in aerospace, national defense, metallurgy, chemical industry and scientific research. The production of fused magnesia mainly depends on a three-phase AC fused magnesium furnace. This furnace is cylindrical and uses AC arc furnace technology with magnesia as raw material. The electrical energy is converted into heat energy through a high current (reaching the 10,000A level), so that the magnesia is melted in the furnace and undergoes a series of purification reactions. Subsequently, the liquid magnesium oxide crystallizes during the cooling process, separates from the impurities, and finally obtains high-purity magnesium oxide crystals.

[0003] The smelting process of the fused magnesium furnace includes furnace start-up, material addition, normal smelting, under-burning and other working conditions. Among them, the under-burning condition usually occurs because the impurities in the raw materials cause the raw materials to not be completely melted, and the resulting bubbles cause the local temperature in the furnace to be too high. If the under-burning condition is not discovered and handled in time, it will not only greatly reduce the quality of the product, but may also cause major accidents such as furnace wall burning and leakage of molten raw materials, threatening personnel safety. Therefore, timely judgment of the under-burning condition is very important for the preparation of fused magnesia. In actual production, the under-burning condition is mainly judged by manual inspection and observation of the furnace wall and flame status. However, this method relies on the experience of the operator and is prone to misjudgment or omission. At the same time, the production site environment is harsh, and there is a risk to personal safety.

[0004] At present, the research on the identification of under-burned working conditions of electric fused magnesium furnaces can be divided into two categories. The first method is based on the three-phase current change pattern, analyzes the statistical characteristics of historical currents through rule inference algorithms, establishes a set of expert rule bases for working condition judgment, and identifies under-burned working conditions based on the current data collected in real time on site. However, due to the large amount of noise in the current signal, the effect of working condition judgment based on current characteristics alone is not good, so this method is usually only used as an auxiliary means. The second method establishes a working condition perception model by analyzing the monitoring images of the electric fused magnesium furnace, especially the image information of the furnace wall and furnace mouth flame. Conventional research methods, such as supervised learning, rely on a large amount of labeled data to train models with excellent performance. However, as a heavy industrial equipment, the labeling of the working condition data of the electric fused magnesium furnace requires experienced workers. This not only makes the production cost of the data set extremely high, but also the labeled data may have a large number of feature coupling problems, which will reduce the validity of the data. More importantly, feature coupling may also cause overfitting of the model, which in turn affects the robustness of the model and the actual application effect. Summary of the invention

[0005] The technical problem to be solved by the present invention is how to reduce the cost of manual marking and design an efficient and accurate method for identifying the working condition of a fused magnesium furnace.

[0006] The present invention solves the above technical problems through the following technical means:

[0007] The present invention provides a method for identifying working conditions of a molten magnesium furnace based on semi-supervised learning and free threshold, comprising the following steps:

[0008] S1. Collect historical working conditions and video image sample data of the fused magnesium furnace and divide them into labeled data x a and unlabeled data x b ;

[0009] S2, preprocessing the initial image sample data obtained in step S1;

[0010] S3, using the ViT feature extraction network based on the global channel spatial attention module to perform feature enhancement extraction on the preprocessed image data;

[0011] S4. Extract feature images X with labels a_out Conduct CNN model training based on dual-time fusion mechanism of time series;

[0012] S5, the unlabeled image data X preprocessed in step S2 b_out The trained model is fed into the model to predict the category probability, and a free threshold mechanism is used to generate a qualified pseudo-label dataset.

[0013] S6, update the training data set, re-train the model according to the method described in step S4, and continuously iterate the training by calculating the training loss until the model performance meets the preset requirements;

[0014] S7. The final model is used to output the working condition discrimination result using the verification data set to verify the model performance.

[0015] Furthermore, the step S1 is specifically as follows:

[0016] First, the historical operating data and video sample data of the fused magnesium furnace are collected, and the video sequence is represented in the multivariate time series as T = {(x i ,y i )|1≤i≤N}; where x i =[x 1,i ,x 2,i ,...,x M,i ], corresponding to the M-dimensional feature vector representation of the i-th frame image, y i Mark for working condition;

[0017] Then, the initial data samples are labeled in small quantities, and the classification categories are four categories: normal working condition, underburning working condition, exhaust working condition, and abnormal working condition; the initial data samples are divided into labeled data and unlabeled data, which are: x a and x b ; where x a is the labeled data, x b is unlabeled data.

[0018] Furthermore, the step S2 comprises the following steps:

[0019] S21. Use histogram equalization to count the pixel values ​​of the initial image sample to obtain the frequency of each gray level, that is, the number of times the gray level appears in the image; calculate its cumulative distribution function CDF(x), as follows:

[0020]

[0021] Where: h(i) is the number of pixels with gray level i, N is the total number of pixels in the image; x is the upper limit of the gray level;

[0022] S22, normalize the cumulative distribution function and map it to a new gray value range; the normalization formula is as follows:

[0023]

[0024] Among them, S(x) is the new gray level, CDF min is the minimum non-zero cumulative distribution function value, and L is the total number of gray levels;

[0025] S23, replacing the grayscale value of each pixel in the original image with a new grayscale value to generate a balanced image.

[0026] Furthermore, step S3 includes the following steps:

[0027] S31, using an image encoder of a feedforward neural network FNN to perform encoding operation on the image data, in which the global dependency in the feature map is captured by a global channel and a spatial attention mechanism GCA, so that each Transformer layer data layer performs independent processing at each position through an FNN; the operation of the image encoder is as follows:

[0028]

[0029] Among them, X is the original input image, LNorm represents the layer normalization operation, GCA is the global channel space attention mechanism, and FNN is the feedforward neural network;

[0030] S32, sending the feature map obtained in step S31 to a decoder; the decoder is composed of multiple fully connected layers and outputs a vector whose dimension matches the number of categories of the target classification task, and the formula is as follows:

[0031] X out =softmax(W c Xspatial+b c ) (7)

[0032] Among them, X out is the output feature vector, softmax is the activation function, W c and b c are weights and biases, and Xspatial is the feature map.

[0033] Furthermore, the step S31 of capturing the global dependency in the feature map through the global channel and spatial attention mechanism GCA includes the following steps:

[0034] (1) The equalized image is sent to the channel attention submodule. In the channel attention submodule, the dimension is first permuted from C×H×W to W×H×C. Then, the dependency between channels is captured by two layers of multi-layer perceptrons (MLP). The first layer of MLP reduces the number of channels to (1 / 4) times the original number, and then introduces nonlinearity through the ReLU activation function. The second layer of MLP restores the number of channels to the original dimension. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The equalized image and the channel attention map are multiplied element by element to obtain the enhanced feature map, as follows:

[0035] X channel =σ(MLP(Permute(X input )))e X input (4)

[0036] Where: X channel is the enhanced feature map, σ is the Sigmoid activation function, e is the element-by-element multiplication, X input is the original input image, MLP is a two-layer multilayer perceptron, and Permute is a permutation operation;

[0037] (2) Divide the enhanced feature map into 4 groups, each group contains C / 4 channels; perform a transposition operation on the grouped feature maps to disrupt the order of channels in each group; then, restore the disrupted feature map to its original shape C×H×W, as follows:

[0038] X shuffle =ChannelShuffle(X channel ) (5)

[0039] Among them, X shuffle is the shuffled feature map, ChannelShuffle is the shuffle operation;

[0040] (3) In the spatial attention submodule, the input feature map passes through a 7x7 convolution layer, and the number of channels is reduced to 1 / 4 of the original; it undergoes nonlinear transformation through batch normalization and ReLU activation function; then, the number of channels is restored to the original dimension C through a second 7x7 convolution layer, and then passes through a batch normalization layer; finally, the spatial attention map is generated through the Sigmoid activation function, and the shuffled feature map and the spatial attention map are multiplied element by element to obtain the final output feature map, as follows:

[0041] Xspatial=σ(Conv(BN(ReLU(Conv(Xshuffle)))))e Xshuffle (6)

[0042] Among them, Xspatial is the feature map after the spatial attention module, Xshuffle is the shuffled feature map, σ is the Sigmoid activation function, Conv is the convolution layer, BN is batch normalization, ReLU is the ReLU function, and e is element-by-element multiplication.

[0043] Furthermore, the step S4 comprises the following steps:

[0044] S41, the input contains feature maps T1 and T2 at two different times as follows:

[0045]

[0046] Among them, H and W are the spatial dimensions of the feature map, C is the number of channels, and R is a three-dimensional real matrix with a specific height, width, and number of channels;

[0047] S42. Convolution operations are performed on the input feature maps at two moments using three different sizes of convolution kernels: 3×3, 5×5, and 7×7 to capture features of different sizes, as follows:

[0048]

[0049] Among them, F ms (T1) and F ms (T2) is a multi-scale feature map, F 3x3 (T1)~F 7x7 (T1) and F 3x3 (T2)~F 7x7 (T2) is the feature map at time T1 and T2 after 3×3, 5×5, 7×7 convolution operations;

[0050] S43. For multi-scale feature maps, global average pooling and global maximum pooling are performed in the channel dimension to aggregate features in the entire space; meanwhile, average pooling and maximum pooling are performed in the spatial dimension and concatenated. The formula is as follows:

[0051]

[0052] in: and T 1 ,T 2 The average pooling and maximum pooling results of channel pooling, and is the average pooling and maximum pooling results of spatial pooling, Concat is the concatenation operation, and F ch (T1),F ch (T2),F sp (T1),F sp (T2) is the new feature vector result of feature concatenation;

[0053] S44. Perform attention calculation on two channels. The formula is as follows:

[0054]

[0055] Among them: Conv1D and Conv2D are one-dimensional convolution operations and two-dimensional convolution operations, Softmax is a normalization operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the attention calculation result;

[0056] S45, perform attention fusion and output, perform weighted multiplication on T1 and T2, add the results, and fuse the feature maps of the two moments. The formula is as follows:

[0057]

[0058] Among them: e is the multiplication operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the weighted attention, T1 ch_weighted ,T2 ch_weighted ,T1 sp_weighted ,T2 sp_weighted is the channel weighted and spatial weighted result, and Output is the final classification result output by the model.

[0059] Furthermore, the step S5 comprises the following steps:

[0060] S51, based on the initial model parameters obtained in step S4, initialize the global threshold τ according to the global average confidence t ; and calculate the global threshold τ by exponential moving average EMA t , the calculation formula is as follows:

[0061]

[0062] Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, and its value is between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data;

[0063] S52. Use the current model to predict the unlabeled data and calculate the confidence q of each sample. i , and according to the global threshold τ t and the local threshold τ c Generate qualified pseudo labels if the confidence q i Greater than the local threshold τ c , the prediction result is taken as a qualified pseudo label; the local threshold τ c The confidence of different categories is normalized and the formula is as follows:

[0064]

[0065] in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τt is the global threshold, τ min Minimum threshold to prevent the threshold from being too low.

[0066] Furthermore, step S6 includes the following steps:

[0067] S61, merging the qualified pseudo-label data set obtained in step S5 with the labeled data set to form an extended training set; using the extended training set to re-perform model training according to the CNN model training based on the dual-time fusion mechanism of time series described in step S4;

[0068] S62. During model training, calculate the training loss; for labeled data, use the traditional cross entropy loss L s , for unlabeled data, use pseudo-label and unlabeled loss function L u , the formula is as follows:

[0069]

[0070] Where: L u is the loss function, B is the batch size of unlabeled data, q i is the pseudo label confidence, τ c is the local threshold, is a pseudo label, is the cross entropy loss function;

[0071] The total loss function of the model is expressed as follows:

[0072] L=L s +w u L u (16)

[0073] Where: L is the total loss function, L s is the cross entropy loss of labeled data, w u is the unlabeled loss weight, L u is the no-label loss;

[0074] S63, repeat steps S5 and S6, iteratively update the global threshold τ t and the local threshold τ c , until the calculation result of the total loss function meets the preset requirements, the iteration ends and the model training is completed.

[0075] The present invention also provides a semi-supervised learning and free threshold electric magnesium furnace working condition identification system, the system adopts the above method when executing, and includes the following modules:

[0076] The data acquisition module is used to collect the historical working conditions and video image sample data of the fused magnesium furnace, and divide it into labeled data x aand unlabeled data x b ;

[0077] A data preprocessing module, used to preprocess the initial image sample data obtained by the data acquisition module;

[0078] A feature extraction module is used to perform feature enhancement extraction on the preprocessed image data using a ViT feature extraction network based on a global channel spatial attention module;

[0079] Model training module, used to extract features from labeled images X a_out Conduct CNN model training based on dual-time fusion mechanism of time series;

[0080] The pseudo-label annotation module is used to label the unlabeled image data X preprocessed by the data preprocessing module. b_out The trained model is fed into the model to predict the category probability, and a free threshold mechanism is used to generate a qualified pseudo-label dataset.

[0081] The model optimization module is used to update the training data set, retrain the model according to the model training module, and continuously iterate until the model performance meets the preset requirements by calculating the training loss;

[0082] The model verification module is used to output the working condition discrimination results of the final model using the verification data set to verify the model performance.

[0083] Furthermore, the pseudo-label marking module includes the following units:

[0084] The global threshold calculation unit is used to initialize the global threshold τ based on the initial model parameters obtained by the model training module and the global average confidence. t , and calculate the global threshold τ by exponential moving average EMA t , the calculation formula is as follows:

[0085]

[0086] Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, and its value is between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data;

[0087] The local threshold calculation unit is used to use the current model to predict unlabeled data and calculate the confidence q of each sample. i , and according to the global threshold τ t and the local threshold τ c Generate qualified pseudo labels if the confidence qi Greater than the local threshold τ c , the prediction result is taken as a qualified pseudo label; the local threshold τ c The confidence of different categories is normalized and the formula is as follows:

[0088]

[0089] in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τ t is the global threshold, τ min Minimum threshold to prevent the threshold from being too low.

[0090] The advantages of the present invention are:

[0091] (1) A global channel-spatial attention module is used to extract image features. In traditional convolutional neural networks, the relationship between channels may be ignored, but these relationships are very important for capturing global information. If the dependencies between channels are not considered, the model may not be able to fully utilize all the information in the feature map, resulting in insufficient capture of global features. This module combines channel attention, channel shuffling, and spatial attention mechanisms to capture global dependencies in feature maps.

[0092] (2) A model based on a time series dual-time fusion module is used. In computer vision tasks such as time series feature analysis and change detection, the performance of the model depends on its ability to effectively fuse features from different moments. However, traditional simple addition or multiplication feature fusion methods have significant limitations in dealing with these tasks. These methods are easily affected by noise, have difficulty dynamically adjusting the importance of features, and cannot fully utilize time series information. The multi-scale dual-time fusion module combines multi-scale feature extraction with channel and spatial attention mechanisms to improve the effect and robustness of feature fusion. This module captures feature information of different scales through multi-scale convolution, and dynamically adjusts the importance of features through the attention mechanism to ensure that key features receive higher weights during the fusion process.

[0093] (3) In the semi-supervised learning strategy, a free threshold qualified pseudo-label generation method is adopted. When the data annotation cost is high and the labeled data is very scarce, this mechanism can maximize the use of unlabeled data and improve model performance. By dynamically adjusting the global threshold and local threshold, the threshold of pseudo-label generation is adaptively adjusted as the learning state of the model changes, thereby ensuring appropriate pseudo-label allocation for different categories of unlabeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1It is a flow chart of a method for identifying working conditions of a fused magnesium furnace using semi-supervised learning and free threshold according to an embodiment of the present invention;

[0095] Figure 2 A schematic diagram of a global channel spatial attention mechanism flow chart of an embodiment of the present invention;

[0096] Figure 3 The figure is a schematic diagram of the dual-time fusion mechanism flow based on time sequence according to an embodiment of the present invention. DETAILED DESCRIPTION

[0097] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0098] Example 1

[0099] This embodiment provides a method for identifying the working conditions of a molten magnesium furnace using semi-supervised learning and free thresholds. Figure 1 As shown, the specific steps are as follows:

[0100] S1. Collect the historical working conditions and video image sample data of the fused magnesium furnace. The video sequence is represented in the multivariate time series as T = {(x i ,y i )|1≤i≤N}; where x i =[x 1,i ,x 2,i ,...,x M,i ], corresponding to the M-dimensional feature vector representation of the i-th frame image, y i Mark for working condition.

[0101] The initial data samples are labeled in small quantities, and the classification categories are divided into four categories: normal working condition, underburning working condition, exhaust working condition, and abnormal working condition. The initial data samples are divided into labeled data and unlabeled data, which are: x a and x b ; where x a is the labeled data, x b is unlabeled data.

[0102] S2: Preprocess the initial image samples and extract the image features. The specific process is as follows:

[0103] S21. Preprocess the initial image samples: Histogram equalization is used to count the pixel values ​​of the image to obtain the frequency of each gray level, that is, the number of times the gray level appears in the image. For example, for an 8-bit image, the gray value ranges from 0 to 255, so the histogram has 256 entries, each of which represents the number of pixels at that gray level. The cumulative distribution function CDF(x) is calculated as follows: Formula (1), which is used to represent the proportion of pixels in the image that are less than or equal to a certain gray value. The CDF can be obtained by accumulating the gray level frequencies item by item.

[0104]

[0105] Where: h(i) is the number of pixels with gray level i, N is the total number of pixels in the image, and x is the upper limit of the gray level.

[0106] S22, normalize the cumulative distribution function and map it to a new grayscale value range. For example, for an 8-bit image, the new grayscale range is 0 to 255. The normalization formula (2) is:

[0107]

[0108] Where: S(x) is the new gray level, CDF min is the minimum non-zero CDF value, and L is the total number of gray levels (for 8-bit images, L is 256).

[0109] S23, replace the grayscale value of each pixel in the original image with a new grayscale value to generate a balanced image. This process is actually to redistribute the grayscale values ​​of the pixels so that the pixels originally concentrated in a certain grayscale range are evenly distributed in the entire grayscale range.

[0110] S3. Use the ViT feature extraction network based on the global channel spatial attention module to perform feature enhancement extraction on the preprocessed image data. The specific steps are as follows:

[0111] S31, using an image encoder of a feedforward neural network FNN to perform encoding operation on the image data, in which the global dependency in the feature map is captured by a global channel and a spatial attention mechanism GCA, so that each Transformer layer data layer performs independent processing at each position through an FNN; the operation of the image encoder is as follows (3):

[0112]

[0113] Among them: X is the original input graph, LNorm represents the layer normalization operation, GCA is the global channel attention mechanism, and FNN is the feedforward neural network.

[0114] More specifically, if Figure 2 As shown in Figure 2, the specific steps of capturing the global dependencies in the feature map through the global channel and spatial attention mechanism GCA are as follows:

[0115] First, the equalized image is sent to the channel attention submodule. In the channel attention submodule, the input feature map is first dimensionally permuted from C×H×W to W×H×C. Then, the dependencies between channels are captured by two layers of multi-layer perceptrons (MLP). The first layer of MLP reduces the number of channels to (1 / 4) times the original, then introduces nonlinearity through the ReLU activation function, and then restores the number of channels to the original dimension through the second layer of MLP. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The equalized image and the channel attention map are multiplied element by element to obtain the enhanced feature map, which is as follows:

[0116] X channel =σ(MLP(Permute(X input )))e X input (4)

[0117] Where: X channel is the enhanced feature map, σ is the Sigmoid activation function, e is the element-by-element multiplication, X input is the original input image, MLP is a two-layer multi-layer perceptron, and Permute is a permutation operation.

[0118] Then, the channels are shuffled. The enhanced feature maps are divided into 4 groups, each containing C / 4 channels. The grouped feature maps are transposed to disrupt the channel order in each group. Subsequently, the shuffled feature maps are restored to the original shape C×H×W. This method can better mix feature information and enhance feature expression capabilities, as shown in the following formula (5):

[0119] X shuffle =ChannelShuffle(X channel ) (5)

[0120] Where: X shuffle It is the feature map after shuffling, and ChannelShuffle is the shuffling operation.

[0121] Finally, in the spatial attention submodule, the input feature map passes through a 7x7 convolution layer, and the number of channels is reduced to 1 / 4 of the original. It is then subjected to nonlinear transformation through batch normalization and ReLU activation function. Then, the number of channels is restored to the original dimension C through the second 7x7 convolution layer, and then passed through the batch normalization layer. Finally, the spatial attention map is generated through the Sigmoid activation function. The shuffled feature map and the spatial attention map are multiplied element by element to obtain the final output feature map, which is as follows:

[0122] Xspatial=σ(Conv(BN(ReLU(Conv(Xshuffle)))))e Xshuffle (6)

[0123] Among them: Xspatial is the feature map after spatial attention, Xshuffle is the shuffled feature map, σ is the Sigmoid activation function, Conv is the convolution layer, BN is batch normalization, ReLU is the ReLU function, and e is element-by-element multiplication.

[0124] S32, the feature map is sent to the decoder (splitter), the decoder is composed of multiple fully connected layers, and outputs a vector whose dimension matches the number of categories of the target classification task, as shown in the following formula (7):

[0125] X out =softmax(W c Xspatial+b c ) (7)

[0126] Where: X out is the output feature vector, softmax is the activation function, W c and b c are weights and biases, and Xspatial is the feature map.

[0127] S4. Extract feature images X with labels a out Conduct CNN model training based on dual-time fusion mechanism of time series and perform preliminary training of the model.

[0128] First, the CNN convolutional neural network is used as the basic classification network. It extracts and processes features through convolutional layers, pooling layers, and fully connected layers, and uses feature image information at two different times based on temporal information to perform attention fusion at the same time to make full use of temporal information for model training. Figure 3 As shown, the specific steps are as follows:

[0129] The input contains feature maps T1 and T2 at two different times as follows (8):

[0130]

[0131] Among them: H and W are the spatial dimensions of the feature map, C is the number of channels, and R is a three-dimensional real matrix with a specific height, width, and number of channels.

[0132] S42, for the input feature map (T 1 ,T 2 ) uses three different sizes of convolution kernels (3×3, 5×5, 7×7) for convolution operations to capture features of different sizes, as follows (9):

[0133]

[0134] Among them: F ms (T1) and F ms (T2) is a multi-scale feature map, F 3x3 (T1)~F 7x7 (T1) and F 3x3 (T2)~F 7x7 (T2) is the feature map at time T1 and T2 after the feature Figure 3 ×3, 5×5, 7×7 convolution operations.

[0135] S43. For the multi-scale feature map, global average pooling and global maximum pooling are performed in the channel dimension. The result of the pooling operation aggregates the features in the entire space, and the average pooling and maximum pooling in the spatial dimension are performed and concatenated, and the following formula (10) is obtained:

[0136]

[0137] in: and T 1 ,T 2 The average pooling and maximum pooling results of channel pooling, and is the average pooling and maximum pooling results of spatial pooling, Concat is the concatenation operation, and F ch (T1),F ch (T2),F sp (T1),F sp (T2) is the new feature vector result of feature concatenation.

[0138] S44. Perform attention calculation on two channels, as shown in formula (11):

[0139]

[0140] Among them: Conv1D and Conv2D are one-dimensional convolution operations and two-dimensional convolution operations, Softmax is a normalization operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the attention calculation result.

[0141] S45, perform attention fusion and output, perform weighted multiplication on T1 and T2, add the results, and fuse the feature maps of the two moments, as shown in the following formula (12):

[0142]

[0143] Among them: e is the multiplication operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the weighted attention, T1 ch_weighted ,T2 ch_weighted ,T1 sp_weighted ,T2 sp_weighted is the result of channel weighting and space weighting, and Output is the final classification result of the model output, that is, the working conditions of the electric magnesium furnace: normal working condition, under-burning working condition, exhaust working condition, and abnormal working condition.

[0144] S5, the unlabeled image data X preprocessed in step S2 b_out The pre-processed unlabeled image data X b_out Input into the model for prediction, get the category probability of each picture, and set a threshold to select the prediction results with higher confidence. Only when the model's prediction probability for a sample is higher than this threshold, it is considered a qualified pseudo-label. Finally, for all samples with probabilities higher than the threshold, assign corresponding category labels to form a qualified pseudo-label dataset X. B This embodiment adopts a free threshold qualified pseudo label generation method, and the specific steps are as follows:

[0145] S51, based on the initial model parameters obtained in step S4, initialize the global threshold τ according to the global average confidence t , the global threshold is a unified threshold for all categories, reflecting the overall learning state of the model. The global threshold τ t As the training progresses, the gradual changes are calculated by exponential moving average (EMA), which can smoothly capture the changes in the confidence of the model on unlabeled data. The calculation formula is as follows (13):

[0146]

[0147] Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, usually between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data.

[0148] S52. Use the current model to predict the unlabeled data and calculate the confidence q of each sample. i , and generate qualified pseudo labels based on the global threshold and local threshold. If the confidence q i Greater than the local threshold τ c , the prediction result is regarded as a qualified pseudo-label. The local threshold normalizes the confidence of different categories to ensure that complex categories have appropriate pseudo-label generation thresholds, avoiding the generation of a large number of pseudo-labels for simple categories and insufficient utilization of complex category data. The local threshold has the following formula (14):

[0149]

[0150] in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τ t is the global threshold, τ min Minimum threshold to prevent the threshold from being too low.

[0151] S6, update the training data set, retrain the model according to the method described in step S4, and continuously iterate the training by calculating the training loss until the model performance meets the preset requirements. The specific steps are as follows:

[0152] S61, merging the qualified pseudo-label data set obtained in step S5 with the labeled data set to form an extended training set; using the extended training set to re-perform model training according to the CNN model training based on the dual-time fusion mechanism of time series described in step S4;

[0153] S62. During model training, calculate the training loss; for labeled data, use the traditional cross entropy loss L s , for unlabeled data, use pseudo-label and unlabeled loss function L u , specifically the following formula (15):

[0154]

[0155] Where: L u is the loss function, B is the batch size of unlabeled data, q i is the pseudo label confidence, τc is the local threshold, is a pseudo label, is the cross entropy loss function.

[0156] The total loss function of the model can be expressed as formula (16):

[0157] L=L s +w u L u (16)

[0158] Where: L is the total loss function, L s is the cross entropy loss of labeled data, w u is the unlabeled loss weight, L u is the unlabeled loss.

[0159] S63, repeat steps S5 and S6, iteratively update the global threshold τ t and the local threshold τ c , until the calculation result of the total loss function meets the preset requirements, the iteration ends and the model training is completed.

[0160] S7. The obtained model uses a verification data set to output its working condition discrimination result, so as to judge whether the electric magnesium furnace is in an abnormal state at the current moment according to the discrimination result. If it is abnormal, an abnormal working condition alarm should be issued at this time.

[0161] Example 2

[0162] It should be further explained that, based on the same inventive concept, this embodiment also provides a semi-supervised learning and free threshold electric magnesium furnace working condition identification system, and the system executes the method described in embodiment 1 when running, including the following modules:

[0163] The data acquisition module is used to collect the historical working conditions and video image sample data of the fused magnesium furnace, and divide it into labeled data x a and unlabeled data x b ;

[0164] A data preprocessing module, used to preprocess the initial image sample data obtained by the data acquisition module;

[0165] A feature extraction module is used to perform feature enhancement extraction on the preprocessed image data using a ViT feature extraction network based on a global channel spatial attention module;

[0166] Model training module, used to extract features from labeled images X a_out Conduct CNN model training based on dual-time fusion mechanism of time series;

[0167] The pseudo-label annotation module is used to label the unlabeled image data X preprocessed by the data preprocessing module. b_out The trained model is fed into the model to predict the category probability, and a free threshold mechanism is used to generate a qualified pseudo-label dataset.

[0168] The model optimization module is used to update the training data set, retrain the model according to the model training module, and continuously iterate until the model performance meets the preset requirements by calculating the training loss;

[0169] The model verification module is used to output the working condition discrimination results of the final model using the verification data set to verify the model performance.

[0170] The pseudo-label marking module includes the following units:

[0171] The global threshold calculation unit is used to initialize the global threshold τ based on the initial model parameters obtained by the model training module and the global average confidence. t , the global threshold τ t As the training progresses, it gradually changes and is calculated by exponential moving average EMA to smoothly capture the changes in the model's confidence on unlabeled data. The calculation formula is as follows:

[0172]

[0173] Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, and its value is between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data;

[0174] The local threshold calculation unit is used to use the current model to predict unlabeled data and calculate the confidence q of each sample. i , and according to the global threshold τ t and the local threshold τ c Generate qualified pseudo labels if the confidence q i Greater than the local threshold τ c , the prediction result is taken as a qualified pseudo label. The local threshold τ c The confidence of different categories is normalized and the formula is as follows:

[0175]

[0176] in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τ t is the global threshold, τmin Minimum threshold to prevent the threshold from being too low.

[0177] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying the working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold, characterized in that: The following steps are involved: S1. Collect historical working conditions and video image sample data of the fused magnesium furnace and divide them into labeled data x a and unlabeled data x b ; S2, preprocessing the initial image sample data obtained in step S1; S3, using the ViT feature extraction network based on the global channel spatial attention module to perform feature enhancement extraction on the preprocessed image data; S4. Extract feature images X with labels a_out Conduct CNN model training based on dual-time fusion mechanism of time series; S5, the unlabeled image data X preprocessed in step S2 b_out The trained model is fed into the model to predict the category probability, and a free threshold mechanism is used to generate a qualified pseudo-label dataset. S6, update the training data set, re-train the model according to the method described in step S4, and continuously iterate the training by calculating the training loss until the model performance meets the preset requirements; S7. The final model is used to output the working condition discrimination result using the verification data set to verify the model performance.

2. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 1 is characterized in that: The step S1 is specifically as follows: First, the historical operating data and video sample data of the fused magnesium furnace are collected, and the video sequence is represented in the multivariate time series as T = {(x i ,y i )|1≤i≤N}; Among them, x i =[x 1,i ,x 2,i ,...,x M,i ], corresponding to the M-dimensional feature vector representation of the i-th frame image, y i Mark for working condition; Then, the initial data samples are labeled in small quantities, and the classification categories are four categories: normal working condition, underburning working condition, exhaust working condition, and abnormal working condition; the initial data samples are divided into labeled data and unlabeled data, which are: x a and x b ; where x a is the labeled data, x b is unlabeled data.

3. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 1 is characterized in that: The step S2 comprises the following steps: S21. Use histogram equalization to count the pixel values ​​of the initial image sample to obtain the frequency of each gray level, that is, the number of times the gray level appears in the image; calculate its cumulative distribution function CDF(x), as follows: Where: h(i) is the number of pixels with gray level i, N is the total number of pixels in the image; x is the upper limit of the gray level; S22, normalize the cumulative distribution function and map it to a new gray value range; the normalization formula is as follows: Among them, S(x) is the new gray level, CDF min is the minimum non-zero cumulative distribution function value, L is the total number of gray levels; S23, replacing the grayscale value of each pixel in the original image with a new grayscale value to generate a balanced image.

4. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 3 is characterized in that: The step S3 comprises the following steps: S31, using an image encoder of a feedforward neural network FNN to perform encoding operation on the image data, in which the global dependency in the feature map is captured by a global channel and a spatial attention mechanism GCA, so that each Transformer layer data layer performs independent processing at each position through an FNN; the operation of the image encoder is as follows: Among them, X is the original input image, LNorm represents the layer normalization operation, GCA is the global channel space attention mechanism, and FNN is the feedforward neural network; S32, sending the feature map obtained in step S31 to a decoder; the decoder is composed of multiple fully connected layers and outputs a vector whose dimension matches the number of categories of the target classification task, and the formula is as follows: X out =softmax(W c Xspatial+b c ) (7) Among them, X out is the output feature vector, softmax is the activation function, W c and b c are weights and biases, and Xspatial is the feature map.

5. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 4 is characterized in that: The step S31 of capturing the global dependency in the feature map through the global channel and spatial attention mechanism GCA includes the following steps: (1) The equalized image is sent to the channel attention submodule. In the channel attention submodule, the dimension is first permuted from C×H×W to W×H×C. Then, the dependency between channels is captured by two layers of multi-layer perceptrons (MLP). The first layer of MLP reduces the number of channels to (1 / 4) times the original number, and then introduces nonlinearity through the ReLU activation function. The second layer of MLP restores the number of channels to the original dimension. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The equalized image and the channel attention map are multiplied element by element to obtain the enhanced feature map, as follows: X channel =σ(MLP(Permute(X input )))e X input (4) Where: X channel is the enhanced feature map, σ is the Sigmoid activation function, e is the element-by-element multiplication, X input is the original input image, MLP is a two-layer multilayer perceptron, and Permute is a permutation operation; (2) Divide the enhanced feature map into 4 groups, each group contains C / 4 channels; perform a transposition operation on the grouped feature maps to disrupt the order of channels in each group; then, restore the disrupted feature map to its original shape C×H×W, as follows: X shuffle =ChannelShuffle(X channel ) (5) Among them, X shuffle is the shuffled feature map, ChannelShuffle is the shuffle operation; (3) In the spatial attention submodule, the input feature map passes through a 7x7 convolution layer, and the number of channels is reduced to 1 / 4 of the original; it undergoes batch normalization and ReLU activation function for nonlinear transformation; then, it passes through a second 7x7 convolution layer to restore the number of channels to the original dimension C, and then passes through a batch normalization layer; finally, the spatial attention map is generated through the Sigmoid activation function, and the shuffled feature map and the spatial attention map are multiplied element by element to obtain the final output feature map, as follows: Xspatial=σ(Conv(BN(ReLU(Conv(Xshuffle)))))e Xshuffle (6) Among them, Xspatial is the feature map after the spatial attention module, Xshuffle is the shuffled feature map, σ is the Sigmoid activation function, Conv is the convolution layer, BN is batch normalization, ReLU is the ReLU function, and e is element-by-element multiplication.

6. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 5 is characterized in that: The step S4 comprises the following steps: S41, the input contains feature maps T1 and T2 at two different times as follows: Among them, H and W are the spatial dimensions of the feature map, C is the number of channels, and R is a three-dimensional real matrix with a specific height, width, and number of channels; S42. Convolution operations are performed on the input feature maps at two moments using three different sizes of convolution kernels: 3×3, 5×5, and 7×7 to capture features of different sizes, as follows: Among them, F ms (T1) and F ms (T2) is a multi-scale feature map, F 3x3 (T1)~F 7x7 (T1) and F 3x3 (T2)~F 7x7 (T2) is the feature map at time T1 and T2 after 3×3, 5×5, 7×7 convolution operations; S43. For multi-scale feature maps, global average pooling and global maximum pooling are performed in the channel dimension to aggregate features in the entire space; meanwhile, average pooling and maximum pooling are performed in the spatial dimension and concatenated. The formula is as follows: in: and is the average pooling and maximum pooling results of channel pooling of T1 and T2. and is the average pooling and maximum pooling results of spatial pooling, Concat is the concatenation operation, and F ch (T1),F ch (T2),F sp (T1),F sp (T2) is the new feature vector result of feature concatenation; S44. Perform attention calculation on two channels. The formula is as follows: Among them: Conv1D and Conv2D are one-dimensional convolution operations and two-dimensional convolution operations, Softmax is a normalization operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the attention calculation result; S45, perform attention fusion and output, perform weighted multiplication on T1 and T2, add the results, and fuse the feature maps of the two moments. The formula is as follows: Among them: e is the multiplication operation, W ch (T1), W ch (T2), W sp (T1), W sp (T2) is the weighted attention, T1 ch_weighted ,T2 ch_weighted ,T1 sp_weighted ,T2 sp_weighted is the channel weighted and spatial weighted result, and Output is the final classification result output by the model.

7. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 6 is characterized in that: The step S5 comprises the following steps: S51, based on the initial model parameters obtained in step S4, initialize the global threshold τ according to the global average confidence t ; and calculate the global threshold τ by exponential moving average EMA t , the calculation formula is as follows: Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, and its value is between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data; S52. Use the current model to predict the unlabeled data and calculate the confidence q of each sample. i , and according to the global threshold τ t and the local threshold τ c Generate qualified pseudo labels if the confidence q i Greater than the local threshold τ c , the prediction result is taken as a qualified pseudo label; the local threshold τ c The confidence of different categories is normalized and the formula is as follows: in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τ t is the global threshold, τ min Minimum threshold to prevent the threshold from being too low.

8. The method for identifying working conditions of a fused magnesium furnace based on semi-supervised learning and free threshold according to claim 7 is characterized in that: The step S6 comprises the following steps: S61, merging the qualified pseudo-label data set obtained in step S5 with the labeled data set to form an extended training set; using the extended training set to re-perform model training according to the CNN model training based on the dual-time fusion mechanism of time series described in step S4; S62. During model training, calculate the training loss; for labeled data, use the traditional cross entropy loss L s , for unlabeled data, use pseudo-label and unlabeled loss function L u , the formula is as follows: Where: L u is the loss function, B is the batch size of unlabeled data, q i is the pseudo label confidence, τ c is the local threshold, is a pseudo label, is the cross entropy loss function; The total loss function of the model is expressed as follows: L=L s +w u L u (16) Where: L is the total loss function, L s is the cross entropy loss of labeled data, w u is the unlabeled loss weight, L u is the no-label loss; S63, repeat steps S5 and S6, iteratively update the global threshold τ t and the local threshold τ c , until the calculation result of the total loss function meets the preset requirements, the iteration ends and the model training is completed.

9. Semi-supervised learning and free threshold electric magnesium furnace working condition recognition system, characterized by: The system, when executed, adopts the method described in any one of claims 1 to 8, and includes the following modules: The data acquisition module is used to collect the historical working conditions and video image sample data of the fused magnesium furnace, and divide it into labeled data x a and unlabeled data x b ; A data preprocessing module, used to preprocess the initial image sample data obtained by the data acquisition module; A feature extraction module is used to perform feature enhancement extraction on the preprocessed image data using a ViT feature extraction network based on a global channel spatial attention module; Model training module, used to extract features from labeled images X a_out Conduct CNN model training based on dual-time fusion mechanism of time series; The pseudo-label annotation module is used to label the unlabeled image data X preprocessed by the data preprocessing module. b_out The trained model is fed into the model to predict the category probability, and a free threshold mechanism is used to generate a qualified pseudo-label dataset. The model optimization module is used to update the training data set, retrain the model according to the model training module, and continuously iterate until the model performance meets the preset requirements by calculating the training loss; The model verification module is used to output the working condition discrimination results of the final model using the verification data set to verify the model performance.

10. The semi-supervised learning and free threshold electric magnesium furnace working condition identification system according to claim 9 is characterized in that: The pseudo-label marking module includes the following units: The global threshold calculation unit is used to initialize the global threshold τ based on the initial model parameters obtained by the model training module and the global average confidence. t , and calculate the global threshold τ by exponential moving average EMA t , the calculation formula is as follows: Where: τ t is the global threshold of the tth round of training, C is the number of categories, λ is the momentum coefficient used to smooth the change of the threshold, and its value is between (0,1), q i is the prediction confidence of the i-th unlabeled sample, and B is the batch size of unlabeled data; The local threshold calculation unit is used to use the current model to predict unlabeled data and calculate the confidence q of each sample. i , and according to the global threshold τ t and the local threshold τ c Generate qualified pseudo labels if the confidence q i Greater than the local threshold τ c , the prediction result is taken as a qualified pseudo label; the local threshold τ c The confidence of different categories is normalized and the formula is as follows: in: is the average prediction confidence of category c, is the highest confidence among all categories, used for normalization, τ t is the global threshold, τ min Minimum threshold to prevent the threshold from being too low.

Citation Information

Cited By

  • Bearing life prediction method based on combination of pooling-multi-scale convolutional neural network and Transform

    CN119598287A

  • Method for detecting anomaly of electrical smelting furnace for magnesia based on knowledge distillation of calibration teacher model

    CN120997587A

  • An anomaly detection method for fused magnesium furnaces based on knowledge distillation using a calibrated teacher model

    CN120997587B

  • RCD multi-field coupling simulation method for marine soft soil with pebbles

    CN121072401A

  • Intelligent operation monitoring method and device of PEM fuel cell

    CN121097136A