A tundish bright spot detection method based on target detection
By combining lightweight convolutional neural networks and long short-term memory networks, the problems of delay and accuracy in bright spot detection during steelmaking were solved, enabling accurate tracking and prediction of bright spot ranges, thereby improving steelmaking quality and output.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2026-03-27
AI Technical Summary
Existing steelmaking processes suffer from problems such as high latency, low accuracy, and lack of transferability, which cannot meet the real-time quality control requirements of complex environments in steel plants.
We employ a combination of lightweight convolutional neural networks and long short-term memory networks. Bright areas are extracted using residual modules and self-attention modules. By combining traditional edge detection operators and brightness thresholds, we utilize LSTM networks to predict bright changes, achieving self-deep learning and semi-automatic annotation.
It enables precise tracking and change prediction of bright areas in a short time, improves the automation efficiency of steelmaking quality and output, and allows for timely addition of covering agents.
Smart Images

Figure CN116486301B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video image processing technology, and in particular to a method for detecting bright spots in intermediate packets based on target detection. Background Technology
[0002] Based on steelmaking videos captured by cameras in steelmaking ladle systems, a bright spot detection method enhanced by image processing technology is developed. This method is semi-automatic, self-annotating, and capable of self-learning. Currently, the technological requirements for steelmaking processes and steel yield are increasingly sophisticated, and the level of bright spots largely reflects the real-time quality of steelmaking. To date, many researchers have attempted to analyze traditional steelmaking processes, but research using emerging information technologies, such as CNN-based image processing, is scarce. A relevant work, such as "Tao H and Lu X, Smoke vehicle detection based on spatiotemporal bag-of-features and professional convolutional neural network[J].IEEE Transactions on Circuits and Systems for Video Technology,2020,30(10):3301-3316," uses a large model structure with high latency, making it unsuitable for steel plant environments with stringent latency requirements. Traditional object detection methods can be based on "Huang Zi. Research on Object Detection Model Based on Convolutional Neural Network [D]. Shanghai: Control Science and Engineering, Shanghai Jiao Tong University, 2014." This paper proposes an efficient spatial attention model, ESA-Net, to achieve bright-eye detection by focusing on the region of interest and ignoring invalid and redundant information. However, this bright-eye detection method lacks adaptability to the high-brightness, low-coverage bright-eye conditions in steelmaking. Furthermore, traditional simple edge detection algorithms struggle to cope with the complex conditions of steel mills. Single feature extraction cannot meet the plant's requirements, and steel mill cameras often change, rendering previously learned features unusable. The complex and ever-changing environment results in low bright-eye detection accuracy for such algorithms, failing to meet actual production requirements. In reality, research in this area, both domestically and internationally, is still relatively limited. A more credible study is Chatterjee S and Chattopadhyay K's article, "Physical modelling of slag eye in an inert gasshrouded un tundish using dimensional analysis," published in Metallurgical and Materials Transactions B. This article is based on a local steel plant in Canada. However, there are many differences between the two countries in terms of raw materials, processes, and equipment, and the model lacks certain transferability and cannot be applied to the current state of steelmaking in my country. Summary of the Invention
[0003] This invention proposes a lightweight, object-detection-based method for detecting bright spots in intermediate packages. This method, implemented using convolutional neural networks and long short-term memory neural networks, aims to address the problem of timely detection of bright spots during steelmaking production, thereby improving the automation efficiency of the production process. It utilizes traditional image processing for semi-automated annotation, enabling self-deep learning for target recognition.
[0004] The technical solution of this invention is as follows: A method for detecting bright eyes in intermediate packages based on target detection, the specific steps of which are as follows:
[0005] A method for detecting bright eyes in intermediate packages based on object detection, the specific steps of which are as follows:
[0006] Step 1: Using a bright video surveillance camera as the source of video data, a CNN network including a residual module and a self-attention module is used to perform preliminary extraction on the video data to obtain candidate bright areas;
[0007] Step 2: Use traditional edge detection operators to extract bright edges from candidate bright regions; further extract the bright range image by setting a bright brightness threshold, calculate the area of the bright range image, and use the area and brightness of the bright range image as features to predict bright changes.
[0008] Step 3: Use an LSTM network to process the extracted bright area image and predict its changing trend;
[0009] The object detection-based intermediate packet bright detection method is implemented based on an object detection-based intermediate packet bright detection network, which consists of four layers: a CNN network layer, late pooling, two temporal convolutional layers, and three LSTM network layers.
[0010] The CNN network layer is divided into two parts. The first part includes a 3D convolutional layer (conv1), a ReLU layer, and a max pooling layer. The second part includes an Inception Block, a ReLU layer, and a max pooling layer.
[0011] late_pooling takes the output of a CNN network layer, inputs it into a fully connected layer, and then performs pooling again.
[0012] The temporal convolutional layer consists of 128 3×3 temporal convolutional kernels;
[0013] When the LSTM network layer is trained, the training output value of each frame is compared with the next frame of the video. After the gradient is backpropagated, the correct next frame is added to the training set.
[0014] After passing through the intermediate packet bright spot detection network based on object detection, the bright spot area image is processed as follows;
[0015] X t =[X t,1 ,…,X t,i (1)
[0016]
[0017] out = W i ·X t +W h ·tan(c t )+W c ·c t (3)
[0018] Equation (1) is the feature vector obtained by the convolutional processing of the image by the CNN network in step one, denoted by X. t,i Let W represent the column vector of the image at time t; Equation (2) is the initial and iterative calculation method of the cell state and hidden state in LSTM; Equation (3) is the output result, where W t It is the neuron-to-neuron weight matrix, W h W is the weight matrix of the memory gates, tan(·) is the nonlinear part in LSTM, and W c These are the weight matrices between cell states. These three matrices are the matrices that LSTM needs to perform gradient backpropagation iterations on.
[0019] In the CNN network layers, the size of the two-dimensional convolutional kernel in the convolutional layer is 3×3. The feature extraction convolutional kernel in the spatial dimension is also 3×3, the same as that in the two-dimensional convolutional network. The feature extraction convolutional kernel in the temporal dimension is 5×5.
[0020] Finally, an assessment of the coverage of the bright spot detection was conducted.
[0021]
[0022]
[0023]
[0024] Equation (6) is the accuracy index obtained by summing the areas of the ROC (Receiver Operating Characteristic) image composed of equations (4) and (5); TP (number of correct samples), FN (number of incorrect samples), FP (number of incorrect samples), TN (number of incorrect samples), m (number of samples), y (TPR value), x (FPR value).
[0025] Traditional edge detection operators, such as the Canny operator, extract edges. Edge extraction is a relatively mature image processing technique; however, due to the complexity of the dimensional environment, there is a high probability of interference. By adding a brightness threshold, the influence of low brightness on edge detection can be effectively reduced.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention aims to solve the problem of bright spots affecting the quality of steelmaking during the steelmaking process. Through the method of the present invention, the range of bright spots can be tracked more accurately in a short time and the subsequent changes in bright spots can be predicted, so that steel mills can add covering agents in a timely and accurate manner, thereby improving the quality and output of steelmaking. Attached Figure Description
[0027] Figure 1 This is a flowchart of a method for detecting bright spots in intermediate packages based on target detection;
[0028] Figure 2 This is a schematic diagram of the CNN network model and LSTM network model in this invention. Detailed Implementation
[0029] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.
[0030] The specific experimental dataset for this embodiment is frontline data from a branch of a key steelmaking plant in a certain city. The fixed camera orientation changes daily depending on the workers. This approach selects the best-positioned angles and angles, and preprocesses 50 videos. The test set consists of 3 videos, with a total duration of approximately 10 hours.
[0031] This implementation example Figure 1 As shown, the specific steps are as follows:
[0032] Step 1, Brilliant CNN Data Preprocessing:
[0033] Based on the collected steelmaking videos, firstly through... Figure 2 The CNN network shown extracts the specific locations of prominent bright spots in the video, tracks these bright spots through object detection, and continuously updates their size and location. This method captures the changing process of bright spots, creating a bright spot video dataset. The non-black smoke vehicle video dataset is created in the same way.
[0034] Step 2, the impressive LSTM variation model:
[0035] When the impressive output of CNN is put into such Figure 2Before implementing the LSTM model shown, the output needs some processing. First, the output is fed into a fully connected layer and then pooled to enhance the inter-frame features. Next, to further enhance the features at the temporal scale, a temporal convolutional layer consisting of 128 3×3 convolutional kernels can be added to capture local differences between frames through a smaller temporal window.
[0036] Next, the output can be fed into the LSTM model. Note that supervised learning can be performed during each training session to avoid gradient explosion after multiple training sessions.
[0037] Step 3, on-site steelmaking testing:
[0038] To verify the real-time performance of the model in this invention, we tested it at a steelmaking site. The conditions at the steelmaking site differed significantly from our expectations, and there were issues with the transmission of camera video files, leading to inherent latency in the network. Furthermore, considering the fleeting nature of critical moments in steelmaking, this placed even higher demands on our network. Therefore, the network fell short of our initial expectations. It was relatively successful in detecting bright targets, accurately and quickly locating them. However, its performance in predicting changes in brightness was less satisfactory. While it correctly identified the location of bright targets, it was not very sensitive to their size. This resulted in inaccurate timing of additive addition. Based on this observation, we further optimized the LSTM network by increasing the LSTM step size, which effectively improved the prediction of bright target size.
Claims
1. A targe detection based tundish light detection method, characterized in that, The specific steps are as follows, Step 1, taking a bright eye video monitor as a video data source, the video data is preliminarily extracted through a CNN network including a residual module and a self-attention module to obtain a candidate bright eye area; Step 2, using a traditional edge detection operator, the bright eye edge is extracted from the candidate bright eye area; by setting a bright eye brightness threshold, a bright eye range image is further extracted, the area of the bright eye range image is calculated, and the area of the bright eye range image and the brightness of the bright eye range image are used as features for predicting the change of the bright eye; Step 3, using an LSTM network to process the bright eye range image extracted above to predict the change trend thereof; The intermediate bright eye detection method based on target detection is implemented based on an intermediate bright eye detection network based on target detection, and the intermediate bright eye detection network based on target detection is divided into four layers, which are a CNN network layer, a late_pooling, two time domain convolution layers and three LSTM network layers in sequence; The CNN network layer is divided into two parts, the first part includes a three-dimensional convolution layer conv1, a ReLU layer and a maximum pooling layer pooling; the second part includes an Inception Block, a ReLU layer, a maximum pooling layer pooling; The late_pooling inputs the output result of the CNN network layer into a full connection layer and then performs pooling once; The time domain convolution layer includes 128 3*3 time domain convolution kernels; When the LSTM network layer is trained, the training output value of each frame is compared with the next frame of the video, and after the gradient is back propagated, the correct next frame is added to the training set; After the intermediate bright eye detection network based on target detection, the bright eye range image is processed as follows; X t = [X t,1 ,…,X t,i ](1) out = W i • X t + W h • tan(c t ) + W c • c t (3) Formula (1) is the feature vector obtained by the convolution processing of the image by the first step CNN network, denoted as X t,i , which represents the column vector of the image at time t; formula (2) is the initial and iterative calculation method of the cell state and hidden state in the LSTM; formula (3) is the output result, wherein W t is the weight matrix of neurons to neurons, W h is the weight matrix of the memory gate, tan(·) is the nonlinear part in the LSTM, W c is the weight matrix between the cell states, and the three matrices are the matrices that need to be iterated in the gradient direction of the LSTM.
2. The target detection-based tundish light spot detection method according to claim 1, characterized in that, In the CNN network layer, the size of the two-dimensional convolution kernel in the convolution layer is 3*3, the feature extraction convolution kernel in the spatial dimension is also 3*3, and the feature extraction convolution kernel in the time dimension is 5*5.
Citation Information
Patent Citations
Driver fatigue detection based on the long-term and short-term memory network
CN109886241A
Medical image classification method based on graph network time sequence
CN115222688A