A Zinc Flotation Condition Recognition Method Based on Key Frame Attention and Bi-GRU
The dynamic timing characteristics of bubble videos are extracted through the Bi-GRU model and the keyframe attention mechanism, which solves the problem that single bubble images are difficult to portray the dynamic changes in flotation, improves the accuracy of flotation conditions recognition and reduces the calculation cost.
Patent Information
- Application Number
- CN202211077061.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-05
AI Technical Summary
The existing flotation condition recognition method based on machine vision mainly relies on a single foam image, and it is difficult to fully characterize the dynamic changes in the flotation process. The foam video data is large and information is redundant, resulting in low recognition accuracy and high calculation cost.
The dynamic timing characteristics of bubble videos are extracted using the Bi-GRU model, and the degree of change of frames is calculated through the keyframe attention mechanism, giving higher weight to frames with large changes, reducing redundant information interference, and extracting key timing characteristics.
It improves the accuracy of flotation conditions recognition and reduces calculation costs, achieving a more comprehensive characterization and more efficient identification of the flotation process.
Smart Images

Figure CN115457439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of froth flotation, and particularly to a zinc flotation condition recognition method based on key-frame attention and Bi-GRU. Background Art
[0002] Ore dressing is a technical science for enriching useful mineral components in ores, and plays an important role in industrial fields such as metallurgy, chemical engineering, and building materials. Among them, froth flotation is one of the most important ore dressing methods, which mainly realizes the enrichment of useful minerals by utilizing the hydrophobicity differences between minerals.
[0003] At present, most flotation plants still rely on manual methods to monitor and control the flotation process. Due to the harsh environment in the flotation industrial site, and the manual monitoring and control methods are affected by human operation levels, subjective consciousness, etc., it is easy to cause problems such as working condition fluctuations, reagent waste, and unqualified concentrate quality. Since the flotation process is related to many process variables, there are diversity and coupling between process variables, and the flotation process technology and dynamics are very complex. Therefore, it is difficult to establish an effective mathematical model based on mechanism. The monitoring method based on machine vision has advantages such as operability, real-time performance, and accuracy. Therefore, studying the flotation condition recognition method based on machine vision is beneficial to the automatic control of flotation reagent addition and the stability of flotation performance indicators (such as concentrate grade and recovery rate).
[0004] Most of the existing methods for condition recognition based on machine vision use a single foam image to recognize the flotation condition. Since a single foam image only shows the foam state at a certain moment, it is difficult to comprehensively depict the flotation state. Compared with a single foam image, a foam video shows the dynamic change process of froth flotation and can more comprehensively depict the flotation state. However, video data has characteristics such as large data volume and information redundancy. In order to ensure the accuracy of flotation condition recognition and reduce the calculation amount, it is very necessary to propose an effective and efficient foam video feature extraction method. Summary of the Invention
[0005] The purpose of the present invention is to provide a zinc flotation condition recognition method based on key-frame attention and Bi-GRU. First, aiming at the fact that a single foam image is difficult to comprehensively reflect the dynamic flotation process, Bi-GRU is used to extract the dynamic time-series features of the sequence frames in the foam video, which can more comprehensively represent the dynamic froth flotation process. Then, aiming at the problem that there is more redundant information in the foam video data, the attention mechanism is used to calculate the change degree of each frame in the sequence frames, and the frames with large change degrees are regarded as key frames and given higher weights, which can reduce the interference of redundant information and further promote the extraction of key time-series features.
[0006] The specific steps of the technical solution adopted by the present invention are as follows:
[0007] Step 1: Install an industrial camera at the foam flotation site to collect foam videos of the zinc flotation process. According to the concentrate grade value, the foam videos are divided into 4 types of working conditions, and the concentrate grade intervals corresponding to the 4 types of working conditions are (51, 53], (53, 54], (54, 55], and (55, 57]. First, each foam video is 20 seconds long. Foam images are sampled from each foam video every 0.5 seconds to obtain 40 consecutive frames of foam images. Then, the average value x1 and standard deviation x2 of the bubble size, the contrast x3, correlation x4, entropy x5, homogeneity x6 of the foam texture, and the average gray value x7 of the foam color are extracted from each frame of foam image. These features are concatenated together to form a feature vector X t = [x1, x2, x3, x4, x5, x6, x7]. Then, the consecutive multiple frames of foam images sampled in the foam video are characterized as a time feature sequence {X t}, t = 1, 2, …, N, where N is the number of sampled image frames in each video. In the present invention, N is selected as 40;
[0008] Step 2: Construct a Bi-GRU model based on key-frame attention, and use the time feature sequence representing the foam video as the input to extract the dynamic key time series features of the foam video. The time feature sequence {Xt}, t = 1, 2, …, N is input into the Bi-GRU. The Bi-GRU is composed of GRUs for forward and backward transmission. Among them, the GRU for forward transmission reads the input sequence from X1 to XN: In the t-th GRU unit, the current feature vector X t and the hidden layer unit state at the previous moment are used as the input to obtain the hidden layer unit state at the current moment
[0009]
[0010] where GRUUnit(·) is the GRU unit function. The obtained hidden layer unit state continues the forward transmission and finally obtains the last hidden layer unit state The calculation formula is:
[0011]
[0012] Therefore, the time feature sequence {X t}, t = 1, 2, …, N is encoded by the forward-transmitted GRU into a series of hidden unit states with strong temporal correlation The backward-transmitted GRU reads the time series from X N to X1: In the t-th GRU unit, the current feature vector X t and the hidden layer unit state at the current moment As the input, obtain the hidden layer unit state at the previous moment
[0013]
[0014] The obtained hidden layer unit state Continue the backward transmission, and finally obtain the state of the first hidden layer unit The calculation formula is:
[0015]
[0016] Then the time feature sequence {X t}, t = 1, 2, …, N is also encoded by the backward-transmitted GRU into a series of hidden unit states with strong temporal correlation
[0017] Meanwhile, key-frame attention is added to the GRUs for forward and backward transmission respectively. By calculating the degree of change of each frame in the sequence frames of the foam video, the frames with a larger degree of change are regarded as key frames and given larger weights. In the key-frame attention of the forward-transmitted GRU, since the hidden unit state at the previous moment contains important information of all previous sequences, the correlation between the hidden unit state at the previous moment and the feature vector at the current moment can be used to infer the degree of change of the current frame, that is, the smaller the correlation, the greater the degree of change of the current frame. Use a multi-layer perceptron (MLP) to calculate the hidden unit state at the previous moment and the correlation of the feature vector X t at the current moment:
[0018]
[0019] where α t is the correlation of the t-th frame, W t , U t , b t , V t are the parameters of the key-frame attention, and tanh(·) is the activation function. Then the correlations of all frames are calculated in the same way and input into the SoftMax layer for normalization:
[0020]
[0021] where SoftMax(·) is the SoftMax function. Since the smaller the correlation indicates the greater the degree of change of the current frame, and the frames with a greater degree of change contain more useful information, the frames with a smaller correlation should be assigned higher weights. Take as the weight coefficient of the corresponding feature vector and accumulate the weighted feature vectors:
[0022]
[0023] Among them is the weighted cumulative feature vector obtained by the forward-transmission GRU. Similarly, the weighted cumulative feature vector of the backward-transmission GRU can be calculated The weighted cumulative feature vectors of the forward and backward GRUs and are added to obtain the final dynamic key timing feature This dynamic key timing feature contains the time-domain change information of the foam video and focuses on the important information of the foam video.
[0024] Step 3: Input the dynamic key timing feature X into the fully connected layer to map it to the label space, and calculate the probability belonging to each working condition category through SoftMax:
[0025]
[0026] where W f and b f are the parameters of the fully connected layer, and the category with the highest probability is output as the final recognition result of the working condition. During the training process, cross-entropy is selected as the loss function for the backpropagation algorithm:
[0027]
[0028] where y and refer to the true label and the predicted label respectively.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows: Most of the traditional methods extract the features of a single foam image to realize the recognition of the flotation working condition. Compared with a single foam image, the foam video shows the dynamic change process of the flotation process and can better reflect the flotation performance. In order to improve the accuracy of flotation working condition recognition, the present invention uses Bi-GRU to extract the dynamic timing features of the foam video, which can more comprehensively characterize the dynamically changing flotation process; Secondly, a key frame attention module is added to adaptively calculate the change degree of the sequence frames in the foam video, regard the frames with a larger change degree as key frames and assign higher weights, which solves the problems of information redundancy and high computational cost in video data, and can promote the extraction of the dynamic key timing features of the foam video. Description of the Drawings
[0030] Figure 1 is the overall network structure diagram of the present invention;
[0031] Figure 2 is the key frame attention module. Detailed Embodiments
[0032] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments.
[0033] Step 1: Install an industrial camera at the foam flotation site to collect foam videos of the zinc flotation process. Divide the foam videos into 4 types of working conditions according to the concentrate grade value. The concentrate grade intervals corresponding to the 4 types of working conditions are (51, 53], (53, 54], (54, 55], and (55, 57]. First, each foam video is 20 seconds long. Sample foam images from each foam video every 0.5 seconds to obtain 40 consecutive frames of foam images. Then, extract the average value x1 and standard deviation x2 of the bubble size, the contrast x3, correlation x4, entropy x5, homogeneity x6 of the foam texture, and the average gray value x7 of the foam color from each frame of foam image. Concatenate these features together to represent a feature vector X t = [x1, x2, x3, x4, x5, x6, x7]. Then, the consecutive multiple frames of foam images sampled in the foam video are characterized as a time feature sequence {X t}, t = 1, 2, …, N, where N is the number of sampled image frames in each video. In the present invention, N is selected as 40;
[0034] Step 2: Construct a Bi-GRU model based on key-frame attention, and use the time feature sequence representing the foam video as the input to extract the dynamic key time series features of the foam video. As Figure 1 shown, Bi-GRU is composed of GRUs for forward and backward transmission. Input the time feature sequence {X t}, t = 1, 2, …, N into Bi-GRU. Among them, the GRU for forward transmission reads the input sequence from X1 to X N : In the t-th GRU unit, the current feature vector X t and the hidden layer unit state at the previous moment are used as inputs to obtain the hidden layer unit state at the current moment
[0035]
[0036] where GRUUnit(·) is the GRU unit function. The obtained hidden layer unit state continues to be transmitted forward, and finally the last hidden layer unit state is obtained. The calculation formula is:
[0037]
[0038] Therefore, the time feature sequence {X t}, t = 1, 2, …, N is encoded by the GRU for forward transmission into a series of hidden unit states with strong temporal correlation The backward-transmitting GRU reads the time series from X N to X1: In the t-th GRU cell, the current feature vector X t and the hidden layer cell state at the current moment are used as inputs to obtain the hidden layer cell state at the previous moment
[0039]
[0040] The obtained hidden layer cell state continues the backward transmission and finally obtains the state of the first hidden layer cell The calculation formula is:
[0041]
[0042] Then the time feature sequence {X t}, t = 1, 2, …, N is also encoded by the backward-transmitting GRU into a series of hidden unit states with strong temporal correlations
[0043] At the same time, key frame attention is added to the forward and backward-transmitting GRUs respectively. By calculating the degree of change of each frame in the sequence of frames of the foam video, the frames with a larger degree of change are regarded as key frames and given larger weights. Figure 2 For the key frame attention in the forward-transmitting GRU, since the hidden unit state at the previous moment contains important information of all sequences at previous moments, the correlation between the hidden unit state at the previous moment and the feature vector at the current moment can be used to infer the degree of change of the current frame by calculating the key frame attention, that is, the smaller the correlation, the greater the degree of change of the current frame. The hidden unit state at the previous moment is calculated using a multi-layer perceptron (MLP) and the feature vector X at the current moment t correlation:
[0044]
[0045] where α t is the correlation of the t-th frame, W t , U t , b t , V t are the parameters of the key frame attention, and tanh(·) is the activation function. Then the correlations of all frames are calculated in the same way and input into the SoftMax layer for normalization:
[0046]
[0047] Among them, SoftMax(·) is the SoftMax function. Since the smaller the correlation, the greater the degree of change of the current frame, and the more useful information the frame with a greater degree of change contains, frames with smaller correlations should be assigned higher weights. Take it as the weight coefficient of the corresponding feature vector and accumulate the weighted feature vectors:
[0048]
[0049] Where is the weighted accumulated feature vector obtained by the forward-transmitted GRU. Similarly, the weighted accumulated feature vector of the backward-transmitted GRU can be calculated Add the weighted accumulated feature vectors of the forward and backward GRUs and to obtain the final dynamic key timing feature This dynamic key timing feature contains the time-domain change information of the foam video and focuses on the important information of the foam video.
[0050] Step 3: Input the dynamic key timing feature X into the fully connected layer to map it to the label space, and calculate the probability belonging to each working condition category through SoftMax:
[0051]
[0052] Where W f and b f are the parameters of the fully connected layer, and the category with the highest probability is output as the final recognition result of the working condition. During the training process, the cross-entropy is selected as the loss function for the backpropagation algorithm:
[0053]
[0054] Where y and refer to the true label and the predicted label respectively.
[0055] Taking the recognition accuracy, precision, recall rate, and F1 as the evaluation criteria, the recognition results of the method of the present invention and other methods for the flotation working condition are shown in the following table:
[0056] Table 1 Recognition results of the method of the present invention and other methods for the flotation working condition
[0057]
[0058] As can be seen from Table 1, compared with the Bi-GRU method, the method of the present invention adds key-frame attention to Bi-GRU, and its recognition performance is greatly improved. Compared with the other three methods, the method of the present invention achieves better results in terms of accuracy, precision, recall, and F1. Therefore, the method of the present invention combining Bi-GRU and key-frame attention has a good effect on the recognition of flotation conditions.
Claims
1. A zinc flotation condition recognition method based on key-frame attention and Bi-GRU, characterized in that The following steps are involved: Step 1: Since foam videos can show the dynamic flotation process, multiple consecutive foam image frames are sampled at equal time intervals from the foam video and manual features are extracted from each foam image, including bubble size features, foam texture features, and foam color features. The multiple consecutive foam image frames in the foam video are then represented as a temporal feature sequence. Step 2: Build a Bi-GRU model based on keyframe attention, and use the temporal feature sequence representing the foam video as input to extract the dynamic key temporal features of the foam video: Since Bi-GRU is composed of forward and backward GRU, it can read time series information from both the forward and backward directions, thereby more fully mining the temporal correlation in the time series and using Bi-GRU to extract the dynamic temporal features of the foam video; Input the time feature sequence { X t} t=1,2,…,N into the Bi-GRU. Among them, the GRU for forward transmission reads the input sequence from X 1 to X N and encodes it into a series of hidden unit states with strong temporal correlation ; while the GRU for backward transmission reads the input sequence from X N to X 1 and encodes it into a series of hidden unit states with strong temporal correlation ; Meanwhile, considering the feature that the attention mechanism can adaptively focus on the important parts of the data, we utilize this feature to construct key-frame attention and add it to the Bi-GRU to facilitate the extraction of the dynamic key temporal features of the foam video. Since the frames with large variations in the foam video sequence contain more useful information, the key-frame attention selects the key frames by calculating the degree of variation of each frame in the sequence and assigns higher weights to them. In the forward GRU, since the hidden unit state at the previous moment contains all the important information in the sequence at the previous moments, calculating the correlation between the hidden unit state at the previous moment and the feature vector at the current moment using the key-frame attention can infer the degree of variation of the current frame, that is, the smaller the correlation, the greater the degree of variation of the current frame. Use a multi-layer perceptron to calculate the hidden unit state at the previous moment and the feature vector at the current moment X t correlation: Formula (1) Among them is the correlation of the t th frame, is the parameter of the key-frame attention, tanh (·) is the activation function; then the correlations of all frames are calculated in the same way and input into the SoftMax layer for normalization: Equation (2) where SoftMax(·) is the SoftMax function; since the smaller the correlation indicates the greater the degree of change of the current frame, and the more useful information the frame with a greater degree of change contains, frames with smaller correlations should be assigned higher weights, and used as the weight coefficient of the corresponding feature vector and accumulate the weighted feature vectors: Formula (3) Among them, is the weighted cumulative feature vector obtained by the forward GRU; similarly, the weighted cumulative feature vector of the backward GRU can be calculated , and the weighted cumulative feature vectors of the forward and backward GRUs and are added to obtain the final dynamic key timing feature, which contains the time-domain change information of the foam video and focuses on the important information of the foam video; Step 3: The extracted dynamic key timing features are input into the fully connected layer to identify the working conditions of zinc flotation, and the probability of belonging to each working condition category is calculated through SoftMax. The category with the highest probability is output as the final identification result of the working condition.
2. The zinc flotation condition recognition method based on key-frame attention and Bi-GRU according to claim 1, wherein In the step 1, each foam video is 20 seconds long, and one foam image frame is sampled every 0.5 seconds, so 40 foam image frames are sampled for each foam video.
3. The zinc flotation condition recognition method based on key frame attention and Bi-GRU according to claim 1, characterized in that In the step 2: the number of neurons in the hidden layer of Bi-GRU is set to 64; during the training process, the Adam optimizer is used, the learning rate is set to 0.001, and the mini-batch is set to 100.
4. A zinc flotation condition recognition method based on key-frame attention and Bi-GRU according to claim 1, characterized in that In the step 3, cross entropy loss is used as the loss function and back propagation algorithm is used for training.
Citation Information
Patent Citations
Soft measurement method for key indexes in complex industrial process
CN111222798A
Lightweight video action recognition method
CN113673307A