A Concentrate Grade Prediction Method Based on Dual-Modal CNN Secondary Transfer Learning
Through the dual-modal CNN secondary transfer learning method, combined with the RGB-D data set and the dosing state data set, the transfer learning is adopted by dual-modal ISE-DenseNet and adaptive DTAE-KELM, which solves the problems of data lag and overfitting in the flotation concentrate grade detection technology, and achieves high-precision grade prediction.
Patent Information
- Application Number
- CN202310546945.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-05-16
AI Technical Summary
The existing flotation concentrate grade detection technology has problems such as data lag, poor real-time performance and overfitting, making it difficult to achieve high-precision prediction under small-scale training set conditions.
The dual-modal CNN secondary transfer learning method is adopted to train the dual-modal ISE-DenseNet network model through RGB-D large-scale data set, and transfer learning is performed on the small-scale dosing state data set, and retransfer learning is performed in combination with adaptive DTAE-KELM instead of the full connection layer and softmax.
It effectively expands the degree of difference in image features of adjacent grade levels, reduces the misidentification rate, solves the overfitting problem of single transfer learning, and improves the accuracy and recall rate of concentrate grade prediction.
Smart Images

Figure CN116503378B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of flotation concentrate grade detection, in particular to a method for predicting the concentrate grade by dual-modal CNN secondary transfer learning. Background Art
[0002] Currently, most of the detections of flotation concentrate grade adopt regular sampling on-site manually, and then obtain the corresponding grade through off-line chemical analysis and later calculation in the laboratory. The lag of detection data seriously affects the timely adjustment of production variables, resulting in the flotation quality not reaching the optimal level. In recent years, aiming at the difficulty of online detection of flotation concentrate grade and the poor real-time performance of manual chemical analysis method and the inability to give corresponding guiding information in time, domestic and foreign scholars have carried out some research work on the prediction of concentrate grade based on machine vision technology. The research results show that establishing a prediction model of concentrate grade according to the visual characteristics of the foam surface is an effective and feasible method. However, these methods randomly set parameters during the model establishment process, which is easy to cause overfitting phenomenon and difficult to ensure the optimal generalization performance. Moreover, there are great difficulties in accurately extracting the foam color characteristics and bubble stability. In recent years, deep learning has entered a stage of rapid development. In view of the superior feature extraction ability of the deep convolutional neural network, its application in the field of pattern recognition such as speech, action, and image has achieved breakthrough research results, which has attracted extensive attention and in-depth research of many scholars in the academic community. The convolutional neural network has also been applied in the feature extraction and recognition of flotation foam images, and the recognition effect is significantly better than that of the traditional artificial neural network. However, these methods only consider extracting the static depth features of the foam image, while the flotation concentrate grade is closely related to the movement characteristics of the foam surface. Researchers proposed a dual-stream feature extraction model based on deep learning, which extracts the appearance and movement characteristics of the foam to establish a prediction model of concentrate grade, with high prediction accuracy. However, the network structure is complex and the operation efficiency is low. The above methods need a large number of samples for training to obtain an excellent network structure. However, the on-site working environment of the flotation plant is harsh, and it is difficult to establish a large-scale sample. Transfer learning can effectively solve the problem of the large demand for training sample data volume and improve the classification accuracy of the convolutional neural network in the application of small sample datasets. In the existing transfer learning process of foam images, it is necessary to repeatedly iterate and train the fully connected layer and the classification algorithm. There are many parameters to be adjusted, and the parameter setting has a certain randomness, which is easy to fall into problems such as local minimum and overfitting, and it is difficult to ensure the optimal generalization performance. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method for predicting the concentrate grade by dual-modal CNN secondary transfer learning, which can effectively expand the difference degree of adjacent grade image features and reduce the misrecognition rate under the condition of a small-scale training set, effectively solve the overfitting problem of single transfer learning, and has high prediction accuracy and recall rate.
[0004] To achieve the above object, the present invention adopts the following technical solution: A method for predicting the concentrate grade by dual-modal CNN secondary transfer learning, comprising the following steps:
[0005] Step 1: Collect foam dual-modal images under three dosing states, namely normal, excessive, and insufficient. According to the dosing states and the corresponding concentrate grade data provided by the on-site laboratory, construct a small-scale dataset of dual-modal images under three dosing states;
[0006] Step 2: Use the RGB-D large-scale dataset to train the dual-modal ISE-DenseNet network model. Then, transfer the pre-trained model and freeze the first convolutional layer, the first three ISE-Dense Blocks, and the corresponding three transition layers and SE layers of the pre-training;
[0007] Step 3: Respectively use the small-scale datasets of dual-modal foam images under three dosing states to perform transfer learning training on the dual-modal ISE-DenseNet network model, and perform training and learning on the ISE-Dense Block4, fully connected layer, and softmax of the transferred model to obtain the dual-modal ISE-DenseNet pre-trained models under three dosing states, namely normal, excessive, and insufficient;
[0008] Step 4: Perform secondary transfer learning training on the dual-modal ISE-DenseNet pre-trained models under three dosing states, freeze the first convolutional layer, the first four ISE-Dense Blocks, and the corresponding transition layers and SE layers of the pre-trained models, and use the adaptive DTAE-KELM to replace the fully connected layer and softmax for transfer learning training;
[0009] Step 5: During the transfer learning training process, use the quantum wolf pack algorithm to adaptively optimize the L, C, and σ parameters of DTAE-KELM, and use the recognition accuracy of the training set as the fitness. Finally, obtain the concentrate grade prediction models under three dosing states;
[0010] Step 6: Real-time collect visible light and infrared images of the foam on the surface of the flotation cell. According to different dosing states, if it is a fault state, directly output the result; otherwise, use the model under the corresponding dosing state to predict the concentrate grade.
[0011] In a preferred embodiment, constructing the dual-modal ISE-DenseNet network model of the foam image specifically is:
[0012] Perform infrared thermal imaging on the foam on the surface of the flotation cell. The infrared image of the foam contains the dynamic characteristic information of the foam. Infrared thermal imaging can directly show the bubbles that collapse and merge. According to the flotation production conditions, the concentrate grade is divided into six grades: excellent, good, medium, qualified, poor, and abnormal. Comprehensively extract the bimodal image features of the foam visible light and infrared thermal imaging as the driving features for predicting the concentrate grade.
[0013] Improve SE-DenseNet using the Inception-v3 network structure. Perform an asymmetric operation on the convolution in the Dense Block, replace the 1×1 convolution and 3×3 convolution with the 1×3 and 3×1 convolution forms, embed SENet into the Dense Block, and add an SE module after the 1×3 and 3×1 convolution layers in the Dense Block to fuse and obtain ISE-DenseBlock.
[0014] Extract the depth features of the visible light and infrared images of the foam comprehensively. Based on the ISE-Dense Block, construct a bimodal ISE-DenseNet network model. The DenseNet network requires a 224×224 image with 3 channels as input. Decompose the bimodal 256×256 image into a low-frequency image and high-frequency scale images through NSST. Then, interpolate the original image, low-frequency image, and high-frequency scale images into 3 224×224 images as the input of DenseNet. After image decomposition, performing CNN feature extraction can fully extract the contour, texture, and edge detail information of the image. The constructed bimodal ISE-DenseNet network model contains the ISE-DenseNet network with upper and lower channels. Remove the pooling layer after the first convolutional layer of the original DenseNet, use the ISE-Dense Block to replace the Dense Block structure in DenseNet, and add an SE layer after the transition layer of each ISE-Dense Block block to make each channel have different weights. Each of the upper and lower ISE-DenseNet channels contains four ISE-Dense Block blocks, as well as three transition layers and three SE layers. Pool the last ISE-DenseBlock blocks of the two channels and then perform a fully connected operation, and concatenate them into FC0. Then, perform feature fusion and learning through the fully connected layers FC1 and FC2, and finally use softmax for multi-classification. According to the idea of transfer learning, directly transfer some parameters of the source domain training model to the target domain model. The RGB-D dataset contains RGB and depth-of-field modality images. The task of foam bimodal image recognition is similar to that of RGB-D image recognition, and pre-train the bimodal ISE-DenseNet network model. Pre-train the bimodal ISE-DenseNet model using the RGB-D large dataset, then transfer some model structures and parameters to the foam concentrate grade prediction model, and then perform secondary training on the model.
[0015] In a preferred embodiment, the concentrate grade prediction based on the secondary transfer learning of the bimodal ISE-DenseNet is specifically as follows:
[0016] First, construct a bimodal ISE-DenseNet network model based on ISE-DenseNet and pre-train the model with a large RGB-D dataset; secondly, use a small-scale bimodal dataset under three dosing states to perform transfer learning on the initial pre-trained model, and re-train the ISE-Dense Block4, fully connected layer, and softmax of the transferred model to obtain a bimodal ISE-DenseNet pre-trained model under three dosing states; then, use an adaptive deep kernel extreme learning machine to replace the fully connected layer and softmax for another transfer learning to obtain a bimodal ISE-DenseNet concentrate grade prediction model under three dosing states; finally, according to the dosing state recognition result, select the corresponding bimodal ISE-DenseNet secondary transfer learning model to predict the concentrate grade;
[0017] Connect multiple layers of extreme learning machine autoencoders in series as the feature learning network of KELM. The extreme learning machine autoencoder makes the input equal to the output, and completes high-level feature extraction through a feedforward neural network to construct a double-hidden-layer autoencoder extreme learning machine, and the number of nodes in both hidden layers is N h , set N h to be greater than the number of input connection points, and randomly generate the input weight vectors w 1 、w 2 and the bias b 1 ′, b 2 ′. Calculate the output matrix of the first hidden layer through the input X, w 1 and b 1 ′, and then calculate the output matrix H of the second hidden layer through the output matrix of the first hidden layer, w 2 and b 2 ′. Calculate the output weight matrix β i of each double-hidden-layer autoencoder extreme learning machine through formula (1): i The value of:
[0018]
[0019] Connect multiple double-hidden-layer autoencoder extreme learning machines in series with the kernel extreme learning machine KELM to form a deep double-hidden-layer autoencoder kernel extreme learning machine DTAE-KELM. The input weight of each hidden node H i is the transpose β i T of the previous output weight. Through formula (2), the original input data is abstracted and extracted layer by layer through L double-hidden-layer autoencoder extreme learning machines, and then mapped to a higher-dimensional space through KELM for decision-making;
[0020]
[0021] Perform transfer learning on the pre-trained bimodal ISE-DenseNet model. Freeze the first convolutional layer, the first three Dense Blocks in the pre-training: ISE-Dense Block1 to ISE-Dense Block3, and the corresponding three transition layers and SE layers, and only train the last ISE-Dense Block4; after model transfer, only train and learn ISE-Dense Block4 and the DTAE-KELM model; use the quantum wolf pack algorithm to adaptively optimize the number L of double-hidden layer autoencoder extreme learning machines, the penalty coefficient C of KELM, and the kernel function parameter σ of the DTAE-KELM model.
[0022] Compared with the prior art, the present invention has the following beneficial effects: Under the condition of a small-scale training set, the bimodal feature extraction method of the present invention can effectively expand the difference degree of adjacent grade image features and reduce the misrecognition rate. The secondary transfer learning can effectively solve the overfitting problem of single transfer learning. The concentrate grade prediction method combined with the dosing state recognition has higher accuracy (P RE ) and recall rate (R EC ). The recognition accuracy and stability of the concentrate grade prediction of the present invention are improved to a certain extent compared with the existing foam image deep learning method. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 The foam infrared thermal images of different concentrate grades in the preferred embodiment of the present invention. Among them, (a) is the visible light image, (b) is the infrared thermal image, (c) is the bubble collapse, (d) is the bubble merger, (e) is the excellent grade, (f) is the good grade, (g) is the medium grade, (h) is the qualified grade, (i) is the poor grade, and (j) is the abnormal.
[0024] Figure 2 The schematic diagram of the improved SE-Dense Block structure in the preferred embodiment of the present invention;
[0025] Figure 3 The bimodal ISE-DenseNet network model diagram in the preferred embodiment of the present invention;
[0026] Figure 4 The deep double-hidden layer autoencoder kernel extreme learning machine network model diagram in the preferred embodiment of the present invention;
[0027] Figure 5 The bimodal ISE-DenseNet adaptive transfer learning model diagram in the preferred embodiment of the present invention;
[0028] Figure 6 The implementation flow diagram of the concentrate grade prediction in the preferred embodiment of the present invention;
[0029] Figure 7 It is the change curve of the accuracy and loss value of the transfer learning process in the preferred embodiment of the present invention. Among them, (a) is the change curve of the accuracy and loss value of the training set, and (b) is the change curve of the accuracy and loss value of the test set;
[0030] Figure 8 It is the effect of the second transfer learning of four models in the preferred embodiment of the present invention. Among them, (a) is the change curve of fitness, and (b) is the accuracy curve of the test set;
[0031] Figure 9 It is the test result of different numbers of training samples in the preferred embodiment of the present invention;
[0032] Figure 10 It is the prediction result and comparison of the concentrate grade in the preferred embodiment of the present invention. Among them, (a) is the effect of the dual-modal single transfer learning under the dosing state, (b) is the effect of the single-modal second transfer learning under the dosing state, (c) is the effect of the direct dual-modal second transfer learning, and (d) is the effect of the dual-modal second transfer learning under the dosing state;
[0033] Figure 11 It is the prediction accuracy and recall rate of the concentrate grade in the preferred embodiment of the present invention. Among them, (a) is the P RE value of the prediction of each concentrate grade, and (b) is the R EC value. Detailed implementation manners
[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0036] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0037] The present invention proposes a method for predicting the concentrate grade by double-modal CNN secondary transfer learning. First, a deep learning network model for foam double-modal images based on improved SE-DenseNet is constructed, and the model is pre-trained with the RGB-D large dataset. Secondly, a small-scale dataset under different dosing states is constructed to retrain the last convolutional layer, fully connected layer and softmax of the transferred model. Finally, an adaptive deep kernel extreme learning machine is used to replace the fully connected layer and softmax for further transfer learning to obtain a concentrate grade prediction model under various dosing states. The method of the present invention can effectively expand the difference degree of image features of adjacent grade levels and reduce the misrecognition rate under the condition of a small-scale training set, effectively solving the overfitting problem of single transfer learning, and having high prediction accuracy and recall rate.
[0038] The specific technical solutions are as follows:
[0039] 1. Construction of the foam image double-modal ISE-DenseNet network model
[0040] At present, many domestic and foreign researchers have used foam visual feature parameters as the input of the model to train the concentrate grade prediction model, and the modeling methods have developed from traditional algorithms such as clustering algorithms, least square methods, and support vector machines to deep learning methods using convolutional neural networks. Previous studies only stayed at the processing of foam visible light images. The present invention performs infrared thermal imaging on the foam on the surface of the flotation cell, such as Figure 1As shown, Figure (a) is the visible light image of the foam, and Figure (b) is the corresponding infrared thermal image. When bubbles collapse or merge, heat is released, and the temperature is higher than that of other bubbles. After thermal imaging, a highlighted yellow area appears. The infrared image of the foam has a certain display effect on the collapsing and merging bubbles. In Figure (c), two bubbles collapse, and in Figure (d), three small bubbles merge into one large bubble. It can be seen that the infrared image of the foam contains dynamic characteristic information of the foam. The visible light image can only show static characteristics such as apparent color, size, shape, and distribution, while the corresponding infrared thermal image can directly display the collapsing and merging bubbles. According to the flotation production conditions, the concentrate grade can be divided into 6 grades: excellent, good, medium, qualified, poor, and abnormal. Figures (e) to (j) show the thermal images of the foam under the 6 grades: the fewer the collapsing and merging bubbles, the higher the stability of the bubbles, the better the bearing capacity, and the higher the concentrate grade; when the bubbles are too small and too many, some areas are cotton-like, the flow order is chaotic, and the temperature distribution difference is large, the concentrate grade is relatively poor; when abnormal conditions occur, the bubbles are severely hydrated, the rolling speed is fast, the heat is released quickly, and the overall image tends to be highlighted yellow. Therefore, the infrared thermal image of the foam can effectively characterize dynamic characteristics such as the collapse and merger of bubbles, and has a certain discrimination under different concentrate grade levels. Therefore, the dual-modal image features of the visible light and infrared thermal images of the foam can be comprehensively extracted as the driving features for predicting the concentrate grade.
[0041] In the present invention, DenseNet with relatively excellent current performance is selected to construct a concentrate grade prediction model. DenseNet can effectively solve the problem of gradient disappearance during the training of deep networks. However, in the Dense Block, the dense connection from any layer to subsequent layers easily leads to a certain redundancy in the extracted features. By leveraging the feature screening ability of SENet, SENet is embedded into DenseNet to obtain SE-DenseNet, so as to enhance favorable features and suppress redundant features, and fuse the advantages of both to improve the network robustness. SE-DenseNet can selectively enhance favorable features using global feature information and suppress unimportant features, alleviating the impact of feature redundancy. The present invention improves SE-DenseNet by referring to the Inception-v3 network structure, performing an asymmetric operation on the convolution in the Dense Block, and replacing the 1×1 convolution and 3×3 convolution with 1×3 and 3×1 convolution forms to improve the feature extraction efficiency and expressiveness of the model. The improved SE-Dense Block (Improved SE-Dense Block, ISE-Dense Block) is as Figure 2As shown in the figure, SENet is embedded into the Dense Block, and an SE module is added after the 1×3 and 3×1 convolutional layers in the Dense Block to fuse and obtain the ISE-Dense Block. The fused network can not only achieve lossless transmission of feature information, but also judge the nature of the feature information of each channel, thereby enhancing beneficial features and suppressing redundant features, and enhancing the robustness of the network.
[0042] In order to comprehensively extract the deep features of foam visible light and infrared images, based on the ISE-Dense Block, a bimodal ISE-DenseNet network model is constructed. The DenseNet network requires a 224×224 image with 3 channels as input. Therefore, the bimodal 256×256 image is decomposed into a low-frequency image and high-frequency scale images by NSST, and then the original image, low-frequency image, and high-frequency scale images are interpolated into 3 224×224 images as the input of the DenseNet. After the image decomposition, CNN feature extraction can fully mine the contour, texture, and edge detail information of the image, and also expand the number of samples, which is beneficial to improving the classification accuracy of the image. The structure of the bimodal ISE-DenseNet network model constructed in the present invention is as Figure 3 shown, and it includes an ISE-DenseNet network with upper and lower channels. The pooling layer after the first convolutional layer of the original DenseNet is removed. To prevent the loss of shallow features caused by the pooling operation, the ISE-Dense Block is used to replace the DenseBlock structure in the DenseNet, and an SE layer is added after the transition layer of each ISE-Dense Block block to make each channel have different weights. Each of the upper and lower ISE-DenseNet channels includes four ISE-Dense Block blocks, as well as three transition layers and three SE layers. After pooling the last ISE-Dense Block blocks of the two channels, a fully connected layer is performed, and they are cascaded and spliced into FC0, and then feature fusion and learning are performed through the fully connected layers FC1 and FC2. Finally, softmax is used for multi-classification. According to the transfer learning idea, if the tasks of the source domain and the target domain are similar, some parameters of the source domain training model can be directly transferred to the target domain model. The RGB-D dataset contains two-modal images of RGB and depth of field. The task of foam bimodal image recognition is similar to that of RGB-D image recognition. The bimodal ISE-DenseNet network model can be pre-trained with the relatively large-scale RGB-D dataset. The present invention pre-trains the bimodal ISE-DenseNet model with the existing RGB-D large dataset, then transfers some model structures and parameters to the foam concentrate grade prediction model, and then performs secondary training on the model to reduce the required hardware resources and improve the training efficiency.
[0043] 2. Prediction of Concentrate Grade Based on Dual-Modal ISE-DenseNet Secondary Transfer Learning
[0044] Under different dosing states, in order to improve the prediction effect of the CNN feature-driven model under a small-scale training set, a prediction method for concentrate grade based on dual-modal ISE-DenseNet secondary transfer learning is proposed. First, a dual-modal ISE-DenseNet network model based on ISE-DenseNet is constructed, and the model is pre-trained with a large RGB-D dataset; second, transfer learning is performed on the initial pre-trained model using a small-scale dual-modal dataset under 3 dosing states, and the ISE-Dense Block4, fully connected layer, and softmax of the transferred model are re-trained to obtain dual-modal ISE-DenseNet pre-trained models under 3 dosing states; then, an adaptive deep kernel extreme learning machine is used to replace the fully connected layer and softmax for further transfer learning to obtain dual-modal ISE-DenseNet concentrate grade prediction models under 3 dosing states; finally, according to the dosing state recognition result, the corresponding dual-modal ISE-DenseNet secondary transfer learning model is selected to predict the concentrate grade.
[0045] To reduce the influence of the penalty coefficient C and the kernel function parameter σ and improve the generalization performance of the kernel extreme learning machine (KELM), the present invention connects multiple levels of extreme learning machine autoencoders in series as the feature learning network of KELM. The extreme learning machine autoencoder makes the input equal to the output and completes high-level feature extraction through a feedforward neural network. To extract higher-dimensional sparse features, the present invention constructs a dual-hidden layer autoencoder extreme learning machine on this basis, as Figure 4 shown, the number of nodes in both hidden layers is N h , to achieve sparse feature representation, N h can be set to a value greater than the number of input connection points, and the input weight vectors w 1 , w 2 and the biases b 1 ′, b 2 ′ of the first and second hidden layer nodes are randomly generated. The output matrix of the first hidden layer is calculated through the input X, w 1 and b 1 ′, and then the output matrix H 2 of the second hidden layer is calculated through the output matrix of the first hidden layer, w 2 ′. The output weight matrix β i of each dual-hidden layer autoencoder extreme learning machine can be calculated through Equation (1): i value:
[0046]
[0047] The present invention draws on the construction idea of a deep learning network, and cascades multiple double-hidden-layer autoencoder extreme learning machines and a kernel extreme learning machine KELM, as Figure 4 shown, to form a Deep Two HiddenLayer Autoencoder Kernel Extreme Learning Machine (DTAE-KELM). The input weight of each hidden node H i is the transpose β of the output weight of the previous layer. i T The original input data is abstracted layer by layer through L double-hidden-layer autoencoder extreme learning machines via Equation (2), and then mapped to a higher-dimensional space through KELM for decision-making, which is beneficial to improving the recognition accuracy and generalization performance.
[0048]
[0049] The present invention uses an existing large-scale RGB-D dataset to train a bimodal ISE-DenseNet network model, and then performs transfer learning on the pre-trained bimodal ISE-DenseNet model, as Figure 5 shown. The first convolutional layer, the first three Dense Blocks: ISE-Dense Block1 to ISE-Dense Block3, and the corresponding three transition layers and SE layers of the pre-training are frozen, and only the last ISE-Dense Block4 is trained. The pre-trained bimodal ISE-DenseNet classifies the fully connected features through a softmax classifier and uses backpropagation to train all network parameters, which is prone to problems such as falling into local minima and overfitting. The present invention replaces the fully connected layer of the pre-trained model with multiple cascaded double-hidden-layer autoencoder extreme learning machines, and a kernel extreme learning machine (KELM) replaces the original softmax. After model transfer, only the ISE-Dense Block4 and the DTAE-KELM model need to be trained. To obtain the optimal fitting performance, a quantum wolf pack algorithm is used to adaptively optimize the number L of double-hidden-layer autoencoder extreme learning machines, the penalty coefficient C of KELM, and the kernel function parameter σ of the DTAE-KELM model during the training process.
[0050] 3. Specific implementation process and steps
[0051] The quality of the flotation chemical addition state directly affects the grade of the concentrate: under normal chemical addition conditions, the foam has good stability and viscosity, the bubbles have a strong adsorption force on the minerals, the flotation effect is good, the grade of the concentrate is relatively high, and it is basically stable at the excellent, good, and medium levels; in the excessive state, the foam has high viscosity, and both minerals and impurities adhere to the surface of the bubbles, resulting in a decrease in the concentrate grade, which is generally at the lower levels of medium, qualified, and poor; in the under-dose state, the foam has strong fluidity and a high bubble collapse rate, and the useful minerals easily sink to the bottom of the flotation cell, and the grade of the concentrate is not high, mainly concentrated in the medium, qualified, and poor levels. Therefore, in the present invention, secondary transfer learning is respectively performed on the dual-modal ISE-DenseNet under each chemical addition state to obtain prediction models for the concentrate grade under three chemical addition states, and on the basis of identifying the chemical addition state, the concentrate grade is further predicted. The implementation process is as Figure 6 shown, and the specific steps are as follows:
[0052] Step1: Collect dual-modal images of the foam under three chemical addition states, and construct a small-scale dataset of dual-modal images under three chemical addition states according to the chemical addition state and the corresponding concentrate grade data provided by the on-site laboratory.
[0053] Step2: Use the RGB-D large-scale dataset to train the dual-modal ISE-DenseNet network model, and then transfer the pre-trained model, freezing the first convolutional layer, the first three ISE-Dense Blocks, and the corresponding three transition layers and SE layers of the pre-training.
[0054] Step3: Respectively use the small-scale datasets of dual-modal foam images under three chemical addition states to perform transfer learning training on the dual-modal ISE-DenseNet network model, and perform training and learning on the ISE-Dense Block4, fully connected layer, and softmax of the transferred model to obtain pre-trained models of dual-modal ISE-DenseNet under three chemical addition states of normal, excessive, and under-dose.
[0055] Step4: Perform transfer learning training again on the pre-trained models of dual-modal ISE-DenseNet under three chemical addition states, freeze the first convolutional layer, the first four ISE-Dense Blocks, and the corresponding transition layers and SE layers of the pre-trained models, and use the adaptive DTAE-KELM to replace the fully connected layer and softmax for transfer learning training.
[0056] Step5: During the transfer learning training process, use the quantum wolf pack algorithm to adaptively optimize the L, C, and σ parameters of DTAE-KELM, and use the recognition accuracy of the training set as the fitness, and finally obtain prediction models for the concentrate grade under three chemical addition states.
[0057] Step 6: Collect visible light and infrared images of the foam on the surface of the flotation cell in real time. According to different chemical addition states, if it is a fault state, directly output the result; otherwise, use the model under the corresponding chemical addition state to predict the concentrate grade.
[0058] 4 Specific embodiments and descriptions
[0059] To verify the effectiveness of the method of the present invention, the foam images collected in the lead-zinc ore flotation plant of Fujian Jindong Mining Co., Ltd. are used as experimental samples. The hardware platform for the experiment is Intel(R) Core(TM) i7-9800X CPU@3.80GHz, NVIDIA GeForce RTX 3080Ti, 128GB RAM, and the software running environment is Windows 10, Matlab2019a, Python3.7, Pytorch1.7. The method proposed in the present invention is verified through experiments.
[0060] To verify the effectiveness of the concentrate grade prediction method in this paper, during the period from October 12, 2020 to December 20, 2020, a Flir T620 infrared thermal imager is used to collect bimodal images of the foam on the surface of the lead ore cleaning cell II. Under the normal chemical addition state, 5000×3 groups of bimodal foam images under the excellent, good, and medium grade levels are selected. Similarly, under the over-dose and under-dose chemical addition states, 5000×3 groups of bimodal foam images under the medium, qualified, and poor grade levels are selected to create a training data set for the concentrate grade prediction model under three chemical addition states. In addition, 1000×6 groups of bimodal images corresponding to the roughing cell and the cleaning cell II under six concentrate grade levels are collected as a test data set.
[0061] (1) Construction of the dual-modal ISE-DenseNet network model. First, the RGB-D dataset published by Lai et al. was used to pre-train the dual-modal ISE-DenseNet network model. 70% of the samples were used as the training set data, and 30% of the samples were used as the test set data. During the training process, the learning rate was set to 0.001, the learning momentum was set to 0.9, the weight decay coefficient was set to 0.0003, the loss function was selected as Cross-Entropy, the batch size was set to 32, and the training iteration times were set to 5000. To select the appropriate network depth, DenseNet and ISE-DenseNet with different numbers of layers were respectively selected to construct the dual-modal network for training and testing. The classification accuracy of the test set increased with the increase of the network depth. The classification accuracy of ISE-DenseNet under each network layer was higher than that of DenseNet. When the number of layers of ISE-DenseNet reached 52 layers, the accuracy began to level off, and the accuracy reached the highest and remained stable starting from 58 layers. Therefore, in this paper, the 58-layer ISE-DenseNet was selected to construct the dual-modal network as the initial pre-training model.
[0062] (2) The first transfer learning of the dual-modal ISE-DenseNet network model. The initial dual-modal ISE-DenseNet pre-training model was transferred. The first convolutional layer, the first three ISE-Dense Blocks, and the corresponding 3 transition layers and SE layers of the pre-training model were frozen, and the ISE-Dense Block4, the fully connected layer, and the softmax of the transferred model were trained. 5000×3 groups of samples in 3 dosing states were used as the training data, and 1000×6 groups of samples in 6 concentrate grade levels were used as the test data. After 2000 iterations of training, the dual-modal ISE-DenseNet pre-training models in 3 dosing states of normal, excessive, and insufficient were obtained. For comparative analysis, 5000×3 groups of samples in 3 dosing states were mixed together as the training data, and the initial dual-modal ISE-DenseNet pre-training model was directly transferred for learning training to obtain an overall dual-modal ISE-DenseNet pre-training model. The transfer learning training processes of the 4 models are as Figure 7As shown in the figure: During the training process of the four models, the fluctuations of the training set accuracy and loss value are relatively large. The accuracy of the over-dose and under-dose model training processes is relatively high, and the loss value is relatively low; the test set accuracy of the over-dose and under-dose models reaches more than 90%, the loss value is lower than 0.4, and the change is relatively stable, and the model performance is better; the test set accuracy of the normal model is slightly higher and the loss value is slightly lower than that of the overall model, and the performance is slightly better than that of the overall model; the accuracy and loss value of the overall model fluctuate more than those of the other three models in the later stage of training, and the model performance is poor. Therefore, using the samples under three dosing states as training data, the performance of the three dual-modal ISE-DenseNet pre-trained models, namely normal, over-dose, and under-dose, is better than that of the overall trained model.
[0063] (3) Second transfer learning of the dual-modal ISE-DenseNet network model. According to Figure 7 the learning and training process, it can be seen that different degrees of overfitting phenomena occur in the training processes of the four dual-modal ISE-DenseNet pre-trained models. This is because the full connection features are classified by softmax and backpropagation is used to train all network parameters, which easily falls into local minima and causes overfitting. Therefore, in the present invention, L serially connected dual-hidden layer autoencoder extreme learning machines are used to replace the full connection layer of the pre-trained model, and KELM is used to replace the original softmax. The four dual-modal ISE-DenseNet pre-trained models are transferred and learned again. The first convolutional layer, four ISE-Dense Blocks, and the corresponding transition layers and SE layers of the pre-trained model are frozen, and only the DTAE-KELM of the transferred model is trained and learned. To obtain the optimal classification performance, the quantum wolf pack algorithm is used to optimize parameters such as L, C, and σ during the training process. The ranges of the three parameters are: 1 ≤ L ≤ 10, 0.01 ≤ C ≤ 1000, 0.01 ≤ σ ≤ 100. The three parameters are used as the genes of the artificial wolves, and the recognition error of the training set is used as the fitness. In the experiment, the wolf pack size M = 500, the gene length h = 20, the distance factor ω = 500, the step factor τ = 1000, the update factor γ = 6, the exploration wolf ratio factor δ = 4, the maximum wandering times T max = 20, the nonlinear exponent λ = 1.6, the basic rotation angle Δθ = 0.3π, and the maximum iteration times K max = 500. The iterative training processes of the four models are as Figure 8As shown in the figure: During the training process of DTAE-KELM for the 4 models, the recognition error was used as the fitness. During the iterative training process, the recognition error gradually decreased. The convergence efficiency of the normal and under-quantity models was relatively high. After the number of iterations exceeded 300 times, the recognition errors of the 4 models on the training set stabilized between 1.0% and 1.8%. The accuracy of the normal and under-quantity models on the test set increased rapidly and changed relatively smoothly. Finally, the accuracy of the test set stabilized at about 95%. The accuracy of the over-quantity model on the test set was relatively low in the initial stage of iteration, but it increased significantly in the later stage of iteration, and the fluctuation of the accuracy of the test set was the smallest. Finally, the accuracy of the test set stabilized at about 96%. The accuracy of the test set of the overall model fluctuated more than the other three models during the training process, and the accuracy of the test set was also slightly lower than that of the other 3 models, and finally stabilized at about 94%. Therefore, after re-transfer learning the 4 pre-trained models of dual-modal ISE-DenseNet, the performance of the models has been improved to a certain extent, and the overfitting phenomenon has been effectively solved. Finally, the performance of the 3 prediction models for concentrate grade in the normal, over-quantity, and under-quantity states is better than that of the models trained as a whole.
[0064] To verify the influence of the number of training samples on the recognition accuracy of the 3 prediction models for concentrate grade in the normal, over-quantity, and under-quantity states, the training set was experimented by increasing 500×3 groups of samples each time, and the number of test samples was 1 / 3 of the number of training samples. The test results are as Figure 9 shown: When the number of samples was less than 1000×3 groups, the recognition accuracies of the 3 prediction models for concentrate grade were all relatively low, and the recognition accuracy of the prediction model for concentrate grade in the over-quantity state was the lowest. As the number of samples increased, the recognition accuracies of the 3 prediction models gradually increased, and the accuracy improvement of the prediction model for concentrate grade in the over-quantity state was the largest. When the number of samples exceeded 2500×3 groups, the 3 grade prediction models all had relatively high recognition accuracies, and the recognition accuracy exceeded 90%. When the number of samples exceeded 4000×3 groups, the recognition accuracies of the 3 grade prediction models tended to be stable and reached the highest value.
[0065] (4) Prediction effect and comparative analysis of concentrate grade. To verify the transfer learning performance of the concentrate grade prediction model of the present invention, 4000×3 groups of dual-modal images in 3 dosing states were used as training samples, and 1000×6 groups of dual-modal images in 6 concentrate grade levels were used as test samples to test the model. To compare the effect of transfer learning, the pre-trained model of dual-modal ISE-DenseNet with single transfer learning, the single-modal secondary transfer learning model trained only with visible light images, the dual-modal secondary transfer learning model directly trained without dosing state recognition, and the dual-modal secondary transfer learning model in the dosing state of this article were trained and tested using the same data set. The recognition confusion matrix of the test set is as Figure 10As shown, the six grids on the diagonal represent the correct recognition quantities of six grade levels, and the remaining grids represent the quantities of the actual grade levels on the X-axis misrecognized as the grade levels on the Y-axis. According to Figure 10 's recognition confusion matrix, the prediction accuracy and recall rate of each model are statistically analyzed, and the results are as Figure 11 shown: The visible light and infrared image information changes greatly under the fault state, and the fault recognition accuracy and recall rate of the four models are relatively high; because the "medium" grade of concentrate is included in the normal, over-dose, and under-dose dosing states, the recognition accuracy and recall rate of the "medium" grade of several models are the lowest; due to the certain differences between RGB-D images and foam dual-modal images, directly using the dual-modal ISE-DenseNet pre-training model with single-time transfer learning for prediction, the prediction accuracy and recall rate of the concentrate grade are relatively low; the average values of the accuracy (P RE ) and recall rate (R EC ) of the single-modal secondary transfer learning model trained with visible light images can reach 90%, but the standard deviations of P RE and R EC are relatively large, and P RE and R EC also need to be further improved; if the dosing state recognition is not used and the overall secondary transfer learning is directly performed on the dual-modal ISE-DenseNet initial pre-training model, the average values of P RE and R EC are both 92.44%, and the standard deviations of P RE and R EC are 3.34% and 3.37%; for the dual-modal secondary transfer learning model under the dosing state of the present invention, the average values of P RE and R EC are 94.79% and 94.77%, and the standard deviations of P RE and R EC are 2.50% and 2.39%. The P RE and R EC have been greatly improved and the standard deviation is the smallest. The experimental results show that: using the dual-modal image CNN feature extraction method can effectively expand the difference degree of adjacent grade image features and reduce the misrecognition rate between adjacent grade levels; secondary transfer learning can effectively solve the overfitting problem of single-time transfer learning, further improve the average values of P RE and R EC , and reduce the standard deviations of P RE and R EC ; performing secondary transfer learning under three dosing states respectively, the performance of the fused prediction model under the three states is better than that of the overall trained model.
[0066] The results of on-site tests show that: under the condition of a small-scale training set, the training efficiency and test accuracy of ISE-DenseNet are better than those of SE-DenseNet. The dual-modal feature extraction method can effectively expand the difference degree of image features of adjacent grade levels and reduce the misrecognition rate. The secondary transfer learning can effectively solve the overfitting problem of single transfer learning, improve the recognition accuracy and stability. The concentrate grade prediction method combined with the chemical addition state recognition has higher precision (P RE ) and recall rate (R EC ). The average values of P RE and R EC for the concentrate grade prediction of the present invention are 94.79% and 94.77% respectively. The standard deviations of P RE and R EC are 2.50% and 2.39% respectively. The prediction accuracy and stability are improved to a certain extent compared with the existing deep learning methods for foam images.
[0067] The working environment of the flotation site is harsh and it is difficult to establish a large-scale sample. To improve the prediction effect of the flotation concentrate grade driven by CNN features under a small-scale training set, and conduct deep learning on both foam visible light and infrared images at the same time, the foam dual-modal images and transfer learning are introduced into the construction of the model, and a concentrate grade prediction method based on dual-modal CNN secondary transfer learning is proposed. Under the condition of a small-scale training set, the dual-modal feature extraction method of the method of the present invention can effectively expand the difference degree of image features of adjacent grade levels and reduce the misrecognition rate. The secondary transfer learning can effectively solve the overfitting problem of single transfer learning. The concentrate grade prediction method combined with the chemical addition state recognition has higher precision (P RE ) and recall rate (R EC ). The recognition accuracy and stability of the concentrate grade prediction of the present invention are improved to a certain extent compared with the existing deep learning methods for foam images.
Claims
1. A prediction method for concentrate grade based on dual-modal CNN secondary transfer learning, characterized in that, it includes the following steps: Step 1: Collect foam dual-modal images under three dosing states of normal, excessive, and insufficient dosing. According to the dosing states provided by the on-site laboratory and the corresponding concentrate grade data, construct a small-scale dataset of dual-modal images under the three dosing states; Step 2: Use the RGB-D large-scale dataset to train the dual-modal ISE-DenseNet network model, and then transfer the pre-trained model, freezing the first convolutional layer, the first three ISE-Dense Blocks, and the corresponding three transition layers and SE layers of the pre-training; Step 3: Respectively use the small-scale datasets of dual-modal foam images under the three dosing states to perform transfer learning training on the dual-modal ISE-DenseNet network model, and perform training and learning on the ISE-Dense Block4, fully connected layer, and softmax of the transferred model to obtain the dual-modal ISE-DenseNet pre-trained models under the three dosing states of normal, excessive, and insufficient dosing; Step 4: Perform secondary transfer learning training on the dual-modal ISE-DenseNet pre-trained models under the three dosing states, freezing the first convolutional layer, the first four ISE-Dense Blocks, and the corresponding transition layers and SE layers of the pre-trained models, and using the adaptive DTAE-KELM to replace the fully connected layer and softmax for transfer learning training; Step 5: During the transfer learning training process, use the quantum wolf pack algorithm to adaptively optimize the L, C, and σ parameters of DTAE-KELM, with the recognition accuracy of the training set as the fitness, and finally obtain the prediction models for the concentrate grade under the three dosing states; Step 6: Real-time collect visible light and infrared images of the foam on the surface of the flotation cell. According to different dosing states, if it is a fault state, directly output the result, otherwise use the model under the corresponding dosing state to predict the concentrate grade; Specifically, constructing the dual-modal ISE-DenseNet network model for foam images is as follows: Perform infrared thermal imaging on the foam on the surface of the flotation cell. The infrared image of the foam contains the dynamic characteristic information of the foam, and the infrared thermal imaging can directly show the bubbles that collapse and merge. According to the flotation production conditions, the concentrate grade is divided into six grades: excellent, good, medium, qualified, poor, and abnormal; comprehensively extract the dual-modal image features of the visible light and infrared thermal imaging of the foam as the driving features for predicting the concentrate grade; Improve SE-DenseNet using the Inception-v3 network structure, perform asymmetric operations on the convolutions in the Dense Block, replace the 1×1 convolution and 3×3 convolution with 1×3 and 3×1 convolution forms, embed SENet into the DenseBlock, and add an SE module after the 1×3 and 3×1 convolution layers in the Dense Block to fuse and obtain the ISE-DenseBlock; Extract the depth features of the visible light and infrared images of the foam comprehensively. Based on the ISE-Dense Block, construct a bimodal ISE-DenseNet network model. The DenseNet network requires 224×224 images with 3 channels as input. Decompose the bimodal 256×256 images into low-frequency images and high-frequency scale images through NSST, and then interpolate the original image, low-frequency image, and high-frequency scale image into 3 224×224 images as the input of DenseNet. After image decomposition, CNN feature extraction can fully mine the contour, texture, and edge detail information of the images; the constructed bimodal ISE-DenseNet network model contains the ISE-DenseNet network with upper and lower channels. Remove the pooling layer after the first convolutional layer of the original DenseNet, use the ISE-Dense Block to replace the Dense Block structure in DenseNet, and add an SE layer after the transition layer of each ISE-Dense Block block to make each channel have different weights; each of the upper and lower ISE-DenseNet channels contains four ISE-Dense Block blocks, as well as three transition layers and three SE layers. Pool the last ISE-DenseBlock blocks of the two channels and then perform a fully connected operation, and concatenate them into FC0, and then perform feature fusion and learning through the fully connected layers FC1 and FC2, and finally use softmax for multi-classification; according to the idea of transfer learning, directly transfer some parameters of the source domain training model to the target domain model. The RGB-D dataset contains RGB and depth-of-field modal images, and pre-train the bimodal ISE-DenseNet network model; pre-train the bimodal ISE-DenseNet model using the RGB-D large dataset, and then transfer some model structures and parameters to the foam concentrate grade prediction model, and then perform secondary training on the model; The prediction of the concentrate grade based on the secondary transfer learning of the bimodal ISE-DenseNet is specifically as follows: First, construct a dual-modal ISE-DenseNet network model based on ISE-DenseNet and pre-train the model with a large-scale RGB-D dataset; second, perform transfer learning on the initial pre-trained model using a small-scale dual-modal dataset under three dosing states, and re-train the ISE-Dense Block4, fully connected layer, and softmax of the transferred model to obtain a dual-modal ISE-DenseNet pre-trained model under three dosing states; then, use an adaptive deep kernel extreme learning machine to replace the fully connected layer and softmax for further transfer learning to obtain a dual-modal ISE-DenseNet concentrate grade prediction model under three dosing states; finally, according to the dosing state recognition result, select the corresponding dual-modal ISE-DenseNet secondary transfer learning model to predict the concentrate grade; A feature learning network of KELM is constructed by cascading multiple layers of extreme learning machine autoencoders. The extreme learning machine autoencoder makes the input equal to the output and completes high-level feature extraction through a feedforward neural network. A double-hidden-layer autoencoder extreme learning machine is constructed, and the number of nodes in both hidden layers is N h , set N h to be greater than the number of input nodes, and randomly generate the input weight vectors w 1 , w 2 and biases b 1 ′, b 2 ′. Calculate the output matrix of the first hidden layer through the input X, w 1 and b 1 ′, and then calculate the output matrix H 2 of the second hidden layer through the output matrix of the first hidden layer, w 2 and b i ′. Calculate the output weight matrix β i of each double-hidden-layer autoencoder extreme learning machine according to Equation (1): A number of double-hidden-layer autoencoder extreme learning machines are cascaded with a kernel extreme learning machine KELM to form a deep double-hidden-layer autoencoder kernel extreme learning machine DTAE-KELM. The input weights of each hidden node H i are the transpose of the output weights of the previous layer The original input data is abstracted layer by layer through L double-hidden-layer autoencoder extreme learning machines by formula (2), and then mapped to a higher-dimensional space through KELM for decision-making; Perform transfer learning on the pre-trained dual-modal ISE-DenseNet model, freeze the first convolutional layer, the first three Dense Blocks of the pre-training: ISE-Dense Block1 to ISE-Dense Block3, and the corresponding three transition layers and SE layers, and only train the last ISE-Dense Block4; after model transfer, only train and learn the ISE-Dense Block4 and DTAE-KELM model; Use the quantum wolf pack algorithm to adaptively optimize the number L of the double-hidden layer auto-encoder extreme learning machines of the DTAE-KELM model, the penalty coefficient C of KELM, and the kernel function parameter σ.
Citation Information
Patent Citations
CNN and transfer learning based disease intelligent identification method and system
AU2020103613A4
Human behavior identification method based on 3D deep convolutional network
CN107506712A