A Lead Flotation Foam Image Classification Method and System Based on Deep Autonomous Learning

By integrating EfficientNetB6 and Transformer networks into an image classification network, and combining autonomous learning to select information-rich data from unlabeled data for labeling, the problem of high manpower consumption in the flotation process is solved, and precise control and optimization of the flotation process are achieved.

CN119600339BActive Publication Date: 2025-11-14CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411643708.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-11-14
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing technologies suffer from significant individual differences and uncertainties in subjective decision-making during the flotation process, leading to fluctuations in flotation results. Furthermore, deep learning models require large labeled datasets for training, which consumes a great deal of manpower and time.

Method used

An image classification network based on the fusion of EfficientNetB6 and Transformer network structures is adopted. It combines autonomous learning to select the data with the most information from unlabeled data for manual annotation. A novel loss function is used to optimize model training, reduce the need for manual annotation and improve classification accuracy.

Benefits of technology

It significantly reduces labor and time costs, improves the accuracy of image classification and training efficiency, and enables precise control and optimization of the flotation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600339B_ABST
    Figure CN119600339B_ABST
Patent Text Reader

Abstract

This invention relates to computer vision technology, and particularly to a method and system for classifying lead flotation foam images based on deep autonomous learning. The method includes constructing an image classification network that fuses EfficientNetB6 and Transformer network structures; acquiring labeled and unlabeled image data; preprocessing the images; sampling m samples from the labeled image data to initialize the image classification network for training; manually labeling n data points with the highest information content from the unlabeled image data using autonomous learning; and further training the initialized data using the selected data to obtain the fully trained image classification network. This invention significantly improves the model's training speed and ability to classify complex foam images, while reducing the number of samples required for labeling, thereby reducing labor and time costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision technology, and in particular to a method and system for classifying lead flotation foam images based on deep autonomous learning. Background Technology

[0002] Flotation is a widely used separation technology in mineral processing. It is a mature method for effectively separating minerals in aqueous slurry, relying on the surface characteristics of mineral particles. During flotation, target mineral particles selectively adhere to rising air bubbles after the addition of specific flotation reagents, and float to the surface to form a mineral-rich froth layer. This froth layer not only carries the concentrate rich in useful minerals but also serves as a window for evaluating flotation performance. By observing its color, stability, flow rate, and bubble size, real-time feedback on the flotation process can be obtained. Concentrate grade, tailings grade, and final recovery rate are key indicators determining the economic benefits of the flotation process, directly affecting the break-even point of mining enterprises. Therefore, the scientific and precise execution, monitoring, and control of the flotation process have become ongoing technical challenges for engineers. In traditional operations, experienced operators rely on their intuitive judgment of the froth state to make operational decisions in the flotation cell. However, the individual differences and uncertainties in subjective decision-making brought about by this method lead to fluctuations in flotation results, often resulting in problems such as insufficient utilization of raw materials, excessive consumption of reagents, and low resource recovery efficiency.

[0003] The introduction of machine vision technology has shown great potential in extracting surface features of flotation foam, opening up new avenues for accurately identifying the working conditions of minerals. With continuous technological advancements, especially the sustained breakthroughs in deep learning, deep convolutional neural networks (CNNs) have demonstrated outstanding performance in image recognition, classification, and segmentation. CNN-based models not only improve the accuracy of image processing but also rapidly identify different types of foam images, effectively guiding the flotation process. However, a significant challenge is that deep learning models typically require massive labeled datasets for training. The accurate labeling of this data usually relies on the extensive experience and knowledge of domain experts, a process that consumes substantial human and time resources. Summary of the Invention

[0004] To improve the accuracy of image classification while significantly reducing the overall cost and labor input of annotation, this invention proposes a lead flotation foam image classification method based on deep autonomous learning. An image classification network is constructed by fusing EfficientNetB6 and Transformer network structures. The trained image classification network is then used for image classification. The training process of the image classification network includes:

[0005] Acquire image data with and without classification labels, and preprocess the images;

[0006] Initialize and train the image classification network by sampling m samples from image data with classification labels;

[0007] Through autonomous learning, n data points with the highest information content are selected from unlabeled image data and manually labeled.

[0008] The selected data is used to further train the initial training data, resulting in a fully trained image classification network.

[0009] Furthermore, the image classification network consists of cascaded convolutional layers, EfficientNetB6, a Transformer network, and linear layers. EfficientNetB6 consists of seven cascaded convolutional modules, each of which consists of one or more convolutional units.

[0010] Furthermore, the seven convolutional modules that make up EfficientNetB6 are respectively composed of one convolutional unit, two convolutional units cascaded together, two convolutional units cascaded together, three convolutional units cascaded together, three convolutional units cascaded together, four convolutional units cascaded together, and one convolutional unit.

[0011] Furthermore, the data processing process of the convolutional unit includes: the input feature map is processed by cascaded convolutional layers, depthwise separable convolutional layers, self-attention modules, convolutional layers, and dropout layers to obtain a feature map, and this feature map is added to and fused with the input feature map through a skip connection layer as the output of the convolutional module.

[0012] Furthermore, the data processing of the self-attention module includes: the input feature map is processed by a cascaded global average pooling layer and two fully connected layers to obtain a feature map, which is then multiplied and fused with the input feature map by a skip connection layer to serve as the output of the self-attention module.

[0013] Furthermore, the Transformer network structure consists of four cascaded sub-layers with the same structure. Each sub-layer includes a multi-head attention layer and a location-based feedforward network, and each sub-layer employs residual connections.

[0014] Furthermore, the data processing of the linear layer includes: processing the feature map output by the Transformer network structure sequentially through a global average pooling layer, a fully connected layer, and a softmax classifier.

[0015] This invention also proposes a lead flotation foam image classification system based on deep autonomous learning, used to implement a lead flotation foam image classification method based on deep autonomous learning. The system includes using a trained image classification network to perform a lead flotation foam image classification task. The training process of the image classification network includes:

[0016] Acquire labeled and unlabeled image data, and preprocess the images;

[0017] Initialize and train the image classification network by sampling m samples from image data with classification labels;

[0018] The image classification network trained after initialization is used to classify and predict unlabeled image data, and the n samples with the lowest classification prediction confidence are manually labeled.

[0019] The manually labeled image data is added to the image data with classification labels, and then used to train the image classification network again.

[0020] Compared with existing technologies, the technical solution of this invention has the following significant advantages:

[0021] This invention proposes a novel deep autonomous learning algorithm and system for classifying lead flotation foam images, particularly suitable for processing flotation foam images captured by cameras in industrial settings. It innovatively constructs a deep autonomous learning framework by employing a novel loss function specifically designed to address class imbalance. This framework adjusts and integrates the features of EfficientNetB6 and Transformer, creating a fusion convolutional neural network model as the driving force for foam image classification. The model utilizes its convolutional neural network portion to deeply learn the unique features of flotation foam, while the autonomous learning portion focuses on effectively labeling unlabeled data. The most informative samples are intelligently selected and integrated into the labeled dataset for model training and updates. This approach significantly reduces the burden of manual labeling while achieving accurate information capture. The combination of deep learning and autonomous learning methods significantly improves the model's training speed and ability to classify complex foam images, while reducing the number of samples required for labeling, thereby lowering labor and time costs. Attached Figure Description

[0022] Figure 1 This is a complete flowchart of a novel deep autonomous learning-based lead flotation foam image classification algorithm in one embodiment of the present invention.

[0023] Figure 2 These are flotation foam images under five different working conditions in a flotation process according to an embodiment of the present invention;

[0024] Figure 3 This is a structural diagram of the novel fusion convolutional neural network model proposed in this invention;

[0025] Figure 4 This is a framework diagram of deep autonomous learning proposed in this invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] This invention proposes a deep autonomous learning-based image classification method for lead flotation foam. It constructs an image classification network that fuses EfficientNetB6 and Transformer network structures, and uses the trained image classification network for image classification. The training process of the image classification network includes:

[0028] Acquire image data with and without classification labels, and preprocess the images;

[0029] Initialize and train the image classification network by sampling m samples from image data with classification labels;

[0030] Through autonomous learning, n data points with the highest information content are selected from unlabeled image data and manually labeled.

[0031] The selected data is used to further train the initial training data, resulting in a fully trained image classification network.

[0032] The core objective of this invention is to introduce a novel lead ore flotation froth image classification algorithm and corresponding system. Accurate classification of flotation froth images facilitates real-time analysis of slurry conditions, enabling precise control of flotation operation variables and achieving continuous optimization and real-time management of the entire flotation process. Addressing the issues of class imbalance and high annotation workload in current froth image classification, this invention proposes a method combining the powerful feature extraction capabilities of deep learning with autonomous learning to reduce the need for manual annotation. This method intelligently selects samples that contribute most to improving model performance for annotation, significantly reducing annotation workload and effectively improving image classification accuracy and training efficiency, providing strong technical support for efficient lead ore processing and resource recovery. Figure 1 This invention involves data augmentation of the collected data, followed by training of the model. The trained model is then used for prediction of unlabeled samples. The method specifically includes the following steps:

[0033] Step 1: Dataset Construction;

[0034] Foam images were collected during the lead plating process, and a foam image dataset was constructed. This dataset consists of two parts: one part is a set of clearly labeled samples (X...). L Y L The other part is the unlabeled sample set X to be labeled. U ;

[0035] Step 2: Data preprocessing and augmentation;

[0036] The collected foam images undergo a series of preprocessing measures, including flipping, rotating, and brightness adjustment, to achieve data augmentation and expand and enrich the sample image set. Then, all foam images are divided into training set and test set according to the proportion. In this embodiment, the data in the training set is used for model training, and the data in the test set is used to verify the effectiveness of the invention.

[0037] Step 3: Construct a deep network architecture;

[0038] The deep network architecture is based on the fusion of EfficientNetB6 and Transformer network models, which includes one 3×3 convolutional layer, seven MBConv modules, and one Transformer encoder block. The Transformer encoder block is placed after the seven MBConv modules to better handle complex patterns and relationships in the image.

[0039] Step 4: Perform initial training of the network model;

[0040] From the labeled sample set (X) L Y L The initial training set L = (x1, x2, ..., xn) is formed by randomly selecting m samples from the given data. m Then these samples are input into the network model to be trained to begin the initial training of the model;

[0041] Step 5: Sample selection and model update;

[0042] Through autonomous learning, the unlabeled sample set X U The most informative samples are selected and manually labeled, and then these newly labeled samples are used to further train and adjust the current model.

[0043] Step Six: Obtain the fusion model and perform image classification;

[0044] After completing all training and adjustments, a highly efficient convolutional neural network model integrating the Transformer and EfficientNetB6 architectures is obtained. The foam image data to be identified is then input into this fusion model for classification and recognition, producing the final classification decision for the foam image.

[0045] As an optional implementation, this embodiment divides the images in the foam image dataset into five different categories based on their characteristics, such as... Figure 2 They are labeled as Class I, Class II, Class III, Class IV, and Class V, respectively, and are denoted as abnormal, qualified, medium, good, and excellent. The quality value ranges are (-, 43], (43, 44], (44, 45], (45, 46], and (46, -) respectively.

[0046] like Figure 3 In this embodiment, the image classification network that fuses EfficientNetB6 and Transformer network structures includes the following structure:

[0047] The convolutional (Conv) layer receives a 512×512 pixel bubble image as input from the data layer. It uses a 3×3 convolutional kernel with a stride of 2 to convolve the input data. After batch normalization, ReLU activation function, and max pooling layer, 32 feature maps of size 224×224 are obtained and passed to the first convolutional (MBConv1) module.

[0048] The scaling factor in the MBConv1 module is 1. The MBConv1 module consists of one convolutional unit (MBConv). One MBConv includes, from front to back, one 1×1 Conv layer, one Depwise Conv layer (3×3 convolution), one SE module, one 1×1 Conv layer, one Dropout layer, and a shortcut connection. The SE module consists of one global average pooling and two fully connected layers, which finally extract 16 feature maps of size 112×112, which are then passed to the second convolutional module (MBConv2).

[0049] The scaling factor in the MBConv2 module is 6. The MBConv2 module consists of two MBConv modules. The convolution kernel size used by DepthwiseConv is 3×3. Finally, 24 feature maps of size 112×112 are extracted and passed to the third convolution (MBConv3) module.

[0050] The scaling factor in the MBConv3 module is 6. The MBConv3 module consists of two MBConv modules. The convolution kernel size used in DepthwiseConv is 5×5. Finally, 40 feature maps of size 224×224 were extracted. The fourth convolution (MBConv4) module.

[0051] The scaling factor in the MBConv4 module is 6. The MBConv4 module consists of 3 MBConv modules. The convolution kernel size used by DepthwiseConv is 3×3. Finally, 80 feature maps of size 112×112 are extracted and passed to the fifth convolution (MBConv5) module.

[0052] The scaling factor in the MBConv5 module is 6. The MBConv5 module consists of 3 MBConv modules. The convolution kernel size used by DepthwiseConv is 5×5. Finally, 112 feature maps of size 64×64 are extracted and passed to the sixth convolution (MBConv6) module.

[0053] The scaling factor in the MBConv6 module is 6. The MBConv6 module consists of 4 MBConv modules. The convolution kernel size used by DepthwiseConv is 5×5. Finally, 192 feature maps of size 32×32 were extracted and passed to the seventh convolution (MBConv7) module.

[0054] The scaling factor in the MBConv7 module is 6. The MBConv7 module consists of one MBConv module. The convolution kernel size used by DepthwiseConv is 3×3. Finally, 320 feature maps of size 32×32 are extracted and passed to the Transformer encoder module.

[0055] The Transformer encoder module consists of four identical layers stacked together, each layer having two sub-layers. The first sub-layer is a multi-head self-attention convergence network; the second sub-layer is a position-based feedforward network.

[0056] When calculating the self-attention of the encoder, the query, key, and value all come from the output of the previous encoder layer, and each sub-layer uses residual connections; finally, 1280 feature maps of size 32×32 were extracted; the output of the Transformer layer first passes through a global average pooling layer, which is responsible for averaging the feature maps to reduce the dimensionality, thereby producing a one-dimensional feature vector.

[0057] This vector is then fed into a fully connected layer. In this layer, to prevent the model from overfitting due to too much training data, the Dropout technique is used to randomly discard a certain percentage of the neuron outputs.

[0058] After processing by the fully connected layer, the final output is fed into the softmax classifier, which is responsible for outputting the final category result.

[0059] In this invention, the process of manually labeling n data points with the highest information content from unlabeled image data through autonomous learning includes the following steps:

[0060] First, an autonomous learning strategy is used to evaluate the unlabeled sample set X. U The information content of each unlabeled sample is used to rank the samples from highest to lowest according to their information content.

[0061] Next, the top K samples with the most information are selected and manually labeled to generate new sample label pairs (x, k, y). * y * );

[0062] These newly labeled sample pairs will be merged into the already labeled sample set (X). L Y L This data is then used as a new dataset to optimize and adjust the training model.

[0063] By iteratively executing this process, the model continuously learns and adapts until it reaches the predetermined performance standard or all unlabeled sample sets have been labeled; finally, the finely tuned convolutional neural network model is saved.

[0064] like Figure 4 The autonomous learning strategy employs a built-in loss prediction mechanism. This mechanism adds a loss prediction module to the deep learning model, aiming to estimate the possible loss value of unlabeled samples, thereby quantifying the information content of each sample in the unlabeled sample pool. Through this method, the model can identify and select the top K samples with the highest predicted loss values, then label these samples and add them to the training set for further model training.

[0065] During the model training phase, a weighted loss function that integrates class weights and loss prediction functionality is adopted. This loss function is designed to fully consider the weight differences between different classes and incorporates the output of the loss prediction module, thus forming a comprehensive loss function expression. It is defined as follows:

[0066]

[0067] The first part of the formula is the weighted loss based on the different importance of each category, and the second part is the prediction loss; N represents the number of samples in a mini-batch, C represents the total number of categories, and w c y represents the weight of the c-th category. c Represents the actual category value, This represents the predicted class value; α is the weight. Indicate l c and The predicted loss between, l c Let represent the target loss for the c-th class sample (this loss is obtained by calculating the difference between the true label and the predicted result). This represents the predicted loss value for the c-th category sample (in this embodiment, during training, the feature vector used for prediction in this invention is fitted to the distribution of the true loss through an auxiliary network (e.g., a shallow neural network), and the loss value obtained by the auxiliary network fitting is used as the predicted loss value).

[0068]

[0069]

[0070] in, Describes the target loss for a sample x in class c. The target loss of another sample y in category c. The differences between them This represents the predicted loss value for a sample x in category c. δ represents the predicted loss value of a sample y in category c; δ represents a pre-defined positive boundary value.

[0071] Table 1 Simulation Results

[0072]

[0073] To verify the effectiveness of this invention, the classification method of this invention is compared with other image classification methods in the prior art. The accuracy (the proportion of correctly predicted samples out of the total number of samples), precision (the proportion of samples predicted as positive by the model that are actually positive), recall, and F1 score are used to evaluate the classification results of each classification model on lead flotation foam images. The simulation results shown in Table 1 show that the accuracy, precision, recall, and F1 score of this invention are significantly better than other existing image classification models in the prior art.

[0074] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for classifying lead flotation foam images based on deep autonomous learning, characterized in that, An image classification network integrating EfficientNetB6 and Transformer network structures is constructed. The trained image classification network is used for image classification. The network consists of cascaded convolutional layers, EfficientNetB6, a Transformer network, and linear layers. EfficientNetB6 comprises seven cascaded convolutional modules, each consisting of one or more convolutional units. The seven convolutional modules forming EfficientNetB6 are respectively composed of one, two, two, or three convolutional units. The network consists of three cascaded convolutional units, four cascaded convolutional units, and one convolutional unit. The convolutional unit processes data as follows: the input feature map is processed through cascaded convolutional layers, depthwise separable convolutional layers, a self-attention module, another convolutional layer, and a dropout layer to obtain a new feature map. This new feature map is then fused with the input feature map through a skip connection layer to become the output of the convolutional module. The self-attention module processes data as follows: the input feature map is processed through cascaded global average pooling layers and two fully connected layers to obtain a new feature map, which is then multiplied with the input feature map through a skip connection layer to become the output of the self-attention module. The training process of the image classification network includes: Acquire image data with and without classification labels, and preprocess the images; Initialize and train the image classification network by sampling m samples from image data with classification labels; Through autonomous learning, n data points with the highest information content are selected from unlabeled image data and manually labeled. The manually labeled data is used to update the image data with classification labels, and the updated image dataset is used to further train the initialized image classification network. The image classification network, after further training, was used to perform the task of classifying lead flotation foam images.

2. The method for classifying lead flotation foam images based on deep autonomous learning according to claim 1, characterized in that, The Transformer network structure consists of four cascaded sub-layers with the same structure. Each sub-layer includes a multi-head attention layer and a location-based feedforward network, and each sub-layer uses residual connections.

3. The lead flotation foam image classification method based on deep autonomous learning according to claim 1, characterized in that, The linear layer's data processing involves sequentially passing the feature map output by the Transformer network structure through a global average pooling layer, a fully connected layer, and a softmax classifier.

4. The method for classifying lead flotation foam images based on deep autonomous learning according to claim 1, characterized in that, The loss function of an image classification network is expressed as: Where N represents the size of a mini-batch of samples, and C represents the total number of categories. This represents the weight of category c. This represents the true class value of the c-th class sample. This represents the predicted class value for the c-th class sample; As weight, This represents the target loss for the c-th class of samples. This represents the predicted loss value for the c-th class sample. express and The predicted loss between Describes the target loss for a sample x in class c. The target loss of another sample y in category c. The differences between them This represents the predicted loss value for a sample x in category c. This represents the predicted loss value for a sample y in category c; It represents a pre-defined positive boundary value.

5. A lead flotation foam image classification system based on deep autonomous learning, characterized in that, To implement the lead flotation foam image classification method based on deep autonomous learning as described in claim 1, the method includes performing a lead flotation foam image classification task using a trained image classification network. The training process of the image classification network includes: Acquire labeled and unlabeled image data, and preprocess the images; Initialize and train the image classification network by sampling m samples from image data with classification labels; The image classification network trained after initialization is used to classify and predict unlabeled image data, and the n samples with the lowest classification prediction confidence are selected for manual labeling. The manually labeled image data is added to the image data with classification labels, and then used to train the image classification network again.

Citation Information

Patent Citations

  • Image classification method based on semi-supervised self-paced learning cross-task deep network

    CN108764281A

  • Small-scale target component segmentation method based on deep learning

    CN115482381A