An intelligent classification method for grade based on ore image features

By constructing a residual neural network model with multi-branch structure and introducing an ECA attention mechanism, the problem of overfitting the deep learning ore grade classification model is solved, and efficient and accurate ore grade level recognition is achieved.

CN119273958BActive Publication Date: 2025-06-06YICHUN JIANGLI LITHIUM BATTERY NEW ENERGY IND RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411123126.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-06-06
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

The existing ore grade classification methods based on deep learning are prone to overfitting, resulting in insufficient reliability of ore grade classification and may miss important grading characteristics, affecting the grading accuracy.

Method used

A residual neural network model with multi-branch structure is adopted, combining transfer learning and ECA attention mechanism, and the model is trained to reduce overfitting and improve the generalization and comprehensiveness of feature extraction.

Benefits of technology

It realizes rapid and accurate identification of ore grade levels, improves the recognition performance and real-time performance of the model, and ensures recognition accuracy and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273958B_ABST
    Figure CN119273958B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent classification method for grade based on ore image features, belonging to the field of machine vision technology, the method comprising: constructing an ore image data set, dividing the ore image data set into a training set, a validation set and a test set, and performing preprocessing operations on the ore image data in each subset; constructing a residual neural network model with a multi-branch structure as an ore grade detection model; training the ore grade detection model based on the ore image data set; and using the trained ore grade detection model to detect the grade of the ore to be tested. The present invention constructs an ore grade classification model based on a multi-branch structure and a residual structure, which can improve the intelligence level of mines in the ore dressing link, provide guidance for production ore matching, and has important practical application value and significance for improving the operation efficiency of mines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of machine vision technology, and in particular to an intelligent classification method for grade based on ore image features. Background Art

[0002] In mining production, the beneficiation stage is one of the sources of mineral mining and processing. Its sorting accuracy and production efficiency are closely related to the grade information of the ore, and the quality of sorting directly affects the quality and market competitiveness of the final product. By optimizing the sorting process, timely detecting and obtaining the grade information of the ore and classifying it, it can provide guidance for subsequent production and obtain mineral products that are more suitable for different specific application needs, thereby meeting the needs of different industries. Among them, grade refers to the proportion of key substances in the ore. Taking the carbon grade of graphite ore as an example, it is the proportion of carbon content. At present, the traditional grade detection method is mainly carried out through chemical experiments or physical instruments. Using this method, technicians need to conduct a large number of sampling and testing regularly every day. The advantage is that the analysis results are accurate and effective, but this technology faces two significant disadvantages: 1) Large randomness: At present, the mining workshop uses manual sampling to determine the grade. The sampling has a large randomness and is greatly affected by human participation factors; 2) Low timeliness: The entire "sampling-sample preparation-testing" analysis chain is long, and there is usually a delay of several hours. After the test results came out, the ore of the test batch was mixed with the ore of the subsequent shifts, making it difficult to grasp the average grade of the ore in time. This made it difficult to effectively guide the ore discharge process at the mining site and the production plan in the future.

[0003] With the rapid development of computer vision and deep learning technology, it has become possible to use deep learning technology to directly apply to ore images for grade classification. The efficient convolutional neural network uses continuous convolution of convolution kernels of different sizes. Each convolution is performed on the basis of the previous convolution, so as to gradually extract various features on the ore. These features are combined with the corresponding grade information to achieve comprehensive classification of grade. The efficiency of this method is reflected in the real-time and accuracy of its classification, which can realize the rapid recording and quantitative evaluation of the ore grade information, and reduce labor costs and process complexity. The advantage of this method is that it can classify the ore grade very quickly. After collecting a relatively sufficient set of ore images, the ore grade classification can be achieved within a millisecond response time by training the artificial intelligence algorithm, and it can be deployed on the ore transportation production line in combination with the hardware and software system to achieve real-time sorting of the ore. The disadvantage of this method is that the model is prone to overfitting when the number of samples is limited, making it difficult to ensure the reliability of ore grade classification. At the same time, due to the small difference in ore characteristics between different grades, and the grade attributes of ore belong to the high-level semantic features of the image, which in turn rely on various complex low-level features (such as texture, cleavage, color, etc.), it is difficult to achieve such high-level semantic expression by relying solely on continuous convolution operations to extract low-level features of the image, and some important grading features may be omitted, resulting in the extracted features being not comprehensive and sufficient, affecting the improvement of grading accuracy; on the contrary, if there are a small number of samples, too much attention to high-level semantics leads to neglect of object details and overfitting, which will affect the model judgment in subsequent classification, resulting in insignificant improvement in effect, and thus lack of generalization and versatility. Secondly, deep learning models, especially convolutional neural networks, may require significant computing resources for training and inference, which means that traditional deep learning models require an effective and relatively high number of layers. In the case of a small number of samples, the improvement effect is limited, and the cost of the improvement is a surge in the number of model parameters and model complexity. This is a limiting factor in limited hardware resources, especially in mobile or embedded environments such as ore transportation production lines, which may affect the performance of the model in real-time applications.

[0004] In summary, the existing deep learning-based ore grade classification method is prone to overfitting, which makes it difficult to ensure the reliability of ore grade classification, and may miss some important classification features, resulting in the extracted features being not comprehensive and sufficient, affecting the classification accuracy. In this regard, how to use fewer layers to achieve better results on fewer sample data is an important direction for exploring how to improve model performance. Therefore, in order to achieve the fast, accurate and reliable characteristics of the deep learning-based ore grade classification method, in the case of limited samples of several ore images, it is a technical problem that needs to be solved to suppress the overfitting tendency of the model by technical means, guide the model to find features with better generalization and not miss any useful classification features, so as to ensure the recognition accuracy and versatility on different ores. Summary of the invention

[0005] The present invention provides an intelligent classification method for grade based on ore image features to solve the technical problems that the existing ore grade classification method based on deep learning is prone to overfitting, making it difficult to ensure the reliability of ore grade classification, and may miss some important classification features, resulting in the extracted features being not comprehensive and sufficient, thus affecting the classification accuracy.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] On the one hand, the present invention provides a method for intelligently classifying grades based on ore image features, the method comprising:

[0008] Constructing an ore image data set, dividing the ore image data set into a training set, a validation set and a test set, and performing preprocessing operations on the ore image data in the training set, the validation set and the test set respectively;

[0009] Construct a multi-branch residual neural network model as an ore grade detection model;

[0010] Training the ore grade detection model based on the ore image data set;

[0011] The trained ore grade detection model is used to detect the grade of the ore to be tested.

[0012] Furthermore, an ore image dataset is constructed, and the ore image dataset is divided into a training set, a validation set, and a test set, including:

[0013] Sample randomly distributed ores in the mine, take images of the ores, and measure their grade levels to build an ore image dataset with sufficient quantity, uniform grade distribution, and grade level labels;

[0014] The ore image dataset is divided into three subsets, namely a training set, a validation set and a test set, in a ratio of 6:2:2 to ensure that each subset does not contain image data from other subsets. Image hashing technology is used to compare the image data in each subset, and a secondary check is performed to determine whether the same image data exists in each subset to avoid data contamination.

[0015] Furthermore, the ore image data in the training set, validation set and test set are preprocessed respectively, including:

[0016] The ore image data in the training set are randomly scaled and cropped, and the ore image data are randomly transformed in brightness or contrast and randomly flipped horizontally with a probability of 80%;

[0017] The ore image data in the validation set and test set are only scaled and center cropped.

[0018] Furthermore, the ore grade detection model includes a backbone block, 5 IR_A-MBConv-block blocks, a first downsampling block, 10 IR_B-block blocks, a second downsampling block, 5 IR_C-block blocks, a 1×1 convolution block, a global average pooling layer and a fully connected layer, which are connected in sequence.

[0019] Furthermore, the trunk block includes three 3×3ConvBNAct blocks connected in sequence, a MaxPool layer with a stride of 2, a 1×1 convolution block, a 3×3MaxPool layer with a stride of 2, and multiple branches, wherein branch one in the trunk block is a 1×1 convolution block, branch two includes a 1×1 convolution block and a 5×5 convolution block connected in sequence, branch three includes a 1×1 convolution block and two MBConv blocks connected in sequence, branch four includes a 3×3AvgPool layer with a stride of 1 and a padding of 1 and a 1×1 convolution block connected in sequence, and the feature data output by each branch are finally feature spliced ​​according to the channel dimension.

[0020] Furthermore, the IR_A-MBConv-block, the IR_B-block and the IR_C-block are all multi-branch residual structures, each including a trunk for maintaining the original output and a plurality of branches; wherein,

[0021] Branch 1 in the IR_A-MBConv-block is a 1×1 convolution block, branch 2 contains a 1×1 convolution block and an MBConv block connected in sequence, branch 3 contains a 1×1 convolution block and two MBConv blocks connected in sequence, and branch 4 contains a 3×3 average pooling layer with a stride of 1 and a padding of 1 and a 1×1 convolution block connected in sequence. The feature data output by each branch is concatenated according to the channel dimension, passed through a 1×1 convolution layer, and then residually connected to the backbone block before being input into the activation function;

[0022] Compared with the IR_A-MBConv-block, the IR_B-block removes branch 2 in the IR_A-MBConv-block and adjusts the two MBConv blocks of branch 3 in the IR_A-MBConv-block to a 1×7 convolution block and a 7×1 convolution block;

[0023] Compared with the IR_A-MBConv-block, the IR_C-block removes branch 2 in the IR_A-MBConv-block, and adjusts the two MBConv blocks of branch 3 in the IR_A-MBConv-block to a 1×3 convolution block and a 3×1 convolution block; wherein, the last of the five IR_C-blocks in the ore grade detection model does not use an activation function after the last 1×1 convolution layer.

[0024] Furthermore, both the first downsampling block and the second downsampling block are multi-branch structures; wherein,

[0025] Branch 1 in the first downsampling block is a 1×1 convolution block, branch 2 contains a 1×1 convolution block and two MBConv blocks connected in sequence, and branch 3 is a 3×3 maximum pooling layer with a stride of 2 and a padding of 1. The feature data output by each branch are finally concatenated according to the channel dimension;

[0026] Both branches 1 and 2 in the second downsampling block consist of a 1×1 convolution block and an MBConv block connected in sequence. Branch 3 contains a 1×1 convolution block and two MBConv blocks connected in sequence. Branch 4 is a 3×3 maximum pooling layer with a stride of 2 and a padding of 1. The feature data output by each branch are finally concatenated according to the channel dimension.

[0027] Furthermore, the convolutional block contains a convolutional layer, a BN normalization layer, and an activation function connected in sequence;

[0028] The MBconv block consists of a dilated convolutional layer with sequentially connected skip connections, a 3×3 depthwise separable convolutional layer, a compression and excitation module, and a pointwise convolutional layer;

[0029] The activation function is the LeakyReLU function.

[0030] Furthermore, the ore grade detection model is trained based on the ore image dataset, including:

[0031] Using the transfer learning strategy, the constructed ore grade detection model is first fully trained on the large dataset ImageNet-1k to obtain the training parameters;

[0032] Rebuild an ore grade detection model with an ECA attention mechanism module;

[0033] Initializing the reconstructed ore grade detection model using the training parameters, and then training the model and adjusting parameters using the preprocessed training set and validation set ore image data based on a learning rate descent strategy combined with a weight decay mechanism until the optimal situation is reached;

[0034] The ore image data in the test set is input into the trained model to test the generalization and robustness of the model, and the grade of each ore image data in the test set is classified.

[0035] Furthermore, the ECA attention mechanism module is added between the IR_C-block and the 1×1 convolution block.

[0036] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0037] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the above method.

[0038] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0039] 1. A model for intelligent classification of ore grades was constructed based on multi-branch structure (Inception) and residual structure (Residual), which achieved rapid and accurate recognition of various ore grade classifications; by understanding the multi-layer and multi-angle image features and combining the transfer learning theory training model, the efficient transfer and reuse of multi-level feature knowledge of the image was achieved, which effectively reduced the training complexity of the model and improved the recognition performance of the model;

[0040] 2. The Inception-ResNet_V2 model was reasonably improved, some convolution operations were reasonably improved, some branch feature extraction functions were added, and efficient channel attention ECA was used to focus on the small feature differences between adjacent ore grades, which not only reduced the number of parameters and calculations, but also improved the classification accuracy and real-time performance of the model;

[0041] 3. Use the cosine annealing learning rate descent strategy and automatic mixed precision training techniques. This smooth learning rate curve descent strategy effectively speeds up the training progress and reduces memory usage, fully releasing the performance potential of the model;

[0042] 4. While using basic data enhancement methods such as random cropping, flipping, rotation, brightness and contrast transformation, we increase the probability of random brightness or contrast transformation by 80%, which effectively simulates the on-site lighting transformation in actual applications, increases the novelty of data during model training, prevents the model from overfitting when the number of data sets is insufficient, and ensures the robustness and versatility of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 It is a technical principle diagram of the intelligent classification method of grade based on ore image features provided by an embodiment of the present invention;

[0045] Figure 2 It is a schematic diagram of the execution flow of the intelligent classification method of grade based on ore image features provided by an embodiment of the present invention;

[0046] Figure 3 is a diagram of the overall network structure provided by an embodiment of the present invention;

[0047] Figure 4 It is a schematic diagram of the structure of a convolution block (ConvBNAct) provided in an embodiment of the present invention;

[0048] Figure 5 is a structural diagram of the MBConv module and the SE submodule provided in an embodiment of the present invention;

[0049] Figure 6 It is a schematic diagram of the main block structure and main branch splicing method provided by an embodiment of the present invention;

[0050] Figure 7is a schematic diagram of the ECA attention mechanism provided by an embodiment of the present invention;

[0051] Figure 8 is a training loss curve provided by an embodiment of the present invention;

[0052] Fig. 9 is a block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0054] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0055] First embodiment

[0056] In view of the technical problems that the existing deep learning-based ore grade classification method is prone to overfitting, making it difficult to ensure the reliability of ore grade classification, and may miss some important classification features, resulting in the extracted features being not comprehensive and sufficient, affecting the classification accuracy, this embodiment provides an intelligent grade classification method based on ore image features, such as Figure 1 As shown, after sampling randomly distributed ores in the mine, the method takes high-resolution images and determines the corresponding grade to construct an ore image data set; the ore image data in the data set is strictly divided into training set, validation set and test set, and a series of preprocessing operations are performed on the training set and validation set; a multi-branch structured residual neural network model is specially constructed, and a transfer learning strategy is used to first fully train on the large dataset ImageNet-1k, and the training parameters of the feature extractor are obtained for model parameter initialization. At the same time, the preprocessed training set and validation set data are input into the model for training and parameter adjustment until the optimal situation; the image data in the test set is input into the model in small batches, and the generalization and robustness of the model are tested, and each ore image therein is graded. The method can be implemented by an electronic device, which can be a terminal or a server. Specifically, the execution flow of the method is as follows. Figure 2 As shown, the following steps are included:

[0057] S1, constructing an ore image data set, dividing the ore image data set into a training set, a validation set and a test set, and performing preprocessing operations on the ore image data in the training set, the validation set and the test set respectively;

[0058] It should be noted that for the grade intelligent classification task based on ore image features, it is first necessary to construct a qualified ore image dataset, and the construction process is as follows:

[0059] Randomly distributed ores are sampled in the mine, high-resolution images are taken and the corresponding grade is determined, and a sufficient number of ore image datasets with uniform grade distribution and grade labels are constructed; the ore image data in the dataset is strictly divided into training set, validation set and test set in a ratio of 6:2:2, ensuring that there is no data from other subsets in the three subsets, and using image hashing technology to compare the image content in the dataset to check whether there is the same image data to avoid data contamination. A series of preprocessing operations are performed on each subset: including random resizing and cropping of the training set, random brightness or contrast transformation with an 80% probability, and random horizontal flipping; only resizing and center cropping are performed on the validation set and test set, and all subsets are guaranteed to have a final input size of 299×299 pixels.

[0060] S2, construct a residual neural network model with a multi-branch structure as an ore grade detection model;

[0061] It should be noted that, while acquiring the data set, this embodiment preliminarily constructs a residual neural network model with a multi-branch structure. For this, this embodiment uses the Inception-ResNet_V2 model as the backbone, and improves it from several angles, such as improving the classification accuracy of the model and reducing the amount of calculation and the amount of parameters. The model is constructed based on a module that combines a multi-branch structure (Inception) and a residual structure (Residual). This module is denoted as IR-Block. The classification model is obtained by continuously stacking multiple such modules and other auxiliary modules. The improved classification model is denoted as IRv2EffiNet. Figure 3 As shown in the figure, its main structure sequence includes: backbone block, 5 IR_A-MBConv-blocks, the first downsampling block, 10 IR_B-blocks, the second downsampling block, 5 IR_C-blocks, 1×1 convolution block, global average pooling layer and fully connected layer.

[0062] Among them, the convolution block is one of the basic structures of this classification model, and its structure is as follows Figure 4As shown in the figure, it contains a normal convolution layer, a BN normalization layer, and an activation function, which is denoted as ConvBNAct. At the same time, the MBconv block is used as another basic structure, which can be easily applied to various computing resource-constrained environments, such as mobile devices, embedded systems, etc. It plays the role of replacing the 3×3ConvBNAct block in all sub-branch structures, which can effectively reduce the number of parameters and improve the classification and recognition speed while ensuring high efficiency and high accuracy. Figure 5 As shown in the figure, the MBconv block consists of an expansion convolution layer (Expansion Layer) containing a skip connection (SkipConnection), a 3×3 depthwise separable convolution layer (DepthwiseConvolution), a squeeze and excitation module (Squeeze-and-Excitation, referred to as SE module), and a pointwise convolution layer (Pointwise Convolution); the SE module weights each channel through the channel attention mechanism to enhance the network's ability to capture important features. This module mainly consists of two parts: the compression part (Squeeze), which uses the global average pooling layer to compress the input feature map (size W×H×C) into a 1×1×C vector; the excitation part (Excitation), which generates channel weights through two fully connected layers, the first fully connected layer is used for dimensionality reduction, and the second fully connected layer is used to restore the original number of channels. There is also an activation function between the two fully connected layers; finally, the 1×1×C weight vector output by the excitation part is multiplied by the original feature map channel by channel to weight it according to the importance of the channel. It is worth adding that the activation functions in this embodiment all use the LeakyReLU function, which avoids the neuron death problem of the ReLU function by introducing a small non-zero slope when the input is negative, while maintaining the advantages of high computational efficiency and simple implementation; compared with the Swish function, the LeakyReLU function has more advantages in computational complexity and is suitable for resource-constrained environments.

[0063] The LeakyReLU function can be expressed by formula (1):

[0064] (1)

[0065] Here, x is the input and α is a small positive number (usually 0.01) that controls the slope of the negative input.

[0066] Furthermore, the detailed structure of the classification model used in this embodiment is as follows: the backbone block is composed of three 3×3ConvBNAct blocks, a MaxPool layer with a step size of 2, a 1×1ConvBNAct block, a 3×3MaxPool layer with a step size of 2, and multiple branches in sequence, wherein branch one is a 1×1ConvBNAct block, branch two contains a 1×1ConvBNAct block, a 5×5ConvBNAct block, branch three contains a 1×1ConvBNAct block, two MBConv blocks, branch four contains a 3×3AvgPool layer with a step size of 1 and a padding of 1 and a 1×1ConvBNAct block, and finally feature splicing is performed according to the channel dimension; the IR-Block modules are all multi-branch residual structures, including a backbone (Identity) that maintains the original output and several branches (Branch), wherein the IR_A-MBConv-block module structure is as follows: branch one is a 1×1ConvBNAct block, branch two contains a 1×1ConvBNA ct block, an MBConv block, branch three contains a 1×1ConvBNAct block, two MBConv blocks, branch four contains a 3×3 average pooling (AvgPool) layer with a stride of 1 and a padding of 1 and a 1×1ConvBNAct block. After feature concatenation according to the channel dimension, it passes through a 1×1 ordinary convolution layer and is then residually connected to the trunk before inputting into the LeakyReLU activation function; compared with the IR_A-MBConv-block module, the IR_B-block module removes branch two and adjusts the two MBConv blocks of branch three to a 1×7ConvBNAct block and a 7×1ConvBNAct block; compared with the IR_A-MBConv-block module, the IR_C-block module removes branch two and adjusts the two MBConv blocks of branch three to a 1×3ConvBNAct block and a 3×1ConvBNAct block; among them, the last of the five IR_C-block modules does not use an activation function after passing through the 1×1 ordinary convolution layer at the end.Both downsampling blocks are multi-branch structures; the first downsampling block has the following structure: branch 1 is a 1×1ConvBNAct block, branch 2 is a 1×1ConvBNAct block and two MBConv blocks, branch 3 is a 3×3 MaxPool layer with a stride of 2 and padding of 1, and finally features are spliced ​​according to the channel dimension; the second downsampling block has the following structure: branch 1 and branch 2 are both composed of a 1×1ConvBNAct block and an MBConv block, branch 3 is a 1×1ConvBNAct block and two MBConv blocks, branch 4 is a 3×3MaxPool layer with a stride of 2 and padding of 1, and finally features are spliced ​​according to the channel dimension. The feature splicing method can be expressed by formula (2):.

[0067] (2)

[0068] Among them, n represents the maximum number of branches, and the output feature map of each branch is represented by F 1 , F 2 ,…, F n Indicates that F concat is the feature map after splicing. The main block structure and the main branch splicing method of the present invention are as follows Figure 5 shown.

[0069] It is worth noting that in the above structure, the ConvBNAct block can increase the nonlinear representation ability of the model and better capture complex features because it contains a nonlinear activation function (LeakyReLU). The 1×1ConvBNAct block is used to maintain the same feature map size in the spatial dimension, but adjusts the width of the feature map by increasing or decreasing the number of channels. While maintaining the spatial resolution of the input feature map, the number of parameters and the amount of calculation are relatively small, which is very suitable for use in places where low computational costs are required. In the neural network used in this embodiment, the convolution operations on different branches extract different types of features. Generally speaking, a larger receptive field can capture rich semantic information, while a smaller receptive field can better retain details, which is achieved by controlling the convolution kernel size of the ConvBNAct block. The feature extraction effects of various branches are as follows: the Identity trunk retains the original features and original receptive field of the ore image to the greatest extent; various convolution branches are used to extract richer semantic information and local receptive fields in the ore image; the average pooling branch is used to extract global background features; the maximum pooling branch is used to extract local features. By reasonably fusing features from different branches, these features can be fully utilized to improve the performance of the model. This fusion strategy is crucial to improving the accuracy and robustness of the model.

[0070] S3, training the ore grade detection model based on the ore image data set;

[0071] It should be noted that after the model IRv2EffiNet of this embodiment is reasonably and fully constructed, this embodiment adopts the following steps to obtain the final model for ore grade detection.

[0072] S31, first adopts the transfer learning strategy, first uses the large dataset ImageNet-1k to fully train on IRv2EffiNet, and the obtained training parameters are saved as binary files for subsequent use.

[0073] S32, reconstruct an IRv2EffiNet model with an ECA attention mechanism. The ECA attention mechanism module (Efficient Channel Attention, referred to as ECA module) is added between the IR_C-block and the 1×1 convolution block. It should be noted that, unlike the SE module, ECA is an improved channel attention mechanism that aims to capture cross-channel dependencies at a lower computational cost. Its structure mainly includes a global average pooling layer and a one-dimensional convolution (1D convolution). A 1D convolution layer (usually with a kernel size of 3) is used to capture local interactions between channels. The fully connected layer is not used for dimensionality reduction and dimensionality increase, which greatly reduces the number of parameters and computation, making it suitable for lightweight models and real-time applications. The weight vector (1×1×C) output by the 1D convolution layer is used for channel weighting, enhancing important features, suppressing unimportant features, and multiplying the original feature map channel by channel to obtain the weighted feature map. The convolution kernel size of the 1D convolution can be adjusted to meet the needs of different network structures, capture richer channel dependencies, and effectively enhance the feature expression ability between channels. Therefore, it can pay good attention to the small feature differences between ore images of adjacent grades. In short, ECA avoids significantly increasing computational overhead while effectively improving model performance. The structure of the ECA attention mechanism is as follows: Figure 7 shown.

[0074] S33, initialize the parameters of the IRv2EffiNet model newly constructed in S32. The specific operation is: add the parameter file obtained by pre-training in S31 to the model, and randomly initialize the newly added ECA module.

[0075] S34, train the complete IRv2EffiNet model using the ore dataset constructed in S1. Before starting training, you need to determine the loss function, set the initial learning rate, plan the number of samples processed in each batch (batch size), the predetermined maximum number of iterations, select a suitable optimizer, use a learning rate reduction strategy to fully unleash the potential of the model, deploy a weight decay mechanism to avoid overfitting, and combine the data enhancement solution designed in S1 to further improve the model's generalization ability for input data.

[0076] S35, monitor the performance of the model by continuously using the validation set. At the end of each training cycle, immediately use the validation data set to perform performance evaluation to observe the accuracy dynamics and loss value changes of the model on unseen data. During this process, if the model shows a continuous improvement in accuracy and a steady decrease in the loss function value, continue training; once these indicators reverse, consider stopping the training process immediately to avoid resource waste and possible overfitting.

[0077] S36, after all training iterations, the trained IRv2EffiNet model and the trained weight file are obtained. The computational complexity and parameter volume of the model are reduced by 26.36% and 28.30% respectively compared with the basic model, which reduces the complexity to a certain extent and can meet the initial requirements of embedded deployment. The test set is used to verify the accuracy and classification speed of the model for ore grade classification. If the model shows a high accuracy and classification speed, it means that the model has good generalization and robustness and can meet the requirements of practical applications.

[0078] S4, using the trained ore grade detection model to detect the grade of the ore to be tested.

[0079] In summary, this embodiment provides a method for intelligent classification of grade based on ore image features with higher accuracy and better real-time performance. This method constructs a multi-branch structure, sets different feature extraction schemes on each branch, extracts features at different scales such as the receptive field, size space, and chromaticity channel of the ore image, reasonably integrates the features extracted from different branch features, and introduces an attention mechanism to focus on the subtle differences between adjacent grade image features, which can make full use of features and improve model performance. Ultimately, it can achieve rapid and accurate classification of the grade of new and unknown grade ores, and improve the level of intelligence in the mineral processing stage. This is particularly important for improving classification efficiency and expanding the range of ore application types.

[0080] Below, taking graphite primary ore as an example, the implementation process of the method of the present invention is further described, and the steps are as follows:

[0081] 1. The above IRv2EffiNet model was built based on Python 3.8.16 programming language and Pytorch 1.13.1 deep learning framework, and the number of output categories was set to 4.

[0082] 2. Pre-train the IRv2EffiNet model on the ImageNet-1K dataset, and let the model achieve a Top-1 accuracy of more than 82% on the validation set. During training, the image pixels are standardized, and the mean and standard deviation values ​​used are calculated based on the pixel values ​​of all image samples in the training set. The pre-trained model parameters are saved in a binary file.

[0083] 3. Construct a graphite primary ore image dataset with four categories of carbon grade, namely [0,5) carbon grade, [5,10) carbon grade, [10,15) carbon grade and [15,20) carbon grade. Strictly divide them into training set, validation set and test set in a ratio of 6:2:2 to ensure that there is no data from other subsets in the three subsets. Use image hashing technology to compare the image content in the dataset to check whether there is the same image data to avoid data contamination.

[0084] 4. Reconstruct the IRv2EffiNet model with the ECA attention mechanism added and determine it as the complete IRv2EffiNet model. It can be found that the computational complexity and parameter amount of the model are reduced by 26.36% and 28.30% respectively compared with the backbone model.

[0085] 5. Use the model pre-training parameter file obtained in step 2 to load into the complete IRv2EffiNet model. Before starting training, determine the loss function as the cross entropy loss function, set the initial learning rate to 5e-4, plan the number of samples processed per batch (batch size) to 32, and the maximum number of iterations to 300. Use AdamW as the optimizer, use the cosine annealing learning rate decline strategy (as low as 5e-6) to fully release the potential of the model, and deploy the weight decay mechanism to avoid overfitting. Use the Automatic Mixed Precision (AMP) training tool, which is a method to accelerate deep learning training by using lower precision (such as FP16) for calculation while maintaining higher precision (such as FP32) gradient accumulation, thereby increasing training speed and reducing memory usage. During training, the training set is randomly resized and cropped, randomly transformed with brightness or contrast with 80% probability, and randomly flipped horizontally; the validation set and test set are only resized and center cropped, and all subsets are guaranteed to have a final input size of 299×299 pixels. The image pixels are also standardized, and the mean and standard deviation values ​​used are calculated based on the pixel values ​​of all image samples in the training set. The training loss curve is as follows: Figure 8 shown.

[0086] 6. After each training cycle on the training data set, the model uses the validation set to test the accuracy and loss function change trend of the model. If the accuracy keeps increasing or the loss function keeps decreasing, the model training continues, otherwise the training is terminated. After the training is completed, the parameter file is saved, and after reloading it into the model, the generalization and robustness of the model are tested using the test set, and the grade grading test is also performed.

[0087] 7. After training, the IRv2EffiNet model that can intelligently classify ore grades was obtained. Its recognition accuracy on the validation set reached 96.23%. While maintaining high accuracy on the test set, the image processing speed was at the millisecond level, realizing a high-accuracy real-time intelligent classification function for ore grades.

[0088] 8. The above application uses a graphite oxide ore dataset with the same grade but smaller quantity repeatedly, and the recognition accuracy on the validation set reaches 90.11%. The above application uses a lithium ore dataset with 6 grades and smaller quantity, and the recognition accuracy on the validation set also reaches 88.71%, which reflects the versatility of the model.

[0089] Second embodiment

[0090] This embodiment provides an electronic device, such as Fig. 9 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment. In addition, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0091] Next, combine Fig. 9 The following is a detailed introduction to the various components of the electronic device:

[0092] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor may execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0093] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Fig. 9 The CPU0 and CPU1 shown in the figure are, of course, only exemplary.

[0094] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0095] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Fig. 9 The processor is coupled to the processor (not shown), which is not specifically limited in this embodiment of the present invention.

[0096] The transceiver may include a receiver and a transmitter ( Fig. 9 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Fig. 9 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0097] In addition, it should be noted that Fig. 9 The structure of the electronic device shown in the figure does not constitute a limitation on the device, and the actual device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they are not repeated here.

[0098] Third embodiment

[0099] This embodiment provides a computer-readable storage medium, which stores at least one instruction, and the instruction is loaded and executed by a processor to implement the method of the first embodiment. The computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein may be loaded by a processor in a terminal to execute the method.

[0100] In addition, it should be noted that the present invention can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present invention can be in the form of a full or partial hardware embodiment, a full or partial software embodiment or an embodiment combining software and hardware. Moreover, when implemented using software, the embodiment of the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more available media sets. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.

[0101] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0102] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0103] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. In addition, the term "and / or" is only an association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or plural.

[0104] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0105] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0106] In several embodiments provided by the present invention, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of functional modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0107] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0108] Finally, it should be noted that the above is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiment of the present invention has been described, for ordinary technicians in this technical field, once the basic creative concept of the present invention is known, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. Therefore, the attached claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. An intelligent classification method for grade based on ore image features, characterized in that: include: Constructing an ore image data set, dividing the ore image data set into a training set, a validation set and a test set, and performing preprocessing operations on the ore image data in the training set, the validation set and the test set respectively; Construct a multi-branch residual neural network model as an ore grade detection model; Training the ore grade detection model based on the ore image data set; Use the trained ore grade detection model to detect the grade of the ore to be tested; The ore grade detection model includes a backbone block, 5 IR_A-MBConv-block blocks, a first downsampling block, 10 IR_B-block blocks, a second downsampling block, 5 IR_C-block blocks, a 1×1 convolution block, a global average pooling layer and a fully connected layer connected in sequence; IR_A-MBConv-block, IR_B-block and IR_C-block are all multi-branch residual structures, each containing a trunk that maintains the original output and several branches; Branch 1 in the IR_A-MBConv-block is a 1×1 convolution block, branch 2 contains a 1×1 convolution block and an MBConv block connected in sequence, branch 3 contains a 1×1 convolution block and two MBConv blocks connected in sequence, and branch 4 contains a 3×3 average pooling layer with a stride of 1 and a padding of 1 and a 1×1 convolution block connected in sequence. The feature data output by each branch is concatenated according to the channel dimension, passed through a 1×1 convolution layer, and then residually connected to the backbone block before being input into the activation function; Compared with the IR_A-MBConv-block, the IR_B-block removes branch 2 in the IR_A-MBConv-block and adjusts the two MBConv blocks of branch 3 in the IR_A-MBConv-block to a 1×7 convolution block and a 7×1 convolution block; Compared with the IR_A-MBConv-block, the IR_C-block removes branch 2 in the IR_A-MBConv-block and adjusts the two MBConv blocks of branch 3 in the IR_A-MBConv-block into a 1×3 convolution block and a 3×1 convolution block; among them, the last of the five IR_C-blocks in the ore grade detection model does not use an activation function after the final 1×1 convolution layer.

2. The method for intelligent classification of grade based on ore image features according to claim 1, characterized in that: Constructing an ore image dataset, dividing the ore image dataset into a training set, a validation set, and a test set, including: Sample randomly distributed ores in the mine, take images of the ores, and measure their grade levels to build an ore image dataset with sufficient quantity, uniform grade distribution, and grade level labels; The ore image dataset is divided into three subsets, namely a training set, a validation set and a test set, in a ratio of 6:2:2 to ensure that each subset does not contain image data from other subsets. Image hashing technology is used to compare the image data in each subset, and a secondary check is performed to determine whether the same image data exists in each subset to avoid data contamination.

3. The method for intelligent classification of grade based on ore image features according to claim 1, characterized in that: The ore image data in the training set, validation set and test set are preprocessed separately, including: The ore image data in the training set are randomly scaled and cropped, and the ore image data are randomly transformed in brightness or contrast and randomly flipped horizontally with a probability of 80%; The ore image data in the validation set and test set are only scaled and center cropped.

4. The method for intelligent classification of grade based on ore image features according to claim 1, characterized in that: The trunk block includes three 3×3ConvBNAct blocks connected in sequence, a MaxPool layer with a stride of 2, a 1×1 convolution block, a 3×3MaxPool layer with a stride of 2, and multiple branches, wherein branch one in the trunk block is a 1×1 convolution block, branch two includes a 1×1 convolution block and a 5×5 convolution block connected in sequence, branch three includes a 1×1 convolution block and two MBConv blocks connected in sequence, branch four includes a 3×3AvgPool layer with a stride of 1 and a padding of 1 and a 1×1 convolution block connected in sequence, and the feature data output by each branch are finally feature spliced ​​according to the channel dimension.

5. The method for intelligent classification of grade based on ore image features according to claim 1, characterized in that: The first downsampling block and the second downsampling block both have a multi-branch structure; wherein, Branch 1 in the first downsampling block is a 1×1 convolution block, branch 2 contains a 1×1 convolution block and two MBConv blocks connected in sequence, and branch 3 is a 3×3 maximum pooling layer with a stride of 2 and a padding of 1. The feature data output by each branch are finally concatenated according to the channel dimension; Both branches 1 and 2 in the second downsampling block consist of a 1×1 convolution block and an MBConv block connected in sequence. Branch 3 contains a 1×1 convolution block and two MBConv blocks connected in sequence. Branch 4 is a 3×3 maximum pooling layer with a stride of 2 and a padding of 1. The feature data output by each branch are finally concatenated according to the channel dimension.

6. The method for intelligent classification of grade based on ore image features according to claim 5, characterized in that: The convolution block contains a convolution layer, a BN normalization layer, and an activation function connected in sequence; The MBconv block consists of a dilated convolutional layer with sequentially connected skip connections, a 3×3 depthwise separable convolutional layer, a compression and excitation module, and a pointwise convolutional layer; The activation function is the LeakyReLU function.

7. The method for intelligent classification of grade based on ore image features according to claim 1, characterized in that: The ore grade detection model is trained based on the ore image data set, including: Using the transfer learning strategy, the constructed ore grade detection model is first fully trained on the large dataset ImageNet-1k to obtain the training parameters; Rebuild an ore grade detection model with an ECA attention mechanism module; Initializing the reconstructed ore grade detection model using the training parameters, and then training the model and adjusting parameters using the preprocessed training set and validation set ore image data based on a learning rate descent strategy combined with a weight decay mechanism until the optimal situation is reached; The ore image data in the test set is input into the trained model to test the generalization and robustness of the model, and the grade of each ore image data in the test set is classified.

8. The method for intelligent classification of grade based on ore image features according to claim 7, characterized in that: The ECA attention mechanism module is added between the IR_C-block and the 1×1 convolution block.

Citation Information

Patent Citations

  • Graphite ore carbon grade prediction method based on deep learning

    CN116797905A