A lightweight copper ore sorting method based on dual-energy X-rays
The lightweight copper ore sorting method using dual-energy X-ray imaging and knowledge distillation strategies addresses computational limitations in mining environments, enhancing classification accuracy and reducing resource consumption.
Patent Information
- Application Number
- CN202510579302.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The real-time and resource consumption problems of deep learning models at the mine site have not been completely solved. The mining environment is complex and the computing power of on-site equipment is limited, which affects the degree of automation of copper ore sorting and the comprehensiveness of feature extraction.
Using a lightweight copper ore sorting method based on dual energy X-rays, a lightweight convolutional neural network is embedded in the CBAM module, combined with the ResNet-50 model and feature pyramid, a comprehensive knowledge distillation strategy is designed, including predicted distribution distillation, feature distillation and attention distillation, and optimizing model structure and parameters.
It improves the accuracy of copper ore sorting, alleviates the limitations of computing resources, enhances the real-time and adaptability of the model, and improves the degree of automation of copper ore sorting and the comprehensiveness of feature extraction.
Smart Images

Figure CN120105920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of copper ore separation, and specifically to a lightweight copper ore separation method based on dual-energy X-rays. Background Art
[0002] Ore resources are rich in variety and large in reserves. As an important strategic resource in China, copper resources are core resources in industries such as infrastructure construction, electronics and power, high-tech, and new energy. Like most other ores, copper ores have problems such as poor resource endowment, few high-grade copper ores, and mainly medium- and low-grade ores. At the same time, copper ores mostly exist in the form of polymetallic ores, symbiotic with associated minerals such as lead, zinc, iron, and molybdenum, and the mineral distribution is complex. Although it is of great significance for comprehensive development and utilization, it brings a not-small problem to the separation of ores.
[0003] Machine learning algorithms have achieved remarkable results in ore classification, but their way of relying on manual feature extraction limits the automation degree of the model and the comprehensiveness of feature extraction. To solve this problem, in recent years, the application of deep learning technology in ore classification has developed rapidly. Through a multi-layer neural network structure, deep learning models can automatically extract deep features of data and are suitable for classification tasks of large-scale and high-dimensional ore data. Convolutional Neural Network (CNN) is a representative algorithm in the field of deep learning. It performs excellently in the field of image classification and is widely used in the classification research of ore image data. In addition to convolutional neural networks, Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) based on time series data have also been gradually applied to ore classification research. In addition, the introduction of the Transformer structure in recent years has provided a new idea for ore classification. Due to its excellent performance in natural language processing and image classification, the Transformer model has been gradually applied to the classification of ore spectral data and multi-modal data. Research shows that the Transformer model can model complex multi-dimensional features in ore data, and its classification performance is better than traditional machine learning and convolutional neural network methods.
[0004] Although deep learning models have demonstrated powerful capabilities in ore classification, their practical applications still face certain challenges. The real-time performance and resource consumption problems of deep learning models in the mine site have not been completely solved. The mine environment is complex, and the computing power of on-site equipment is limited, while deep learning models usually require high-performance computing resource support. Therefore, researchers are working on developing lightweight deep learning models and ore classification systems combined with edge computing to meet the actual needs of the mine site. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a lightweight copper ore sorting method based on dual-energy X-rays, aiming to solve the problems in the background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: A lightweight copper ore sorting method based on dual-energy X-rays, comprising the following steps:
[0007] Step S1: Collect the dual-energy X-ray transmission images of copper ore, where the dual-energy X-ray transmission images of copper ore include high-energy X-ray transmission images and low-energy X-ray transmission images;
[0008] Step S2: Construct a lightweight convolutional neural network, embed the CBAM module in the lightweight convolutional neural network, and use the L2 regularization method to penalize the weights of the CBAM-Net model to obtain the CBAL2-Net model;
[0009] Step S3: Construct a ResNet-50 model, add a feature pyramid and a CBAM module to the ResNet-50 model to obtain a ResNet-50-FPN-CBAM model, and set a dimensionality reduction module and a knowledge aggregation module at the output end of the ResNet-50-FPN-CBAM model to obtain an improved ResNet-50 model;
[0010] Step S4: Take the CBAL2-Net model as the student model and the improved ResNet-50 model as the teacher model, and design a comprehensive knowledge distillation strategy to perform distillation operations on the student model and the teacher model to obtain a Res-CBAL2-Net model;
[0011] The comprehensive knowledge distillation strategy includes a prediction distribution distillation strategy, a feature distillation strategy, and an attention distillation strategy;
[0012] Step S5: Input the dual-energy X-ray transmission images of copper ore into the Res-CBAL2-Net model to sort the copper ore.
[0013] Further, the ResNet-50-FPN-CBAM model includes an initial stage STAGE0, a first stage STAGE1, a second stage STAGE2, a third stage STAGE3, a fourth stage STAGE4, a second CBAM module, a third CBAM module, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a dimensionality reduction module, and a knowledge aggregation module;
[0014] The processing flow of the ResNet-50-FPN-CBAM model is as follows: First, the input passes through the initial stage, the first stage, the second stage, the first CBAM module, the third stage, the second CBAM module, and the fourth stage in sequence. The output of the fourth stage passes through the first convolutional layer and the second convolutional layer in sequence to obtain the fifth fusion feature. The output of the third stage passes through the third convolutional layer. The output of the third convolutional layer is concatenated with the upsampled output of the first convolutional layer to obtain the first concatenated feature. The first concatenated feature is input into the fourth convolutional layer to obtain the fourth fusion feature. The output of the second stage passes through the fifth convolutional layer. The first concatenated feature after upsampling is concatenated with the output of the fifth convolutional layer to obtain the second concatenated feature. The second concatenated feature is input into the sixth convolutional layer to obtain the third fusion feature. The output of the first stage passes through the seventh convolutional layer. The second concatenated feature after upsampling is concatenated with the output of the seventh convolutional layer to obtain the third concatenated feature. The third concatenated feature is input into the eighth convolutional layer to obtain the second fusion feature. ;
[0015] Put 、 、 and into the dimensionality reduction module for dimensionality reduction to obtain the second dimensionality reduction feature , the third dimensionality reduction feature , the fourth dimensionality reduction feature and the fifth dimensionality reduction feature ;
[0016] Put 、 、 and into the knowledge aggregation module to obtain the second alignment feature , the third alignment feature , the fourth alignment feature , the fifth alignment feature ;
[0017] Concatenate the second alignment feature , the third alignment feature , the fourth alignment feature , the fifth alignment feature to obtain the concatenated feature ;
[0018] Compress the number of channels of and generate the attention weights for each feature , denotes the corresponding attention weight, Represent the corresponding attention weight, Represent the corresponding attention weight, Represent the corresponding attention weight;
[0019] Multiply with the corresponding attention weight to obtain the final fused feature which is the output of the ResNet-50-FPN-CBAM model.
[0020] Furthermore, the comprehensive knowledge distillation strategy includes:
[0021] Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution;
[0022] During the comprehensive knowledge distillation strategy, the knowledge generated by the teacher model is hierarchically transferred to the student model, and the student model generates corresponding outputs for alignment. At the same time, calculate the prediction distribution distillation loss of the student model , the classification task loss of the student model , the feature distillation loss and the attention distillation combined loss , and perform joint optimization;
[0023] The knowledge generated by the teacher model includes the original prediction values output by the teacher model, the final fused features output by the teacher model , the channel attention map and the spatial attention map generated by the teacher model;
[0024] Define the total loss of the comprehensive knowledge distillation strategy , expressed as:
[0025] ;
[0026] In the formula, , and are all hyperparameters, which are respectively used to balance the weights of the prediction distribution distillation strategy, the feature distillation strategy and the attention distillation strategy;
[0027] Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
[0028] Furthermore, the prediction distribution distillation strategy includes:
[0029] Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution;
[0030] Calculate the prediction probability distribution of the teacher model :
[0031] Let the final output of the teacher model be the original predicted values for two classes. After passing the original predicted values through function, the predicted probability distribution is obtained:
[0032] ;
[0033] In the formula, represents the final output of the teacher model; represents the original predicted value of the th class in ; represents the original predicted value of the th class in ; represents the number of classes;
[0034] Calculate the predicted probability distribution of the student model :
[0035] Let the final output of the student model be the original predicted values for two classes. After passing the original predicted values through the softmax function, the predicted probability distribution is obtained:
[0036] ;
[0037] In the formula, represents the final output of the student model; represents the original predicted value of the th class in ; represents the original predicted value of the
[0038] Set the prediction distribution distillation loss of the student model . The prediction distribution distillation loss measures the difference between the probability distribution of the student model output and the probability distribution of the teacher model output through the KL divergence, achieving the alignment of the prediction distributions;
[0039] ;
[0040] In the formula, represents the temperature parameter; and respectively represent the predicted probability distributions of the teacher model and the student model for the th class; Denote the KL divergence;
[0041] Set the classification task loss of the student model , and the classification task loss of the student model adopts the cross-entropy loss;
[0042] Set the total loss of the prediction distribution distillation strategy :
[0043] ;
[0044] In the formula, denotes the weight hyperparameter for adjusting the prediction distribution distillation intensity;
[0045] Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
[0046] Furthermore, the feature distillation strategy includes:
[0047] Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution;
[0048] Align the final fused features output by the teacher model with the features output by the pooling layer after the CBAM module in the student model ;
[0049] Design a feature distillation loss :
[0050] ;
[0051] Combine the feature distillation loss with the classification task loss of the student model to form the total loss of the feature distillation strategy :
[0052] ;
[0053] In the formula, denotes the weight hyperparameter for adjusting the feature distillation intensity;
[0054] Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
[0055] Furthermore, the attention distillation strategy includes two parts: channel attention distillation and spatial attention distillation. Channel attention distillation is used to align the channel attention distributions of the teacher model and the student model, and spatial attention distillation is used to align the spatial attention distributions of the teacher model and the student model. The specific process is as follows:
[0056] Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution;
[0057] Set the channel attention maps of the teacher model and the student model; there are two channel attention maps of the teacher model, which are generated by the channel attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 respectively, denoted as the second teacher channel attention map and the third teacher channel attention map ; the channel attention map of the student model is generated by the channel attention module in the first CBAM module of the student model, denoted as the student channel attention map ;
[0058] Set the spatial attention maps of the teacher model and the student model; there are two spatial attention maps of the teacher model, which are generated by the spatial attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 respectively, denoted as the second teacher spatial attention map and the third teacher spatial attention map ; the spatial attention map of the student model is generated by the spatial attention module in the first CBAM module of the student model, denoted as the student spatial attention map ;
[0059] Set the channel attention distillation loss , which measures the difference between the channel attention map of the teacher model and the channel attention map of the student model;
[0060] ;
[0061] Set the spatial attention distillation loss , which measures the difference between the spatial attention map of the teacher model and the spatial attention map of the student model;
[0062] ;
[0063] Set the attention distillation combined loss :
[0064] ;
[0065] wherein, represents the weight hyperparameter used to balance channel attention distillation and spatial attention distillation;
[0066] Combine the attention distillation combined loss with the classification task loss of the student model to define the total attention distillation strategy loss :
[0067] ;
[0068] In the formula, is a weight hyperparameter for adjusting the intensity of the attention distillation loss;
[0069] The Adam optimizer is used, and the cosine annealing learning rate adjustment strategy is adopted.
[0070] Furthermore, the first CBAM module includes a channel attention module and a spatial attention module. The input first passes through the channel attention module to generate channel attention weights, and the channel attention weights are applied to the input of the first CBAM module to obtain weighted features. The weighted features are passed through the spatial attention module to generate spatial attention weights, and the spatial attention weights are applied to the weighted features to obtain the output of the first CBAM module;
[0071] The structures of the second CBAM module and the third CBAM module are the same as that of the first CBAM module.
[0072] Furthermore, the specific process of step S2 is as follows: A lightweight convolutional neural network with a simple structure is selected as the base model by the method of selecting the best through multiple experiments. The first CBAM module is embedded in different layers of the base model, and the optimal embedding position and number of the first CBAM module are determined through experiments to obtain the CBAM-Net model. Finally, the L2 regularization method is used to penalize the weights of the CBAM-Net model to obtain the CBAL2-Net model;
[0073] Among them, the CBAM-Net model is composed of an input layer, a feature extraction layer, a batch normalization layer, a first CBAM module, a pooling layer, a fully connected layer, and an output layer connected in sequence.
[0074] Compared with the existing technologies, the present invention has the following beneficial effects:
[0075] (1) The present invention uses the constructed CBAL2-Net model as the student model and the constructed improved ResNet-50 model as the teacher model, and designs a comprehensive strategy that integrates three different distillation strategies for knowledge transfer of the model to obtain the target model Res-CBAL2-Net. Compared with the existing mainstream models, the accuracy of copper ore sorting is greatly improved.
[0076] (2) The improved ResNet-50 model constructed in the present invention uses ResNet-50 as the base model, adds a feature pyramid structure to ResNet-50, enabling ResNet-50 to output feature maps of multiple scales, alleviating the negative effect of inconsistent intermediate feature map sizes on distillation; secondly, the CBAM module is used to generate attention maps in subsequent attention distillation; finally, the multi-scale feature maps are fused into a single-scale feature map with multi-scale information through a dimensionality reduction module and a knowledge aggregation module, facilitating subsequent feature distillation.
[0077] (3) The CBAL2-Net model constructed in the present invention self-built a simple original model through the method of model structure design, added the lightweight module CBAM to the model, significantly improved the classification accuracy, and further improved the classification accuracy while alleviating overfitting by using L2 regularization. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is the flowchart of the method of the present invention.
[0079] Figure 2 It is the structure diagram of the improved ResNet-50 model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] As Figure 1 shown, the present invention provides a technical solution: a lightweight copper ore sorting method based on dual-energy X-ray, including the following steps:
[0081] Step S1: Collect the dual-energy X-ray transmission images of copper ore, where the dual-energy X-ray transmission images of copper ore include high-energy X-ray transmission images and low-energy X-ray transmission images.
[0082] Step S2: Construct a lightweight convolutional neural network, embed the CBAM module in the lightweight convolutional neural network, and use the L2 regularization method to penalize the weights of the CBAM-Net model to obtain the CBAL2-Net model.
[0083] Step S3: Construct a ResNet-50 model, add a feature pyramid and a CBAM module to the ResNet-50 model to obtain the ResNet-50-FPN-CBAM model, and set a dimensionality reduction module and a knowledge aggregation module at the output end of the ResNet-50-FPN-CBAM model to obtain the improved ResNet-50 model.
[0084] Step S4: Use the CBAL2-Net model as the student model and the improved ResNet-50 model as the teacher model, and design a comprehensive knowledge distillation strategy to perform distillation operations on the student model and the teacher model to obtain the Res-CBAL2-Net model.
[0085] The comprehensive knowledge distillation strategy includes the prediction distribution distillation strategy, the feature distillation strategy, and the attention distillation strategy. The comprehensive knowledge distillation strategy integrates the prediction distribution distillation strategy, the feature distillation strategy, and the attention distillation strategy into one training process. By simultaneously transmitting the output distribution, feature representation, and attention mechanism of the teacher model, it comprehensively guides the learning process of the student model.
[0086] Step S5: Input the dual-energy X-ray transmission image of copper ore into the Res-CBAL2-Net model to sort the copper ore.
[0087] Among them, the specific process of step S2 is as follows: Select a lightweight convolutional neural network with a simple structure as the basic model through the method of selecting the best through multiple experiments. Embed the first CBAM module in different layers of the basic model, and determine the optimal embedding position and number of the first CBAM module through experiments to obtain the CBAM-Net model. Finally, use the L2 regularization method to penalize the weights of the CBAM-Net model to alleviate overfitting and improve generalization ability, and obtain the CBAL2-Net model.
[0088] Among them, the CBAM-Net model consists of an input layer, a feature extraction layer, a batch normalization layer, the first CBAM module, a pooling layer, a fully connected layer, and an output layer connected in sequence.
[0089] The CBAL2-Net model improves the classification accuracy and largely alleviates the overfitting phenomenon by adding the CBAM attention mechanism and L2 regularization to the custom basic model, and also solves the problem of gradient explosion or gradient disappearance during the training process through the use of batch normalization operation, effectively improving the model performance.
[0090] Among them, the first CBAM module includes a channel attention module and a spatial attention module. The input first passes through the channel attention module to generate channel attention weights, applies the channel attention weights to the input to obtain weighted features, passes the weighted features through the spatial attention module to generate spatial attention weights, and applies the spatial attention weights to the weighted features to obtain the output of the first CBAM module.
[0091] Among them, construct the improved ResNet-50 model:
[0092] Since the standard input size of the ResNet-50 model is 224×224, while the input size of the CBAL2-Net model is 96×96, this resolution difference results in a significant scale mismatch of feature maps in the spatial dimension. Although knowledge distillation can be performed operationally, the distillation effect will be greatly reduced or even have a counterproductive effect. In order to enable the teacher model to provide multi-scale feature representations more suitable for the student model to learn, it is necessary for the teacher model to learn the features of the same picture at different scales. Therefore, it is necessary to add a multi-scale feature extraction structure to the ResNet-50 model.
[0093] The Feature Pyramid Network (FPN) is a technique that focuses on enhancing multi-scale feature representation and can integrate two techniques, multi-scale feature extraction and feature fusion, on the basis of the ResNet-50 model.
[0094] The overall architecture of the ResNet-50 model includes an initial stage (STAGE0) for extracting low-level features, followed by four stages: the first stage (STAGE1), the second stage (STAGE2), the third stage (STAGE3), and the fourth stage (STAGE4), each consisting of residual blocks.
[0095] As Figure 2 shown, the ResNet-50-FPN-CBAM model includes the initial stage (STAGE0), the first stage (STAGE1), the second stage (STAGE2), the third stage (STAGE3), the fourth stage (STAGE4), the second CBAM module, the third CBAM module, the first convolutional layer (1×1), the second convolutional layer (1×1), the third convolutional layer (1×1), the fourth convolutional layer (1×1), the fifth convolutional layer (1×1), the sixth convolutional layer (1×1), the seventh convolutional layer (1×1), and the eighth convolutional layer (1×1), a dimensionality reduction module, and a knowledge aggregation module.
[0096] Among them, the second CBAM module, the third CBAM module, and the first CBAM module have the same structure and will not be elaborated here.
[0097] The processing flow of the ResNet-50-FPN-CBAM model is as follows: First, the input passes through the initial stage, the first stage, the second stage, the first CBAM module, the third stage, the second CBAM module, and the fourth stage in sequence. The output of the fourth stage passes through the first convolutional layer and the second convolutional layer in sequence to obtain the fifth fused feature. , the output of the third stage is passed through the third convolutional layer, and the output of the third convolutional layer is concatenated with the upsampled output of the first convolutional layer to obtain the first concatenated feature. The first concatenated feature is input into the fourth convolutional layer to obtain the fourth fused feature. , the output of the second stage is passed through the fifth convolutional layer, and the first concatenated feature is upsampled and then concatenated with the output of the fifth convolutional layer to obtain the second concatenated feature. The second concatenated feature is input into the sixth convolutional layer to obtain the third fused feature. , the output of the first stage is passed through the seventh convolutional layer, and the second concatenated feature is upsampled and then concatenated with the output of the seventh convolutional layer to obtain the third concatenated feature. The third concatenated feature is input into the eighth convolutional layer to obtain the second fused feature. .
[0098] The , , and are input into the dimensionality reduction module for dimensionality reduction to obtain the second dimensionality reduction feature , the third dimensionality reduction feature , the fourth dimensionality reduction feature and the fifth dimensionality reduction feature ; the specific operation of dimensionality reduction is to pass , , and through a 1×1 convolution for processing, then perform batch normalization, and finally pass through the ReLU activation function to obtain , , and .
[0099] The main purpose of the dimensionality reduction module is to simplify the multi-scale features output by the teacher model to make it more adaptable to the learning ability of the student model. However, the dimensionality reduction module does not solve the problem of fusing multi-scale features.
[0100] The , , and are input into the knowledge aggregation module. First, , , and are aligned in resolution to obtain the aligned features, denoted as:
[0101] (1);
[0102] (2);
[0103] (3);
[0104] (4);
[0105] In the formula, represents the second alignment feature; represents the third alignment feature; represents the fourth alignment feature; represents the fifth alignment feature; represents the result after resolution alignment; represents the result after resolution alignment; represents the result after resolution alignment; represents the result after resolution alignment; is the target resolution; represents an upsampling or downsampling operation.
[0106] Concatenate the alignment features:
[0107] (5);
[0108] In the formula, represents the feature obtained after concatenation; represents the concatenation operation in the channel dimension.
[0109] Compress the number of channels of through a convolutional layer and generate the attention weights for each feature :
[0110] (6);
[0111] In the formula, represents the generated attention weight, represents the corresponding attention weight, represents the corresponding attention weight, represents the corresponding attention weight, represents the corresponding attention weight; represents function.
[0112] Multiply by the corresponding attention weight to obtain the final fused feature That is, the output of the ResNet-50-FPN-CBAM model, expressed as:
[0113] (7).
[0114] Based on the specific structures of the teacher model and the student model, an efficient prediction distribution distillation strategy was designed and implemented, aiming to use the output distribution of the teacher model to guide the student model to learn richer category information; the prediction distribution distillation strategy includes:
[0115] 1. Set the model input: The teacher model and the student model respectively receive dual-energy X-ray transmission images of copper ore with input resolutions of 224×224 and 96×96. Here, for sufficient training, to ensure the consistency of the input distribution.
[0116] 2. Calculate the predicted probability distribution of the teacher model :
[0117] Let the final output of the teacher model be the raw predicted values (logits) for two categories. After passing the raw predicted values through the function, the predicted probability distribution is obtained:
[0118] (8);
[0119] In the formula, represents the final output of the teacher model; represents the raw predicted value of the th category in ; the raw predicted value of the th category in represents the number of categories.
[0120] 3. Calculate the predicted probability distribution of the student model :
[0121] Let the final output of the student model be the raw predicted values (logits) for two categories. After passing the raw predicted values through the softmax function, the predicted probability distribution is obtained:
[0122] (9);
[0123] In the formula, represents the final output of the student model; represents the raw predicted value of the th category in represents The original predicted value of the th category.
[0124] 4. Set the prediction distribution distillation loss of the student model , and the prediction distribution distillation loss measures the difference between the probability distribution output by the student model and the probability distribution output by the teacher model through the Kullback-Leibler (KL) divergence to achieve the alignment of the prediction distributions;
[0125] (10);
[0126] In the formula, represents the temperature parameter; and respectively represent the predicted probability distributions of the teacher model and the student model for the th category; represents the KL divergence.
[0127] Among them, the temperature parameter is the key to prediction distribution distillation. When , the predicted probability distribution obtained after passing the original predicted value through the softmax function will be smoother, and the relative relationship between categories will be more significant, which helps the student model better learn the knowledge of the teacher model. In this embodiment, is selected as the optimal setting.
[0128] 5. Set the classification task loss of the student model , and the classification task loss of the student model uses cross-entropy loss.
[0129] 6. Set the total loss of the prediction distribution distillation strategy :
[0130] (11);
[0131] In the formula, represents the weight hyperparameter that adjusts the prediction distribution distillation intensity.
[0132] 7. Use the Adam optimizer with an initial learning rate of 0.001 and adopt the cosine annealing learning rate adjustment strategy (cosineannealing).
[0133] Among them, the feature distillation strategy includes:
[0134] 1. Set the model input: The teacher model and the student model respectively receive dual-energy X-ray transmission images of copper ore with input resolutions of 224×224 and 96×96.
[0135] 2. For the final fused features output by the teacher model Align with the features output by the pooling layer after the CBAM module in the student model to help the student model optimize the feature representation of the intermediate layer (the CBAM module in the student model).
[0136] 3. The goal of the feature distillation strategy is to make the features in the student model as close as possible to the final fused features output by the teacher model ; To achieve this goal, a feature distillation loss is designed :
[0137] (12).
[0138] 4. In order to balance the classification ability and feature learning ability of the student model during the feature distillation strategy, the feature distillation loss needs to be combined with the classification task loss of the student model to form the total loss of the feature distillation strategy :
[0139] (13);
[0140] wherein, represents the weight hyperparameter for adjusting the feature distillation intensity, set to 1.
[0141] 5. Use the Adam optimizer with an initial learning rate of 0.001 and adopt the cosine annealing learning rate adjustment strategy (cosineannealing).
[0142] Among them, the attention distillation strategy includes:
[0143] The attention distillation strategy helps the student model learn the attention distribution of the teacher model by minimizing the difference between the attention maps generated by the student model and the teacher model. The attention distillation strategy is divided into two parts: channel attention distillation and spatial attention distillation. Channel attention distillation is used to align the channel attention distributions of the teacher model and the student model, and spatial attention distillation is used to align the spatial attention distributions of the teacher model and the student model.
[0144] 1. Set the model input: The teacher model and the student model respectively receive dual-energy X-ray transmission images of copper ore with input resolutions of 224×224 and 96×96.
[0145] 2. Set the channel attention map of the teacher model and the channel attention map of the student model; The teacher model has two channel attention maps, which are respectively generated by the channel attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 (i.e., the generated channel attention weights), denoted as the second teacher channel attention map and the third teacher channel attention map ; The channel attention map of the student model is generated by the channel attention module in the CBAM module in the student model, denoted as the student channel attention map .
[0146] 3. Set the spatial attention map of the teacher model and the spatial attention map of the student model; There are two spatial attention maps of the teacher model, which are respectively generated by the spatial attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 (i.e., the generated spatial attention weights), denoted as the second teacher spatial attention map and the third teacher spatial attention map ; The spatial attention map of the student model is generated by the spatial attention module in the first CBAM module in the student model, denoted as the student spatial attention map .
[0147] 4. Set the channel attention distillation loss , which measures the difference between the channel attention map of the teacher model and the channel attention map of the student model;
[0148] (14).
[0149] 5. Set the spatial attention distillation loss , which measures the difference between the spatial attention map of the teacher model and the spatial attention map of the student model;
[0150] (15).
[0151] 6. Set the attention distillation combined loss :
[0152] (16);
[0153] In the formula, represents the weight hyperparameter used to balance channel attention distillation and spatial attention distillation.
[0154] 7. In order to balance the accuracy of the classification task and the attention distillation effect during the attention distillation strategy process, combine the attention distillation combined loss with the classification task loss of the student model to define the total attention distillation strategy loss :
[0155] (17);
[0156] In the formula, is the weight hyperparameter used to adjust the intensity of the attention distillation loss, and is set to 1.
[0157] 8. Use the Adam optimizer with an initial learning rate of 0.001 and adopt the cosine annealing learning rate adjustment strategy (cosineannealing).
[0158] Among them, the comprehensive knowledge distillation strategy includes:
[0159] 1. Set the model input: The teacher model and the student model respectively receive dual-energy X-ray transmission images of copper ore with input resolutions of 224×224 and 96×96.
[0160] 2. During the comprehensive knowledge distillation strategy, the knowledge generated by the teacher model (the original predicted values (logits) output by the teacher model, the final fused features output by the teacher model , the channel attention map and the spatial attention map generated by the teacher model) are hierarchically transmitted to the student model. The student model generates corresponding outputs for alignment. At the same time, calculate the total loss of the prediction distribution distillation strategy, the total loss of the feature distillation strategy, and the total loss of the attention distillation strategy, and perform joint optimization.
[0161] 3. Define the total loss of the comprehensive knowledge distillation strategy :
[0162] (18);
[0163] In the formula, , and are all hyperparameters, which are used to balance the weights of the prediction distribution distillation strategy, the feature distillation strategy, and the attention distillation strategy respectively. = = = 1.
[0164] 4. Use the Adam optimizer with an initial learning rate of 0.001 and adopt the cosine annealing learning rate adjustment strategy (cosineannealing).
[0165] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A lightweight copper ore sorting method based on dual-energy X-ray, characterized in that It includes the following steps: Step S1: Collect the dual-energy X-ray transmission images of copper ore, where the dual-energy X-ray transmission images of copper ore include high-energy X-ray transmission images and low-energy X-ray transmission images; Step S2: Construct a lightweight convolutional neural network, embed the CBAM module in the lightweight convolutional neural network, and use the L2 regularization method to penalize the weights of the CBAM-Net model to obtain the CBAL2-Net model; Step S3: Construct a ResNet-50 model, add a feature pyramid and a CBAM module to the ResNet-50 model to obtain the ResNet-50-FPN-CBAM model, and set a dimensionality reduction module and a knowledge aggregation module at the output end of the ResNet-50-FPN-CBAM model to obtain an improved ResNet-50 model; Step S4: Use the CBAL2-Net model as the student model and the improved ResNet-50 model as the teacher model, and design a comprehensive knowledge distillation strategy to perform distillation operations on the student model and the teacher model to obtain the Res-CBAL2-Net model; The comprehensive knowledge distillation strategy includes a prediction distribution distillation strategy, a feature distillation strategy, and an attention distillation strategy; Step S5: Input the dual-energy X-ray transmission images of copper ore into the Res-CBAL2-Net model to sort the copper ore.
2. The lightweight copper ore sorting method based on dual-energy X-ray according to claim 1, characterized in that: The ResNet-50-FPN-CBAM model includes an initial stage STAGE0, a first stage STAGE1, a second stage STAGE2, a third stage STAGE3, a fourth stage STAGE4, a second CBAM module, a third CBAM module, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, and an eighth convolutional layer, a dimensionality reduction module, and a knowledge aggregation module; The processing flow of the ResNet-50-FPN-CBAM model is as follows: First, the input passes through the initial stage, the first stage, the second stage, the first CBAM module, the third stage, the second CBAM module, and the fourth stage in sequence. The output of the fourth stage passes through the first convolutional layer and the second convolutional layer in sequence to obtain the fifth fusion feature. , the output of the third stage passes through the third convolutional layer, and the output of the third convolutional layer is concatenated with the upsampled output of the first convolutional layer to obtain the first concatenated feature. The first concatenated feature is input into the fourth convolutional layer to obtain the fourth fusion feature. , the output of the second stage passes through the fifth convolutional layer, and the first concatenated feature after upsampling is concatenated with the output of the fifth convolutional layer to obtain the second concatenated feature. The second concatenated feature is input into the sixth convolutional layer to obtain the third fusion feature. , the output of the first stage passes through the seventh convolutional layer, and the second concatenated feature after upsampling is concatenated with the output of the seventh convolutional layer to obtain the third concatenated feature. The third concatenated feature is input into the eighth convolutional layer to obtain the second fusion feature. ; Input , , and into the dimensionality reduction module for dimensionality reduction to obtain the second dimensionality reduction feature , the third dimensionality reduction feature , the fourth dimensionality reduction feature and the fifth dimensionality reduction feature ; Input , , and into the knowledge aggregation module to obtain the second alignment feature , the third alignment feature , the fourth alignment feature , the fifth alignment feature ; For the second alignment feature , the third alignment feature , the fourth alignment feature , the fifth alignment feature are spliced to obtain a spliced feature ; Pair Compress the number of channels and generate the attention weights for each feature , Denote the corresponding attention weight, Denote the corresponding attention weight, Denote the corresponding attention weight, Denote the corresponding attention weight; Multiply with the corresponding attention weights to obtain the final fused feature which is the output of the ResNet-50-FPN-CBAM model.
3. A lightweight copper ore sorting method based on dual-energy X-ray according to claim 2, characterized in that: The comprehensive knowledge distillation strategy includes: Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution; During the comprehensive knowledge distillation strategy process, the knowledge generated by the teacher model is hierarchically transferred to the student model, and the student model generates corresponding outputs for alignment. Meanwhile, the distillation loss of the prediction distribution of the student model, the classification task loss of the student model, the feature distillation loss, and the attention distillation combined loss are jointly optimized. The knowledge generated by the teacher model includes the original predicted values output by the teacher model, the final fused features output by the teacher model , the channel attention map and spatial attention map generated by the teacher model; Define the total loss of the comprehensive knowledge distillation strategy , denoted as: ; wherein, , and are all hyperparameters, which are respectively used to balance the weights of the prediction distribution distillation strategy, the feature distillation strategy and the attention distillation strategy; Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
4. A lightweight copper ore sorting method based on dual-energy X-rays according to claim 3, characterized in that: The prediction distribution distillation strategy includes: Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution; Calculate the predicted probability distribution of the teacher model : Let the final output of the teacher model be the raw prediction values for two categories. After passing the raw prediction values through the function, we obtain the predicted probability distribution : ; In the formula, represents the final output of the teacher model; represents the original predicted value of the th category in represents the original predicted value of the th category in represents the number of categories; represents function; Calculate the predicted probability distribution of the student model : Let the final output of the student model be the raw prediction values for two categories. After passing the raw prediction values through the softmax function, the predicted probability distribution is obtained : ; In the formula, represents the final output of the student model; represents the original predicted value of the th category in represents the original predicted value of the th category in Set the prediction distribution distillation loss of the student model , and the prediction distribution distillation loss measures the difference between the probability distributions output by the student model and the teacher model through the KL divergence to achieve the alignment of the prediction distributions; ; In the formula, represents the temperature parameter; and respectively represent the predicted probability distributions of the teacher model and the student model for the th category; represents the KL divergence; Set the classification task loss of the student model , and the classification task loss of the student model uses cross-entropy loss; Set the total loss of the prediction distribution distillation strategy : ; wherein, denotes a weight hyperparameter for adjusting the intensity of predictive distribution distillation; Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
5. A lightweight copper ore sorting method based on dual-energy X-rays according to claim 4, characterized in that: The feature distillation strategy includes: Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution; The final fused features output by the teacher model are aligned with the features output by the pooling layer after the CBAM module in the student model ; Design a feature distillation loss : ; Combine the feature distillation loss with the classification task loss of the student model to form the total loss of the feature distillation strategy : ; In the formula, represents the weight hyperparameter for adjusting the feature distillation intensity; Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
6. The lightweight copper ore sorting method based on dual-energy X-rays according to claim 5, characterized in that: The attention distillation strategy includes two parts: channel attention distillation and spatial attention distillation. Channel attention distillation is used to align the channel attention distributions of the teacher model and the student model, and spatial attention distillation is used to align the spatial attention distributions of the teacher model and the student model. The specific process is as follows: Set the model input: The teacher model and the student model respectively receive the dual-energy X-ray transmission images of copper ore with a preset resolution; Set the channel attention maps of the teacher model and the student model; there are two channel attention maps of the teacher model, which are generated by the channel attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 respectively, denoted as the second teacher channel attention map and the third teacher channel attention map ; the channel attention map of the student model is generated by the channel attention module in the first CBAM module in the student model, denoted as the student channel attention map ; Set the spatial attention maps of the teacher model and the student model; there are two spatial attention maps of the teacher model, which are generated by the spatial attention modules in the second CBAM module connected after STAGE2 and the third CBAM module connected after STAGE3 respectively, denoted as the second teacher spatial attention map and the third teacher spatial attention map ; the spatial attention map of the student model is generated by the spatial attention module in the first CBAM module in the student model, denoted as the student spatial attention map ; Set the channel attention distillation loss , which measures the difference between the channel attention map of the teacher model and that of the student model; ; Set the spatial attention distillation loss , which measures the difference between the spatial attention map of the teacher model and that of the student model; ; Set the attention distillation combined loss : ; In the formula, represents the weight hyperparameter for balancing channel attention distillation and spatial attention distillation; Combine the attention distillation combined loss with the classification task loss of the student model to define the total loss of the attention distillation strategy as follows: ; In the formula, is a weight hyperparameter used to adjust the intensity of the attention distillation loss; Use the Adam optimizer and adopt the cosine annealing learning rate adjustment strategy.
7. A lightweight copper ore sorting method based on dual-energy X-rays according to claim 6, characterized in that: The first CBAM module includes a channel attention module and a spatial attention module. The input first passes through the channel attention module to generate channel attention weights, which are applied to the input of the first CBAM module to obtain weighted features. The weighted features are then passed through the spatial attention module to generate spatial attention weights, which are applied to the weighted features to obtain the output of the first CBAM module; The structures of the second CBAM module and the third CBAM module are the same as that of the first CBAM module.
8. The lightweight copper ore sorting method based on dual-energy X-ray according to claim 7, characterized in that: The specific process of step S2 is as follows: A lightweight convolutional neural network is selected as the basic model through the method of selecting the best one through multiple experiments. The first CBAM module is embedded in different layers of the basic model, and the best embedding position and number of the first CBAM module are determined through experiments to obtain the CBAM-Net model. Finally, the L2 regularization method is used to penalize the weights of the CBAM-Net model to obtain the CBAL2-Net model; Among them, the CBAM-Net model consists of an input layer, a feature extraction layer, a batch normalization layer, a first CBAM module, a pooling layer, a fully connected layer, and an output layer connected in sequence.
Citation Information
Patent Citations
Ore sorting model training method and system based on multi-scale feature fusion
CN116597258A
Lightweight metal surface defect detection method based on double-source knowledge distillation
CN117540779A