Iron ore image recognition method based on improved ShufflenetV2 network
The iron ore recognition model is constructed through the improved ShufflenetV2 network, which solves the problems of manual dependence and high equipment costs in traditional identification methods, and achieves rapid and accurate iron ore classification, improves production efficiency and reduces costs.
Patent Information
- Application Number
- CN202510330959.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional iron ore recognition methods rely on artificial vision and are susceptible to artificial experience, and modern recognition equipment is costly and unstable, making it difficult to meet the needs of fast and real-time production.
The improved ShufflenetV2 network is adopted to build an iron ore recognition model through an attention mechanism, combining data augmentation and feature transformation to improve the robustness and accuracy of the model.
It realizes accurate and rapid identification of iron ore species, improves production efficiency and saves production costs.
Smart Images

Figure CN120259750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of iron ore identification, and particularly to an iron ore image recognition method based on an improved ShuffleNetV2 network. Background Art
[0002] The iron ore sorting process is cumbersome, and the technological process makes the entire sorting process progress slowly, which is not suitable for fast and real-time production situations. Traditional iron ore identification mainly relies on manual visual identification, which is easily affected by subjective factors of manual experience. Although the current iron ore identification methods have made significant progress compared with traditional methods, there are still some technical deficiencies. Many modern iron ore identification methods, such as X-ray fluorescence spectroscopy (XRF), X-ray diffraction (XRD), laser-induced breakdown spectroscopy (LIBS), etc., although they can provide high-precision mineral composition analysis, have high equipment costs and require specialized technical personnel for operation and maintenance; moreover, in the mine environment, there may be dust, humidity changes, and temperature fluctuations, which can affect the sample surface due to oxidation, weathering, or pollution, resulting in unstable analysis results. Therefore, the present invention proposes an iron ore image recognition method based on an improved ShuffleNetV2 network. Summary of the Invention
[0003] The purpose of the present invention is to provide an iron ore image recognition method based on an improved ShuffleNetV2 network to achieve real-time and fast iron ore identification and classification for the above technical problems.
[0004] To achieve the above purpose, the present invention provides the following solution:
[0005] An iron ore image recognition method based on an improved ShuffleNetV2 network, comprising:
[0006] Obtain an iron ore image to be detected;
[0007] Input the iron ore image to be detected into a preset iron ore recognition model, and output an iron ore category recognition result, wherein the iron ore recognition model is obtained by training with a training set, the training set includes a number of sample pictures of hematite, magnetite, siderite, and chlorite, and the iron ore recognition model is constructed using a ShuffleNetV2 network improved based on an attention mechanism.
[0008] Optionally, before training the iron ore recognition model with the training set, it further includes expanding the training set, and the expansion includes:
[0009] Construct a first training set based on the sample pictures of hematite, magnetite, siderite, and chlorite;
[0010] Perform geometric transformations on the first training set, and add the sample images after geometric transformation to the first training set to construct a second training set, where the geometric transformations include flipping and rotation.
[0011] Perform color transformation on the second training set, and add the sample images after color transformation to the second training set to construct a third training set, where the color transformation is performed by adding salt-and-pepper noise.
[0012] Optionally, the iron ore recognition model includes n processing modules connected in sequence, which are used to extract information layer by layer from the input image and output the recognition result.
[0013] Optionally, the n processing modules include a first processing module, a second processing module, a third processing module, a fourth processing module, and a fifth processing module. Among them, the first processing module includes a 3×3 convolutional layer and a 3×3 depthwise separable convolutional layer connected in sequence; the second processing module, the third processing module, and the fourth processing module all include a ShuffleNetV2 Unit2 and a ShuffleNetV2 Unit1 improved by the attention mechanism and connected in sequence; the fifth processing module includes a 1×1 convolutional layer, a global pooling layer, and a fully connected layer connected in sequence.
[0014] Optionally, improving the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 by the attention mechanism includes:
[0015] Add a hybrid attention layer after the Channel Shuffle layer of the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1, where the hybrid attention layer uses the SK attention mechanism and the ECA attention mechanism in series fusion.
[0016] Optionally, the 3×3 convolutional layer, the 1×1 convolutional layer, and the convolutional layers in the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 all use the H-Swish activation function.
[0017] Optionally, before inputting the iron ore image to be detected into the iron ore recognition model, it also includes normalizing and cropping the iron ore image to be detected.
[0018] The beneficial effects of the present invention are:
[0019] The present invention constructs an iron ore recognition model through the ShuffleNetV2 network improved by the attention mechanism. The model can accurately, quickly and real-time predict the types of iron ore, saving production costs while improving production efficiency. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 It is the original ShuffleNetV2 network structure diagram of the embodiment of the present invention;
[0022] Figure 2 It is the SE-ShuffleNetV2 network structure diagram with the introduction of the fusion attention mechanism of the embodiment of the present invention;
[0023] Figure 3 It is the improved SEH-ShuffleNetV2s network structure diagram of the embodiment of the present invention;
[0024] Figure 4 It is the flowchart of the iron ore image recognition method based on the improved ShufflenetV2 network of the embodiment of the present invention;
[0025] Figure 5 It is the recognition effect diagram of magnetite of the embodiment of the present invention;
[0026] Figure 6 It is the recognition effect diagram of hematite of the embodiment of the present invention;
[0027] Figure 7 It is the recognition effect diagram of siderite of the embodiment of the present invention;
[0028] Figure 8 It is the recognition effect diagram of chlorite of the embodiment of the present invention;
[0029] Figure 9 It is the accuracy comparison curve graph of the original ShuffleNetV2 network model, SE-ShuffleNetV2 network model and SEH-ShuffleNetV2s network model. Detailed implementation manners
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0031] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] This embodiment provides an iron ore image recognition method based on an improved ShuffleNetV2 network, as Figure 4 shown, including:
[0033] Obtain the iron ore image to be detected;
[0034] Input the iron ore image to be detected into a preset iron ore recognition model, and output the iron ore category recognition result. Among them, the iron ore recognition model is obtained by training with a training set, the training set includes a number of sample pictures of hematite, magnetite, siderite, and chlorite, and the iron ore recognition model is constructed using a ShuffleNetV2 network improved based on the attention mechanism.
[0035] Specifically, in this embodiment, an iron ore recognition model is constructed by a ShuffleNetV2 network improved by the attention mechanism. The model can accurately, quickly, and real-time predict the types of iron ore, saving production costs while improving production efficiency.
[0036] Furthermore, training set preparation:
[0037] Before training the iron ore recognition model with the training set, it also includes expanding the training set. The expansion includes:
[0038] Construct a first training set based on the sample pictures of hematite, magnetite, siderite, and chlorite;
[0039] Perform geometric transformation on the first training set, and add the geometrically transformed sample images to the first training set to construct a second training set. Among them, the geometric transformation includes flipping and rotation.
[0040] Perform color transformation on the second training set, and add the color-transformed sample images to the second training set to construct a third training set. Among them, the color transformation is performed by adding salt and pepper noise.
[0041] Specifically, in this embodiment, sample pictures of hematite, magnetite, siderite, and chlorite are used, and following the principle of diversity of different iron ore characteristics, as much sample data as possible is selected. To improve the accuracy of the network, the dataset is expanded and divided into a training set and a test set. The traditional data expansion method is used to expand the dataset. The specific method is as follows:
[0042] 1. Geometric transformation:
[0043] 1) Flip: Mirror and swap the left and right parts of the image with the vertical central axis of the image as the center. Assume the original image has a height of h and a width of w, and a certain pixel point P(x, y) in the original image becomes P(w - 1 - x, y) after horizontal transformation;
[0044] 2) Rotation: Rotate by 45°, 90°, and 135° respectively on the axis.
[0045] 2. Color transformation: Use salt-and-pepper noise to perform color transformation on the image. Adding an appropriate amount of noise can enhance the learning ability of the model. The specific method is as follows:
[0046] 1) Randomly select the signal-to-noise ratio SNR between [0, 1]. In the present invention, SNR = 0.5 is selected;
[0047] 2) Calculate the total number of pixels SP, and obtain the number of pixels NP to be added with noise, NP = SP * (1 - SNR);
[0048] 3) Randomly obtain the pixel positions P(i, j) to be added with noise;
[0049] 4) Specify the pixel value as 255;
[0050] 5) Repeat steps 3) and 4) until all NP pixels are completed;
[0051] 6) Output the image after adding noise.
[0052] In this embodiment, the addition of salt-and-pepper noise randomly generates black and white pixel points on the image, making the image more robust.
[0053] Furthermore, model construction:
[0054] The iron ore recognition model includes n processing modules connected in sequence, which are used to extract information layer by layer from the input image and output the recognition result.
[0055] Specifically, it includes a first processing module, a second processing module, a third processing module, a fourth processing module, and a fifth processing module. Among them, the first processing module includes a 3×3 convolutional layer and a 3×3 depthwise separable convolutional layer connected in sequence; the second processing module, the third processing module, and the fourth processing module all include the improved ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 with an attention mechanism connected in sequence; the fifth processing module includes a 1×1 convolutional layer, a global pooling layer, and a fully connected layer.
[0056] Improving the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 with an attention mechanism includes:
[0057] A hybrid attention layer is added after the Channel Shuffle layers of the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1, where the hybrid attention layer adopts the cascaded fusion of the SK attention mechanism and the ECA attention mechanism.
[0058] The 3×3 convolutional layer, the 1×1 convolutional layer, and the convolutional layers in the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 all adopt the H-Swish activation function.
[0059] Specifically, as Figure 1 shown, in this embodiment, the ShuffleNetV2 network is adopted to perform multi-scale fusion of shallow surface features and deep abstract features, and on this basis, a fusion attention mechanism is adopted to further enhance the network's ability to capture image details. The specific improvement methods are as follows:
[0060] (1) As Figure 2 shown, first, the SK attention mechanism and the ECA attention mechanism are cascaded and fused according to their structural characteristics to construct a SE-ShuffleNetV2 network model. First, the SK module is used to select different convolutional kernels according to the input information to adjust the receptive field size, and then the ECA module is used to adaptively adjust the weights of the channel features so as to better focus on important features and suppress unimportant features, thereby enhancing the network's representation ability without significantly increasing parameters and computational costs. Integrating the fused module into the network can enhance the generalization ability and robustness of the model.
[0061] (2) The H-Swish activation function proposed by MobieNetV3 can effectively replace the ReLU activation function, which helps to improve the model's non-linear modeling ability. H-Swish has better smoothing performance at the boundary compared to ReLU, which helps to reduce the problem of gradient disappearance, and is also more efficient in terms of computational volume, which improves the overall training speed and stability. Since it is simple to implement in hardware, it can effectively improve the inference speed, especially for edge computing scenarios, especially in resource-constrained situations, it can reduce energy consumption and improve the response speed while ensuring computational efficiency and improving the overall performance.
[0062] (3) In this embodiment, four different types of iron ore images are classified. The classification task is relatively simple, and the depth of the required network model does not need to be too deep. Therefore, in order to reduce the consumption of the number of parameters and the amount of computation, the stacking numbers of Unit1 and Unit2 in stages 2, 3, and 4 of the original network are reduced to 1, so that the ratio of the stacking unit numbers of Unit1 to Unit2 is 1:1. After the original network passes through a 3×3 ordinary convolution, a max-pooling layer is used for downsampling to reduce the dimension and the number of parameters. Finally, the improved SEH-ShuffleNetV2s network model is formed. The improved network uses a depthwise separable convolution with a stride of 2 and a size of 3×3 with a small number of parameters to replace the max-pooling layer, as shown in Figure 3 , this change can not only avoid information loss, but also capture the subtle features in the image more precisely.
[0063] Furthermore, model training and testing:
[0064] The original ShuffleNetV2 network model, the SE-ShuffleNetV2 network model in (1) above, and the SEH-ShuffleNetV2s network model in (3) are trained with the same parameter configuration respectively.
[0065] Training:
[0066] Experimental environment configuration: The network is trained and tested using the Pytorch framework. Hardware environment: AMD Ryzen 5 5600H with Radeon Graphics @ 3.30GHz processor, NVIDIA GeForce GTX1060 Ti, 16GB of memory. GPU software environment: Windows10 64bit system, CUDA11.7, CUDNN8.6.0, PyTorch2.3.1, Pycharm2023.
[0067] The random gradient descent method is used to accelerate the model. The training cycle Epoch = 100 is set, and the selection of the batch size Batch_size is determined by the video memory size. Here, Batch_size = 32 is selected. The SGD optimizer is used, and the learning rates lr = 0.1, 0.01, and 0.001 are set respectively for comparative experiments. It is concluded that when lr = 0.01, the experimental effect is the best, and the cross-entropy is used to calculate the loss.
[0068] Testing:
[0069] Using the above-trained original ShuffleNetV2 network model, SE-ShuffleNetV2 network model, and SEH-ShuffleNetV2s network model, the recognition results are compared and analyzed through the test set. The accuracy comparison curve is asFigure 9 As shown, the dotted horizontal line is the recognition accuracy curve of the SEH-ShuffleNetV2s network model adopted in this embodiment. It can be seen that as the number of training times increases, the accuracy gradually stabilizes, and compared with the original ShuffleNetV2 network model (solid line) and the SE-ShuffleNetV2 network model that only introduces the attention mechanism (dotted line), the SEH-ShuffleNetV2s network model has a higher accuracy.
[0070] Furthermore, model application:
[0071] Before inputting the iron ore image to be detected into the iron ore recognition model, it also includes normalizing and cropping the iron ore image to be detected.
[0072] Specifically, use the above-trained SEH-ShuffleNetV2s network model, that is, the iron ore recognition model for prediction. Normalize the iron ore image to be detected to a size of 224*224, perform the CenterCrop operation to crop the picture without distortion, load the model and weights. To prevent errors when predicting grayscale images, convert the image to an RGB image, and detect the picture. The output result is the category of the iron ore. The recognition effect of some parts is as shown in Figure 5 、 6 、7, 8, and the final recognition accuracy can reach up to 96.3%.
[0073] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An iron ore image recognition method based on an improved ShufflenetV2 network, characterized in that, Including: Obtain an iron ore image to be detected; Input the iron ore image to be detected into a preset iron ore recognition model, and output an iron ore category recognition result. Among them, the iron ore recognition model is obtained by training with a training set, the training set includes a number of sample pictures of hematite, magnetite, siderite and chlorite, and the iron ore recognition model is constructed using a ShuffleNetV2 network improved based on the attention mechanism.
2. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 1, wherein Before training the iron ore recognition model with the training set, it also includes augmenting the training set. The augmentation includes: Construct a first training set based on the sample pictures of hematite, magnetite, siderite and chlorite; Perform geometric transformation on the first training set, and add the geometrically transformed sample images to the first training set to construct a second training set. Among them, the geometric transformation includes flipping and rotation. Perform color transformation on the second training set, and add the color-transformed sample images to the second training set to construct a third training set. Among them, the color transformation is performed by adding salt and pepper noise.
3. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 1, characterized in that, The iron ore recognition model includes n processing modules connected in sequence, which are used to extract information layer by layer from the input image and output the recognition result.
4. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 3, wherein The n processing modules include a first processing module, a second processing module, a third processing module, a fourth processing module, and a fifth processing module. Among them, the first processing module includes a 3×3 convolutional layer and a 3×3 depthwise separable convolutional layer connected in sequence; the second processing module, the third processing module, and the fourth processing module all include a ShuffleNetV2Unit2 and a ShuffleNetV2 Unit1 improved with the attention mechanism connected in sequence; the fifth processing module includes a 1×1 convolutional layer, a global pooling layer, and a fully connected layer.
5. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 4, wherein, Improving the ShuffleNetV2 Unit2 and ShuffleNetV2Unit1 with the attention mechanism includes: Adding a hybrid attention layer after the Channel Shuffle layer of the ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1, where the hybrid attention layer uses the SK attention mechanism and the ECA attention mechanism in series fusion.
6. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 4, wherein The 3×3 convolutional layer, the 1×1 convolutional layer, and the convolutional layers in the ShuffleNetV2Unit2 and ShuffleNetV2 Unit1 all use the H-Swish activation function.
7. The iron ore image recognition method based on the improved ShufflenetV2 network according to claim 1, characterized in that, Before inputting the iron ore image to be detected into the iron ore recognition model, it also includes normalizing and cropping the iron ore image to be detected.
Citation Information
Patent Citations
Improved ShuffleNetV2-based ventricular premature beat identification method
CN114886437A
Citrus disease identification method based on improved ShuffleNetV2
CN117830821A
Document image direction recognition and model training
WO2021174962A1