A Coal-Rock Image Classification Method, Device and Storage Medium Based on Deep Learning

Multi-scale features of coal rock images are extracted through large-core parallel convolution and attention residual network (Att-Resnet), which solves the problem of feature extraction and classification accuracy in coal rock image classification, and achieves more efficient coal rock image recognition.

CN119295819BActive Publication Date: 2025-07-22LINYI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411367514.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-07-22
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

The prior art has problems such as difficulty in feature extraction, low efficiency of classification algorithms, data set problems and category imbalance in coal rock image classification, resulting in insufficient classification accuracy.

Method used

The attention residual network (Att-Resnet) that combines the attention mechanism with the attention mechanism is adopted. By constructing a feature extraction module and the residual attention module, multi-scale spatial information and channel dependencies are extracted, and coal rock image classification is performed.

Benefits of technology

The accuracy and efficiency of coal rock image classification are improved, the fixed convolution kernel limitation of feature extraction is solved, the gradient disappearance and explosion problems are avoided, and a richer feature map is provided for downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119295819B_ABST
    Figure CN119295819B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and specifically relates to a coal-rock image classification method, device and storage medium based on deep learning. The steps are as follows: construct a coal-rock image data set and perform preprocessing operations on the data set; construct an attention residual network Att-Resnet, which includes a feature extraction module A1 and multiple residual attention modules A2; input the images in the coal-rock image data set into the attention residual network Att-Resnet for feature extraction to obtain feature maps; input the feature maps into a classifier for classification to obtain the probability values of each category corresponding to the feature maps, and then select the category with the highest probability value as the predicted category. The present invention adopts a large-kernel parallel convolution mechanism to provide richer feature maps for downstream tasks. By using residual connections and introducing an attention mechanism, it can effectively extract finer-grained multi-scale spatial information, and at the same time establish longer-range channel dependencies, providing a prerequisite for the improvement of the classification accuracy of coal-rock images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a coal-rock image classification method based on deep learning. Background Art

[0002] Coal is the most economical fossil energy in the world and also the main energy source in China, playing a crucial role in China's energy security and economic and social development. The intelligent and unmanned mining of coal can improve coal production while reducing coal mine accidents, especially face accidents. Coal-rock image recognition is the process of automatically recognizing and classifying coal-rock images using computer vision technology. The background technology of coal-rock image recognition mainly includes aspects such as image preprocessing, feature extraction, classification algorithms, and model evaluation. Through these technologies, the automatic recognition and classification of coal-rock images can be achieved. However, due to the usually complex texture and structural features of coal-rock images, such as the longitudinal and transverse textures of coal seams, the contact surfaces between rock layers, etc., an effective feature extraction method is the key to solving the problem of coal-rock image classification. Commonly used feature extraction methods include texture features, shape features, color features, etc. In summary, the background technology and problems in the field of coal-rock image classification include feature extraction, classification algorithms, dataset problems, class imbalance problems, and real-time requirements, etc. Solving these problems requires the comprehensive application of technologies in fields such as computer vision, machine learning, and data processing, and targeted research and optimization in combination with the characteristics of coal-rock images. Summary of the Invention

[0003] In view of the above problems, the present invention provides a coal-rock image classification method, device, and storage medium based on deep learning. The large kernel parallel convolution mechanism is used to provide richer feature maps for downstream tasks. At the same time, residual connections and the introduction of an attention mechanism are adopted, which can effectively extract finer-grained multi-scale spatial information, and at the same time can establish longer-range channel dependencies, providing richer feature maps for the classification task, and also providing a premise for the improvement of the classification accuracy of coal-rock images.

[0004] The technical solution adopted by the present invention to overcome its technical problems is as follows:

[0005] In the first aspect, the present invention provides a coal-rock image classification method based on deep learning, including the following steps:

[0006] S1. Collect coal-rock images and construct a coal-rock image dataset, and then perform preprocessing operations on the dataset;

[0007] S2. Construct an attention residual network Att-Resnet, where Att-Resnet includes a feature extraction module A1 and multiple residual attention modules A2;

[0008] S3. Construct the feature extraction module A1, which consists of large-kernel parallel convolution modules. Input the images in the processed dataset into the feature extraction module A1, and extract the feature maps of coal-rock images through the large-kernel parallel convolution modules. ;

[0009] S4. Construct multiple residual attention modules A2. The residual attention module A2 is a variant residual module that replaces the 3×3 convolution in the classical residual module with an attention mechanism. Input the feature maps into multiple residual attention modules A2 for further feature extraction to obtain the final feature maps ;

[0010] S5. Input the final feature maps into the classifier for classification to obtain the probability values corresponding to each category of the feature maps , and then select the category with the highest probability value as the predicted category.

[0011] S1 is specifically as follows:

[0012] Scale the coal-rock images in the dataset to a unified size, and then further enhance the images through random cropping and rotation operations to obtain the preprocessed dataset , , represents the th coal-rock image in the processed dataset represents the number of coal-rock images in the preprocessed dataset , represents any coal-rock image in the preprocessed dataset .

[0013] S3 is specifically as follows:

[0014] The large-kernel parallel convolution module includes four branches. The feature maps pass through a 1×1 convolution in the first branch, a convolution in the second branch, a 3×3 convolution in the third branch, directly pass through the fourth branch, then add the results of the four branches, and then successively pass through a normalization layer and an activation function, and finally output the feature maps , and the calculation process is as follows:

[0015] ,

[0016] Among them, represents the operation of the activation function , is the normalization layer, , and respectively represent the 3×3 convolution, convolution, and 1×1 convolution operations in the large kernel parallel convolution module, represents the size of the convolution kernel in the convolution, and the 2D convolution 3×3 convolution, convolution, and 1×1 convolution in the large kernel parallel convolution module are replaced by depth over-parameterized convolution , The convolution calculation is as follows:

[0017] ,

[0018] where, is a two-dimensional tensor, The number of input channels of is and the height is , represents the depth convolution trainable kernel, , represents the depth multiplier in the depth convolution, represents the traditional convolution trainable kernel, , represents the number of output channels, represents the conventional convolution operator, represents the depth convolution operator, represents the output convolution result.

[0019] S4 is specifically as follows:

[0020] S4.1 Att-Resnet includes residual attention modules A2, and the residual attention module A2 sequentially includes the first 1×1 convolution, the attention module, and the second 1×1 convolution;

[0021] S4.2 Input the feature map into the first residual attention module A2. After passing through the first 1×1 convolution in the residual attention module A2, the number of channels of the feature map is reduced to be passable through the attention module, and then the feature map with the reduced number of channels is input into the attention module. The attention module first divides the image channels of the feature map, and the number of divided channels is n. Then, feature extraction is performed on each channel respectively, and n feature maps are correspondingly obtained, denoted as , represents the feature map output from the th channel, , and the calculation process is as follows:

[0022] ,

[0023] in, Indicates The channel is the convolution kernel Convolution;

[0024] Then, the feature map of each channel output is extracted through the attention mechanism in the attention module The attention weight , the calculation process is as follows:

[0025] ,

[0026] in, represents the attention mechanism operation, and Represents the height and width of the input feature map, The feature map Any point in , Representatives passed the The feature map output by each channel, Represents the Sigmoid function operation, Represents the relu activation function operation, and are two different weight parameters in the attention mechanism, , , is the reduction ratio, is the number of input channels;

[0027] Then use The activation function is used to recalibrate the weights of the channel attention information. The calculation process is as follows:

[0028] ,

[0029] in, Indicates the image The attention weights after recalibration of the channels, express Operation of activation function;

[0030] S4.3 Feature map Attention weights recalibrated with corresponding weights Perform channel-wise multiplication to obtain the feature map , the calculation process is as follows:

[0031] ,

[0032] in, Denote multiplication at the channel-wise level;

[0033] Then we will obtain Perform dimensional concatenation to obtain a feature map after multi-scale feature information attention weighting , and the calculation process is as follows:

[0034] ,

[0035] Among them, Denote the concatenation operation;

[0036] Finally, the feature map Restores the number of channels of the feature map through the second 1×1 convolution operation to obtain the feature map ;

[0037] S4.4. Use the feature map output by the first residual attention module A2 as the input of the second residual attention module A2, repeat the operations in steps S4.2 and S4.3 to obtain the output feature map of the second residual attention module A2. Use the output of the previous residual attention module A2 as the input of the next residual attention module A2, and iterate in turn to obtain the output feature map of the th residual attention module A2;

[0038] S4.5. The feature map passes through the bn layer, relu layer, bn layer in sequence, and then adds it to the feature map through residual connection and then passes through a relu layer to obtain the final feature map , and the calculation process is as follows:

[0039] .

[0040] S5 is specifically as follows:

[0041] Input the feature map into the classifier for classification. The classifier includes an average pooling layer and a fully connected layer. The feature map first passes through the average pooling layer to obtain a feature space vector, and then passes through the mapping of the fully connected layer to obtain the class with the highest probability value. The calculation process is as follows:

[0042] ,

[0043] Among them, Denote the prediction result, Denote the operation of the average pooling layer, Denote the operation of the fully connected layer.

[0044] In a second aspect, the present invention provides a coal-rock image classification device based on deep learning, including the following units:

[0045] (1) Preprocessing unit: used to construct a coal-rock image data set and preprocess the data set;

[0046] (2) Feature extraction module A1: extracts features of coal-rock images by using a large kernel parallel convolution module;

[0047] (3) Residual attention module A2: composed of a variant residual module and an attention module, used to extract channel attention of feature maps at different scales, obtain channel attention vectors at each different scale, and output a feature map after attention weighting of multi-scale feature information;

[0048] (4) Classification unit: used to classify the feature map.

[0049] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes a coal-rock image classification method based on deep learning.

[0050] The advantages of the present invention are as follows:

[0051] By using a large kernel parallel convolution module to extract features of the original coal-rock images, the limitations of the fixed convolution kernel size and receptive field are avoided, and the extracted features are more rich. Moreover, the large kernel parallel convolution module is a plug-and-play module and can be embedded into any convolutional neural network architecture to improve performance. The network structure of the present invention uses a residual module as the backbone network, avoiding problems such as gradient disappearance and gradient explosion caused by the network being too deep. In addition, the present invention also embeds an attention mechanism in the backbone network, performs dimensional splicing on the feature map after obtaining the attention weights, and finally outputs a feature map with richer multi-scale information, which more effectively provides a strong guarantee for the downstream classification task. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0053] Figure 1 It is a schematic diagram of the Att-Resnet structure of the present invention.

[0054] Figure 2 It is a heat map of the region of interest in coal-rock image feature extraction.

[0055] Figure 3This is the schematic diagram of the calculation principle of Do-conv convolution. Detailed implementation manners

[0056] In order to better understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0057] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0058] Embodiment 1

[0059] This method is aimed at the complex underground coal and rock environment, uses a coal and rock image classification network based on deep learning to solve the problem of low classification accuracy of coal and rock images. At the same time, based on the characteristics of coal and rock images, channel attention of feature maps with different scales is extracted on each channel, and a feature map after attention weighting of multi-scale feature information is output. At the same time, the large kernel parallel convolution block can handle the problem of long-distance spatial position information dependence.

[0060] The following combines Figure 1 to specifically describe a coal and rock image classification method based on deep learning implemented by the present invention. The specific steps are as follows:

[0061] S1. Collect coal and rock images and construct a coal and rock image data set, and then perform preprocessing operations on the data set;

[0062] S2. Construct an attention residual network Att-Resnet, and Att-Resnet includes a feature extraction module A1 and multiple residual attention modules A2;

[0063] S3. Construct a feature extraction module A1, and the feature extraction module A1 is composed of large kernel parallel convolution modules. Input the images in the processed data set into the feature extraction module A1, and extract the feature maps of coal and rock images through the large kernel parallel convolution modules ;

[0064] S4. Construct multiple residual attention modules A2. The residual attention module A2 is a variant residual module that replaces the 3×3 convolution in the classical residual module with an attention mechanism. Input the feature maps into multiple residual attention modules A2 for further feature extraction to obtain the final feature maps ;

[0065] S5. Use the final feature maps Input it into the classifier for classification to obtain the feature map The probability values corresponding to each category are obtained, and then the category with the highest probability value is selected as the predicted category.

[0066] S1 is specifically as follows:

[0067] Scale the coal-rock images in the dataset to a unified size, and then further enhance the images through random cropping and rotation operations to obtain the preprocessed dataset , , denotes the th coal-rock image in the processed dataset denotes the number of coal-rock images in the preprocessed dataset , denotes any coal-rock image in the preprocessed dataset .

[0068] S3 is specifically as follows:

[0069] The large-kernel parallel convolution module includes four branches. The feature map passes through a 1×1 convolution in the first branch, a convolution in the second branch, a 3×3 convolution in the third branch, directly passes through the fourth branch, then adds the results of the four branches, and then successively passes through a normalization layer and an activation function, and finally outputs the feature map . The calculation process is as follows:

[0070] ,

[0071] where denotes the operation of the activation function . is the normalization layer , and respectively denote the operations of the 3×3 convolution, convolution, and 1×1 convolution in the large-kernel parallel convolution module denotes the size of the convolution kernel in the convolution, and the 2D convolution 3×3 convolution, convolution, and 1×1 convolution in the large-kernel parallel convolution module are replaced by the depth over-parameterized convolution . The convolution calculation is as follows:

[0072] ,

[0073] where is a two-dimensional tensor The number of input channels is , the height is , the width is , represents the depth convolutional trainable kernel, , represents the depth multiplier in the depth convolution, represents the traditional convolutional trainable kernel, , represents the number of output channels, represents the conventional convolution operator, represents the depth convolution operator, represents the convolutional result of the output.

[0074] S4 is specifically as follows:

[0075] S4.1 Att-Resnet includes residual attention modules A2. The residual attention module A2 sequentially includes the first 1×1 convolution, the attention module, and the second 1×1 convolution;

[0076] S4.2 Input the feature map into the first residual attention module A2. After passing through the first 1×1 convolution in the residual attention module A2, the channels of the feature map are reduced to be passable through the attention module. Then, the feature map with the reduced number of channels is input into the attention module. The attention module first divides the image channels of the feature map. The number of divided channels is n. Then, feature extraction is performed on each channel respectively, and n feature maps are correspondingly obtained, denoted as , represents the feature map output from the th channel, , and the calculation process is as follows:

[0077] ,

[0078] Among them, represents that the th channel is a convolution with a convolution kernel of ;

[0079] Then, the attention weights of the feature map output from each channel are extracted through the attention mechanism in the attention module, and the calculation process is as follows:

[0080] ,

[0081] Among them, represents the attention mechanism operation, and denote the height and width of the input feature map, is the feature map at any point in , represents the feature map output after passing through the th channel, represents the Sigmoid function operation, represents the relu activation function operation, and are two different weight parameters in the attention mechanism, , , is the reduction ratio, is the number of input channels;

[0082] Then use activation function to recalibrate the weights of the channel attention information, and the calculation process is as follows:

[0083] ,

[0084] Among them, represents the attention weight after recalibration of the th channel of the image, represents the operation of the activation function;

[0085] S4.3 Multiply the feature map with the corresponding attention weight after weight recalibration at the channel-wise level to obtain the feature map , and the calculation process is as follows:

[0086] ,

[0087] Among them, represents the channel-wise multiplication;

[0088] Then splice the obtained to obtain a feature map after multi-scale feature information attention weighting , and the calculation process is as follows:

[0089] ,

[0090] Among them, represents the splicing operation;

[0091] Finally, the feature map restores the number of channels of the feature map through the second 1×1 convolution operation to obtain the feature map ;

[0092] S4.4. Use the feature map output by the first residual attention module A2 as the input of the second residual attention module A2, and repeat the operations in steps S4.2 and S4.3 to obtain the output feature map of the second residual attention module A2 . Use the output of the previous residual attention module A2 as the input of the next residual attention module A2, and iterate in turn to obtain the output feature map of the th residual attention module A2 ;

[0093] S4.5. The feature map passes through a bn layer, a relu layer, and a bn layer in sequence, and then is added to the feature map through a residual connection and then passes through a relu layer to obtain the final feature map . The calculation process is as follows:

[0094] .

[0095] S5 is specifically as follows:

[0096] Input the feature map into the classifier for classification. The classifier includes an average pooling layer and a fully connected layer. The feature map first passes through the average pooling layer to obtain a feature space vector, and then passes through the mapping of the fully connected layer to obtain the class with the highest probability value. The calculation process is as follows:

[0097] ,

[0098] where represents the prediction result, represents the operation of the average pooling layer, represents the operation of the fully connected layer.

[0099] To further prove that the present invention can more effectively provide strong guarantee for the downstream classification task, the classification results of coal-rock images were analyzed based on this method. The experimental dataset was divided into seven categories (coal, basalt, granite, limestone, marble, quartzite, and sandstone), with a total of 1,503 training images and 377 test images. Through the method in the present invention, data preprocessing operations were uniformly performed on the above data. In this small dataset, the classification accuracy of the model reached 90.343% after only 60 rounds of training, and experimental comparisons were made with existing image classification methods. The comparison experimental results are shown in Table 1. Att-Resnet50 is the method in the present invention, and EPSANet50 and PreActResNet34 are existing methods. Under the same training conditions, the method in the present invention has a higher accuracy rate.

[0100] Table 1 Comparison table of recognition accuracies between the present method and existing methods on this small dataset

[0101]

[0102] Example 2

[0103] The present invention provides a coal-rock image classification device based on deep learning, including the following units:

[0104] (1) Preprocessing unit: used to construct a coal-rock image dataset and preprocess the dataset;

[0105] (2) Feature extraction module A1: used to extract features of coal-rock images by using a large-kernel parallel convolution module;

[0106] (3) Residual attention module A2: composed of a variant residual module and an attention module, used to extract channel attention of feature maps at different scales, obtain channel attention vectors at each different scale, and output a feature map after multi-scale feature information attention weighting;

[0107] (4) Classification unit: used to classify the feature map.

[0108] Example 3

[0109] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes a coal-rock image classification method based on deep learning.

[0110] The storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks or optical discs.

[0111] Example 4

[0112] As Figure 2 shown, the present invention uses training parameters to implement a visualization heat map of the attention area of the network for the feature map. The heat map can display key points or attention areas in tasks such as object detection, image classification, or face recognition in computer vision. Figure 2 The blue area in [description of the reference] belongs to the key attention area of the network. This method can widely focus on the coal-rock area of the image in the coal-rock image feature extraction task, providing effective features for the downstream classification task.

[0113] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A coal-rock image classification method based on deep learning, characterized in that, It includes the following steps: S1. Collect coal-rock images and construct a coal-rock image dataset, and then perform preprocessing operations on the dataset; S2. Construct an attention residual network Att-Resnet, and Att-Resnet includes a feature extraction module A1 and multiple residual attention modules A2; S3. Construct the feature extraction module A1, which consists of a large-kernel parallel convolution module. Input the images in the processed dataset into the feature extraction module A1, and extract the feature maps of coal-rock images through the large-kernel parallel convolution module ; S4. Construct multiple residual attention modules A2. The residual attention module A2 is a variant residual module that replaces the 3×3 convolution in the classical residual module with an attention mechanism, and input the feature map into multiple residual attention modules A2 for further feature extraction to obtain the final feature map ; S4 is specifically as follows: S4.1 The Att-Resnet includes two residual attention modules A2. The residual attention module A2 sequentially includes a first 1×1 convolution, an attention module, and a second 1×1 convolution; S4.2 Input the feature map into the first residual attention module A2. Through the first 1×1 convolution in the residual attention module A2, reduce the number of channels of the feature map to a number that can pass through the attention module. Then input the feature map with the reduced number of channels into the attention module. The attention module first divides the image channels of the feature map, and the number of divided channels is n. Then, feature extraction is performed on each channel respectively, and n feature maps are correspondingly obtained, denoted as , indicating the feature map output from the th channel, ; Then, the attention mechanism in the attention module is used to extract the feature maps output by each channel of the attention weights ; Then, use the activation function to recalibrate the weights of the channel attention information; S4.3 Multiply the feature map with the attention weights after corresponding weight recalibration at the channel-wise level to obtain the feature map . The calculation process is as follows: , Among them, represents multiplication at the channel-wise level; Then, the obtained is dimensionally concatenated to obtain a feature map after attention weighting of multi-scale feature information , and the calculation process is as follows: , Among them, represents a splicing operation; Final feature map The number of channels of the feature map is restored through the second 1×1 convolution operation to obtain the feature map ; ; S4.

4. Use the feature map output by the first residual attention module A2 as the input of the second residual attention module A2, repeat the operations in steps S4.2 and S4.3 to obtain the output feature map of the second residual attention module A2 . Use the output of the previous residual attention module A2 as the input of the next residual attention module A2, and iterate sequentially to obtain the output feature map of the nth residual attention module A2 ; S4.

5. Feature Map It passes through the BN layer, the ReLU layer, the BN layer in turn, and then connects to the feature map through the residual connection After adding, it passes through a relu layer to get the final feature map , the calculation process is as follows: ; S5. Input the final feature map into the classifier for classification to obtain the feature map corresponding to the probability values of each category, and then select the category with the highest probability value as the predicted category.

2. The coal-rock image classification method based on deep learning according to claim 1, wherein, S1 is specifically as follows: Scale the coal-rock images in the dataset to a unified size, and then further enhance the images through random cropping and rotation operations to obtain the preprocessed dataset , , denotes the th coal-rock image in the processed dataset denotes the number of coal-rock images in the preprocessed dataset , denotes any coal-rock image in the preprocessed dataset .

3. The method for classifying coal-rock images based on deep learning according to claim 2, characterized in that, S3 is specifically as follows: The large-core parallel convolution module includes four branches, and the feature map In the first branch, a 1×1 convolution is performed. In the second branch, a convolution is performed. In the third branch, a 3×3 convolution is performed. The fourth branch passes directly through. Then, the results of the four branches are added together, and then successively passed through a normalization layer and an activation function, and finally the feature map is output. The calculation process is as follows: , Among them, represents the operation of the activation function . is the normalization layer, , and respectively represent the operations of 3×3 convolution, convolution and 1×1 convolution in the large kernel parallel convolution module, represents the size of the convolution kernel in the convolution, and the 2D convolution 3×3 convolution, convolution and 1×1 convolution in the large kernel parallel convolution module are replaced by the depth over-parameterized convolution . The convolution calculation is as follows: , Among them, is a two-dimensional tensor, has an input channel count of , a height of , and a width of . represents the depth convolution trainable kernel, . represents the depth multiplier in depth convolution, represents the traditional convolution trainable kernel, . represents the output channel count, represents the conventional convolution operator, represents the depth convolution operator, represents the output convolution result.

4. A coal-rock image classification method based on deep learning according to claim 3, characterized in that The calculation process in S4.2 is specifically as follows: , Among them, indicates that the th channel is a convolution with a convolution kernel of . , Among them, represents the attention mechanism operation, and represent the height and width of the input feature map, is an arbitrary point in the feature map , , represents the feature map output through the th channel, represents the Sigmoid function operation, represents the relu activation function operation, and are two different weight parameters in the attention mechanism, , , is the reduction ratio, is the number of input channels; , Among them, represents the attention weight after recalibration of the th channel of the image, represents the operation of the activation function.

5. A coal-rock image classification method based on deep learning according to claim 4, characterized in that S5 is specifically as follows: Input the feature map into the classifier for classification. The classifier includes an average pooling layer and a fully connected layer. The feature map first passes through the average pooling layer to obtain a feature space vector, and then passes through the mapping of the fully connected layer to obtain the class with the highest probability value. The calculation process is as follows: , Among them, represents the prediction result, represents the operation of the average pooling layer, represents the operation of the fully connected layer.

6. A coal-rock image classification device based on deep learning, which executes a coal-rock image classification method based on deep learning as described in any one of claims 1-5, characterized in that, It includes the following units: (1) Preprocessing unit: used to construct a coal-rock image dataset and perform preprocessing on the dataset; (2) Feature extraction module A1: uses a large-kernel parallel convolution module to extract the features of coal-rock images; (3) Residual attention module A2: consists of a variant residual module and an attention module, used to extract the channel attention of feature maps at different scales, obtain the channel attention vectors at each different scale, and output a feature map after attention weighting of multi-scale feature information; (4) Classification unit: used to classify the feature maps.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the deep learning-based coal-rock image classification method according to any one of claims 1 to 5.