CNN-Transform combination-based coal rock image classification method and apparatus, and electronic device
By combining CNN and Transformer, the local and global features of coal rock images are extracted, and the model structure is optimized, the problem of the difficulty of taking into account the accuracy and efficiency of the existing technology in coal rock image classification is solved, and the coal rock image classification effect with high accuracy and low computational complexity is achieved.
Patent Information
- Application Number
- CN202510068770.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The prior art is difficult to effectively combine convolutional networks and Transformer in coal rock image classification tasks, resulting in a difficult trade-off between accuracy and efficiency, especially on low-latency mobile devices.
By building a deep network model based on CNN-Transformer combination, combining the advantages of convolutional layers and Transformer, the local and global features of the image are extracted, and by optimizing the model structure, such as replacing the convolutional layer and using RefConv, the amount of parameters and computational redundancy is reduced.
High accuracy and robustness in low-precision coal rock image classification are achieved, and the model parameter quantity and calculation complexity are reduced, making it possible to classify coal rock image on mobile devices.
Smart Images

Figure CN119942210A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a coal rock image classification method, device and electronic equipment based on CNN-Transformer combination. Background Art
[0002] Coal is the most economical fossil energy in the world and plays a decisive role in world energy security and social development. However, due to the complex environment, heavy dust and poor lighting conditions in coal mines, high-noise and low-quality coal images are often collected in this environment. Therefore, it is difficult to extract useful information from these low-quality images, which seriously restricts the application of image and video technology in coal mine intelligence. Among them, coal rock image classification is a challenging task. Since coal rock images usually have complex texture and structural features, such as the longitudinal and transverse textures of coal seams and the contact surface between rock layers, effective feature extraction methods are the key to solving the coal rock image classification problem. Simple convolutional neural networks have fallen behind emerging technologies such as Transformer in a large number of coal rock image classification tasks. In addition, the underground environment of coal mines is complex, and the demand for high-precision coal rock image classification on low-latency mobile devices is also growing.
[0003] In order to deal with these difficulties, image classification tasks usually require the use of high-performance computer vision technologies, such as the introduction of Transformers into the field of image classification, but they still lag behind the most advanced convolutional networks. This work shows that although Transformers tend to have larger model capacity, their generalization may be worse than convolutional networks due to the lack of correct inductive bias, resulting in poor performance in coal and rock image classification tasks. At the same time, the introduction of Transformers leads to excessive computational parameters, which is not conducive to the support of mobile devices. Convolutional layers often have better generalization and faster convergence due to their strong inductive bias priors, while attention layers have higher model capacity and can benefit from larger data sets. The combination of convolutional layers and attention layers can achieve better generalization and capacity. However, a key challenge here is how to effectively combine them to achieve a better trade-off between accuracy and efficiency.
[0004] Therefore, the present invention provides a coal rock image classification method, device and electronic device based on CNN-Tramsformer combination to solve the above problems. Summary of the invention
[0005] In response to the above problems, the present invention provides a coal rock image classification method, device and electronic equipment based on the combination of CNN-Transformer. By combining the respective advantages of CNN and Transformer, it can extract both local features and global features of the image, providing a powerful tool for the field of geological research.
[0006] The present invention is achieved through the following technical solutions: In a first aspect, the present invention provides a coal rock image classification method based on CNN-Transformer combination, which specifically comprises the following steps: S1. Construct the initial deep network model CNN-Transformer, input the data in the texture library dataset into the initial CNN-Transformer model, and obtain the basic weights. ; S2. Collect coal and rock images and filter them to obtain a data set , for the dataset Preprocessing is performed to obtain the preprocessed data set , and then the preprocessed data set The image is processed preliminarily to obtain the feature map ; S3. Optimize the initial CNN-Transformer model by replacing a standard convolution and a depth-separable convolution layer in the initial CNN-Transformer model to obtain an optimized CNN-Transformer model. S4. Feature map Input into the optimized CNN-Transformer model to obtain the feature map; S5. Input the results of the optimized CNN-Transformer model into the classifier to obtain the classification result.
[0007] S1 is as follows: The structure of the initial deep network model CNN-Transformer is as follows: Block RBConv shallow feature extraction module, Block Sim-Transformer deep feature extraction module; Among them, the RBConv shallow feature extraction module includes: a standard convolution layer, a depth-separable convolution layer, and a SE residual connection module; The Sim-Transformer deep feature extraction module includes: Sim-Attntion module, feed-forward neural network FFN; The texture library dataset is an existing dataset in which images have been classified and labeled. After the data in the texture library dataset is input into the initial CNN-Transformer model, the feature map of the image in the texture library dataset is first extracted, and then the extracted feature map is processed by the initial CNN-Transformer model to complete the training of the initial CNN-Transformer model. Finally, the convolution kernel finally output by the initial CNN-Transformer model is frozen as the basis weight .
[0008] S2 is as follows: For the dataset Preprocess the images in the dataset, scale them to the same size, and increase the number of images in the dataset by random cropping and rotation operations. Finally, the preprocessed dataset is obtained. ; Dataset Expressed as , the preprocessed dataset Expressed as ,in, , Representation dataset The number of images in Represents the preprocessed dataset The number of images in Representation dataset Middle images, Represents the preprocessed dataset Middle images, Represents the preprocessed dataset Any image in After preprocessing the data set The image is processed initially. Input to two 3×3 convolutions for feature map extraction to obtain the feature map .
[0009] S3 is as follows: The optimized model is obtained based on the initial model by replacing one standard convolution layer and one depthwise separable convolution layer in the initial model with RefConv convolution; The optimized CNN-Transformer model consists of Block RBConv shallow feature extraction module and The block Sim-Transformer deep feature extraction modules are stacked and connected in series; Among them, the RBConv shallow feature extraction module includes: RefConv re-parameterized refocusing convolution and SE residual connection module; The SE residual connection module includes: a global average pooling layer, a fully connected layer, and an activation function for generating channel weights; The Sim-Transformer module includes: Sim-Attntion module and feed-forward neural network FFN.
[0010] S4 is as follows: S4.1. Feature Map Before being input into the optimized CNN-Transformer model, it is first passed through a Convolution, the feature map The number of channels C is expanded by 4 times, and the feature map The original number of channels is C, and the number of channels after expansion is 4C; S4.2. Feature map after channel expansion Input to the RBConv shallow feature extraction module of the CNN-Transformer model, feature map Output feature map after RefConv ; Then the basis weights of the initial CNN-Transformer model output Perform refocusing transformation to obtain conversion weight , the calculation process is as follows: , in, represents the refocusing transformation, Representation feature map The trainable parameters of Then the feature map Input to the SE residual module, and the input feature map is processed by the SE module Perform global average pooling, the calculation process is as follows: , in, Indicates channel The features of the global average pooling output are of size 1×1×4C. Indicates height, Indicates width, Representation feature map Height × width × channels, Representation feature map The index in the height direction, , Representation feature map Index in the width direction, ,aisle The number of channels is 4C; The features of the global average pooling output Input to the fully connected layer, through the fully connected layer The activation function gets the channel The weight value , the calculation process is as follows: , The input feature map Channel Multiply it with the corresponding weight to generate a new feature map , the calculation process is as follows: , Among them, the feature map The dimensions are H×W×4C; Final feature map Then a 1×1 convolution is used to reduce the number of channels to 1 / 4, and the number of channels after reduction is C; S4.3. Feature map Input to Sim-Transformer module, feature map It is a matrix composed of multiple neurons, each neuron corresponds to a receptive field, and the energy function is defined for finding important neurons , the calculation process is as follows: , , , , in, Representation feature map The target neuron of a single channel in Indicates that except for the target god general External feature map Other neurons in a single channel, represents the number of energy functions, , , represents the total number of energy functions in each channel, represents the transformation weight, represents the transformation bias, and Represents the feature maps Remove target neurons from a single channel All other neurons The mean and variance of Calculate minimum energy , according to the minimum energy Determine the importance, The lower the target neuron The greater the difference from the surrounding neurons, the higher the importance and the minimum energy The calculation process is as follows: , in, and denote the mean and variance of the neuron respectively, represents the regularization coefficient; Then the feature map The energy function value The form propagates downward, and the energy values of all neurons form a set ; Continue to map the feature map Input to the Sim-Attntion module, through Neuron weights and feature maps obtained by activation function Perform dot product operation and then transform the feature map The input value is mapped to the range of (0,1), and the output is a refined feature map focusing on the key target perception area. , the calculation process of the Sim-Attntion module is as follows: , The feature map Input to the feedforward neural network FFN, feature map First input to the fully connected layer to get the feature map , the calculation process is as follows: , in, Represents the activation function Operation, represents the feature dimension, represents fixed parameters; Then Perform layer normalization to obtain the final output feature map , the calculation formula is as follows: , in, Represents the feature map pair Each layer channel is normalized one by one.
[0011] S5 is as follows: The feature map is processed by the average pooling layer and the fully connected layer. To classify; For the input feature map Perform global average pooling on each channel of: , Among them, the input feature map The size is × × C ,in and are the height and width of the feature map, respectively. It is a channel The features of the global average pooling output, channel The number of channels is , size is 1×1× C ; The average pooled vector Through the fully connected layer, it is mapped to the category space to obtain the probability distribution of each category, from which the category label with the highest probability is selected as the final category label of the coal rock image. Assuming that there are The specific process is as follows: , in, Represents the weight matrix of the fully connected layer, with size , represents the bias vector, Represents the score of each category, with size , , Indicates The scores of the categories, ; Then apply the output of the fully connected layer Activation function to obtain the probability distribution of each category: in, Indicates that the input image belongs to the category probability; The classification probability of each category is calculated respectively, and the category corresponding to the highest probability value is recorded as the category of the input image.
[0012] In a second aspect, the present invention further provides a coal-rock image classification device based on CNN-Transformer combination, which implements a coal-rock image classification method based on CNN-Transformer combination, comprising the following units: 1) Pre-training learning unit: The initial CNN-Transformer model is used to process the existing data with labeled image types to complete the pre-training of the initial CNN-Transformer model; 2) Coal and rock image acquisition unit: collects data used to construct coal and rock image datasets and preprocesses the datasets; 3) Optimization unit: replace part of the structure in the initial CNN-Transformer model to obtain the optimized CNN-Transformer model; 4) Image data processing unit: The collected image is processed by the optimized CNN-Transformer model and the parameters output by the initial CNN-Transformer model to obtain the feature map of the input image; 5) Classification unit: classifies the input image according to the feature map obtained by the image data processing unit.
[0013] In the third aspect, the present invention also provides an electronic device, including a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, a coal rock image classification method based on the CNN-Transformer combination is performed.
[0014] The advantages of the present invention are: Through the technical solution of the present invention, the classification of low-precision coal and rock images has high accuracy and strong robustness. It cleverly combines the convolutional network and Transformer. The convolutional layer often has better generalization and faster convergence speed due to its strong inductive bias prior, while the attention layer in the Transformer has a higher model capacity and can benefit from a larger data set. The combination of the convolutional layer and the attention layer can obtain better generalization and capacity; and the use of Refconv can greatly reduce the parameters while reducing computational redundancy and memory access, making it possible to realize the recognition of coal and rock images on mobile devices, which can play an important role in complex coal mine environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0016] Figure 1 It is a schematic diagram of the process of the present invention.
[0017] Figure 2 The figure is a comparative thermal map of the focus areas of the method of the present invention and CoAtnet-0 in coal and rock image feature extraction. DETAILED DESCRIPTION
[0018] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0019] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0020] Example 1 Combine the following Figure 1 A coal rock image classification method based on CNN-Transformer combination implemented by the present invention is specifically described.
[0021] A coal rock image classification method based on CNN-Transformer combination includes the following steps: S1. Construct the initial deep network model CNN-Transformer, input the data in the texture library dataset into the initial CNN-Transformer model, and obtain the basic weights. ; S2. Collect coal and rock images and filter them to obtain a data set , for the dataset Preprocessing is performed to obtain the preprocessed data set , and then the preprocessed data set The image is processed preliminarily to obtain the feature map ; S3. Optimize the initial CNN-Transformer model by replacing a standard convolution and a depth-separable convolution layer in the initial CNN-Transformer model to obtain an optimized CNN-Transformer model. S4. Feature map Input into the optimized CNN-Transformer model to obtain the feature map; S5. Input the results of the optimized CNN-Transformer model into the classifier to obtain the classification result.
[0022] S1 is as follows: The structure of the initial deep network model CNN-Transformer is as follows: Block RBConv shallow feature extraction module, Block Sim-Transformer deep feature extraction module; Among them, the RBConv shallow feature extraction module includes: a standard convolution layer, a depth-separable convolution layer, and a SE residual connection module; The Sim-Transformer deep feature extraction module includes: Sim-Attntion module, feed-forward neural network FFN; The texture library dataset is an existing dataset in which images have been classified and labeled. After the data in the texture library dataset is input into the initial CNN-Transformer model, the feature map of the image in the texture library dataset is first extracted, and then the extracted feature map is processed by the initial CNN-Transformer model to complete the training of the initial CNN-Transformer model. Finally, the convolution kernel finally output by the initial CNN-Transformer model is frozen as the basis weight .
[0023] S2 is as follows: For the dataset Preprocess the images in the dataset, scale them to the same size, and increase the number of images in the dataset by random cropping and rotation operations. Finally, the preprocessed dataset is obtained. ; Dataset Expressed as , the preprocessed dataset Expressed as ,in, , Representation dataset The number of images in Represents the preprocessed dataset The number of images in Representation dataset Middle images, Represents the preprocessed dataset Middle images, Represents the preprocessed dataset Any image in After preprocessing the data set The image is processed initially. Input to two 3×3 convolutions for feature map extraction to obtain the feature map .
[0024] S3 is as follows: The optimized model is obtained based on the initial model by replacing one standard convolution layer and one depthwise separable convolution layer in the initial model with RefConv convolution; The optimized CNN-Transformer model consists of Block RBConv shallow feature extraction module and The block Sim-Transformer deep feature extraction modules are stacked and connected in series; Among them, the RBConv shallow feature extraction module includes: RefConv re-parameterized refocusing convolution and SE residual connection module; The SE residual connection module includes: a global average pooling layer, a fully connected layer, and an activation function for generating channel weights; The Sim-Transformer module includes: Sim-Attntion module and feed-forward neural network FFN.
[0025] S4 is as follows: S4.1. Feature Map Before being input into the optimized CNN-Transformer model, it is first passed through a Convolution, the feature map The number of channels C is expanded by 4 times, and the feature map The original number of channels is C, and the number of channels after expansion is 4C; S4.2. Feature map after channel expansion Input to the RBConv shallow feature extraction module of the CNN-Transformer model, feature map Output feature map after RefConv ; Then the basis weights of the initial CNN-Transformer model output Perform refocusing transformation to obtain conversion weight , the calculation process is as follows: , in, represents the refocusing transformation, Representation feature map The trainable parameters of Then the feature map Input to the SE residual module, and the input feature map is processed by the SE module Perform global average pooling, the calculation process is as follows: , in, Indicates channel The features of the global average pooling output are of size 1×1×4C. Indicates height, Indicates width, Representation feature map Height × width × channels, Representation feature map The index in the height direction, , Representation feature map Index in the width direction, ,aisle The number of channels is 4C; The features of the global average pooling output Input to the fully connected layer, through the fully connected layer The activation function gets the channel The weight value , the calculation process is as follows: , The input feature map Channel Multiply it with the corresponding weight to generate a new feature map , the calculation process is as follows: , Among them, the feature map The dimensions are H×W×4C; Final feature map Then a 1×1 convolution is used to reduce the number of channels to 1 / 4, and the number of channels after reduction is C; S4.3. Feature map Input to Sim-Transformer module, feature map It is a matrix composed of multiple neurons, each neuron corresponds to a receptive field, and the energy function is defined for finding important neurons , the calculation process is as follows: , , , , in, Representation feature map The target neuron of a single channel in Indicates that except for the target god general External feature map Other neurons in a single channel, represents the number of energy functions, , , represents the total number of energy functions in each channel, represents the transformation weight, represents the transformation bias, and Represents the feature maps Remove target neurons from a single channel All other neurons The mean and variance of Calculate minimum energy , according to the minimum energy Determine the importance, The lower the target neuron The greater the difference from the surrounding neurons, the higher the importance and the minimum energy The calculation process is as follows: , in, and denote the mean and variance of the neuron respectively, represents the regularization coefficient; Then the feature map The energy function value The form propagates downward, and the energy values of all neurons form a set ; Continue to map the feature map Input to the Sim-Attntion module, through Neuron weights and feature maps obtained by activation function Perform dot product operation and then transform the feature map The input value is mapped to the range of (0,1), and the output is a refined feature map focusing on the key target perception area. , the calculation process of the Sim-Attntion module is as follows: , The feature map Input to the feedforward neural network FFN, feature map First input to the fully connected layer to get the feature map , the calculation process is as follows: , in, Represents the activation function Operation, represents the feature dimension, represents fixed parameters; Then Perform layer normalization to obtain the final output feature map , the calculation formula is as follows: , in, Represents the feature map pair Each layer channel is normalized one by one.
[0026] S5 is as follows: The feature map is processed by the average pooling layer and the fully connected layer. To classify; For the input feature map Perform global average pooling on each channel of: , Among them, the input feature map The size is × × C ,in and are the height and width of the feature map, respectively. It is a channel The features of the global average pooling output, channel The number of channels is , size is 1×1× C ; The average pooled vector Through the fully connected layer, it is mapped to the category space to obtain the probability distribution of each category, from which the category label with the highest probability is selected as the final category label of the coal rock image. Assuming that there are The specific process is as follows: , in, Represents the weight matrix of the fully connected layer, with size , represents the bias vector, Represents the score of each category, with size , , Indicates The scores of the categories, ; Then apply the output of the fully connected layer Activation function to obtain the probability distribution of each category: , in, Indicates that the input image belongs to the category probability; The classification probability of each category is calculated respectively, and the category corresponding to the highest probability value is recorded as the category of the input image.
[0027] The present invention combines CNN and Transformer. The convolution layer often has better generalization and faster convergence speed due to its strong inductive bias prior, while the attention layer in Transformer has a higher model capacity and can benefit from a larger data set. The combination of the convolution layer and the attention layer can obtain better generalization and capacity, and adopts Refconv, which can greatly reduce the parameters while reducing the computational redundancy and memory access, and has a high classification accuracy that exceeds similar ones, making it possible to realize the recognition of coal and rock images on mobile devices, and can play an important role in the complex coal mine environment.
[0028] In order to verify the effectiveness of the method, the method of the present invention is compared with the existing method under the same conditions. As shown in Table 1, this paper compares thirteen models, which can be divided into three categories of classification methods, namely, image classification models with pure convolutional (CNN) architecture (ResNet-101 and NFNet-F3); variant image classification models of Vision Transformer (ViT) (DeiT-B and ViT-L / 16); and various models using CNN-Transformer hybrid architecture (MobileViT-S, ConFormer-S, T2T-ViT-24, FastVit-SA36, Next-ViT-S, DeepMAD-50M, CoAtNet-0), including the reference model CoAtNet-0 (removing the RBConv shallow feature extraction module and the Sim-Transformer deep feature extraction module); in terms of Params model parameter quantity, FLOPs floating-point operation number and Top-1 Accuracy The three aspects of accuracy are compared. From the visualization results, it can be seen that the model of the present invention achieves 91.4% accuracy without pre-training, achieving the highest accuracy, which is 7.4% higher than the baseline. While the number of parameters has not increased significantly, FLOPs has been reduced by 19.0%. After pre-training, the model of the present invention has achieved the highest accuracy of 95.72%, which indicates a significant improvement in data efficiency and computing efficiency. Table 1 Performance comparison results of our method and existing methods on coal and rock datasets Example 2 A coal-rock image classification device based on CNN-Transformer combination, which implements a coal-rock image classification method based on CNN-Transformer combination, includes the following units: 1) Pre-training learning unit: The initial CNN-Transformer model is used to process the existing data with labeled image types to complete the pre-training of the initial CNN-Transformer model; 2) Coal and rock image acquisition unit: collects data used to construct coal and rock image datasets and preprocesses the datasets; 3) Optimization unit: replace part of the structure in the initial CNN-Transformer model to obtain the optimized CNN-Transformer model; 4) Image data processing unit: The collected image is processed by the optimized CNN-Transformer model and the parameters output by the initial CNN-Transformer model to obtain the feature map of the input image; 5) Classification unit: classifies the input image according to the feature map obtained by the image data processing unit.
[0029] Example 3 An electronic device comprises a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the coal-rock image classification method based on the combination of CNN and Transformer is executed; The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0030] Example 4 In order to further prove the effectiveness of the method of the present invention in extracting feature information from coal rock images, Figure 2The focus and attention information heat maps of the method of the present invention and the existing method CoAtNet-0 on the test set are respectively shown, wherein the blue area belongs to the key attention area of the network. It can be clearly seen from the heat map that the method of the present invention can accurately focus on the key feature area in the image and effectively extract detailed information for analysis. In contrast, although CoAtNet-0 also demonstrates a certain feature capture capability, its accuracy in local features is slightly lower than that of the model of the present invention. The above observation results further demonstrate the advantages of the method of the present invention in coal and rock image analysis, and provide strong support for future practical applications.
[0031] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A coal-rock image classification method based on CNN-Transformer combination, characterized in that , the steps are as follows: S1. Construct the initial deep network model CNN-Transformer, input the data in the texture library dataset into the initial CNN-Transformer model, and obtain the basic weights. ; S2. Collect coal and rock images and filter them to obtain a data set , for the dataset Preprocessing is performed to obtain the preprocessed data set , and then the preprocessed data set The image is processed preliminarily to obtain the feature map ; S3. Optimize the initial CNN-Transformer model by replacing a standard convolution and a depth-separable convolution layer in the initial CNN-Transformer model to obtain an optimized CNN-Transformer model. S4. Feature map Input into the optimized CNN-Transformer model to obtain the feature map; S5. Input the results of the optimized CNN-Transformer model into the classifier to obtain the classification result.
2. According to the coal-rock image classification method based on CNN-Transformer combination as described in claim 1, it is characterized in that: S1 is as follows: The structure of the initial deep network model CNN-Transformer is as follows: Block RBConv shallow feature extraction module, Block Sim-Transformer deep feature extraction module; Among them, the RBConv shallow feature extraction module includes: a standard convolution layer, a depth-separable convolution layer, and a SE residual connection module; The Sim-Transformer deep feature extraction module includes: Sim-Attntion module, feed-forward neural network FFN; The texture library dataset is an existing dataset in which images have been classified and labeled. After the data in the texture library dataset is input into the initial CNN-Transformer model, the feature map of the image in the texture library dataset is first extracted, and then the extracted feature map is processed by the initial CNN-Transformer model to complete the training of the initial CNN-Transformer model. Finally, the convolution kernel finally output by the initial CNN-Transformer model is frozen as the basis weight .
3. According to claim 2, a coal-rock image classification method based on CNN-Transformer combination is characterized in that: S2 is as follows: For the dataset Preprocess the images in the dataset, scale them to the same size, and increase the number of images in the dataset by random cropping and rotation operations. Finally, the preprocessed dataset is obtained. ; Dataset Expressed as , the preprocessed dataset Expressed as ,in, , Representation dataset The number of images in Represents the preprocessed dataset The number of images in Representation dataset Middle images, Represents the preprocessed dataset Middle images, Represents the preprocessed dataset Any image in After preprocessing the data set The image is processed initially. Input to two 3×3 convolutions for feature map extraction to obtain the feature map .
4. According to claim 3, a coal-rock image classification method based on CNN-Transformer combination is characterized in that: S3 is as follows: The optimized model is obtained based on the initial model by replacing one standard convolution layer and one depthwise separable convolution layer in the initial model with RefConv convolution; The optimized CNN-Transformer model consists of Block RBConv shallow feature extraction module and The block Sim-Transformer deep feature extraction modules are stacked and connected in series; Among them, the RBConv shallow feature extraction module includes: RefConv re-parameterized refocusing convolution and SE residual connection module; The SE residual connection module includes: a global average pooling layer, a fully connected layer, and an activation function for generating channel weights; The Sim-Transformer module includes: Sim-Attntion module and feed-forward neural network FFN.
5. According to claim 4, a coal-rock image classification method based on CNN-Transformer combination is characterized in that: S4 is as follows: S4.
1. Feature Map Before being input into the optimized CNN-Transformer model, it is first passed through a Convolution, the feature map The number of channels C is expanded by 4 times, and the feature map The original number of channels is C, and the number of channels after expansion is 4C; S4.
2. Feature map after channel expansion Input to the RBConv shallow feature extraction module of the CNN-Transformer model, feature map Output feature map after RefConv ; Then the basis weights of the initial CNN-Transformer model output Perform refocusing transformation to obtain conversion weight , the calculation process is as follows: , in, represents the refocusing transformation, Representation feature map The trainable parameters of Then the feature map Input to the SE residual module, and the input feature map is processed by the SE module Perform global average pooling, the calculation process is as follows: , in, Indicates channel The features of the global average pooling output are of size 1×1×4C. Indicates height, Indicates width, Representation feature map Height × width × channels, Representation feature map The index in the height direction, , Representation feature map Index in the width direction, ,aisle The number of channels is 4C; The features of the global average pooling output Input to the fully connected layer, through the fully connected layer The activation function gets the channel The weight value , the calculation process is as follows: , The input feature map Channel Multiply it with the corresponding weight to generate a new feature map , the calculation process is as follows: , Among them, the feature map The dimensions are H×W×4C; Final feature map Then a 1×1 convolution is used to reduce the number of channels to 1 / 4, and the number of channels after reduction is C; S4.
3. Feature map Input to Sim-Transformer module, feature map It is a matrix composed of multiple neurons, each neuron corresponds to a receptive field, and the energy function is defined for finding important neurons , the calculation process is as follows: , , , , in, Representation feature map The target neuron of a single channel in Indicates that except for the target god general External feature map Other neurons in a single channel, represents the number of energy functions, , , represents the total number of energy functions in each channel, represents the transformation weight, represents the transformation bias, and Represents the feature maps Remove target neurons from a single channel All other neurons The mean and variance of Calculate minimum energy , according to the minimum energy Determine the importance, The lower the target neuron The greater the difference from the surrounding neurons, the higher the importance and the minimum energy The calculation process is as follows: , in, and denote the mean and variance of the neuron respectively, represents the regularization coefficient; Then the feature map The energy function value The form propagates downward, and the energy values of all neurons form a set ; Continue to map the feature map Input to the Sim-Attntion module, through Neuron weights and feature maps obtained by activation function Perform dot product operation and then transform the feature map The input value is mapped to the range of (0,1), and the output is a refined feature map focusing on the key target perception area. , the calculation process of the Sim-Attntion module is as follows: , The feature map Input to the feedforward neural network FFN, feature map First input to the fully connected layer to get the feature map , the calculation process is as follows: , in, Represents the activation function Operation, represents the feature dimension, represents fixed parameters; Then Perform layer normalization to obtain the final output feature map , the calculation formula is as follows: , in, Represents the feature map pair Each layer channel is normalized one by one.
6. According to claim 7, a coal-rock image classification method based on CNN-Transformer combination is characterized in that: S5 is as follows: The feature map is processed through the average pooling layer and the fully connected layer. To classify; For the input feature map Perform global average pooling on each channel of: , Among them, the input feature map The size is × × C ,in and are the height and width of the feature map, respectively. It is a channel The features of the global average pooling output, channel The number of channels is , size is 1×1× C ; The average pooled vector Through the fully connected layer, it is mapped to the category space to obtain the probability distribution of each category, from which the category label with the highest probability is selected as the final category label of the coal rock image. Assuming that there are The specific process is as follows: , in, Represents the weight matrix of the fully connected layer, with size , represents the bias vector, Represents the score of each category, with size , , Indicates The scores of the categories, ; Then apply the output of the fully connected layer Activation function to obtain the probability distribution of each category: in, Indicates that the input image belongs to the category The probability of The classification probability of each category is calculated respectively, and the category corresponding to the highest probability value is recorded as the category of the input image.
7. A coal-rock image classification device based on CNN-Transformer combination, characterized in that: The method for classifying coal and rock images based on CNN-Transformer combination according to any one of claims 1 to 6 comprises the following units: 1) Pre-training learning unit: The initial CNN-Transformer model is used to process the existing data with labeled image types to complete the pre-training of the initial CNN-Transformer model; 2) Coal and rock image acquisition unit: collects data used to construct coal and rock image datasets and preprocesses the datasets; 3) Optimization unit: replace part of the structure in the initial CNN-Transformer model to obtain the optimized CNN-Transformer model; 4) Image data processing unit: The collected image is processed by the optimized CNN-Transformer model and the parameters output by the initial CNN-Transformer model to obtain the feature map of the input image; 5) Classification unit: classifies the input image according to the feature map obtained by the image data processing unit.
8. An electronic device, characterized in that: It includes a processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, a coal rock image classification method based on the CNN-Transformer combination as described in any one of claims 1 to 6 is performed.
Citation Information
Patent Citations
Image classification method and device in combination with CNN and Transform, and computer storage medium
CN116912552A
Lightweight maritime complex target detection method for unmanned ship
CN118366031A
Transform-based coal rock image super-resolution reconstruction method and device
CN118411291A
Image classification model construction method, image classification method, image classification device, image classification equipment and storage medium
CN118675000A
Joint modeling method and apparatus for enhancing local features of pedestrians
WO2024060321A1