Drilling rock core image recognition method based on improved ResNet
By improving the ResNet network, combining the three-dimensional attention mechanism and self-calibration convolution module, the problem of insufficient accuracy and generalization performance in core image recognition is solved, and efficient and accurate core image recognition is achieved, supporting efficient geological surveying.
Patent Information
- Application Number
- CN202510499086.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing core image recognition methods have poor accuracy under different geological environments, and there are small differences between core image categories and large intra-class differences, resulting in inaccurate recognition effects and poor generalization performance.
The improved ResNet network is adopted, and a self-calibrated convolution module and MAM pooling method with a three-dimensional attention mechanism are introduced to enhance the model's adaptability to complex features of core images, and to improve the recognition performance through data augmentation and multi-scale feature pooling.
It significantly improves the classification performance of core images, assists geological surveyors to accurately identify cores, unify geological cataloging standards, improves geological survey efficiency and accuracy, and reduces the impact of artificial differences.
Smart Images

Figure CN120339715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of image recognition and geological exploration, and specifically relates to a method for identifying drilling core images based on an improved ResNet. Background Art
[0002] During geological exploration and mineral resource development, cores are important carriers for obtaining underground geological information. By analyzing drilling cores, the structure, composition, and physical and chemical properties of underground rock formations can be accurately understood, providing a scientific basis for mineral deposit evaluation, engineering construction, and resource planning. However, traditional core analysis mainly relies on manual observation and recording, which is not only time-consuming and laborious but also easily affected by subjective factors and human differences, resulting in inaccurate identification results.
[0003] With the rapid development of deep learning technology, rock analysis methods based on image recognition have gradually attracted attention. In early studies on rock and lithology recognition, the principal component analysis method (PCA) was mostly used. By reducing the dimension of lithology parameters, effective features related to lithology were extracted. However, the PCA method cannot effectively handle non-linear and complex lithology features, resulting in limited recognition effects in complex geological environments. Convolutional neural networks (CNNs) can capture local features of images through structures such as convolutional layers, pooling layers, and fully connected layers, and can also extract abstract high-level features layer by layer, thus achieving relatively significant results in various rock recognition tasks. Residual networks (ResNets) solve the problem of gradient disappearance and degradation in deep neural networks by introducing residual connections, improving the training efficiency and accuracy of the network. The Faster R-CNN architecture based on ResNet effectively retains the detailed information of the original image using residual learning, improving the detection accuracy of rock images. The deep network based on Inception-v3 can effectively identify the texture features of rock images and achieve efficient classification of rocks. The modular design of the Branch Module based on the AlexNet network not only improves the accuracy of rock recognition but also has higher computational efficiency.
[0004] However, in the identification of drilling core images, many challenges still remain. First, the core image data collected in different geological environments vary greatly. Second, core images have the characteristics of small differences between categories and large differences within categories. Finally, the data categories of core images are unbalanced. Therefore, when extending the above lithology analysis methods to core image recognition, problems such as poor image recognition accuracy and poor generalization performance caused by insufficient model pertinence occur. Recently, some researchers have proposed a joint recognition method based on multiple training models, that is, an ensemble learning method, to identify drilling cores. The recognition effect has been improved to a certain extent compared with a single model, but the training and test core data only come from 220 cores of 15 categories. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method for identifying drilling core images based on an improved ResNet. By introducing a self-calibrating convolution module based on a three-dimensional attention mechanism and a MAM pooling method, the ResNet network structure is improved, making up for the deficiencies of the original ResNet in effectively capturing high-level semantic information due to being limited by fixed convolution kernels and receptive fields when extracting subtle features, and the difficulty in coordinating between spatial and channel attention. It significantly enhances the adaptability of the model to the complex features of core images and the classification performance, providing efficient and reliable technical support for practical application scenarios such as geological exploration and resource evaluation.
[0006] To achieve the above object, the present application is implemented through the following technical solutions:
[0007] The present application is a method for identifying drilling core images based on an improved ResNet, and the method specifically includes the following steps:
[0008] Step 1, collect and store the image dataset: Take drilling core images through an image acquisition device, and upload and store the captured drilling core images to the server in real time;
[0009] Step 2, classify the drilling core images collected in Step 1, assign a unique code to each type of drilling core category, and establish a preliminary sample annotation set;
[0010] Step 3, construct a drilling core dataset: Delete invalid samples or fuzzy samples from the sample annotation set, and perform data augmentation operations on the drilling core images after deleting invalid samples or fuzzy samples, including image rotation and cropping, and then annotate the data-augmented drilling core images and construct a high-quality drilling core dataset;
[0011] Step 4, construct a drilling core image recognition model. The drilling core image recognition model is to embed a self-calibrating convolution module based on a three-dimensional attention mechanism in the basic recognition model ResNet50, and introduce a multi-scale feature pooling method in the feature extraction stage;
[0012] Step 5, train the drilling core image recognition model constructed in Step 4, and test the trained drilling core image recognition model, that is, input the test set into the drilling core image recognition model, and examine the transfer ability and generalization ability of the trained drilling core image recognition model to ensure that the drilling core image recognition model is not only suitable for the training dataset, but also applicable to other core datasets.
[0013] A further improvement of the present application lies in: The specific classification method in Step 2 includes the following steps:
[0014] Step 2.1: Divide the drilling core images into clay, silt, sand and gravel, and sandstone according to the mineral composition, texture characteristics and structure of the drilling cores. Define the set of drilling core categories C = {C1, C2, …, C n}, where n is the number of types of drilling cores;
[0015] Step 2.2: Identify the sub-categories of each type of drilling core. For each type of drilling core C i ∈C, 1 ≤ i ≤ n, divide it into the set of sub-categories of drilling cores C i = {C i1 , C i2 , …, C im}, where m is the number of sub-categories of drilling cores;
[0016] Step 2.3: Encode all the categories of drilling cores. The number of core types is n × m.
[0017] A further improvement of this application is that the construction of the drilling core data set in Step 3 specifically includes the following steps:
[0018] Step 3.1: Denote the sample annotation set as ImageSet = {image1, image2, …, image n};
[0019] Step 3.2: Preprocess the sample annotation set imageSet in Step 3.1, that is, delete invalid images or blurred images from the sample annotation set ImageSet;
[0020] Step 3.3: Enhance the sample annotation set ImageSet by means of image rotation, image pixel value normalization and color adjustment.
[0021] A further improvement of this application is that in Step 3.2, preprocessing the sample annotation set ImageSet, that is, deleting invalid images or blurred images from the sample annotation set ImageSet, specifically includes the following steps:
[0022] Step 3.2.1: Images in the sample annotation set ImageSet with a resolution lower than the set threshold are invalid images. Set the threshold and delete the invalid images with a resolution lower than the set threshold;
[0023] Step 3.2.2: Uniformly convert the drilling core images in the sample annotation set ImageSet after deleting the invalid images into RGB color images and store them in a common image format. If the drilling core image is damaged or cannot be converted, delete the damaged or unconverted drilling core image;
[0024] Step 3.2.3: Set the blurriness threshold vague_thres, and perform grayscale conversion on the i-th drilling core image image i ;
[0025] Step 3.2.4: Calculate the Laplacian variance value of the i-th drilling core image image i as the blurriness v i . If the blurriness v i <vague_thres, then delete the blurry drilling core image imge i .
[0026] A further improvement of the present application is that the construction method of the drilling core image recognition model in step 4 specifically includes the following steps:
[0027] Step 4.1: Select LeakyReLU as the activation function to replace the activation function of the basic recognition model ResNet50;
[0028] Step 4.2: On the basis of step 4.1, embed a self-calibrating convolution module based on a three-dimensional attention mechanism to enhance the attention of the drilling core image recognition model to the key region features, improve the network feature transformation ability without adding additional learnable parameters, divide the convolution kernel into multiple parts, and send them into two different scale spaces for feature transformation respectively. Use the message passing mechanism inside the features to break the conventional mechanism of using small-size kernels for information fusion and extraction in the channel and spatial dimensions, so as to more effectively capture the context information around each spatial position;
[0029] Step 4.3: On the basis of step 4.2, introduce a three-dimensional attention module combined with the SimAM mechanism to make up for the problem of ignoring the interaction between channels and spaces when serially connecting the spatial dimension and the channel dimension in the three-dimensional attention mechanism, and avoid losing important span information;
[0030] Step 4.4: Introduce multi-scale feature pooling in the feature extraction stage, comprehensively analyze the multi-scale information in the drilling core image through pooling operations of different scales, make the exploration core image recognition model adapt to the diverse features of complex images, and improve the recognition ability for different sizes and morphological features.
[0031] A further improvement of the present application is that step 4.2 specifically includes the following steps:
[0032] Step 4.2.1: The feature map size of the input drilling core image is C×H×W. It is split into two image features X1 and X2 with a size of C / 2×H×W. The obtained image feature X1 is sent into the original feature space and the latent space respectively, and the image feature X2 is sent into the original feature space to collect different types of context information. The dimension of the convolution kernel K is C×H×W, and there are 4 convolution kernels in total, denoted as K1, K2, K3, and K4 respectively. The dimension of each convolution kernel is C / 2×H×W.
[0033] Step 4.2.2: In the latent space, first, average pooling is used to downsample the image feature X1 by 4 times to obtain the downsampled image T1: T1 = AvgtPool r (X1). Then, bilinear interpolation operation is used to perform upsampling to map the small-scale space to the original feature space to obtain the weight X'1 for calibration:
[0034] X′1 = Up(F2(T1)) = Up(T1*K2),
[0035] where * represents the convolution operation, and F2 = T1*K2;
[0036] Secondly, a residual structure is constructed and added, and after passing through the sigmoid activation function σ, the latent space feature M1 is obtained:
[0037] M1 = σ(X1 + X′1)
[0038] In the original space, the image feature K1 is sent into the K3 convolution kernel, and the latent space feature M1 obtained from the latent space is used to calibrate the image feature X1 sent into the K3 convolution kernel to obtain the calibrated final scale space output feature Y1:
[0039] Y1 = F4(Y′1) = K4*Y′1,
[0040] Y′1 = F3(X1)·M1 = (X1*K3)·M1
[0041] where F3(X1) = X1K3, and · is element-wise multiplication.
[0042] Step 4.2.3: Processing in the original space: The image feature X2 is convolved through the K1 convolution kernel to obtain the scale space output feature Y2 = K1*X2;
[0043] Step 4.2.4: The final scale space output feature Y1 in Step 4.2.2 and the scale space output feature Y2 in Step 4.2.3 are concatenated to obtain the final output feature Y = Cat(Y1, Y2), where Cat represents the concatenation operation of the feature maps.
[0044] A further improvement of the present application lies in: in step 4.3, on the basis of step 4.2, a three-dimensional attention module combined with the SimAM mechanism is introduced, which specifically includes the following steps:
[0045] Step 4.3.1: Define a function to evaluate the importance of neurons:
[0046]
[0047] Among them, M represents the number of neurons in each channel, and M = H × W, t represents the target neuron, and x i represents other neurons on the same channel as the target neuron t, represents performing a linear transformation on the target neuron t, represents performing a i linear transformation on x w t , b t are the weights and biases assigned during the linear transformation, and y o and y t represent binary labels;
[0048] Step 4.3.2: After obtaining the importance of neurons, use the SimAM attention mechanism to enhance the image feature X1 to obtain a new feature map
[0049]
[0050]
[0051] Among them, E represents the set of all energy values in the input feature map, ⊙ represents dot product calculation, and a Sigmoid function is added to limit the impact of overly large values in E on the whole. A further improvement of the present application lies in: step 4.4 specifically includes the following steps:
[0052] Step 4.4.1: Design a multi-scale pooling method: If the mean of the pooling window is higher than the weighted mean, that is, the weighted value of the mean plus the standard deviation, then choose max pooling to retain the more prominent features within the window. If the mean of the pooling window is lower than the weighted value of the mean minus the standard deviation, then use min pooling to better capture subtle features; when the mean of the pooling window is between the two, then use average pooling to synthesize the feature information within the window; this pooling strategy can flexibly adjust the pooling method according to the performance of different features, extract more accurate and comprehensive image features, and only increase a small amount of computational overhead, thereby improving the robustness and accuracy of the model;
[0053] Step 4.4.2: In the feature extraction stage, adopt the multi-scale feature pooling method designed in step 4.4.1, and combine the results of the multi-scale pooling operation with the new feature map in a weighted manner Perform fusion.
[0054] A further improvement of this application lies in: The training in step 5 specifically includes the following steps:
[0055] Step 5.1: Divide the drilling core dataset in step 3 into a training set, a validation set, and a test set. The training set is used for training the drilling core image recognition model, the validation set is used to evaluate the training effect of the drilling core image recognition model to prevent overfitting, and the test set is used to detect the transfer ability and generalization ability of the drilling core image recognition model.
[0056] Step 5.2 Initialize the parameters of the drilling core image recognition model, and input the training set and the validation set to optimize the parameters of the drilling core image recognition model through learning.
[0057] The beneficial effects of this application are: An identification model based on the improved ResNet proposed in this application has achieved remarkable core image classification performance, and can be used to assist geological surveyors to efficiently and accurately identify cores.
[0058] Using the identification model of this application can not only prevent the uneven accuracy of the identification results caused by the differences in the professional levels of surveyors, thereby unifying the geological logging standards and improving the accuracy of stratigraphic division; but also avoid the difficulty in effectively tracing and proofreading the accuracy of geological stratification caused by the timeliness of core retention, thereby improving the risk control level of geological surveys; and can also improve the survey efficiency, providing efficient and reliable technical support for data processing and decision-making in geological surveys. Description of the Drawings
[0059] Figure 1 is the flow chart of the identification method of this application.
[0060] Figure 2 is the structural diagram of the self-calibrating convolution module based on the three-dimensional attention mechanism of this application.
[0061] Figure 3 is the structural diagram of the drilling core image recognition model of this application.
[0062] Figure 4 is the verification accuracy curve diagram of the drilling core image recognition model of this application on the validation dataset. Detailed Implementation Modes
[0063] The embodiments of the present invention will be disclosed below with reference to the drawings. For the sake of clarity, many practical details will be described together in the following description. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are unnecessary. In addition, for the sake of simplifying the drawings, some conventional structures and components will be shown in the drawings in a simple schematic manner.
[0064] As Figure 1 shown, the present application is a method for identifying drilling core images based on an improved ResNet. The method for identifying drilling core images specifically includes the following steps:
[0065] Step 1, collecting and storing an image data set: Taking drilling core images through an image acquisition device, and uploading and storing the taken drilling core images to a server in real time. On-site survey personnel take pictures of the drilled cores section by section through professional auxiliary devices such as high-definition cameras or smartphones carried. Each photo needs to contain the complete structure of the core, and then upload the taken core images to the server in the data center in real time. The server automatically generates a unique number for each image and stores it in the database, and at the same time records meta-information such as the upload time and the location of the survey point to ensure data traceability.
[0066] Step 2, classifying the drilling core images collected in Step 1, assigning a unique code to each type of drilling core category, and having technicians conduct a preliminary manual classification of the collected core images to establish a preliminary sample annotation set to provide standard data for subsequent model training. The specific classification method includes the following steps:
[0067] Step 2.1, dividing the drilling core images into clay, silt, sand and gravel, and sandstone according to the mineral composition, texture characteristics and structure of the drilling core, and defining the set of drilling core categories C = {C1, C2, …, C n}, where n is the number of types of drilling cores;
[0068] Step 2.2, clarifying the sub-categories of each type of drilling core. For each type of drilling core C i ∈ C, 1 ≤ i ≤ n, it is divided into the set of sub-categories of drilling cores C i = {C i1 , C i2 , …, C im}, where m is the number of sub-categories of drilling cores;
[0069] Step 2.3, encoding all drilling core categories, and the number of core types is n × m.
[0070] Step 3: Construct a drilling core dataset. Delete invalid or blurred samples from the collected drilling core images, and perform data augmentation operations on the drilling core images after deleting invalid or blurred samples, including image rotation and cropping. Then annotate the data-augmented drilling core images and construct a high-quality drilling core dataset. The construction method is as follows:
[0071] First, based on the core image samples in different geological environments in the preliminary sample annotation set, the images include core images at different time periods and different cross-sections after drilling. By recording the performance of the core under different humidity, angles, and background interferences, it is ensured that the dataset covers the diversity in the actual collection environment, with a total of 3,106 images.
[0072] Second, to ensure image quality, eliminate invalid or blurred images. Specifically, delete images with an original resolution lower than 1024×1024; then uniformly convert the images to RGB color images and store them as.jpg or.png. If the image is damaged or cannot be converted, delete it; then set the blurriness threshold vague_th res = 120, and perform grayscale conversion on the image image i ; calculate the Laplacian variance value of image i as the blurriness v i ; if v i < 120, then delete the picture image i .
[0073] Enhance the image set by image rotation, color adjustment, and adding noise. First, rotate the core image 10 degrees clockwise or counterclockwise, then horizontally flip the core image symmetrically left and right, then adjust the brightness, contrast, and saturation of the image through the scaling factor, and finally add Gaussian noise to the core image to form an enhanced image set, with a total of 15 categories, including 2-1A, 2-1B, 3-1, 3-2, etc., and the number of each category is about 200.
[0074] Step 4: Construct a drilling core image recognition model. The drilling core image recognition model is to embed a self-calibrating convolution module based on a three-dimensional attention mechanism in the basic recognition model ResNet50 and introduce a multi-scale feature pooling method in the feature extraction stage.
[0075] Select a basic image recognition network suitable for core images. Respectively, use AlexNet, GoogleNet, Vgg16, Vision-Transformer, and ResNet networks with different depths to test their recognition effects on the drilling core image dataset using precision. The specific comparison tests are shown in Table 1.
[0076] Table 1
[0077]
[0078] It is found from the comparative experiment that the effect of ResNet50 is the best, so ResNet50 is selected as the basic recognition model.
[0079] Improve the ResNet50 model. First, a comparative experiment is conducted on the combination of the basic residual network and different activation functions, including ReLU, PreLU, ELU, GELU, and LeakyReLU. The experimental results are shown in Table 2.
[0080] Table 2
[0081]
[0082] It is found from the experimental results that the present application selects LeakyReLU as the activation function
[0083] Specifically, as Figure 2 shown, the method for constructing a drilling core image recognition model includes the following steps:
[0084] Step 4.1: Select LeakyReLU as the activation function to replace the activation function of the basic recognition model ResNet50;
[0085] Step 4.2: While retaining the original 3×3 standard convolution branch of ResNet50, a self-calibration branch is added. This branch adopts a multi-path design. Path 1 obtains low-resolution but larger receptive field context information through downsampling to form a latent space representation; Path 2 maintains the original resolution for local feature extraction. The outputs of the two paths are added and then fused through convolution to generate a calibrated feature map. To suppress the interference of background noise during the self-calibration process, the SimAM attention mechanism is introduced into the latent space of Path 1, and the energy function is calculated through three-dimensional cross-dimensional interaction to dynamically weight key features and suppress irrelevant information. Finally, the module adds the outputs of the standard convolution branch and the self-calibration branch to form enhanced features with both local details and global context. This structure significantly improves the feature discriminability through dual-path complementarity and attention screening. The module structure is shown in the appendix Figure 2 Compare the model performance before and after the embedding module; replace the pooling method of the intermediate feature extraction layer with the MAM pooling method to enhance the adaptability to features of different sizes in the core image through the fusion of multiple-scale pooling operations; the overall model architecture is shown in the appendix Figure 3 Specifically, it includes the following steps:
[0086] Step 4.2.1: The feature map of the input drilling core image has a size of C×H×W. It is split into two image features X1 and X2 with a size of C / 2×H×W. The obtained image feature X1 is sent into the original feature space and the latent space respectively, and the image feature X2 is sent into the original feature space to collect different types of context information. The dimension of the convolutional kernel K is C×H×W, and there are 4 convolutional kernels in total, denoted as K1, K2, K3, and K4 respectively. The dimension of each convolutional kernel is C / 2×H×W;
[0087] Step 4.2.2: In the latent space, first, the image feature X1 is downsampled by a factor of 4 using average pooling to obtain the downsampled image T1: T1 = AvgPool r (X1). Then, bilinear interpolation operation is used for upsampling to map the small-scale space to the original feature space, obtaining the calibrated weight X'1:
[0088] X′1 = Up(F2(T1)) = Up(T1 * K2),
[0089] where * represents the convolution operation, F2 = T1 * K2X′1;
[0090] Secondly, a residual structure is constructed and added, and through the sigmoid activation function σ, the latent space feature M1 is obtained:
[0091] M1 = σ(X1 + X′1)
[0092] In the original space, the image feature K1 is sent into the K3 convolutional kernel, and the latent space feature M1 obtained from the latent space is used to calibrate the image feature X1 sent into the K3 convolutional kernel, obtaining the calibrated final scale space output feature Y1:
[0093] Y1 = F4(Y′1) = K4 * Y′1,
[0094] Y′1 = F3(X1)·M1 = (X1 * K3)·M1
[0095] where F3(X1) = X1K3, and · is the element-wise multiplication.
[0096] Step 4.2.3: Processing in the original space: The image feature X2 is convolved through the K1 convolutional kernel to obtain the scale space output feature Y2 = X2 * K1;
[0097] Step 4.2.4: The final scale space output feature Y1 in Step 4.2.2 and the final scale space output feature Y2 in Step 4.2.3 are concatenated to obtain the final output feature Y = Cat(Y1, Y2), where Cat represents the concatenation operation of the feature maps.
[0098] Step 4.3. On the basis of Step 4.2, introduce a 3D attention module combined with the SimAM mechanism to make up for the problem in the 3D attention mechanism that ignores the interaction between channels and space when serially connecting the spatial dimension and the channel dimension, and avoid losing important span information. The specific steps are as follows:
[0099] Step 4.3.1. Define a function to evaluate the importance of neurons:
[0100]
[0101] where M represents the number of neurons in each channel, and M = H × W, t represents the target neuron, and x i represents other neurons on the same channel as the target neuron t, represents performing linear transformation on the target neuron t, represents x i performing linear transformation on, w t , b t are the weights and biases assigned during the linear transformation, and y o and y t represent binary labels;
[0102] Step 4.3.2. After obtaining the importance of neurons, use the SimAM attention mechanism to enhance the image feature X1 to obtain a new feature map
[0103]
[0104] where E represents the set of all energy values in the input feature map, ⊙ represents dot product calculation, and a Sigmoid function is added to limit the impact of overly large values in E on the whole. Compare the performance of the basic ResNet model and the improved model, and record key indicators such as accuracy and recall.
[0105] Step 4.4. Introduce multi-scale feature pooling in the feature extraction stage. Through pooling operations at different scales, comprehensively analyze the multi-scale information in the drilling core image, so that the drilling core image recognition model can adapt to the diverse features of complex images and improve the recognition ability for different sizes and morphological features. The specific steps are as follows:
[0106] Step 4.4.1. Design a multi-scale pooling method: If the mean value of the pooling window is higher than the weighted mean value, which is the weighted value of the mean plus the standard deviation, then select max pooling to retain the more prominent features within the window. If the mean value of the pooling window is lower than the weighted mean value minus the weighted value of the standard deviation, then use min pooling to better capture subtle features. When the mean value of the pooling window is between the two, then use average pooling to synthesize the feature information within the window. This pooling strategy can flexibly adjust the pooling method according to the performance of different features, extract more accurate and comprehensive image features, and only increase a small amount of computational overhead, thereby improving the robustness and accuracy of the model.
[0107] Step 4.4.2. In the feature extraction stage, adopt the multi-scale feature pooling method designed in Step 4.4.1, and fuse the results of the multi-scale pooling operation with the new feature map in a weighted manner. Verify the role of multi-scale pooling in improving the model performance through ablation experiments.
[0108] Step 5. Train the drill core image recognition model constructed in Step 4, and test the trained drill core image recognition model, that is, input the test set into the drill core image recognition model to examine the transfer ability and generalization ability of the trained drill core image recognition model, and ensure that the drill core image recognition model is not only suitable for the training dataset, but also applicable to other core datasets. The training specifically includes the following steps:
[0109] Step 5.1. Divide the drill core dataset in Step 3 into a training set, a validation set, and a test set according to the ratio of 8:1:1. Among them, there are 2456 pictures in the training set, 307 pictures in the validation set, and 343 pictures in the test set. The training set is used for the training of the drill core image recognition model, the validation set is used to evaluate the training effect of the drill core image recognition model to prevent overfitting, and the test set is used to detect the transfer ability and generalization ability of the drill core image recognition model.
[0110] Step 5.2. Initialize the parameters of the drill core image recognition model, and input the training set and the validation set to optimize the parameters of the drill core image recognition model through learning. The optimizer adopts Adam, the loss function is cross-entropy loss, the batch size batch_size is 32, the learning rate lrs is 0.0001, and the number of iterations epoch is set to 200. The settings of the model parameters can be adjusted according to the actual training data. Input the training set and the validation set into the model to optimize the model parameters. The accuracy curve of the validation set is shown in Figure 3 , and the recognition accuracy is shown in Table 3 below.
[0111] Table 3
[0112]
[0113] Step 6: Test the core image model, that is, input the test set into the model, input the test data set to evaluate the recognition effect of the network, and the test effect is shown in Table 4 below.
[0114] Table 4
[0115]
[0116] The method of this application can accurately identify core images, thus effectively assisting geological surveyors in geological logging, improving the accuracy and efficiency of geological survey tasks, and providing efficient and reliable technical support for geological survey data processing and decision-making.
[0117] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. An image recognition method for drilling core based on improved ResNet, characterized in that: The described method for identifying drilling core images specifically includes the following steps: Step 1, collect and store the image dataset: Use an image acquisition device to capture drilling core images, and upload and store the captured drilling core images to the server in real time; Step 2, classify the drilling core images collected in Step 1, assign a unique code to each type of drilling core, and establish a preliminary sample annotation set; Step 3, construct the drilling core dataset: Delete invalid samples or fuzzy samples from the sample annotation set, and perform data augmentation operations on the drilling core images after deleting invalid samples or fuzzy samples, including image rotation and cropping. Then, annotate the data-augmented drilling core images and construct the drilling core dataset; Step 4, construct a drilling core image recognition model. The drilling core image recognition model is to embed a self-calibrating convolution module based on a three-dimensional attention mechanism in the basic recognition model ResNet50, and introduce a multi-scale feature pooling method in the feature extraction stage; Step 5, train the drilling core image recognition model constructed in Step 4, and test the trained drilling core image recognition model, that is, input the test set into the drilling core image recognition model to examine the transfer ability and generalization ability of the trained drilling core image recognition model.
2. The method for identifying drilling core images based on improved ResNet according to claim 1, wherein: The specific classification method in Step 2 includes the following steps: Step 2.1: Divide the drill core images into clay, silt, sand and gravel, and sandstone according to the mineral composition, texture characteristics and structure of the drill cores. Define the drill core category set C = {C1, C2, …, C n}, where n is the number of types of drill cores; Step 2.2: Define the subcategories of each type of drilling core. For each type of drilling core C i ∈ C, where 1 ≤ i ≤ n, divide it into a set of subcategories of drilling core C i ={C i1 , C i2 , …, C im}, where m is the number of subcategories of the drilling core; Step 2.3, encode all types of drilling cores. The number of core types is n×m.
3. The method for identifying drilling core images based on the improved ResNet according to claim 1, characterized in that: The specific steps for constructing the drilling core dataset in Step 3 include the following steps: Step 3.1: Denote the sample annotation set as ImageSet = {image1, image2, …, image n}; Step 3.2, preprocess the sample annotation set imageSet in Step 3.1, that is, delete invalid images or fuzzy images from the sample annotation set ImageSet; Step 3.3, enhance the sample annotation set ImageSet by means of image rotation, image pixel value normalization, and color adjustment.
4. The method for identifying drilling core images based on the improved ResNet according to claim 3, wherein: The preprocessing of the sample annotation set ImageSet in Step 3.2, that is, deleting invalid images or fuzzy images from the sample annotation set ImageSet, specifically includes the following steps: Step 3.2.1, an image with a resolution lower than the set threshold in the drilling core images of the sample annotation set ImageSet is an invalid image. Set the threshold and delete the invalid images lower than the set threshold; Step 3.2.2, uniformly convert the drilling core images in the sample annotation set ImageSet after deleting invalid images into RGB color images and store them in a common image format. If a drilling core image is damaged or cannot be converted, delete the damaged or unconverted drilling core image; Step 3.2.
3. Set the blurriness threshold vague_thres, and grayscale the i-th drilling core image image i Gray-scale it; Step 3.2.4, calculate the Laplacian variance value of the \(i\)-th drilling core image image i as the blur degree \(v\) i , if the blur degree \(v\) i < \(vague\_thres\), then delete the blurred drilling core image image i .
5. A method for identifying drilling core images based on an improved ResNet according to claim 1, characterized in that: The specific construction method of the drilling core image recognition model in Step 4 includes the following steps: Step 4.1, select LeakyReLU as the activation function to replace the activation function of the basic recognition model ResNet50; Step 4.2: On the basis of Step 4.1, embed a self-calibrated convolution module based on a three-dimensional attention mechanism to enhance the attention of the drilling core image recognition model to the features of key regions. Split the convolution kernel into multiple parts and send them into two different-scale spaces for feature transformation respectively. The message passing mechanism inside the features captures the context information around each spatial position; Step 4.3: On the basis of Step 4.2, introduce a three-dimensional attention module combined with the SimAM mechanism to make up for the problem that the interaction between channels and space is ignored when serially connecting the spatial dimension and the channel dimension in the three-dimensional attention mechanism, and avoid losing important span information; Step 4.4: Introduce multi-scale feature pooling in the feature extraction stage. Through pooling operations of different scales, comprehensively analyze the multi-scale information in the drilling core image, so that the drilling core image recognition model can adapt to the diverse features of complex images and improve the recognition ability for different sizes and morphological features.
6. The method for identifying drilling core images based on the improved ResNet according to claim 5, characterized in that: The specific steps of Step 4.2 are as follows: Step 4.2.1: The size of the feature map of the input drilling core image is C×H×W. Split it into two image features X1 and X2 with the size of C / 2×H×W. Send the obtained image feature X1 into the original feature space and the latent space respectively, and send the image feature X2 into the original feature space; The dimension of the convolution kernel K is C×H×W, and there are 4 convolution kernels in total, denoted as K1, K2, K3, K4 respectively, and the dimension of each convolution kernel is C / 2×H×W; Step 4.2.2: In the latent space, first perform 4-fold downsampling on the image feature X1 using average pooling to obtain the downsampled image T1: T1 = AvgPool r (X1), and then use bilinear interpolation operation to perform upsampling, mapping the small-scale space to the original feature space to obtain the weight X'1 for calibration: X′1 = Up(F2(T1)) = Up(T1 * K2), where, * represents the convolution operation, and F2 = T1 * K2; Secondly, construct the addition of the residual structure and obtain the latent space feature M1 through the sigmoid activation function σ: M1 = σ(X1 + X′1) In the original space, send the image feature X1 into the K3 convolution kernel, and use the latent space feature M1 obtained from the latent space to calibrate the image feature X1 sent into the K3 convolution kernel to obtain the calibrated final-scale space output feature Y1: Y1 = F4(Y′1) = K4 * Y1, Y′1 = F3(X1)·M1 = (X1 * K3)·M1 where, F3(X1) = X1K3, and · is the element-wise multiplication; Step 4.2.3: Processing in the original space: Perform a convolution operation on the image feature X2 through the K1 convolution kernel to obtain the scale space output feature Y2 = X2 * K1; Step 4.2.4: Concatenate the final-scale space output feature Y1 in Step 4.2.2 and the scale space output feature Y2 in Step 4.2.3 to obtain the final output feature Y = Cat(Y1, Y2), where Cat represents the concatenation operation of the feature maps.
7. A method for identifying drilling core images based on improved ResNet according to claim 1, characterized in that: In Step 4.3, on the basis of Step 4.2, introduce a three-dimensional attention module combined with the SimAM mechanism, which specifically includes the following steps: Step 4.3.1: Define a function to evaluate the importance of neurons: Among them, M represents the number of neurons in each channel, and M = H × W. t represents the target neuron, and x i represents the neuron on the same channel as the target neuron t, represents performing a linear transformation on the target neuron t, represents performing i a linear transformation on x, and w t , b t are the weights and biases assigned during the linear transformation. y o and y t represent binary labels; Step 4.3.
2. After obtaining the importance of the neurons, use the SimAM attention mechanism to enhance the image feature X1 to obtain a new feature map where, E represents the set of all energy values in the input feature map, ⊙ represents the dot product calculation, and the Sigmoid function is added to limit the influence of excessive values in the set E on the whole.
8. A method for identifying drilling core images based on an improved ResNet according to claim 7, characterized in that: Step 4.4 specifically includes the following steps: Step 4.4.1, Design a multi-scale pooling method: If the mean value of the pooling window is higher than the weighted mean value, that is, the weighted value of the mean plus the standard deviation, then select the maximum pooling to retain the prominent features within the window. If the mean value of the pooling window is lower than the weighted mean value minus the weighted value of the standard deviation, then use the minimum pooling to capture subtle features; when the mean value of the pooling window is between the two, then use average pooling to synthesize the feature information within the window; Step 4.4.2: In the feature extraction stage, adopt the multi-scale feature pooling method designed in Step 4.4.1, and fuse the results of the multi-scale pooling operation with the new feature map in a weighted manner. for fusion.
9. A method for identifying drilling core images based on an improved ResNet according to claim 1, characterized in that: Step 5 for training specifically includes the following steps: Step 5.1, Divide the drilling core dataset in Step 3 into a training set, a validation set, and a test set. The training set is used for training the drilling core image recognition model, the validation set is used to evaluate the training effect of the drilling core image recognition model to prevent overfitting, and the test set is used to detect the transfer ability and generalization ability of the drilling core image recognition model; Step 5.2 Initialize the parameters of the drilling core image recognition model, and input the training set and the validation set to optimize the parameters of the drilling core image recognition model through learning.