Rock slice identification method, apparatus and device, and storage medium

The proposed neural network architecture with multi-head self-attention mechanisms addresses the limitations of existing models by improving classification accuracy and identifying complex rock subtypes, enhancing the practicality and industrial applicability of AI-based rock classification systems.

CN120318578APending Publication Date: 2025-07-15INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510414790.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing deep learning models face the problem of insufficient training samples and significant reduction in classification accuracy when rock types are diversified in rock sheet image recognition, making it difficult to achieve high-precision rock subclass recognition.

Method used

A recognition model including convolutional layer, flattening layer, multi-head self-attention layer, fully connected layer and SoftMax layer is built. Pre-trained model and transfer learning technology are used, combined with multi-head self-attention mechanism to improve the classification performance of the model under a small data set.

Benefits of technology

The model's identification accuracy and generalization ability of complex rock subclasses has been significantly improved, and the automated analysis ability of rock sheet images has been enhanced, providing technical support for the practicalization and industrialization of artificial intelligence lithologic recognition systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318578A_ABST
    Figure CN120318578A_ABST
Patent Text Reader

Abstract

The invention provides a rock slice identification method and device, equipment and a storage medium. Relates to the technical field of image processing. The method comprises the following steps: constructing a first data set and a second data set; wherein the first data set and the second data set both comprise different types of rock slice images; training a neural network model by using the first data set to obtain a feature weight; constructing an identification model, wherein the identification model comprises a convolution layer, a flattening layer, a multi-head self-attention layer, a full connection layer, a SoftMax layer and a classification layer which are connected in sequence; wherein the weight of the convolutional layer is a feature weight; and training the identification model by using the second data set to obtain a trained identification model, the trained identification model being used for identifying the rock slice category in the rock slice image. According to the invention, the precision of fine classification and prediction of the rock slice image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a method, device, equipment and storage medium for rock thin section recognition. Background Art

[0002] The accurate identification of rock types is of crucial significance in fields such as geological engineering, rock mechanics, mining engineering, and oil and gas and mineral resource exploration. Traditional rock identification methods mainly rely on the experience of geologists. Through means such as naked-eye observation, magnifying glass inspection, and thin section microscopic analysis, the rock type is determined based on the composition, content, color, grain size, structure, and texture characteristics of minerals. These methods are not only time-consuming and laborious, but also highly subjective, making it difficult to achieve quantitative analysis and automated processing, which limits their efficiency and consistency in large-scale applications.

[0003] With the rapid development of image acquisition technology and computer vision technology, automatic rock type identification systems based on image analysis have gradually emerged. These systems achieve automatic lithology identification by accurately extracting apparent features related to rock types, effectively overcoming the limitations of traditional methods. However, existing research mainly focuses on the identification of macroscopic rock photos. Although good training results have been achieved with the support of large-scale datasets, there are still many challenges in the processing of rock thin section images. Due to the relatively small amount of rock thin section image data and the difficulty of fine differentiation of lithology, existing models still need to be improved in terms of classification accuracy and generalization ability.

[0004] The development of artificial intelligence in the field of rock thin section image recognition has gone through three main stages: manual feature extraction, deep learning, and transfer learning. The methods in the manual feature extraction stage classify by combining multiple features, but the accuracy of the test set is usually low. Although the convolutional neural network model introduced in the deep learning stage has improved the classification accuracy, when faced with scarce data or large changes in data distribution, the transferability and generalization ability of the model are still insufficient. The recently developed transfer learning technology has made some progress using pre-trained models on small datasets, but when the number of rock types increases or when facing complex rock subclasses, the effectiveness of the model will still decrease significantly.

[0005] Therefore, in the case of insufficient training images, how to improve the model's fine recognition ability for rock subclasses has become a key issue in current research. Solving this problem is of great significance for promoting the practicalization and industrialization of artificial intelligence lithology recognition systems. In response to the above challenges, there is an urgent need to develop a new method that can achieve high-precision rock thin section image classification under limited sample conditions to meet the needs of geological research and engineering applications. Summary of the Invention

[0006] The present application provides a method, device, equipment and storage medium for rock thin section recognition, aiming to improve the accuracy of fine classification and prediction of rock thin section images.

[0007] In a first aspect, the present application provides a method for rock thin section recognition, including:

[0008] Construct a first data set and a second data set; wherein, both the first data set and the second data set include rock thin section images of different categories;

[0009] Use the first data set to train a neural network model to obtain feature weights;

[0010] Construct a recognition model, the recognition model includes a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer and a classification layer connected in sequence; wherein, the weights of the convolutional layer are the feature weights;

[0011] Use the second data set to train the recognition model to obtain a trained recognition model, and the trained recognition model is used to identify the rock thin section category in the rock thin section image.

[0012] In a possible design, in response to an input rock thin section image, the recognition model uses the convolutional layer to extract image features, and after passing through the flattening layer, converts the image features into a one-dimensional vector. The multi-head self-attention layer processes the one-dimensional vector as an input sequence to obtain self-attention features. The fully connected layer and the SoftMax layer are used to integrate and normalize the self-attention features to obtain normalized features. The classification layer performs classification output based on the normalized features to obtain the rock thin section category.

[0013] In a possible design, the data processing process of the multi-head self-attention layer is:

[0014] Obtain an input sequence \(X\in R\) n×d ; where \(R\) is the real number field, and \(n\) and \(d\) are the dimensions of the input sequence;

[0015] Calculate a query matrix \(Q\), a key matrix \(K\) and a value matrix \(V\) through the following formula:

[0016] \(Q = XW\) Q , \(K = XW\) k , \(V = XW\) v (1)

[0017] where \(W\) Q , \(W\) k and \(W\) v are all learnable weight matrices;

[0018] Calculate an attention weight matrix \(A\) through the following formula:

[0019]

[0020] Among them, is the scaling factor, softmax is the normalization function, and T is the matrix transpose;

[0021] Based on the attention weight matrix A, calculate the self-attention output through the following formula:

[0022] Attention(Q, K, V) = AV (3)

[0023] Based on the self-attention output, calculate the self-attention feature through the following formula:

[0024] MultiHead(Q, K, V) = Concat(head1, head2, … head i , …, head h )W o (4)

[0025] Among them, MultiHead(Q, K, V) is the self-attention feature, and head i = Attention(Q i , K i , V i ), Q i , K i and V i respectively represent the i-th query matrix, key matrix, and value matrix, and head1, head2, head i and head h respectively represent the self-attention outputs of the 1st, 2nd, i-th, and h-th attention heads, and W o is the output weight matrix.

[0026] In one possible design, the rock thin section categories include sedimentary rocks, igneous rocks, and metamorphic rocks.

[0027] In one possible design, based on the stochastic gradient descent optimizer with momentum, use the second dataset to train the recognition model to obtain the trained recognition model.

[0028] In one possible design, when using the second dataset to train the recognition model, perform a validation every 3 mini-batches to closely track the performance of the model.

[0029] In a second aspect, the present application provides a rock thin section recognition device, and the device includes:

[0030] A data acquisition module, configured to construct a first data set and a second data set; wherein, both the first data set and the second data set include rock thin section images of different categories;

[0031] A first training module, configured to train a neural network model using the first data set to obtain feature weights;

[0032] A model construction module, constructing an identification model, the identification model including a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer, and a classification layer connected in sequence; wherein, the weights of the convolutional layer are the feature weights;

[0033] A second training module, using the second data set to train the identification model to obtain a trained identification model, the trained identification model being used to identify the categories of rock thin sections in rock thin section images.

[0034] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the rock thin section identification method described in the first aspect above and various possible designs of the first aspect.

[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when a processor executes the computer execution instructions, the rock thin section identification method described in the first aspect above and various possible designs of the first aspect are implemented.

[0036] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the rock thin section identification method described in the first aspect above and various possible designs of the first aspect are implemented.

[0037] The rock thin section identification method, device, equipment, and storage medium provided by the present application have at least the following beneficial effects:

[0038] Through the improved network architecture design and innovative training strategy of the present application, the learning ability of the model under small sample conditions is effectively improved, and the recognition accuracy of complex rock subcategories is significantly enhanced. This technological breakthrough provides a new solution for the automated analysis of rock thin section images, thereby not only improving the classification performance of the model under small data sets, but also significantly enhancing the recognition ability of complex rock subcategories. The present application provides technical support for the practical application and industrialization of the artificial intelligence lithology recognition system, and is expected to be widely used in fields such as mineral exploration, oil and gas development, and geological engineering. Description of the Drawings

[0039] The accompanying drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with this application, and are used together with the description to explain the principles of this application.

[0040] Figure 1 It is a flowchart of a method for identifying rock thin sections provided in an embodiment of this application;

[0041] Figure 2 It is an architecture diagram of an identification model provided in an embodiment of this application;

[0042] Figure 3 It is a flowchart for training the identification model provided in an embodiment of this application;

[0043] Figure 4 It is a schematic diagram of evaluation metrics for different models provided in an embodiment of this application;

[0044] Figure 5 It is a comparison chart of the AUC values of the deep learning model provided in an embodiment of this application, exemplifying the mean squared error MSE of four models, which are AlexNet, MSA-AlexNet, VGG16, and MSA-VGG16 respectively. The optimal AUC for each category is 1, indicating completely accurate prediction;

[0045] Figure 6 It is a schematic structural diagram of a rock thin section identification device provided in an embodiment of this application.

[0046] Through the above accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed Embodiments

[0047] Exemplary embodiments will be described in detail here, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0048] In the technical solution of this application, the collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data or user data all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0049] It should be noted that in the embodiments of the present application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0050] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the accompanying drawings.

[0051] The embodiments of the present application provide a method for identifying rock thin sections, aiming to solve the technical problem that the classification accuracy of existing deep learning models significantly decreases when facing insufficient training samples and diverse rock types. The core improvements of this method include the following aspects: 1) Utilize pre-trained models and transfer learning techniques to make full use of the general features learned on large-scale data sets, effectively alleviating the problems of insufficient rock thin section image samples and insufficient generalization performance. 2) Introduce a multi-head self-attention mechanism to improve the traditional convolutional neural network, capture the features of lithologic thin section images from both the global and local aspects, and enhance the fine recognition ability of images; this method not only improves the classification performance of the model on small data sets, but also significantly enhances the recognition ability of complex rock subclasses, providing technical support for the practical application and industrialization of artificial intelligence lithology recognition systems, and is expected to be widely used in fields such as mineral exploration, oil and gas development, and geological engineering.

[0052] Specifically, Figure 1 is a flowchart of a method for identifying rock thin sections provided by the embodiments of the present application. As Figure 1 shown, this method for identifying rock thin sections includes steps S10 to S40.

[0053] S10: Construct a first data set and a second data set; wherein, both the first data set and the second data set include rock thin section images of different categories.

[0054] It should be noted that the first data set and the second data set can be exactly the same data set or different data sets. The description of the first data set and the second data set here in this embodiment aims to distinguish the roles of the two data sets. The first data set is used to train a traditional neural network model to obtain feature weights, and based on these feature weights, the convolutional layer in the recognition model is configured, while the second data set is used to train the recognition model.

[0055] S20: Use the first data set to train the neural network model to obtain feature weights.

[0056] It should be noted that the neural network model is an existing convolutional neural network, such as traditional convolutional neural network models like CNNs models, AlexNet, and VGG16.

[0057] S30: Construct an identification model, where the identification model includes a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer, and a classification layer connected in sequence; among them, the weights of the convolutional layer are the feature weights.

[0058] Figure 2 This is an architecture diagram of an identification model provided by an embodiment of the present application. As Figure 2 shown, the identification model MSA-CNNs includes a convolutional layer 201, a flattening layer 202, a multi-head self-attention layer 203, a fully connected layer 204, a SoftMax layer 205, and a classification layer 206 connected in sequence. The identification model responds to the input thin section image of the rock, extracts image features using the convolutional layer 201, converts the image features into a one-dimensional vector through the flattening layer 202, and the multi-head self-attention layer 203 processes the one-dimensional vector as an input sequence to obtain self-attention features. The fully connected layer 204 and the SoftMax layer 205 are used to integrate and normalize the self-attention features to obtain normalized features, and the classification layer 206 performs classification output based on the normalized features to obtain the category of the thin section of the rock.

[0059] Specifically, Figure 2 in, taking the traditional CNNs model as the neural network model, through training on the ImageNet dataset, a multi-layer convolutional structure with weight sharing is used to extract general features of the image. The identification model is specifically optimized for rock images. By adding a flattening layer 202 and a multi-head self-attention layer 203 on the basis of the pre-trained traditional CNNs model, the ability of the model to process features is enhanced. In this embodiment, the fully connected part of the traditional models (AlexNet and VGG16) is modified, the layers after the fully connected layer are removed, and a convolutional layer 201, a flattening layer 202, a multi-head self-attention layer 203, a fully connected layer 204, a SoftMax layer 205, and a classification layer 206 are added. This structural improvement enables the identification model to show higher efficiency and accuracy when processing specific images with complex textures and structures.

[0060] In some embodiments, considering that rock thin sections usually exhibit complex texture and structural features, the interrelationships between different regions may be extremely important. Therefore, by focusing on different regions or combinations of features in the image, complex patterns in the rock thin section can be captured more effectively, enabling fine differentiation of lithologic subclasses. The multi-head attention mechanism is an advanced technology widely used in fields such as natural language processing and computer vision. It mainly functions by enhancing the model performance. This mechanism was initially introduced through the "Transformer" model. In the multi-head attention mechanism, attention is divided into multiple "heads", each of which independently calculates the weighted representation of the input, and then these representations are combined. This approach allows the model to capture information in different representation subspaces, thereby enhancing the model's expressiveness and complex processing ability. This embodiment explores the application of the multi-head self-attention mechanism (MSA) in the rock thin section image classification task. MSA significantly improves the feature capture ability by parallel processing of multi-regions or feature combinations, and can simulate the dependencies between any two points in the image, which is crucial for identifying complex textures and structures in rock thin sections. Compared with traditional convolutional neural networks, MSA can process the overall context information in a single calculation, rather than just local regions. This self-attention mechanism not only improves the model's feature extraction efficiency but also supports geologists to understand the model decision-making process more deeply and further analyze key features. The multi-head self-attention mechanism is configured in the multi-head self-attention layer 203, and its data processing process is as shown in steps S301 - S304.

[0061] S301: Calculate self-attention.

[0062] For a given input sequence X ∈ R n×d , where R is the real number field, and n and d are the dimensions of the input sequence; first, calculate the query (Q) matrix, the key (K) matrix, and the value (V) matrix, which are different representations of the input data:

[0063] Q = XW Q , K = XW k , V = XW v (1)

[0064] where W Q , W k and W v are learnable weight matrices.

[0065] S302: Calculate the attention weight matrix A, and the calculation formula is:

[0066]

[0067] where, is the scaling factor, used to avoid too large inner product values and prevent gradient disappearance. Softmax is the normalization function, and T is the matrix transpose.

[0068] S303: Calculate the output of self-attention, and the calculation formula is:

[0069] Attention(Q, K, V) = AV (3)

[0070] S304: Since multi-head attention (Multihead Attention) performs the above self-attention calculation in parallel through multiple heads, each head uses a different weight matrix, and calculates the self-attention feature through the following formula:

[0071] MultiHead(Q, K, V) = Concat(head1, head2, … head i , …, head h )W o (4)

[0072] where, MultiHead(Q, K, V) is the self-attention feature, head i = Attention(Q i , K i , V i ), Q i , K i and V i respectively represent the i-th query matrix, key matrix and value matrix. head1, head2, head i and head h respectively represent the self-attention outputs of the 1st, 2nd, i-th and h-th attention heads, and W o is the output weight matrix.

[0073] Finally, the output of multi-head attention will pass through linear transformation and subsequent layer processing, such as a feed-forward neural network, etc., to generate the final feature representation for classification.

[0074] S40: Use the second data set to train the recognition model to obtain a trained recognition model, and the trained recognition model is used to identify the rock thin section category in the rock thin section image.

[0075] In some embodiments, Figure 3This is the flow chart for training the recognition model provided by the embodiments of this application. The second data set was divided into a training set and a test set in a ratio of 7:3. The project was developed and implemented using the deep learning toolbox on MATLAB R2023a. The experiment was conducted on a Lenovo P920 graphics workstation equipped with an Intel Xeon Gold 6226R CPU@2.90GHz and 64GB of memory. In this embodiment, thin rock slice images were used to train the proposed model through the Stochastic Gradient Descent with Momentum (SGDM) algorithm. All existing methods were trained with the same parameter settings. In this embodiment, the Stochastic Gradient Descent with Momentum (SGDM) optimizer was adopted to enhance the training stability of the model. The mini-batch size of the model was set to 64, the maximum number of training epochs was 30, and the initial learning rate was 0.001 to ensure the progressive convergence of the training process and reduce the risk of falling into local extrema. In addition, validation was performed after every 3 mini-batches to closely track the performance of the model and make timely adjustments to avoid overfitting. Meanwhile, 2 multi-head self-attention mechanisms were introduced into the recognition model to improve the sensitivity to different features, thereby enhancing the overall learning ability.

[0076] To test the performance of the thin rock slice recognition method provided by the embodiments of this application, in some embodiments, the results were evaluated through five metrics: accuracy, precision, recall, F1 score, and ROC-AUC. The evaluation metrics are shown in Table 1.

[0077] Table 1 Evaluation Metrics

[0078]

[0079] In Table 1,

[0080] Accuracy: The proportion of correct judgments among all predictions;

[0081] Precision: The proportion of actual correct ones among the predicted positive examples;

[0082] Recall: The proportion of actual positive examples that are correctly identified;

[0083] F1 scores: A comprehensive evaluation metric for precision and recall;

[0084] ROC curves: A curve graph used to display the performance of a classifier, reflecting the relationship between the true positive rate and the false positive rate;

[0085] AUC: The area under the ROC curve, used to evaluate the overall performance of the model;

[0086] TP: Correctly identified positive examples;

[0087] TN: Correctly identified negative examples;

[0088] FP: False positives, which are negative examples misidentified as positive examples.

[0089] TPR: True positive rate, which is also the recall rate.

[0090] FPR: False negative rate.

[0091] Area under the ROC curve: The area under the ROC curve, which reflects the overall ability of the model to distinguish between positive and negative examples.

[0092] The thin-section images of rocks in the dataset used in this embodiment include thin sections of sedimentary rocks, igneous rocks, and metamorphic rocks, a total of 34 subcategories. The thin-section images of rocks are micrographs, and all micrographs are in RGB format with a resolution of 1280×1024 or 4908×3264 pixels.

[0093] Integrate MSA into traditional convolutional neural network models such as AlexNet and VGG16 to form MSA-AlexNet and MSA-VGG16 (i.e., two recognition models constructed according to the method proposed in this application), and comprehensively evaluate their performance. As Figure 4 shown, the experimental results show that the models integrated with MSA are significantly superior to the existing models in terms of classification accuracy. Specifically, the test set accuracy of the MSA-AlexNet model reaches 90.82%, an increase of 4.81% compared to the original AlexNet's 86.01%. Similarly, the test accuracy of the MSA-VGG16 model is 96.43%, an increase of 2.3% compared to the original VGG16's 94.13%. The MSA model is significantly superior to the traditional models in terms of average precision, average recall rate, and average F1-score. Specifically, the precision of MSA_AlexNet is 93.11% (an increase of 1.15%), the recall rate is 91.35% (an increase of 6.24%), and the F1-score is 91.82% (an increase of 4.78%), all higher than those of the traditional AlexNet. The precision of MSA_VGG16 is 96.75% (an increase of 2.14%), the recall rate is 95.90% (an increase of 4.13%), and the F1-score is 96.06% (an increase of 3.82%), also superior to the traditional VGG16. It should be noted that the F1-score, as the harmonic mean of precision and recall rate, can more comprehensively reflect the model performance, especially important when dealing with imbalanced datasets. Among the 34 lithologies with sample imbalance, the F1-score of MSA-AlexNet has the highest increase of 34% (metamorphic rocks), while that of MSA-VGG16 has the highest increase of 48% (granulite), fully demonstrating the superiority of MSA in fine classification.

[0094] As Figure 5As shown, this embodiment evaluates the AUC values of AlexNet, VGG16 and their variants. The results show that the models after introducing the MSA mechanism are generally better than the original versions. Specifically, the ROC-AUC value of MSA-AlexNet is higher than that of the original AlexNet in 62% of the categories; while MSA-VGG16 is better than VGG16 in 97% of the categories, and the ROC-AUC values of most categories are close to 1, indicating that its classification performance is close to the ideal state. To quantify the model performance, we calculated the mean square relative error (MSE) of each model relative to the ideal ROC-AUC value, and the results are: MSA-VGG16 < VGG16 < MSA-AlexNet < AlexNet. This further confirms the significant effect of the MSA mechanism in improving the model performance.

[0095] This embodiment of the present application also provides a thin rock slice recognition device, as Figure 6 shown, the thin rock slice recognition device includes:

[0096] A data acquisition module 601, configured to construct a first data set and a second data set; wherein, both the first data set and the second data set include thin rock slice images of different categories;

[0097] A first training module 602, configured to train a neural network model using the first data set to obtain feature weights;

[0098] A model construction module 603, constructing a recognition model, the recognition model includes a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer and a classification layer connected in sequence; wherein, the weights of the convolutional layer are the feature weights;

[0099] A second training module 604, training the recognition model using the second data set to obtain a trained recognition model, and the trained recognition model is used to recognize the thin rock slice category in the thin rock slice image.

[0100] In some embodiments, in response to an input thin rock slice image, the recognition model extracts image features using the convolutional layer, converts the image features into a one-dimensional vector through the flattening layer, the multi-head self-attention layer processes the one-dimensional vector as an input sequence to obtain self-attention features, the fully connected layer and the SoftMax layer are used to integrate and normalize the self-attention features to obtain normalized features, and the classification layer performs classification output based on the normalized features to obtain the thin rock slice category.

[0101] In some embodiments, the data processing process of the multi-head self-attention layer is:

[0102] Obtain the input sequence X ∈ R n×d; where R is the real number field, and n and d are the dimensions of the input sequence;

[0103] The query matrix Q, the key matrix K, and the value matrix V are calculated through the following formulas:

[0104] Q = XW Q , K = XW k , V = XW v (1)

[0105] where W Q , W k , and W v are all learnable weight matrices;

[0106] The attention weight matrix A is calculated through the following formula:

[0107]

[0108] where is the scaling factor, softmax is the normalization function, and T represents matrix transpose;

[0109] Based on the attention weight matrix A, the self-attention output is calculated through the following formula:

[0110] Attention(Q, K, V) = AV (3)

[0111] Based on the self-attention output, the self-attention feature is calculated through the following formula:

[0112] MultiHead(Q, K, V) = Concat(head1, head2, … head i , …, head h )W o (4)

[0113] where MultiHead(Q, K, V) is the self-attention feature, head i = Attention(Q i , K i , V i ), Q i , K i , and V i respectively represent the i-th query matrix, key matrix, and value matrix, head1, head2, head i , and head h respectively represent the self-attention outputs of the 1st, 2nd, i-th, and h-th attention heads, and W o is the output weight matrix.

[0114] In some embodiments, the thin rock slice categories include sedimentary rocks, igneous rocks, and metamorphic rocks.

[0115] In some embodiments, the second training module is further configured to train the recognition model using the second data set based on a stochastic gradient descent optimizer with momentum to obtain a trained recognition model.

[0116] In some embodiments, the second training module is further configured to perform a validation every 3 mini-batches when training the recognition model using the second data set to closely track the performance of the model.

[0117] Embodiments of the present application provide an electronic device. The electronic device may include: a processor and a memory, where the processor and the memory can communicate; exemplarily, the processor and the memory communicate through a communication bus.

[0118] The processor executes the computer-executable instructions stored in the memory, causing the processor to execute the solutions in the above embodiments. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0119] The communication bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The transceiver is used to implement communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include a random access memory (RAM), and may also include a non-volatile memory.

[0120] The electronic device provided by the embodiments of the present application may be the terminal device in the above embodiments.

[0121] An embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer, the computer is enabled to execute the technical solution of the rock thin section recognition method in the above embodiment.

[0122] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solution of the rock thin section recognition method in the above embodiment can be implemented.

[0123] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0124] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0125] In addition, the functional modules in each embodiment of the present application can be integrated in a processing unit, or each module exists physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.

[0126] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in each embodiment of the present application.

[0127] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor.

[0128] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0129] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus in the attached drawings of this application is not limited to only one bus or one type of bus.

[0130] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0131] An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.

[0132] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying thin rock slices, characterized in that, The method includes: Constructing a first data set and a second data set; wherein, both the first data set and the second data set include rock thin section images of different categories; Training a neural network model using the first data set to obtain feature weights; Constructing an identification model, the identification model includes a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer, and a classification layer connected in sequence; wherein, the weights of the convolutional layer are the feature weights; Training the identification model using the second data set to obtain a trained identification model, and the trained identification model is used to identify the categories of rock thin sections in rock thin section images.

2. The rock thin section identification method according to claim 1, characterized in that In response to an input rock thin section image, the identification model extracts image features using the convolutional layer, converts the image features into a one-dimensional vector through the flattening layer, the multi-head self-attention layer processes the one-dimensional vector as an input sequence to obtain self-attention features, the fully connected layer and the SoftMax layer are used to integrate and normalize the self-attention features to obtain normalized features, and the classification layer performs classification output based on the normalized features to obtain the categories of rock thin sections.

3. The rock thin section identification method according to claim 2, characterized in that, The data processing process of the multi-head self-attention layer is as follows: Obtain the input sequence \(X\in\mathbb{R}\) n×d ; where \(\mathbb{R}\) is the real number field, and \(n\) and \(d\) are the dimensions of the input sequence. Calculating a query matrix Q, a key matrix K, and a value matrix V through the following formula: Q = XW Q , K = XW k , V = XW v (1) Among them, W Q , W k and W v are all learnable weight matrices; Calculating an attention weight matrix A through the following formula: Among them, is the scaling factor, softmax is the normalization function, and T is the matrix transpose; Based on the attention weight matrix A, calculating a self-attention output through the following formula: Attention(Q,K,V)=AV (3) Based on the self-attention output, calculating self-attention features through the following formula: MultiHead(Q,K,V)=Concat(head1,head2,…head i ,…,head h )W o (4) Among them, MultiHead(Q, K, V) is the self-attention feature, and head i = Attention(Q i , K i , V i ), Q i , K i and V i respectively represent the i-th query matrix, key matrix, and value matrix. head1, head2, head i and head h respectively represent the self-attention outputs of the 1st, 2nd, i-th, and h-th attention heads, and W o is the output weight matrix.

4. The rock thin section identification method according to claim 1, wherein The categories of rock thin sections include sedimentary rocks, igneous rocks, and metamorphic rocks.

5. The rock thin section identification method according to claim 1, wherein Based on the stochastic gradient descent optimizer with momentum, training the identification model using the second data set to obtain a trained identification model.

6. The rock thin section identification method according to claim 5, characterized in that, When training the identification model using the second data set, perform a validation every 3 mini-batches to closely track the performance of the model.

7. A rock thin section identification device, characterized in that, The device includes: A data acquisition module configured to construct a first data set and a second data set; wherein, both the first data set and the second data set include rock thin section images of different categories; A first training module configured to train a neural network model using the first data set to obtain feature weights; A model construction module that constructs an identification model, the identification model includes a convolutional layer, a flattening layer, a multi-head self-attention layer, a fully connected layer, a SoftMax layer, and a classification layer connected in sequence; wherein, the weights of the convolutional layer are the feature weights; A second training module that trains the identification model using the second data set to obtain a trained identification model, and the trained identification model is used to identify the categories of rock thin sections in rock thin section images.

8. An electronic device, characterized in that, Including: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the rock thin section identification method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which when executed by a processor are used to implement the rock thin section recognition method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, which when executed by a processor, implements the rock thin section recognition method according to any one of claims 1-6.

Citation Information

Patent Citations

  • TransUNet-based rock slice image granularity identification method, electronic equipment and storage medium

    CN116543256A

  • Lithology identification method and system, computer equipment and storage medium

    CN116935115A

  • Lithology identification method and device and storage medium

    CN118247579A

  • Rock lithology classification and identification method based on multi-scale convolutional neural network

    CN118587702A