Rock slice detail identification and slice text description generation method

By fine-tuning the CN_CLIP comparison learning model, the rock sheet image is matched with the text description, which solves the efficiency and scalability problems of existing rock identification methods, and realizes efficient rock details recognition and text description generation.

CN120047939APending Publication Date: 2025-05-27NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510107931.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing rock identification methods have problems with artificial under-micro identification methods that consume high time and labor, as well as poor scalability of deep learning visual models and single tasks.

Method used

The fine-tuned pre-trained CN_CLIP comparison learning model is used to match the content in the picture with the rock name, mineral, and polarization microscope type to realize the details recognition of rock sheet images and the generation of text descriptions.

Benefits of technology

The matching of rock sheet images and corresponding text descriptions is achieved, the efficiency and scalability of rock identification are improved, and the lithophago description can be directly generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047939A_ABST
    Figure CN120047939A_ABST
Patent Text Reader

Abstract

The invention provides a rock slice detail identification and slice text description generation method, and relates to the technical field of geological analysis. On the basis of constructing a picture and corresponding text description database under the rock polarizing microscope, matching the image under the rock polarizing microscope with a rock name, a mineral composition and a polarizing microscope type by utilizing a fine-tuned CNCLIP contrast learning network; and a text generation model is combined to directly generate slice description from the picture under the rock polarizing microscope. By adopting the method disclosed by the invention, the organization, classification and identification of the rock picture data under the polarizing microscope can be effectively realized, the rock slice image under the microscope can be directly read by the method constructed by the invention, and the content in the picture is matched with the rock name, the mineral and the polarizing microscope type; and on the basis, a text generation model is combined, and finally, the function of automatically carrying out text description on the image under the rock polarizing microscope is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological analysis, and in particular to a method for identifying rock thin slice details and generating thin slice text descriptions. Background Art

[0002] Rock classification can provide the most basic information for regional geological surveys, building three-dimensional geological models, understanding tectonic evolution, delineating mineral prospecting areas, mineralization mechanisms, etc. In order to identify rock types, especially uncommon rocks, detailed analysis of texture and mineral composition information is usually required, including grain size and shape, chemical composition, mineral hardness, luster, color and cleavage.

[0003] The existing methods of rock identification and classification mainly include traditional manual microscopic identification and deep learning visual model classification. Among them, manual microscopic identification is to identify the type of rock thin sections and name the rocks based on the structure, texture, interference color, cleavage and other characteristics of the minerals themselves under a single polarizing microscope and an orthogonal polarizing microscope; the deep learning visual model classification method is that with the development of deep learning technology in recent years, many researchers have adopted deep learning models such as VGG16, ResNet50, and EfficientNetB0 to classify rock types.

[0004] The two existing rock identification methods have different problems when used for rock classification and identification. First, the traditional manual microscopic identification method requires a lot of manual operation and experience, which results in a lot of time and labor consumption when manually identifying and naming rocks; second, when using deep learning visual models for classification, they all mine the feature relationship between images and labels, ignoring the content of the image. This results in the trained model only having the ability to classify and identify specific lithologies, and poor scalability. In addition, after deep learning visual models are trained, a model can often only adapt to one task, such as: a model that distinguishes rock names cannot identify mineral names. Summary of the invention

[0005] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and provide a method for identifying rock thin section details and generating thin section text descriptions. Based on the fine-tuned pre-trained CN_CLIP contrast learning model, the content in the image is matched with the rock name, mineral, and polarizing microscope type, thereby realizing detail recognition of rock thin section images, matching of rock thin section images with corresponding text descriptions, and ultimately generating lithofacies descriptions directly from rock thin sections.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for identifying rock thin section details and generating thin section text descriptions comprises the following steps:

[0008] Step 1: Construct a rock thin section image description sample database;

[0009] Step 2: Extract rock text description feature vectors based on the enhanced bidirectional encoder representation method, and extract rock thin section image feature vectors based on the visual converter;

[0010] Step 3: Calculate the similarity between the rock thin section image feature vector and the rock text description feature vector, and fine-tune the CN_CLIP model;

[0011] Step 4: Save the fine-tuned CN_CLIP model; based on the fine-tuned CN_CLIP model and the text generation model, finally output the text description corresponding to the rock thin section image.

[0012] Furthermore, the specific steps of step 1 include:

[0013] Step 1.1: Take or collect images of igneous rock, sedimentary rock and metamorphic rock thin sections under single polarization microscope and orthogonal polarization microscope; perform image enhancement and normalization on the images taken or collected under polarization microscope of rock thin sections, and use two-dimensional cubic interpolation to reduce the image to a size of 224×224, as shown in formula (1);

[0014]

[0015] Where (x, y) is the point to be solved; i and j represent the row and column indices of the known grid points; m and n represent the offset indices of the surrounding data; B(m) and B(n) are cubic interpolation basis vectors;

[0016] The rock polarizing microscope images are converted into base64 format and numbered individually, and finally saved as .tsv files. Each line of the tsv file includes an image number and the corresponding base64 image data.

[0017] Step 1.2: Provide a text description of the images of the igneous rock, sedimentary rock and metamorphic rock thin sections taken or collected under a single polarizing microscope and an orthogonal polarizing microscope; the text description shall at least include the comprehensive name of the rock, the mineral composition, the mineral content and the type of polarizing microscope used to take the image; the text description corresponding to the thin sections under the rock polarizing microscope shall also be numbered separately;

[0018] Step 1.3: Manually match the individually numbered rock polarizing microscope thin section images and text descriptions, and save the matching result file in .jsonl format. Each line in the jsonl file includes the image number, the corresponding text number, and the text content.

[0019] Furthermore, the step 2 specifically includes the following steps:

[0020] Step 2.1: Extract feature vectors of all rock text descriptions based on the Transformer in the enhanced bidirectional encoder representation method;

[0021] For a single rock description text overall T R , consisting of multiple words t i Composition, namely T R =[t 1 ,t 2 ,,,,t a ], extract the text feature vector Feature(T R ), as shown in formula (2);

[0022]

[0023] Among them, E(t a ) is the input word t a The embedding vector of ; P(a) is the positional encoding of the rock text description, which is used to preserve the word order information;

[0024] Step 2.2: Extract feature vectors of all rock slice images based on the visual converter;

[0025] For a single rock slice image with a pixel size of 224×224 after step 1, it is divided into N non-repeating small image blocks, as shown in formula (3);

[0026]

[0027] Where P is the side length of the non-repeating image block;

[0028] Then these N image blocks are regarded as sequences and input into Transformer for processing, and finally the feature vector Feature(IR) of the rock slice image is extracted, as shown in formula (4);

[0029]

[0030] Among them, y b represents a P×P image block; Flatten is the image block flattened into a one-dimensional vector, and Q(b) is the position encoding of the rock thin section image, which is used to retain the image sequence information.

[0031] Furthermore, the step 3 specifically includes the following steps:

[0032] Step 3.1: Set the training hyperparameters of fine-tuning the pre-trained CN_CLIP contrastive learning model;

[0033] During the training process, the training batch size is defined as 16 and the learning rate is set to 5×10-7 , using adam as the optimizer and the 1.7.0 version of the dependent library lmdb, the overall model is trained under the Linux system;

[0034] Step 3.2: Map all the rock thin section image feature vectors and rock text description feature vectors extracted in step 2 into the same high-dimensional sample space, calculate the Euclidean distance between the rock thin section image feature vector and the rock text description, and shorten the Euclidean distance between related images and texts in this space, while lengthening the Euclidean distance between unrelated images and texts, as shown in formula (5);

[0035]

[0036] Where M is the number of classes; Y represents a binary label, Y = 0 means that the paired samples are similar, and Y = 1 means that the paired samples are dissimilar; D c is the Euclidean distance between the cth pair of samples in the embedding space; m r is a preset boundary. If the Euclidean distance between two samples is less than m r , then the two samples are considered similar; max(0, m r -D c ) 2 If D c Less than m r , then the loss of this part is 0, otherwise the result is m r -D c .

[0037] Furthermore, in step 3, mixed precision is used during the training process, so that the parameters modified by the model during the training process are saved as single precision, and the parameters not modified during the model training process are still retained as half precision.

[0038] Furthermore, the step 4 specifically includes the following steps:

[0039] Step 4.1: Output the accuracy curve and loss function curve of the rock thin section image to text description and text description to rock thin section image matching training process, observe the convergence status of the three training curves, and save the fine-tuned CN_CLIP model after all three training curves converge;

[0040] Step 4.2: Based on the fine-tuned CN_CLIP model, extract the visual feature vector of the rock thin section image, and organize the visual feature vector generated by the CN_CLIP model through the text generation model, and finally output the text description corresponding to the rock thin section image.

[0041] Furthermore, the step 4.2 specifically includes the following steps:

[0042] Step 4.2.1: Construct a feature vector converter to solve the problem that the visual feature vector input by the CN_CLIP model cannot be directly input into the text generation model; the feature vector converter uses a cross-modal interaction mechanism to convert visual features into text feature descriptions T suitable for the text generation model. Q , as shown in formula (6);

[0043]

[0044] Where W represents the mapping of attention output to the embedding space of the language model; Q is the position of the query vector; K is the value of the Q query position; V represents the visual feature vector; softmax is the attention weight; is the scaling factor used to stabilize the gradient;

[0045] Step 4.2.2: Input the visual feature vector processed by the feature vector converter into the text generation model to finally generate the text description corresponding to the rock slice.

[0046] The beneficial effect of adopting the above technical solution is that the rock thin section detail recognition and thin section text description generation method provided by the present invention, on the basis of constructing a database of rock polarizing microscope images and corresponding text descriptions, uses the fine-tuned CN_CLIP comparative learning network to finally achieve the matching of rock polarizing microscope images with rock names, mineral compositions and polarizing microscope types, and provides a research basis for directly generating thin section descriptions from rock polarizing microscope images. The method of the present invention can effectively realize the organization of rock image data under a polarizing microscope, and perform classification and recognition, reflecting that the method constructed by the present invention can directly read the rock thin section image under a microscope, and match the content in the image with the rock name, mineral, and polarizing microscope type. The present invention is simple to implement, has significant effects, and meets the application requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A flow chart of a method for identifying rock thin section details and generating thin section text descriptions provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0049] like Figure 1 As shown, the method of this embodiment is described as follows.

[0050] A method for identifying rock thin section details and generating thin section text descriptions comprises the following steps:

[0051] Step 1: Construct a rock thin section image description sample (Thin images-Text description) database.

[0052] The detail recognition method of rock thin section images proposed in this embodiment is different from the common deep learning method that mines the relationship between images and labels, but mines the relationship between images and text descriptions. Therefore, the method of constructing the Thin images-Text description database is different from the common deep learning database construction method. The specific steps of constructing the Thin images-Text description database include:

[0053] Step 1.1: Take or collect images of igneous rock, sedimentary rock and metamorphic rock thin sections under single polarization microscope and orthogonal polarization microscope. Perform image enhancement and normalization on the images taken or collected under polarization microscope of rock thin sections, and use two-dimensional cubic interpolation to reduce the image to a size of 224×224, as shown in formula (1);

[0054]

[0055] Among them, (x, y) is the point to be solved; i and j represent the row and column indices of the known grid points; m and n represent the offset indices of the surrounding data; B(m) and B(n) are the cubic interpolation basis vectors.

[0056] The rock polarizing microscope images were converted into base64 format and numbered individually, and finally saved as .tsv files. Each line of the tsv file included an image number and the corresponding base64 image data.

[0057] Step 1.2: Provide a text description of the images of the igneous rock, sedimentary rock and metamorphic rock thin sections taken or collected under a single polarizing microscope and an orthogonal polarizing microscope. The text description required in this embodiment at least includes the comprehensive naming of the rock, the mineral composition, the mineral content and the type of polarizing microscope used when the image was taken. In addition, the text description corresponding to the thin sections under the rock polarizing microscope also needs to be numbered separately.

[0058] Step 1.3: Manually match the individually numbered rock polarizing microscope thin section images and text descriptions, and save the matching result file in .jsonl format. Each line in the jsonl file includes the image number, the corresponding text number, and the text content.

[0059] Step 2: Extract rock text description feature vectors based on the enhanced bidirectional encoder representation method (RoBERTa), and extract rock thin section image feature vectors based on the visual transformer (ViT). Specifically, the following steps are included:

[0060] Step 2.1: Extract feature vectors of all rock text descriptions based on the Transformer in the Enhanced Bidirectional Encoder Representation Method (RoBERTa).

[0061] For a single rock description text overall T R , is considered to consist of multiple words t i Composition, namely T R =[t 1 ,t 2 ,,,,t a ], extract the text feature vector Feature(T R ), as shown in formula (2):

[0062]

[0063] Among them, E(t a ) is the input word t a ; P(a) is the positional encoding of the rock text description, which is used to preserve the word order information.

[0064] Step 2.2: Extract feature vectors of all rock thin section images based on the visual transformer (ViT).

[0065] The single rock slice image with a pixel size of 224×224 after step 1 is divided into N non-repeating small image blocks, as shown in formula (3);

[0066]

[0067] Where P is the side length of the non-repeating image block.

[0068] Then these N image blocks are regarded as sequences and input into Transformer for processing, and finally the feature vector Feature (I R ), as shown in formula (4).

[0069]

[0070] Among them, y b represents a P×P image block; Flatten is the image block flattened into a one-dimensional vector, and Q(b) is the position encoding of the rock thin section image, which is used to retain the image sequence information.

[0071] Step 3: Calculate the similarity between the rock thin section image feature vector and the rock text description feature vector, and fine-tune the CN_CLIP model. The specific steps include:

[0072] Step 3.1: Set the training hyperparameters for fine-tuning the pre-trained CN_CLIP contrastive learning model.

[0073] In the training process constructed in this embodiment, the training batch size is defined as 16 and the learning rate is set to 5×10 -7 , using adam as the optimizer and the 1.7.0 version of the dependent library lmdb, the overall model is trained under the Linux system.

[0074] Step 3.2: Map all the rock thin section image feature vectors and rock text description feature vectors extracted in step 2 into the same high-dimensional sample space, calculate the Euclidean distance between the rock thin section image feature vector and the rock text description, and shorten the Euclidean distance between related images and texts in this space, while increasing the Euclidean distance between unrelated images and texts, as shown in formula (5).

[0075]

[0076] Where M is the number of classes; Y represents a binary label, Y = 0 means that the paired samples are similar, and Y = 1 means that the paired samples are dissimilar; D c is the Euclidean distance between the cth pair of samples in the embedding space; m r is a preset margin. If the Euclidean distance between two samples is less than m r , then the two samples are considered similar; max(0, m r -D c ) 2 If D c Less than m r , then the loss of this part is 0, otherwise the result is m r -D c .

[0077] In order to avoid the fine-tuned CN_CLIP model saved in step 3 from being too large, mixed precision (amp) is used during the training process so that the parameters modified by the model during the training process are saved as single precision (FP32), and the parameters that are not modified during the model training process are still retained as half precision (FP16).

[0078] Step 4: Save the fine-tuned CN_CLIP model; based on the fine-tuned CN_CLIP model and the text generation model, finally output the text description corresponding to the rock thin section image. Specifically, the following steps are included:

[0079] Step 4.1: Output the accuracy curve and loss function curve of the rock thin section image to text description and text description to rock thin section image matching training process, observe the convergence status of the three training curves, and save the fine-tuned CN_CLIP model after all three training curves converge.

[0080] Step 4.2: Based on the fine-tuned CN_CLIP model, extract the visual feature vector of the rock thin section image, organize the visual feature vector generated by the CN_CLIP model through the text generation model, and finally output the text description corresponding to the rock thin section image. Specifically, it includes the following steps:

[0081] Step 4.2.1: Construct a feature vector converter, which is used to solve the problem that the visual feature vector input by the CN_CLIP model cannot be directly input into the text generation model. The core of the feature vector converter is to use the cross-modal interaction mechanism to convert visual features into text feature descriptions suitable for the text generation model. Q , as shown in formula (6).

[0082]

[0083] Where W represents the mapping of attention output to the embedding space of the language model; Q is the position of the query vector; K is the value of the Q query position; V represents the visual feature vector; softmax is the attention weight; is a scaling factor used to stabilize the gradient.

[0084] Step 4.2.2: Input the visual feature vector processed by the feature vector converter into the text generation model to finally generate the text description corresponding to the rock slice.

[0085] Based on the construction of a database of rock polarizing microscope images and corresponding text descriptions, this embodiment uses a fine-tuned CN_CLIP comparative learning network to ultimately achieve matching of rock polarizing microscope images with rock names, mineral compositions, and polarizing microscope types. On this basis, based on a text generation model, it is ultimately possible to directly generate thin-section text descriptions from rock polarizing microscope images. The method of this embodiment is simple to implement, has significant effects, and meets the requirements of the application.

[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for identifying rock thin section details and generating thin section text description, characterized by: The method comprises the following steps: Step 1: Construct a rock thin section image description sample database; Step 2: Extract rock text description feature vectors based on the enhanced bidirectional encoder representation method, and extract rock thin section image feature vectors based on the visual converter; Step 3: Calculate the similarity between the rock thin section image feature vector and the rock text description feature vector, and fine-tune the CN_CLIP model; Step 4: Save the fine-tuned CN_CLIP model; based on the fine-tuned CN_CLIP model and the text generation model, finally output the text description corresponding to the rock thin section image.

2. A method for identifying rock thin section details and generating thin section text description according to claim 1, characterized in that: The specific steps of step 1 include: Step 1.1: Take or collect images of igneous rock, sedimentary rock and metamorphic rock thin sections under single polarization microscope and orthogonal polarization microscope; perform image enhancement and normalization on the images taken or collected under polarization microscope of rock thin sections, and use two-dimensional cubic interpolation to reduce the image to a size of 224×224, as shown in formula (1); Where (x, y) is the point to be solved; i and j represent the row and column indices of the known grid points; m and n represent the offset indices of the surrounding data; B(m) and B(n) are cubic interpolation basis vectors; The rock polarizing microscope images are converted into base64 format and numbered individually, and finally saved as .tsv files. Each line of the tsv file includes an image number and the corresponding base64 image data. Step 1.2: Provide a text description of the images of the igneous rock, sedimentary rock and metamorphic rock thin sections taken or collected under a single polarizing microscope and an orthogonal polarizing microscope; the text description shall at least include the comprehensive name of the rock, the mineral composition, the mineral content and the type of polarizing microscope used to take the image; the text description corresponding to the thin sections under the rock polarizing microscope shall also be numbered separately; Step 1.3: Manually match the individually numbered rock polarizing microscope thin section images and text descriptions, and save the matching result file in .jsonl format. Each line in the jsonl file includes the image number, the corresponding text number, and the text content.

3. A method for identifying rock thin section details and generating thin section text description according to claim 2, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Extract feature vectors of all rock text descriptions based on the Transformer in the enhanced bidirectional encoder representation method; For a single rock description text overall T R , consisting of multiple words t i Composition, namely T R =[t1,t2,,,,t a ], extract the text feature vector Feature(T R ), as shown in formula (2); Among them, E(t a ) is the input word t a The embedding vector of ; P(a) is the positional encoding of the rock text description, which is used to preserve the word order information; Step 2.2: Extract feature vectors of all rock slice images based on the visual converter; For a single rock slice image with a pixel size of 224×224 after step 1, it is divided into N non-repeating small image blocks, as shown in formula (3); Where P is the side length of the non-repeating image block; Then these N image blocks are regarded as sequences and input into Transformer for processing, and finally the feature vector Feature (I R ), as shown in formula (4); Among them, y b represents a P×P image block; Flatten is the image block flattened into a one-dimensional vector, and Q(b) is the position encoding of the rock thin section image, which is used to retain the image sequence information.

4. The method for identifying rock thin section details and generating thin section text description according to claim 3 is characterized by: The step 3 specifically comprises the following steps: Step 3.1: Set the training hyperparameters of fine-tuning the pre-trained CN_CLIP contrastive learning model; During the training process, the training batch size is defined as 16 and the learning rate is set to 5×10 -7 , using adam as the optimizer and the 1.7.0 version of the dependent library lmdb, the overall model is trained under the Linux system; Step 3.2: Map all the rock thin section image feature vectors and rock text description feature vectors extracted in step 2 into the same high-dimensional sample space, calculate the Euclidean distance between the rock thin section image feature vector and the rock text description, and shorten the Euclidean distance between related images and texts in this space, while lengthening the Euclidean distance between unrelated images and texts, as shown in formula (5); Where M is the number of classes; Y represents a binary label, Y = 0 means that the paired samples are similar, and Y = 1 means that the paired samples are dissimilar; D c is the Euclidean distance between the cth pair of samples in the embedding space; m r is a preset boundary. If the Euclidean distance between two samples is less than m r , then the two samples are considered similar; max(0,m r -D c ) 2 If D c Less than m r , then the loss of this part is 0, otherwise the result is m r -D c .

5. A method for identifying rock thin section details and generating thin section text description according to claim 4, characterized in that: In step 3, mixed precision is used during the training process, so that the parameters modified by the model during the training process are saved as single precision, and the parameters not modified during the model training process are still retained as half precision.

6. The method for identifying rock thin section details and generating thin section text description according to claim 4, characterized in that: The step 4 specifically comprises the following steps: Step 4.1: Output the accuracy curve and loss function curve of the rock thin section image to text description and text description to rock thin section image matching training process, observe the convergence status of the three training curves, and save the fine-tuned CN_CLIP model after all three training curves converge; Step 4.2: Based on the fine-tuned CN_CLIP model, extract the visual feature vector of the rock thin section image, and organize the visual feature vector generated by the CN_CLIP model through the text generation model, and finally output the text description corresponding to the rock thin section image.

7. A method for identifying rock thin section details and generating thin section text description according to claim 6, characterized in that: The step 4.2 specifically includes the following steps: Step 4.2.1: Construct a feature vector converter to solve the problem that the visual feature vector input by the CN_CLIP model cannot be directly input into the text generation model; the feature vector converter uses a cross-modal interaction mechanism to convert visual features into text feature descriptions T suitable for the text generation model. Q , as shown in formula (6); Where W represents the mapping of attention output to the embedding space of the language model; Q is the position of the query vector; K is the value of the Q query position; V represents the visual feature vector; softmax is the attention weight; is the scaling factor used to stabilize the gradient; Step 4.2.2: Input the visual feature vector processed by the feature vector converter into the text generation model to finally generate the text description corresponding to the rock slice.

Citation Information

Cited By

  • Intelligent geotechnical investigation method and system based on AI

    CN120852341A