Mineral particle classification method based on multi-mode polarized light image

By designing the MMGC-Net model, using the feature fusion technology of multimodal polarized images, the problems of complex steps, low efficiency and large influence of subjective factors in mineral particle classification and identification are solved, and more efficient and accurate mineral particle classification is achieved.

CN120014632APending Publication Date: 2025-05-16SICHUAN UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311530601.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems in the classification and identification of mineral particles with complex steps, low efficiency, inability to qualitative quantitative analysis, and the results are greatly affected by subjective factors. At the same time, it is difficult to accurately identify mineral particles in single-modal image features.

Method used

A mineral particle classification method based on multimodal polarized images is designed, called MMGC-Net. This model extracts the features of single-polarized and orthogonal polarized images through a 2D backbone network with parameter sharing, and performs feature fusion through in-class feature fusion branches and inter-class feature fusion modules to ultimately achieve accurate classification of mineral particles.

Benefits of technology

It improves the accuracy and efficiency of mineral particle recognition, reduces the problems of information loss and feature misalignment, and can more accurately capture the spatio-temporal characteristics of mineral particles, and enhances the qualitative and quantitative analysis capabilities of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a mineral particle classification method based on a multi-mode polarized light image. Comprising the following steps: firstly, extracting features of two types of polarized light images by using a parameter-shared 2D backbone network so as to ensure feature alignment; then, two parallel intra-class feature fusion branches are designed to ensure independent optimization of features; wherein one branch is used for integrating the characteristics of the single polarized light image, and the other branch is used for further extracting the spatio-temporal characteristics from the extracted characteristics of the orthogonal polarized light sequence image. And then, fusing the two types of features through a customized multi-modal inter-class feature fusion module. And finally, predicting the category of the polarized image of the mineral particles by using a classification module. Compared with the prior art, the method provided by the invention has the advantages that the four classification evaluation indexes are obviously improved, the accurate classification of the mineral particle multi-mode polarized light images is realized, and the method has a wide application prospect in geological research and oil exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention designs a mineral particle classification method based on multimodal polarized images and a new network model MMGC-Net, which are related to the problem of determining the lithological characteristics of underground rock formations in geological research and petroleum exploration, and belong to the field of computer vision and intelligent information processing. Background Art

[0002] Classification and identification of mineral particles is an extremely important technical means in practical research such as analysis of sedimentary characteristics of complex lithology oil and gas reservoirs, diagenesis and comprehensive reservoir evaluation. In the traditional manual identification process, geologists usually comprehensively analyze the morphology, cleavage, absorption and interface information of mineral particles under a single polarizing microscope, and combine the extinction, interference, twinning and other characteristics under an orthogonal polarizing microscope to determine the category of mineral particles. However, on the one hand, mineral particle identification has problems such as complex steps, low efficiency and inability to conduct qualitative and quantitative analysis. On the other hand, due to the limitations of the professional knowledge and work experience of the identification workers, the results of mineral particle identification are greatly interfered by subjective factors. With the in-depth development of geological exploration, people have introduced computer image automation analysis technology to improve the efficiency of mineral particle classification and identification.

[0003] In recent years, with the rapid development of pattern recognition and deep learning technology in the field of image, many researchers have tried to classify and identify single-modal mineral grain images from the characteristics of rock texture, absorbance and morphology. However, due to the complexity of mineral composition, different mineral grains may present similar characteristics when mineral grains are imaged due to factors such as environment, light source, measurement equipment and grain size. Therefore, the reliability of mineral grain classification and identification based on single-modal image features is challenged. In order to improve the accuracy of mineral grain identification, Rob et al. used the multimodal features of mineral grains in a single polarized image and an orthogonal polarized image to segment and identify mineral grains, thereby improving the identification accuracy. Under orthogonal polarization conditions, the extinction characteristics of mineral grains are the key elements in the identification process. However, this characteristic presents a changing process, and it is difficult for a single orthogonal polarized image to capture this key feature, so it is necessary to use orthogonal polarization sequence images to fully present it. This means that this method is quite different from the mineral grain identification method used by actual geologists. In practical applications, geologists usually select a single polarization image at one angle (0 degrees) within an extinction period (0 degrees-90 degrees) and orthogonal polarization images at seven different angles (0 degrees, 15 degrees, 30 degrees, 45 degrees, 60 degrees, 75 degrees, 90 degrees) to classify and identify mineral grain images. This strategy can reduce the amount of data while obtaining more comprehensive mineral grain information. In this paper, we added orthogonal polarization images at two angles of 105 degrees and 120 degrees to obtain more comprehensive grain information.

[0004] It is well known that combining the complementary features in single polarized images and multi-angle orthogonal polarized images can improve the accuracy of mineral particle identification. However, when processing polarized images with different numbers of mineral particles, the complexity of feature extraction and fusion will increase significantly. In addition, when existing 2D-DNN-based and 3D-DNN-based general image multimodal classification methods are applied to single polarized images and orthogonal polarized sequence images of mineral particles, they still face problems such as information loss, feature misalignment, and difficulty in extracting spatiotemporal features. To meet this challenge, we propose a mineral particle classification method based on multimodal polarized images. Summary of the invention

[0005] The present invention designs a mineral particle classification method based on multimodal polarized images, called MMGC-Net, which can extract multimodal features from two types of polarized images and effectively integrate them to achieve accurate classification of mineral particles. Specifically, the model first uses a parameter-sharing 2D backbone network to efficiently extract features from two types of polarized images to ensure feature alignment. Then, we designed two parallel intra-class feature fusion branches to ensure independent optimization of features. Among them, one branch is used to integrate the features of a single polarized image, while the other branch further refines the spatiotemporal features from the extracted features of the orthogonal polarized sequence images. Finally, we fuse these two types of features through a customized inter-class feature fusion module to improve the accuracy of classification.

[0006] A mineral particle classification method based on multimodal polarization images comprises the following steps:

[0007] (1) Use the ResNet-34 baseline to extract the features of mineral particles;

[0008] (2) A Plane-Polarized Intra-Modal Feature Fusion Module (PIFM) was constructed;

[0009] (3) A Cross-Polarized Intra-Modal Feature Fusion Module (CIFM) was constructed;

[0010] (4) Constructed an Inter-Modal Feature Fusion Module (IFM);

[0011] (5) A classification prediction module was constructed. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a block diagram of a mineral particle classification method based on multimodal polarization images of the present invention.

[0013] Figure 2 It is a structural diagram of the CIFM of the present invention. DETAILED DESCRIPTION

[0014] The following is combined with Figure 1 and attached Figure 2 The present invention will be further described:

[0015] The network structure and principle of the MMGC-Net model are as follows:

[0016] The MMGC-Net proposed in the present invention is a deep neural network model for mineral particle classification based on multimodal polarized images, which is a classification model designed specifically for mineral particle polarized images. Figure 1 As shown in the figure, MMGC-Net consists of three parts: Backbones, Multimodal Feature Fusion Module and Classification Module.

[0017] (1) Backbones

[0018] The consistency and comparability of features not only help the subsequent spatiotemporal feature extraction and multimodal feature fusion, but also reduce the number of parameters and increase computational efficiency. Therefore, parameter-sharing Backbones is used to extract the features of polarized sequence images. Figure 2 As shown, two parameter-sharing ResNet-34s are used to extract the single polarization images of mineral particles I P ∈R 224×224×3 Features of M P ∈R 1×512×8×8 and orthogonal polarization sequence image I C ={C1,...,C9}∈R 9×224×224×3 Features of M C ∈R 9 ×512×8×8 .

[0019] (2) Multimodal feature fusion module

[0020] like Figure 1As shown in the figure, the multimodal feature fusion module is mainly composed of three modules: single polarization modality feature fusion module (PIFM), orthogonal polarization modality feature fusion module (CIFM) and inter-polarization modality feature fusion module (IFM). PIFM and CIFM can eliminate the noise and redundant information in the single polarization and orthogonal polarization modality features respectively, and extract more compact and abstract features of each. IFM is committed to achieving efficient integration of features between the two modalities.

[0021] PIFM consists of a convolutional layer, a batch normalization layer (BN), and a Relu activation layer. The convolutional layer maps high-dimensional features to more compact and abstract features, which can reduce the parameters that the model needs to learn, avoid overfitting, and enhance generalization ability. The BN layer is used to stabilize the distribution of features. Finally, Rule is used to enhance the nonlinear expression of features. Figure 1 As shown, M P ∈R 1×512×8×8 As the input of PIFM, the output F of PIFM P ∈R 1×128×8×8 , the specific expression is as follows:

[0022] F P =PIFM(M P )=Relu(BN(Conv3(M P )))

[0023] Among them, Conv3(·) represents a 3×3 convolution operation.

[0024] like Figure 2 As shown in the figure, CIFM contains nine FsuLSTMCell units. Each FsuLSTMCell contains a fusion separation unit (Fsu) and a long short-term memory unit (LSTM); the main function of Fsu is to integrate high-dimensional input features and map them into four independent features. Subsequently, these features are used as input and assigned to the three gates and cell states of LSTM to further capture the dependencies of long sequence data; by cascading nine FsuLSTMCells, CIFM can efficiently extract the spatiotemporal features of 9 orthogonal polarization image features; given FsuLSTMCell t-1 The output (h) at time (t-1) t-1 ,c t-1 ), the input x at time t t , then at time t FsuLSTMCell t Output (h t ,c t ) can be defined as:

[0025] (h t ,c t)=FsuLSTMCell t (h t-1 ,c t-1 ,x t )

[0026] Specifically, the following relationship holds:

[0027] h t =σ(z o )·tanh(c t )

[0028] c t =σ(z f )·c t-1 +σ(z i )·tanh(z g )

[0029] Among them, (z f ,z i ,z g ,z o ) is given by the following formula:

[0030] (z f ,z i ,z g ,z o )=Fsu(h t-1 ,x t )

[0031] =Split(Conv1(Concat(Relu(BN(Conv3(x t ))),h t-1 )))

[0032] Among them, z f , z i , z g and z o Respectively represent the input of the forget gate, input gate, unit state and output gate in LSTM, Concat(·) represents the channel stacking operation, Split(·) represents the division of feature vectors according to the number of channels, sigma(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function, and Conv1(·) represents the 1x1 convolution operation;

[0033] Finally, for the input M of CIFM C , output F C It can be expressed as:

[0034] F C =CIFM(M C )=FsuLSTMCell 9 (MC )

[0035] Among them, FsuLSTMCell 9 (·) indicates nine FsuLSTMCells are cascaded.

[0036] IFM: In order to maintain the integrity and diversity of the original features, F P and F C The stacking operation is performed to fuse into a richer feature set; then, the PIFM module is used to fuse to obtain F PC ∈R 1×128×8×8 , the specific expression is as follows:

[0037] F PC =IFM(F P ,F C )=Relu(BN(Conv3(Concat(F P ,F C ))))

[0038] Among them, Concat(·) represents the channel superposition operation.

[0039] (3) Classification module

[0040] As attached Figure 1 As shown, the classification module first uses the pooling function to classify the feature F PC Perform spatial aggregation, then pass the result to the fully connected layer (FC(·)) to generate decision values, and finally convert them into probability outputs P of various categories through the Softmax function. results , the specific expression is as follows:

[0041] P results =Softmax(FC(GAP(F PC )))

[0042] Here, GAP(·) represents the global average pooling operation.

[0043] The MMGC-Net of the present invention is implemented using the PyTorch 1.9 framework on a single RTX3090 GPU. The relevant hyperparameter settings of the experiment are shown in Table 1. We also used data enhancement techniques such as random image flipping, rotation, cropping, and a combination of these techniques to help the method achieve better performance. The implementation codes of the relevant comparison methods can be downloaded from their respective websites and executed using default settings in all tables.

[0044] According to the needs of scientific research, this paper uses the manual cropping technology of ellipse and polygon interaction to crop polarized images of 5 main mineral particles from the polarized sequence images of rocks. After identification by geological experts, these images construct the original data set of polarized images of mineral particles required for this experiment. The five mineral particles involved are: quartz, mica, plagioclase, alkali feldspar and rock fragments. There are 3000 groups of each mineral particle, each group consists of a 0-degree polarized image and 9 orthogonal polarized images at different angles. In the experiment, the data set is divided into training set, validation set and test set in a ratio of 7:1.5:1.5.

[0045] The present invention conducted comparative tests of different algorithms on the mineral particle polarization image dataset, and studied the influence of network structure ablation on the classification performance of MMGC-Net, using Accuracy, Precision, Recall and F1score as evaluation indicators, and the larger the value, the better the classification effect. The relevant experimental results are shown in Table 2, Table 3, Table 4 and Table 5.

[0046] Table 1. MMGC-Net training hyperparameter settings

[0047]

[0048] Table 2 Comparison of MMGC-Net and traditional single-modal classification methods on four evaluation indicators on the mineral particle polarization image dataset

[0049]

[0050] Table 3 Comparison of MMGC-Net and 3D-DNN-based classification models on the mineral particle polarization image dataset on four evaluation indicators

[0051]

[0052]

[0053] Table 4 Comparison of MMGC-Net and 2D-DNN-based classification models on the mineral particle polarization image dataset on four evaluation indicators

[0054]

[0055] Table 5 Ablation study of MMGC-Net on mineral particle polarization image dataset

[0056]

[0057] In Table 2, ResNet-34-P indicates that the ResNet-34 model is used to classify mineral particles using only single polarization images, and 3D ResNet-34-C indicates that the 3D ResNet-34 model is used to classify mineral particles using only orthogonal polarization sequence images. In Table 5, MMGC-Net-PIFM, MMGC-Net-CIFM, and MMGC-Net-IFM respectively represent the models after MMGC-Net removes PIFM, CIFM, and IFM.

[0058] The results in Tables 2, 3 and 4 show that MMGC-Net has better classification effects than single-modal, 3D-DNN-based and 2D-DNN-based classification models; the results in Table 5 show that PIFM, CIFM and IFM modules can improve the classification performance of MMGC-Net.

Claims

1. A mineral particle classification method based on multimodal polarization images, characterized by the following steps: (1) Use the ResNet-34 baseline to extract features of polarized images of mineral particles; (2) A Plane-Polarized Intra-Modal Feature Fusion Module (PIFM) was constructed; (3) A Cross-Polarized Intra-Modal Feature Fusion Module (CIFM) was constructed; (4) Constructed an Inter-Modal Feature Fusion Module (IFM); (5) A classification module was constructed.

2. The method according to claim 1, characterized in that The multimodal features constructed in step (1) are constructed as follows: two parameter-sharing ResNet-34s are used to extract the single polarization images of mineral particles respectively. Features and orthogonal polarization sequence images Features 3. The method according to claim 1, characterized in that In step (2), a feature fusion module within the single polarization modality is constructed. The construction method is as follows: a convolutional layer, a batch normalization layer (BN), and a Rule activation layer are used to construct the network; the convolutional layer maps the high-dimensional feature fusion to a more compact and abstract feature, which can reduce the parameters that the model needs to learn, avoid overfitting, and enhance the generalization ability; the BN layer is used to stabilize the distribution of features; finally, the Rule is used to enhance the nonlinear expression of features; As the input of PIFM, the output of PIFM The specific expression is as follows: F P =PIFM(M P )=Relu(BN(Conv3(M P ))) Among them, Conv3(·) represents a 3×3 convolution operation.

4. The method according to claim 1, characterized in that In step (3), an orthogonal polarization modality feature fusion module is constructed, and the construction method is as follows: nine FsuLSTMCell units are used to construct the CIFM; each FsuLSTMCell contains a fusion separation unit (Fsu) and a long short-term memory unit (LSTM) unit; the main function of Fsu is to integrate high-dimensional input features and map them into four independent features; then, these features are used as input and assigned to the three gates and unit states of LSTM to further capture the dependencies of long sequence data; by cascading nine FsuLSTMCells, CIFM can efficiently extract the spatiotemporal features from the nine orthogonal polarization image features; given an FsuLSTMCell t-1 The output (h) at time (t-1) t-1 ,c t-1 ), the input x at time t t , then at time t FsuLSTMCell t Output (h t ,c t ) can be defined as: (h t ,c t )=FsuLSTMCell t (h t-1 ,c t-1 ,x t ) Specifically, the following relationship holds: h t =σ(z o )·tanh(c t ) c t =σ(z f )·c t-1 +σ(z i )·tanh(z g ) Among them, (z f ,z i ,z g ,z o ) is given by the following formula: (z f ,z i ,z g ,z o )=Fsu(h t-1 ,x t ) =Split(Conv1(Concat(Relu(BN(Conv3(x t ))),h t-1 ))) Among them, z f , z i , z g and z o Respectively represent the input of the forget gate, input gate, unit state and output gate in LSTM, Concat(·) represents the channel stacking operation, Split(·) represents the division of feature vectors according to the number of channels, sigma(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function, and Conv1(·) represents a 1×1 convolution operation; Finally, for the CIFM input M C , output F C It can be expressed as: F C =CIFM(M C )=FsuLSTMCell 9 (M C ) Among them, FsuLSTMCell 9 (·) indicates nine FsuLSTMCells are cascaded.

5. The method according to claim 1, characterized in that Step (4) constructs the inter-modal feature fusion module. The construction method is as follows: In order to maintain the integrity and diversity of the original features, F P and F C Perform stacking operations to merge into a richer feature set; Then, the PIFM module is used for fusion to obtain The specific expression is as follows: F PC =IFM(F P ,F C )=Relu(BN(Conv3(Concat(F P ,F C )))) Among them, Concat(·) represents the channel superposition operation.

6. The method according to claim 1, characterized in that Step (5) constructs the classification prediction. The construction method is as follows: First, use the pooling function to transform the feature F PC Perform spatial aggregation, then pass the result to the fully connected layer (FC(·)) to generate decision values, and finally convert them into probability outputs P of various categories through the Softmax function. results , the specific expression is as follows: P results =Softmax(FC(GAP(F PC ))) Here, GAP(·) represents the global average pooling operation.

Citation Information

Cited By

  • Reservoir oiliness evaluation method and device based on multi-modal data fusion, electronic equipment and storage medium

    CN120822036A

  • Reservoir oiliness evaluation method and device based on multi-modal data fusion, electronic equipment and storage medium

    CN120822036B

  • Intelligent particle identification method for lunar soil particle screening

    CN121147586A