Method, system and product for identifying coal rock type based on combination of machine learning and logging
By integrating the majority voting method of KNN, SVM and DT algorithms and combining logging data to identify coal rock types, the limitations and accuracy problems of traditional methods are solved, efficient and accurate identification of coal rock types is achieved, and coalbed methane exploration and development is supported.
Patent Information
- Application Number
- CN202510320375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology has high cost in the identification of coal rock types, incomplete core acquisition, and lack of universal applicability of traditional methods, and the decision tree algorithm is prone to fall into local optimal solutions, affecting the prediction accuracy.
Using a machine learning-based method, combining well logging data, three classification algorithms: KNN, SVM and DT, the model prediction results are integrated through the majority voting method, and the number of iterations and prediction accuracy thresholds are set to build a comprehensive classification model.
It improves the accuracy and applicability of macroscopic coal rock type identification, reduces the uncertainty of information interpretation, and significantly improves research efficiency.
Smart Images

Figure CN120296618A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of coal and rock geophysical logging, and relates to a method, system and product for identifying coal and rock types, in particular to a method, system and product for identifying coal and rock types based on machine learning combined with logging. Background Art
[0002] The exploration and development of coalbed methane are restricted by multiple geological and engineering factors, among which the macroscopic coal and rock type is one of the decisive control factors. Research shows that coal samples of different macroscopic coal and rock types show significant differences in adsorption-desorption characteristics, and this difference directly restricts the occurrence state and migration mechanism of coalbed methane. Accurately predicting the macroscopic coal and rock type not only has important guiding significance for optimizing the exploration and development strategy of coalbed methane, but also has profound theoretical value and practical significance for the accurate prediction and prevention of gas disasters in coal mines.
[0003] At present, the identification of the macroscopic coal and rock type of coal mainly relies on the core description method. However, this method has limitations such as high cost and incomplete core acquisition. With the rapid development of geophysical logging technology, its high efficiency and applicability have gradually made it an indispensable technical means in the exploration and development of coalbed methane. Scholars have carried out prediction research on the macroscopic coal and rock type by using methods such as multiple regression analysis, principal component analysis and Haar wavelet transform based on logging curve data. However, these traditional methods are often limited by regional geological characteristics and lack universal applicability. In recent years, machine learning algorithms have shown significant advantages in the field of identifying macroscopic coal and rock types. Classification algorithms such as K-nearest neighbor (KNN) and decision tree (DT) are superior to traditional methods in terms of identification efficiency and accuracy. However, the existing methods still have obvious limitations. Taking the decision tree algorithm as an example, its feature selection mechanism based on local optimality easily leads the model to fall into a local optimal solution, thus affecting the prediction accuracy. Therefore, developing an efficient and accurate method for identifying the macroscopic coal and rock type has important theoretical value and practical significance for improving the exploration and development efficiency of coalbed methane. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a method, system and product for identifying coal and rock types based on machine learning combined with logging.
[0005] The technical solution adopted by the method of the present invention is: a method for identifying coal and rock types based on machine learning combined with logging, comprising the following steps:
[0006] Step 1: Obtain the core to be identified and logging curve data;
[0007] Step 2: Perform outlier rejection processing on the core and logging curve data, and further perform normalization processing;
[0008] Step 3: Input the processed data into the trained classification model to identify the coal and rock types;
[0009] The trained classification models, namely the KNN model, the SVM model and the DT model, integrate the prediction results of the three classification models through the majority voting method, and set the number of iterations and the prediction accuracy threshold to obtain a comprehensive classification model.
[0010] Preferably, in step 1, the core and logging curve data include macro coal rock type, natural gamma (GR, API), resistivity (RT, Ω·m), acoustic time difference (AC, μs / m) and density (DEN, g / cm 3 ).
[0011] Preferably, in step 2, the data is processed to remove outliers using the 3σ criterion;
[0012]
[0013] In the formula, σ is the standard deviation of the data, reflecting the degree of dispersion of the data set; n is the data sample x i The number of is the average value of the sample characteristics; if a data point falls outside the range of [μ-3σ, μ+3σ], the data point is considered to be an outlier, where μ is the mean deviation of the data.
[0014] Preferably, in step 3, the trained classification model is a KNN classification model; first, the processed data is divided into a training set and a test set, the training set is used for model training and development, and the test set is used for model testing; during training, the distance d(x, y) between the target sample and all samples in the training set is calculated to find the K samples with the closest distance, and the category of the target sample is predicted based on the label information of these neighbors;
[0015]
[0016] In the formula, x i is the sample feature vector; d(x,y) is the sample distance, n is the number of samples, x i ,y i are the corresponding feature vectors in different dimensions respectively.
[0017] Preferably, in step 3, the trained classification model is an SVM classification model; first, the processed data is divided into a training set and a test set, the training set is used for model training and development, and the test set is used for model testing; during training, samples of different categories are separated by finding an optimal hyperplane, and the classification interval is maximized; the calculation formula is as follows:
[0018]
[0019] In the formula, x is the feature vector of the sample, and x i is the feature of the i-th sample; y i is the class label, and y i ∈{-1, 1}), when it is equal to 1, it represents a positive output, and when it is equal to -1, it represents a negative output; sign(x) is the sign function, which is used to determine the sample class; K(x, x) is the kernel function; α i is the Lagrange multiplier, which is used to represent the importance of the sample; b is the bias term; n is the number of samples.
[0020] Preferably, in step 3, the trained classification model is a DT classification model; first, the processed data is divided into a training set and a test set. The training set is used for model training and development, and the test set is used for model testing; during training, starting from the root node, the optimal feature is selected for division; according to the value of the feature, the samples are assigned to the child nodes; the above process is repeated for each child node until the stopping condition is met;
[0021] Among them, the decision tree classifies based on the information gain criterion, and the information gain is based on the information entropy Entropy, which is used to measure the improvement of sample purity before and after division; the calculation formula of the information entropy is as follows:
[0022]
[0023] In the formula, S is the sample set; c is the number of classes; P i is the proportion of the i-th class of samples; the information gain is defined as:
[0024]
[0025] In the formula, A is the feature; Values(A) is the value set of the feature A; S v is the sample subset where the value of the feature A is v.
[0026] Preferably, in step 3, the calculation formula of the majority voting method is:
[0027]
[0028] In the formula, M is the number of classifiers; C is the set of classes c, C = {c1, c2... c k}, k is the number of classes; y m is the prediction result of the m-th classifier, and Ⅱ(y m = c) is the indicator function, which is 1 when y m = c, and 0 otherwise.
[0029] The technical solution adopted by the system of the present invention is: a system for identifying coal and rock types based on machine learning combined with well logging, including:
[0030] One or more processors;
[0031] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method for identifying coal and rock types based on machine learning combined with well logging.
[0032] The technical solution adopted by the product of the present invention is: a product for identifying coal and rock types based on machine learning combined with well logging, which when the computer program instructions run on a computer cause the computer to execute the method for identifying coal and rock types based on machine learning combined with well logging.
[0033] Compared with the prior art, the beneficial effects of the present invention include:
[0034] Based on core data and well logging curve data, the present invention proposes a multi-model integration strategy by setting the number of iterations and the prediction accuracy threshold. This strategy integrates three machine learning algorithms, namely KNN, SVM, and DT, and uses the voting method to integrate the prediction results of the models to achieve high-precision identification of macroscopic coal and rock types. Compared with traditional methods, this method has significant advantages: First, by integrating multiple high-performance algorithms, it effectively avoids the feature preference problem that may exist in a single model; Second, it minimizes the uncertainty of information interpretation and improves the reliability of the prediction results; Third, this method has strong applicability and can significantly improve the research efficiency. This integration strategy shows excellent performance in the task of identifying macroscopic coal and rock types, providing a new technical means for the research in the field of coalbed methane exploration and development. Brief Description of the Drawings
[0035] The following uses examples and specific implementation manners to further illustrate the technical solution of the present invention. In addition, some drawings are also used in the process of describing the technical solution. For those skilled in the art, without creative efforts, other drawings and the intention of the present invention can also be obtained based on these drawings.
[0036] Figure 1 It is the flowchart of the method of the embodiment of the present application;
[0037] Figure 2 It is the confusion matrix diagram of the classification results of different algorithms in the embodiment of the present application; Labels 1, 2, 3, and 4 are dull coal, semi-dull coal, semi-bright coal, and bright coal respectively, and the color card represents the number of samples;
[0038] Figure 3 It is the comprehensive bar chart of the comparative analysis of the prediction of macroscopic coal and rock types by different algorithms in the embodiment of the present application. Specific Embodiment
[0039] In order to facilitate ordinary technicians in the field to understand and implement the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0040] Please see Figure 1 This embodiment provides a method for identifying coal and rock types based on machine learning combined with well logging, including the following steps:
[0041] Step 1: Obtain the core and logging curve data to be identified;
[0042] In one embodiment, the core and logging data include macro coal type, natural gamma (GR, API), resistivity (RT, Ω·m), acoustic time difference (AC, μs / m) and density (DEN, g / cm 3 ).
[0043] Step 2: performing outlier elimination processing on the core and logging curve data, and further performing normalization processing;
[0044] In one implementation, the data is normalized after processing to satisfy the normal distribution as much as possible. The calculation formula of σ is as follows:
[0045]
[0046] Where σ is the standard deviation of the data, reflecting the degree of dispersion of the data set. n is the number of samples, is the sample feature average.
[0047] The 3σ criterion is based on the characteristics of normal distribution (Gaussian distribution). For data that obeys normal distribution, its values are mainly concentrated near the mean, and the probability of data points that deviate from the mean is lower. Specifically:
[0048] The probability that the data falls within the range [μ-σ, μ+σ] is about 68.27%;
[0049] The probability that the data falls within the range [μ-2σ, μ+2σ] is about 95.45%;
[0050] The probability that the data falls within the range [μ-3σ, μ+3σ] is about 99.73%.
[0051] Where μ is the mean deviation of the data. According to the 3σ criterion, if a data point falls outside the range of [μ-3σ, μ+3σ], the data point is considered to be an outlier.
[0052] Step 3: Input the processed data into the trained classification model to identify the coal and rock types;
[0053] The trained classification models, namely the KNN model, SVM model, and DT model, integrate the prediction results of the three classification models through the majority voting method, and set the number of iterations and the prediction accuracy threshold to obtain a comprehensive classification model.
[0054] In one implementation, during training, first, data preparation: clear the working environment variables, load the preprocessed dataset, and ensure data integrity and consistency.
[0055] Then, perform data partitioning and label encoding: perform label encoding on the macroscopic coal petrographic types, and divide the dataset into a training set and a test set. Using the random sampling method, 70% of the samples are used as the training set, and the remaining 30% are used as the test set to ensure the independence of model training and evaluation.
[0056] In one implementation, during the training of the KNN classification model, by calculating the distance d(x, y) between the target sample and all samples in the training set, find the K samples with the closest distance, and predict the class of the target sample according to the label information of these neighbors;
[0057]
[0058] where x i is the sample feature vector; d(x, y) is the sample distance, n is the number of samples, x i , y i are the corresponding feature vectors in different dimensions respectively.
[0059] In one implementation, during the training of the SVM classification model, by finding an optimal hyperplane to separate samples of different classes and maximize the classification margin; the calculation formula is as follows:
[0060]
[0061] where x is the feature vector of the sample, x i is the feature of the i-th sample; y i is the class label, y i ∈{-1, 1}), equal to 1 indicates a positive output, equal to -1 indicates a negative output; sign(x) is the sign function used to determine the sample class; K(x, x) is the kernel function; α i is the Lagrange multiplier used to represent the sample importance; b is the bias term; n is the number of samples.
[0062] In one implementation, during the training of the DT classification model, starting from the root node, select the optimal feature for partitioning; allocate samples to child nodes according to the values of the features; repeat the above process for each child node until the stopping condition is met;
[0063] Among them, the decision tree classifies based on the information gain criterion, and the information gain is based on the information entropy Entropy, which is used to measure the improvement of the sample purity before and after division; the calculation formula of the information entropy is as follows:
[0064]
[0065] In the formula, S is the sample set; c is the number of categories; P i is the proportion of the i-th type of samples; the information gain is defined as:
[0066]
[0067] In the formula, A is the feature; Values(A) is the value set of the feature A; S v is the sample subset where the value of the feature A is v.
[0068] In one implementation, the majority voting method is used to integrate the prediction results of the three models of KNN, SVM, and DT. The maximum number of iterations is set to 100 times, and the prediction accuracy threshold is set to 80%. The model performance is evaluated by the overall classification accuracy. If the preset accuracy threshold is not reached, it will return to re-partition the training set and iterate the training until the accuracy requirement is met.
[0069] Among them, the calculation formula of the majority voting method is:
[0070]
[0071] In the formula, M is the number of classifiers; C is the set of categories c, C = {c1, c2... c k}, k is the number of categories; y m is the prediction result of the m-th classifier, Ⅱ(y m = c) is the indicator function, which is 1 when y m = c, otherwise it is 0.
[0072] Based on the final voting result, a confusion matrix is established to visualize the model result.
[0073] During the model training process, when the number of iterations reaches 23 times, the performance of each classification model tends to be stable. Specifically, the accuracies of the KNN, SVM, and DT models are 79.73%, 77.03%, and 77.03% respectively, while the accuracy of the integrated voting method reaches 83.78%, meeting the preset accuracy requirements. As Figure 2As shown in the figure, the comparison and analysis of the prediction performance of different algorithms through the confusion matrix shows that the voting method has advantages in the identification of various types of coal and rock: the recognition accuracy of dull coal (label 1), semi-dark coal (label 2), semi-bright coal (label 3) and bright coal (label 4) are 82%, 77%, 79% and 95% respectively. Among them, the voting method has achieved the best performance in the identification of dull coal and bright coal. Although the accuracy of the voting method in the identification of semi-dark coal (77%) is slightly lower than that of the DT model (85%), it is significantly better than the KNN (69%) and SVM (54%) models; at the same time, its semi-bright coal recognition accuracy (79%) is comparable to that of the DT model and better than the KNN model (72%). Although the voting method is not optimal in the recognition accuracy of some coal and rock types, it effectively reduces the possible feature deviations of a single model through ensemble learning, thereby significantly improving the generalization ability of the model. In addition, as Figure 3 As shown in the figure, the comprehensive bar chart based on the prediction results of different algorithms further verifies the superiority of the voting method in the identification of macro coal and rock types, and intuitively demonstrates the differences in the prediction accuracy of each algorithm.
[0074] The present invention proposes an intelligent identification method for coal and rock types based on the combination of machine learning and logging data. The method collects core data and logging data through the system, including macro coal and rock types, natural gamma (GR, API), resistivity (RT, Ω·m), acoustic time difference (AC, μs / m), and density (DEN, g / cm 3 ). First, the 3-Sigma (3σ) criterion is used to remove outliers from the original data to ensure data quality. Subsequently, based on the preprocessed logging data, three classification algorithms, namely K nearest neighbor (KNN), support vector machine (SVM) and decision tree (DT), are used to construct prediction models, and the prediction results of each model are fused through the voting mechanism of ensemble learning to obtain the final recognition result. In addition, the present invention achieves high-precision prediction of macroscopic coal and rock types by setting the number of iterations and prediction accuracy thresholds, providing reliable technical support for coalbed methane exploration and development.
[0075] It should be understood that the embodiments described above are part of the embodiments of the present invention, rather than all of the embodiments. In addition, the technical features in the various embodiments or single embodiments provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the sequence of steps and / or the structural composition mode, but must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0076] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection claimed by the present invention shall be subject to the appended claims.
Claims
1. A method for identifying coal and rock types based on machine learning combined with well logging, characterized in that, It includes the following steps: Step 1: Obtain the core to be identified and well logging curve data; Step 2: Remove outliers from the core and well logging curve data, and further perform normalization processing; Step 3: Input the processed data into the trained classification model for coal and rock type identification; The trained classification model, namely the KNN model, SVM model and DT model, integrates the prediction results of the three classification models by the majority voting method, and sets the number of iterations and the prediction accuracy threshold to obtain a comprehensive classification model.
2. The method for identifying coal and rock types based on machine learning combined with well logging according to claim 1, wherein: In Step 1, the core and well logging curve data include maceral type, gamma ray (GR, API), resistivity (RT, Ω·m), acoustic transit time (AC, μs / m), and density (DEN, g / cm 3 ).
3. The method for identifying coal and rock types based on machine learning combined with well logging according to claim 1, characterized in that: In Step 2, the 3σ criterion is used to remove outliers from the data; where σ is the standard deviation of the data, reflecting the degree of dispersion of the data set; n is the number of data samples x i , is the average value of the sample features; if the data point falls outside the range of [μ - 3σ, μ + 3σ], the data point is considered an outlier, where μ is the mean deviation of the data.
4. The method for identifying coal and rock types based on machine learning combined with well logging according to claim 1, characterized in that: In Step 3, the trained classification model is the KNN classification model; First, divide the processed data into a training set and a test set. The training set is used for model training and development, and the test set is used for model testing; During training, by calculating the distance d(x,y) between the target sample and all samples in the training set, find the K nearest samples, and predict the category of the target sample according to the label information of these neighbors; where x i is the sample feature vector; d(x, y) is the sample distance, n is the number of samples, x i , y i are the corresponding feature vectors in different dimensions respectively.
5. The method for identifying coal and rock types based on machine learning combined with well logging according to claim 1, characterized in that: In Step 3, the trained classification model is the SVM classification model; First, divide the processed data into a training set and a test set. The training set is used for model training and development, and the test set is used for model testing; During training, find an optimal hyperplane to separate samples of different categories and maximize the classification margin; the calculation formula is as follows: where \(x\) is the feature vector of the sample, \(x\) i is the feature of the \(i\)-th sample; \(y\) i is the class label, \(y\) i \(\in\{-1, 1\}\), and when it is equal to 1, it represents a positive output, and when it is equal to -1, it represents a negative output; \(sign(x)\) is the sign function used to determine the sample class; \(K(x, x)\) is the kernel function; \(\alpha\) i is the Lagrange multiplier used to represent the importance of the sample; \(b\) is the bias term; \(n\) is the number of samples.
6. The method for identifying coal and rock types based on machine learning combined with well logging according to claim 1, wherein: In Step 3, the trained classification model is the DT classification model; First, divide the processed data into a training set and a test set. The training set is used for model training and development, and the test set is used for model testing; During training, start from the root node, select the optimal feature for division; allocate samples to child nodes according to the values of the features; repeat the above process for each child node until the stopping condition is met; Among them, the decision tree classifies based on the information gain criterion, and the information gain is based on the information entropy Entropy, which is used to measure the improvement of sample purity before and after division; the information entropy calculation formula is as follows: where S is the sample set, c is the number of categories, and P i is the proportion of the samples in the i-th category. The information gain is defined as: In the formula, A is the feature; Value s(A) is the set of values of feature A; S v is the subset of samples where the value of feature A is v.
7. The method for identifying coal and rock types based on machine learning combined with well logging according to any one of claims 1-6, characterized in that: In Step 3, the calculation formula of the majority voting method is: Where M is the number of classifiers; C is the set of classes c, C = {c1, c2... c k}, k is the number of classes; y m is the prediction result of the m-th classifier, Ⅱ(y m = c) is an indicator function, which is 1 when y m = c, and 0 otherwise.
8. A system for identifying coal and rock types based on machine learning combined with well logging, characterized in that, It includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method for identifying coal and rock types based on machine learning combined with well logging as described in any one of claims 1 to 7.
9. A product for identifying coal and rock types based on machine learning combined with well logging, characterized in that: When the computer program instructions run on a computer, cause the computer to execute the method for identifying coal and rock types based on machine learning combined with well logging as described in any one of claims 1 to 7.