Coal body structure identification method based on unbalanced logging data
The uneven coal structure logging data is processed through k-means clustering and ADASYN resampling technology, and the recognition model is constructed using the random forest algorithm, which solves the model learning bias caused by data imbalance and improves the accuracy and stability of coal structure recognition.
Patent Information
- Application Number
- CN202510226391.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
When identifying coal structures, the logging data is uneven, resulting in model learning bias, which affects classification performance and accuracy.
The uneven logging data was processed by k-means clustering and ADASYN resampling technology. The optimal cluster number was determined through k-means clustering analysis, and synthetic samples were generated using ADASYN to make the data volume of each coal structure evenly distributed. Subsequently, a coal structure recognition model was constructed based on the random forest algorithm.
It effectively eliminates the model learning bias caused by uneven data distribution, improves the accuracy of coal structure recognition and model stability, and provides a scientific basis for coalbed methane development.
Smart Images

Figure CN120067735A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of coalbed methane development, and in particular relates to a coal body structure identification method based on non-equilibrium logging data. Background Art
[0002] my country's coal-bearing basins have undergone multiple tectonic movements. Low-strength coal has undergone varying degrees of deformation and destruction under the action of geological stress, forming a variety of coal structures. These structures directly affect the development of coal seam fractures, porosity, permeability, gas content and mechanical properties, and are of great significance to coalbed methane development and coal resource utilization. Geophysical logging has become one of the common methods for identifying coal structure due to its low cost, easy operation and vertical continuity. However, there is a large overlap in the values of logging curves of different coal structures, resulting in a nonlinear and complex relationship between the two. As an important classification and regression algorithm, random forest is suitable for processing high-dimensional data, with the advantages of fast training speed, low overfitting, strong generalization ability and high prediction accuracy. Therefore, random forest can be used to construct a classification model of coal structure based on logging data. It should be noted that the amount of data of different coal structures in coal seams is often unbalanced, which will cause the model to be biased towards the majority class during learning, thereby affecting the classification performance of the minority class and reducing the accuracy and stability of the model. Therefore, we try to use k-means clustering and Adaptive Synthetic Sampling (ADASYN) to resample the imbalanced dataset before modeling to eliminate the model learning bias caused by the imbalanced distribution of data. Summary of the invention
[0003] In order to solve the related problems existing in the prior art and to more quickly and accurately identify the coal structure according to geophysical logging data, the present invention provides a coal structure identification method based on non-equilibrium logging data.
[0004] In order to solve the above technical problems, the present invention adopts the following technical solution: a method for identifying coal structure based on non-equilibrium logging data, comprising the following steps:
[0005] S1. Obtain logging data of different coal structures in multiple wells;
[0006] S2, outlier removal and standardization of logging data;
[0007] S3, k-means cluster analysis;
[0008] S4, ADASYN hybrid resampling;
[0009] S5. Evaluation of data equalization processing effect based on KS test and information entropy;
[0010] S6. Construction of a coal structure recognition model based on the random forest algorithm.
[0011] Further, the specific process of step S1 is as follows: According to the coal structure classification standard of GB / T 30050-2013, the coal samples obtained by drilling and coring are divided into four categories: primary structure coal, fragmented coal, crushed coal, and mylonitized coal; at the same time, the corresponding logging data are collected. The extraction interval of the logging data is 0.125 m, and the logging types include deep lateral resistivity (LLD), shallow lateral resistivity (LLS), spontaneous potential (SP), natural gamma (GR), density (DEN), neutron (CNL), acoustic wave (AC), and dual caliper logging (CALX, CALY); according to the burial depth (DEPTH), the numerical values of various logging curves are saved in the form of data columns.
[0012] Further, the specific process of step S2 is as follows:
[0013] To ensure the quality and consistency of the logging data of different coal structures, abnormal data need to be removed, including the parting in the coal seam and the logging data within 1 m from the top and bottom coal seams; this is because the parting is usually composed of claystone, carbonaceous mudstone, siltstone, etc., and its logging value is significantly different from that of coal; while the logging data of the coal seam too close to the top and bottom coal seams will be affected by the adjacent rock strata, and the amplitude changes greatly, resulting in abnormal distribution of the coal structure logging data; in addition, different types of logging data have different dimensions, and the Z-score method needs to be used to standardize the original data of various logging curves. The formula is as follows:
[0014]
[0015] In formulas (1)-(3), M refers to the type of logging curve, and m refers to the m-th type of logging curve; n represents the number of samples of each logging curve, and i refers to the logging value of the i-th sample; Z mi refers to the normalized value of the i-th sample of the m-th type of logging curve; x mi refers to the original logging value of the i-th sample of the m-th type of logging curve; μ m and σ m respectively refer to the mean and standard deviation of the m-th type of logging curve.
[0016] Further, to avoid the phenomenon that the new samples generated in the adaptive synthetic sampling (ADASYN) process are overly concentrated in certain categories, thus changing the distribution characteristics of the original data set and increasing the uncertainty of the data, the original data of each coal structure are subjected to k-means clustering analysis, the optimal number of clusters is selected by using the silhouette coefficient, and then ADASYN is used to independently resample each cluster; the k-means clustering analysis in step 3 specifically includes the following steps:
[0017] ①Divide the coal body structure data into p groups according to the coal body structure categories
[0018] D = {X j} j = 1, 2, ..., p (4)
[0019] X j = {x ij} i = 1, 2, ..., n p (5)
[0020] Wherein, D refers to the coal body structure data set, X j refers to the j-th coal body structure data set, x ij refers to the i-th sample in the j-th coal body structure, n p refers to the number of samples of the p-th coal body structure;
[0021] ②For each type of coal body structure data, randomly select k initial clustering centers;
[0022]
[0023] Wherein, refers to the set of initial clustering centers in the j-th coal body structure, refers to the k-th initial clustering center in the j-th coal body structure;
[0024] ③For any data point x ij , calculate the Euclidean distance from each sample to the k clustering centers, and assign the sample to the nearest clustering center according to the principle of non-intersection;
[0025]
[0026] Wherein, refers to the sample of the u-th clustering category in the j-th coal body structure, and respectively refer to the distances from any data point x ij in the j-th coal body structure to the u-th and v-th clustering centers;
[0027] ④For the classification results calculated in formula (3), recalculate the clustering centers using the category means;
[0028]
[0029] Wherein, n uj refers to the number of samples in the u-th clustering category in the j-th coal body structure, refers to the recalculated clustering center in the j-th coal body structure, refers to the k-th clustering center in the recalculated clustering centers in the j-th coal body structure;
[0030] ⑤ Repeat steps ③ and ④ to continuously update the cluster centers until stop the iteration, and take the result updated at the t-th time as the clustering result of the coal seam structure data;
[0031] ⑥ Select the silhouette coefficient characterizing the clustering effect to optimize the optimal number of clusters; change the number of clusters k, calculate the silhouette coefficient of the clustering result for each number of clusters k, and select the number of clusters with the largest silhouette coefficient as the optimal number of clusters, providing the basic data for ADASYN resampling;
[0032]
[0033] In the formula, S(k) refers to the silhouette coefficient of the coal seam structure data set under the k-th number of clusters, S(x ij ) refers to the silhouette coefficient of the sample x ij , n k refers to the maximum value of the number of cluster centers, n refers to the number of samples of all coal seam structure data, a(ij) refers to the average distance from the sample x ij to other samples within the category, and b(ij) refers to the average distance from the sample x ij to samples of other categories.
[0034] Furthermore, ADASYN in step S4 is a resampling method for imbalanced data sets, which has an adaptive characteristic for the generation of minority class samples. By generating minority class samples, the data volume of each coal seam structure is evenly distributed to improve the performance of the classification model. Step S4 specifically includes the following steps:
[0035] (1) According to the data volume n j (j = 1, 2,..., p) of each coal seam structure data, resample all coal seam structure data with a smaller data volume according to the maximization principle, so that the sample data volume of each coal seam structure data after resampling is the same as the sample data volume of the coal seam structure with the largest data volume;
[0036]
[0037] In the formula, G refers to the data volume that each coal seam structure data requiring resampling needs to reach after resampling, and n j (max) refers to the number of samples of the coal seam structure with the largest data volume;
[0038] (2) For the k-means clustering results of each coal seam structure data, count the number of samples in each category, and record the minority samples as m s , and obtain the synthetic data volume g;
[0039]
[0040] In the formula, β refers to the parameter of the synthetic data, β ∈ [0, 1], and β = 1 indicates that the balanced data set is generated;
[0041] (3) For each category of coal seam structure data samples clustered in (2), use the Euclidean distance to identify the s nearest neighbor samples of the sample, and calculate its ratio r ui , and perform normalization processing on it;
[0042]
[0043] In the formula, Δui refers to the number of the majority samples among the s nearest neighbor samples of the sample x ui , n s refers to the total number of samples of the minority category, refers to r ui normalized result;
[0044] (4) Based on the calculated synthetic sample quantity and the normalized result, calculate the synthetic sample quantity g of each minority category coal seam structure sample data ui , and on this basis, generate the synthetic data s of each minority category sample x ui ; ui ;
[0045]
[0046] s ui = x ui +(x zi -x ui )λ(17)
[0047] In the formula, x ui refers to the i-th sample of the u-th minority category, x zi refers to the minority sample randomly selected from the s nearest neighbor samples of x ui , and λ refers to a random number, λ ∈ [0, 1];
[0048] For each type of coal seam structure data, repeat steps (2)-(4) to generate synthetic samples that meet the data volume requirements, so that the data volume of each type of coal seam structure after resampling is n j (max).
[0049] Furthermore, the specific process of step S5 is as follows:
[0050] The KS (Kolmogorov-Smirnov) test and information entropy can test whether the equalization processing has changed the original distribution characteristics of the sample data and examine whether it has improved the information volume of the sample data; the specific calculation methods are as follows:
[0051] Ⅰ. KS Test
[0052] According to the maximum difference between the empirical distribution functions of the original coal body structure data and the resampled coal body structure data, analyze whether the distribution of the data after equalization processing has changed. Generally, it is considered that if the P - value of the KS test statistic \(d_{max}\) is greater than the significance level of 0.05, it can be considered that the equalization processing has not changed the distribution of the original data;
[0053] \(d_{max}=\max|F′(x)-F(x)| (18)\)
[0054] In the formula, \(F(x)\) and \(F'(x)\) respectively refer to the empirical distribution functions of the original data and the sample data after equalization processing, and \(d_{max}\) refers to the statistic of the KS test;
[0055] Ⅱ. Information Entropy
[0056] Construct the probability density functions of the coal body structure sample data before and after equalization processing, and then calculate the information entropy of the two respectively; since the coal body structure data is discrete and the distribution is unknown, refer to relevant research and select the kernel density estimation method to estimate the probability density function of the coal body structure sample;
[0057]
[0058] In the formula, refers to the estimated probability density function, \(H(X)\) refers to the information entropy, \(N\) refers to the total number of samples, \(h\) refers to the bandwidth, and \(K()\) refers to the kernel function.
[0059] Furthermore, the main idea of the random forest algorithm is to use the method of random sampling with replacement to extract multiple samples from the original samples to form a series of new sample sets, and then perform decision tree modeling on each sample set. The prediction values of multiple decision trees are combined by voting or taking the average to obtain the final prediction result. It is an important classification and regression algorithm in data mining; the construction of the coal body structure recognition model based on the random forest algorithm in step S6 specifically includes the following steps:
[0060] A. Division of the test set and the training set: Randomly select 75% of the data from all logging data as training data for model construction, and the remaining 25% of the data as test data for testing the model performance;
[0061] B. The variance inflation factor (VIF) and the maximal information coefficient (MIC) are used to analyze the correlation between different logging data indicators and the correlation between logging data indicators and coal structure respectively, and the logging types suitable for coal structure identification are optimized;
[0062] An overly high value of VIF indicates that the collinearity degree between different logging data indicators is too high, which may lead to the model relying too much on certain correlated variables, thus affecting the generalization performance. Generally, it is considered that when VIF > 10, there is multicollinearity;
[0063]
[0064] In the formula, VIF(X i ) refers to the VIF value of the i-th independent variable, refers to the independent variable X i and the coefficient of determination when regressing with other independent variables;
[0065] MIC is a statistic used to evaluate the strength of the relationship and non-linear dependence between two variables. It calculates the correlation degree between variables based on mutual information. First, the sample E = {x i , y i} i = 1, 2,..., N composed of variables X and Y is segmented along the x-axis and y-axis into a grid G of a×b. The probability I(E| G , a, b) that the points in the set E fall into the grid G is used as the mutual information between the independent variable X and the dependent variable Y in this segmentation case; according to different combinations G of the grid division, the mutual information in different cases is calculated, and the maximum value is selected as the final MIC;
[0066] I(X, Y) = I(E| G , a, b) (22)
[0067]
[0068] In the formula, a and b are the number of grid divisions, B is the grid upper limit, generally taking B = N 0.6 ;
[0069] C. Combining grid search and ten-fold cross-validation to optimize the model hyperparameters; referring to the parameter value ranges of the random forest algorithm in relevant research and the logging data characteristics of coal structure, the search space of the model hyperparameters is determined as shown in Table 1. Each hyperparameter searches in the search space starting from the initial value and increasing upward by an incremental value each time until the maximum value;
[0070] Table 1 Search space of model hyperparameters
[0071] Hyperparameter Starting value Maximum value Increment value Number of decision trees 10 1000 10 Maximum depth of the tree 2 100 1 Minimum number of samples per node 2 50 1
[0072] D. The accuracy (Accuracy, A), precision (Precision, P), recall (Recall, R), and F1-score (F-score, F1) are used to evaluate the performance of the model in coal seam structure recognition; the accuracy reflects the overall accuracy of the model classification, and is calculated as the ratio of the number of correctly classified samples to the total number of samples; the precision focuses on the proportion of samples that are actually positive among the samples predicted as positive by the model to evaluate the accuracy of the model in predicting positive classes; the recall (sensitivity or true positive rate) is used to measure the ability of the model to correctly identify positive class samples, that is, among all samples that are actually positive, the proportion of positive class samples successfully detected; the F1-score, as the harmonic mean of precision and recall, aims to achieve a balance between these two metrics; the formulas are as follows:
[0073]
[0074] In the formulas, TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives;
[0075] E. Test set model verification: Based on the constructed random forest model for coal seam structure recognition, the test samples are input into the model to obtain the prediction results of the coal seam structure.
[0076] Adopting the above technical solutions, compared with the prior art, the present invention has the following advantages and beneficial effects:
[0077] The present invention provides a method for identifying coal seam structures based on imbalanced logging data. After resampling the logging data set of the minority-class coal seam structures by using k-means clustering and ADASYN, a random forest model is constructed to identify various coal seam structures. This method can effectively eliminate the model learning bias caused by the imbalanced data distribution, improve the accuracy of the machine learning model in coal seam structure recognition, and provide a scientific basis for the prediction of sweet spots for coalbed methane extraction and the efficient development of coalbed methane resources. Description of the Drawings
[0078] Figure 1 It is a schematic diagram of the silhouette coefficient under different numbers of clusters;
[0079] Figure 2 It is a schematic diagram of the performance analysis of the constructed model in identifying different types of coal seam structures in the training set. Detailed Embodiments
[0080] Next, in combination with the above invention content, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0081] A method for identifying coal body structure based on unbalanced logging data of the present invention includes the following steps:
[0082] S1. Obtain logging data of different coal body structures in multiple wells: According to the coal body structure classification standard of GB / T 30050-2013, the coal samples taken by drilling cores are divided into four categories: primary structure coal, fragmented coal, granular coal, and mylonite coal. At the same time, collect the corresponding logging data. The extraction interval of the logging data is 0.125m, and the logging types include deep lateral resistivity (LLD), shallow lateral resistivity (LLS), spontaneous potential (SP), natural gamma (GR), density (DEN), neutron (CNL), acoustic wave (AC), and dual caliper logging (CALX, CALY). According to the burial depth (DEPTH), the numerical values of various logging curves are saved in the form of data columns. Taking a certain coal seam of a certain well in the southern margin of the Junggar Basin as an example, as shown in Table 2 below:
[0083] Table 2 Logging data sets of different coal body structures
[0084]
[0085] S2. Eliminate and standardize the outliers in the logging data: According to the existing lithology interpretation data, eliminate the outliers in the logging data, including the parting in the coal seam and the logging data of the coal seam within 1 meter from the top and bottom plates. Taking a certain coal seam of a certain well in the southern margin of the Junggar Basin as an example, as shown in Tables 3 and 4 below. Standardize the original logging data according to the above formulas (1)-(3).
[0086] Table 3 Table of eliminating outliers in coal seam logging data
[0087]
[0088]
[0089] Table 4 Table of eliminating outliers in parting logging data of coal seam
[0090]
[0091] S3. k-means clustering analysis: According to the algorithms of formulas (4)-(11), use the scikit-learn library of python to perform k-means clustering analysis on the coal body structure logging data, and generate the optimal number of clusters with the maximization of the silhouette coefficient as the goal. The results are as Figure 1As shown in the figure, the optimal number of clusters is 3. Therefore, each type of coal structure is clustered into 3 categories using the k-means method, providing basic data for ADASYN resampling.
[0092] S4. ADASYN hybrid resampling: There are a total of 680 groups of original data for coal structure modeling selected this time, including 235 groups of data for primary structure coal, 316 groups of data for fragmented coal, 96 groups of data for granular coal, and 33 groups of data for mylonitized coal. There is a significant imbalance in the data volume of different coal structures. To ensure the accuracy and stability of the model, taking the fragmented coal with the largest data volume as the standard, according to the optimal number of clusters, adaptive new sample synthesis is performed on the data of the remaining primary structure coal, granular coal, and mylonitized coal. According to the algorithms of formulas (12)-(17), the present invention uses the Imbalanced-learn library of Python to implement resampling of the coal structure of the minority category, so that the data of each coal structure is 316 groups.
[0093] S5. Evaluation of the data equalization processing effect based on KS test and information entropy: According to the coal structure data obtained by ADASYN hybrid resampling, calculate the KS test and information entropy before and after equalization processing according to formulas (18)-(20). In this paper, the SciPy and NumPy libraries of Python are used to implement the KS test and information entropy calculation of coal structure data equalization. The P values of the KS test for different coal structures are shown in Table 5, all of which are greater than 0.05, indicating that the distribution of coal structure data before and after resampling has not changed. The original value of information entropy is 2.02, and the value after resampling is 2.04. The information entropy has a small increase, indicating that the amount of original data information has increased after resampling. The above proves that the proposed method combining k-means clustering and ADASYN hybrid resampling can effectively balance the data.
[0094] Table 5 List of P values of KS test for different coal structures
[0095] Category Primary structural coal Fragmented coal Granular coal Mylonitized coal Average KS test P-value 0.68 1 0.16 0.20 0.51
[0096] S6. Construction of a coal structure recognition model based on the random forest algorithm: Randomly extract 25% of the data of each coal structure after expansion as the validation set, and the remaining data as the training set of the random forest model. Calculate the VIF and MIC values of each logging curve according to formulas (21)-(23), and analyze the correlation between different logging data indicators and the correlation between logging data indicators and coal structures. Then, use the method of grid search combined with ten-fold cross-validation to optimize the model hyperparameters. According to the results of running the random forest code, the optimal number of decision trees for the model parameters is 100, and the maximum depth of the tree is 50. According to formulas 24-27, calculate the coal structure recognition effect of the random forest model constructed based on the optimal hyperparameters under the training sample event, such asFigure 2 As shown, the values of A, P, R, and F1 are all above 0.83, indicating that the model can achieve good coal structure recognition under the training samples. Under the test samples, the values of A, P, R, and F1 for coal structure recognition are all above 0.94, indicating that the constructed model can generally recognize the coal structure more accurately. As Figure 2 shown.
[0097] The above embodiments illustrate the basic principles and features of the present invention. However, the above only illustrates the preferred embodiments of the present invention and is not limited by the embodiments. Those of ordinary skill in the art, inspired by this patent, can also make many forms of deformation and improvement without departing from the spirit of the present invention and the scope protected by the claims. These all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be subject to the appended claims.
Claims
1. A method for identifying coal structure based on non-equilibrium logging data, characterized in that: The following steps are involved: S1. Obtain logging data of different coal structures in multiple wells; S2, outlier removal and standardization of logging data; S3, k-means cluster analysis; S4, ADASYN hybrid resampling; S5. Evaluation of data equalization processing effect based on KS test and information entropy; S6. Construction of coal structure identification model based on random forest algorithm.
2. The method for identifying coal structure based on non-equilibrium logging data according to claim 1, characterized in that: The specific process of step S1 is as follows: according to the coal body structure classification standard of GB / T 30050-2013, the coal samples cored from drilling are divided into four categories: primary structure coal, crushed coal, crushed coal and mylonitic coal; at the same time, the corresponding logging data are collected, the logging data extraction interval is 0.125m, and the logging types include deep lateral resistivity (LLD), shallow lateral resistivity (LLS), spontaneous potential (SP), natural gamma (GR), density (DEN), neutron (CNL), acoustic wave (AC) and dual-caliper logging (CALX, CALY); according to the burial depth (DEPTH), the values of various logging curves are saved in the form of data columns.
3. The method for identifying coal structure based on non-equilibrium logging data according to claim 2, characterized in that: The specific process of step S2 is as follows: In order to ensure the quality and consistency of logging data of different coal structures, it is necessary to eliminate abnormal data, including interlayers in coal seams and logging data within 1 meter from the roof and floor plates; this is because interlayers are usually composed of clay rock, carbonaceous mudstone and siltstone, and their logging values are significantly different from those of coal; and the logging data of coal seams that are too close to the roof and floor plates will be affected by the adjacent rock formations, and the amplitude changes greatly, resulting in abnormal distribution of logging data of coal structure; in addition, different types of logging data have different dimensions, and the Z-score method needs to be used to standardize the original data of various logging curves. The formula is as follows: In formulas (1)-(3), M refers to the type of logging curve, m refers to the mth type of logging curve; n represents the number of samples of each logging curve, i refers to the logging value of the i-th sample; Z mi Refers to the standardized value of the i-th sample of the m-th type of logging curve; x mi Refers to the original logging value of the i-th sample of the m-th type of logging curve; μ m and σ m They refer to the mean and standard deviation of the mth type of logging curve respectively.
4. The method for identifying coal structure based on non-equilibrium logging data according to claim 3, characterized in that: In order to avoid the phenomenon that the new samples generated in the adaptive synthetic sampling (ADASYN) process are overly concentrated in certain categories, thereby changing the distribution characteristics of the original data set and increasing the uncertainty of the data, k-means clustering analysis is performed on the original data of each coal structure, and the silhouette coefficient is used to select the optimal number of clusters, and then ADASYN is used to resample each cluster independently; the k-means clustering analysis in step 3 specifically includes the following steps: ① Divide the coal structure data into p groups according to the coal structure category D={X j }j=1,2,...,p (4) X j ={x ij }i=1,2,...,n p (5) Where D refers to the coal structure dataset, X j refers to the j-th coal structure data set, x ij refers to the i-th sample in the j-th coal structure, n p It refers to the number of samples of the p-th coal structure; ② For each type of coal structure data, k initial cluster centers are randomly selected; In the formula, Refers to the set of initial cluster centers in the j-th coal structure, It refers to the kth initial cluster center in the jth coal structure; ③For any data point x ij , calculate the Euclidean distance from each sample to the k cluster centers, and assign the sample to the nearest cluster center according to the principle of non-intersection; In the formula, Refers to the samples of the uth cluster category in the jth coal structure, and They refer to any data point x in the j-th coal structure. ij The distance to the u-th and v-th cluster centers; ④ For the classification results calculated in formula (3), the cluster center is recalculated using the category mean; Where n uj Refers to the number of samples in the u-th cluster category in the j-th coal structure, refers to the recalculated cluster center in the j-th coal structure, It refers to the kth cluster center among the cluster centers recalculated in the jth coal structure; ⑤ Repeat steps ③ and ④ to continuously update the cluster center until The iteration is stopped when t is reached, and the clustering result of the coal body structure data is taken as the tth update; ⑥ Select the silhouette coefficient that represents the clustering effect to select the optimal number of clusters; change the number of clusters k, calculate the silhouette coefficient of the clustering results for each number of clusters k, and select the number of clusters with the largest silhouette coefficient as the optimal number of clusters to provide basic data for ADASYN resampling; In the formula, S(k) refers to the silhouette coefficient of the coal structure data set under the kth cluster number, S(x ij ) refers to the sample x ij The silhouette coefficient, n k refers to the maximum number of cluster centers, n refers to the number of samples of all coal structure data, and a(ij) refers to the number of samples x in the jth coal structure. ij The average distance to other samples in the category, b(ij) refers to the distance of sample x in the j-th coal structure. ij The average distance to samples of other categories.
5. The method for identifying coal structure based on non-equilibrium logging data according to claim 4, characterized in that: ADASYN in step S4 is a resampling method for unbalanced data sets, which has adaptive characteristics for the generation of minority class samples. By generating minority class samples, the data volume of each coal body structure is evenly distributed to improve the performance of the classification model. Step S4 specifically includes the following steps: (1) According to the number of data of each coal structure n j (j=1,2,...,p), resample all coal structure data with smaller data volume according to the maximization principle, so that the sample data volume of each coal structure data after resampling is the same as the sample data volume of the coal structure with the largest data volume; In the formula, G refers to the amount of data that needs to be achieved after resampling for each type of coal structure data that needs to be resampled, n j (max) refers to the number of coal structure samples with the maximum amount of data; (2) Based on the k-means clustering results of each type of coal structure data, the number of samples in each category is counted, and the minority sample is recorded as m s , obtain the number of synthetic data g; In the formula, β refers to the parameter of synthetic data, β∈[0,1], β=1 means that the balanced data set is generated; (3) For each type of coal structure data sample obtained by clustering each type of coal structure in (2), the Euclidean distance is used to identify the s nearest neighbor samples of the sample and calculate their ratio r ui , and normalize it; Where Δui refers to the sample x ui The number of majority samples among the s nearest neighbor samples, n s refers to the total number of samples of the minority class, refers to r ui The normalized result of (4) Based on the calculated number of synthetic samples and the normalized results, the number of synthetic samples g for each minority class of coal structure sample data is calculated. ui , on this basis, generate each minority class sample x ui The synthetic data ui ; s ui =x ui +(x zi -x ui )λ(17) In the formula, x ui refers to the i-th sample of the u-th minority class, x zi Refers to x ui A few samples are randomly selected from the s nearest neighbor samples, λ refers to a random number, λ∈[0,1]; For each type of coal structure data, repeat steps (2)-(4) to generate synthetic samples that meet the data volume requirements, so that the data volume of each coal structure after resampling is n j (max).
6. The method for identifying coal structure based on non-equilibrium logging data according to claim 5, characterized in that: The specific process of step S5 is as follows: The KS (Kolmogorov-Smirnov) test and information entropy can test whether the equalization process has changed the original distribution characteristics of the sample data and whether it has increased the amount of information in the sample data. The specific calculation method is as follows: Ⅰ. KS test According to the maximum difference of the empirical distribution function between the original coal structure data and the resampled coal structure data, it is analyzed whether the distribution of the data has changed after the equalization treatment. It is generally believed that if the P value of the KS test statistic d_{max} is greater than the significance level of 0.05, it can be considered that the equalization treatment has not changed the distribution of the original data; d_{max}=max|F′(x)-F(x)| (18) In the formula, F(x) and F'(x) refer to the empirical distribution functions of the original data and the sample data after equalization processing, respectively, and d_{max} refers to the statistic of the KS test; II. Information Entropy The probability density function of the coal structure sample data before and after equalization is constructed, and then the information entropy of the two is calculated respectively; since the coal structure data is discrete and the distribution is unknown, the kernel density estimation method is selected with reference to relevant research to estimate the probability density function of the coal structure sample; In the formula, refers to the estimated probability density function, H(X) refers to the information entropy, N refers to the total number of samples, h refers to the bandwidth, and K() refers to the kernel function.
7. The method for identifying coal structure based on non-equilibrium logging data according to claim 6, characterized in that: The main idea of the random forest algorithm is to use a random sampling method with replacement to extract multiple samples from the original sample to form a series of new sample sets, and then perform decision tree modeling on each sample set, and combine the prediction values of multiple decision trees by voting or taking the average value to obtain the final prediction result. It is an important classification regression algorithm in data mining; the construction of the coal structure recognition model based on the random forest algorithm in step S6 specifically includes the following steps: A. Division of test set and training set: 75% of the data are randomly selected from all the well logging data as training data for model construction, and the remaining 25% of the data are used as test data to test model performance; B. Variance Inflation Factor (VIF) and Maximum Information Coefficient (MIC) are used to analyze the correlation between different logging data indicators and the correlation between logging data indicators and coal body structure, and the logging type suitable for coal body structure identification is selected; Too high a value of VIF indicates that the degree of collinearity between different logging data indicators is too high, which may cause the model to be overly dependent on certain related variables, thus affecting the generalization performance. It is generally believed that when VIF>10, it indicates the existence of multicollinearity; In the formula, VIF(X i ) refers to the VIF value of the ith independent variable, Refers to the independent variable X i The coefficient of determination when regressing with other independent variables; MIC is a statistic used to evaluate the strength of the relationship and nonlinear dependence between two variables. It is based on mutual information to calculate the correlation between variables. First, data segmentation is used to divide the sample E composed of variables X and Y into {x i ,y i }i=1,2,...,N, divided into a×b grid G along the x-axis and y-axis, with the probability I(E| G ,a,b) as the mutual information between the independent variable X and the dependent variable Y in this segmentation situation; according to different combinations of grid division G, the mutual information in different situations is calculated, and the maximum value is selected as the final MIC; I(X,Y)=I(E| G ,a,b)(22) In the formula, a and b are the number of grid divisions, and B is the upper limit of the grid, which is generally taken as B = N 0.6 ; C. Combine grid search and ten-fold cross validation to optimize the model hyperparameters; refer to the parameter value range of the random forest algorithm in related studies and the logging data characteristics of coal body structure, determine the model hyperparameter search space as shown in Table 1, and each hyperparameter in the search space starts from the starting value and increases upward each time according to the incremental value until the maximum value; Table 1 Search space of model hyperparameters D. Use accuracy (A), precision (P), recall (R) and F-score (F1) to evaluate the performance of the model in coal structure identification; accuracy reflects the overall accuracy of model classification, calculated as the ratio of the number of correctly classified samples to the total number of samples; precision focuses on the proportion of samples that are actually positive among the samples predicted by the model to be positive, in order to evaluate the accuracy of the model in predicting the positive class; recall (sensitivity or true positive rate) is used to measure the ability of the model to correctly identify positive samples, that is, the proportion of positive samples that are successfully detected among all samples that are actually positive; F1 score is the harmonic mean of precision and recall, aiming to achieve a balance between these two indicators; the formula is as follows: Where TP is the number of true positive examples, TN is the number of true negative examples, FP is the number of false positive examples, and FN is the number of false negative examples; E. Test set model verification: Based on the constructed random forest model for coal structure identification, the test samples were input into the model to obtain the prediction results of the coal structure.