Rock-soil mass mechanical parameter clustering inversion method based on drilling data
By performing zero-mean, logistic normalization, and principal component analysis on borehole data, combined with fuzzy C-means clustering, the accuracy and efficiency issues of soil and rock mass parameter inversion were resolved, enabling efficient and accurate prediction of soil and rock mass parameters and supporting intelligent decision-making in geotechnical engineering.
Patent Information
- Application Number
- CN202511847822.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-01-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies rely on field sampling and laboratory testing in geotechnical engineering investigations, resulting in a large workload, long time consumption, and delayed data acquisition. Furthermore, traditional regression methods are not accurate enough when processing highly nonlinear data of soil and rock masses, making it difficult to accurately invert soil and rock parameters.
The borehole data were simplified using zero-mean, logistic normalization, and principal component analysis. The soil and rock samples were grouped using fuzzy C-means clustering to establish multiple sub-models. Predictions were made by adding the membership weights and combining the borehole parameters to retrieve the soil and rock parameters.
It improves the accuracy and engineering applicability of soil and rock parameters inversion, reduces reliance on traditional prior information, and supports intelligent decision-making in soil and rock exploration.
Smart Images

Figure CN121365604A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of geotechnical engineering investigation, and particularly relates to a geotechnical body mechanics parameter clustering inversion method based on drilling data. BACKGROUND
[0002] Under the joint action of geological structure and long-term evolution, the geotechnical body shows non-uniformity, discontinuity and anisotropy with its complex structural characteristics, making it one of the most complex media on the earth at present. The geotechnical body medium exhibits different physical and mechanical properties due to different components and structures, and these properties can usually be characterized by mechanical parameters such as strength, integrity, hardness and abrasiveness.
[0003] Generally, the measurement of geotechnical body parameters through geological investigation is a necessary step in the process of geotechnical engineering construction. Cases of improper engineering design, delayed construction period, cost waste, and manpower increase caused by unknown or inaccurate geotechnical body parameters are common, and even more, catastrophic consequences such as foundation instability, slope sliding, and tunnel collapse often accompany significant economic losses and casualties. Therefore, accurate geotechnical body parameter investigation and testing are the premise of ensuring safe and efficient construction of projects.
[0004] However, at present, geotechnical engineering investigation relies on field sampling and indoor mechanical testing, which is time-consuming and has a lag in data acquisition. In this regard, the drilling-while-measuring technology in the field of oil exploration provides a reliable approach to quickly obtaining engineering geotechnical body mechanical parameters. This method integrates sensor measurement instruments on the drilling machine to obtain electromagnetic, vibration and other signals of the geotechnical body during drilling, as well as drilling pressure, torque and other parameters. These drilling data are closely related to the mechanical parameters of the geotechnical body in the drilled stratum. If this relationship can be captured through regression calculation or machine learning methods, it is expected to accurately invert the mechanical parameters of the geotechnical body in the drilled stratum, providing accurate and practical prior geological information for the scientific design and efficient construction of geotechnical engineering.
[0005] Under the existing technology, the main difficulties in predicting geotechnical body parameters based on drilling data are as follows: 1. The internal structure of the geotechnical body is complex, and its parameters have strong discreteness, which makes the relationship between drilling parameters and geotechnical body parameters also show strong nonlinearity. However, current geotechnical body parameter inversion models mostly rely on traditional regression methods, which have limited effect on strong nonlinear data sets, resulting in limited inversion accuracy of geotechnical body parameters. Given the excellent performance of machine learning and other artificial intelligence technologies in handling nonlinear data, using artificial intelligence technology to establish a relationship model between drilling parameters and geotechnical body parameters is a feasible approach to improving the inversion accuracy of geotechnical body parameters.
[0006] 2、In its complex rock mass condition background, the correlation between drilling data and rock mass parameters is influenced by many factors such as lithology, rock mass composition, groundwater, etc., and single mapping is difficult to apply to all geological conditions. Multiple mapping based on grouped rock mass conditions has higher focusing degree than single mapping, and can always increase the step of classifying rock mass conditions by experience. The clustering algorithm of machine learning can be used as a more scientific classification method to provide a good reference for predicting rock and soil parameters. SUMMARY
[0007] To solve the problems of the prior art, the present application provides a rock and soil mass mechanical parameter clustering inversion method based on drilling data, which simplifies drilling data by using zero mean, logical normalization and principal component analysis method for complex drilling data parameters, groups the field samples into multiple clusters by using clustering method, takes drilling parameters as input and rock mass parameters as output, trains each sub-model by the membership weight of the sample, obtains multiple sub-models, and multiplies the predicted results of all sub-models by the membership weight and adds them up to record the predicted results of the to-be-predicted sample.
[0008] The technical scheme adopted by the present application is as follows: A rock and soil mass mechanical parameter clustering inversion method based on drilling data, comprising the following steps: A. Data arrangement: Drilling data is obtained through monitoring and data system, rock mass conditions are divided into different clusters based on the data, multiple sub-models are obtained by using samples in different clusters, and the sub-models are regarded as known quantities, while rock mass parameters are regarded as unknown quantities for prediction, the unknown quantities are clustered into clusters according to the grouping standard of the known quantities; B. Normalization of drilling parameters: Zero mean normalization is used to eliminate the size and dimension of drilling data, so that the original features are dimensionless and the large values are eliminated, and then logical normalization is used to balance the fluctuation difference between features, amplify the floating strength of some features near some constants, and reduce the negative influence of abnormal data; Unlike the conventional excavation, blasting and other construction processes, the drilling machine drilling is a construction process similar to piston movement, which has significant regular characteristics. Correspondingly, drilling data often has high-frequency cyclic characteristics. Conventional engineering data usually simply uses mean value to describe these high-frequency cyclic data, resulting in a large amount of data information loss. In this method, the mean value, frequency and amplitude of each drilling data are used to describe the three sub-parameters; C. Dimension reduction of drilling data based on principal component analysis; The principal component analysis technology is used to select the characteristic value with high average absolute value, and the corresponding characteristic vector is regarded as a new independent variable used in clustering, so as to reduce the number of features and simplify the drilling data; D. Clustering and prediction method based on fuzzy C-means algorithm The field samples are grouped into multiple clusters by using fuzzy C-means clustering, multiple sub-models are obtained by using samples in different clusters; the results of all sub-models are multiplied by membership weight and added to obtain a membership matrix, each sub-model is trained by the membership weight of the sample, the drilling parameters are taken as input, the rock mass parameters are taken as output, and the prediction results of the to-be-predicted samples are recorded.
[0009] Preferably, the rock mass parameters include but are not limited to uniaxial / triaxial compressive strength, joint frequency, cohesion, internal friction angle, and elastic modulus.
[0010] The rock mass strength index is the rock strength index of the complete rock section in the borehole, which generally corresponds to the test strength of various rock samples, but there may still be a phenomenon that the actual strength of the rock mass is greatly different from the rock strength index of the sampling section.
[0011] Preferably, the drilling data include but are not limited to drilling distance, effective axial pressure, drilling tool rotation speed, torque, pressure, and drilling rate.
[0012] Preferably, the sampling of the field samples aims to prepare geotechnical samples similar to the physical properties of the field geotechnical body, to prepare original test materials for indoor tests, and the means include but are not limited to field drilling coring, standard penetration testing, point load testing, and lithology identification.
[0013] Preferably, the indoor test aims to obtain geotechnical physical and mechanical parameters and penetration test data with a corresponding relationship, including in-situ tests and numerical simulation.
[0014] Preferably, the field samples are data sets collected on site, and each sample contains drilling parameters and rock mass parameters of the same borehole.
[0015] Preferably, zero-mean normalization is that only data whose values conform to a standard normal distribution with a mean of 0 and a variance of 1 are considered suitable for clustering; Zero-mean normalization is that only data whose values conform to a standard normal distribution with a mean of 0 and a variance of 1 are considered suitable for clustering, but most of its normal values are close to 0, which makes it extremely susceptible to abnormal values, and logical normalization should be used to amplify the influence of these features.
[0016] The abnormal values do not occur in real situations and can be divided into the following three cases: The drilling rate is reasonable, while the rotation speed, thrust, or torque is close to 0; The data exceeds the critical value or appears an abnormal negative value; The thrust and torque are normal, while the drilling rate is close to 0.
[0017] The technical scheme provided by the present application brings the beneficial effects that: The rock mass mechanical parameter clustering inversion method based on drilling data provided by the present application can effectively solve the volume and structural differences between rock mass parameters and drilling data and the acceptable rock mass parameter precision prediction, and in view of the uncertain negative influence of abnormal data on the prediction result, three simple standards are adopted to judge whether the data is abnormal, and a machine algorithm is used to quickly organize the data, and at the same time, the statistical result can also be used to predict the rock mass parameters newly encountered in the rock mass survey process.
[0018] The rock mass mechanical parameters obtained by the present application reduce the excessive dependence of traditional inversion on prior information to a certain extent, improve the parameter identification precision and engineering applicability, and provide theoretical support for intelligent decision-making of rock mass survey. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0020] Figure 1 The overall method flowchart of the rock mass mechanical parameter clustering inversion method based on drilling data of the present application is shown in the figure. Figure 2 In the rock mass mechanical parameter clustering inversion method based on drilling data of the present application, the sub parameter table of each drilling parameter is shown in the figure. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings. Embodiment one
[0022] As shown in the figure, the present embodiment provides a rock mass mechanical parameter clustering inversion method based on drilling data, which includes the following steps: Figure 1 A, data organization; the related parameters needed include drilling parameters and rock mass parameter data.
[0023] The borehole data often contains a large amount of redundant information, and the size thereof does not match the amount of rock mass data, so the data needs to be arranged to eliminate redundancy and match the borehole data with the rock mass data. In various types of rock mass, multiple boreholes are tested, the borehole data collected within the borehole range is recorded, and the rock mass parameters of the corresponding types of rock mass are obtained through indoor testing to establish a key-value type database matching the rock mass parameters with the borehole parameters. Among them, the borehole data sampling frequency is high, and the data characteristics fluctuate with time, so the average of each drilling parameter within the borehole range is taken as the sample feature, and the following formula is used for calculation: (1) In the formula, is the result of the jth feature of the ith sample, represents the value measured at the kth second, and n is the total number of seconds spent by the ith part; B, normalization of borehole parameters: The zero-mean normalization can eliminate the size and dimension of the borehole data, making the original features dimensionless and eliminating large values, and the following formula is used for calculation: (2) In the formula, μ is the mean, σ is the variance, is the mean and variance variable, and is the normalization result; The logical normalization can balance the fluctuation difference between features, amplify the floating strength of some features near some constants, and reduce the negative impact of abnormal data, and the following formula is used for calculation: (3) C, dimensionality reduction of borehole data based on principal component analysis; The principal component analysis method uses orthogonal transformation to reorganize the original variables to form new independent variables, calculates the covariance between the processed features based on the normalized data to form a covariance matrix, selects the feature value with a higher average absolute value and its eigenvector, and regards the eigenvector as a new independent variable in clustering. Taking the data with m samples and n features as an example, the following formula is used: (4) (5) In the formula, is the covariance between the ith and jth features, λ is the eigenvalue, and μ is the eigenvector corresponding to the eigenvalue; The absolute value of the covariance represents the mutual influence of the ith and jth features, and the influence degree is positively correlated with the absolute value μ; The n-dimensional matrix C contains n eigenvalues and corresponding eigenvectors, which can be represented as multiple solutions of formula (5), The absolute value is used to measure the amount of information contained in the processed feature by the i-th feature vector; D. Clustering and prediction methods based on fuzzy C-means algorithm; First, initial centroids are randomly selected from the samples. An equation is constructed using the silhouette coefficients of K-means clustering. For each value of c, five fuzzy C-means clustering operations are performed using different random initial centroids. The number of clusters c is determined based on the characteristics of the dataset, aiming to maximize the distance between different clusters and minimize the distance between samples within the same cluster, calculated using the following formula: (6) (7) (8) In the formula, Let be the average distance between the j-th sample and the other samples in the same group; Let j be the average distance between the j-th sample and samples from different groups; Let be the number of samples in the i-th cluster. This is the average silhouette coefficient of the entire dataset with n samples.
[0024] Since samples within the same group obtained by fuzzy C-means clustering are relatively similar, multiple machine learning mappings are used to address the issue of samples belonging to different groups. Each sample partially belongs to all groups and has a certain membership degree, but does not completely belong to a specific group. The grouping result is represented by a membership matrix, with each sample having a corresponding membership vector. Taking the division of N samples into K groups as an example, the grouping result is a... The matrix is obtained through the following objective function: (9) (10) In the formula, Establish the objective function for the sub-model of the i-th sample group. For the data of the j-th sample, Let be the center point of the i-th group, and m be a fuzzification parameter greater than 1. The weights are between 0 and 1 and are randomly generated before iteration. The sum of the weights for the same sample is 1.
[0025] The center point of each group is calculated by minimizing the objective function during the iteration process. Once the center point is determined, the weight matrix can be updated according to the following formula: (11) (12) During multiple iterations, the objective function value continuously decreases, and the iteration continues until the decrease is less than a specific value each time. Then, the iteration stops, and K sub-models are established. The prediction results of all sub-models are multiplied by the membership weights and summed to obtain the prediction result of the sample to be predicted.
[0026] The silhouette coefficient is a metric used to evaluate the effectiveness of clustering algorithms. It measures the density and separation of data points in a cluster, and the reasonableness of the clustering result is reflected in how close the value is to 1.
[0027] The prediction results are expressed as a range of percentages for each group of data. The i-th and j-th percentages are selected as the prediction results. Based on the clustering results, the i-th percentile value and the j-th percentile value of each group are obtained to obtain the prediction results of the soil and rock parameters.
[0028] The i-th percentile indicates that i% of the samples in each group have lower values than its soil and rock parameters. The values of i and j need to be selected according to actual needs. The larger the difference, the worse the prediction range, and the greater the probability that the soil and rock parameters are within the prediction range.
[0029] In this embodiment, logical normalization is a linear representation method that makes the original positive data with a large absolute value closer to 1, and the original negative data closer to 0. Features close to 0 have a larger gradient. By amplifying the fluctuation intensity of some features near a certain constant, the negative impact of outlier data is reduced.
[0030] Principal component analysis reorganizes the original variables through orthogonal transformations, selects eigenvalues with high average absolute values, and uses the corresponding eigenvectors as new independent variables for clustering.
[0031] Fuzzy C-means clustering optimizes cluster centers and membership degrees of data points by minimizing an objective function based on preprocessed and refined data. Instead of classifying samples into specific clusters, it uses the membership degree of each sample to each cluster to represent the result.
[0032] Suppose a dataset contains multiple sets of borehole data, which are decomposed into n clusters using cluster analysis. Each of these n clusters can be used to establish multiple sub-models for predicting soil and rock parameters, and the resulting predictions are denoted as y1, y2, ..., y... n; The membership vector of a certain test sample is denoted as [w1, w2, ..., w n That is, its membership degree to the first cluster is w1, and its membership degree to the nth cluster is w n Then its prediction result y is:
[0033] The membership is obtained according to the clustering result, and the specificity of the sub-model to the samples in the corresponding cluster can be improved, the higher the membership is, the closer the sample is to the data characteristics of the group, and the greater the influence on the mapping established according to the samples in the group is.
[0034] The machine learning algorithm used in the embodiment includes but is not limited to zero mean normalization, principal component analysis technology, fuzzy C-means clustering and the like, and also includes traditional t-distributed stochastic neighbor embedding algorithm, artificial neural network, support vector regression and other algorithms.
[0035] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application, and any modification, equivalent replacement, improvement and the like made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for clustering and inversion of rock mass mechanical parameters based on borehole data, comprising the following steps: A. Data arrangement: Obtain borehole data through monitoring and data system, divide rock mass conditions into different clusters based on the data, obtain multiple sub-models by using samples in different clusters, and take them as known quantities, while rock mass parameters as unknown quantities for prediction, cluster unknown quantities into clusters according to known quantities as grouping standard; B. Normalization of borehole parameters: Use zero-mean normalization to eliminate the size and dimension of borehole data, so that the original features are dimensionless and the large values are eliminated, and then use logical normalization to balance the fluctuation difference between features; C. Dimension reduction of borehole data based on principal component analysis: Use principal component analysis technology to select feature values with high average absolute value, and take the corresponding eigenvectors as new independent variables used in clustering, so as to reduce the number of features and simplify the borehole data; D. Clustering and prediction method based on fuzzy C-means algorithm: Use fuzzy C-means clustering to group field samples into multiple clusters, obtain multiple sub-models by using samples in different clusters, multiply the prediction results of all sub-models by membership weight and add them up to obtain membership matrix, train each sub-model through the membership weight of the sample, take borehole parameters as input and rock mass parameters as output, and record the prediction results of the predicted samples.
2. The method according to claim 1, wherein, The rock mass parameters include but are not limited to uniaxial / triaxial compressive strength, joint frequency, cohesion, internal friction angle and elastic modulus.
3. The method according to claim 1, wherein, The borehole data includes but is not limited to drilling distance, effective axial pressure, drilling tool rotation speed, torque, pressure and drilling rate.
4. The method of claim 1, wherein, The sampling of field samples aims to prepare rock mass samples similar to the physical properties of field rock mass, and prepare original test materials for laboratory tests, which includes but is not limited to field drilling coring, standard penetration test, point load test and lithology identification.
5. The method of claim 4, wherein, The laboratory test aims to obtain rock mass physical and mechanical parameters and penetration test data with corresponding relationship, including in-situ test and numerical simulation.
6. The method according to claim 1 or 4, characterized in that, The field samples are data sets collected on site, each of which contains borehole parameters and rock mass parameters of the same borehole.
7. The method of claim 1, wherein, The zero-mean normalization is considered suitable for clustering only when the data value obeys standard normal distribution, the mean value is 0 and the variance is 1; The logical normalization is a linear representation method, which makes the original positive data with large absolute value close to 1, and the original negative data close to 0; The principal component analysis method reorganizes the original variables through orthogonal transformation, selects feature values with high average absolute value, and takes the corresponding eigenvectors as new independent variables used in clustering; The fuzzy C-means clustering optimizes the cluster center and the membership of each data point by minimizing the objective function on the basis of pre-processing and refining data, and uses the membership of each sample to each cluster to represent the result, instead of dividing the sample into a specific cluster.