A method and system for predicting single pile self-sinking depth based on machine learning
By employing machine learning and deep interval feature encoding, the accuracy and adaptability issues of traditional methods in predicting the self-sinking depth of single piles are resolved, achieving high-precision and stable self-sinking depth prediction applicable to complex geological conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN PORT ENG INST LTD OF CCCC FIRST HARBOR ENG
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional methods are not accurate enough in predicting the self-sinking depth of a single pile, are difficult to adapt to complex strata, have weak generalization ability, cannot effectively handle soil resistance fluctuations caused by thin hard layers or soft interlayers, and cannot automatically learn complex patterns from multidimensional characteristic engineering data.
Machine learning methods are employed to collect historical engineering data, construct a dataset, encode deep interval features, train models such as XGBoost, and output prediction results of sinking depth to adapt to different geological conditions.
It achieves sub-meter level high-precision prediction, significantly improves the reliability of construction control, adapts to complex geological conditions, and reduces engineering risks and costs.
Smart Images

Figure CN121525533B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning technology, and in particular relates to a method and system for predicting the self-sinking depth of a single pile based on machine learning. Background Technology
[0002] As marine engineering progresses and its scale expands, the weight and size of marine structures also increase. Large-diameter monopile foundations, due to their superior bearing capacity and relatively mature construction technology, have become one of the most common foundation types in current marine engineering. During monopile installation, the first stage is "self-sinking," where the pile, after being hoisted into place, penetrates the seabed under its own weight and the force of the installation hammer, without the application of additional driving energy.
[0003] The penetration depth during this self-sinking stage (also known as the "free depth into the mud") has a significant impact on construction safety and subsequent procedures. On the one hand, if the penetration depth is insufficient, the free length of the pile above the mud surface will be excessive, forming a cantilever structure with a high slenderness ratio. Under the action of hammer weight and environmental loads, this structure is prone to buckling instability, endangering structural safety. On the other hand, as the diameter of a single pile and the weight of the hammer increase, the self-weight of the pile-hammer system increases significantly. In complex seabed strata, this may lead to unexpected "pile slippage," where the pile penetrates rapidly and continuously during the self-sinking stage. If not properly controlled, this could result in damage to construction equipment or even personnel accidents. Therefore, accurately predicting the self-sinking depth of a single pile is not only the foundation for optimizing installation techniques and controlling the construction period and costs, but also a crucial prerequisite for ensuring construction safety and mitigating engineering risks.
[0004] Currently, the main method for predicting the self-sinking depth of a single pile in the industry relies on traditional semi-empirical or analytical models. The core principle is based on static equilibrium, which compares the weight of the pile hammer system with the calculated total soil resistance.
[0005] One application of total soil resistance is to assess the ultimate bearing capacity of the soil on the pile by calculating the sum of the pile side friction and the pile end resistance, thereby determining the self-settlement depth. The specific steps are as follows:
[0006] Calculate the total resistance of the pile side Q s Discretize the soil layers penetrated by the pile along its depth. For each soil layer, calculate the unit side friction resistance using the corresponding empirical formula based on its soil type (cohesive or non-cohesive). f s Multiply by the corresponding pile side surface area of that layer. A s Finally, sum over all soil layers. For cohesive soil, the formula is: f s = α × S u ;inα The empirical adhesion coefficient, S u For soil, the undrained shear strength is given by the formula: f s =K× in v0 ×tan d Where K is the lateral earth pressure coefficient, in v0 For vertical effective stress, d The angle of friction between the pile and the soil.
[0007] Calculate the total resistance at the pile end Q p Based on the soil type at the pile tip, the unit end resistance is calculated using the corresponding empirical formula. q p Multiply by the pile tip area A p For cohesive soil, q p = N c × S u For non-cohesive soils, q p = N q × in v0 ;in, N c and N q This is the bearing capacity coefficient.
[0008] Determine the self-sinking depth: Starting from the seabed surface, gradually accumulate the total soil resistance at different depths ( R = Q s + Q p When the total resistance R The depth at which the weight of the pile hammer system first equals or exceeds the weight of the pile hammer system is the predicted self-sinking depth.
[0009] Besides the API method, other more sophisticated calculation methods exist in the industry, such as ICP-05 and UWA-05. However, these methods require detailed static cone penetration test (CPT) data for specific sites, limiting their application to data availability. Therefore, the API method obtains soil parameters (such as soil type, soil structure, etc.) through conventional laboratory tests. S u , d In practice, it is still widely regarded as standard practice and is commonly used.
[0010] Although traditional prediction methods, represented by the API method, provide a mature engineering framework, they have significant limitations when dealing with complex real-world geological conditions, mainly in the following four aspects:
[0011] Prediction accuracy is limited and error is relatively large: traditional methods rely on a series of simplified, generalized empirical coefficients (such as...). α K N c , N q These coefficients are derived from the summary of specific engineering experience and fail to fully reflect the complex pile-soil interaction mechanism on specific sites, especially in areas with large spatial variability of soil layers.
[0012] It is difficult to effectively handle complex layered strata: Marine foundations are typically composed of alternating layers of cohesive and non-cohesive soils, whose mechanical properties vary significantly vertically. Traditional methods use a simple superposition approach after layered calculations, which is essentially a "layer averaging" treatment. This approach cannot capture the dramatic fluctuations in soil resistance over short distances caused by the presence of thin hard layers or soft interlayers, nor can it model nonlinear effects such as interlayer interactions. Therefore, the reliability of predictions decreases in layered or stratified soil foundations.
[0013] Model simplification leads to insufficient generalization ability: Traditional models are based on a series of idealized assumptions (such as homogeneous soil, isotropic soil, and fixed failure modes). When applied to new sites with geological conditions significantly different from those at the time the model was established, the model struggles to adapt and exhibits weak generalization ability. This results in poor stability of its predictions across different geographical regions or under complex geological conditions, increasing the uncertainty and risk of engineering projects.
[0014] Traditional methods build models based on fixed physical experience, and their performance is limited by the assumptions of the formula itself: the method is difficult to automatically learn deep, non-explicit complex laws and relationships from accumulated multi-dimensional feature engineering data (such as pile parameters and continuously changing soil layer information), resulting in insufficient ability to mine the implicit laws of the data. Summary of the Invention
[0015] To address the aforementioned technical problems, this invention provides a machine learning-based method and system for predicting the self-sinking depth of a single pile, aiming to solve the problems of insufficient accuracy and poor adaptability to complex geological formations in traditional empirical methods.
[0016] To achieve the above-mentioned objectives, the first objective of this invention is to provide a method for predicting the self-sinking depth of a single pile based on machine learning, comprising:
[0017] S1. Collect historical engineering data and construct a dataset. Each data sample includes: pile parameters, soil profile data, and the actual self-settlement depth as the target variable; the pile parameters include the weight of the pile. w,length l ,diameter D Wall thickness t The data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. c′ Undrained shear strength S u internal friction angle f, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile;
[0018] S2. Perform depth interval feature encoding on the dataset. Discretize the continuous soil exploration profiles of varying lengths according to fixed depth intervals. Assign a complete set of physical and mechanical parameters to each discrete depth point, thereby transforming the soil layer sequence data of varying lengths into a fixed-length, structured feature vector. The complete set of physical and mechanical parameters includes soil type, state, and strength index.
[0019] S3. Divide the dataset into training and test sets. Train the machine learning model and optimize hyperparameters based on the encoded training set. Validate the performance of the machine learning model using an independent test set. If the performance meets the requirements, proceed to the next step; otherwise, adjust the machine learning model and re-validate.
[0020] S4. For the new single pile project to be predicted, collect its pile parameters and soil profile data obtained from the site survey, input them into the trained machine learning model, and output the predicted result of its self-sinking depth.
[0021] Preferably, in S1, the data samples undergo the following preprocessing:
[0022] For missing values, parameters that do not exist due to soil properties are filled with an identifier value that is outside the normal physical range.
[0023] Categorical variable encoding: For multiple categorical variables, one-hot encoding is used to convert them into numerical features;
[0024] For handling missing continuous values, the median of the same type of soil is used to fill in the occasional missing continuous parameters.
[0025] Feature standardization involves standardizing all numerical features so that their mean is 0 and their variance is 1.
[0026] Preferably, S2 includes: S201, determining the depth range and resolution: setting an analysis range that covers all possible self-submersion depths, discretizing this analysis range at fixed high-resolution intervals to generate a series of discrete depth points.
[0027] Preferably, S2 includes: S202, constructing the depth-feature matrix: for each single pile sample, traversing each discrete depth point. i :
[0028] Determine the depth point based on the original soil layer data of the pile. i The actual soil layer in which it is located;
[0029] Copy all soil parameters of the actual soil layer and assign them to this depth point. i; The soil parameters include coded soil type, soil state, and buoyancy unit. c′ Undrained shear strength S u internal friction angle f .
[0030] Preferably, S2 includes: S203, generating a fixed-length feature vector: concatenating the fixed parameters of the pile itself with the soil parameters of all discrete depth points to form a fixed-dimensional feature vector X; the soil parameters are arranged in depth order.
[0031] Preferably, S3 includes:
[0032] S301. Model Selection: A regression algorithm from supervised learning will be adopted.
[0033] S302. Dataset partitioning: Divide the dataset into a training set and an independent test set in an 8:2 ratio;
[0034] S303, Machine Learning Model Training and Hyperparameter Optimization: Based on the above training set, cross-validation is used to divide it into a training subset and a validation subset. The machine learning model is trained using the training subset, and the hyperparameters of the machine learning model are iteratively adjusted based on the performance feedback on the validation subset through grid search, random search, or Bayesian optimization methods to determine the optimal combination of hyperparameters.
[0035] S304. Model Evaluation: Evaluate the trained machine learning model using an independent test set, and quantify the prediction accuracy using root mean square error and mean absolute error.
[0036] A second objective of this invention is to provide a machine learning-based prediction system for the self-sinking depth of a single pile, comprising:
[0037] The dataset module collects historical engineering data and constructs a dataset. Each data sample includes: pile parameters, soil profile data, and the actual self-settlement depth as the target variable; the pile parameters include the pile weight. w ,length l ,diameter D Wall thickness tThe data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. c′ Undrained shear strength S u internal friction angle f, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile;
[0038] The encoding module performs depth interval feature encoding on the training dataset, discretizing continuous soil exploration profiles of varying lengths according to fixed depth intervals, and assigning a complete set of physical and mechanical parameters of the soil layer to each discrete depth point, thereby transforming the soil layer sequence data of varying lengths into a fixed-length, structured feature vector; the complete set of physical and mechanical parameters includes soil type, state, and strength index.
[0039] The machine learning model divides the dataset into training and test sets. The machine learning model is trained and hyperparameters are optimized based on the encoded training set. The performance of the machine learning model is verified using an independent test set. If the performance meets the requirements, the next step is taken; otherwise, the machine learning model is adjusted and verified again.
[0040] The prediction output module collects the pile parameters and soil profile data obtained from site surveys for the new single pile project to be predicted, inputs them into the trained machine learning model, and outputs the predicted result of its self-sinking depth.
[0041] Preferably, in the dataset module, the data samples undergo the following preprocessing:
[0042] For missing values, parameters that do not exist due to soil properties are filled with an identifier value that is outside the normal physical range.
[0043] Categorical variable encoding: For multiple categorical variables, one-hot encoding is used to convert them into numerical features;
[0044] For handling missing continuous values, the median of the same type of soil is used to fill in the occasional missing continuous parameters.
[0045] Feature standardization involves standardizing all numerical features so that their mean is 0 and their variance is 1.
[0046] A third objective of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned machine learning-based method for predicting the self-sinking depth of a single pile.
[0047] A fourth objective of this invention is to provide a computer program product, including a computer program that, when executed by a processor, implements the aforementioned machine learning-based method for predicting the self-sinking depth of a single pile.
[0048] The advantages and positive effects of this application are:
[0049] This invention significantly improves the accuracy of pile foundation self-settlement prediction, achieving high-precision prediction with sub-meter error, and greatly enhancing the reliability and predictive ability of construction control.
[0050] Specifically, by introducing a data-driven machine learning model, this invention significantly reduces the mean absolute error (MAE) of prediction from the meter level (e.g., 2.85m) of the traditional API method to the sub-meter level (e.g., 0.89m), improving accuracy by more than 68%, and demonstrating significant advantages in practical engineering applications.
[0051] This technological breakthrough primarily stems from the organic combination of the machine learning model and deep interval coding mechanism employed in this invention. Traditional API methods rely on empirical formulas and fixed coefficients, making it difficult to capture the complex nonlinear interactions and geological variability in the pile-soil system, thus limiting their generalization and accuracy. In contrast, the machine learning algorithms applied in this invention (such as XGBoost and Random Forest) possess powerful fitting and generalization capabilities, automatically learning and approximating the true physical mapping relationship during the pile's self-sinking process from historical construction data, thereby achieving accurate modeling of pile driving behavior under different geological conditions.
[0052] Experimental verification fully demonstrates the superior performance of this invention. On a test set constructed based on data from 52 actual engineering piles, the proposed optimized model (XGBoost) achieved a mean absolute error (MAE) of 0.89 m and a root mean square error (RMSE) of 1.09 m in completely independent samples, significantly outperforming the traditional API method (MAE=2.85 m, RMSE=3.68 m). This result not only reflects a substantial reduction in error indicators but also highlights the stability and practical value of this invention under complex geological conditions.
[0053] This invention can effectively capture and process the characteristics of complex, layered strata, significantly improving the accuracy and reliability of self-sinking depth prediction. Specifically, this invention is particularly suitable for heterogeneous strata conditions such as thin interbedded layers and soft-hard interlayers, and can more reliably predict the self-sinking depth during penetration behavior, overcoming the accuracy limitations of traditional methods due to the use of "layer averaging" processing.
[0054] This advantage is primarily due to the core technology employed in this invention—the depth interval feature encoding method. This method systematically discretizes the entire soil profile at high resolution (e.g., 0.5m intervals), constructing a continuous and detailed image of soil resistance distribution along depth, providing a comprehensive and detailed data foundation for the model. With this encoding method, the model can identify drastic changes in resistance at specific depths caused by abrupt changes in soil properties (such as a sudden transition from soft clay to dense sand), thereby more accurately determining the penetration termination point.
[0055] In contrast, traditional methods typically divide strata into several layers based on their thickness and calculate the average resistance value for each layer. This approach tends to smooth out key abrupt changes and obscure the true mechanical behavior at thin layers or interfaces, leading to biased predictions. The coding strategy of this invention, however, effectively preserves and highlights these subtle yet crucial features, enabling the model to identify, learn, and fully utilize high-frequency variations in the strata, ultimately achieving more accurate self-sinking depth predictions.
[0056] This invention possesses stronger model generalization ability and adaptability, effectively addressing the diverse environments and data conditions encountered in real-world engineering projects. Specifically, the prediction model trained by this invention can be widely applied to new engineering projects in different geographical regions and geological backgrounds, reducing the reliance on model recalibration due to site condition differences, thereby improving engineering efficiency and reducing resource consumption.
[0057] This advantage stems from the combined support of the inherent generalization properties of machine learning algorithms and robust feature preprocessing strategies. Tree ensemble models, such as XGBoost and Random Forest, exhibit strong resistance to overfitting and good ability to handle unseen data due to their ensemble learning mechanism and diversity principle. In this invention, cross-validation and hyperparameter optimization methods are used to further enhance the stability and generalization performance of this type of model.
[0058] In addition, regarding the common problem of missing values in geotechnical engineering data (such as the missing undrained shear strength in sandy soil), S u This invention designs a special identifier value processing strategy, using specific values (such as -999) to fill in missing attributes. This method not only avoids the biases that may be introduced by traditional imputation methods, but also enables the model to actively learn "when..." S u The decision rule that "when the value is invalid, the internal friction angle φ should be relied upon more for judgment" significantly improves the model's adaptability to incomplete data. This mechanism makes the model not only applicable to ideal data conditions, but also tolerant of data noise, missing data, and distribution shifts common in real-world projects, further enhancing its universality and practicality across different scenarios.
[0059] This invention represents a paradigm shift from "experience-driven" to "data-driven" approaches. This transformation significantly improves the accuracy and reliability of engineering predictions, enabling the uncovering of deep-seated patterns and overcoming the limitations of traditional methods. Specifically, this invention fully utilizes a large amount of accumulated engineering data and, through advanced machine learning algorithms, uncovers the complex, non-explicit interactions among multiple factors influencing self-settlement depth. These relationships include non-linear interactions between soil parameters, pile foundation characteristics, and environmental factors, which are difficult to express or often overlooked by traditional empirical formulas. This is precisely the essential advantage of machine learning methods: their ability to process high-dimensional data and automatically identify patterns.
[0060] During the machine learning model training process, the system can automatically evaluate the importance of hundreds of features and their combined effects, such as the changing trends of parameters like soil strength, cohesion, and internal friction angle at different depths. It can discover the influence of combination patterns of soil strength within specific depth ranges or the nonlinear coupling between pile diameter and friction angle of specific soil types on the final depth, thus providing a more refined prediction model. Furthermore, the model can optimize parameter weights through iterative learning, ensuring high generalization ability and adaptability in practical engineering applications, providing a scientific basis for self-settlement depth control.
[0061] This invention makes it possible to transform the "tacit knowledge" in these data into a reusable "explicit predictive model".
[0062] This invention is highly practical in engineering, easy to integrate and apply, and can significantly improve the efficiency and reliability of geotechnical engineering design. Specifically, this invention only requires inputting standard parameters already present in conventional geotechnical engineering investigation reports, such as soil type and undrained shear strength. S u internal friction angle f Effective severe c′ This invention achieves common indicators without relying on special or expensive field tests (such as detailed static cone penetration tests, CPT), significantly reducing the technical application threshold and project costs. The input feature design of this invention fully aligns with industry best practices, taking into full account data availability and operational convenience in actual engineering scenarios. Through a built-in feature encoding module, the system can automatically process various input parameters and convert them into standardized feature vectors, ensuring data consistency and model identifiability. This design enables the invention to seamlessly integrate with existing engineering workflows, is compatible with conventional design software and data processing platforms, greatly facilitating rapid application and promotion by engineers, and further promoting the implementation of intelligent methods in geotechnical engineering. Attached Figure Description
[0063] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0064] Figure 1 A flowchart of a preferred embodiment of the present invention is shown;
[0065] Figure 2 A diagram showing the pile body dimension parameters in a preferred embodiment of the present invention is provided.
[0066] Figure 3 A summary diagram of the self-sinking depth in a preferred embodiment of the present invention is shown;
[0067] Figure 4 A comparison chart of prediction results from machine learning and traditional API methods in a preferred embodiment of the present invention is shown. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Please see Figure 1 The first embodiment, a method for predicting the self-sinking depth of a single pile based on machine learning, mainly includes:
[0070] S1. Data Acquisition and Dataset Construction; specifically including: collecting historical engineering data and constructing a dataset, each data sample including: pile parameters, soil profile data, and actual self-settlement depth as the target variable; the pile parameters include the pile weight. w ,length l ,diameter D Wall thickness t The data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. c′ Undrained shear strength S u internal friction angle f, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile;
[0071] S2. Depth interval coding; specifically, it includes: performing depth interval feature coding on the dataset, discretizing continuous soil exploration profiles of varying lengths according to fixed depth intervals, and assigning a complete set of physical and mechanical parameters of the soil layer to each discrete depth point, thereby transforming the soil layer sequence data of varying lengths into a fixed-length, structured feature vector; the complete set of physical and mechanical parameters includes soil type, state, and strength index.
[0072] Interval feature coding is a soil profile feature coding method based on depth interval discretization. This method addresses the characteristics of soil exploration profiles, which typically exhibit strong continuity, varying lengths, and irregular structures. By setting fixed high-resolution depth intervals (e.g., one sampling point every 0.5m), the entire profile is systematically discretized. At each discrete depth point, the system assigns a complete set of physical and mechanical parameters to the soil layer, including soil type, physical state, strength indices, and other relevant engineering properties. Through this transformation process, the originally variable-length, non-uniformly structured soil layer sequence data is converted into fixed-length, highly structured feature vectors, thus exhibiting good machine readability and model input adaptability.
[0073] This encoding strategy effectively solves the problem that traditional geological survey data, due to inconsistent sample lengths and complex structures, is difficult to directly apply to traditional machine learning models (such as support vector machines, random forests, or neural networks). It not only achieves a standardized representation of soil profile data, but also provides machine learning models with high-resolution, depth-continuously distributed feature images, which can be figuratively understood as a "depth-distribution image of soil resistance," significantly improving the model's ability to identify and predict soil mechanical behavior.
[0074] S3. Machine learning model training and validation, specifically including: dividing the dataset into training set and test set, the ratio of which depends on the needs, such as 8:2; training the machine learning model and optimizing hyperparameters based on the encoded training set; and validating the performance of the machine learning model using an independent test set. If the performance meets the requirements, proceed to the next step; otherwise, adjust the machine learning model and revalidate it.
[0075] S4. New site application, specifically including: for new single pile projects to be predicted, collecting pile parameters and soil profile data obtained from site surveys, inputting them into a trained machine learning model, and outputting the predicted results of its self-sinking depth.
[0076] The core of this invention lies in constructing a data-driven prediction method that integrates deep interval feature encoding and machine learning models. To better understand the technical solution of this invention, the various components, implementation steps, and principles of the technical solution will be described in detail below.
[0077] The method of this invention mainly includes four steps. Its core idea is to transform the pile-soil system in the project into a structured machine learning problem and learn the complex mapping relationship between self-sinking depth and multi-dimensional features through a data-driven approach.
[0078] The detailed technical solution and steps are as follows:
[0079] Step 1: Collect historical engineering data and construct a dataset; this mainly includes data preparation and preprocessing.
[0080] Data preparation includes:
[0081] Data Acquisition: Collect historical or on-site data on single-pile installation projects to form a training dataset. Each data sample should include:
[0082] Pile parameters: weight of the pile w ,length l ,diameter D Wall thickness t And the type of hammer used (categorical variable).
[0083] Soil profile data: Information on continuous soil layers from top to bottom, starting from the seabed surface (mud surface). Each soil layer should include: depth of the bottom layer, soil type (e.g., clay, sand, etc.), and buoyant unit weight. c′ Undrained shear strength S u (For cohesive soils) Angle of internal friction f (For sandy soil) and descriptions of soil condition (such as loose, plastic, dense, etc.).
[0084] The actual self-sinking depth as the target variable: the actual observed self-sinking depth of the pile.
[0085] Preprocessing includes data preprocessing and quality processing; data preprocessing includes:
[0086] Handling missing values: For parameters that do not exist due to soil properties (e.g., sandy soil has no undrained shear strength). S u There is no internal friction angle in clay. f The missing data is filled with a special, out-of-normal physical value (such as -999) to inform the model that this information is missing and has physical significance, rather than using conventional mean or median imputation. This allows algorithms such as tree models to recognize this "physical missing data" and learn it as valid classification information.
[0087] Categorical variable encoding: For categorical variables such as "soil type", "soil state", and "hammer type", one-hot encoding is used to convert them into numerical features to avoid introducing incorrect order relationships.
[0088] Handling missing continuous values: For individual continuous parameters (such as buoyancy), c′ For occasional missing data, the median of the same type of soil was used for filling.
[0089] Quality processing includes feature standardization: all numerical features (including pile parameters and soil parameters) are standardized to have a mean of 0 and a variance of 1 to eliminate the influence of dimensions, accelerate model convergence, and ensure that each feature is treated equally during training.
[0090] Step 2: Perform deep interval feature encoding on the training dataset; this mainly includes feature engineering and deep interval encoding.
[0091] To address the issue of varying soil profile lengths that prevent direct input into standard machine learning models, this embodiment employs a depth-range feature encoding method:
[0092] First, determine the depth range and resolution: Set an analysis range covering all possible self-sinking depths, for example, from 0m to 45m from the seabed. Discretize this range at fixed high-resolution intervals (e.g., 0.5m), generating a series of discrete depth points (e.g., 0m, 0.5m, 1.0m, ..., 45m). It should be noted that other values can also be used for the depth discretization interval, such as 0.2m, 1m, etc.
[0093] Then, construct the depth-feature matrix: for each single-pillar sample, iterate through each of the above discrete depth points. i :
[0094] Determine the depth point based on the original soil layer data of the pile. i The actual soil layer in which it is located.
[0095] All soil parameters of the actual soil layer (including coded soil type, soil state, and buoyancy unit) c′ Undrained shear strength S u internal friction angle f ,in S u or f (Possibly an identifier value of -999) Copy and assign to this depth point. iIn this way, each depth point carries complete characteristic information of the soil layer in which it is located. It should be noted that when assigning soil layer parameters to depth points, more complex interpolation methods (such as linear interpolation) can be used to handle the changes in parameters with depth within the same soil layer, rather than simply copying them.
[0096] Finally, a fixed-length feature vector is generated: the fixed parameters of the pile itself (such as...) w, l, D, t The soil parameters (including the encoded hammer type) are concatenated with the soil parameters of all discrete depth points (arranged in depth order) to form a unified, fixed-dimensional feature vector X. This can be formally expressed as:
[0097] X=[x_pile,x_soil_0,x_soil_0.5,...,x_soil_45]
[0098] Where x_pile represents the pile parameter vector, and x_soil_i represents the depth. i Soil parameter vector at a distance of meters.
[0099] The role and principle of encoding: This method transforms a variable-length sequence of soil layers into a fixed-length, high-resolution "soil profile image". It enables machine learning models to perceive and analyze soil resistance at every depth along the potential penetration path with a consistent structure, thereby identifying key soil layers or interfaces that lead to penetration termination.
[0100] Step 3: Divide the dataset into training and test sets. Train the machine learning model and optimize hyperparameters based on the encoded training set, and verify the performance of the machine learning model using an independent test set. If the performance meets the requirements, proceed to the next step; otherwise, adjust the machine learning model and re-verify. This mainly includes machine learning model construction and training, and hyperparameter optimization.
[0101] Model Selection: Regression algorithms from supervised learning are used for modeling. In practical applications, preferred algorithms mainly include tree ensemble models, such as Random Forest, Gradient Boosting Decision Tree (GBDT), and efficient implementations such as XGBoost, Lightweight Gradient Boosting Machine, Gaussian Process Regression, or shallow neural networks. In addition, Support Vector Machine (SVM) can also be considered as an alternative machine learning algorithm, especially suitable for small sample sizes and high-dimensional scenarios. These algorithms not only have strong robustness and generalization ability but also effectively capture the intricate nonlinear relationships and interaction effects between high-dimensional features. For example, tree models handle feature combinations through splitting rules, while SVM uses kernel functions to map to a high-dimensional space to achieve nonlinear fitting; both demonstrate good predictive performance in practical applications.
[0102] This invention can construct model ensemble strategies, such as weighted averaging of prediction results from XGBoost, random forest, and support vector machine, to further improve prediction stability and accuracy.
[0103] Machine learning model training:
[0104] The dataset (feature vector X and target value) obtained after preprocessing and encoding is divided into a training set and a validation set, with a ratio of 8:2 between the training set and the validation set.
[0105] The selected machine learning model is trained using the training set data, and the internal parameters of the machine learning model are adjusted so that it learns a mapping function from feature X to target value.
[0106] Hyperparameter optimization: To achieve optimal model performance, hyperparameter tuning is necessary. Methods such as grid search, random search, or Bayesian optimization, combined with cross-validation, are used to find the optimal combination of hyperparameters. These settings help prevent overfitting while maintaining accuracy.
[0107] Model evaluation: The trained machine learning model is evaluated using a validation set, and the prediction accuracy is quantified using metrics such as root mean square error (RMSE) and mean absolute error (MAE).
[0108] Step 4: Collect pile parameters and soil profile data obtained from site surveys for the new single-pile project to be predicted, and use a machine learning model to determine the self-settlement depth; this mainly includes:
[0109] For new monopile projects to be predicted:
[0110] Collect the pile parameters and soil profile data obtained from the site survey.
[0111] Perform data preprocessing and depth interval encoding following the exact same process as steps one and two to generate the feature vector X_new corresponding to the stake.
[0112] Input X_new into the trained and saved optimal machine learning model.
[0113] The model output is the predicted self-sinking depth of the single pile.
[0114] Please see Figure 2 to Figure 4 Example: Prediction of the self-sinking depth of a single pile based on data from two offshore wind farms;
[0115] S11. Data Sources and Preprocessing
[0116] Data Collection: Fifty-two large-diameter monopiles from two existing offshore wind farms in China (designated Site A and Site B) were selected as the research subjects. Site A contains 28 monopiles, and Site B contains 24 monopiles. The following data were collected for each monopil:
[0117] Pile parameters: weight w ,length l ,diameter D Wall thickness t The model of the installation hammer (MHU 3500S for site A, IHC S3000 for site B) is specified. The pile parameter distribution results for both sites are as follows: Figure 2 As shown.
[0118] Soil profile: derived from engineering geological survey report, including bottom depth of each soil layer, soil type, and buoyancy unit weight. c′ Undrained shear strength S u internal friction angle f Soil condition description. Soil profiles were selected under one pile at each site A and site B. The soil parameters are shown in Tables 1 and 2 below:
[0119] Table 1 shows the soil parameters for site A.
[0120]
[0121] Table 2 shows the soil parameters for site B.
[0122]
[0123] Target value: The actual self-sinking depth measured on-site. The summarized self-sinking depths for both sites are as follows: Figure 3 As shown;
[0124] Data preprocessing:
[0125] Special identifier for missing values: For values missing in the sandy soil layer S u Values, and missing values in the clay layer f The value should be uniformly filled with the special value "-999".
[0126] Categorical variable coding: Hot coding is used to perform unique hot coding for "soil type" (e.g., clay, sand), "soil state" (e.g., plastic, dense) and "hammer type".
[0127] Continuous value imputation: filling in missing values for buoyancy. c′ The values are grouped according to the same "soil type" and filled with the median of that group.
[0128] Feature standardization: Standard deviation standardization is used for all numerical features (such as...) w, l , D , c′ , S u , f Standardize the expression so that its mean is 0 and its standard deviation is 1.
[0129] S22, Depth Range Coding
[0130] Determine the encoding parameters: Based on the maximum measured self-sinking depth in the data (31.9m) and with a margin, determine the analysis depth range as 0m (mud surface) to 45m. Select 0.5m as the depth discretization interval, and generate a total of 91 discrete depth points (0, 0.5, 1.0, ..., 45.0).
[0131] Execution coding: For the first j For each pile, the original soil layer data is [(bottom depth of soil layer 1, soil parameter 1), (bottom depth of soil layer 2, soil parameter 2), ...]. Write a program to iterate through each discrete depth point d_i (i=0 to 90):
[0132] Determine which original soil layer d_i is located in.
[0133] All soil parameters (pre-processed) of this soil layer are used as the feature x_soil_i at depth d_i.
[0134] Generating the feature vector: The standardized pile body parameter x_pile (dimension 5) is concatenated with all 91 x_soil_i (each dimension being the encoded soil parameter dimension) in depth order to form the final feature vector X_j. In this example, the final feature vector dimension is 5 + 91 = 96 (encoded soil parameter dimension).
[0135] S33, Machine Learning Model Training and Optimization
[0136] Dataset splitting: The 52 samples were randomly shuffled and divided into a training set (41 stakes) and an independent test set (11 stakes) in a ratio of 80:20.
[0137] Model selection and hyperparameter tuning: On the training set, 5-fold cross-validation combined with grid search was used to optimize the hyperparameters of the following four models:
[0138] Random Forest: Optimize n_estimators, max_depth, min_samples_split, and min_samples_leaf.
[0139] Gradient boosting regression tree: Optimize n_estimators, max_depth, learning_rate, and min_samples_leaf.
[0140] XGBoost: Optimize n_estimators, max_depth, learning_rate, subsample, colsample_bytree.
[0141] Support Vector Machine: Optimize kernel function (RBF), penalty parameter C, and kernel parameter gamma.
[0142] Optimal parameters: After optimization, the optimal set of hyperparameters for the XGBoost model is: n_estimators=50, max_depth=4, learning_rate=0.05, subsample=0.8, colsample_bytree=0.8.
[0143] Machine learning model training: Retrain the four final models using the optimal hyperparameters and the full training set data.
[0144] S44. Prediction and Performance Evaluation (Comparative Analysis)
[0145] Comparison method: Select the API recommendation method widely used in the industry as the comparison example.
[0146] Evaluation metrics: The root mean square error and mean absolute error were used to evaluate the performance of all models on the same test set (11 piles).
[0147] Implement the API method: According to the API RP 2GEO specification, using the exact same pile parameters and layered soil data as the machine learning model, manually calculate the predicted self-settlement depth of each test pile.
[0148] Experimental results: The prediction performance comparison on the test set is as follows. Figure 4 As shown in Table 3.
[0149] Table 3 compares the prediction results of machine learning and traditional API methods.
[0150]
[0151] Effect Analysis:
[0152] All machine learning models showed significantly higher prediction accuracy than traditional API methods.
[0153] The preferred model of this invention, XGBoost, achieves the best performance with a MAE of 0.89m, which is about 68.8% lower than the MAE (2.85m) of the API method, realizing sub-meter level high-precision prediction.
[0154] Figure 4(Comparison of prediction results) This visually demonstrates that the XGBoost model's predicted values are in high agreement with the measured values, while the API method's predicted points are discretely distributed on both sides of the measured values, resulting in a significantly larger error.
[0155] This embodiment verifies the effectiveness of the technical solution of the present invention. By structuring complex geological information through "depth interval coding" and using machine learning algorithms such as XGBoost for learning, the self-sinking depth of a single pile can be predicted with extremely high accuracy, providing a reliable new method for solving key uncertainties in marine engineering construction.
[0156] A machine learning-based system for predicting the self-settlement depth of a single pile, used to implement the machine learning-based method for predicting the self-settlement depth of a single pile in the above embodiments, the system comprising:
[0157] The dataset module collects historical engineering data and constructs a dataset. Each data sample includes: pile parameters, soil profile data, and the actual self-settlement depth as the target variable; the pile parameters include the pile weight. w ,length l ,diameter D Wall thickness t The data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. c′ Undrained shear strength S u internal friction angle f, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile;
[0158] The encoding module performs depth interval feature encoding on the training dataset, discretizing continuous soil exploration profiles of varying lengths according to fixed high-resolution depth intervals, and assigning a complete set of physical and mechanical parameters of the soil layer to each discrete depth point, thereby transforming the soil layer sequence data of varying lengths into a fixed-length, structured feature vector; the complete set of physical and mechanical parameters includes soil type, state, and strength index.
[0159] The machine learning model divides the dataset into training and test sets. The machine learning model is trained and hyperparameters are optimized based on the encoded training set. The performance of the machine learning model is verified using an independent test set. If the performance meets the requirements, the next step is taken; otherwise, the machine learning model is adjusted and verified again.
[0160] The prediction output module collects the pile parameters and soil profile data obtained from site surveys for the new single pile project to be predicted, inputs them into the trained machine learning model, and outputs the predicted result of its self-sinking depth.
[0161] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned machine learning-based method for predicting the self-sinking depth of a single pile.
[0162] A computer program product includes a computer program that, when executed by a processor, implements the aforementioned machine learning-based method for predicting the self-sinking depth of a single pile.
[0163] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented, in whole or in part, as a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line, or wireless (e.g., infrared, wireless, microwave, etc.) means). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0164] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for predicting the self-sinking depth of a single pile based on machine learning, characterized in that, include: S1. Collect historical engineering data and construct a dataset. Each data sample includes: pile parameters, soil profile data, and the actual self-settlement depth as the target variable; the pile parameters include the weight of the pile. w ,length l ,diameter D Wall thickness t The data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. γ′ Undrained shear strength S u internal friction angle φ, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile; S2. Perform depth interval feature encoding on the dataset. Discretize the continuous soil exploration profiles of varying lengths according to fixed depth intervals, and assign a complete set of physical and mechanical parameters of the soil layer to each discrete depth point. This transforms the soil layer sequence data of varying lengths into a fixed-length, structured feature vector. The complete set of physical and mechanical parameters includes soil type, state, and strength index. S2 includes: S201. Determine the depth range and resolution: Set an analysis range that covers all possible self-submersion depths, and discretize this analysis range at fixed high-resolution intervals to generate a series of discrete depth points. S202. Constructing the Depth-Feature Matrix: For each single-pillar sample, traverse each discrete depth point. i : Determine the depth point based on the original soil layer data of the pile. i The actual soil layer in which it is located; Copy all soil parameters of the actual soil layer and assign them to this depth point. i; The soil parameters include coded soil type, soil state, and buoyancy unit. γ′ Undrained shear strength S u internal friction angle φ ; S203. Generate a fixed-length feature vector: Concatenate the fixed parameters of the pile itself with the soil parameters of all discrete depth points to form a fixed-dimensional feature vector X; the soil parameters are arranged in depth order. S3. Divide the dataset into training and test sets. Train the machine learning model and optimize hyperparameters based on the encoded training set. Validate the performance of the machine learning model using an independent test set. If the performance meets the requirements, proceed to the next step; otherwise, adjust the machine learning model and re-validate. S4. For the new single pile project to be predicted, collect its pile parameters and soil profile data obtained from the site survey, input them into the trained machine learning model, and output the predicted result of its self-sinking depth.
2. The method for predicting the self-sinking depth of a single pile based on machine learning according to claim 1, characterized in that, In S1, the data samples are preprocessed as follows: For missing values, parameters that do not exist due to soil properties are filled with an identifier value that is outside the normal physical range. Categorical variable encoding: For multiple categorical variables, one-hot encoding is used to convert them into numerical features; For handling missing continuous values, the median of the same type of soil is used to fill in the occasional missing continuous parameters. Feature standardization involves standardizing all numerical features so that their mean is 0 and their variance is 1.
3. The method for predicting the self-sinking depth of a single pile based on machine learning according to claim 2, characterized in that, S3 include: S301. Model Selection: A regression algorithm from supervised learning will be adopted. S302. Dataset partitioning: Divide the dataset into a training set and an independent test set in an 8:2 ratio; S303, Machine Learning Model Training and Hyperparameter Optimization: Based on the above training set, cross-validation is used to divide it into a training subset and a validation subset. The machine learning model is trained using the training subset, and the hyperparameters of the machine learning model are iteratively adjusted based on the performance feedback on the validation subset through grid search, random search, or Bayesian optimization methods to determine the optimal combination of hyperparameters. S304. Model Evaluation: Evaluate the trained machine learning model using an independent test set, and quantify the prediction accuracy using root mean square error and mean absolute error.
4. A machine learning-based prediction system for the self-sinking depth of a single pile, characterized in that, include: The dataset module collects historical engineering data and constructs a dataset. Each data sample includes: pile parameters, soil profile data, and the actual self-settlement depth as the target variable; the pile parameters include the pile weight. w ,length l ,diameter D Wall thickness t The data includes the type of installation hammer; the soil profile data includes: continuous soil layer information from top to bottom starting from the seabed; each soil layer information must include: bottom depth of the soil layer, soil type, and buoyancy unit weight. γ′ Undrained shear strength S u internal friction angle φ, And a description of the soil condition; the target variable includes: the actual observed self-settlement depth of the pile; The encoding module performs depth interval feature encoding based on the dataset. It discretizes continuous soil exploration profiles of varying lengths at fixed depth intervals, assigning each discrete depth point a complete set of physical and mechanical parameters for its corresponding soil layer. This transforms the varying length soil layer sequence data into a fixed-length, structured feature vector. The complete set of physical and mechanical parameters includes soil type, state, and strength indices. The encoding module includes: Determine the depth range and resolution: Set an analysis range that covers all possible self-submersion depths, and discretize this analysis range at fixed high-resolution intervals to generate a series of discrete depth points; Constructing the depth-feature matrix: For each single pile sample, traverse each discrete depth point. i : Determine the depth point based on the original soil layer data of the pile. i The actual soil layer in which it is located; Copy all soil parameters of the actual soil layer and assign them to this depth point. i; The soil parameters include coded soil type, soil state, and buoyancy unit. γ′ Undrained shear strength S u internal friction angle φ ; Generate a fixed-length feature vector: Concatenate the fixed parameters of the pile itself with the soil parameters of all discrete depth points to form a fixed-dimensional feature vector X; the soil parameters are arranged in depth order. The machine learning model divides the dataset into training and test sets. The machine learning model is trained and hyperparameters are optimized based on the encoded training set. The performance of the machine learning model is verified using an independent test set. If the performance meets the requirements, the next step is taken; otherwise, the machine learning model is adjusted and verified again. The prediction output module collects the pile parameters and soil profile data obtained from site surveys for the new single pile project to be predicted, inputs them into the trained machine learning model, and outputs the predicted result of its self-sinking depth.
5. The machine learning-based prediction system for the self-sinking depth of a single pile according to claim 4, characterized in that, In the dataset module, the data samples are preprocessed as follows: For missing values, parameters that do not exist due to soil properties are filled with an identifier value that is outside the normal physical range. Categorical variable encoding: For multiple categorical variables, one-hot encoding is used to convert them into numerical features; For handling missing continuous values, the median of the same type of soil is used to fill in the occasional missing continuous parameters. Feature standardization involves standardizing all numerical features so that their mean is 0 and their variance is 1.
6. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the machine learning-based method for predicting the self-sinking depth of a single pile as described in any one of claims 1-3.
7. A computer-readable storage medium storing a computer program, characterized in that, When executed by the processor, the program implements the machine learning-based method for predicting the self-sinking depth of a single pile as described in any one of claims 1-3.
Citation Information
Patent Citations
Method, device and equipment for monitoring pile sinking of pile foundation and storage medium
CN119672101A
Single pile bearing capacity prediction method based on XGBoost machine learning algorithm
CN120930455A