Method of training a machine learning model, machine learning model trained by same method and method of predicting geotechnical parameters from geophysical measurement data

A machine learning model combining geotechnical and geophysical data with weighted sampling and geo-statistical methods addresses spatial sampling limitations, enabling accurate 3D prediction of geotechnical parameters and reducing ground risk in infrastructure projects.

WO2025242579A1PCT designated stage Publication Date: 2025-11-27FNV IP BV

Patent Information

Application Number
PCT/EP2025/063639
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-21
Filing Date
2025-05-19
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing geotechnical surveys face challenges in accurately predicting geotechnical parameters due to limited spatial sampling and budget constraints, leading to uncertainty and ground risk in infrastructure projects.

Method used

A machine learning model is trained using a weighted training matrix that combines geotechnical and geophysical data, with geophysical data further away from geotechnical data having smaller weights, and incorporates geo-statistical methods like ordinary kriging interpolation and vertical anisotropy to enhance prediction accuracy.

Benefits of technology

The method allows for accurate prediction of geotechnical parameters in three dimensions, reducing uncertainty and improving ground risk management by integrating spatial continuity and autocorrelation, thus enhancing infrastructure design and engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063639_27112025_PF_FP_ABST
    Figure EP2025063639_27112025_PF_FP_ABST
Patent Text Reader

Abstract

A method of training a machine learning model representing association between geotechnical parameters and geophysical measurement data is disclosed. The method comprises the steps of: obtaining geotechnical training data for a first area; obtaining geophysical training data for a second area, wherein the second area is larger than and surrounds the first area; assigning sample weights to a training matrix depending on distance between the geophysical and geotechnical training data such that geophysical training data further away from the geotechnical training data has smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data; and training the machine learning model using the weighted training matrix.Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD OF TRAINING A MACHINE LEARNING MODEL, MACHINE LEARNING MODEL TRAINED BY SAME METHOD AND METHOD OF PREDICTING GEOTECHNICAL PARAMETERS FROM GEOPHYSICAL MEASUREMENT DATAFIELD OF THE INVENTION

[0001] The present disclosure generally relates to foundation engineering, and more specifically to a method of predicting geotechnical parameters from geophysical measurement data using artificial intelligence. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.BACKGROUND OF THE INVENTION

[0002] Geotechnical properties of the ground play a major role in infrastructure projects and direct methods form the bulk of most site investigation / site characterisation programmes. Geotechnical surveys typically involve direct measurements and sampling of soil and rock properties at specific locations. The well-known Cone Penetration Testing, CPT, investigations are used in soils to derive physical measurements of cone resistance (t / c) and sleeve friction (4). These measurements are then converted through semi-empirical means to estimate geotechnical stiffness and strength parameters used in infrastructure foundation design. This "ground truth" data is essential for understanding the properties of the materials directly beneath the survey locations.

[0003] CPT measurements are locally accurate but cannot evaluate ground properties in 3D due to limited spatial sampling in one dimensional (depth) and limited budget and schedule. This means gaps in the model space that result in uncertainty that translates into ground risk. Ground risk, if left improperly managed, can lead to unwanted engineering business outcomes impacting cost and schedule.

[0004] It follows that by reducing uncertainty, that is, filling the otherwise empty model space with information, ground risk is better managed. Geophysics can be used to fill the gaps between direct in situ geotechnical investigation by measuring geophysical signals in two dimensions or sometimes three dimensions.

[0005] Geophysical surveys utilize non-invasive techniques such as seismic imaging, electrical resistivity tomography, ground-penetrating radar, and others, to gather data about subsurface structures and properties over a larger area. These methods can detect variations in subsurface properties without the need for physical excavation or drilling. Geophysical surveys provide a broader perspective of the subsurface and can detect features and anomalies that may not be apparent from geotechnical data alone.

[0006] Therefore, the possible combination of geotechnical and geophysical measurements is of an interest, as it combines the local ground truth from a geotechnical survey with global non-invasive data from a geophysical survey to derive a 3D model of ground properties.

[0007] The combination of geotechnical and geophysical measurements involves utilizing correlations or empirical relationships between geophysical measurements and geotechnical properties. This generally involves analysing the relationship between geophysical data and geotechnical parameters, which can make use of statistical analysis to identify correlations, trends, or patterns in the data. For example, there may be relationships between electrical resistivity values from ERT data and soil properties that influence the cone resistance and the sleeve friction. Thereafter predictive models can be developed based on the identified correlations.

[0008] In practice, the link between geophysical signal and ground properties is sometimes difficult to establish. Moreover, different geophysical methods sensitive to different ground parameters can be applied on the same study site, which helps to bring complementary information that may be necessary to interpret as a unique information matrix to refine the ground model. Handle several sources of information in the objective of a better interpretation is most of the time complex and require lot of processing time. Additionally, emergence of new acquisition system in earth science domain led to deal with much more information that need more processing capacity to treat.

[0009] In consideration of the above, there is a need for an improved method for establishing association between geotechnical parameters and geophysical measurement data, thereby allowing accurate prediction of geotechnical parameters from geophysical measurement data.BRIEF SUMMARY OF THE INVENTION

[0010] In one aspect of the invention there is provided a method of training a machine learning model representing an association between geotechnical parameters and geophysical measurement data, comprising the steps of:

[0011] - obtaining geotechnical training data for a first area;

[0012] - obtaining geophysical training data for a second area, wherein the second area is larger than and surrounds the first area;

[0013] assigning sample weights to a training matrix depending on distance between the geophysical and geotechnical training data such that geophysical training data further away from the geotechnical training data has smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data, wherein the training matrix comprises the geotechnical training data and the geophysical training data; and

[0014] - training the machine learning model using the weighted training matrix.

[0015] The present disclosure is based on the inventor’s insight that data science approach can be used effectively to derive link between geotechnical data and geophysical data, considering the large amount of data involved in the combination of geotechnical and geophysical measurements. Specifically, the machine learning approach represents a mean to achieve an objective of integrating multiple input data, including measured geophysical data and geotechnical data, into a unique processing phase, which is associated with the capacity to integrate large amount of data.

[0016] Based on the above insight, the present disclosure proposes a method of training a machine learning model representing association between geotechnical parameters and geophysical measurement data based on the obtained geotechnical and geophysical training data. As can be understood by those skilled in the art, the geophysical training data is obtained over a larger area than for the geotechnical training data.

[0017] Considering the principles of spatial continuity and spatial autocorrelation in geosciences, sample weights are assigned to the training matrix comprising the geotechnical training data and the geophysical training data. By using the sample weights, the geophysical training data further away from the geotechnical training data has smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data.

[0018] This is in line with the fact that values measured at nearby locations are more likely to be similar to each other than values measured at distant locations. The machine learningmodel thus trained can therefore represents the association between geotechnical parameters and geophysical measurement data in a more accurate way.

[0019] In an example of the present disclosure, the geotechnical training data comprises 3D geotechnical data derived from geotechnical measurement data using a geo-statistical, GS, approach, such as the ordinary kriging interpolation.

[0020] In practice, land site characterization may suffer from lack of data to properly and efficiently trained a ML model. One way of tackling this problem is to increase the number of input information to train the model. Following this observation, the present disclosure proposes to use the geo-statistical, GS, approach to increase the number of training data.

[0021] The use of the GS approach, such as the ordinary kriging interpolation process, makes possible larger number of connections between geophysical and geotechnical data through interpolated CPT values. By “geo-statistically” increasing the number of input information, the ML model will be better trained to predict the geotechnical parameters from measured geophysical data. In particular, a general trend instead of precise expected value at a specific site location can be predicted.

[0022] In an example of the present disclosure, the geotechnical measurement data is cone penetrometer test, CPT, data.

[0023] As can be understood by those skilled in the art, CPT, as a well-established and known geotechnical method, can be used in the present method as the basis for the ML model to be trained.

[0024] In an example of the present disclosure, an exponential semi-variogram model is used in the ordinary kriging interpolation.

[0025] The semi-variogram model describes how the variability or correlation of a variable, such as soil properties, changes with distance or lag. An exponential semi-variogram model specifically describes this change as an exponential decay with distance. As the distance between sample points increases, the correlation between their values decreases exponentially.

[0026] In an example of the present disclosure, vertical anisotropy factor is used in the ordinary kriging interpolation.

[0027] Anisotropy refers to directional dependence or variability. In this context, vertical anisotropy factor implies that the variability of the geotechnical parameters, such as normalized cone resistance and friction ration, with depth is different from the variability in the lateral direction. By introducing a vertical anisotropy factor, the model accounts for the fact that the parameters change more rapidly or exhibit stronger variability with depth compared to their lateral variation.

[0028] By using an exponential semi-variogram model associated with a vertical anisotropy factor, the model accounts for the specific spatial characteristics of normalized cone resistance and friction ration parameters, emphasizing their variation with depth while considering the lateral variation as well. This approach helps in accurately capturing the spatial variability of these parameters in geotechnical applications.

[0029] In an example of the present disclosure, the training step comprises a series of traintest validation procedure.

[0030] In the training phase, the ML model is trained on a portion of the available dataset called the "training set." During training, the model learns the relationships and patterns in the geotechnical and geophysical training data by adjusting its parameters to minimize the prediction error or loss function.

[0031] After the model has been trained, it is evaluated on another portion of the dataset called the "test set" that was not used during training. The model's performance is assessed by making predictions of the geotechnical parameter based on the test set and comparing them to the actual values. This is used to estimate how well the model generalizes to new, unseen data.

[0032] In addition to the training and testing phases, a validation phase may also be included. In this phase, the model is further evaluated using a separate validation set. The validation set helps tune hyperparameters, such as regularization strength or learning rate, and assess the model's performance on data that it has not been directly trained on or tested against.

[0033] The train-test validation process helps assess model performance, prevent overfitting, and optimize hyperparameters.

[0034] In an embodiment of the present disclosure, the machine learning model comprises one or more algorithms selected from the group consisting of CatBoost algorithm, Random Forest algorithm, Gradient Boosting algorithm, AdaBoost algorithm, and XGBoost algorithm.

[0035] The listed models can be conveniently used in the present disclosure to represent the association between geotechnical parameters and geophysical measurement data.

[0036] In a second aspect of the present disclosure, there is provided a machine learning model trained according to the first aspect of the present disclosure.

[0037] In a third aspect of the present disclosure, there is provided a method of predicting geotechnical parameters from geophysical measurement data, the method comprising:

[0038] - predicting geotechnical parameters in an area using the machine learning model according to the second aspect of the present disclosure by using geophysical measurement data as input data.

[0039] The machine learning model trained according to the present disclosure can be used conveniently to predict geotechnical parameters from geophysical measurement data.

[0040] In an example of the present disclosure, the geotechnical parameters comprise CPT parameters, including cone resistance and sleeve friction.

[0041] Parameters such as soil behaviour type index may also be predicted using the training ML model. It is also possible to predict all wireline soil parameters recording such as gamma ray for example.

[0042] In an example of the present disclosure, the geophysical measurement data are obtained using one or more of Electrical Resistivity Tomography, Multichannel Analysis of Surface Waves (Vs), seismic refraction tomography (Vp), Ground-Penetrating Radar technology.

[0043] These methods are non-invasive and much less costly than the geotechnical methods, therefore these methods can be conveniently used to obtain the geophysical measurement data, which will then be used to predict the geotechnical parameters.

[0044] In an example of the present disclosure, the geophysical training data comprises geophysical measurement data, geophysical data derived values and spatial features.

[0045] Using geophysical data derived values, it is possible to capture more complex relationships and patterns in the data that may not be apparent in the original features alone. By transforming or combining existing features, derived features can provide additional information that enhances the model's ability to learn and make accurate predictions.

[0046] In an example of the present disclosure, the predicting step comprises predicting trend of the geotechnical parameters.

[0047] The method of the present disclosure will possibly be better trained to predict a general trend instead of precise expected value at a specific site location, this is in consideration that the training data includes data derived using the geo-statistical approach.

[0048] In an example of the present disclosure, the method further comprises building at least one of a 2D and 3D ground model of the area based on the predicted geotechnical parameters.

[0049] A detailed ground model provides valuable information for geotechnical design and engineering. Engineers can use the model to assess soil stability, analyse foundation behaviour, and design appropriate earthworks, retaining structures, and foundations.

[0050] In a fourth aspect of the present disclosure, there is provided a computer program product, comprising a computer readable storage medium storing instructions which, whenexecuted on at least one processor, cause the at least one processor to carry out the method according to the first aspect or according to the third aspect of the present disclosure.

[0051] The above mentioned and other features and advantages of the disclosure will be best understood from the following description referring to the attached drawings. In the drawings, like reference numerals denote identical parts or parts performing an identical or comparable function or operation.BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to describe the manner in which the above-recited and other advantages and features of the disclosure can be obtained, a more particular description of the principles briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the appended drawings. Understanding that these drawings depict only exemplary embodiments of the disclosure and are therefore not to be considered to be limiting of its scope, the principles herein are described and explained with additional specificity and detail through the use of the accompanying drawings in which:

[0053] FIG. 1 schematically illustrate, in a flow chart, a method of training a machine learning model representing association between geotechnical parameters and geophysical measurement data, sensor in accordance with the present disclosure;

[0054] FIG. 2 schematically illustrates an area under study and indication of relevant geophysical and geotechnical measurements;

[0055] FIG. 3 presents 2D ERT and MASW sections of line LI 10 of FIG. 2 (in the center of the study site);

[0056] FIG. 4 shows CPT results for Qtnand Frfor CPT number 235 of FIG. 2;

[0057] FIG. 5 schematically illustrates configuration cases N°1 andN°2 in accordance with the present disclosure;

[0058] FIGS. 6 and 7 respectively show an example of the weighting matrix in 2D and 3D with value from 0.252to 1;

[0059] FIG. 8 schematically illustrates variation of the weight as a function of distance from borehole / CPT;

[0060] FIG. 9 shows the comparison between measured and interpolated values CPT N°6;

[0061] FIG. 10 presents the 3D weighting matrix visualisation that is associated with the training matrix;

[0062] FIGS. 11 and 12 respectively show measured and predicted Qtnand Frvalues for CPT N°232 and CPT N°238 of the first sites configuration;

[0063] FIGS. 13 and 14 respectively show measured and predicted Qtnand Frvalues for CPT N°234 and CPT N°238 of the second sites configuration;

[0064] FIG. 15 schematically illustrates a Dynamic Moving Average, DMA, being applied to the CPT data; and

[0065] FIG. 16 presents the CPT result with DMA.DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS

[0066] Embodiments contemplated by the present disclosure will now be described in more detail with reference to the accompanying drawings. The disclosed subject matter should not be construed as limited to only the embodiments set forth herein. Rather, the illustrated embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.

[0067] The present disclosure describes an approach of using machine learning to couple localised geotechnical measurements with more global 3D measurements derived from geophysical methods. A machine learning model trained according to the method of the present disclosure can be used to predict geotechnical parameters in three dimensions, which allows a high fidelity 3D overview of the subsurface to be obtained.

[0068] The method of the present disclosure offers an effective means to capture all type of geophysical information associated with geotechnical data at a site scale to build a 2D or 3D ground model.

[0069] Figure 1 schematically illustrate, in a flow chart, a method 10 of training a machine learning model representing association between geotechnical parameters and geophysical measurement data, in accordance with the present disclosure.

[0070] At step 11, geotechnical training data for a first area is obtained. Geotechnical data typically include measurements from in-situ tests like Cone Penetration Test, CPT, Standard Penetration Test, SPT, or laboratory tests on soil samples. After the measurement data is collected, preprocessing steps can be performed to clean and prepare the geotechnical training data to be used in the training of the machine learning model. This may involve handling missing values, removing outliers, and so on.

[0071] At step 12, geophysical training data for a second area is obtained. As can be understood by those skilled in the art, the geophysical measurement can be performed for an area lager than where the geotechnical measurements are done. In this sense, the second area from where the geophysical training data are derived is larger than and surrounds the first area where the geotechnical data is obtained.

[0072] Geophysical measurements may include data collected using techniques such as Electrical Resistivity Tomography, ERT, Seismic Refraction, Ground Penetrating Radar, GPR, Multichannel Analysis of Surface Waves, Seismic reflection or any other relevant method. The collected geophysical measurement can be pre-processed to clean and prepare the geophysical training data to be used in training the machine learning model. Normalization or standardization is performed as range of seismic values are completely different from range of ERT values for example.

[0073] A training matrix comprising the geotechnical training data, the geophysical training data and geospatial data associated with the training data is generated, which will be used to train the machine learning model.

[0074] At step 13, sample weights are assigned to the training matrix depending on distance between the geophysical and geotechnical training data such that geophysical training data further away from the geotechnical training data has smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data.

[0075] The step will be described in more details in the following.

[0076] At step 14, the machine learning model is trained using the weighted training matrix. The trained machine model represents relationship between geotechnical parameters and geophysical measurement data, which can then be used to predict geotechnical parameters from geophysical measurements.

[0077] The method will be described in more detail in the following with reference to examples.

[0078] The machine learning model of the present disclosure will be discussed and tested based on data from a study site following two prediction configurations. One configuration will consist of making geotechnical ground parameter predictions on an area located inside the geophysical acquisition volume used for the ML training phase. Another case will consist of making predictions using geophysical data that have not been part of the ML training phase.

[0079] For both configurations, eight CPTs were performed to locally capture geomechanical properties of the site under study. CPT measurements consisted of cone resistance, sleeve friction and pore pressure with a vertical resolution of two centimetres. Two CPTmeasurements were excluded from the training phase to test prediction quality of the result. These two configurations allow us to evaluate prediction quality in the context of “filling the gap” between CPT acquisition and “complete ground properties mapping” outside the investigated CPT area.

[0080] Figure 2 schematically illustrates an area 20 under study and indication of relevant geophysical and geotechnical measurements. It will be understood by those skilled in the art the area as shown in Figure 2 is for illustrative purposes only and therefor does not correspond strictly to real-line scenario.

[0081] The area 20 is about two hundred and twenty -five meters. The ground investigation campaign consists of a series of forty-five parallel geophysical lines and spaced every five meters. Geophysical acquisitions were performed to collect both shear wave velocity and electrical resistivity values down to twenty and thirty meters respectively. Thicker lines 230 in Figure 2 represent MASW measurement lines and thinner lines L0 to L220 represent ERT measurement lines.

[0082] In addition to the geophysical investigations, eight CPTs indicated by block dots 231-238 are performed to locally capture geo-mechanical properties of the ground in the area 20. The measured CPT parameters comprise cone resistance, sleeve friction and pore pressure with a vertical resolution of two centimetres.

[0083] Figure 3 presents 2D ERT and MASW sections of line LI 10 (in the center of the study site). The position of the top of the bedrock (black dashed line) seems to be clearly identifiable on both resistivity and Vs results (high resistivity and high Vs values) at similar depths.

[0084] In general, 2D ERT sections show resistivity variations in the range 5 to 300 Qm with higher resistivity value at the ground surface and lower resistivity value at depth (see the bottom profile in Figure 3) to the exception of the bedrock with high resistivity value. However, resistivity variations over the site are not strictly ID (only vertical variation) and significant resistivity variations occur along each line from West to East (left to right of Figure 3) and from South to North (bottom to top of Figure 3) indicating a more complex 3D (3 Dimensional) ground resistivity model.

[0085] Regarding MASW (see the top profile in Figure 3), Vs values are in the range 250 to 1000 m.s'1with lower values at the top and higher values at depth. The vertical Vs variation is quite linear with Vs increasing with depth, apart between +3 and -3m where Vs increases and decreases inducing a velocity inversion.

[0086] A total of eight CPTs 231-238 were performed, centered on the MASW lines and covering the site from South to North. The CPT measurements consisted of cone resistance qc(in MPa), sleeve friction fs(in MPa) and pore pressure u (in kPa) acquisition. Because qcand fsmeasurements are strongly affected by the effective overburden pressure, raw data were normalized from total and effective in situ vertical stress to calculate the normalized cone penetration resistance Qtn(dimensionless) and friction ratio Fr(in %).

[0087] As an example, Figure 4 shows CPT results for Qtnand Frfor CPT number 235. The dashed lines 51 and 52 represent the general trend i.e., Qtnand Frvariation with depth. This linear trend helps to identify vertical Qtnand Frvariation that deviates from the expected increase in values with depth for a homogenous subsurface. When strong lateral variations are present, this should show as strong variation in these trends across the site.

[0088] Machine learning will be used to test the capacity to predict geotechnical parameters (Qtn and F) from geophysical data. For this purpose, a machine learning model representing association between geotechnical parameters and geophysical measurement data will be trained based on measurement data such as the ones described above.

[0089] In the present disclosure, training dataset will be based on the geophysical and geotechnical measurements. The training dataset can be constituted of standardised data, that is, a method of rescaling is applied such that values are centered around the mean with a unit standard deviation. The training dataset is divided into five equal parts of which four parts will be used to train the model and the rest to test it.

[0090] As an example, the present disclosure considers three types of features: geophysical features such as processed geophysical data, such as Vs and resistivity described above, derived features (such as the logarithm of the resistivity and the vertical gradient of Vs) and spatial features such as XYZ coordinates of geophysical measurement data.

[0091] Following the “train-test validation” procedure explained above, CPT data excluded from the training matrix will be used to check the ability of the ML model to predict Qtnand Fr. Quality of prediction will make possible conclusion on processing procedure and possible improvement discussion.

[0092] As explained above, two different prediction configurations are tested during the ML study, which are respectively referred to case N°1 and case N°2. For configuration case N°l, CPT N° 232 and N° 238 are excluded from the ML training phase while for configuration case N°2, CPT N° 234 and N° 238 have been excluded.

[0093] Figure 5 schematically illustrate configuration cases N°1 and N°2. In Figure 5(a), the excluded CPT N° 232 and N° 238 are marked by white dot, which means the training matrix is based on all geophysical lines and all CPTs to the except of CPT N°2 and N°8.

[0094] In Figure 5(b), the excluded CPT N° 234 and N° 238 are marked by white dot. That is, the training matrix is based on geophysical lines from L0 to the one next to CPT N° 237, which is L140, and all CPTs except of CPT N° 234 and N° 238.

[0095] Case N°1 is processed as if CPT N° 232 and N° 238 were not drilled during field acquisition. In opposition to case N°2 these two CPTs are located inside the investigated ground volume were geophysical and geotechnical data were collected.

[0096] Case N°2 is processed as if CPT N° 234 and CPT N° 238 were not already drilled on site during field acquisition. The training phase has consequently been performed using geophysical data close to CPT N° 237 (geophysical data inside a radius of 5 m around each CPT). Geophysical data close to CPT N° 234 and N° 238 (data corresponding to line L180 and L185 for CPT N° 238 and lines L205 and L210 for CPT N° 234) are used by the ML trained model to predict Qtnand Fr.

[0097] Note that the above geophysical lines are not indicated with reference numbers, the description is based on training dataset used in the exemplary training procedures on which the present disclosure is based.

[0098] Regarding these two configurations (case N°1 and N°2), case N°2 is the more sensitive case. Indeed, for case N°2, if geotechnical site variability is not well captured by geophysical measurements used during the training phase, then prediction will be affected by the lack of information. For the case N°l, geophysical information is collected inside the prediction volume, not implying a lack of data to describe 3D variability of geotechnical parameters.

[0099] In accordance with the present disclosure, when preparing the training dataset for training the machine learning, spatial proximity of geophysical measurements to the location of geotechnical measurements is considered. This is in consideration of the following.

[0100] In the first place, geological and geophysical properties often exhibit spatial continuity, meaning that values measured at nearby locations are more likely to be similar to each other than values measured at distant locations. This is particularly true for properties that vary gradually over space, such as soil composition, density, or groundwater flow.

[0101] Moreover, geotechnical properties are influenced by underlying geological features and structures, which can extend over relatively large spatial scales. Geophysicalmeasurements taken closer to the geotechnical measurement location are more likely to capture information about these geological features and provide insights into the subsurface conditions.

[0102] Besides, while spatial continuity is generally observed, there can also be local variability in geological and geophysical properties due to factors such as lithology changes, depositional environments, or anthropogenic activities. By prioritizing geophysical data taken closer to the geotechnical measurement location, this local variability can be better captured and the accuracy of predictions can be improved.

[0103] Based on the above considerations, sample weights are assigned to the training matrix depending on distance between the geophysical and geotechnical training data such that geophysical training data further away from the geotechnical training data has smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data.

[0104] Table 1 schematically illustrates an exemplary training matrix in which sample weights are assigned based on the distance between the geophysical and geotechnical training data.

[0105] Table 1 training matrix with sample weights

[0106] By using the training matrix with sample weights, it allows the link or association between the geotechnical and geophysical data, when both data are not located at the same place (i.e., same GPS coordinates), to vary depending on the distance between the geophysical and geotechnical data. The idea is that geophysical data which is obtained at a location closer to where the CPT is done is given a higher sample weight such that it is considered as more closely related to the geotechnical parameters.

[0107] In an example of the present disclosure, a circle can be drawn around each ID geotechnical CPT and if there are geophysical data inside the circle, those geophysical data is“connected” to CPT data but with an associated weight calculated from the distance between the two data. It will be understood by those skilled in the art that the circle as described above is for illustrative purpose only. The geophysical data chosen to be connected the geotechnical data may be selected from a square surrounding the CPT. Moreover, the radius of the circle or the edge of the square may be determined as needed, depending on the available computational resources and the expected prediction accuracy.

[0108] For example, if geophysical data is located 5m away from the CPT a weight of 1 / 52is assigned. If the geophysical data is located at the exact same GPS position as the CPT, then the weight = 1.

[0109] As an implementation, a “search radius” can be specified, through for example a user interface, around each CPT to connect CPT data with geophysical data. If geophysical data are located inside the radius, then we keep those data and associate it with a weight corresponding to the distance from the CPT (Weight = f(distance)).

[0110] An alternative option is also available to shape the weighting matrix as a cone because at the ground surface the resolution is higher compared to deeper data.[OHl] It will be understood by those skilled in the art that geospatial data is implicitly included in the training matrix as each geophysical data and geotechnical data has corresponding coordinates.

[0112] Figures 6 and 7 respectively show an example of the weighting matrix in 2D and 3D with value from 0.252to 1.

[0113] Figure 8 schematically illustrates variation of the weight as a function of distance from borehole / CPT.

[0114] ML models require a sufficient amount of data to learn meaningful patterns and relationships within the dataset. In land site characterization, the available data may be limited due to factors such as budget constraints, time limitations, or the complexity of data collection procedures. With a small sample size, ML models may struggle to capture the full range of variability in soil properties, geological features, or environmental conditions, leading to less reliable predictions.

[0115] One possible solution for the problem of lack of available data is therefore to increase the number of input information to train the model. In the present disclosure, ordinary kriging interpolation process is used to make possible larger number of connections between geophysical and geotechnical data through interpolated CPT values.

[0116] In statistics, originally in geostatistics, Kriging, which also known as Gaussian process regression, is a method of interpolation based on Gaussian process governed by priorcovariances. Under suitable assumptions of the prior, kriging gives the best linear unbiased prediction (BLUP) at unsampled locations. Interpolating methods based on other criteria such as smoothness (e.g., smoothing spline) may not yield the BLUP.

[0117] The method is widely used in the domain of spatial analysis and computer experiments. Kriging predicts the value of a function at a given point by computing a weighted average of the known values of the function in the neighbourhood of the point. The method is closely related to regression analysis. Both theories derive a best linear unbiased estimator based on assumptions on covariances, make use of Gauss-Markov theorem to prove independence of the estimate and error, and use very similar formulae. Even so, they are useful in different frameworks: kriging is made for estimation of a single realization of a random field, while regression models are based on multiple observations of a multivariate data set.

[0118] The kriging estimation may also be seen as a spline in a reproducing kernel Hilbert space, with the reproducing kernel given by the covariance function. The difference with the classical kriging approach is provided by the interpretation: while the spline is motivated by a minimum-norm interpolation based on a Hilbert-space structure, kriging is motivated by an expected squared prediction error based on a stochastic model.

[0119] Kriging with polynomial trend surfaces is mathematically identical to generalized least squares polynomial curve fitting.

[0120] Kriging can also be understood as a form of Bayesian optimization. Kriging starts with a prior distribution over functions. This prior takes the form of a Gaussian process:

[0121] N samples from a function will be normally distributed, where the covariance between any two samples is the covariance function (or kernel) of the Gaussian process evaluated at the spatial location of two points. A set of values is then observed, each value associated with a spatial location. Now, a new value can be predicted at any new spatial location by combining the Gaussian prior with a Gaussian likelihood function for each of the observed values. The resulting posterior distribution is also Gaussian, with a mean and covariance that can be simply computed from the observed values, their variance, and the kernel matrix derived from the prior.

[0122] By “geo-statistically” increasing the number of input information, the ML model will be better trained to predict a general trend instead of precise expected value at a specific site location. Adding prior geo-statistical, GS, procedure can therefore be considered as a more conservative way (less risky) to predict ground parameters.

[0123] Based on the above, 3D interpolation processes were performed for case N°1 and case N°2. For each case CPTs data used to evaluate the prediction quality (CPT N°232 andN°238 for caseN°l and CPT N°234 andN°238 for case N°2) were exclude from the 3D kriging interpolation process.

[0124] An exponential semi-variogram model was used, associated with a vertical anisotropy factor. Vertical anisotropy factor is set to ensure that both Qtnand Frstrongly vary with depth compared to lateral variation. To consider the difference between vertical and longitudinal variation, an anisotropy factor is automatically set through a stopping criteria based on the correlation coefficient (R2) between measured and interpolated CPT values.

[0125] The vertical anisotropy factor is 18 for friction ratio and 12 for Qtnwith an associated correlation coefficient R2of 0.7 and 0.9 respectively.

[0126] Figure 9 shows the comparison between measured and interpolated values, respectively solid and dashed line for CPT N° 236.

[0127] The use of prior 3D kriging interpolation process helps to increase the size of the training matrix, which will help to avoid as much as possible the overfitting problem.

[0128] In the example of the present disclosure, the 3D kriging interpolation process enables to pass from a training matrix size of around 20 000 lines to more than one million lines for case N°1 and around 800 000 lines for case N°2. Note that training matrix size for case N°1 is higher because all geophysical lines are included in the ML process. The amount of input information therefore is increased by a factor of 50 and 40 respectively for case N'T and case N°2.

[0129] Interpolated CPT value in 3D can also be associated with a complete 3D weighting matrix (no more radius parameters) to increase by more than one order of magnitude the size of the training matrix. In that case the weighting matrix is associated to each interpolated CPT parameter and take the form of a 3D volume as illustrated in Figure 10, which presents the 3D weighting matrix visualisation that is associated with the training matrix. An adapted weighting matrix is built to give more weight to the training matrix lines associated with a smaller spatial distance from CPT data.

[0130] When the training matrix is ready, a ML model is built using all possible geophysical information collected on site (and not only geophysical data close to CPT locations) and the proven 3D GS approach is an interesting mean to include the possibility of global trend prediction, thus avoiding totally wrong prediction. Obviously, using interpolated value decreases the model sensitivity to determine local ground singularity that may correspond to project objectives.

[0131] Instead of limiting the ML model to geophysical data collected only at specific locations (such as close to the CPT locations), the model incorporates all available geophysicalinformation collected across a much wider area over the site. This approach allows the model to consider a broader range of spatial variations and patterns in the geophysical data, potentially capturing more comprehensive insights into the subsurface conditions.

[0132] By using the 3D GS approach and incorporating all available geophysical information, the ML model can avoid making "totally wrong" predictions by capturing the broader trends and patterns in the data. Instead of relying solely on localized geophysical measurements, which may not fully represent the spatial variability of the subsurface, the model considers the overall trends and structures revealed by the 3D GS approach. This helps to reduce the risk of making erroneous predictions based on incomplete or localized information.

[0133] While using interpolated values from the 3D GS approach may decrease the model's sensitivity to determine local ground singularities (such as abrupt changes or anomalies in soil properties), it also provides a more comprehensive understanding of the overall subsurface conditions. While the model may not capture every localized variation in the ground properties, it can still provide valuable insights into the broader trends and characteristics of the site, which may be sufficient for many project objectives.

[0134] It is discussed above that spatial features are considered in the present disclosure. The addition of spatial features into the training process brings important information related to spatial connectivity. However, using the coordinates as spatial variables may generate considerable overfitting because they are highly correlated. The increasing of training matrix size also helps to overcome this difficulty by considering more geophysical data in the learning process and making advantage of spatial connectivity information.

[0135] Various machine learning algorithm including CatBoost algorithm, Random Forest, Gradient Boosting, AdaBoost, XGBoost can be used in the present disclosure. Catboost ML algorithm will be described as an exemplary ML algorithm in the following.

[0136] Catboost ML algorithm uses gradient boosting on decision trees. Gradient boosting is a machine learning technique used for regression and classification tasks. It works by sequentially combining weak learners (usually decision trees) to create a strong predictive model. In gradient boosting, each new tree is trained to correct the errors made by the ensemble of existing trees, with a focus on minimizing the residual errors of the previous models.

[0137] Decision trees are a type of supervised learning algorithm used for both classification and regression tasks. They recursively partition the input space into regions based on feature values, with each region corresponding to a leaf node in the tree. Decision trees areknown for their simplicity, interpretability, and ability to capture non-linear relationships in the data.

[0138] CatBoost is a gradient boosting algorithm that specifically targets categorical features in tabular data. It builds on the principles of gradient boosting and incorporates several optimizations to handle categorical variables more effectively. CatBoost uses decision trees as base learners and iteratively improves the model's performance by minimizing the gradient of a specified loss function.

[0139] CatBoost can handle categorical features directly, without the need for preprocessing such as one-hot encoding or label encoding. It uses an efficient algorithm to convert categorical variables into numerical values during training.

[0140] CatBoost implements various optimizations to improve training speed and reduce memory usage, making it especially suitable for large-scale datasets with high-dimensional feature spaces.

[0141] The ML model such as one using CatBoost is trained from a series of train-test validation process. The model with the highest R2scores between measured and predicted values is selected.

[0142] Assuming that Qtnand Frare correlated to the same geology, a multi -regression approach was selected. Indeed, combining several regressions into a single multi-target model make possible of exploiting correlations among targets.

[0143] In the present disclosure, a parametric ML is performed to explore several combinations of input features, including for example geophysical data, geophysical data derived values and spatial features.

[0144] Figures 11 and 12 respectively show measured and predicted Qtnand Frvalues for CPT N°232 and CPT N°238 of the first site configuration. Figures 11 and 12 present prediction values (in thick solid line) of Fr, indicated by reference numerals 111 and 121, and prediction values of Qtn, indicated by reference numerals 112 and 122, using feature combination Vsand Elevation (Z) without kriging process. Figures 11 and 12 further presents prediction values (in thick light gray dashed line) ofTL, indicated by reference numerals 113 and 123, and prediction values of Qtn, indicated by reference numerals 114 and 124, using Vsand XYZ with prior 3D kriging process. To highlight the benefit of adding geophysical data into the learning process, predictions performed only using spatial features XYZ are presented in thin dashed line. Moreover, to point out overfitting problem, the Qtn prediction for CPT N°2 without GS preprocessing step and using Vs+ XYZ is also presented in thin solid lines.

[0145] Figures 13 and 14 respectively show measured and predicted Qtnand Frvalues for CPT N°234 and CPT N°238 of the second site configuration. Lines in same patterns indicate same parameters.

[0146] It can be seen from Figures 11 to 14 that having geophysical data into the learning process is important. For all CPT, Vsinformation makes possible a better Qtnprediction. Indeed, using only spatial features XYZ, ML predictions result give Qtnestimation values globally 100 MPa less than reals measured Qtn values and predictions including Vsinformation (shift of around 100 MPa between thick solid lines and thin dashed lines presented in Figures 11 to 14).

[0147] For Fr, similar observation can be done at depth where Frreach 5% for CPT N°232 and N°234. For these CPTs, Frpredictions are better using Vsinstead of using only spatial features.

[0148] The influence of prior GS process upon prediction quality is also seen from Figures 11 to 14. Broadly, the prediction quality is not altered. Vertical variations are weaker and prior kriging process seems to give a more conservative way of predicting CPT values. Even if the dataset do not present strong 3D variations (which could have helped to better identify efficiencies of presented ML methodology), in first order, having a first set of prediction without coordinates information (and consequently a learning phase more focussed on geophysical data) and a second set of prediction less hazardous in terms of prediction robustness (using prior GS process including spatial XY features) seems to be a good compromise.

[0149] The latest point seems to be confirmed by CPT N°232 Qtndata prediction from the input feature combination “K + XYZ” without prior GS phase (see thick black dashed line in Figure 11). In this case, overfitting process seems to be the cause of the bad Qt„ data prediction using a model trained with spatial features.

[0150] In addition to the above, as can be understood by those skilled in the art, geophysical data resolution decreases exponentially with depth. That is, as it goes deeper into the subsurface, the resolution or detail of the geophysical data decreases. This could be due to factors such as attenuation of signals, increased noise, or limitations of the measurement technology.

[0151] Based on the above, CPT data can be smoothed accordingly. As an example, a Dynamic Moving Average (DMA) step can be implemented, where when the depth goes deeper then the size of the moving average window increases (see Figure 15).

[0152] The purpose of smoothing the CPT data using DMA is to adjust for the decreasing resolution of geophysical data with depth. By increasing the size of the moving average window as it goes deeper, more data points can get effectively averaged to obtain a smoother representation of the CPT data. This helps to preserve the high resolution of the CPT data near the ground surface, where it may be more reliable, while gradually smoothing the data with increasing depth where geophysical data resolution is lower.

[0153] By smoothing the CPT data more with depth, the sensitivity to minor variations or features in the subsurface is reduced and it helps to focus more on major structures or trends. This approach acknowledges that deeper subsurface features are likely to be captured more accurately by geophysical methods, which have inherently lower resolution at depth.

[0154] The end result (see Figure 16) is that we almost keep the high CPT resolution at ground surface but then we smooth data more and more with depth because with geophysical data we never have CPT resolution with depth, and we will only be sensitive to major structures. In Figure 16, the smoothed qcvalues are shown with the doted line while the original values are shown with the solid line.

[0155] The trained model as described above can be used to predict geotechnical parameters using geophysical measurements.

[0156] The geotechnical parameters can comprise CPT parameters, including cone resistance and sleeve friction.

[0157] Parameters such as soil behaviour type index may also be predicted using the training ML model. It is also possible to predict all wireline soil parameters recording such as gamma ray for example.

[0158] As described above, the geophysical measurement data are obtained using one or more of Electrical Resistivity Tomography, Multichannel Analysis of Surface Waves (Vs), seismic refraction tomography (Vp), Ground-Penetrating Radar.

[0159] These methods are non-invasive and much less costly than the geotechnical methods, therefore these methods can be conveniently used to obtain the geophysical measurement data, which will then be used to predict the geotechnical parameters.

[0160] The predicting step comprises predicting trend of the geotechnical parameters.

[0161] The method of the present disclosure will possibly be better trained to predict a general trend instead of precise expected value at a specific site location, this is in consideration that the training data includes data derived using the geo-statistical approach.

[0162] The method of the present disclosure further comprises building at least one of a 2D and 3D ground model of the area based on the predicted geotechnical parameters.

[0163] A detailed ground model provides valuable information for geotechnical design and engineering. Engineers can use the model to assess soil stability, analyse foundation behaviour, and design appropriate earthworks, retaining structures, and foundations.

[0164] The invention has been described by reference to certain embodiments discussed above. It will be recognized that these embodiments are susceptible to various modifications and alternative forms well known to those of skill in the art.

[0165] Further modifications in addition to those described above may be made to the structures and techniques described herein without departing from the spirit and scope of the invention. Accordingly, although specific embodiments have been described, these are examples only and are not limiting upon the scope of the invention.

Claims

CLAIMS1. A method of training a machine learning model representing an association between geotechnical parameters and geophysical measurement data, comprising the steps of: obtaining geotechnical training data for a first area; obtaining geophysical training data for a second area, wherein the second area is larger than and surrounds the first area; assigning sample weights to a training matrix depending on distance between the geophysical and geotechnical training data such that geophysical training data further away from the geotechnical training data has a smaller weight during a learning phase compared to geophysical training data closer to the geotechnical training data, wherein the training matrix comprises the geotechnical training data and the geophysical training data; training the machine learning model using the weighted training matrix.

2. The method according to claim 1, wherein the geotechnical training data comprises 3D geotechnical data derived from geotechnical measurement data using a geo-statistical, GS, approach, such as ordinary kriging interpolation.

3. The method according to claim 2, wherein the geotechnical measurement data is cone penetrometer test data.

4. The method according to claim 2 or 3, wherein an exponential semi-variogram model is used in the ordinary kriging interpolation.

5. The method according to any of the previous claims 2 to 4, wherein vertical anisotropy factor is used in the ordinary kriging interpolation.

6. The method according to any of the previous claims, wherein the training step comprises series of train-test validation procedure.

7. The method according to any of the previous claims, wherein the machine learning model comprises one or more algorithms selected from the group consisting of CatBoostalgorithm, Random Forest algorithm, Gradient Boosting algorithm, AdaBoost algorithm, and XGBoost algorithm.

8. A machine learning model trained according to any of the previous claims 1 to 7.

9. A method of predicting geotechnical parameters from geophysical measurement data, the method comprising: predicting geotechnical parameters in an area using the machine learning model of claim 8 by using geophysical measurement data as input data.

10. The method according to claim 9, wherein the geotechnical parameters comprises cone penetrometer test parameters, including cone resistance and sleeve friction.

11. The method according to claim 9 or 10, wherein the geophysical measurement data are obtained using one or more of Electrical Resistivity Tomography, Multichannel Analysis of Surface Waves (Vs), seismic refraction tomography (Vp), Ground-Penetrating Radar analysis technology.

12. The method according to any of the previous claims 9 to 11, wherein the geophysical training data comprises geophysical measurement data, geophysical data derived values and spatial features.

13. The method according to any of the previous claims 9 to 12, wherein the predicting step comprises predicting trend of the geotechnical parameters.

14. The method according to any of the previous claims 9 to 13, further comprising building at least one of a 2D and 3D ground model of the area based on the predicted geotechnical parameters.

15. A computer program product, comprising a computer readable storage medium storing instructions which, when executed on at least one processor, cause the at least one processor to carry out the method according to any of claims 1 to 7 or according to any of claims 9 to 14.

Citation Information

Patent Citations

  • Machine-learning integration for 3D reservoir visualization based on information from multiple wells

    US20220122320A1

Cited By

  • Method and electronic device of inferring ground properties based on digital twin model

    KR102973325B1