Geotechnical Sampling Data Augmentation Method Based on Variational Autoencoder
The VAE-based soil sampling data enhancement method addresses data scarcity and model inefficiencies by generating high-quality soil data, enhancing model performance and reducing experimental costs.
Patent Information
- Application Number
- CN202410802341.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-06-20
AI Technical Summary
In existing soil mechanics research, traditional experimental methods are costly, long periods and difficult to obtain data. Machine learning methods are prone to overfitting and lack generalization capabilities in small data sets, making it difficult to effectively capture the nonlinear relationships of the soil.
The geotechnical sampling data augmentation method based on variational autoencoder (VAE) is adopted to expand the data set by generating the model, combining physical property loss terms and consistency loss terms, update the total loss function of the VAE generation model, and generate more new data that conforms to the original data distribution, which is used to train machine learning models.
It improves the scale and diversity of data sets, enhances the generalization ability of machine learning models, improves the prediction accuracy of models, reduces the demand for field testing and laboratory experiments, and improves data security and model stability.
Smart Images

Figure CN118708967B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of geotechnical data augmentation, and particularly to a method for enhancing geotechnical sampling data based on a variational autoencoder. Background Art
[0002] Existing soil mechanics technologies mainly rely on traditional experimental methods and statistical models to evaluate soil properties. These methods include direct shear tests and triaxial shear tests. Although they can measure the shear strength of soil, due to the long experimental period, high cost, and sensitivity to environmental conditions, it is difficult to obtain data and the amount of experimental data is insufficient. Traditional statistical models, such as linear regression, often perform poorly when dealing with complex soil characteristics because they cannot capture non-linear relationships. In addition, current machine learning methods also face some challenges when dealing with soil data, such as data scarcity, model overfitting, and insufficient generalization ability.
[0003] To address these problems in soil mechanics research, recent studies have attempted to use machine learning technologies such as artificial neural networks, support vector machines, and Gaussian process regression. However, these methods are prone to overfitting when dealing with small datasets and have limited performance when facing diverse soil characteristics. Summary of the Invention
[0004] In view of the above deficiencies in the prior art, the method for enhancing geotechnical sampling data based on a variational autoencoder provided by the present invention has the problems of difficult data acquisition and high cost in the prior art, which obtains a large amount of soil data through experiments for machine learning.
[0005] To achieve the above object of the invention, the technical solution adopted by the present invention is as follows:
[0006] Provide a method for enhancing geotechnical sampling data based on a variational autoencoder, which includes the steps of:
[0007] S1. Obtain a dataset including a number of samples, and preprocess the soil data in the dataset; each sample includes various types of geotechnical data;
[0008] S2. Input the preprocessed dataset into a trained VAE generation model for data augmentation to obtain reconstructed data;
[0009] S3. Perform interval limitation on the reconstructed data to update the physical property loss term of the dataset;
[0010] S4. Use any soil data in the reconstructed data as a prediction item, and then input the non-prediction items in the reconstructed data into a trained regression model to obtain regression prediction data;
[0011] S5. Update the consistency loss term using the prediction item and the regression prediction data, and update the total loss function of the VAE generation model according to the consistency loss term and the physical property loss term;
[0012] S6. Increment the iteration count by one, and determine whether the iteration count is greater than or equal to the preset number. If so, output the reconstructed data at the most recent iteration; otherwise, return to step S2.
[0013] Furthermore, the expression of the physical property loss term is:
[0014]
[0015] where, is the physical property loss term; α is the importance of controlling the constraint; m is the total number of samples in the reconstructed data; x i2 is the prediction item in the i-th sample; x max and x min are the maximum and minimum values respectively among the prediction items of all samples in the reconstructed data; max(.) is the function to take the maximum value;
[0016] The expression of the consistency loss term is:
[0017]
[0018] where, is the consistency loss term; is the regression prediction data of the regression model with respect to the prediction item in the i-th sample.
[0019] The expression of the total loss function of the VAE generation model is:
[0020]
[0021] where, is the total loss; is the VAE damage function.
[0022] Furthermore, the expression of the VAE damage function is:
[0023]
[0024] where β is a hyperparameter for balancing the reconstruction loss and the KL divergence; is the similarity between the dataset and the reconstructed data generated by the decoder; p θ(x|z) is the probability distribution for the decoder to reconstruct the original data from the samples in the latent space; q φ (z|x) is the probability distribution for the encoder to map the input data to the latent space; p(z) is the prior distribution of the embedding space; D KL (qφ (z|x)||p(z)) is q φ The difference degree between (z|x) and p(z); x is the data set; z is the latent variable; are the encoder network parameters; θ are the decoder network parameters; and are the mean and variance of the latent variable respectively; μ θ (z) and are the mean and variance of the reconstructed data respectively; is the standard normal distribution.
[0025] Furthermore, various types of geotechnical data include the depth of measurement points, effective stress, normalized undrained shear strength, overconsolidation ratio, normalized cone tip resistance, normalized effective cone tip resistance, normalized excess pore water pressure, and pore pressure ratio; the prediction item is the normalized undrained shear strength.
[0026] Furthermore, the expressions of the normalized cone tip resistance f1, the normalized effective cone tip resistance f2, the normalized excess pore water pressure q1, and the pore pressure ratio q2 are respectively:
[0027]
[0028] where q t is the cone tip resistance; σ v is the vertical stress; σ v ′ is the effective stress; u2 is the pore water pressure; u0 is the initial pore water pressure.
[0029] Furthermore, the preprocessing of the soil data in the data set includes processing missing values, outliers, and unifying the data format for the samples in sequence.
[0030] Furthermore, the data set is the Clay / 6 / 535 data set in the ISSMGE TC304 soil database; the regression model is the GPR regression model.
[0031] Furthermore, both the encoder and decoder of the VAE generation model include five-layer fully connected neural networks; the outputs of the first four fully connected neural networks of the encoder pass through a batch normalization layer and a Dropout layer in sequence and then enter the next fully connected neural network; the outputs of the first two fully connected neural networks of the decoder pass through a batch normalization layer and then enter the next fully connected neural network.
[0032] The beneficial effects of the present invention are as follows: This solution uses a VAE generative model to address data scarcity, enhance the scale and diversity of the dataset, and improve the generalization ability of machine learning models. By utilizing the VAE generative model, more new data conforming to the original data distribution can be obtained. Training the model with the generated data can improve the prediction accuracy of the model and mitigate the risk of overfitting.
[0033] The diverse data of the VAE generative model helps to more accurately capture the non-linear relationships in soil properties, thereby improving the performance and stability of the model. The data generated by the VAE generative model is similar to the original data distribution, which can better protect the original data during data transmission and processing. This approach can both expand the dataset and avoid disclosing the original data, thus enhancing data security.
[0034] Furthermore, by generating additional data through VAE, the need for field tests and laboratory experiments can be reduced; this means that when training machine learning models, there is no longer a need to invest a large amount of resources in experiments, which is particularly important in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of a geotechnical sampling data enhancement method based on a variational autoencoder.
[0036] Figure 2 is a two-dimensional distribution comparison diagram of the data obtained by using the enhancement method of this solution and the original data.
[0037] Figure 3 is a regression performance comparison diagram of five regression models before and after data enhancement. DETAILED DESCRIPTION OF THE INVENTION
[0038] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0039] Refer to Figure 1 , Figure 1 which shows a flowchart of a geotechnical sampling data enhancement method based on a variational autoencoder; as Figure 1 shown, this method S includes steps S1 to S6.
[0040] In step S1, a data set including a number of samples is obtained, and the soil data in the data set is preprocessed; each sample includes various types of geotechnical data; among them, the various types of geotechnical data include the depth of the measurement point, effective stress, normalized undrained shear strength, overconsolidation ratio, normalized cone tip resistance, normalized effective cone tip resistance, normalized excess pore water pressure, and pore pressure ratio.
[0041] In implementation, this solution preferably uses the Clay / 6 / 535 data set in the ISSMGE TC304 soil database as the data set. This data set contains 535 groups of lightly overconsolidated clay data from 40 sites distributed in Brazil, Canada, Italy, Malaysia, Singapore, Norway, the United States, the United Kingdom, Sweden, the North Sea, and Venezuela.
[0042] In an embodiment of the present invention, preprocessing the soil data in the data set includes successively performing missing value processing, outlier processing, and data format unification on the samples:
[0043] Missing value processing: Missing values may exist in geotechnical data due to measurement difficulties or equipment failures. For continuous data, consider using K-nearest neighbor interpolation, geology-based interpolation methods, or model-based filling (such as using regression models for prediction). For discrete data, the mode or category feature-based mode filling can be used.
[0044] Outlier processing: Since geotechnical properties are affected by various natural factors, outliers may reflect real geological phenomena. It should be carefully processed, assisted by geological expert knowledge for identification, only obvious error records are removed, or judged through box plots, IQR combined with professional knowledge.
[0045] Data format unification: Ensure that geological data from different sources (such as collected by different laboratories and different instruments) are consistent in units, coordinate systems, and measurement standards, and perform conversions if necessary.
[0046] When the samples in the data set can be divided into time data or spatial sequence data, after the above-mentioned preprocessing, the following methods are also required for preprocessing:
[0047] Spatial data processing: For geotechnical data containing spatial information, perform coordinate system calibration, projection conversion, and use GIS tools for spatial interpolation and data smoothing to eliminate geographical information errors.
[0048] Time series data processing: For monitoring data that changes over time, perform time series analysis to identify trends, seasonality, and perform detrending and deseasonalization processing to better reveal data characteristics.
[0049] In step S2, the preprocessed dataset is input into the trained VAE generative model for data augmentation to obtain reconstructed data;
[0050] In implementation, this solution preferably uses a five-layer fully-connected neural network for both the encoder and decoder of the VAE generative model; the outputs of the first four fully-connected neural networks of the encoder pass through a batch normalization layer and a Dropout layer in sequence and then enter the next fully-connected neural network; after the outputs of the first two fully-connected neural networks of the decoder pass through a batch normalization layer, they enter the next fully-connected neural network.
[0051] The data input dimension used by the VAE generative model is 8. The five-layer fully-connected neural network of the encoder maps the data dimension in sequence as 8 -> 10 -> 12 -> 14 -> 16 -> 20, while the decoder maps in the reverse order of the encoder. Given that geotechnical data usually contains information in multiple dimensions such as soil type, physical properties (such as permeability, density), and chemical composition, it is crucial to adopt a multi-channel input design. This allows the model to process input data of different modalities in parallel and improves its ability to recognize and express multi-dimensional features.
[0052] In step S3, interval limitation is performed on the reconstructed data to update the physical property loss term of the dataset:
[0053]
[0054] where, is the physical property loss term; a is to control the importance of the constraint; m is the total number of samples in the reconstructed data; x i2 is the prediction term in the i-th sample; x max and x min are the maximum and minimum values respectively among the prediction terms of all samples in the reconstructed data; max(.) is the function to take the maximum value.
[0055] This solution introduces evaluation metrics based on physical principles, such as the difference metric of key physical parameters such as permeability and compression coefficient, as additional loss terms, forcing the model to respect the physical properties of geotechnical materials during the learning process and ensuring that the generated samples are physically feasible.
[0056] In step S4, any soil data in the reconstructed data is used as the prediction term, and then the non-prediction terms in the reconstructed data are input into the trained regression model to obtain regression prediction data; this solution preferably uses the normalized undrained shear strength as the prediction term.
[0057] Among them, both the regression model and the VAE generative model are trained using small-sample training data (the existing soil dataset), and the regression model is preferably a GPR regression model.
[0058] In step S5, the consistency loss term is updated using the prediction term and the regression prediction data:
[0059]
[0060] Where, is the consistency loss term; is the regression prediction data of the regression model with respect to the prediction term in the i-th sample.
[0061] According to the consistency loss term and the physical property loss term, the total loss function of the VAE generation model is updated:
[0062]
[0063] Where, is the total loss; is the VAE damage function; the expression of the VAE damage function is:
[0064]
[0065] Where, β is a hyperparameter used to balance the reconstruction loss and the KL divergence; is the similarity between the data set and the reconstructed data generated by the decoder; p θ(x|z) is the probability distribution for the decoder to reconstruct the original data from the samples in the latent space; q φ (z|x) is the probability distribution for the encoder to map the input data to the latent space; p(z) is the prior distribution of the embedding space; D KL (q φ (z|x)||p(z)) is the difference degree between q φ (z|x) and p(z); x is the data set; z is the latent variable; are the encoder network parameters; θ are the decoder network parameters; and are the mean and variance of the latent variable respectively; μ θ (z) and are the mean and variance of the reconstructed data respectively; is the standard normal distribution.
[0066] In step S6, the iteration count is incremented by one, and it is determined whether the iteration count is greater than or equal to the preset number of times. If so, the reconstructed data at the most recent iteration is output; otherwise, return to step S2.
[0067] In an embodiment of the present invention, the expressions of the normalized cone tip resistance f1, the normalized effective cone tip resistance f2, the normalized excess pore water pressure q1, and the pore pressure ratio q2 are respectively:
[0068]
[0069] Among them, q t is the cone tip resistance; σ v is the vertical stress; σ v ′ is the effective stress; u2 is the pore water pressure; u0 is the initial pore water pressure.
[0070] To evaluate the effect of the VAE generation model in the data augmentation task of non-image data Clay / 6 / 535, 535 samples were augmented using the data augmentation method of this scheme, and the augmented data retained all the features of the initial data.
[0071] To explore the influence of different data augmentation amounts on the performance of the regression model, when the encoder maps the data to the latent space, the number of sampling times in the latent space starts from 1% of the original data volume and increases by 1% successively until 20%.
[0072] To visualize the changes in the data before and after augmentation, t-sne was used in the experiment to reduce the dimension of the data. Figure 2 Figure shows the two-dimensional distribution of the data with an 8% data augmentation amount and the original data. The circular points in the figure represent the two-dimensional distribution of the original data, and the triangular points represent the two-dimensional distribution of the augmented data.
[0073] From Figure 2 It is not difficult to find that the distribution of the augmented data is similar to that of the original data, only the direction of the data distribution has changed. This is because of the data sampling in the latent space. The sampled data only conforms to the original data distribution and is not a complete reconstruction of the original data, that is, the two distributions are similar but the directions are different.
[0074] The distribution of the augmented data is more compact than that of the original data, indicating that the data generated by the data augmentation method of this scheme is more aggregated and continuous in the latent space. This situation can usually be explained as that the generation model has learned the latent structure of the data and can represent the features of the data in a more continuous and compact way.
[0075] Through experiments, it is found that the data augmentation method of this scheme can not only expand the data set, but also improve the quality of the data and smooth the original data.
[0076] Next, the prediction accuracy of different regression models is used to judge the improvement of the data quality after data augmentation:
[0077] Five regression models were used for experimental verification, namely ANN regression prediction model, SVM regression model, GPR regression model, RFR random forest regression, and XGBoost regression model; compare the differences in the regression performance of the five ML algorithms at each index before and after data augmentation. In the experiment, the data enhanced with different augmentation ratios was input into the five regression models to calculate the MAE, and the obtained multiple MAEs were averaged, which was used as the performance MAE evaluation index on the data after data augmentation, as Figure 3 shown.
[0078] As Figure 3 shown, the regression performance of the GPR model for the Clay / 6 / 535 data was the worst when no data augmentation was performed. When the data was enhanced by VAE, the performance of the RF model was the best. Although the performance of the GPR model was the worst before data augmentation, data augmentation had the greatest impact on the performance of the GPR model, greatly improving the model performance.
[0079] The experiment verified all the indexes of the five regression models on the Clay / 6 / 535 data. In addition to the MAE, the regression model performance evaluation indexes also included: mean square error (MSE), coefficient of determination (R2), and explained variance score (EVS).
[0080] Among them, the value ranges of MAE and MSE are [0, +∞], and the smaller the value, the better the performance of the regression model; the value ranges of R2 and EVS are [-∞, 1], and the larger the value, the better the performance of the regression model; the comparison of the performance evaluation indexes of the five regression models before and after data augmentation refers to Table 1.
[0081] Table 1 Performance evaluation indexes of the model before and after data augmentation
[0082]
[0083] It can be seen from the comparison of the data before and after augmentation in Table 1 that the performance of the regression model was improved under all evaluation indexes after data augmentation. Considering the four evaluation indexes comprehensively, the performance of the GPR regression model was improved most significantly when the data was augmented; that is, when the data was augmented, the quality of the reconstructed data obtained by using the GPR regression model was the best.
Claims
1. A method for enhancing geotechnical sampling data based on a variational autoencoder, characterized in that, Including the steps: S1. Obtain a dataset including a number of samples, and preprocess the soil data in the dataset; each sample includes various types of geotechnical data; S2. Input the preprocessed dataset into the trained VAE generation model for data augmentation to obtain reconstructed data; S3. Perform interval limitation on the reconstructed data to update the physical property loss term of the dataset; S4. Take any soil data in the reconstructed data as the prediction item, and then input the non-prediction items in the reconstructed data into the trained regression model to obtain regression prediction data; S5. Update the consistency loss term using the prediction item and the regression prediction data, and update the total loss function of the VAE generation model according to the consistency loss term and the physical property loss term; the expression of the physical property loss term is: Among them, is the physical property loss term; is the importance of control constraints; m is the total number of samples in the reconstructed data; is the i th predicted item in the sample; and are the maximum and minimum values respectively among the predicted items of all samples in the reconstructed data; is the maximum value function; The expression of the consistency loss term is: Among them, is the consistency loss term; is the regression prediction data of the regression model with respect to the prediction item in the i th sample; The expression of the total loss function of the VAE generation model is: Among them, is the total loss; is the VAE damage function; S6. Increment the iteration count by one, and determine whether the iteration count is greater than or equal to the preset number. If so, output the reconstructed data at the most recent iteration; otherwise, return to step S2.
2. The method for enhancing geotechnical sampling data based on variational autoencoder according to claim 1, wherein The expression of the VAE damage function is: , Among them, is a hyperparameter for balancing the reconstruction loss and the KL divergence; is the similarity between the dataset and the reconstructed data generated by the decoder; is the probability distribution for the decoder to reconstruct the original data from samples in the latent space; is the probability distribution for the encoder to map the input data to the latent space; is the prior distribution of the embedding space; is and the degree of difference between; x is the dataset; z is the latent variable; is the encoder network parameter; is the decoder network parameter; and are the mean and variance of the latent variable respectively; and are the mean and variance of the reconstructed data respectively; is the standard normal distribution.
3. The method for enhancing geotechnical sampling data based on variational autoencoder according to claim 1, wherein The various types of geotechnical data include the depth of the measurement point, effective stress, normalized undrained shear strength, overconsolidation ratio, normalized cone tip resistance, normalized effective cone tip resistance, normalized excess pore water pressure, and pore pressure ratio; the prediction item is the normalized undrained shear strength.
4. The method for enhancing geotechnical sampling data based on variational autoencoder according to claim 3, characterized in that Normalized cone tip resistance 、Normalized effective cone tip resistance 、Normalized excess pore water pressure and pore pressure ratio are expressed as follows: , , , Among them, is the cone tip resistance; is the vertical stress; is the effective stress; is the pore water pressure; is the initial pore water pressure.
5. The method for enhancing geotechnical sampling data based on variational autoencoder according to claim 3, characterized in that Preprocessing the soil data in the dataset includes performing missing value processing, outlier processing, and data format unification on the samples in sequence.
6. The method for enhancing geotechnical sampling data based on a variational autoencoder according to any one of claims 1-5, characterized in that, The dataset is the Clay / 6 / 535 dataset in the ISSMGE TC304 soil database; the regression model is a GPR regression model.
7. The method for enhancing geotechnical sampling data based on variational autoencoder according to any one of claims 1-5, characterized in that Both the encoder and decoder of the VAE generation model include five-layer fully connected neural networks; the outputs of the first four fully connected neural networks of the encoder pass through a batch normalization layer and a Dropout layer in sequence and then enter the next fully connected neural network; the outputs of the first two fully connected neural networks of the decoder pass through a batch normalization layer and then enter the next fully connected neural network.
Citation Information
Patent Citations
Zero-sample image classification method based on regression variation auto-encoder
CN111563554A
Heart data anomaly detection method based on unsupervised adaptive weight
CN114548281A