Soil undrained shear strength rejection method based on out-of-distribution detection

By using out-of-distribution detection technology and edge calculation in soil shear strength prediction, untrusted samples are identified and rejected, the problem of inaccurate soil shear strength prediction in the prior art is solved, the accuracy and reliability of the prediction are improved, and the safety of engineering construction is ensured.

CN119943199APending Publication Date: 2025-05-06中国水利水电第七工程局有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When facing unseen soil samples, the prediction results of existing soil shear strength prediction models are inaccurate and difficult to identify the samples outside the distribution, resulting in safety hazards in engineering construction.

Method used

The soil irresistance shear strength rejection method based on external distribution detection is adopted, and the soil data is pre-processed through edge computing technology, combined with Gaussian mixed model, isolated forest and deep learning method for external distribution detection, identify and reject untrusted samples, and update the predicted model parameters.

Benefits of technology

It improves the accuracy and reliability of soil shear strength prediction, reduces calculation complexity and network resource consumption, enhances users' trust in the prediction results, and ensures the safety of project construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943199A_ABST
    Figure CN119943199A_ABST
Patent Text Reader

Abstract

The invention discloses a soil undrained shear strength rejection method based on out-of-distribution detection, and the method comprises the following steps: S1, collecting soil data, carrying out the preprocessing of the soil data through employing an edge calculation technology, and generating the preprocessed soil data; s2, dividing the preprocessed soil data into a training set, a verification set and a test set, and constructing a prediction model of the undrained shear strength of the soil; s3, acquiring new soil data, preprocessing the new soil data, judging whether the new soil data is out-of-distribution data or not through out-of-distribution detection, if yes, entering S4, and if not, entering S5; s4, sending prompt information and data processing suggestions to the user through a rejection mechanism, recording out-of-distribution data, in response to recording a preset amount of out-of-distribution data within a preset time, updating parameters of the prediction model, and returning to S3; and S5, inputting soil data subjected to out-of-distribution detection into the prediction model to obtain a prediction result of the undrained shear strength of the soil.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of soil and water conservation, and in particular relates to a method for identifying the undrained shear strength of soil based on out-of-distribution detection. Background Art

[0002] In civil engineering and foundation design, the undrained shear strength of soil is one of the important parameters for evaluating soil stability and construction safety. Accurately predicting the undrained shear strength of soil is of great significance to the safety and economy of engineering construction. However, due to the complexity and regionality of soil properties, soil characteristics in different regions and different soil layers often vary greatly, which brings many challenges to existing technologies.

[0003] Existing soil shear strength prediction methods are mainly based on data-driven machine learning models, which usually acquire prediction capabilities by training and testing on existing observation samples. For example, data-driven methods such as artificial neural networks, support vector machines, and regression models are used to model and predict soil samples collected in the laboratory. These methods often show good prediction accuracy on data with the same distribution, especially when the training and test samples are consistent.

[0004] However, in practical applications, the source and properties of soil often vary significantly due to differences in environment and geographical location. For example, the soil composition and physical properties in different regions may be completely different, which makes the assumptions of the pre-trained model inconsistent with the actual situation, leading to significant prediction bias. In particular, when the model is faced with unseen soil samples that are significantly different from the training data, these data-driven models may produce extremely inaccurate predictions. What's more serious is that such errors are often difficult for users to detect or identify in time, which in turn poses serious safety hazards to engineering construction. Especially when soil characteristics affect construction safety, this uncertainty can lead to catastrophic consequences.

[0005] Out-of-Distribution Detection (OOD) is an effective way to solve the above problems. The goal of OOD is to identify input samples that are inconsistent with the distribution of training data. These samples are called out-of-distribution samples. Through effective out-of-distribution detection, samples that do not conform to the model training distribution can be identified, thereby avoiding unreliable predictions on these samples. The rejection strategy is a key component of out-of-distribution detection, that is, for detected out-of-distribution samples, the model chooses not to make predictions, but to mark them as "untrustworthy" to avoid misleading prediction results.

[0006] At present, some methods for out-of-distribution detection have been proposed, such as uncertainty quantification methods based on probability models, feature representation methods based on deep learning, and methods for identifying out-of-distribution samples by measuring the difference between input samples and training samples in feature space. However, these methods still have certain shortcomings when applied to soil shear strength prediction. Existing out-of-distribution detection methods may have problems such as insufficient detection accuracy, high computational complexity, and difficulty in processing multidimensional features in complex soil characteristics and diverse actual engineering environments, which limits their wide application in engineering.

[0007] In the existing technology, the accuracy of soil undrained shear strength prediction is limited by many factors. In particular, when there is a significant difference between the soil sample and the model training data, the prediction results of the existing data-driven model are often unreliable and difficult to provide a reliable basis for engineering construction. The existing technology has the following problems: Insufficient accuracy in identifying out-of-distribution samples: Existing soil shear strength prediction models are difficult to effectively identify out-of-distribution samples that are significantly different from the training data, resulting in inaccurate prediction results on these samples, thereby increasing safety hazards in engineering construction.

[0008] Lack of a mechanism to handle untrustworthy samples: In existing technologies, models usually still make predictions for out-of-distribution samples, but these predictions may mislead users, thus bringing security and economic risks. Therefore, a rejection mechanism is needed that can refuse to give prediction results when untrustworthy samples are detected.

[0009] Computational complexity and operability issues: The applicability of existing out-of-distribution detection methods in complex engineering environments is limited, mainly because of their high computational complexity and difficulty in real-time detection and rejection of unreliable samples. Therefore, an efficient and suitable out-of-distribution detection method for soil shear strength prediction is needed. Summary of the invention

[0010] In view of the above-mentioned deficiencies in the prior art, the soil undrained shear strength rejection method based on out-of-distribution detection provided by the present invention solves the problems of insufficient accuracy in identifying out-of-distribution samples, lack of a processing mechanism for unreliable samples, as well as computational complexity and operability in the prior art.

[0011] In order to achieve the above-mentioned purpose of the invention, the technical solution adopted by the present invention is: a method for rejecting the undrained shear strength of soil based on out-of-distribution detection, comprising the following steps: S1. Collect soil data, use edge computing technology to preprocess the soil data, and generate preprocessed soil data; S2, dividing the preprocessed soil data into a training set, a validation set and a test set, and constructing a prediction model for undrained soil shear strength; S3, obtaining new soil data, preprocessing the new soil data, and determining whether it is out-of-distribution data through out-of-distribution detection. If so, proceed to S4, if not, proceed to S5; S4, sending prompt information and data processing suggestions to the user through the rejection mechanism, and recording the out-of-distribution data, in response to recording a preset amount of out-of-distribution data within a preset time, updating the parameters of the prediction model, and returning to S3; S5. Input the soil data detected outside the distribution into the prediction model to obtain the prediction result of the undrained shear strength of the soil.

[0012] Further: S1 comprises the following sub-steps: S11. Collect soil data under different regions and environmental conditions. The soil data includes physical property factors, chemical property factors and mechanical property factors; Among them, physical property factors include density and porosity, chemical property factors include pH value and mineral composition, and mechanical property factors include shear strength and compressibility; S12. Use edge computing technology to perform standardized preprocessing on the collected soil data to generate preprocessed soil data.

[0013] Further: In S1, the formula of edge computing technology is specifically:

[0014] In the formula, For bandwidth requirements, is the total amount of data, The amount of data reduced after processing by edge devices, is the transmission time; Overall system response time The specific expression is:

[0015] In the formula, is the computing time of the edge device, is the transmission time, The processing time of the central server.

[0016] Further: S2 comprises the following sub-steps: S21, dividing the preprocessed soil data set into a training set, a validation set and a test set according to a preset ratio; S22, inputting the training set into the initial model, learning the relationship between soil properties and shear strength, optimizing the structure of the initial model, and obtaining a trained initial model; S23. Use the validation set to evaluate the trained initial model to ensure that the model can accurately predict the shear strength of the soil; S24. The test set is used to conduct final verification on the initial model that has passed the evaluation to ensure the stability of the model's performance on different data subsets. The initial model that has passed the evaluation is used as a prediction model for the undrained shear strength of soil.

[0017] The beneficial effect of the above further scheme is that by combining advanced machine learning and statistical detection methods, the present invention can effectively identify out-of-distribution soil samples, thereby improving the reliability of the prediction model in different environments, and further improving the accuracy of out-of-distribution sample identification.

[0018] Further: in said S22, the initial model includes a support vector machine, a random forest and an artificial neural network; The method for optimizing the structure of the initial model is specifically to simplify the structure of the neural network by using pruning technology and quantization technology.

[0019] Further: In S3, the method for performing out-of-distribution detection is specifically: checking whether the input new soil data conforms to the distribution characteristics of the training data, the method comprising: The Gaussian mixture model method calculates the likelihood probability of new soil data through the Gaussian mixture model, and determines whether the input new soil data conforms to the distribution characteristics of the training data according to the likelihood probability; The isolation forest method calculates the anomaly score of the new soil data by randomly splitting the depth of isolated features, and determines whether the input new soil data conforms to the distribution characteristics of the training data based on the anomaly score; Deep learning methods, such as generative adversarial networks or variational autoencoders, are used to evaluate whether new soil data conform to the distribution characteristics of training data.

[0020] The beneficial effect of the above further scheme is that the present invention reduces the computational complexity by optimizing the out-of-distribution detection algorithm, so that it has high real-time and operability in engineering applications. The method is not only suitable for soil sample prediction in laboratory environments, but can also run efficiently in complex actual engineering scenarios, reducing computational complexity and improving practicality.

[0021] Further: in said S4, the prompt information includes: Text prompt information provides users with detailed text instructions to inform them of the risk sources of the data, the errors caused, and the impact on model accuracy; Graphical warnings, with color coding to highlight abnormal data points in the chart; Risk map information, generating GIS-based risk maps showing high-risk areas or soil data points that do not meet model assumptions; The specific data processing suggestions are: Provide data processing suggestions to users based on the results of out-of-distribution detection, including resampling and correcting data.

[0022] The beneficial effect of the above further solution is that the present invention provides a rejection mechanism that can refuse to give a prediction result when an out-of-distribution sample is detected to ensure safety during the construction process. In this way, users can more clearly identify the credibility of the prediction result, thereby avoiding misleading decisions and achieving rejection of unreliable samples.

[0023] Further: in said S4, the method for updating the parameters of the prediction model is specifically: dynamically updating the prediction model by using incremental learning and transfer learning; Among them, incremental learning uses stochastic gradient descent to incorporate new samples into the existing prediction model and update the prediction model; Transfer learning updates the prediction model by migrating the original prediction model to a new task and fine-tuning it to adapt to the specific scenario.

[0024] Further: in S5, the prediction result of the undrained soil shear strength obtained includes the predicted value and confidence interval of the undrained soil shear strength.

[0025] The beneficial effects of the present invention are as follows: the present invention provides a rejection method for undrained soil shear strength based on out-of-distribution detection, and proposes a systematic rejection strategy for potential risks in the process of predicting undrained soil shear strength to improve the reliability of prediction and engineering safety. Compared with the soil strength prediction using the existing soil shear strength prediction model, the present invention has the following effects: (1) Improved data security: By adopting a rejection mechanism, the present invention can effectively identify out-of-distribution samples and refuse to make predictions on untrustworthy samples, thus avoiding misleading decisions on abnormal samples and reducing safety hazards caused by incorrect predictions. At the same time, this mechanism makes data processing more robust, helps protect the security of soil sample data, and ensures the reliability of engineering decisions.

[0026] (2) Network resource saving: The present invention uses edge computing technology to transfer some computing tasks to edge devices, thereby reducing dependence on central servers. This not only reduces the bandwidth requirements for network transmission, but also reduces the consumption of cloud computing resources, thereby saving network resources. In addition, the introduction of edge computing also significantly improves the real-time performance of the model, which is particularly suitable for the needs of rapid on-site detection.

[0027] (3) Enhanced prediction reliability and adaptability: Through multi-model fusion technology, the present invention integrates multiple different types of machine learning models to enhance the adaptability to different soil characteristics. Compared with a single model, the fusion model can provide more accurate and reliable prediction results when facing various complex soil samples, thereby greatly improving the prediction ability of the model, especially showing good stability when the environmental conditions change greatly.

[0028] (4) Reduced computational complexity: Compared with the prior art, the present invention reduces the computational complexity of the model by optimizing the structure of the deep neural network. This makes the model more operational in practical engineering applications, can run efficiently on devices with limited resources, reduces dependence on high-performance computing devices, and reduces deployment and operation costs.

[0029] (5) Improved user experience: The present invention provides a graphical user interface that allows users to intuitively view detection and rejection results, and interpret model outputs in combination with confidence information. Compared with existing black box models, the visualization tool of the present invention greatly enhances users’ understanding and trust in model predictions, which helps to make construction decisions more scientific and transparent.

[0030] With the above advantages, the present invention is significantly superior to the existing technology in the prediction of undrained soil shear strength, especially in terms of data security, network resource saving, prediction reliability and adaptability, as well as calculation complexity and user experience, providing strong support for the safe and efficient construction of civil engineering projects. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flow chart of the method for rejecting the undrained shear strength of soil based on out-of-distribution detection of the present invention. DETAILED DESCRIPTION

[0032] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0033] like Figure 1 As shown, in one embodiment of the present invention, a method for identifying undrained shear strength of soil based on out-of-distribution detection comprises the following steps: S1. Collect soil data, use edge computing technology to preprocess the soil data, and generate preprocessed soil data; S2, dividing the preprocessed soil data into a training set, a validation set and a test set, and constructing a prediction model for undrained soil shear strength; S3, obtaining new soil data, preprocessing the new soil data, and determining whether it is out-of-distribution data through out-of-distribution detection. If so, proceed to S4, if not, proceed to S5; S4, sending prompt information and data processing suggestions to the user through the rejection mechanism, and recording the out-of-distribution data, in response to recording a preset amount of out-of-distribution data within a preset time, updating the parameters of the prediction model, and returning to S3; S5. Input the soil data detected outside the distribution into the prediction model to obtain the prediction result of the undrained shear strength of the soil.

[0034] In this embodiment, the basic idea of ​​the present invention is to propose a systematic rejection strategy for the potential risks in the process of soil undrained shear strength prediction through out-of-distribution detection (OOD) technology, so as to improve the reliability of prediction and engineering safety. Specifically, most of the existing soil undrained shear strength prediction models adopt a data-driven approach to train historical observation data, so as to achieve good prediction results on existing data sets. However, in practical applications, the physical properties of the soil may be contrary to the preset assumptions of the model due to changes in regional and environmental conditions, resulting in the model's prediction results deviating from the actual situation, and even causing serious misjudgments and safety hazards. The present invention introduces out-of-distribution detection technology to automatically identify input data that may be inconsistent with the model assumptions, and refuses to directly predict these out-of-distribution data, thereby effectively avoiding major errors caused by data distribution mismatch and ensuring the safety of engineering implementation. Through this method, the present invention not only improves the accuracy of soil strength prediction, but also provides users with an intuitive risk warning mechanism, enhancing the reliability and safety of engineering decision-making.

[0035] The S1 comprises the following sub-steps: S11. Collect soil data under different regions and environmental conditions. The soil data includes physical property factors, chemical property factors and mechanical property factors; Among them, physical property factors include density and porosity, chemical property factors include pH value and mineral composition, and mechanical property factors include shear strength and compressibility; S12. Use edge computing technology to perform standardized preprocessing on the collected soil data to generate preprocessed soil data.

[0036] In this embodiment, the present invention collects soil data under different geographical areas and environmental conditions through a systematic soil sample collection module, covering the physical, chemical and mechanical properties of the soil, such as density, porosity, pH value, shear strength, etc., to ensure the extensiveness and representativeness of data collection, such as year-round sampling across different seasons and multi-point sampling covering different soil types in the target area. Representative samples need to be reasonably selected during the collection process to include typical geological characteristics and environmental conditions to ensure that the samples are sufficiently representative and diverse, thereby supporting the robustness and generalization ability of the model and avoiding distortion of model predictions due to incomplete local data.

[0037] The collected raw data is standardized, such as mean normalization and Z-score normalization, to eliminate the impact of noise and outliers on model training. This step ensures the consistency and high quality of the data and provides stable input for subsequent modeling. The collected data is standardized using the formula:

[0038] In the formula, is the original data, is the sample mean, and are the maximum and minimum values ​​of the samples respectively. Scores were normalized to eliminate the impact of noise and outliers to the greatest extent possible, ensuring data quality and consistency.

[0039]

[0040] In the formula, is the original data, is the sample mean, is the sample standard deviation. During the collection process, the soil profile structure characteristics, especially the stratified sampling data at different depths, need to be recorded in detail. This information plays an important guiding role in describing soil properties and predicting shear strength in subsequent modeling. In addition, the present invention uses advanced sensors and telemetry technology to improve the accuracy and reliability of data collection, enhance the model's understanding of soil physical properties, and provide high-quality input for subsequent modeling. To ensure the comprehensiveness of the data, repeated sampling is required to verify the consistency of the data and reduce the impact of single sampling errors.

[0041] The S2 comprises the following sub-steps: S21, dividing the preprocessed soil data set into a training set, a validation set and a test set according to a preset ratio; In this embodiment, the present invention is divided into a training set, a validation set and a test set according to 7:2:1.

[0042] S22, inputting the training set into the initial model, learning the relationship between soil properties and shear strength, optimizing the structure of the initial model, and obtaining a trained initial model; S23. Use the validation set to evaluate the trained initial model to ensure that the model can accurately predict the shear strength of the soil; S24. The test set is used to conduct final verification on the initial model that has passed the evaluation to ensure the stability of the model's performance on different data subsets. The initial model that has passed the evaluation is used as a prediction model for the undrained shear strength of soil.

[0043] In S22, the initial model includes a support vector machine, a random forest and an artificial neural network; In this embodiment, the present invention uses a variety of advanced machine learning and deep learning methods to build a prediction model, including a support vector machine (SVM), whose objective function is defined as:

[0044] In the formula, is the weight vector, is the bias, is the slack variable, is a regularization parameter used to control the trade-off between model complexity and misclassification.

[0045] Random Forest (RF) improves the robustness and accuracy of prediction by integrating multiple decision trees, and the construction of decision trees is based on the maximization criterion of information gain or Gini index.

[0046] Artificial neural network (ANN) optimizes network parameters through back propagation algorithm (BP). The loss function of neural network is mean square error (MSE), which measures the average difference between the predicted value and the true value, and aims to minimize the prediction error of the model, thereby improving the prediction accuracy. The role of mean square error is to penalize large prediction errors, prompting the model to continuously adjust parameters during training to achieve the best fit. MSE The specific expression is:

[0047] In the formula, is the true value, is the predicted value, is the number of samples. To improve the generalization ability of the model, -fold cross validation strategy to evaluate the model, usually taking To ensure the balance of model performance among different data subsets, thus effectively reducing the risk of overfitting. In addition, hyperparameter tuning methods such as Bayesian optimization and grid search are used to optimize model parameters, such as learning rate and the regularization coefficient , thereby further improving the prediction accuracy and stability of the model. To enhance the interpretability of the model, feature importance analysis (such as SHAP value) is also used to evaluate the impact of various soil characteristics on the prediction results, so that users can better understand the decision basis of the model.

[0048] The method for optimizing the structure of the initial model is specifically to simplify the structure of the neural network by using pruning technology and quantization technology.

[0049] In this embodiment, the present invention reduces the computational complexity of the model by optimizing the structure of the deep neural network. Specifically, pruning and quantization techniques are used to simplify the structure of the neural network. Pruning reduces redundant connections and parameter quantities. Quantization converts weights from floating point representations to representations with lower digits, such as 8-bit, thereby reducing the amount of computation and memory usage. The target loss of optimization during pruning The specific expression is:

[0050] In the formula, is the original loss, is the coefficient of the regularization term, is the parameter weight; During the pruning process, the optimization goal of the network model is not only to ensure the original loss of the pruned model on the task Keep it as unchanged as possible, and also adjust the parameter weights in the network By introducing This term can promote the sparseness of weights, so that more small weights tend to zero, so that they can be removed during pruning. The coefficient of this part of the regularization term is The strength of regularization is controlled. The larger the value, the more obvious the sparsification effect. Therefore, the optimization direction of this objective function can balance the original performance of the model and the effect of sparsification weights after pruning.

[0051] In S3, the method for performing out-of-distribution detection is specifically to check whether the input new soil data conforms to the distribution characteristics of the training data. The method includes: The Gaussian mixture model method calculates the likelihood probability of new soil data through the Gaussian mixture model, and determines whether the input new soil data conforms to the distribution characteristics of the training data according to the likelihood probability; The isolation forest method calculates the anomaly score of the new soil data by randomly splitting the depth of isolated features, and determines whether the input new soil data conforms to the distribution characteristics of the training data based on the anomaly score; Deep learning methods, such as generative adversarial networks or variational autoencoders, are used to evaluate whether new soil data conform to the distribution characteristics of training data.

[0052] In this embodiment, in order to deal with the heterogeneity problem of soil data in practical applications, the present invention designs an out-of-distribution detection module to determine whether the new input data conforms to the distribution characteristics of the training data. The out-of-distribution detection module is implemented by the following methods: Gaussian mixture model (GMM): estimates the probability density function of the training sample based on the Gaussian distribution, by calculating the input data point The likelihood probability and compare it with the threshold Compare to determine whether it is an out-of-distribution sample. The expression of the likelihood probability of the Gaussian mixture model is as follows:

[0053] In the formula, is the mixing coefficient, For the The mean of a Gaussian distribution, K is the total number of Gaussian distributions, For the The covariance matrix of a Gaussian distribution, For is the mean, is the covariance matrix of the normal distribution at the point x The probability density at .

[0054] Isolation forests calculate the anomaly score of data points by randomly splitting the depth of isolated features, thereby determining whether the data is an abnormal sample. The advantage of isolation forests in out-of-distribution detection lies in its high efficiency and the lack of explicit modeling of data, which is particularly suitable for processing multidimensional data.

[0055] The deep learning method uses the softmax probability output by the neural network to evaluate the model's confidence in the new input data. When the confidence is lower than a certain threshold, the data is considered to be an out-of-distribution sample. In addition, variational autoencoders (VAE) and generative adversarial networks (GAN) are used to further improve the accuracy and robustness of out-of-distribution detection. The autoencoder maximizes the log-likelihood of the input data. The probability representation of the sample in a specific distribution is obtained, and the generative adversarial network effectively characterizes the complex data distribution through adversarial training of the generator and the discriminator. The loss function of the autoencoder includes reconstruction error and KL divergence, and the loss function of VAE is The specific expression is:

[0056] In the formula, is the posterior probability, is the prior probability, For the generated reconstructed data, is the reconstruction error, is the KL divergence; In S4, the prompt information includes: Text prompt information provides users with detailed text descriptions of the data’s risk sources, resulting errors, and impact on model accuracy, helping users make informed decisions; Graphical warning information, anomaly data points are marked in the chart through color coding, so that users can intuitively perceive data anomalies. The color selection and marking scheme of the graphic are optimized to ensure that users can respond to data anomalies in the shortest time; Risk map information, generate GIS-based risk maps, display high-risk areas or soil data points that do not meet model assumptions; generate visual maps based on geographic information systems (GIS), mark the distribution of soil samples and their differences from standard distribution, and support regional risk analysis. High-risk areas are intuitively displayed in the form of heat maps, providing engineers with more sophisticated risk assessment tools. The system also provides detailed log reports to record the details of each out-of-distribution detection and rejection process, including the number of out-of-distribution samples detected, reasons for rejection, and treatment suggestions, providing data support for subsequent engineering verification and model improvement.

[0057] The specific data processing suggestions are: Provide data processing suggestions to users based on the results of out-of-distribution detection, including resampling and correcting data.

[0058] In this embodiment, for the detected out-of-distribution data, the present invention adopts a rejection mechanism to refuse to directly predict these data. The rejection mechanism includes various types of prompt information to help users better understand whether the data is suitable for the existing model.

[0059] In S4, the method for updating the parameters of the prediction model is specifically: dynamically updating the prediction model by using incremental learning and transfer learning; Among them, incremental learning uses stochastic gradient descent to incorporate new samples into the existing prediction model and update the prediction model. The expression of updating the parameters of the prediction model using incremental learning is specifically:

[0060] In the formula, are the parameters of the updated prediction model, is the parameter of the prediction model before updating, is the learning rate, is the gradient of the loss function; Transfer learning migrates the original prediction model to the new task, updates the prediction model through fine-tuning to adapt to the specific scenario, and improves learning efficiency and model performance. This adaptive update mechanism enables the model to cope with dynamic environments and changing data characteristics, maintaining a high level of prediction accuracy and robustness. In addition, this embodiment introduces an online verification mechanism to dynamically evaluate the performance of the model on new data and trigger a model update when the performance is below a threshold, thereby ensuring continuous optimization and stable performance of the model.

[0061] In this embodiment, the present invention uses edge computing technology to sink some computing tasks to edge devices, thereby reducing dependence on central servers. Specifically, edge computing technology significantly reduces the need to upload data to the central server by performing soil data preprocessing, dividing soil data, and predicting the model output of soil undrained shear strength on edge devices. The formula can be expressed as:

[0062] In the formula, For bandwidth requirements, is the total amount of data, The amount of data reduced after processing by edge devices, is the transmission time. Through edge computing, Increase, The bandwidth requirement is reduced. In addition, the edge device directly performs some inference calculations, which significantly reduces the response time of the system. Specifically, the overall response time of the system The specific expression is:

[0063] In the formula, is the computing time of the edge device, is the transmission time, is the processing time of the central server. Optimize and reduce , which significantly improves the overall real-time responsiveness and is particularly suitable for the needs of rapid on-site detection.

[0064] In S5, the prediction result of the undrained soil shear strength is obtained, including the predicted value and confidence interval of the undrained soil shear strength. The specific expression is:

[0065] In the formula, is the predicted value, is the critical value of the normal distribution, is the standard deviation, is the sample size.

[0066] In this embodiment, the present invention also provides warnings and detailed explanations of out-of-distribution data. For example, when it is detected that the mineral composition of the soil in a certain area is significantly different from the training data set, the system will issue a warning, pointing out that the data in this area may cause large model prediction errors, and recommending that the user resample or perform model correction. This allows users to have an in-depth understanding of the limitations and application risks of the model. In addition, the system supports multi-language functions, which is convenient for international users to use and enhances the global applicability and promotion value of the system. The system interface also supports customization, and users can adjust the display content and method of the interface according to specific engineering requirements, so that the system can adapt to different application scenarios. In addition, in order to facilitate the comprehensive management of engineering projects, data logging and report generation functions are integrated into the system. Users can trace back and verify the entire process through automatically generated reports to ensure that each step is well documented.

[0067] The beneficial effects of the present invention are as follows: the present invention provides a rejection method for undrained soil shear strength based on out-of-distribution detection, and proposes a systematic rejection strategy for potential risks in the process of predicting undrained soil shear strength to improve the reliability of prediction and engineering safety. Compared with the soil strength prediction using the existing soil shear strength prediction model, the present invention has the following effects: (1) Improved data security: By adopting a rejection mechanism, the present invention can effectively identify out-of-distribution samples and refuse to make predictions on untrustworthy samples, thus avoiding misleading decisions on abnormal samples and reducing safety hazards caused by incorrect predictions. At the same time, this mechanism makes data processing more robust, helps protect the security of soil sample data, and ensures the reliability of engineering decisions.

[0068] (2) Network resource saving: The present invention uses edge computing technology to transfer some computing tasks to edge devices, thereby reducing dependence on central servers. This not only reduces the bandwidth requirements for network transmission, but also reduces the consumption of cloud computing resources, thereby saving network resources. In addition, the introduction of edge computing also significantly improves the real-time performance of the model, which is particularly suitable for the needs of rapid on-site detection.

[0069] (3) Enhanced prediction reliability and adaptability: Through multi-model fusion technology, the present invention integrates multiple different types of machine learning models to enhance the adaptability to different soil characteristics. Compared with a single model, the fusion model can provide more accurate and reliable prediction results when facing various complex soil samples, thereby greatly improving the prediction ability of the model, especially showing good stability when the environmental conditions change greatly.

[0070] (4) Reduced computational complexity: Compared with the prior art, the present invention reduces the computational complexity of the model by optimizing the structure of the deep neural network. This makes the model more operational in practical engineering applications, can run efficiently on devices with limited resources, reduces dependence on high-performance computing devices, and reduces deployment and operation costs.

[0071] (5) Improved user experience: The present invention provides a graphical user interface that allows users to intuitively view detection and rejection results, and interpret model outputs in combination with confidence information. Compared with existing black box models, the visualization tool of the present invention greatly enhances users’ understanding and trust in model predictions, which helps to make construction decisions more scientific and transparent.

[0072] With the above advantages, the present invention is significantly superior to the existing technology in the prediction of undrained soil shear strength, especially in terms of data security, network resource saving, prediction reliability and adaptability, as well as calculation complexity and user experience, providing strong support for the safe and efficient construction of civil engineering projects.

[0073] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.

Claims

1. A method for identifying undrained shear strength of soil based on out-of-distribution detection, characterized in that: The following steps are involved: S1. Collect soil data, use edge computing technology to preprocess the soil data, and generate preprocessed soil data; S2, dividing the preprocessed soil data into a training set, a validation set and a test set, and constructing a prediction model for undrained soil shear strength; S3, obtaining new soil data, preprocessing the new soil data, and determining whether it is out-of-distribution data through out-of-distribution detection. If so, proceed to S4, if not, proceed to S5; S4, sending prompt information and data processing suggestions to the user through the rejection mechanism, and recording the out-of-distribution data, in response to recording a preset amount of out-of-distribution data within a preset time, updating the parameters of the prediction model, and returning to S3; S5. Input the soil data detected outside the distribution into the prediction model to obtain the prediction result of the undrained shear strength of the soil.

2. The method for identifying undrained shear strength of soil based on out-of-distribution detection according to claim 1 is characterized in that: The S1 comprises the following sub-steps: S11. Collect soil data under different regions and environmental conditions. The soil data includes physical property factors, chemical property factors and mechanical property factors; Among them, physical property factors include density and porosity, chemical property factors include pH value and mineral composition, and mechanical property factors include shear strength and compressibility; S12. Use edge computing technology to perform standardized preprocessing on the collected soil data to generate preprocessed soil data.

3. The method for identifying undrained shear strength of soil based on out-of-distribution detection according to claim 2 is characterized in that: In S1, the formula of edge computing technology is specifically: In the formula, For bandwidth requirements, is the total amount of data, The amount of data reduced after processing by edge devices, is the transmission time; Overall system response time The specific expression is: In the formula, is the computing time of the edge device, is the transmission time, The processing time of the central server.

4. The method for identifying undrained soil shear strength based on out-of-distribution detection according to claim 1 is characterized in that: The S2 comprises the following sub-steps: S21, dividing the preprocessed soil data set into a training set, a validation set and a test set according to a preset ratio; S22, inputting the training set into the initial model, learning the relationship between soil properties and shear strength, optimizing the structure of the initial model, and obtaining a trained initial model; S23. Use the validation set to evaluate the trained initial model to ensure that the model can accurately predict the shear strength of the soil; S24. The test set is used to conduct final verification on the initial model that has passed the evaluation to ensure the stability of the model's performance on different data subsets. The initial model that has passed the evaluation is used as a prediction model for the undrained shear strength of soil.

5. The method for identifying undrained soil shear strength based on out-of-distribution detection according to claim 3 is characterized in that: In S22, the initial model includes a support vector machine, a random forest and an artificial neural network; The method for optimizing the structure of the initial model is specifically to simplify the structure of the neural network by using pruning technology and quantization technology.

6. The method for identifying undrained soil shear strength based on out-of-distribution detection according to claim 1, characterized in that: In S3, the method for performing out-of-distribution detection is specifically to check whether the input new soil data conforms to the distribution characteristics of the training data. The method includes: The Gaussian mixture model method calculates the likelihood probability of new soil data through the Gaussian mixture model, and determines whether the input new soil data conforms to the distribution characteristics of the training data according to the likelihood probability; The isolation forest method calculates the anomaly score of the new soil data by randomly splitting the depth of isolated features, and determines whether the input new soil data conforms to the distribution characteristics of the training data based on the anomaly score; Deep learning methods, such as generative adversarial networks or variational autoencoders, are used to evaluate whether new soil data conform to the distribution characteristics of training data.

7. The method for identifying undrained shear strength of soil based on out-of-distribution detection according to claim 1, characterized in that: In S4, the prompt information includes: Text prompt information provides users with detailed text instructions to inform them of the risk sources of the data, the errors caused, and the impact on model accuracy; Graphical warnings, with color coding to highlight abnormal data points in the chart; Risk map information, generating GIS-based risk maps showing high-risk areas or soil data points that do not meet model assumptions; The specific data processing suggestions are: Provide data processing suggestions to users based on the results of out-of-distribution detection, including resampling and correcting data.

8. The method for identifying undrained soil shear strength based on out-of-distribution detection according to claim 1, characterized in that: In S4, the method for updating the parameters of the prediction model is specifically: dynamically updating the prediction model using incremental learning and transfer learning; Among them, incremental learning uses stochastic gradient descent to incorporate new samples into the existing prediction model and update the prediction model; Transfer learning updates the prediction model by transferring the original prediction model to the new task and fine-tuning it to adapt to the specific scenario.

9. The method for identifying undrained soil shear strength based on out-of-distribution detection according to claim 1, characterized in that: In S5, the prediction result of the undrained soil shear strength is obtained, including the predicted value and confidence interval of the undrained soil shear strength.