Sandy soil liquefaction discrimination method based on CPTU and weighted nonlinear similarity

By constructing a weighted nonlinear similarity and uncertainty quantification method based on CPTU data, the problems of prediction accuracy and transparency in sand liquefaction discrimination were solved, achieving high-precision and interpretable discrimination results, and improving the credibility of the model and the safety of engineering applications.

CN120974283AActive Publication Date: 2025-11-18OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202511500900.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing methods for identifying sand liquefaction have limitations in terms of insufficient prediction accuracy, opaque model decision-making process, and lack of uncertainty assessment of results. Traditional empirical methods are highly subjective, while machine learning models rely on numerical simulation and are difficult to interpret due to their black-box nature.

Method used

We employ a weighted nonlinear similarity method based on CPTU data. By constructing prototype vectors and feature weight matrices, and combining radial basis kernel functions and information entropy theory, we achieve high-precision discrimination and uncertainty quantification, providing interpretability and credibility.

Benefits of technology

It improves the accuracy and interpretability of sand liquefaction identification, quantifies the uncertainty of the identification results, and enhances the reliability of the model and the safety of engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974283A_ABST
    Figure CN120974283A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of geotechnical engineering, and particularly relates to a sand liquefaction discrimination method based on CPTU and weighted nonlinear similarity. Comprising the following steps: collecting sand CPTU data containing liquefaction and non-liquefaction cases, and calculating a cyclic stress ratio (CSR) as a derivative feature; preprocessing and standardizing the data and dividing the data into a training set and a test set; respectively calculating prototype vectors of liquefaction and non-liquefaction categories based on the training set; by analyzing the influence of disturbance of each characteristic value on the similarity of the sample and the prototype vector, calculating a characteristic weight to quantify the importance of the characteristic weight; constructing a weighted radial basis kernel function fused with feature weights, calculating the nonlinear similarity between the to-be-tested sample and the two types of prototype vectors, and converting the nonlinear similarity into a liquefaction probability; and based on the confidence coefficient of the information entropy quantification prediction result, uncertainty evaluation is realized. The method has high discrimination precision, strong interpretability and reliable uncertainty quantification capability, and effectively overcomes the defects of strong subjectivity and black box of a machine learning model in a traditional experience method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of sand liquefaction and artificial intelligence, specifically to a method for determining the probability of sand liquefaction. Background Technology

[0002] Soil liquefaction is one of the major geological hazards induced by earthquakes. It refers to the phenomenon where saturated, loose sand, under seismic loads, experiences a sharp increase in pore water pressure and a significant decrease in effective stress, leading to the soil losing its shear strength and transforming from a solid to a liquid state. Soil liquefaction areas typically exhibit phenomena such as sand boils and water seepage, ground subsidence, uneven settlement of buildings, tilting or even complete collapse, and the uplift of underground facilities, causing widespread damage to land and existing infrastructure, resulting in unpredictable losses. Therefore, how to quickly and accurately determine the liquefaction potential of soil has become one of the most challenging problems in the field of geotechnical earthquake engineering, and it has crucial practical significance for site seismic safety assessment, site selection for major projects, and seismic fortification.

[0003] Currently, methods for identifying sand liquefaction mainly fall into two categories. The first is the traditional empirical method, represented by the Seed simplified method. This method uses the ratio of cyclic shear stress ratio (CSR) to soil liquefaction resistance ratio (CRR) as the liquefaction safety factor, and establishes empirical curves to identify sand liquefaction. CSR is usually calculated based on theoretical formulas or numerical simulations, while CRR is determined through physical experiments or in-situ tests. The advantages of this method are its clear physical concepts and simple calculations, making it widely accepted in the geotechnical and earthquake engineering community. The disadvantages are: 1) empirical charts are based on specific historical earthquake data, limiting their generalization ability and accuracy; 2) they are usually based on a single indicator (such as standard penetration test blow count). N Cone tip resistance q c 1) It fails to make full use of the multi-parameter data provided by modern pore pressure static cone penetration test (CPTu); 2) The discrimination process relies on manual map lookup, which has the limitations of strong subjectivity and difficulty in automation.

[0004] Secondly, data-driven methods, such as machine learning algorithms like logistic regression, support vector machines (SVM), random forests, and XGBoost, are widely used for liquefaction detection. These methods can learn complex nonlinear patterns from large amounts of data and fuse multiple features for detection, often achieving higher accuracy than empirical methods. For example, invention patent application CN202510231063.6 discloses a method and system for detecting liquefaction of sandy soil in geotechnical engineering during earthquakes. This method, based on a combination of numerical simulation and field test verification, has a certain physical basis, but still suffers from the following shortcomings: 1) It heavily relies on numerical models, which are based on physical assumptions and simplifications and deviate from actual complex site conditions. Model correction requires field data, but the parameter calibration process is highly subjective and cannot fully capture the nonlinear behavior of the soil, affecting the accuracy of the judgment.

[0005] 2) Numerical simulations rely entirely on "black box" equation solutions, making it difficult to trace the decision-making process and mechanisms. Engineers cannot understand why a specific area is identified as having a high risk of liquefaction, which reduces the credibility of the model.

[0006] 3) Ignoring uncertainty and only outputting deterministic results (such as pore pressure ratio and shear strain) makes it impossible to quantify the confidence level of the prediction results, which is a fatal shortcoming for high-risk engineering decisions. Summary of the Invention

[0007] To address the limitations of current methods for identifying sand liquefaction, such as insufficient prediction accuracy, opaque model decision-making processes, and a lack of uncertainty assessment, this invention proposes a novel method for identifying sand liquefaction. This method, by integrating multi-source features and machine learning techniques, balances high-precision prediction, strong interpretability, and complete uncertainty quantification, aiming to effectively bridge the gap between sand liquefaction theory, in-situ test data, and practical applications of machine learning.

[0008] The objective of this invention can be achieved through the following technical solutions: A method for determining the liquefaction of sandy soil includes the following steps: S1. Collect CPTU data of sand containing liquefied and non-liquefied soils, and calculate the cyclic stress ratio (CSR) as a derived feature. S2. Preprocess and standardize the data, and divide it into training and testing sets; S3. Based on the training set, calculate the prototype vectors for the liquefaction category and the non-liquefaction category respectively; S4. Based on the training set and the prototype vector, the weight of each feature is calculated by analyzing the impact of perturbation of each feature value on the similarity between the sample and the prototype vector, so as to quantify its importance to liquefaction discrimination. S5. Construct a weighted radial basis kernel function that incorporates the weights, calculate the nonlinear similarity between the test sample and the liquefied and non-liquefied prototype vectors, and convert it into liquefaction probability; S6. Uncertainty Quantification: Based on the liquefaction probability output, the confidence level of each prediction is calculated using information entropy to achieve a quantitative assessment of the uncertainty of the judgment result.

[0009] Further, in step S5, the expression for the weighted radial basis kernel function is: Where ⊙ represents element-wise multiplication, γ is the bandwidth parameter of the RBF kernel, w is the feature weight calculated by S4, X is the sample to be tested, and P is the prototype vector.

[0010] Furthermore, in step S5, the probability of liquefaction is converted using the Softmax function: P (y = non-liquefied - X) = 1 - P (y = liquefaction - X) In the formula, y represents the binary classification result of sand liquefaction, that is, the two results corresponding to sample X: liquefaction or non-liquefaction. This is the feature weight vector for the liquefaction category. This is the feature weight vector for the non-liquefiable category. This is the prototype vector for the liquefaction category. This is the prototype vector for the non-liquefied category.

[0011] Furthermore, in step S3, the formula for calculating the prototype vector is as follows: P 液化 = P 非液化 = Where P is the prototype vector. and These are the number of samples in two categories. It is the standardized sample feature vector.

[0012] Furthermore, in step S4, the specific steps for calculating the weight corresponding to each feature include: a. Apply a small perturbation Δ to the j-th feature; b. For each sample within the category Calculate the cosine similarity change ΔSimlarity between feature j and the corresponding class prototype vector P before and after feature j is perturbed. c. Calculate the average similarity change ΔSimlarity-P for all samples on feature j; d. According to the formula Calculate the final weight of feature j, where α is the weight amplification factor.

[0013] Furthermore, in step S1, the formula for calculating the liquefaction-derived parameter cyclic stress ratio (CSR) is as follows: CSR=0.65 r d ( () ) Furthermore, after step S6, a hyperparameter optimization step is also included: using a Bayesian optimization algorithm, with the average AUC value of K-fold cross-validation as the performance index, the optimal combination of kernel width parameter γ and weight amplification factor α is automatically searched and determined.

[0014] The beneficial effects of this invention are: 1. This invention effectively addresses the inherent shortcomings of traditional empirical methods, such as strong subjectivity, and black-box models in machine learning. Compared to empirical methods based on the Seed formula, this invention possesses stronger nonlinear mapping capabilities and higher discrimination accuracy. Compared to existing machine learning models, this invention provides clear physical meaning and decision-making logic through prototype vectors and feature weight matrices, greatly enhancing the interpretability and credibility of the model and overcoming the difficulty of trusting black-box models in engineering applications.

[0015] 2. This invention proposes a nonlinear similarity measurement method based on a weight matrix. By introducing a feature weight matrix and deeply fusing it into a radial basis function (RBF) kernel, a transformation from simple "directional similarity" (cosine similarity) to "weighted spatial proximity" is achieved. This core innovation enables the model to accurately capture the contribution of different features to the liquefaction discrimination difference while maintaining the powerful performance of nonlinear classification, thus achieving the best balance between interpretability and accuracy.

[0016] 3. This invention is the first to quantify the uncertainty of sand liquefaction discrimination results. It innovatively introduces information entropy theory to transform the probability distribution of the model output into an intuitive confidence index. This index can effectively identify and quantify the uncertainty of prediction results, providing crucial risk warning information and supporting the differentiation between deterministic predictions and questionable judgments in major engineering decisions, significantly improving the safety and reliability of AI-based discrimination methods. Attached Figure Description

[0017] The invention will now be further described with reference to the accompanying drawings.

[0018] Figure 1 This is a flowchart of the sand liquefaction discrimination method based on weighted nonlinear similarity and uncertainty quantification of the present invention; Figure 2 This invention provides 3D feature distribution of CPTU data for liquefied and non-liquefied sand. Figure 1 ; Figure 3 This invention provides 3D feature distribution of CPTU data for liquefied and non-liquefied sand. Figure 2 ; Figure 4 This invention provides 3D feature distribution of CPTU data for liquefied and non-liquefied sand. Figure 3 ; Figure 5This is a heatmap of the feature weights of CPTU data for liquefied and non-liquefied sand in this invention; Figure 6 This is the Bayesian algorithm optimization process of the present invention; Figure 7 This is the confusion matrix result of the test set of this invention; Figure 8 This invention provides a test map of the probability distribution of sand liquefaction. Figure 9 This is a confidence analysis diagram of the probability of soil liquefaction in the test set of this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] refer to Figure 1 This invention proposes a method for sand liquefaction discrimination based on weighted nonlinear similarity and uncertainty quantification. This method constructs prototype vectors for sand liquefaction and non-liquefaction categories based on pore pressure static cone penetration test (CPTU) data. It uses a unit increment method to calculate the global feature weight matrix, quantifying the contribution of each feature to the discrimination result. A radial basis function (RBF) kernel function that fuses feature weights is designed, transforming the similarity measurement between the test sample and the prototype vector from "directional similarity" to "weighted spatial proximity." A Bayesian optimization algorithm is used to automatically search for optimal kernel parameters and weight amplification factors, aiming to maximize the AUC value of the five-fold cross-validation. The liquefaction probability is output through the Softmax function, and information entropy theory is innovatively introduced to quantify the prediction confidence level, significantly improving the interpretability, accuracy, and reliability of sand liquefaction discrimination. Specifically, it includes: I. Collect CPTU data on soil liquefaction in the earthquake zone. Based on the Seed simplified formula, calculate the liquefaction-derived parameter, cyclic stress ratio (CSR), as follows: CSR=0.65 r d ( () ) II. Data Preprocessing and Z-score Standardization: To eliminate the influence of different feature dimensions and orders of magnitude, making model training more stable, the preprocessed data is divided into training and test sets. Based on the training set, the feature prototype vector for each soil type is calculated, see... Figure 2 , Figure 3 and Figure 4 .

[0021] The data preprocessing and Z-score standardization formulas are as follows: X scaled = in, μ It is the characteristic mean. σ It is the characteristic standard deviation.

[0022] III. Based on the training set, calculate the prototype vector and find the "center" representative of each class (liquefied / non-liquefied) in the feature space, i.e., the prototype vector P, see [link to documentation]. Figure 5 The calculation formula is as follows: P 液化 = P 非液化 = Where P is the prototype vector. and These are the number of samples in two categories. It is the standardized sample feature vector. This prototype represents the statistical center of this category in the training dataset. It is a representative concept supported by actual data, rather than a mathematical abstraction, which makes the entire model's decision-making process very easy to understand and interpret.

[0023] Furthermore, this invention directly uses the "arithmetic mean of the feature vectors of all samples within the liquefied and non-liquefied ranges" as the prototype vector. The calculation formula is concise and clear. The mean calculation process is simple, requires no iteration, has high computational efficiency, and produces stable results, unlike some clustering algorithms that get stuck in local optima or are sensitive to initial values.

[0024] IV. Based on cosine similarity, feature weights and weighting factors are calculated to quantify the importance or sensitivity of each feature in distinguishing between two types of liquefaction. The higher the weight, the greater the impact of a small change in the feature value on the similarity between the sample and the prototype, i.e., the more important the feature.

[0025] The steps for calculating feature weights and weight factors based on cosine similarity are as follows: a. Unit increment: Apply a small perturbation Δ to the j-th feature (Δ in the algorithm) j =0.1); b. Calculate similarity change: For each sample within both classes... Calculate the cosine similarity change and similarity change between the vector P and the prototype vector P before and after the perturbation. Simlarity P) = Δ Simlarity =Simlarity( Δ, P) - Simlarity ( (P) c. Calculate the average change Δ Simlarity-P : Calculate the average of the similarity changes of all samples on the j-th feature; d. Determine the weights: Amplify the average change and map it to a reasonable range to obtain the final weight calculation function for this feature. ; in, It is an amplification factor.

[0026] The weight amplification factor α is a global optimization parameter used to uniformly control the scaling magnitude and degree of difference of all feature importance weights. The liquefied prototype weight vector and the non-liquefied prototype weight vector are calculated and generated by the same amplification factor α based on the similarity changes of their respective intra-class samples, thereby ensuring the comparability of the two weight systems and simplifying the hyperparameter optimization process.

[0027] This invention determines feature weights through a dynamic perturbation strategy, which directly measures the sensitivity of a feature to similarity calculation. Specifically, by applying a small perturbation to each feature and precisely observing the average magnitude of the change in similarity between the sample and the prototype vector, the weight is defined as a function of this change. This mechanism ensures that the model can keenly capture the relative importance of different features in the discrimination, rather than merely their static numerical statistics.

[0028] V. By incorporating the weight matrix into the distance metric, a radial basis function (RBF) kernel function that fuses feature weights is designed to calculate the nonlinear similarity between the test sample and the two prototype vectors, and then converted into liquefaction probability. Details are as follows: a. Design of the weighted RBF kernel function: Where ⊙ represents element-wise multiplication (Hadamard product), γ is the bandwidth parameter of the RBF kernel, w is the feature weight vector calculated by S4, X is the sample to be tested, and P is the prototype vector.

[0029] The value of γ determines the rate at which similarity decays with distance. A larger γ value results in a "narrower" kernel function, meaning only samples very close to the prototype are considered highly similar, leading to a more complex model decision boundary and a higher risk of overfitting. Conversely, a smaller γ value results in a "wider" kernel function, allowing for higher similarity even with samples at greater distances, resulting in a smoother model decision boundary and a higher risk of underfitting.

[0030] Traditional RBF kernel functions use Euclidean distance, which implicitly assumes that all feature dimensions contribute equally to the distance. This is clearly not in line with engineering practice. For example, in sand liquefaction assessment, a small change in the cyclic stress ratio (CSR) may be far more significant than a similar change in depth. Directly modifying the distance formula easily compromises the mathematical properties of the kernel function, and a key challenge lies in how to reasonably and effectively integrate the calculated feature weights (a set of scalars).

[0031] To address the above, this invention introduces the concept of "weighted Euclidean distance," transforming the distance formula from ||XP|| to ||W⊙XP||, thereby achieving scaling of the feature dimension according to feature weights. It successfully integrates the feature weight matrix, which has clear physical meaning and is obtained based on cosine similarity perturbation analysis, into the metric space of the nonlinear RBF kernel function, achieving a fundamental shift from "equal measurement" to "differentiation-sensitive measurement."

[0032] The calculated weighted similarity is an absolute value. How can it be transformed into a probability value between 0 and 1 that intuitively reflects the possibilities of "liquefaction" and "non-liquefaction"? Furthermore, it must be ensured that this probability is comparable and meaningful between the two classes. This invention uses the Softmax function to normalize the two weighted similarities.

[0033] b. Use the Softmax function to output the liquefaction probability.

[0034] P (y = non-liquefied - X) = 1 - P (y = liquefaction - X) The liquefaction probability of a sample depends not only on its proximity to the "liquefaction prototype" but also on its distance from the "non-liquefaction prototype." If a sample is far from both prototypes, Softmax will give an uncertain probability close to 0.5; if a sample is very close to the liquefaction prototype and far from the non-liquefaction prototype, it will give a high liquefaction probability close to 1. Furthermore, optimal feature weight vectors are calculated separately for the "liquefaction" and "non-liquefaction" classes. This means that when measuring the similarity of a sample to the liquefaction prototype, the feature weights most relevant to the liquefaction class are used; when measuring the similarity to the non-liquefaction prototype, the feature weights most relevant to the non-liquefaction class are used. This makes the comparison more fair and accurate.

[0035] VI. Confidence assessment based on information entropy is used to evaluate the reliability of each prediction made by the liquefaction assessment model, increasing the algorithm's credibility and practicality. The details are as follows: a. Information entropy: measures the uncertainty of a probability distribution; H (P) = Where P(liquefied, non-liquefied) is the predicted probability vector, and T=2 b. Normalized confidence level: H max =log2( T The confidence level of 1 represents the maximum entropy for a binary classification problem. The closer the confidence level is to 1, the more certain the prediction is.

[0036] The information entropy H (P) can distinguish between high-confidence deterministic predictions and low-confidence uncertain predictions, avoid blindly trusting unreliable prediction results, and is deeply integrated with the core algorithm (weighted nonlinear similarity + softmax probability function) to form a unified probability-confidence evaluation system, which enhances the interpretability and credibility of the model and realizes the quantitative measurement of the reliability of prediction results.

[0037] The applicant discovered that simply weighting by distance can distort the feature space. If the weights are set improperly, it can disrupt the data's inherent distribution structure, leading to uncontrolled model complexity—either becoming overly complex and overfitting the training data, or becoming too smooth and underfitting, failing to capture key nonlinear boundaries. This invention uses a "Bayesian hyperparameter optimization framework" to simultaneously and automatically optimize two key parameters: the kernel parameter γ and the weight amplification factor α. By placing α and γ within the same optimization framework, the algorithm automatically finds the optimal balance between them. This allows the model to adaptively adjust its "complexity" and "dependence on feature weights," achieving an optimal balance between the high precision potential brought by introducing weights and the model's generalization ability, ensuring that the final discrimination model is both powerful and robust. The specific steps are as follows: a. Optimization definition of the width parameter γ and characteristic weight amplification factor α of the weighted radial basis (RBF) kernel function; b. Find the optimal combination of hyperparameters λ = (γ, α) such that the objective function f ( λ ) minimize; λ =min f ( λ ) λ ∈Λ Where Λ represents the hyperparameter search space; f (λ) is the objective function constructed based on K-fold cross-validation (K=5 in this invention), and its value is the negative average AUC, i.e. f (λ) = −AUC CV ( λ Maximizing the cross-validation AUC is equivalent to minimizing this objective function.

[0038] c. Construct a proxy model and use a Gaussian process to model the unknown objective function. f (λ) Modeling is performed. The Gaussian process provides a prediction of the function value (mean) and a measure of uncertainty (variance) for each point in the entire search space through the combination of hyperparameters of the evaluated points and their objective function values; d. Design the sampling function, using the Expectation Boosting (EI) function. The EI function trades off sampling in regions where the surrogate model's prediction function value is low (good performance) and sampling in regions where the surrogate model's uncertainty is high.

[0039] e. Iterative evaluation and model update: Finally, from all evaluated hyperparameter combinations, the combination with the smallest objective function value, λ=(γ, α), is selected as the global optimal hyperparameter for the final output.

[0040] like Figure 6 As shown, the Bayesian optimization convergence curve reveals that after the initial few iterations, the AUC value rapidly increases to a very high level, demonstrating its "efficiency." As iterations continue, the AUC value stabilizes at a high level with minimal fluctuations, proving that it "precisely" found the performance peak, rather than oscillating back and forth. The AUC value to which the curve finally converges represents the known or achievable top performance for this problem, indicating that it has found a globally optimal or near-optimal solution, rather than getting trapped in a local optimum.

[0041] Figure 7 The calculation results of the test set data of this invention are presented. The confusion matrix shows that the sand liquefaction discrimination method provided by this invention maintains high accuracy while paying particular attention to engineering safety. The low false negative rate (only 4 liquefaction samples were missed) indicates that the method can effectively identify liquefaction risks and avoid engineering safety hazards caused by missed detections. At the same time, the high accuracy ensures the reliability of the discrimination results, providing a reliable technical basis for engineering design. The various calculation indicators of the model are shown in Table 1. In the calculation results of this invention, the core indicator AUC of the model reaches 0.97 and the accuracy rate is over 90%, which strongly proves that it achieves a balance without sacrificing accuracy while introducing interpretable components.

[0042] Table 1 shows the calculation results of the test set data for the method of this invention.

[0043] like Figure 8 , Figure 9As shown, the reliability and stability of the computational model are evaluated using test set data. The output 3D feature distribution map visually presents the spatial distribution patterns of liquefied and non-liquefied samples and prototypes; the output feature weight heatmap quantitatively reveals the contribution and importance ranking of each feature to the discrimination results; the output confidence analysis chart quantifies and displays the uncertainty distribution of the model's prediction results; based on the model performance, rigorous evaluation through cross-validation and independent test sets, and comprehensive quantification using indicators such as accuracy, precision, recall, F1 score, and AUC, a reliable evaluation system with high-precision probability discrimination results and uncertainty metrics is output.

[0044] This invention provides a paradigm-innovative method for identifying sand liquefaction in marine geotechnical exploration, evolving sand liquefaction identification from experience-based to a data-driven "science," thus achieving transparency and intelligence in exploration decision-making. Regarding early-stage design safety, the model outputs liquefaction probability and innovatively introduces a confidence level index. This represents a leap from qualitative "yes / no" judgments to quantitative assessments of "how great the risk is and how certain it is." A tiered strategy is implemented for regions with different probabilities and confidence levels, thereby precisely managing risks and avoiding over-design or safety hazards. In terms of efficiency and cost, this method achieves "second-level" rapid assessment, far exceeding the efficiency of numerical simulations that take weeks. More importantly, it can accurately identify high-risk and high-uncertainty areas based on preliminary exploration results, guiding subsequent "targeted" exploration drilling, solving the most critical problems with the fewest boreholes, significantly optimizing exploration plans, and saving substantial time and economic costs. In terms of the identification method, this invention opens the "black box" of machine learning models. By using prototype vectors and feature weights, the decision-making process becomes transparent, traceable, and in line with engineering intuition, greatly enhancing the credibility and acceptability of artificial intelligence in major engineering projects. Therefore, this invention transforms scarce marine exploration borehole data into highly reliable risk probabilities, serving as a key tool for realizing the digital and intelligent transformation of marine engineering survey and design, and is of strategic significance for ensuring the long-term safety and economic benefits of major national marine infrastructure.

[0045] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A method for determining sand liquefaction based on CPTU and weighted nonlinear similarity, characterized in that, Includes the following steps: S1. Collect CPTU data of sand containing liquefied and non-liquefied soils, and calculate the cyclic stress ratio (CSR) as a derived feature. S2. Preprocess and standardize the data, and divide it into training and testing sets; S3. Based on the training set, calculate the prototype vectors for the liquefaction category and the non-liquefaction category respectively; S4. Based on the training set and the prototype vector, the weight of each feature is calculated by analyzing the impact of perturbation of each feature value on the similarity between the sample and the prototype vector, so as to quantify its importance to liquefaction discrimination. S5. Construct a weighted radial basis kernel function that incorporates the weights, calculate the nonlinear similarity between the test sample and the liquefied and non-liquefied prototype vectors, and convert it into liquefaction probability; S6. Uncertainty Quantification: Based on the liquefaction probability output, the confidence level of each prediction is calculated using information entropy to achieve a quantitative assessment of the uncertainty of the judgment result.

2. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 1, characterized in that, In step S5, the expression for the weighted radial basis kernel function is: K(X、P、w)= Where ⊙ represents element-wise multiplication, γ is the bandwidth parameter of the RBF kernel, w is the feature weight calculated by S4, X is the sample to be tested, and P is the prototype vector.

3. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 2, characterized in that, In step S5, the probability of liquefaction is converted using the Softmax function: P (y = liquefaction - X) = P (y = non-liquefied - X) = 1 - P (y = liquefaction - X) In the formula, y represents the binary classification result of sand liquefaction. This is the feature weight vector for the liquefaction category. This is the feature weight vector for the non-liquefiable category. This is the prototype vector for the liquefaction category. This is the prototype vector for the non-liquefied category.

4. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 1, characterized in that, In step S3, the formula for calculating the prototype vector is as follows: P 液化 = P 非液化 = Where P is the prototype vector. and These are the number of samples in two categories. It is the standardized sample feature vector.

5. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 1, characterized in that, In step S4, the specific steps for calculating the weight corresponding to each feature include: a. Apply a small perturbation Δ to the j-th feature; b. For each sample within the category Calculate the cosine similarity change ΔSimlarity between feature j and the corresponding class prototype vector P before and after feature j is perturbed. c. Calculate the average similarity change ΔSimlarity-P for all samples on feature j; d. According to the formula Calculate the final weight of feature j, where α is the weight amplification factor.

6. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 1, characterized in that, In step S1, the cyclic stress ratio (CSR) of the liquefaction-derived parameter is calculated using the following formula: CSR = 0.65 r d ( () ), where r d Let a be the stress reduction factor. max σ is the peak ground acceleration, g is the gravitational acceleration, and σ is the peak ground acceleration. v and σ v 'These represent the total overburden stress and the effective overburden stress, respectively.

7. The method for determining sand liquefaction based on CPTU and weighted nonlinear similarity according to claim 1, characterized in that, Following step S6, a hyperparameter optimization step is also included: using a Bayesian optimization algorithm, with the average AUC value of K-fold cross-validation as the performance index, the optimal combination of kernel width parameter γ and weight amplification factor α is automatically searched and determined.

Citation Information

Patent Citations

  • Geotechnical engineering earthquake sand liquefaction discrimination method and system

    CN119717009A

  • Soft soil parameter value analysis and correction method and system based on machine learning

    CN118917163A

  • Number correctness evaluation method and system based on machine learning

    CN119939193A

  • Explanatable Transform medical diagnosis method based on prototype learning

    CN120280122A

  • Interconnecting neural network system, method of constructing interconnecting neural network structure, method of constructing self-organized neural network structure, and construction program of them

    JP2005032216A

Cited By

  • Multi-source geological survey data fusion method based on CPTU

    CN121834720A

  • A CPTU-based multi-source geological exploration data fusion method

    CN121834720B