Sparse Data Probability Distribution Modeling Method Applied to Geotechnical Parameter Data Acquisition
Through a combination of kernel density estimation algorithm and genetic algorithm, bandwidth parameters are optimized to obtain the approximate probability distribution of geotechnical parameters, solving the modeling problem of sparse data and multi-peak characteristics in geotechnical engineering, and improving modeling accuracy and accuracy.
Patent Information
- Application Number
- CN202510220902.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In geotechnical engineering, due to the sparse data and multimodal characteristics, it is difficult for the prior art to effectively describe the variability of geotechnical parameters, making probability distribution modeling challenging.
A method combining kernel density estimation algorithm and genetic algorithm is used to solve the problem of data sparse and multimodal characteristics by optimizing bandwidth parameters.
The accuracy and accuracy of geotechnical parameter probability distribution modeling is improved, without subjective assumption of bandwidth and parameter distribution type, and is suitable for situations where data is insufficient and multi-peak characteristics are multi-peaked.
Smart Images

Figure CN119720809B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geotechnical parameter analysis, and particularly to a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition. Background Art
[0002] Geotechnical parameters, including effective cohesion, effective friction angle, and undrained shear strength, are important parameters for evaluating the safety and stability of geotechnical structures, such as soil slopes, retaining walls, foundation pits, and shallow foundations. However, due to the combined effects of multi-source uncertainties (such as spatial variability, transformation uncertainty, statistical uncertainty, and measurement errors), these parameters usually have uncertainties, among which the variability of geotechnical parameters is one of the most important issues.
[0003] Since the variability of geotechnical parameters has a significant impact on geotechnical engineering analysis and design, how to correctly describe this variability plays a key role in geotechnical engineering. Data statistics (such as mean, standard deviation, and quantile) and probability distributions (such as probability density function and cumulative distribution function) have been widely used to describe the variability of geotechnical parameters. Some classical distributions (such as normal distribution and lognormal distribution) are also often used to fit the data of geotechnical parameters. In this way, the corresponding statistical data and probability distributions of geotechnical parameters can be easily derived.
[0004] In actual geotechnical engineering, due to time, cost, and site limitations, it is impossible to investigate a large amount of data, and the data usually comes from on-site measurements or experimental tests. Sometimes, on-site data must be collected at specific locations, resulting in very sparse and limited data. Geotechnical parameters based on sparse data may not follow the above normal or lognormal distribution. Instead, these parameters may follow other commonly used distributions (for example, Weibull, exponential, Gamma, and extreme value type I distributions, etc.) or irregular distributions, and these distributions may exhibit multi-modal characteristics. The lack of data and multi-modal characteristics make it challenging to reasonably estimate the probability distribution of geotechnical parameters. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition to solve the problem in the prior art that in geotechnical engineering, due to the lack of data and multi-modal characteristics of geotechnical parameters, a complex calculation process is required to describe the variability of geotechnical parameters first in order to obtain the probability distribution model.
[0006] The embodiments of the present invention provide a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition, including:
[0007] Establish a sparse data set according to the collected sample data and the number of samples;
[0008] Based on a sparse dataset, the Gaussian function is used as the kernel function, and the approximate probability density function and the approximate cumulative distribution function are obtained through the kernel density estimation algorithm;
[0009] Obtain the empirical cumulative distribution function of the sparse dataset;
[0010] Obtain the maximum absolute difference through the difference between the approximate cumulative distribution function and the empirical cumulative distribution function;
[0011] Take several random bandwidth parameters as the initial population of the chromosome, and optimize and solve the maximum absolute difference through the genetic algorithm. Obtain the chromosome corresponding to the maximum fitness as the optimal bandwidth parameter of the approximate probability density function and the approximate cumulative distribution function;
[0012] Substitute the optimal bandwidth parameter into the approximate probability density function and the approximate cumulative distribution function to obtain the target approximate probability density function and the target approximate cumulative distribution function.
[0013] Optionally, the approximate probability density function is:
[0014] ;
[0015] The approximate cumulative distribution function is:
[0016] ;
[0017] where x i ∈ the sparse dataset {x1, x2, …, x n}, n is the number of samples in the sparse dataset; h0 is the bandwidth parameter to be optimized; K(*) is the kernel function, x is the input value, represents the approximate probability density function at x; is the standard normal cumulative distribution function.
[0018] Optionally, based on the sparse dataset, obtaining the empirical cumulative distribution function of the sparse dataset includes:
[0019] ;
[0020] ;
[0021] where x i ∈ the sparse dataset {x1, x2, …, x n}; x is the input value; n is the number of samples in the sparse dataset.
[0022] Optionally, obtaining the maximum absolute difference through the difference between the approximate cumulative distribution function and the empirical cumulative distribution function includes:
[0023] ;
[0024] where D max is the objective function to be optimized.
[0025] Optionally, several random bandwidth parameters are used as the initial population of chromosomes, and the maximum absolute difference is optimized and solved through a genetic algorithm to obtain the chromosome corresponding to the maximum fitness as the optimal bandwidth parameter of the approximate probability density function and the approximate cumulative distribution function, including:
[0026] Randomly generate N chromosomes to obtain the initial population {h1, h2, …, h N};
[0027] Take the reciprocal of the objective function to be optimized to obtain the fitness function:
[0028] ;
[0029] Use the fitness function to calculate the fitness values {f1, f2, …, f N} of the initial population {h1, h2, …, h N}.
[0030] Optionally, it further includes:
[0031] Select M chromosomes from the initial population, and the probability that the jth chromosome is selected is:
[0032] ;
[0033] Generate the remaining N - M chromosomes through a crossover operation, set the crossover probability to c1, and perform a mutation operation on the chromosomes obtained through the crossover operation, set the mutation probability to c2, to obtain a new generation of N chromosomes;
[0034] After a preset number of iterations, output the current optimal solution to obtain the optimal bandwidth parameter.
[0035] Optionally, the crossover probability c1 is set to 0.5; the mutation probability c2 is set to 0.5; the preset number of iterations is set to 100 times.
[0036] Optionally, it further includes:
[0037] Before reaching the preset number of iterations, if the output result of the fitness function has reached convergence, stop the iteration.
[0038] Advantages of the present invention:
[0039] 1. A sparse data probability distribution modeling method applied to geotechnical parameter data acquisition provided in this embodiment transforms the problem of selecting the bandwidth parameter of kernel density estimation into an optimization problem. The goal is to minimize the maximum absolute difference between the approximate cumulative distribution function and the empirical cumulative distribution function. By adaptively solving the optimal bandwidth, the modeling accuracy of the probability distribution is effectively improved.
[0040] 2. A sparse data probability distribution modeling method applied to geotechnical parameter data acquisition provided in this embodiment does not require subjective assumptions about the bandwidth and the type of parameter distribution, and has higher accuracy than traditional modeling methods.
[0041] 3. A sparse data probability distribution modeling method applied to geotechnical parameter data acquisition provided in this embodiment can reasonably estimate the probability distribution under data insufficiency and multimodal characteristics, and has higher applicability than traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present invention. In the drawings:
[0043] Figure 1 shows a flowchart of a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in an embodiment of the present invention;
[0044] Figure 2 shows a schematic diagram of the absolute difference between an approximate cumulative distribution function and an empirical cumulative distribution function in an embodiment of the present invention;
[0045] Figure 3 shows an optimization iteration process diagram of the absolute difference between an approximate cumulative distribution function and an empirical cumulative distribution function in an embodiment of the present invention;
[0046] Figure 4 shows a comparison diagram of probability density functions between a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in an embodiment of the present invention and other methods;
[0047] Figure 5 shows a comparison diagram of cumulative distribution functions between a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in an embodiment of the present invention and other methods;
[0048] Figure 6 shows an optimization iteration process diagram of the absolute difference between an approximate cumulative distribution function and an empirical cumulative distribution function in an embodiment of the present invention;
[0049] Figure 7Shows the comparison diagram of the probability density functions between a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in an embodiment of the present invention and other methods;
[0050] Figure 8 Shows the comparison diagram of the cumulative distribution functions between a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in an embodiment of the present invention and other methods. Specific implementation manner
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] The embodiments of the present invention provide a sparse data probability distribution modeling method applied to geotechnical parameter data acquisition, as Figure 1 shown, including:
[0053] Step S10, establish a sparse data set according to the collected sample data and the number of samples.
[0054] In this embodiment, the sample data set {x1, x2,..., x n} is the sparse data set, and the data is obtained through on-site measurement or experimental testing, and the sample size is n.
[0055] Step S20, based on the sparse data set, use the Gaussian function as the kernel function, and obtain the approximate probability density function and the approximate cumulative distribution function through the kernel density estimation algorithm.
[0056] In this embodiment, the approximate probability density function is:
[0057] .
[0058] .
[0059] Among them, x i ∈ sparse data set {x1, x2,..., x n}; h0 is the bandwidth parameter to be optimized, which is an unknown value; K(*) is the kernel function, x is the input value, represents the approximate probability density function at x; is the standard normal cumulative distribution function.
[0060] The approximate probability density function can be rewritten as:
[0061] ;
[0062] In the formula, is the standard normal probability density function.
[0063] The approximate probability density function can be rewritten as:
[0064] .
[0065] In the formula, .
[0066] Step S30: Obtain the empirical cumulative distribution function of the sparse data set.
[0067] In this embodiment, the empirical cumulative distribution function of the sparse data set is:
[0068] .
[0069] .
[0070] where x i ∈ the sparse data set {x1, x2, …, x n}; x is the input value; n is the number of samples in the sparse data set. As can be seen from the above formula, T(0) = 0, T(x n ) = 1.
[0071] Step S40: Obtain the maximum absolute difference by the difference between the approximate cumulative distribution function and the empirical cumulative distribution function.
[0072] In this embodiment, the maximum absolute difference . The selection of the bandwidth parameter is the key parameter affecting the maximum absolute difference and is also an important factor affecting the accuracy of kernel density estimation. Therefore, minimizing the maximum absolute difference D max is used as the objective function and optimized to solve it:
[0073] .
[0074] In a specific embodiment, for a given point x0, the absolute difference between the approximate cumulative distribution function and the empirical cumulative distribution function is as Figure 2 shown.
[0075] Step S50: Use several random bandwidth parameters as the initial population of chromosomes, and optimize and solve the maximum absolute difference through a genetic algorithm to obtain the chromosome corresponding to the maximum fitness as the optimal bandwidth parameter of the approximate probability density function and the approximate cumulative distribution function.
[0076] In this embodiment, the process of optimizing and solving by the genetic algorithm is as follows:
[0077] Randomly generate N chromosomes to obtain the initial population {h1, h2, …, h N}.
[0078] Since the genetic algorithm aims to find the optimal solution with the maximum fitness value, the reciprocal of the objective function to be optimized is taken to obtain the fitness function:
[0079] .
[0080] Calculate the fitness values {f1, f2, …, f N} of the initial population {h1, h2, …, h N} using the fitness function.
[0081] Select M chromosomes from the initial population. The probability that the j-th chromosome is selected is:
[0082] .
[0083] The larger the fitness value f j , the greater the probability that the j-th chromosome h j is selected. That is to say, as the algorithm progresses, chromosomes with small fitness values will be eliminated, and the overall quality of the chromosomes will be improved.
[0084] Generate the remaining N - M chromosomes through the crossover operation. The crossover probability is set to c1, and the chromosomes obtained through the crossover operation are subjected to the mutation operation. The mutation probability is set to c2 to obtain a new generation of N chromosomes. In a specific embodiment, the crossover probability c1 is set to 0.5; the mutation probability c2 is set to 0.5.
[0085] After a preset number of iterations, output the current optimal solution to obtain the optimal bandwidth parameter. In a specific embodiment, the preset number of iterations is set to 100 times.
[0086] In a specific embodiment, if the output result of the fitness function has reached convergence, stop the iteration. In a specific implementation, when the calculation result tends to be stable, it is determined that convergence has been reached.
[0087] Step S60, substitute the optimal bandwidth parameter into the approximate probability density function and the approximate cumulative distribution function to obtain the target approximate probability density function and the target approximate cumulative distribution function.
[0088] A sparse data probability distribution modeling method applied to geotechnical parameter data acquisition in this embodiment transforms the problem of selecting the bandwidth parameter of kernel density estimation into an optimization problem. The goal is to minimize the maximum absolute difference between the approximate cumulative distribution function and the empirical cumulative distribution function. By adaptively solving the optimal bandwidth, the modeling accuracy of the probability distribution is effectively improved. There is no need to make subjective assumptions about the bandwidth and the type of parameter distribution, and it has higher accuracy than traditional modeling methods. In addition, the sparse data probability distribution modeling method applied to geotechnical parameter data acquisition provided in this embodiment can reasonably estimate the probability distribution under data insufficiency and multimodal characteristics, and has higher applicability than traditional methods.
[0089] Table 1 Data of effective cohesion
[0090] ;
[0091] Taking "Efficient sampling of the irregular probability distributions of geotechnical parameters for reliability analysis" (Structural Safety Volume 101, March 2023, 102309) and "Site-specific probability distribution of geotechnical properties" (Computers and Geotechnics Volume 70, October 2015, Pages 159-168) as references, the histogram, as well as the methods proposed by Wang et al. and Jiang et al., are compared and analyzed with the method given in the present invention.
[0092] Taking the effective cohesion of 63 real data as an example, the data is shown in Table 1. Using the sparse data probability distribution modeling method provided in this embodiment, the process of solving the optimization iteration by using the genetic algorithm is as Figure 3 shown, and the optimal bandwidth parameter can be effectively obtained, and finally the probability distribution of the effective cohesion is obtained.
[0093] The comparison results of the probability density functions of different methods are as Figure 4 shown. In addition, for a full comparison and explanation, as Figure 5 shown, the comparison results of the cumulative distribution functions of each method are also given. From Figure 4 and Figure 5 it can be seen that the results between the method given in this embodiment and the methods developed by Wang et al. and Jiang et al. are very close, verifying the effectiveness of the proposed method.
[0094] It can be seen from Figure 3 that the iterative process converges very quickly. After 10 iterations, the calculation results tend to be stable. The results of the first 12 iterations are shown in Table 2 in detail.
[0095] Table 2 Results of the First 12 Iterations Based on the Dataset in Table 1
[0096] ;
[0097] The optimal bandwidth is determined to be h = 0.8369. Substituting n = 63 and h = 0.8369 into the approximate probability density function and the approximate cumulative distribution function in step S20 of this embodiment, we get:
[0098] .
[0099] .
[0100] The scarce data of the 10 effective friction angles in this embodiment come from the target sand layer by Lake Mildred, and the specific values are shown in Table 3.
[0101] Table 3 Data of Effective Friction Angles
[0102] ;
[0103] Through the sparse data probability distribution modeling method provided by this embodiment, the optimization iteration process is as Figure 6 shown in and Table 4.
[0104] Table 4 Results of the First 35 Iterations Based on the Dataset in Table 3
[0105] ;
[0106] It can be seen from Figure 6 and Table 4 that the iterative process converges after 33 iterations, and the optimal bandwidth is determined to be h = 0.3345. Substituting n = 10 and h = 0.3345 into the approximate probability density function and the approximate cumulative distribution function in step S20 of this embodiment, the corresponding approximate probability density function and approximate cumulative distribution function can be obtained.
[0107] The comparison results of the probability density functions and cumulative distribution functions of different methods are respectively as Figure 7 and Figure 8 shown. It can be seen that the sparse data probability distribution modeling method proposed in this embodiment for geotechnical parameter data acquisition can effectively approximate the probability distribution of sparse data.
[0108] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A sparse data probability distribution modeling method for geotechnical parameter data acquisition, characterized in that: include: Establish a sparse data set based on the collected sample data and sample quantity; Based on the sparse data set, a Gaussian function is used as a kernel function, and an approximate probability density function and an approximate cumulative distribution function are obtained by a kernel density estimation algorithm; Obtaining an empirical cumulative distribution function of the sparse data set; Obtaining a maximum absolute difference by the difference between the approximate cumulative distribution function and the empirical cumulative distribution function; Using several random bandwidth parameters as the initial population of chromosomes, optimizing and solving the maximum absolute difference through a genetic algorithm, and obtaining the chromosome corresponding to the maximum fitness as the optimal bandwidth parameter of the approximate probability density function and the approximate cumulative distribution function; Substituting the optimal bandwidth parameter into the approximate probability density function and the approximate cumulative distribution function to obtain a target approximate probability density function and a target approximate cumulative distribution function; Wherein, the approximate probability density function is: ; The approximate cumulative distribution function is: ; Among them, x i ∈ sparse dataset {x1,x2,…,x n }, n is the number of samples in the sparse data set; h0 is the bandwidth parameter to be optimized; K(*) is the kernel function, x is the input value, represents the approximate probability density function at x; is the standard normal cumulative distribution function; Based on the sparse data set, obtaining an empirical cumulative distribution function of the sparse data set includes: ; ; Among them, x i ∈ sparse dataset {x1,x2,…,x n }; x is the input value; n is the number of samples in the sparse data set; Obtaining a maximum absolute difference through a difference between the approximate cumulative distribution function and the empirical cumulative distribution function, comprising: ; Where D max is the objective function to be optimized; Randomly generate N chromosomes to obtain the initial population {h1,h2,…,h N }; The inverse of the objective function to be optimized is taken to obtain the fitness function: ; The fitness function is used to calculate the initial population {h1,h2,…,h N The fitness value of {f1,f2,…,f N }.
2. The sparse data probability distribution modeling method for geotechnical parameter data acquisition according to claim 1 is characterized in that: Also includes: M chromosomes are selected from the initial population, and the probability of the jth chromosome being selected is: ; The remaining NM chromosomes are generated through crossover operation, with the crossover probability set to c1. The chromosomes obtained through crossover operation are mutated, with the mutation probability set to c2, to obtain a new generation of N chromosomes. After a preset number of iterations, the current optimal solution is output and the optimal bandwidth parameter is obtained.
3. The sparse data probability distribution modeling method for geotechnical parameter data acquisition according to claim 2 is characterized in that: The crossover probability c1 is set to 0.5; the mutation probability c2 is set to 0.5; and the preset number of iterations is set to 100.
4. The sparse data probability distribution modeling method for geotechnical parameter data acquisition according to claim 2 is characterized in that: Also includes: Before reaching the preset number of iterations, if the output result of the fitness function has reached convergence, the iteration is stopped.
Citation Information
Patent Citations
Path probability forming method for collision-free environment exploration of mobile robot
CN116642507A
Thin terminal oriented multi-video stream display method and system
CN1595976A