High-dimensional soil data visualization method and system based on density peak clustering

By using the density peak clustering algorithm to reduce the dimensionality of high-dimensional soil and rock data and visualize it, the problem that traditional methods are difficult to capture nonlinear characteristics is solved, and accurate prediction and visualization of soil mechanical behavior are achieved.

CN118708998BActive Publication Date: 2025-11-18中国水利水电第七工程局有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410802343.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-11-18
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

Traditional geotechnical engineering data analysis methods struggle to capture the nonlinear characteristics and multidimensional relationships of high-dimensional data, making it impossible to accurately predict the mechanical behavior of soil.

Method used

A high-dimensional soil data visualization method based on density peak clustering is adopted, which includes data cleaning, standardization, dimensionality reduction, clustering and visualization steps. Clustering is performed in 3D space using the density peak clustering algorithm, and the dimensionality reduction effect is evaluated using ARI, NMI and ACC indicators. Finally, the clustering results are displayed graphically.

Benefits of technology

It improves the ability to analyze high-dimensional data, enabling more accurate prediction of soil mechanical behavior, and enhances the interpretability and applicability of research results through visualization tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118708998B_ABST
    Figure CN118708998B_ABST
Patent Text Reader

Abstract

The application discloses a high-dimensional soil data visualization method and system based on density peak clustering, and the visualization method comprises the following steps: S1, a data set comprising a plurality of samples is acquired, and data cleaning and standardization processing are performed on the samples in the data set, and each sample comprises various types of rock-soil data; S2, the data set after the standardization processing is labeled according to the geographical positions of stations, and then dimension reduction processing is performed on the data set after the label labeling, so as to obtain 3-dimensional reduced data; S3, a density peak clustering algorithm is used to cluster the 3-dimensional reduced data, so as to obtain a clustering result; and S4, the visualization method is used to convert the clustering result into a graphic display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data visualization technology, specifically to a high-dimensional soil data visualization method and system based on density peak clustering. Background Technology

[0002] In geotechnical engineering practice, the collection and analysis of geotechnical data is not only the cornerstone of engineering safety and stability, but also has a profound impact on the project's economy and sustainability. Geotechnical data obtained through meticulous and systematic geological exploration, field tests, in-situ testing, and long-term monitoring cover a series of key parameters, from soil layer distribution and rock structure to dynamic changes in groundwater levels.

[0003] In the practice of traditional geotechnical engineering data analysis, statistical characteristic description methods such as mean, standard deviation, and correlation coefficient are usually relied upon to generate visualizations of basic statistical data so that researchers can intuitively understand geotechnical engineering data. Although the existing information can provide the central trend, dispersion, and linear relationship between variables, it cannot capture the nonlinear characteristics of geotechnical engineering data distribution, multidimensional relationships, and complex interactions between data points.

[0004] Furthermore, geotechnical engineering data typically exhibits high dimensionality and nonlinearity, making it difficult for traditional statistical methods to handle this complexity. For instance, the mechanical behavior of soil can be influenced by a variety of factors, including particle size distribution, mineral composition, pore water pressure, and stress history. The interactions and influences between these factors often exceed the scope of traditional statistical analysis, making it difficult to fully extract information such as the geometric and density structures of geotechnical data, thus hindering accurate prediction of soil mechanical behavior. Summary of the Invention

[0005] To address the aforementioned shortcomings in existing technologies, the high-dimensional soil data visualization method based on density peak clustering provided by this invention solves the problem that existing technologies are unable to fully extract information such as the geometric and density structures of soil and rock data, thus failing to accurately predict the mechanical behavior of soil.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0007] Firstly, a high-dimensional soil data visualization method based on density peak clustering is provided, which includes the following steps:

[0008] S1. Obtain a dataset containing several samples, and perform data cleaning and standardization on the samples in the dataset. Each sample includes multiple types of geotechnical data.

[0009] S2. Label the standardized dataset according to the geographical location of the sites, and then perform dimensionality reduction on the labeled dataset to obtain 3-dimensional reduced data.

[0010] S3. The density peak clustering algorithm is used to cluster the dimensionality-reduced data that has been reduced to 3 dimensions to obtain the clustering results;

[0011] S4. Use visualization methods to transform the clustering results into graphical displays.

[0012] Furthermore, one of the following dimensionality reduction methods is selected: LE dimensionality reduction, PCA dimensionality reduction, UMAP dimensionality reduction, ISOMAP dimensionality reduction, and t-SNE dimensionality reduction. The selected methods include:

[0013] S21. Obtain a standard dataset that includes multiple standard samples;

[0014] S22. Dimensionality reduction of the standard dataset is performed using LE, PCA, UMAP, ISOMAP, and t-SNE methods respectively.

[0015] S23. The density peak clustering algorithm is used to cluster the standard samples without dimensionality reduction and the data samples after dimensionality reduction, respectively, to obtain the clustering results;

[0016] S24. Treat the clustering results of the standard samples without dimensionality reduction as the true classification. Calculate and quantify the true classification Clustering results corresponding to each dimensionality reduction method Consistent ARI, NMI, and ACC metrics;

[0017] S25. Based on the ARI, NMI, and ACC metrics corresponding to all dimensionality reduction methods, select one dimensionality reduction method:

[0018]

[0019] ,

[0020]

[0021]

[0022]

[0023]

[0024] Where max represents the maximum value; These represent the LE dimensionality reduction method, PCA dimensionality reduction method, UMAP dimensionality reduction method, ISOMAP dimensionality reduction method, and t-SNE dimensionality reduction method, respectively. For ARI index functions; For NMI index functions; This is the ACC indicator function; , and The weights for the ARI, NMI, and ACC indicators are respectively. To obtain the minimum value; They are respectively and The Middle i Number of standard samples in a cluster; m The total number of clusters; for and In the i The number of standard samples with the same corresponding position in each cluster; for The Middle i The number of standard samples in a cluster; , and These are used to reflect the magnitudes of the three indicators, ARI, NMI, and ACC, and are all values ​​between [0, 1].

[0025] Furthermore, , and The expressions are as follows:

[0026]

[0027]

[0028]

[0029] in, a for and The number of point pairs belonging to the same cluster in the standard dataset; a point pair is a pair of comparative data consisting of two different standard samples in the standard dataset. b for They belong to the same cluster and The number of point pairs that do not belong to the same cluster; c for They do not belong to the same cluster. The number of point pairs belonging to the same cluster; d In order to be in and The number of point pairs that do not belong to the same cluster; They are respectively and Information entropy; for and Mutual information; Representing the same point and Whether the clustering results are the same, and whether they are assigned to the same cluster; when hour, ,otherwise The values ​​of the ARI, NMI, and ACC indicators all range from [0, 1].

[0030] Furthermore, various types of geotechnical data include the depth of measurement points, effective stress, normalized undrained shear strength, overconsolidation ratio, normalized cone tip resistance, normalized effective cone tip resistance, normalized excess pore water pressure, and pore pressure ratio.

[0031] Furthermore, the standardized cone tip resistance Standardized effective cone tip resistance Standardized ultrastatic pore water pressure And pore pressure ratio The expressions are as follows:

[0032] , , ,

[0033] in, For cone tip resistance; It is vertical stress; Effective stress; Pore ​​water pressure; This represents the initial pore water pressure.

[0034] Furthermore, the dataset is the Clay / 6 / 535 dataset from the ISSMGE TC304 soil database.

[0035] Secondly, a high-dimensional soil data visualization system based on density peak clustering is provided, which includes:

[0036] The preprocessing module is used to acquire a dataset containing several samples and to perform data cleaning and standardization on the samples in the dataset. Each sample includes multiple types of geotechnical data.

[0037] The data labeling and dimensionality reduction module is used to label the standardized dataset according to the geographical location of the sites, and then perform dimensionality reduction on the labeled dataset to obtain 3D dimensionality-reduced data.

[0038] The clustering module is used to cluster the dimensionality-reduced data that has been reduced to 3 dimensions using the density peak clustering algorithm to obtain the clustering results;

[0039] The visualization module uses visualization methods to transform clustering results into graphical displays.

[0040] The beneficial effects of this invention are as follows: This solution can perform low-dimensional mapping on high-dimensional datasets, thereby enabling analysis of the datasets from the perspectives of geometric and density structures. By mining the geometric and density structure information of the data, it can provide better guidance for clustering and visualization. The objective function designed in this solution can automatically select the method with the best dimensionality reduction effect, opening up new paths for the modernization and intelligent development of the geotechnical engineering field, enabling a more comprehensive understanding of geotechnical data, and providing more accurate and reliable support for the planning and implementation of engineering projects.

[0041] The visualization system provided by this solution effectively overcomes the abstract and difficult-to-understand problems that may exist in simple algorithm implementation, enhances the interpretability and practicality of research results, and provides researchers and engineers with a powerful data mining and analysis tool. It has significant application value in many fields involving complex high-dimensional data processing. Attached Figure Description

[0042] Figure 1 This is a flowchart of a high-dimensional soil data visualization method based on density peak clustering. Detailed Implementation

[0043] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0044] refer to Figure 1 , Figure 1 A high-dimensional soil data visualization method based on density peak clustering is shown; the method S includes steps S1 to S4.

[0045] In step S1, a dataset including several samples is obtained, and the samples in the dataset are cleaned and standardized. Each sample includes multiple types of geotechnical data.

[0046] During implementation, this scheme prioritizes various types of geotechnical data, including the depth of measurement points, effective stress, standardized undrained shear strength, overconsolidation ratio, standardized cone tip resistance, standardized effective cone tip resistance, standardized excess pore water pressure, and pore pressure ratio.

[0047] Among them, the standardized cone tip resistance Standardized effective cone tip resistance Standardized ultrastatic pore water pressure And pore pressure ratio The expressions are as follows:

[0048] , , ,

[0049] in, For cone tip resistance; It is vertical stress; Effective stress; Pore ​​water pressure; This represents the initial pore water pressure.

[0050] The preferred dataset is the Clay / 6 / 535 dataset from the ISSMGE TC304 soil database.

[0051] In step S2, the standardized dataset is labeled according to the geographical location of the site. Then, the labeled dataset is dimensionality reduced to obtain 3D dimensionality-reduced data to facilitate observation of the geometric structure and density distribution of the samples.

[0052] In step S3, the density peak clustering algorithm is used to cluster the dimensionality-reduced data to 3 dimensions to obtain the clustering results; the density peak clustering algorithm is described in detail below:

[0053] First, calculate the Euclidean distance between points in the reduced-dimensional data and construct a distance matrix, where each element... d op It is the first o The point and the first p The Euclidean distance between points.

[0054] Then, the data in the distance matrix are sorted in ascending order, and a cutoff distance is selected. The selected cutoff distance should ensure that the average number of samples contained in the neighborhood of each data point accounts for 2% of the total number of samples.

[0055] For data point o, its local density The calculation is performed using a Gaussian kernel, as shown in equation (1):

[0056]

[0057] in, An exponential function refers to... ; The cutoff distance serves as the bandwidth of the Gaussian kernel function;

[0058] Define the distance between the point and the nearest larger density point. As shown in equation (2).

[0059]

[0060] This refers to the distance between point o and the nearest point with the highest density. When the density of point o is at its maximum value, its... It is expressed by equation (3).

[0061]

[0062] in, Let be the Euclidean distance between point p (excluding the point o with the highest density) and the nearest larger density point of point p.

[0063] After obtaining the local density of all data points and its distance from the nearest larger density point, a decision value is constructed using these two indicators to select the cluster center, where the decision value is represented by equation (4).

[0064]

[0065] in, Let O be the density of data point o;

[0066] The top 40 points with the largest decision values ​​are used as cluster centers, and 40 is the number of clusters in the dataset.

[0067] After determining the cluster centers, the remaining points need to be assigned. The Density Peak Clustering (DPC) algorithm uses a one-step assignment method, where each remaining point is assigned to the cluster to which its nearest neighbor with the highest density belongs.

[0068] In step S4, a visualization method is used to transform the clustering results into a graphical display. This visualization method can employ mature existing technologies or software, and the graphics can be scatter plots, heatmaps, or 3D diagrams, allowing users to intuitively observe the distribution of the data after dimensionality reduction, as well as the characteristics and boundaries of each cluster.

[0069] In one embodiment of the present invention, one of the following dimensionality reduction methods is selected: LE dimensionality reduction, PCA dimensionality reduction, UMAP dimensionality reduction, ISOMAP dimensionality reduction, and t-SNE dimensionality reduction; the selection method includes:

[0070] S21. Obtain a standard dataset that includes multiple standard samples; specifically, the standard samples are relatively complete data of various types of geotechnical data, and a small number of relatively accurate and complete standard samples can be obtained through experiments.

[0071] S22. Dimensionality reduction of the standard dataset is performed using LE, PCA, UMAP, ISOMAP, and t-SNE methods respectively.

[0072] S23. The density peak clustering algorithm is used to cluster the standard samples without dimensionality reduction and the data samples after dimensionality reduction, respectively, to obtain the clustering results;

[0073] S24. Treat the clustering results of the standard samples without dimensionality reduction as the true classification. Calculate and quantify the true classification Clustering results corresponding to each dimensionality reduction method Consistent ARI, NMI, and ACC metrics;

[0074] S25. Based on the ARI, NMI, and ACC metrics corresponding to all dimensionality reduction methods, select one dimensionality reduction method:

[0075]

[0076] ,

[0077]

[0078]

[0079]

[0080]

[0081] Where max represents the maximum value; These represent the LE dimensionality reduction method, PCA dimensionality reduction method, UMAP dimensionality reduction method, ISOMAP dimensionality reduction method, and t-SNE dimensionality reduction method, respectively. For ARI index functions; For NMI index functions; This is the ACC indicator function; , and The weights for the ARI, NMI, and ACC indicators are respectively. To obtain the minimum value; They are respectively and The Middle i Number of standard samples in a cluster; m The total number of clusters; for and In the i The number of standard samples with the same corresponding position in each cluster; for The Middle i The number of standard samples in a cluster; , and These values ​​are used to reflect the magnitudes of the three indicators, ARI, NMI, and ACC, and are all values ​​between [0, 1]. When calculating the weight of each indicator, these three values ​​can be used for normalization, and the indicator with the larger value will have a larger weight.

[0082] Based on the objective function innovatively designed in this scheme (all formulas mentioned in step S25), the dimensionality reduction method with the best clustering effect is selected, realizing automatic evaluation and selection of the best dimensionality reduction method to present the optimal visualization form of soil data.

[0083] During implementation, this solution is preferred. , and The expressions are as follows:

[0084]

[0085]

[0086]

[0087] in, a for and The number of point pairs belonging to the same cluster in the standard dataset; a point pair is a pair of comparative data consisting of two different standard samples in the standard dataset. b for They belong to the same cluster and The number of point pairs that do not belong to the same cluster; c for They do not belong to the same cluster. The number of point pairs belonging to the same cluster; d In order to be in and The number of point pairs that do not belong to the same cluster; They are respectively and Information entropy; for and Mutual information; Representing the same point and Whether the clustering results are the same, and whether they are assigned to the same cluster; when hour, ,otherwise The values ​​of the ARI, NMI, and ACC indicators all range from [0, 1].

[0088] and All can be calculated using formula (5). Formula (6) can be used for calculation, specifically:

[0089]

[0090]

[0091] in, and They are respectively l The Middle o and p The percentage of points in each cluster; exist Belonging to the o-th cluster, in The percentage of points belonging to the p-th cluster.

[0092] This solution also provides a high-dimensional soil data visualization system based on density peak clustering, which includes:

[0093] The preprocessing module is used to acquire a dataset containing several samples and to perform data cleaning and standardization on the samples in the dataset. Each sample includes multiple types of geotechnical data.

[0094] The data labeling and dimensionality reduction module is used to label the standardized dataset according to the geographical location of the sites, and then perform dimensionality reduction on the labeled dataset to obtain 3D dimensionality-reduced data.

[0095] The clustering module is used to cluster the dimensionality-reduced data that has been reduced to 3 dimensions using the density peak clustering algorithm to obtain the clustering results;

[0096] The visualization module uses visualization methods to transform clustering results into graphical displays.

[0097] In summary, this solution improves the ability to capture and analyze the nonlinear characteristics of soil data through dimensionality reduction, thereby more accurately predicting the mechanical behavior of soil; visualization can present complex soil data in graphical, image, or other visual formats, helping users to understand the data content more intuitively.

Claims

1. A high-dimensional soil data visualization method based on density peak clustering, characterized in that, Including the following steps: S1. Obtain a dataset containing several samples, and perform data cleaning and standardization on the samples in the dataset. Each sample includes multiple types of geotechnical data. S2. Label the standardized dataset according to the geographical location of the sites, and then perform dimensionality reduction on the labeled dataset to obtain 3-dimensional reduced data. S3. The density peak clustering algorithm is used to cluster the dimensionality-reduced data that has been reduced to 3 dimensions to obtain the clustering results; S4. Use visualization methods to transform the clustering results into graphical displays; Choose one of the following dimensionality reduction methods: LE, PCA, UMAP, ISOMAP, and t-SNE. Selection methods include: S21. Obtain a standard dataset that includes multiple standard samples; S22. Dimensionality reduction of the standard dataset is performed using LE, PCA, UMAP, ISOMAP, and t-SNE methods respectively. S23. The density peak clustering algorithm is used to cluster the standard samples without dimensionality reduction and the data samples after dimensionality reduction, respectively, to obtain the clustering results; S24. Treat the clustering results of the standard samples without dimensionality reduction as the true classification. Calculate and quantify the true classification Clustering results corresponding to each dimensionality reduction method Consistent ARI, NMI, and ACC metrics; S25. Based on the ARI, NMI, and ACC metrics corresponding to all dimensionality reduction methods, select one dimensionality reduction method: Where max represents the maximum value; These represent the LE dimensionality reduction method, PCA dimensionality reduction method, UMAP dimensionality reduction method, ISOMAP dimensionality reduction method, and t-SNE dimensionality reduction method, respectively. For ARI index functions; For NMI index functions; This is the ACC indicator function; , and The weights for the ARI, NMI, and ACC indicators are respectively. To obtain the minimum value; They are respectively and The Middle i Number of standard samples in a cluster; m The total number of clusters; for and In the i The number of standard samples with the same corresponding position in each cluster; for The Middle i The number of standard samples in a cluster; , and These are used to reflect the magnitudes of the three indicators, ARI, NMI, and ACC, and are all values ​​between [0, 1]. They are respectively and Information entropy.

2. The high-dimensional soil data visualization method based on density peak clustering according to claim 1, characterized in that, , and The expressions are as follows: in, a for and The number of point pairs belonging to the same cluster in the standard dataset; a point pair is a pair of comparative data consisting of two different standard samples in the standard dataset. b for They belong to the same cluster and The number of point pairs that do not belong to the same cluster; c for They do not belong to the same cluster. The number of point pairs belonging to the same cluster; d In order to be in and The number of point pairs that do not belong to the same cluster; for and Mutual information; Representing the same point and Whether the clustering results are the same, and whether they are assigned to the same cluster; when hour, ,otherwise The values ​​of the ARI, NMI, and ACC indicators all range from [0, 1].

3. The high-dimensional soil data visualization method based on density peak clustering according to claim 1, characterized in that, Multiple types of geotechnical data include the depth of measurement points, effective stress, normalized undrained shear strength, overconsolidation ratio, normalized cone tip resistance, normalized effective cone tip resistance, normalized excess pore water pressure, and pore pressure ratio.

4. The high-dimensional soil data visualization method based on density peak clustering according to claim 3, characterized in that, Standardized cone tip resistance Standardized effective cone tip resistance Standardized ultrastatic pore water pressure And pore pressure ratio The expressions are as follows: , , , in, For cone tip resistance; It is vertical stress; Effective stress; Pore ​​water pressure; This represents the initial pore water pressure.

5. The high-dimensional soil data visualization method based on density peak clustering according to any one of claims 1-4, characterized in that, The dataset in question is the Clay / 6 / 535 dataset from the ISSMGE TC304 soil database.

6. A high-dimensional soil data visualization system based on density peak clustering, characterized in that, To implement the high-dimensional soil data visualization method based on density peak clustering as described in claim 1, the system comprises: The preprocessing module is used to acquire a dataset containing several samples and to perform data cleaning and standardization on the samples in the dataset. Each sample includes multiple types of geotechnical data. The data labeling and dimensionality reduction module is used to label the standardized dataset according to the geographical location of the sites, and then perform dimensionality reduction on the labeled dataset to obtain 3D dimensionality-reduced data. The clustering module is used to cluster the dimensionality-reduced data that has been reduced to 3 dimensions using the density peak clustering algorithm to obtain the clustering results; The visualization module uses visualization methods to transform clustering results into graphical displays.

Citation Information

Patent Citations

  • Power grid topology automatic identification and construction method and system

    CN118133068A

  • Intelligent real-time updating method and system for stratigraphic framework with geosteering-while-drilling

    US20230333274A1