Geotechnical engineering investigation data analysis method and system based on machine learning

By constructing geotechnical data maps and using graph convolutional neural network (GNN) for information dissemination, the problems of high survey costs, uneven data and difficulty in real-time updates in traditional survey methods are solved, and high-precision inference and dynamic geological adaptation of geotechnical parameters in unknown areas are achieved, and the efficiency and accuracy of engineering surveys are improved.

CN120296351APending Publication Date: 2025-07-11CHONGQING THREE GORGES UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510370215.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional geotechnical engineering survey methods have high cost and long time, and sparse distribution of survey points leads to uneven geological data, making it difficult to obtain comprehensive geotechnical information. Traditional interpolation methods have low prediction accuracy in complex geological areas and cannot be updated in real time. The utilization rate of historical data is low, making it difficult to adapt to dynamic geological changes.

Method used

The geotechnical data map is constructed based on machine learning, and the node relationship is determined through Euclidean distance and geological similarity measurements. The graph convolutional neural network (GNN) is used for information dissemination, and combined with dynamic thresholding method and sparse processing, the graph structure is optimized to achieve high-precision inference of geotechnical parameters in unknown areas.

Benefits of technology

It improves the utilization efficiency and geological prediction accuracy of geotechnical survey data, reduces the survey cost, maintains high prediction accuracy in complex geological areas, adapts to dynamic geological changes, and enhances the accuracy of engineering decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296351A_ABST
    Figure CN120296351A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of geotechnical engineering, and provides a geotechnical engineering investigation data analysis method and system based on machine learning, and the method comprises the steps: obtaining geotechnical engineering investigation data; taking the sampling points as nodes of the rock-soil data graph; according to the spatial proximity relation and the geological similarity, determining edges of the rock-soil data graph, and endowing the edges with weights, so that nodes with similar spaces and similar geological characteristics are connected; carrying out sparse processing on the rock and soil data graph; performing weighted fusion according to the spatial distance weight and the geological similarity weight to form a final rock-soil data graph; a graph convolutional neural network is adopted, and the initial features and the adjacency relation of the nodes are used for information propagation; neighbor information of each node is aggregated through multi-layer graph convolution calculation, and node features are updated step by step; a semi-supervised learning or supervised learning method is adopted to train a GNN model, and prediction error minimization is taken as a target; and deducing rock and soil parameters of an unknown area through the trained GNN.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geotechnical engineering, and particularly relates to a method and system for analyzing geotechnical engineering exploration data based on machine learning. Background Art

[0002] Geotechnical engineering exploration is an important link in civil engineering, infrastructure construction, and geological disaster prevention. Its main purpose is to obtain the underground soil layer structure, geotechnical physical and mechanical properties, and groundwater conditions, providing a scientific basis for engineering design, construction, and safety assessment. Traditional geotechnical engineering exploration methods mainly include drilling and sampling, in-situ tests (such as standard penetration test SPT, cone penetration test CPT), and laboratory tests. While these methods provide high-precision geological data, they are usually accompanied by the following problems:

[0003] Due to the high exploration cost and long operation time, the spatial distribution of exploration points is often sparse, resulting in uneven coverage of geological data in different regions, making it difficult to obtain comprehensive geotechnical information during engineering design.

[0004] In un-sampled areas, traditional interpolation methods such as Kriging interpolation and inverse distance weighting (IDW) are often used to infer geotechnical parameters. However, these methods are mainly based on the spatial smoothing assumption and are difficult to accurately describe complex geological structures, especially in areas with drastic stratigraphic changes or significant groundwater influence, where the prediction accuracy is relatively low.

[0005] Geotechnical exploration data contains geographical information, geotechnical characteristics, mechanical parameters, and environmental influencing factors, with spatial correlation and non-linear relationships. Traditional methods are difficult to fully explore the potential relationships between data, resulting in a large amount of historical exploration data not being effectively utilized.

[0006] As the engineering construction progresses, the geological conditions may change. However, the static characteristics of traditional exploration data make it difficult to adapt to the dynamic adjustment requirements and unable to provide the latest geotechnical information in real time, affecting the accuracy of engineering decisions. Summary of the Invention

[0007] To solve the problems in the prior art, the present invention provides a method for analyzing geotechnical engineering exploration data based on machine learning, which is characterized by including the following steps:

[0008] Obtain geotechnical engineering exploration data, clean, normalize, and standardize the data, remove outliers, and fill in missing data;

[0009] Take the sampling points as the nodes of the geotechnical data map, and the attributes of each node include its geographical location information, soil characteristics, and mechanical parameters;

[0010] Use the Euclidean distance to determine the spatial proximity relationship between adjacent nodes to obtain the spatial distance weight;

[0011] Based on soil layer type, shear strength, and porosity parameters, a geological similarity measure between nodes is established to obtain geological similarity weights.

[0012] According to the spatial proximity relationship and geological similarity, the edges of the geotechnical data graph are determined, and edge weights are assigned so that nodes that are spatially close and have similar geological characteristics are connected.

[0013] Using the dynamic threshold method, the number of neighbors of each node is determined, and the geotechnical data graph is sparsified according to the spatial relationship.

[0014] The sparsified geotechnical data graph is weighted and fused according to the spatial distance weight and geological similarity weight to form the final geotechnical data graph.

[0015] Using a graph convolutional neural network, information propagation is carried out using the initial features and adjacency relationships of the nodes.

[0016] Through multi-layer graph convolution calculations, the neighbor information of each node is aggregated, and the node features are gradually updated so that the geotechnical parameters of unknown areas are obtained through information propagation from known points.

[0017] Using semi-supervised learning or supervised learning methods, the GNN model is trained with the goal of minimizing the prediction error.

[0018] The geotechnical parameters of unknown areas are inferred through the trained GNN.

[0019] Furthermore, the implementation of the dynamic threshold method includes:

[0020] The spatial distance between each pair of nodes in the geotechnical data graph is calculated through the Euclidean distance.

[0021] For each pair of nodes, the geological similarity measure between the nodes is calculated using geological features.

[0022] If the spatial distance between nodes is less than a preset threshold and the geological similarity of the nodes is higher than a second threshold, these nodes are regarded as neighbors.

[0023] If the spatial relationship and geological characteristics of the nodes meet the threshold conditions, they are connected; otherwise, the nodes are not connected.

[0024] Furthermore, the sparsification process includes:

[0025] For each edge in the graph, the weight of the edge is set by calculating the spatial distance and geological similarity weight of the edge. The weight of the edge is a combined function of the spatial distance and geological similarity of the two nodes, and the weight formula is as follows:

[0026]

[0027] Among them, w ij is the edge weight between node i and node j, α and β are adjustment parameters, and d ij is the spatial distance, and s ij is the geological similarity;

[0028] After calculating the edge weights, a sparsification threshold τ is set, and the edges with weights less than this threshold will be removed.

[0029] Furthermore, the spatial distance weight is calculated using the following method:

[0030]

[0031] Among them, w d (i,j) represents the spatial distance weight between node i and node j, and d ij represents the Euclidean distance,

[0032] and σ is an adjustment parameter.

[0033] Furthermore, the geological similarity weight is calculated using the following method:

[0034]

[0035] Among them, w g (i,j) represents the geological similarity weight between node i and j, f i,k , and f j,k respectively represent the values of the node on the kth geological feature, and w k is the weighting factor for each geological feature.

[0036] The present invention also provides a geotechnical engineering investigation data analysis system based on machine learning, which is characterized by including the following modules:

[0037] A data module for obtaining geotechnical engineering investigation data, cleaning, normalizing, and standardizing the data, removing outliers, and filling in missing data;

[0038] A graph construction module is used to take sampling points as nodes of the geotechnical data graph. The attributes of each node include its geographical location information, soil characteristics, and mechanical parameters. The Euclidean distance is used to determine the spatial proximity relationship between adjacent nodes to obtain the spatial distance weight. Based on soil layer type, shear strength, and porosity parameters, a geological similarity measure between nodes is established to obtain the geological similarity weight. According to the spatial proximity relationship and geological similarity, the edges of the geotechnical data graph are determined and edge weights are assigned so that nodes that are spatially close and geologically similar are connected. The dynamic threshold method is used to determine the number of neighbors of each node, and the geotechnical data graph is sparsified according to the spatial relationship. The sparsified geotechnical data graph is weighted and fused according to the spatial distance weight and geological similarity weight to form the final geotechnical data graph.

[0039] A training module is used to adopt a graph convolutional neural network to perform information propagation using the initial features and adjacency relationships of nodes. Through multi-layer graph convolution calculations, the neighbor information of each node is aggregated, and the node features are gradually updated so that the geotechnical parameters of unknown areas are obtained from the information propagation of known points. A semi-supervised learning or supervised learning method is used to train the GNN model with the goal of minimizing the prediction error.

[0040] An inference module is used to infer the geotechnical parameters of unknown areas through the trained GNN.

[0041] Furthermore, the implementation of the dynamic threshold method includes:

[0042] Calculate the spatial distance between each pair of nodes in the geotechnical data graph through the Euclidean distance.

[0043] For each pair of nodes, calculate the geological similarity measure between the nodes using geological features.

[0044] If the spatial distance between nodes is less than a preset threshold and the geological similarity of the nodes is higher than a second threshold, these nodes are regarded as neighbors.

[0045] If the spatial relationship and geological characteristics of the nodes meet the threshold conditions, they are connected; otherwise, the nodes are not connected.

[0046] Furthermore, the sparsification process includes:

[0047] For each edge in the graph, set the weight of the edge by calculating the spatial distance and geological similarity weight of the edge. The weight of the edge is a combined function of the spatial distance and geological similarity of the two nodes. The weight formula is as follows:

[0048]

[0049] where w ij is the edge weight between node i and node j, α and β are adjustment parameters, dij is the spatial distance, s ij is the geological similarity;

[0050] After calculating the edge weights, a sparsification threshold τ is set, and the edges with weights less than this threshold will be removed.

[0051] Furthermore, the spatial distance weight is calculated using the following method:

[0052]

[0053] where w d (i, j) represents the spatial distance weight between node i and node j, and d ij represents the Euclidean distance,

[0054] σ is a tuning parameter.

[0055] Furthermore, the geological similarity weight is calculated using the following method:

[0056]

[0057] where w g (i, j) represents the geological similarity weight between nodes i and j, and f i,k , f j,k respectively represent the values of the node on the k-th geological feature, and w k is the weighting factor for each geological feature.

[0058] The present invention proposes a method for analyzing geotechnical engineering exploration data based on a neural network. By constructing a geotechnical data graph and combining spatial proximity relationships and geological similarities, high-precision inference of geotechnical parameters in unknown areas is achieved. Compared with traditional geotechnical engineering exploration methods, the present invention has the following beneficial effects:

[0059] Through the multi-layer information propagation mechanism of the GNN, the information of the sampled area is fully utilized to infer the geotechnical parameters of the unsampled area.

[0060] By combining the spatial distance weight and the geological similarity weight, a more accurate geotechnical data graph is established, improving the prediction ability of the model in areas with complex geological conditions. Compared with traditional interpolation methods such as Kriging interpolation and inverse distance weighting (IDW), the present invention can more accurately depict the spatial distribution law of geotechnical characteristics and still maintain a high prediction accuracy in areas with large formation changes.

[0061] The dynamic threshold method is used to sparsify the geotechnical data map, optimize the map structure, reduce unnecessary edge connections, and improve the calculation efficiency. In the case of uneven distribution of exploration points, the GNN can learn geological patterns across regions, making the prediction results still highly reliable in data-sparse regions and reducing the impact of data non-uniformity.

[0062] The present invention intelligently analyzes geotechnical engineering exploration data through a graph convolutional neural network (GNN), realizes accurate inference of geotechnical parameters in unknown regions, and effectively solves problems such as uneven data distribution, low accuracy of traditional interpolation methods, low data utilization rate, and difficulty in real-time updating. This method reduces the cost of geotechnical exploration, improves the accuracy of geological prediction, enhances the data utilization efficiency and the interpretability of geological analysis, is applicable to multiple engineering scenarios such as foundation design, underground construction, and geological disaster prevention, and has broad engineering application value and popularization prospects. Brief Description of the Drawings

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0065] Next, a preferred description of the invention will be made in combination with the drawings and specific embodiments.

[0066] In one embodiment of the present invention, as Figure 1 shown, a method for analyzing geotechnical engineering exploration data based on machine learning is provided.

[0067] Geotechnical engineering refers to the research and application of geological conditions involved in engineering construction, including foundation design, slope stability analysis, underground engineering construction, etc., aiming to ensure the safety, stability, and long-term reliability of building structures.

[0068] Exploration data refers to the geotechnical physical and mechanical parameters obtained through on-site drilling, in-situ testing, laboratory testing, etc. during the geotechnical engineering exploration process, including but not limited to soil layer types, shear strength, porosity, groundwater level, permeability coefficient, seismic wave velocity, etc. This data can be used for geological modeling, foundation bearing capacity assessment, and engineering suitability analysis.

[0069] This embodiment constructs a geotechnical data graph and uses a graph neural network (GNN) to perform spatial modeling on the survey data, uses the information of known survey points to infer the geological parameters of unknown areas, and visualizes the analysis results through a geographic information system (GIS), thereby improving the comprehensive utilization of geotechnical engineering survey data and improving the efficiency and accuracy of engineering geological surveys. Specifically, the following steps are included:

[0070] First, geotechnical engineering survey data are obtained, cleaned, normalized and standardized, outliers are removed, and missing data are filled.

[0071] The geotechnical engineering investigation data include but are not limited to drilling exploration data, in-situ test data, laboratory test data and geological survey data, wherein the drilling exploration data include the drilling depth, soil layer distribution and its physical and mechanical parameters, the in-situ test data include the standard penetration test (SPT), the static penetration test (CPT), and the foundation bearing capacity test results, the laboratory test data include shear strength, permeability coefficient, liquid limit, plastic limit and other parameters, and the geological survey data include regional geological structure, fault distribution and groundwater level and other information.

[0072] The geotechnical engineering survey data is cleaned to remove outliers caused by sensor errors, manual recording errors or environmental interference, and missing data is filled based on historical data, engineering experience or statistical methods to ensure the integrity and consistency of the data set.

[0073] The processed data is normalized or standardized. For parameters with different dimensions, such as soil thickness (m), shear strength (kPa), porosity (dimensionless value), etc., maximum and minimum normalization, Z-score standardization or Box-Cox transformation are used to eliminate the influence of numerical scale on model training, so that all feature data are in the same order of magnitude, and the training stability and computational efficiency of the machine learning model are improved.

[0074] Through the above data processing steps, the geotechnical engineering survey data has high reliability and consistency, providing high-quality data input for subsequent spatial correlation modeling, machine learning training and geotechnical property prediction.

[0075] The sampling points are regarded as nodes of the geotechnical data graph, and the attributes of each node include its geographical location information, soil characteristics and mechanical parameters.

[0076] The present invention uses sampling points as nodes of the geotechnical data graph, where:

[0077] A sampling point refers to the geographical location where geotechnical data is obtained during geotechnical engineering investigation through means such as drilling, in-situ testing, or remote sensing measurement. This location corresponds to a specific set of investigation data, and its data sources include, but are not limited to, borehole layout points, static cone penetration test points, standard penetration test points, and other engineering measurement points.

[0078] A geotechnical data graph refers to a graph structure model constructed based on geotechnical investigation data. This model consists of multiple nodes (sampling points) and edges (spatial association relationships). Each node represents the geological information of an investigation point, while the edges represent the spatial correlation or geological similarity between different investigation points, thus forming a mathematical structure that can perform spatial analysis and reasoning.

[0079] A node refers to the basic unit in a geotechnical data graph. Each node corresponds to a sampling point and contains the characteristic information of this sampling point. This characteristic information includes, but is not limited to, geographical location information, soil characteristics, and mechanical parameters.

[0080] Among them, the geographical location information refers to the spatial coordinates of the sampling point, including longitude, latitude, and burial depth, which are used to determine the relative spatial position of the node in the geotechnical data graph.

[0081] Soil characteristics refer to the geotechnical physical properties at the sampling point, including, but not limited to, soil layer type, particle composition, water content, permeability, and other parameters, to characterize the soil or rock composition at this location.

[0082] Mechanical parameters refer to the engineering mechanical properties of the geotechnical mass, including, but not limited to, shear strength, compression modulus, foundation bearing capacity, elastic modulus, etc., which are used to evaluate the deformation, failure, and stability of the geotechnical mass under external loads.

[0083] In this step, by constructing sampling points as nodes of the geotechnical data graph and assigning corresponding geographical locations, soil characteristics, and mechanical parameters to the nodes, geotechnical data can be stored and analyzed in the form of a graph structure, providing basic data support for subsequent spatial association modeling, graph neural network learning, and geotechnical property prediction.

[0084] The Euclidean distance is used to determine the spatial proximity relationship between adjacent nodes, obtaining the spatial distance weight.

[0085] For the nodes in geotechnical engineering investigation data, the Euclidean distance represents the geographical location difference between nodes, that is, the straight-line distance between two points calculated based on the spatial coordinates (longitude, latitude, depth, etc.) of the nodes.

[0086] Adjacent nodes refer to two or more nodes with relatively close distances obtained by the Euclidean distance measurement method in the geotechnical data graph. These nodes are usually located in adjacent or connected areas in space. In the actual geological environment, their physical or geological properties may be relatively similar, so the interaction between them is relatively large. Adjacent nodes are connected to form edges in the geotechnical data graph through spatial distance and geological similarity.

[0087] Spatial proximity relationship refers to the geographical proximity between adjacent nodes and their connection in geological characteristics in the geotechnical data graph. By using the Euclidean distance, the relative positions between nodes can be quantitatively determined, so as to define whether there is a spatial proximity relationship between nodes. When the Euclidean distance between nodes is small and there is a certain geological similarity, their spatial proximity relationship is strong, indicating that they have a greater impact on the prediction of each other's geological parameters.

[0088] Spatial distance weight refers to the weight value between nodes determined according to the Euclidean distance. This weight value quantifies the influence degree of the spatial distance between nodes on the prediction of geological attributes. Generally speaking, the smaller the Euclidean distance, the greater the influence between nodes, and the larger the spatial distance weight; on the contrary, the larger the Euclidean distance, the smaller the influence between nodes, and the smaller the spatial distance weight. The spatial distance weight is usually used as the weight of the edge in the geotechnical data graph for the calculation of the graph convolutional neural network (GNN) model, so as to make more accurate inferences and predictions on geotechnical properties.

[0089] Through the above steps, obtaining the spatial distance weight provides a quantitative basis for the connection between nodes in the geotechnical data graph, enabling the graph neural network to more accurately transmit information through the geographical spatial relationship between nodes, thereby improving the analysis accuracy and efficiency of geotechnical engineering investigation data.

[0090] Based on soil layer type, shear strength, and porosity parameters, a geological similarity measurement between nodes is established to obtain the geological similarity weight.

[0091] In this step, based on soil layer type, shear strength, and porosity parameters, a geological similarity measurement between nodes is established and the geological similarity weight is obtained, where:

[0092] Soil layer type refers to the classification of soil or rock layers at different depths in the geotechnical body, usually divided according to factors such as particle composition, mineral composition, and sedimentary environment. Different soil layer types have different physical properties and mechanical characteristics, such as clay, sand, gravel, rock, etc. The classification of soil layer types helps to evaluate the stability of underground soil, bearing capacity, and its adaptability to engineering construction.

[0093] Shear strength refers to the ability of geotechnical materials to resist shear stress, which is usually jointly determined by the angle of internal friction and cohesion. In soil, shear strength mainly affects its failure behavior under external loads and is a very important parameter in soil mechanics. The higher the shear strength, the better the stability of the soil and the stronger its ability to resist deformation and slip.

[0094] Porosity is the ratio of the pore volume to the total volume in soil or rock, which reflects the permeability, density and pore size of the soil. Porosity has an important impact on the physical properties of soil such as compressive resistance, expansibility and water retention capacity, so it is one of the key parameters for evaluating the engineering properties of geotechnical bodies.

[0095] The measurement of geological similarity between nodes refers to the quantitative evaluation of the geological characteristics between adjacent nodes (sampling points) in a geotechnical data graph to measure their similarity in parameters such as soil layer type, shear strength, porosity, etc. Specifically, using parameters such as soil layer type, shear strength and porosity, through weighted similarity measurement formulas, such as Euclidean distance, cosine similarity or weighted Euclidean distance, etc., to calculate the similarity between adjacent nodes. Through this measurement method, the geological similarity between nodes can be obtained, that is, whether the geological characteristics of two nodes are similar or the degree of difference.

[0096] The geological similarity weight refers to the weight value obtained according to the measurement of geological similarity, which quantifies the similarity of geological characteristics between two nodes. In the present invention, the weighted combination of parameters such as soil layer type, shear strength, porosity, etc. will affect the geological similarity weight between nodes. Generally, nodes with higher similarity will have a larger geological similarity weight, while nodes with lower similarity will have a smaller weight. This weight value reflects the intensity of information transfer between nodes with similar geological characteristics and is used as the edge weight in the training process of the graph convolutional neural network (GNN).

[0097] Through the above calculation of the geological similarity measurement method and weight, the node relationship in the geotechnical data graph can be established more accurately, so that when the graph neural network (GNN) conducts geotechnical engineering data inference, it can effectively use the geological similarity between nodes for information propagation, thereby improving the accuracy and reliability of predicting geotechnical parameters in unknown areas.

[0098] Based on the spatial proximity relationship and geological similarity, determine the edges of the geotechnical data graph and assign edge weights to connect nodes that are spatially close and have similar geological characteristics.

[0099] In this step, based on the spatial proximity relationship and geological similarity, determine the edges of the geotechnical data graph and assign edge weights to connect nodes that are spatially close and have similar geological characteristics. The specific steps are as follows:

[0100] The edges of the geotechnical data graph refer to the connections established between the nodes (sampling points) in the geotechnical data graph through spatial proximity relationships and geological similarity. Specifically, the edges between nodes indicate that they are spatially close and have similarities in geological properties. The existence of these edges suggests the possibility of information propagation between adjacent nodes, forming a graph structure representing geotechnical data relationships.

[0101] Edge weight refers to the connection strength or importance of an edge in the geotechnical data graph, which is jointly determined by spatial proximity relationships and geological similarity. The spatial distance weight and geological similarity weight are comprehensively assigned to each edge through a weighted method, and the magnitude of the weight value reflects the connection strength of the edge. An edge with a larger weight indicates a stronger similarity and stronger influence between nodes, while an edge with a smaller weight indicates a weaker connection and a lower correlation between nodes.

[0102] By comprehensively considering the spatial distance and geological similarity between nodes, this step can more accurately establish the association relationship between nodes, enabling each edge in the geotechnical data graph to truly reflect the similarity of physical and geological characteristics between nodes. This method is more comprehensive and accurate than simply considering spatial distance or a single geological feature.

[0103] The geological characteristics of geotechnical engineering usually exhibit spatial aggregation and local similarity, that is, the geological characteristics of adjacent regions tend to be similar. Through the design of this step, this geological phenomenon can be better captured, making the inference results conform to the actual engineering geological situation.

[0104] Adopt the dynamic threshold method to determine the number of neighbors of each node, and sparsify the geotechnical data graph according to the spatial relationship.

[0105] The dynamic threshold method refers to dynamically adjusting the number of neighbors of each node according to the geographical location, soil characteristics of the nodes in the geotechnical data graph, as well as the spatial distance and geological similarity of adjacent nodes. Different from the static fixed threshold method, the dynamic threshold method can flexibly adjust the connection strength and the number of neighbors according to the specific geological attributes and relative positions of the nodes, so that the connection of each node is more reasonable and accurate. The setting of the dynamic threshold usually takes into account the spatial distance, geological similarity and other relevant parameters between nodes to ensure that the connections between nodes can effectively reflect geological characteristics without causing the graph to be too dense or too sparse.

[0106] Specifically, the implementation of the dynamic threshold method includes:

[0107] First, calculate the spatial distance between each pair of nodes in the geotechnical data graph through the Euclidean distance or other appropriate distance metrics. The spatial distance reflects the degree of geographical proximity, and the formula is as follows:

[0108]

[0109] where d ij is the spatial distance between node i and node j, and x i , y i are the geographical coordinates of node i, and x j , y j are the geographical coordinates of node

[0110] j.

[0111] For each pair of nodes, the geological similarity measure between the nodes is calculated using geological features (such as soil layer type, shear strength, porosity, etc.). Common geological similarity measure methods include eigenvalue-based Euclidean distance, cosine similarity, or weighted distance measure. The geological similarity measure formula is:

[0112]

[0113] where s ij is the geological similarity between node i and node j, f i,k , f j,k are the values of node

[0114] i and node j on the k-th feature respectively, and w k is the weight of this feature.

[0115] According to the spatial distance and geological similarity of each pair of nodes, a dynamic threshold is set. The calculation of the threshold is based on the following rules:

[0116] If the spatial distance between nodes is less than a certain threshold and the geological similarity of the nodes is high, then these nodes can be regarded as neighbors.

[0117] This threshold can be determined by the nearest neighbor algorithm or the K-nearest neighbor algorithm (K-NN), and can be adaptively adjusted for different geological features.

[0118] By comparing the calculated spatial distance with the dynamic threshold, it is judged whether the two nodes are taken as neighbors. If the spatial relationship and geological characteristics of the nodes meet the threshold conditions, they are connected. Otherwise, the nodes are not connected. In this way, the number of neighbors is no longer fixed, but dynamically adjusted according to the specific geological attributes and spatial relationships of the nodes.

[0119] The number of neighbors refers to the number of directly connected nodes of each node in the geotechnical data graph, or the strength of the relationship between each node and other nodes. The determination of the number of neighbors directly affects the propagation range of node information. In this step, the number of neighbors is dynamically adjusted according to the spatial relationship and geological similarity, which can ensure that each node has appropriate connections with neighbors in a similar geological environment.

[0120] Spatial relationships refer to the relative connections established among various nodes in the geotechnical data graph based on geological parameters such as spatial distance, soil layer type, shear strength, porosity, etc. These spatial relationships are quantified by factors such as distance and similarity weights to form the edges and their weights in the graph.

[0121] Sparsification processing means that in the geotechnical data graph, by dynamically adjusting the number of neighbors of each node, unnecessary edge connections are reduced, thereby reducing the density of the graph. The purpose of sparsification processing is to remove some redundant connections or the edges between insignificant nodes in the graph to avoid excessive increase in computational complexity and improve the efficiency of model training. In this step, sparsification processing ensures that only important connections exist between nodes, and these connections reflect the actual geological similarity and spatial relationships between nodes.

[0122] After obtaining the adjacency relationship of the geotechnical data graph, the goal of sparsification processing is to reduce the number of edges in the graph by removing unimportant connections, making the calculation process more efficient while keeping the main structural features of the graph unchanged. The specific implementation steps are as follows:

[0123] For each edge in the graph, the weight of the edge is set by calculating the spatial distance and the geological similarity weight of the edge. The weight of the edge is a combined function of the spatial distance and geological similarity of the two nodes. The weight formula is as follows:

[0124]

[0125] where, w ij is the edge weight between node i and node j, α and β are adjustment parameters, d ij is the spatial distance, and s ij is the geological similarity.

[0126] After completing the calculation of the edge weights, a sparsification threshold τ is set. The edges with weights less than this threshold will be removed. This threshold can be adjusted according to experimental data, the requirements of the model, and the complexity of the graph. A lower threshold will result in more edges being removed and a stronger sparsity of the graph; a higher threshold will retain more connections.

[0127] By comparing the weight of the edge with the set sparsification threshold, for each edge, if its weight is less than the threshold, then the edge is deleted. Through sparsification processing, the number of edges in the graph is significantly reduced, reducing the computational complexity.

[0128] Finally, the geotechnical data graph after sparsification processing only retains important node connections, the structure of the graph is more streamlined, and at the same time, the main geological relevance between nodes is maintained. The geotechnical data graph after sparsification processing is suitable for subsequent training and prediction of graph convolutional neural networks (GNNs).

[0129] Through the dynamic threshold method and sparsification processing, the number of edges in the geotechnical data graph is significantly reduced, thereby reducing the computational burden of the graph convolutional neural network (GNN) and improving the training and inference speeds. Sparsification processing removes unimportant edges in the graph and only retains key node relationships, making the data graph more concise and removing irrelevant node connections. By dynamically adjusting the number of neighbors of nodes, the adjacency relationship is ensured to be more in line with the actual geological characteristics, improving the prediction accuracy of the model.

[0130] The sparsified geotechnical data graph is weighted and fused according to the spatial distance weight and geological similarity weight to form the final geotechnical data graph.

[0131] The sparsified geotechnical data graph refers to the geotechnical data graph that has undergone the dynamic threshold method and sparsification processing, removing redundant edges and optimizing the node connection relationship. This data graph only retains important spatial proximity relationships and node connections with strong geological similarities, thereby reducing the computational complexity and retaining the core structural characteristics of the geotechnical data.

[0132] The spatial distance weight refers to the weight value calculated based on the geographical location information of nodes, and this weight is used to quantify the strength of the spatial relationship between different nodes.

[0133] Furthermore, the spatial distance weight in this embodiment is calculated using the following method:

[0134]

[0135] where, w d (i,j) represents the spatial distance weight between nodes i and j, d ij represents the Euclidean distance,

[0136] σ is a regulation parameter. Nodes with smaller distances have a greater influence on each other, so the weight is higher, while nodes with larger distances have a smaller influence and the weight is relatively lower.

[0137] The geological similarity weight refers to the weight value calculated based on geological parameters such as soil layer type, shear strength, and porosity in geotechnical engineering investigation data, and this weight is used to quantify the similarity in geological attributes between two nodes.

[0138] Furthermore, the geological similarity weight in this embodiment is calculated using the following method:

[0139]

[0140] where, w g (i,j) represents the geological similarity weight between nodes i and j, f i,k , f j,k respectively represent the values of the node on the kth geological feature, wk is the weighting factor for each geological feature.

[0141] Weighted fusion refers to the weighted calculation of the edges in the geotechnical data map by integrating the spatial distance weight and the geological similarity weight, so as to ensure that the final data map can reflect both the influence of spatial position and the contribution of geological similarity. The formula for weighted fusion is as follows:

[0142] W final (i, j) = α · W d (i, j) + β · W g (i, j)

[0143] Among them, W final (i, j) is the final edge weight, W d (i, j) is the spatial distance weight, W g (i, j) is the geological similarity weight,

[0144] α and β are weighting coefficients to balance the influence of the two, and the specific values can be optimized through experiments.

[0145] The final geotechnical data map refers to the optimized geotechnical data map generated after weighted fusion. Compared with the original data map, this data map not only reduces invalid connections, improves the calculation efficiency, but also enhances the ability to express the spatial structure of geotechnical data, enabling more accurate feature inference and prediction during information propagation in the graph neural network (GNN).

[0146] Adopt a graph convolutional neural network to perform information propagation using the initial features and adjacency relationships of nodes.

[0147] A graph convolutional neural network is a neural network model based on graph-structured data, which can learn using the topological structure of the graph and node features, and update the features of each node through information aggregation of adjacent nodes. In the present invention, GNN is used for spatial information inference of geotechnical engineering investigation data to learn the spatial relationships and geological similarities between nodes, and then infer the geological parameters of unsampled points.

[0148] The initial features of a node refer to the attributes possessed by each sampling point in the geotechnical data map, including but not limited to:

[0149] Geographical location information (such as longitude, latitude, depth)

[0150] Soil layer type (such as sand, clay, silt, etc.)

[0151] Mechanical parameters (such as shear strength, compression modulus, foundation bearing capacity)

[0152] Physical parameters (such as porosity, permeability coefficient, water content)

[0153] These initial features are used as input data for the training and prediction of the neural network.

[0154] The adjacency relationship refers to the spatial connection relationship between nodes in the geotechnical data graph, that is, which nodes are interconnected and the edge weights between them. This relationship is jointly determined by the spatial distance weight and the geological similarity weight, representing the degree of information propagation between nodes. In the GNN model, the adjacency relationship is stored and calculated through the adjacency matrix.

[0155] This step aggregates information from neighbor nodes and updates the feature representation of the current node. The update formula is as follows:

[0156]

[0157] Where:

[0158] represents the new feature of node i in the (t + 1)-th layer;

[0159] is the feature of neighbor node j in the t-th layer;

[0160] W (t) is the trainable weight matrix in the t-th layer;

[0161] Z ij is the normalization factor to prevent the numerical value from becoming too large when features accumulate;

[0162] N(i) is the set of neighbor nodes of node i;

[0163] σ(·) is the non-linear activation function

[0164] Through multi-layer information propagation, GNN can make the features of each node contain not only its own information but also the information of its neighbor nodes, thereby enhancing the inference ability for geotechnical data in unknown areas.

[0165] In this step, through the graph convolutional neural network (GNN), information propagation is carried out using the initial features of nodes and the adjacency relationship, so that the geological attributes of each node not only contain its own information but also combine the geological characteristics of the adjacent areas, realizing the efficient inference of geotechnical parameters in unknown areas. This step can improve the utilization efficiency of spatial information in geotechnical exploration data, optimize the prediction results of geological parameters, and provide more accurate data support for geotechnical engineering design and construction.

[0166] Through multi-layer graph convolutional calculations, the neighbor information of each node is aggregated, and the node features are gradually updated, so that the geotechnical parameters in the unknown area are obtained from the information propagation of known points.

[0167] In this step, for each sampled point (node) in the geotechnical data map, initial features such as its geographical location, soil characteristics, and mechanical parameters are assigned. For nodes in unknown areas, the initial features are initialized as null values or filled with the mean value.

[0168] For each node i, geological information is obtained from the set of its directly connected neighbor nodes N(i) and weighted summation is performed.

[0169] By increasing the number of graph convolutional layers, each node can obtain information from more distant neighbors, enabling features to spread across multiple regions.

[0170] After multiple layers of information propagation, the geotechnical parameters in unknown areas are gradually obtained through the information propagation from known sampled points.

[0171] After multiple layers of information propagation, the trained model is used to predict the geotechnical parameters in unknown areas, and the final distribution of geological characteristics is output.

[0172] Through multi-layer graph convolutional calculations, the final representation of each node not only contains its own information but also integrates the geological characteristics of spatially adjacent points, thereby improving the prediction accuracy of geotechnical parameters in unknown areas. Traditional geotechnical engineering exploration methods usually rely on limited drilling points and it is difficult to infer the geological conditions in un-sampled areas. GNN can expand the information from the sampled areas to the un-explored areas through information propagation, thereby improving the spatial utilization rate of data. Since GNN can learn the geological similarity and spatial relationships between nodes, even when there is less exploration data, it can infer the attributes of unknown points through the geological information of known points, thus reducing the cost and time of on-site exploration.

[0173] Semi-supervised learning or supervised learning methods are adopted to train the GNN model with the goal of minimizing the prediction error.

[0174] Semi-supervised learning is a machine learning method that combines labeled data and unlabeled data for training, and is suitable for situations where there is less labeled data but more unlabeled data. In the analysis of geotechnical engineering exploration data, due to the limited number of sampling points and the unknown geotechnical parameters in some areas, semi-supervised learning can make full use of the guiding role of labeled data and at the same time use unlabeled data to improve the generalization ability of the model.

[0175] Minimizing the prediction error means that during the model training process, by optimizing the loss function, the predicted values of the model for geotechnical parameters are made as close as possible to the true values, reducing the prediction deviation. Common loss functions include mean square error (MSE), cross-entropy loss, L1 / L2 regularization, etc.

[0176] The specific implementation of this step includes:

[0177] Data Preparation

[0178] Select some sampling points as labeled data, including known geotechnical parameters (such as soil layer type, shear strength, porosity, etc.).

[0179] Select the unsampled area as unlabeled data, that is, the unknown nodes whose geotechnical characteristics need to be inferred.

[0180] Model Initialization

[0181] Construct a GNN model and set initial parameters, including the number of neural network layers, neighbor aggregation method, activation function, etc.

[0182] Semi-supervised Learning Training (if semi-supervised learning is adopted)

[0183] Adopt methods such as Label Propagation or Pseudo-labeling, and use a small amount of labeled data to guide the model's learning on unlabeled data.

[0184] Through the information propagation mechanism of GNN, let the geotechnical parameters of known sampling points affect the adjacent unknown areas, so as to infer the possible characteristics of unlabeled data.

[0185] Supervised Learning Training (if supervised learning is adopted)

[0186] Directly use the labeled exploration data as the training set, take the geotechnical parameters as labels, and train the GNN model.

[0187] Adopt the Backpropagation algorithm to adjust the neural network weights to reduce the prediction error.

[0188] Infer the geotechnical parameters of unknown areas through the trained GNN. Specifically include:

[0189] In the geotechnical data map, select the nodes in the unknown area. The geographical location information of this node is known, but its geotechnical parameters (such as shear strength, porosity, etc.) are unknown.

[0190] Through the GNN model, use the geological information of adjacent known areas to infer the geotechnical parameters of unknown areas.

[0191] Adopt the trained GNN model to perform multi-layer graph convolution calculations on the nodes in the unknown area, so that the nodes can aggregate geotechnical characteristic information from spatially adjacent areas.

[0192] Through graph convolution calculation, the node features of each unknown area are continuously updated, and finally the optimal geotechnical parameter estimation based on the known area is obtained.

[0193] After multiple layers of information propagation, the GNN model calculates the geotechnical properties of the nodes in the unknown area and standardizes all the inference results to ensure the numerical rationality.

[0194] The credibility of the inference results is evaluated by using cross-validation or probability distribution estimation to ensure that the inference results of the model are highly reliable in engineering applications.

[0195] Due to the uneven distribution of geotechnical exploration data, the sampling points are denser in some areas and sparser in some areas. The GNN inference method of the present invention can make full use of the propagation characteristics of data in the graph structure, learn geological patterns across regions, and enable the inference results to still maintain high reliability in data-sparse areas.

[0196] In one embodiment, the present invention further improves the foregoing geotechnical engineering exploration data analysis system based on machine learning, including:

[0197] A data module for acquiring geotechnical engineering exploration data, cleaning, normalizing, and standardizing the data, removing outliers, and filling in missing data;

[0198] A graph construction module for using the sampling points as the nodes of the geotechnical data graph, where the attributes of each node include its geographical location information, soil characteristics, and mechanical parameters; using the Euclidean distance to determine the spatial proximity relationship of adjacent nodes to obtain the spatial distance weight; based on the soil layer type, shear strength, and porosity parameters, establishing a geological similarity metric between nodes to obtain the geological similarity weight; according to the spatial proximity relationship and geological similarity, determining the edges of the geotechnical data graph and assigning edge weights so that nodes that are spatially close and have similar geological characteristics are connected; using the dynamic threshold method to determine the number of neighbors of each node, and sparsifying the geotechnical data graph according to the spatial relationship; performing weighted fusion on the sparsified geotechnical data graph according to the spatial distance weight and geological similarity weight to form the final geotechnical data graph;

[0199] A training module for using a graph convolutional neural network to perform information propagation using the initial features and adjacency relationships of the nodes; through multi-layer graph convolution calculations, aggregating the neighbor information of each node, gradually updating the node features, so that the geotechnical parameters in the unknown area are obtained from the information propagation of the known points; using semi-supervised learning or supervised learning methods to train the GNN model with the goal of minimizing the prediction error;

[0200] An inference module for inferring the geotechnical parameters of the unknown area through the trained GNN.

[0201] It should be noted that the explanatory descriptions of the foregoing embodiments of the geotechnical engineering exploration data analysis method based on machine learning also apply to the device of the embodiments of the present application, and will not be repeated here.

[0202] To implement the above embodiments, an embodiment of the present application further provides a computer device, which is a schematic structural diagram of the computer device. When the instruction processor in the computer device executes, it implements the method for analyzing geotechnical engineering investigation data based on machine learning in the foregoing embodiments.

[0203] To implement the above embodiments, an embodiment of the present application further provides a non-transitory computer-readable storage medium. A computer program is stored in the non-transitory computer-readable storage medium. When it runs on a computer, it causes the computer to execute the method for analyzing geotechnical engineering investigation data based on machine learning in the foregoing embodiments.

[0204] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0205] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0206] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROM), random access memories (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.

[0207] As described above, it is only the specific implementation manner of the present application. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims. For the module structures that are not specifically defined in the present invention, the content recorded in the prior art shall prevail. The prior art mentioned in the foregoing background art part and specific embodiment part of the present invention can be used as a part of the present invention to understand the meaning of some technical features or parameters.

Claims

1. A data analysis method for geotechnical engineering investigation based on machine learning, characterized in that, The method includes the following steps: Obtain geotechnical engineering investigation data, clean, normalize, and standardize the data, remove outliers, and fill in missing data; Take the sampling points as the nodes of the geotechnical data graph, and the attributes of each node include its geographical location information, soil characteristics, and mechanical parameters; Use the Euclidean distance to determine the spatial proximity relationship between adjacent nodes and obtain the spatial distance weight; Based on soil layer type, shear strength, and porosity parameters, establish a geological similarity measure between nodes and obtain the geological similarity weight; According to the spatial proximity relationship and geological similarity, determine the edges of the geotechnical data graph and assign edge weights so that nodes with similar spatial proximity and geological characteristics are connected; Adopt the dynamic threshold method to determine the number of neighbors of each node and sparsify the geotechnical data graph according to the spatial relationship; Perform weighted fusion on the sparsified geotechnical data graph according to the spatial distance weight and geological similarity weight to form the final geotechnical data graph; Adopt a graph convolutional neural network and use the initial features and adjacency relationships of the nodes for information propagation; Through multi-layer graph convolution calculation, aggregate the neighbor information of each node, gradually update the node features, and obtain the geotechnical parameters of unknown areas through the information propagation of known points; Adopt semi-supervised learning or supervised learning methods to train the GNN model with the goal of minimizing the prediction error; Infer the geotechnical parameters of unknown areas through the trained GNN.

2. The data analysis method for geotechnical engineering investigation based on machine learning according to claim 1, characterized in that The implementation of the dynamic threshold method includes: Calculate the spatial distance between each pair of nodes in the geotechnical data graph through the Euclidean distance; For each pair of nodes, calculate the geological similarity measure between the nodes using geological features; If the spatial distance between nodes is less than a preset threshold and the geological similarity of the nodes is higher than the second threshold, these nodes are regarded as neighbors; If the spatial relationship and geological characteristics of the nodes meet the threshold conditions, connect them; otherwise, do not connect between the nodes.

3. The method for analyzing geotechnical engineering investigation data based on machine learning according to claim 1, wherein, The sparsification process includes: For each edge in the graph, set the weight of the edge by calculating the spatial distance and geological similarity weight of the edge. The weight of the edge is a combined function of the spatial distance and geological similarity of the two nodes, and the weight formula is as follows: Among them, w ij is the edge weight between node i and node j, α and β are adjustment parameters, and d ij is the spatial distance, and s ij is the geological similarity; After completing the edge weight calculation, set a sparsification threshold τ, and remove the edges whose weights are less than this threshold.

4. The data analysis method for geotechnical engineering investigation based on machine learning according to claim 1, characterized in that The spatial distance weight is calculated using the following method: where w d (i, j) represents the spatial distance weight between node i and node j, d ij represents the Euclidean distance, and σ is a tuning parameter.

5. The method for analyzing geotechnical engineering exploration data based on machine learning according to claim 1, characterized in that The geological similarity weight is calculated using the following method: Among them, w g (i, j) represents the geological similarity weight between nodes i and j, f i,k , f j,k respectively represent the values of the node on the k-th geological feature, w k is the weighting factor for each geological feature.

6. A geotechnical engineering investigation data analysis system based on machine learning, characterized in that, The system includes the following modules: A data module for obtaining geotechnical engineering investigation data, cleaning, normalizing, and standardizing the data, removing outliers, and filling in missing data; A graph construction module for taking the sampling points as the nodes of the geotechnical data graph, where the attributes of each node include its geographical location information, soil characteristics, and mechanical parameters; using the Euclidean distance to determine the spatial proximity relationship between adjacent nodes and obtaining the spatial distance weight; Based on soil layer type, shear strength, and porosity parameters, establish a geological similarity measure between nodes and obtain the geological similarity weight; according to the spatial proximity relationship and geological similarity, determine the edges of the geotechnical data graph and assign edge weights so that nodes with similar spatial proximity and geological characteristics are connected; Using the dynamic threshold method, determine the number of neighbors of each node, and sparsify the geotechnical data map according to the spatial relationship; For the sparsified geotechnical data map, perform weighted fusion according to the spatial distance weight and the geological similarity weight to form the final geotechnical data map; The training module is used to adopt a graph convolutional neural network to perform information propagation using the initial features and adjacency relationships of the nodes; through multi-layer graph convolution calculations, aggregate the neighbor information of each node, gradually update the node features, so that the geotechnical parameters of the unknown area are obtained by the information propagation of the known points; adopt semi-supervised learning or supervised learning methods to train the GNN model with the goal of minimizing the prediction error; The inference module is used to infer the geotechnical parameters of the unknown area through the trained GNN.

7. The geotechnical engineering investigation data analysis system based on machine learning according to claim 6, characterized in that, The implementation of the dynamic threshold method includes: Calculate the spatial distance between each pair of nodes in the geotechnical data map through the Euclidean distance; For each pair of nodes, calculate the geological similarity measure between the nodes using geological features; If the spatial distance between the nodes is less than the preset threshold and the geological similarity of the nodes is higher than the second threshold, these nodes are regarded as neighbors; If the spatial relationship and geological characteristics of the nodes meet the threshold conditions, connect them; otherwise, do not connect between the nodes.

8. The data analysis system for geotechnical engineering investigation based on machine learning according to claim 6, wherein, The sparsification process includes: For each edge in the graph, set the weight of the edge by calculating the spatial distance and the geological similarity weight of the edge. The weight of the edge is a combined function of the spatial distance and the geological similarity of the two nodes. The weight formula is as follows: where w ij is the edge weight between node i and node j, α and β are adjustment parameters, d ij is the spatial distance, and s ij is the geological similarity; After completing the edge weight calculation, set a sparsification threshold τ, and the edges with weights less than this threshold will be removed.

9. The geotechnical engineering investigation data analysis system based on machine learning according to claim 6, wherein The spatial distance weight is calculated using the following method: Among them, w d (i, j) represents the spatial distance weight between node i and node j, d ij represents the Euclidean distance, and σ is the adjustment parameter.

10. The geotechnical engineering investigation data analysis system based on machine learning according to claim 6, characterized in that, The geological similarity weight is calculated using the following method: Among them, w g (i, j) represents the geological similarity weight between nodes i and j, f i,k , f j,k respectively represent the values of the node on the k-th geological feature, w k is the weighting factor of each geological feature.

Citation Information

Cited By

  • Geotechnical engineering intelligent reconnaissance system and method based on big data

    CN120742444A

  • Slope deformation monitoring and dynamic early warning method and system based on multi-sensor data

    CN120808544A

  • Geological survey method and system based on artificial intelligence

    CN121637193A