A method for structured conversion of geotechnical data
By using information particle deconstruction and a dual-track reasoning model, geotechnical data is transformed into structured data, solving the problem of chaotic geotechnical data formats, realizing the deep integration of geological information and building information, and improving the efficiency of data sharing and collaborative decision-making in engineering design and construction.
Patent Information
- Application Number
- CN202511163558.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In the current field of geological and mining engineering, geotechnical data exists in an unstructured form, with data scattered and in a chaotic format, making it difficult to use directly for modeling and analysis. There are barriers between the data formats of geological models and BIM models, resulting in the separation of geological information and architectural information, which affects the efficiency of data sharing and collaborative decision-making in engineering design and construction.
By employing information granular deconstruction, adaptive community discovery algorithm, and dual-track inference model, geotechnical log files are converted into structured data. By constructing a geotechnical knowledge graph and mapping it to the extended IFC format, the matching of geological entities with BIM models is achieved.
It has achieved standardized conversion of geotechnical data, broken through data format barriers, improved the efficiency of data sharing and collaborative decision-making in engineering design and construction, and reduced modeling errors and engineering risks.
Smart Images

Figure CN120723832B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data processing, and particularly relates to a geotechnical data structured conversion method. BACKGROUND
[0002] In the current field of geology and mining engineering, a large amount of geotechnical data exists in the form of log files and other unstructured forms, the data is scattered and the formats are chaotic, and it is difficult to be directly used for modeling and analysis; at the same time, there are barriers between the data formats of geological models and BIM models, the existing IFC format lacks special entities adapted to the properties of strata, and the geological core data (geometry, attributes, spatial relationship) is difficult to be effectively mapped to the BIM model, which causes the separation of geological information and building information, seriously affects the data sharing and collaborative decision-making efficiency in engineering design and construction, and increases the modeling error and engineering risk. SUMMARY
[0003] The application proposes a geotechnical data structured conversion method aiming at the technical problems existing in the above background technology.
[0004] In order to achieve the above purpose, the technical scheme adopted by the application comprises the following steps:
[0005] An unstructured data set of geotechnical log files is obtained, information granules are deconstructed on the data set, the data is disassembled into minimum information units containing entity clues, attribute clues and relationship clues, and each information unit is marked with clue type and confidence;
[0006] A clue association network is constructed, the association weight between information units is calculated based on clue co-occurrence frequency and semantic consistency, and a weighted network topology is formed;
[0007] An adaptive community discovery algorithm is used to cluster the weighted network topology, and a plurality of clustering communities are obtained;
[0008] A double-track reasoning model is constructed, and the model comprises a clue aggregation layer and a knowledge completion layer:
[0009] The clue aggregation layer performs feature fusion on the information units in each clustering community to generate an entity candidate set and a relationship candidate set;
[0010] The knowledge completion layer is used for logical verification and missing information completion on the entity candidate set and the relationship candidate set, and outputs standardized entity results and relationship results;
[0011] A geotechnical knowledge graph is constructed based on the entity results and the relationship results, and the knowledge graph is mapped to a structured data format;
[0012] An initial geological model is constructed using structured data, core data in the initial geological model is extracted, including stratum geometric data, attribute data and spatial relationship data, and is converted into an intermediate data file conforming to an extended IFC format, the extended IFC format adding geological special entities to adapt to stratum specific attributes;
[0013] The intermediate data file is associated with a BIM model, geological special entities are matched with component attributes in the BIM model through data mapping rules, and data capable of generating a BIM model containing geological information is obtained.
[0014] As preferred, information granular deconstruction is performed on the data set, and the data is disassembled into the smallest information unit containing entity clues, attribute clues and relationship clues, including:
[0015] Natural language processing technology is used to perform word segmentation and part-of-speech tagging on the unstructured data of the geotechnical log file, and noun phrases in the text are identified as entity clue candidates, quantity phrases and adjective phrases are identified as attribute clue candidates, and preposition phrases and verb phrases are identified as relationship clue candidates;
[0016] A pre-trained BERT model in the field of geotechnical engineering is used to classify the candidate clues and determine the clue type;
[0017] A confusion matrix is used to calculate the confidence of clue classification, which is obtained by the ratio of the number of correctly classified clues to the total number of clues, the total number of clues being the sum of the number of entity clues, attribute clues and relationship clues.
[0018] As preferred, the clue association network is constructed, and the association weight between information units is calculated based on clue co-occurrence frequency, semantic consistency and field logical association degree, and the specific implementation of forming a weighted network topology is:
[0019] First, determine the network nodes: the smallest information unit obtained by information granular deconstruction is used as the network node, and each node carries its clue type and confidence label;
[0020] Calculate the clue co-occurrence frequency: count the number of times that any two information units co-occur in the same text paragraph, table row or adjacent sentences in the unstructured data of the geotechnical log file, and use the ratio of this number to the total number of paragraphs in the data set as the co-occurrence frequency;
[0021] Calculate the semantic consistency: based on the pre-trained BERT model, extract the context vector corresponding to the text content of each information unit, and calculate the semantic consistency of two information units through cosine similarity;
[0022] Weighted fusion of clue co-occurrence frequency and semantic consistency is performed to obtain the association weight between information units;
[0023] A clue-based network is constructed using information units as nodes and association weights as edges. If the association weight of two information units is greater than or equal to a set weight threshold, a connection is established in the network, ultimately resulting in a weighted network topology containing association relationships.
[0024] Preferably, the specific implementation of using an adaptive community detection algorithm to cluster the weighted network topology to obtain several clustered communities is as follows:
[0025] First, initial seed nodes are generated based on weight density. The weight density of a node is the sum of the association weights between that node and its neighboring nodes, calculated as follows: ,in, Let i be the i-th information unit node in the network. Represents a node The set of adjacent nodes, Represents a node With adjacent nodes The correlation weight between them Represented as nodes Weight density;
[0026] Calculate the mean weight density of all nodes in the weighted network. and standard deviation Set dynamic threshold Weight density greater than or equal to The node is used as a candidate seed node;
[0027] Redundancy removal is performed on candidate seed nodes. If two candidate seed nodes... Association weight between If the number of nodes is greater than or equal to the set strong correlation threshold, then the nodes with higher weight density are retained as the initial cluster centers, ultimately forming the seed node set. Where m is the initial number of cluster centers;
[0028] compute nodes With cluster center The weighted distance is calculated as follows: ,in, The center node of the k-th cluster. For nodes With cluster center The correlation weight between them The average of the association weights among all nodes within the k-th cluster is used to divide the nodes based on the aforementioned weighted distance. Assign to the cluster with the smallest distance;
[0029] Define the k-th cluster as a community, and calculate the community cluster cohesion using the following method: ,in, a cohesion degree of a kth community cluster, a correlation weight of a node , a sum of total weights of the network, a correlation weight of a node and all adjacent nodes, a penalty coefficient, a number of nodes of the kth community;
[0030] if the cohesion degree of the kth community cluster is less than zero, performing a splitting operation according to a maximum weight difference, calculating a correlation weight difference between all nodes in the community, selecting a pair of nodes with the maximum difference as a new center, and splitting the original community into two sub-communities; if the cohesion degree of the community cluster after merging of two adjacent communities is improved to a set percentage threshold, performing a merging operation; repeating the splitting and merging process until a set iteration number is reached, and then stopping iteration.
[0031] As preferred, the thread aggregation layer performs feature fusion on information units in each clustered community, and the implementation of generating an entity candidate set and a relationship candidate set includes:
[0032] extracting core features of information units in each clustered community, the core features including text semantic features of entity threads, numerical and description features of attribute threads, and semantic association features of relationship threads, and introducing thread type labels and confidence of each information unit as auxiliary features;
[0033] based on the correlation weight between information units in the clustered community, adopting a weighted feature fusion strategy to combine entity threads and closely correlated attribute threads to form a preliminary entity candidate, and to match relationship threads and correlated entity threads to form a relationship candidate, thereby generating an entity candidate set and a relationship candidate set corresponding to each clustered community.
[0034] As preferred, the knowledge completion layer is configured to perform logical verification and missing information completion on the entity candidate set and the relationship candidate set, and the implementation of outputting standardized entity results and relationship results includes:
[0035] performing logical verification on the entity candidate set and the relationship candidate set, wherein the entity candidate set is checked for rationality of entity attributes, and the relationship candidate set is verified for logicality of relationships between entities;
[0036] based on high-confidence correlation threads in the thread correlation network and existing knowledge, reasoning and completing missing information found after verification;
[0037] outputting the standardized entity results and relationship results.
[0038] As preferred, the specific implementation of constructing an initial geological model by using structured data, extracting core data in the initial geological model, including stratum geometry data, attribute data and spatial relationship data, and converting into an intermediate data file conforming to the extended IFC format is as follows:
[0039] The extended IFC format is designed, and a geological special entity is newly added, wherein the geological special entity includes a stratum unit entity, a geological attribute set entity and a spatial topology entity.
[0040] A mapping relationship between the core data and the extended IFC format is established, the stratum geometry data is mapped to a geometric parameter of the stratum unit entity, the attribute data is mapped to a characteristic parameter of the geological attribute set entity, and the spatial relationship data is mapped to a correlation parameter of the spatial topology entity.
[0041] According to the mapping relationship, the extracted core data is converted into the intermediate data file conforming to the extended IFC format.
[0042] As preferred, the association between the intermediate data file and the BIM model includes: importing the intermediate data file into the BIM platform by using an IFC standard interface, and unifying the coordinate system of the geological model and the coordinate system of the BIM model by using a coordinate conversion algorithm to realize spatial position alignment.
[0043] Compared with the prior art, the advantages and positive effects of the present application are that: by information particle deconstruction and a double-track reasoning model, unstructured rock-soil log data is converted into standardized entities and relationships, the problems of scattered and chaotic data and difficulty in direct modeling are solved, and a structured basis is provided for geological analysis; the extended IFC format is innovatively designed, and a geological special entity is newly added, the data barrier between the geological model and the BIM model is broken, and effective mapping of stratum geometry, attribute and spatial relationship data is realized; through coordinate unification and attribute matching, geological information and building information are deeply integrated, data sharing and collaborative decision-making efficiency in engineering design and construction is improved, modeling error and engineering risk are reduced, and a new solution is provided for efficient application of geology and mining engineering data. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1 It is a structural flowchart of a rock-soil data structured conversion method. DETAILED DESCRIPTION
[0046] In order to enable the above-mentioned objects, features and advantages of the present application to be more clearly understood, further description will be made in conjunction with the accompanying drawings and examples. It should be noted that the examples of the present application and the features in the examples can be combined with each other without conflict.
[0047] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure aspects of the present application.
[0048] Embodiment, in the current field of geotechnical engineering, data management and application face two core problems. On the one hand, a large amount of data exists in the form of unstructured data such as log files and survey reports. These data contain key information such as stratum distribution, geotechnical mechanical properties, and geological structure. However, due to the chaotic format and inconsistent expression, it is difficult to directly use these data for modeling and analysis. On the other hand, there is a natural barrier between the data formats of geological models and building information models (BIM). The existing industry foundation class (IFC) standard lacks special entities that adapt to stratum-specific properties (such as permeability and compression modulus), resulting in the inability to effectively map geological core data (stratum geometry, physical and mechanical properties, and spatial topological relationships) to BIM models. Based on this, a method for structuring rock and soil data is proposed.
[0049] First, the unstructured data set of the rock and soil log file is obtained, the information particle of the data set is deconstructed, the data is disassembled into the smallest information unit containing entity clues, attribute clues and relationship clues, and each information unit is marked with clue type and confidence. The natural language processing technology is used to divide the unstructured data of the rock and soil log file and to mark the word class, the noun phrase in the text is identified as the entity clue candidate, the quantity phrase and the adjective phrase are identified as the attribute clue candidate, and the preposition phrase and the verb phrase are identified as the relationship clue candidate; the pre-trained BERT model in the field of geotechnical engineering is used to classify the candidate clues and determine the clue type; the confidence of the clue classification is calculated by using the confusion matrix, which is obtained by the ratio of the number of correctly classified clues to the total number of clues, and the total number of clues is the sum of the number of entity clues, attribute clues and relationship clues. Specifically, first, the unstructured text of the rock and soil log file is preprocessed to remove redundant symbols (such as redundant punctuation and repeated spaces) and standardize unit expressions. The word segmentation tool is used in combination with the professional dictionary in the field of rock and soil to perform word segmentation, and the HanLP tool is used to perform part-of-speech tagging. The tagging results include noun (n), quantity word (m), adjective (a), preposition (p), and verb (v) part-of-speech tags. Based on the tagging results, the clue candidates are identified: the fragments with noun or noun phrase (such as “n+n” and “n+ns”) part-of-speech tags are selected as entity clue candidates, such as “silty clay” and “drill hole ZK3”; quantity phrases (“m+unit”, such as “28%” and “5×10- The adjectives and adverbial phrases (e.g., “a” or “a+a”, such as “loose” “hard plastic”) are selected as attribute clue candidates; preposition phrases (e.g., “p+ noun phrase”, such as “under the loose fill”) and verb phrases (e.g., “v+ noun phrase”, such as “contacting the pebble layer”) are selected as relationship clue candidates. A pre-trained BERT model is selected, and 5000 rock-soil log samples labeled with “entity clue” “attribute clue” “relationship clue” labels are used for fine-tuning: the candidate clue text is converted into a word vector, and the classification probability is output through a fully connected layer after inputting the model, and the highest probability category is taken as the clue type (e.g., “silty clay” is classified as an entity clue). The confusion matrix is used to count the classification results: the matrix rows represent the actual categories (entity, attribute, relationship), and the columns represent the predicted categories, and the diagonal elements are the number of correctly classified clues. The confidence calculation formula is “the number of correctly classified clues ÷ (the number of entity clues + the number of attribute clues + the number of relationship clues)”. For example, if the total number of clues is 150, and 138 are correctly classified, then the confidence is 138 ÷ 150 = 0.92.
[0050] The information unit association weight is calculated based on the clue co-occurrence frequency and semantic consistency to form a weighted network topology. The specific implementation of constructing the clue association network, calculating the association weight between information units based on clue co-occurrence frequency, semantic consistency and field logical association degree, and forming a weighted network topology is as follows: first, determine the network nodes: the minimum information unit obtained by information granular decomposition is taken as the network node, and each node carries its clue type and confidence label; calculate the clue co-occurrence frequency: count the number of times that any two information units co-occur in the same text paragraph, table row or adjacent sentence in the unstructured data of the geotechnical log file, and take the ratio of the number of times to the total number of paragraphs in the data set as the co-occurrence frequency; calculate the semantic consistency: based on the pre-trained BERT model, extract the context vector corresponding to the text content of each information unit, and calculate the semantic consistency of the two information units by cosine similarity; weight and fuse the clue co-occurrence frequency and semantic consistency to obtain the association weight between information units; construct the clue association network with information units as nodes and association weights as edges: if the association weight of two information units is greater than or equal to a set weight threshold, a connection is established in the network, and finally a weighted network topology containing association relationships is obtained. Specifically, the minimum information unit obtained by information granular decomposition is taken as the network node, and the minimum information unit is an independent semantic fragment obtained after word segmentation, part-of-speech tagging and clue classification. Each node needs to carry two types of labels: one is the clue type label, which is clearly labeled as "entity clue", "attribute clue" or "relationship clue", such as the "soil filling" node labeled as "entity clue" and the "water content 28%" node labeled as "attribute clue"; the other is the confidence label, which is the confidence value of the information unit after classification. The node storage form adopts a structured data format (such as JSON), and an example is as follows: {"node ID": "N1", "text content": "silty clay", "clue type": "entity clue", "confidence": 0.93}. By traversing all the decomposition results, a node set containing all information units is generated to ensure that there are no duplicate or redundant nodes. Then, the clue co-occurrence frequency and semantic consistency are calculated, and the co-occurrence frequency and semantic consistency are fused by weighted summation. The weight coefficients are set based on domain experience, with a co-occurrence frequency weight of 0.6 and a semantic consistency weight of 0.4. The association weight calculation formula is: association weight = co-occurrence frequency weight x co-occurrence frequency + semantic consistency weight x semantic consistency. Finally, a weight threshold is set, and the association weight matrix is traversed. For any two nodes, if their association weight is greater than or equal to the threshold, a connection is established between them, and the weight value of the edge is the association weight. The network topology is stored in the form of an adjacency list, and each node entry contains its connected node ID and corresponding weight. The final generated weighted network topology needs to be visualized and verified to ensure that strongly associated nodes are correctly connected and weakly associated nodes are not redundantly connected.
[0051] Then an adaptive community discovery algorithm is used to cluster the weighted network topology to obtain a plurality of clustering communities; the specific implementation of using the adaptive community discovery algorithm to cluster the weighted network topology to obtain a plurality of clustering communities is as follows: first, an initial seed node is generated based on weight density, the weight density of the node is the sum of the association weights of the node and adjacent nodes, and the calculation method is wherein, is the i th information unit node in the network, denotes the adjacent node set of the node , denotes the association weight between the node and the adjacent node , denotes the weight density of the node ; the mean value and the standard deviation of the weight density of all nodes in the weighted network are calculated, a dynamic threshold is set, and the node with a weight density greater than or equal to is taken as a candidate seed node; the candidate seed node is de-duplicated, if the association weight between two candidate seed nodes is greater than or equal to a set strong association threshold, the node with a higher weight density is reserved as an initial clustering center, and finally a seed node set is formed, wherein m is the number of initial clustering centers; the weighted distance between the node and the clustering center is calculated, and the calculation method is wherein, is the center node of the k th cluster, is the association weight between the node and the clustering center , is the average value of the association weights between all nodes in the k th cluster, the node is assigned to the cluster with the smallest distance based on the above weighted distance; the k th cluster is defined as a community, and the community clustering cohesion is calculated, and the calculation method is wherein, denotes the cohesion of the k th community cluster, is the association weight of the node , is the sum of the total weights of the network, is the sum of the association weights of the node and all adjacent nodes, is a penalty coefficient, is the number of nodes of the k th community; if If the cohesion degree of the community cluster is less than zero, the community is split according to the maximum weight difference, the correlation weight difference between all nodes in the community is calculated, and the pair of nodes with the maximum difference is selected as the new center to split the original community into two sub-communities; if the cohesion degree of the community cluster after the two adjacent communities are merged is improved to a set percentage threshold, the merging operation is performed; the splitting and merging process is repeated until a set iteration number is reached, and then the iteration is stopped. If the cohesion degree of the community cluster is less than zero, the community is split according to the maximum weight difference, the correlation weight difference between all nodes in the community is calculated, and the pair of nodes with the maximum difference is selected as the new center to split the original community into two sub-communities; if the cohesion degree of the community cluster after the two adjacent communities are merged is improved to a set percentage threshold, the merging operation is performed; the splitting and merging process is repeated until a set iteration number is reached, and then the iteration is stopped.
[0052] A double-track reasoning model is constructed, which includes a clue aggregation layer and a knowledge completion layer. The clue aggregation layer fuses the features of information units in each clustering community to generate an entity candidate set and a relationship candidate set. The knowledge completion layer is used for logical verification and missing information completion of the entity candidate set and the relationship candidate set, and outputs standardized entity results and relationship results. The implementation of the clue aggregation layer fusing the features of information units in each clustering community to generate an entity candidate set and a relationship candidate set includes: extracting core features of information units in each clustering community, the core features including text semantic features of entity clues, numerical and description features of attribute clues, and semantic association features of relationship clues, and introducing clue type labels and confidence of each information unit as auxiliary features; based on the association weight between information units in the clustering community, a weighted feature fusion strategy is adopted to combine entity clues and closely associated attribute clues to form a preliminary entity candidate, and to combine relationship clues and associated entity clues to form a relationship candidate, to generate an entity candidate set and a relationship candidate set corresponding to each clustering community. Specifically, for information units in the clustering community, core features are extracted according to clue types and auxiliary features are fused. For entity clues, a pre-trained BERT model is used to extract a 768-dimensional vector at the [CLS] position as a text semantic feature after inputting the entity text, which accurately captures the unique semantics of the entity in the field. For attribute clues, numerical values are standardized to generate a 1-dimensional numerical feature, and description types are extracted by a BERT model to generate a 768-dimensional semantic vector as a description feature. For relationship clues, the context containing the entity pair is input, and a 768-dimensional vector output by BERT is extracted as a semantic association feature, which reflects the directionality of the relationship. At the same time, the clue type label (one-hot encoding of entity [1, 0, 0], attribute [0, 1, 0], and relationship [0, 0, 1]) and the confidence are introduced, and the core features are concatenated to form a 772-dimensional fusion feature vector, realizing the comprehensive integration of features. Based on the association weight of information units in the clustering community, a weighted strategy is used to generate a candidate set. For entity clues, an association weight threshold of 0.4 is set, the setting method of the association weight threshold is to calculate the weight distribution of a large number of known effective association and unrelated information unit samples, select a value that balances the association recognition accuracy and recall rate as the threshold, and select attribute clues with a weight greater than or equal to 0.4 (for example, the weight of “gravel layer (particle size 20-100mm, slightly dense)” is 0.53, and the weight of “slightly dense” is 0.47), and the entity and attribute features are weighted and fused according to the weight normalization value (0.53 / (0.53+0.47)=0.53) to form a preliminary entity candidate of “gravel layer (particle size 20-100mm, slightly dense)”, which includes entity name and attribute key-value pairs. For relationship clues, the associated entity pair (such as “silty clay” and “gravel layer”, with weights of 0.49 and 0.51) is located through an association network, the cosine similarity between the relationship feature and the two entity features is calculated (greater than or equal to 0.6 is considered as matching), and a relationship candidate of “(silty clay, in contact with, gravel layer)” is formed.Finally, the duplicate candidates are removed to generate the entity candidate set and the relationship candidate set of each community.
[0053] The knowledge completion layer is used for logical verification and missing information completion of the entity candidate set and the relationship candidate set, and the implementation of outputting the standardized entity result and the relationship result includes: performing logical verification on the entity candidate set and the relationship candidate set, wherein the rationality of the entity attribute is checked for the entity candidate set, and the logicality of the relationship between entities is verified for the relationship candidate set; for the missing information found after the verification, reasoning and completion are performed based on the high-confidence association clues in the clue association network and the existing knowledge; and the standardized entity result and the relationship result are outputted.
[0054] Specifically, in the logical verification stage, the attribute rationality of the entity candidate set is checked from three aspects. The first step is to perform numerical range verification according to a reasonable data range to check whether the attribute value is within a reasonable range. The second step is to perform attribute correlation verification to verify whether the attributes of the same entity are logically consistent. For example, the "clay content" of the "sandy soil" entity should be ≤ 3%. If "clay content 15%" is detected, it is determined to be contradictory. When the entity is labeled as "slightly dense" (degree of compaction), the "porosity ratio" should be between 0.7 and 0.9. If it exceeds this range, it is considered to be an attribute conflict. Finally, type matching detection is performed to match the attributes with the entity type. For example, "weathering degree" is an exclusive attribute of rock. If the "silt" entity carries this attribute, it is determined to be mismatched. For the relationship candidate set, the logicality of the relationship between entities is verified. First, the stratigraphic chronology principle is combined to check the stratigraphic sequence rationality to ensure that the entity relationship conforms to the sedimentary rule that old strata are below and new strata are above. At the same time, spatial topological consistency is verified to ensure that the relationship of the same entity pair is unique and conflict-free. For example, if entity A is labeled as being above B and also labeled as being horizontally adjacent to B without geological structure interpretation, it is considered to be a topological conflict and needs to be confirmed whether it conforms to the entity spatial physical correlation rule. In addition, it is also necessary to match the pre-set rock-soil field relationship library, such as "water-bearing layer must be adjacent to water-resisting layer" and "strongly weathered rock layer is mostly located above moderately weathered rock layer". If the relationship conflicts with the records in the library, it is marked as to be verified. All abnormal relationships need to be further reviewed in combination with the geological background to finally ensure that the relationship between entities conforms to the objective geological rules and logical rules. Then, for the missing information found after verification, high-confidence clues with an association weight greater than or equal to a set threshold are first screened from the clue association network. The confidence is obtained by the ratio of the number of correctly classified clues to the total number of clues. The total number of clues is the sum of the number of entity clues, attribute clues, and relationship clues. The existing knowledge is historical case data and industry specification data. The reasoning and completion based on high-confidence association clues and existing knowledge are as follows. First, high-confidence clues are screened from the clue association network. Then, structured historical case data and industry specification data are called according to the high-confidence clues. The industry specification provides mandatory constraints to define the boundary of the completion range. The historical data form a statistical rule library to determine the specific completion value of the completion result through probability analysis.
[0055] Based on the entity result and the relationship result, a geotechnical knowledge graph is constructed, and the knowledge graph is mapped to a structured data format. Specifically, the standardized entity result is taken as a node, each node containing an entity unique ID, a name, and a property set (such as the "plain fill" node containing ID "E1", the properties "buried depth 0-3m" and "dry density 1.6g / cm³"); the relationship result is taken as a directed edge, each edge containing a relationship type, a source entity ID, and a target entity ID (such as the "under" relationship corresponding to the edge: source "E2 (silty clay)", target "E1 (plain fill)"). The nodes and edges are stored through a graph database, forming a geotechnical knowledge graph containing entities, properties, and relationships. When mapped to a structured data, an entity table and a relationship table are generated: the entity table contains "entity ID", "name", and "property JSON" fields (such as the E1 corresponding row: "E1", "plain fill", "{'buried depth': '0-3m', 'dry density': '1.6g / cm³'}"); the relationship table contains "relationship ID", "source entity ID", "target entity ID", and "relationship type" fields (such as the R1 corresponding row: "R1", "E2", "E1", "under"), realizing accurate conversion of the graph to a structured format.
[0056] The application discloses a method for constructing an initial geological model by using structured data, extracting core data in the initial geological model, including stratum geometry data, attribute data and spatial relationship data, and converting the core data into an intermediate data file conforming to an extended IFC format, wherein the extended IFC format adds geological special entities to adapt to stratum specific attributes. The specific implementation of constructing the initial geological model by using the structured data, extracting the core data in the initial geological model, including the stratum geometry data, the attribute data and the spatial relationship data, and converting the core data into the intermediate data file conforming to the extended IFC format is as follows: designing the extended IFC format, adding the geological special entities, wherein the geological special entities include a stratum unit entity, a geological attribute set entity and a spatial topology entity; establishing a mapping relationship between the core data and the extended IFC format, mapping the stratum geometry data to geometric parameters of the stratum unit entity, mapping the attribute data to characteristic parameters of the geological attribute set entity, and mapping the spatial relationship data to associated parameters of the spatial topology entity; and converting the extracted core data into the intermediate data file conforming to the extended IFC format according to the mapping relationship. Specifically, when designing the extended IFC format, three types of special entities are added by focusing on the geological data characteristics. The stratum unit entity is used for representing the physical form of the stratum and contains basic identification information such as layer number, stratum name and formation age, as well as geometric description parameters such as layer boundary defined by a three-dimensional coordinate point set, layer thickness and dip angle, so as to accurately define the spatial range of a single stratum. The geological attribute set entity is specially used for carrying rock and soil mechanics and physical characteristics, and covers quantitative attributes such as water content, porosity ratio, compression modulus and internal friction angle, and qualitative attributes such as soil description, and is bound to the stratum unit entity through an attribute set association mechanism. The spatial topology entity is used for describing the spatial relationship between strata and contains contact types of adjacent strata, upper and lower stacking sequences and spatial association with special geological structures, and records the relationship direction and strength through a topology node and an edge structure. The mapping relationship construction needs to realize accurate correspondence between the core data and the special entity parameters. In the stratum geometry data, three-dimensional coordinates (X, Y, Z) of a layer surface are mapped to a "boundary coordinate set" parameter of the stratum unit entity, a layer thickness value is mapped to a "thickness" parameter, and a layer surface elevation is mapped to an "elevation reference" parameter, so as to ensure digital restoration of the geometric form. In the attribute data, physical properties (such as natural density 1.8 g / cm³) of rocks and soils are mapped to a "density" field of the geological attribute set entity, mechanical properties are mapped to a "compression modulus" field, and soil class names are mapped to a "soil type" field, so as to form a complete attribute parameter correspondence chain. In the spatial relationship data, the relationship between strata is mapped to a "vertical sequence" parameter of the spatial topology entity, a contact type (such as "gradual contact") is mapped to a "contact characteristic" parameter, and a spatial distance from a fault is mapped to a "structure association distance" parameter, so as to realize structured expression of spatial logic.The conversion process is executed in three steps: first, the core data is extracted from the initial geological model, classified according to the "stratum geometry-attribute-spatial relationship", and the original data is standardized and the coordinate data is converted to the WGS84 coordinate system; second, based on the preset mapping relationship, the classified data is filled into the corresponding field of the extended IFC format one by one, for example, the "layer boundary coordinates" of a certain silty clay layer are filled into the "boundary coordinate set" of the stratum unit entity, and the "water content 25%" is filled into the "water content" field of the geological attribute set entity; finally, according to the syntax rules of the IFC format (such as entity instantiation, attribute assignment statement), the data is organized, and an intermediate data file containing three types of entities of stratum unit, attribute set and topological relationship is generated, the file structure integrity is verified by a format checking tool to ensure that there is no field missing or format error.
[0057] The intermediate data file is associated with the BIM model, the intermediate data file is imported into the BIM platform by adopting an IFC standard interface, the coordinate system of the geological model is unified with the coordinate system of the BIM model by adopting a coordinate conversion algorithm, and spatial position alignment is realized. The geological special entity and the component attribute in the BIM model are matched by adopting a data mapping rule, and data capable of generating a BIM model containing geological information are obtained. Specifically, associating the intermediate data file with the BIM model needs to be realized in three steps. First, data import is completed by adopting the IFC standard interface: the intermediate data file in the extended IFC format is directly imported by adopting the IFC general interface supported by the BIM platform (such as Revit and Bentley). The interface automatically parses the structural information of the stratigraphic unit entity, the geological attribute set entity and the spatial topology entity in the file, ensures that the hierarchical relationship (such as the binding of the stratum and the attribute and the pointing relationship of the topological relationship) of the geological data is completely retained in the BIM environment, and avoids data loss or structural disorder. Second, spatial position alignment is realized by adopting the coordinate conversion algorithm: the geodetic coordinate system is usually adopted by the geological model, and the construction coordinate system (such as the site relative coordinate system) is usually adopted by the BIM model, so the seven-parameter coordinate conversion method is needed to unify the coordinate systems. First, three or more common points (such as the drilling position and the site boundary corner point) are selected, the coordinate values of the common points in the two coordinate systems are obtained, the translation amount, the rotation angle and the scaling factor are calculated; and then the three-dimensional coordinates of the geological model are batch-converted based on the parameters, so that the spatial position (such as the buried depth and the strike) of the stratum and the building component in the BIM model accurately correspond in the same coordinate system, and it is ensured that there is no deviation when the two are spatially superimposed. Finally, the entity and the component attribute matching are completed by adopting the data mapping rule: the attribute mapping table of the geological special entity and the BIM component is formulated, for example, the geometric parameters such as the “layer thickness” and the “dip angle” of the stratigraphic unit entity are mapped to the “geometric feature” attribute group of the BIM geological component; the parameters such as the “compression modulus” and the “water content” of the geological attribute set entity are mapped to the “physical and mechanical attribute” field of the BIM component; and the relationships such as the “overlying” and the “contact” of the spatial topology entity are mapped to the “spatial correlation” attribute of the BIM. The mapping rule is executed by adopting a script tool (such as Dynamo), the geological information is embedded into the attribute structure of the BIM model, and finally the BIM model containing complete geological data is generated.
[0058] The above merely describes preferred embodiments of the present application and is not intended to limit the present application in other forms. Any skilled person in the art can modify or change the above disclosed technical content to equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made to the above embodiments without departing from the technical solution of the present application, according to the technical essence of the present application, still belongs to the protection scope of the technical solution of the present application.
Claims
1. A method for structured conversion of geotechnical data, characterized by, The method comprises the following steps: obtaining a geotechnical log file unstructured data set, deconstructing information particles of the data set, and decomposing the data into minimum information units containing entity clues, attribute clues and relationship clues, each information unit being marked with a clue type and a confidence level; constructing a clue association network, calculating the association weight between information units based on clue co-occurrence frequency and semantic consistency, and forming a weighted network topology; using an adaptive community discovery algorithm to cluster the weighted network topology to obtain a plurality of clustering communities; constructing a double-track reasoning model, the model comprising a clue aggregation layer and a knowledge completion layer: the clue aggregation layer fuses the features of the information units in each clustering community to generate an entity candidate set and a relationship candidate set; the knowledge completion layer is used for logical verification and missing information completion of the entity candidate set and the relationship candidate set, and outputs standardized entity results and relationship results; constructing a geotechnical knowledge graph based on the entity results and the relationship results, and mapping the knowledge graph into a structured data format; constructing an initial geological model using the structured data, extracting core data from the initial geological model, including stratigraphic geometric data, attribute data and spatial relationship data, and converting the core data into an intermediate data file in an extended IFC format, wherein the extended IFC format adds geological special entities to adapt to the properties of strata; associating the intermediate data file with a BIM model, matching the geological special entities with the component properties in the BIM model through data mapping rules to obtain data that can generate a BIM model containing geological information; the specific implementation of using an adaptive community discovery algorithm to cluster the weighted network topology to obtain a plurality of clustering communities is: First, initial seed nodes are generated based on weight density. The weight density of a node is the sum of the association weights between that node and its neighboring nodes, calculated as follows: ,in, Let i be the i-th information unit node in the network. Represents a node The set of adjacent nodes, Represents a node With adjacent nodes The correlation weight between them Represented as nodes Weight density; Calculate the mean of weight density of all nodes in the weighted network and the standard deviation , set a dynamic threshold , and take the nodes with weight density greater than or equal to as candidate seed nodes; The candidate seed nodes are de-duplicated, and if the association weight between two candidate seed nodes is greater than or equal to a set strong association threshold, the node with a higher weight density is reserved as an initial clustering center to finally form a seed node set where m is the number of initial clustering centers. Computing node The weighted distance between a node and a cluster center is calculated as wherein, is the center node of the kth cluster, is the node is the association weight between the node and the cluster center, is the average of the association weights between all nodes in the kth cluster, and the node is assigned to the cluster with the smallest distance based on the weighted distance. Define the k-th cluster as a community, and calculate the community cluster cohesion using the following method: ,in, Let the cohesion of the k-th community cluster be denoted as . For nodes Association weights, The sum of the total network weights. For nodes The sum of the association weights with all adjacent nodes. The penalty coefficient is... Let be the number of nodes in the k-th community; If less than zero, split according to the maximum weight difference, calculate the correlation weight difference between all nodes in the community, select the pair of nodes with the largest difference as the new center, and split the original community into two sub-communities; if the cohesion degree of the community cluster after merging two adjacent communities increases to the set percentage threshold, the merging operation is performed; repeat the above splitting and merging process until the set number of iterations is reached, and then stop iteration.
2. The method of claim 1, wherein, deconstructing information particles of the data set, and decomposing the data into minimum information units containing entity clues, attribute clues and relationship clues, including: using natural language processing technology to segment and tag the geotechnical log file unstructured data, identifying noun phrases in the text as entity clue candidates, quantity phrases and adjective phrases as attribute clue candidates, and preposition phrases and verb phrases as relationship clue candidates; classifying the candidate clues based on a pre-trained BERT model in the field of geotechnical engineering to determine the clue type; calculating the confidence level of clue classification using a confusion matrix, which is obtained by the ratio of the number of correctly classified clues to the total number of clues, wherein the total number of clues is the sum of the number of entity clues, attribute clues and relationship clues.
3. The method of claim 1, wherein, The specific implementation of constructing a clue association network, calculating the association weight between information units based on clue co-occurrence frequency, semantic consistency and field logical association degree, and forming a weighted network topology is: firstly determining network nodes: the minimum information units obtained by information particle deconstruction are used as network nodes, each node carrying a clue type and a confidence level label; calculating the clue co-occurrence frequency, counting the number of co-occurrences of any two information units in the same text paragraph, table row or adjacent sentences of the geotechnical log file unstructured data, and taking the ratio of the number to the total number of paragraphs in the data set as the co-occurrence frequency; The semantic consistency is calculated, a pre-trained BERT model is used to extract a context vector corresponding to the content of each information unit text, and the semantic consistency of two information units is calculated by using a cosine similarity; The clue co-occurrence frequency and the semantic consistency are weighted and fused to obtain an association weight between the information units; A clue association network is constructed by taking the information units as nodes and the association weights as edges, a connection is established in the network if the association weight of two information units is greater than or equal to a set weight threshold, and finally a weighted network topology containing an association relationship is obtained.
4. The method of claim 1, wherein, The implementation of the clue aggregation layer for generating an entity candidate set and a relationship candidate set includes: Core features of the information units in each clustering community are extracted, the core features include text semantic features of entity clues, numerical and description features of attribute clues, and semantic association features of relationship clues, and a clue type label and a confidence of each information unit are introduced as auxiliary features; Based on the association weight between the information units in the clustering community, a weighted feature fusion strategy is used to combine the entity clues and the closely associated attribute clues to form a preliminary entity candidate, and the relationship clues and the associated entity clues are matched to form a relationship candidate, so as to generate an entity candidate set and a relationship candidate set corresponding to each clustering community.
5. The method of claim 1, wherein, The implementation of the knowledge completion layer for performing logical verification and missing information completion on the entity candidate set and the relationship candidate set and outputting standardized entity results and relationship results includes: Logical verification is performed on the entity candidate set and the relationship candidate set, wherein the rationality of entity attributes is checked for the entity candidate set, and the logicality of the relationship between entities is verified for the relationship candidate set; For missing information found after verification, reasoning and completion are performed based on high-confidence association clues in the clue association network and existing knowledge; The standardized entity results and relationship results are output.
6. The method of claim 1, wherein, The specific implementation of constructing an initial geological model by using structured data, extracting core data in the initial geological model, including stratigraphic geometric data, attribute data and spatial relationship data, and converting the core data into an intermediate data file conforming to an extended IFC format includes: An extended IFC format is designed, and a geological special entity is newly added, the geological special entity includes a stratigraphic unit entity, a geological attribute set entity and a spatial topology entity; A mapping relationship between the core data and the extended IFC format is established, the stratigraphic geometric data is mapped to a geometric parameter of the stratigraphic unit entity, the attribute data is mapped to a feature parameter of the geological attribute set entity, and the spatial relationship data is mapped to an association parameter of the spatial topology entity; According to the mapping relationship, the extracted core data is converted into an intermediate data file conforming to the extended IFC format.
7. The method of claim 1, wherein, Associating the intermediate data file with the BIM model includes: importing the intermediate data file into the BIM platform by using an IFC standard interface, and aligning the spatial positions by converting the coordinate system of the geological model to the coordinate system of the BIM model.
Citation Information
Patent Citations
Railway engineering unfavorable geology knowledge base construction method and system
CN119599120A
Medical text big data intelligent labeling and knowledge graph construction method and system
CN119851968A