Geological condition engineering risk auxiliary analysis method and system based on large language model
By analyzing and integrating multimodal information of geological exploration data, combining engineering information and expert experience, and using a large language model to construct a geological risk knowledge map, the problem of junior engineers having difficulty identifying geological risks has been solved, achieving more accurate and standardized geological risk analysis and possessing self-learning capabilities.
Patent Information
- Application Number
- CN202510861972.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
AI Technical Summary
In existing technologies, it is difficult for junior engineers to effectively identify potential geological risks in geological survey reports, which may lead to design rework or safety hazards. In addition, there is a lack of a systematic method to effectively introduce large language models into geological engineering risk identification.
By performing multimodal information analysis and fusion on geological exploration data, structured geological data is generated. By combining engineering information and expert experience, a geological risk knowledge graph is constructed. Large language models are used for automatic risk identification and reasoning, and the risk identification and reasoning results are output.
It improves the accuracy and standardization of geological risk analysis, makes up for the lack of experience of junior engineers, has scalability and self-learning capabilities, and reduces the probability of design errors and accidents caused by geological misjudgment.
Smart Images

Figure CN120764677A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of geotechnical engineering risk identification, and particularly relates to a geological condition engineering risk auxiliary analysis method and system based on a large language model. BACKGROUND
[0002] In infrastructure construction, the complexity of geological conditions has a decisive influence on engineering safety. Geological risk analysis usually relies on geologists or geotechnical engineers with rich experience, and needs to be combined with stratum structure, hydrological conditions, geological structure, etc. for comprehensive judgment. However, the current large number of preliminary drafts of geological exploration reports are usually written by junior engineers, who have limited professional experience and often have difficulty in identifying potential geological risks in a timely manner, resulting in subsequent design rework or safety hazards. In recent years, large language models (such as GPT models) have made breakthroughs in text understanding and knowledge reasoning, providing the possibility of assisting in complex logical judgment. However, how to introduce large models into the task of geological engineering risk identification still lacks effective and systematic methods. SUMMARY
[0003] The present application provides a geological condition engineering risk auxiliary analysis method and system based on a large language model, which uses large language model technology to intelligently analyze geological information in the survey report, and combines engineering type, specification knowledge and expert experience to assist in identifying key geological risks, improve the accuracy and standardization of risk analysis, and make up for the lack of engineer experience.
[0004] According to a first aspect, a geological condition engineering risk auxiliary analysis method based on a large language model is provided in an embodiment, and the method comprises:
[0005] Step S1, multi-modal information analysis and fusion of geological exploration data is performed to obtain structured geological data;
[0006] Step S2, engineering information and structured geological data are combined to generate an engineering-geological situation label;
[0007] Step S3, a geological risk knowledge graph is constructed, and knowledge graph matching is performed according to the engineering-geological situation label to obtain a knowledge graph matching result;
[0008] Step S4, a prompt word is constructed in combination with the knowledge graph matching result, and a large language model is used to perform automatic identification and reasoning of geological risks according to the constructed prompt word, and an output of the geological risk identification and reasoning result.
[0009] Further, the step S1 specifically comprises:
[0010] Step S11, text information is analyzed, including text preprocessing, geological entity recognition and entity relationship extraction;
[0011] Step S12, parsing table information, including table preprocessing, line frame detection and cell positioning, cell text extraction and structured output;
[0012] Step S13, parsing image information, including image preprocessing, feature extraction, feature fusion and lithology classification;
[0013] Step S14, according to the preset fusion rule, the text, table and image analysis results are fused, and the preset format structured geological data is output.
[0014] Further, the step S14 specifically includes:
[0015] Based on the confidence score voting mechanism, the multi-modal information including text, table and image is fused, and the unified format structured data is output, including project number, drill hole number, drill hole coordinates, and layer parameters.
[0016] Further, the step S2 specifically includes:
[0017] Step S21, obtaining engineering information, including engineering type, excavation depth, site category, distance from surrounding structures or underground structures, and preprocessing, generating engineering context embedding vector according to the processed engineering information;
[0018] Step S22, obtaining different stratum information from structured geological data, including layer thickness, lithology type, hydrology, and standard penetration value of each layer;
[0019] Step S23, dynamically combining engineering parameters and geological horizon features through rule matching or clustering method to generate engineering-geological situation label.
[0020] Further, the step S3 specifically includes:
[0021] Step S31, integrating specification provisions, expert experience and historical cases to build a geological risk knowledge graph containing geological situation, risk type and control strategy triples;
[0022] Step S32, according to the engineering-geological situation label, similarity calculation and screening with the nodes in the knowledge graph are performed to obtain the matched triple result.
[0023] Further, the step S32 specifically includes:
[0024] If the similarity calculation result exceeds the preset threshold, it is considered to be matched, and the Top-K similar risk items are returned and sorted by score, and if it is not matched, the large language model is prompted for autonomous risk identification and reasoning.
[0025] Further, the step S4 specifically comprises:
[0026] The step S41 creates a prompt word template, which includes engineering information, formation information and knowledge graph guide language. The knowledge graph guide language is constructed according to the matched knowledge graph triple result and used to provide reference and guidance for risk analysis of the large language model.
[0027] The step S42 embeds the actual engineering information, formation information and matched triple information into the prompt word template to obtain a prompt word.
[0028] The step S43 uses the large language model to perform automatic geological risk identification and reasoning according to the obtained prompt word, and outputs a result including risk type, risk level, reasoning reason, control countermeasure suggestion and cited standard provisions.
[0029] According to a second aspect, an embodiment provides a geological condition engineering risk auxiliary analysis system based on a large language model, which comprises:
[0030] A data analysis module is configured to analyze and fuse multi-modal information of geological exploration data to obtain structured geological data.
[0031] A label generation module is configured to generate an engineering-geological situation label in combination with engineering information and structured geological data.
[0032] A knowledge graph matching module is configured to construct a geological risk knowledge graph, match the knowledge graph according to the engineering-geological situation label, and obtain a knowledge graph matching result.
[0033] A large model reasoning module is configured to construct a prompt word in combination with the knowledge graph matching result, use a large language model to perform automatic geological risk identification and reasoning according to the constructed prompt word, and output a geological risk identification and reasoning result.
[0034] According to a third aspect, an embodiment provides an electronic device, which comprises a processor and a memory.
[0035] The memory is configured to store one or more program instructions.
[0036] The processor is configured to run the one or more program instructions to perform the steps of the geological condition engineering risk auxiliary analysis method based on a large language model according to any one of the above aspects.
[0037] According to a fourth aspect, an embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the geological condition engineering risk auxiliary analysis method based on a large language model according to any one of the above aspects are implemented.
[0038] The application provides a geological condition engineering risk auxiliary analysis method and system based on a large language model, which is suitable for intelligent auxiliary decision-making in the stages of geological exploration, engineering design and preliminary feasibility study, and has the following advantages
[0039] Beneficial effects:
[0040] (1) The problem of insufficient risk identification experience of junior engineers is solved;
[0041] (2) The powerful knowledge reasoning ability of the large model is utilized to realize "person + AI" collaborative analysis;
[0042] (3) The large model has scalability and self-learning ability, and can continuously adapt to different engineering scenarios;
[0043] (4) The systematization, accuracy and standard consistency of geological risk analysis are improved;
[0044] (5) The probability of design errors and engineering accidents caused by geological judgment errors is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 A flowchart of a geological condition engineering risk auxiliary analysis method based on a large language model is provided for an embodiment of the application;
[0046] Figure 2 A specific implementation flowchart of a geological condition engineering risk auxiliary analysis method based on a large language model is provided for an embodiment of the application;
[0047] Figure 3 A structural diagram of a geological condition engineering risk auxiliary analysis system based on a large language model is provided for an embodiment of the application. DETAILED DESCRIPTION
[0048] The application will be further described in detail through specific embodiments combined with the drawings. In different embodiments, similar elements are associated with similar element labels. In the following embodiments, many details are described in order to make the application better understood. However, those skilled in the art can easily recognize that some features can be omitted or replaced by other elements, materials or methods in different cases. In some cases, some operations related to the application are not shown or described in the specification in order to avoid the core part of the application being overwhelmed by too much description, and it is not necessary to describe these related operations in detail for those skilled in the art according to the description in the specification and general technical knowledge in the art.
[0049] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various embodiments. Meanwhile, the steps or actions in the method description can also be sequentially changed or adjusted in a manner that can be apparent to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment, and do not mean that the sequence is necessary, unless otherwise stated that a certain sequence must be followed.
[0050] The first embodiment of the present application provides a geological condition engineering risk auxiliary analysis method based on a large language model (LLM). The following will be described in detail in combination with Figure 1 and Figure 2 .
[0051] As shown in Figure 1 , step S1, multi-modal information analysis and fusion of geological exploration data is performed to obtain structured geological data.
[0052] This step aims to comprehensively analyze the multi-modal information (text, table, drawing) in the original geological exploration data and convert it into structured geological parameters, which is the input basis for the subsequent modules.
[0053] The information processing sequence is as follows:
[0054] Text information extraction: obtain descriptive stratigraphic information;
[0055] Table information analysis: standard drilling data supplement and verification;
[0056] Drawing recognition and fusion: graphical information verification and correction;
[0057] Multi-source fusion and anomaly detection: output unified structure data.
[0058] The above steps specifically include:
[0059] Step S11, text information analysis, including text preprocessing, geological entity recognition and entity relationship extraction.
[0060] Text information analysis specifically includes the following contents:
[0061] 1. Text preprocessing and word segmentation
[0062] The original description in the geological exploration report is divided into: unstructured paragraphs (according to chapters, drilling numbers, etc.); self-defined dictionary word segmentation (such as "silty clay" "interlayer silt"); eliminate unit ambiguity (such as "m" "meter" "millimeter" unified to SI unit); syntactic normalization, which is convenient for subsequent extraction.
[0063] 2. Geological entity recognition (NER)
[0064] The BERT-BiLSTM-CRF joint model is used to identify keywords such as "stratum name", "lithology", "thickness", and "hydrological state".
[0065] The model principle is as follows:
[0066] Input text sequence: x = x1, x2, …, x n ;
[0067] Each word is encoded by BERT into a context embedding vector h i ;
[0068] BiLSTM further captures long dependency information, outputting h' i ;
[0069] Finally, the CR layer is used for sequence labeling.
[0070] The formula is as follows:
[0071] h i = BERT(xi) h i = BERT(x i ) h i = BiLSTM(h i )
[0072] h' i = BiLSTM(h i )
[0073]
[0074] Where:
[0075] x i : the i-th word;
[0076] h i : BERT encoding vector;
[0077] h' i : BiLSTM output vector;
[0078] y i : the candidate label of the i-th word (such as B-lithology, I-lithology), which is the label sequence predicted by the model, belonging to one of all possible label paths;
[0079] y * : the label finally predicted by the model;
[0080] label transition weight;
[0081] P(y i |h' i ): the label probability of the word in the current context.
[0082] The output entity includes:
[0083]
[0084] 3. Relation extraction (RE)
[0085] The relation scoring function is used to quantify whether a specific semantic relation (such as "rock = silt clay") exists between two candidate entities. In the implementation of the present embodiment, the relation extraction process is divided into two steps:
[0086] a. Dependency syntax analysis is used to determine possible attribute-value structures (such as "thickness 6.5 meters" -> thickness = 6.5);
[0087] b. Calculate the relation scoring function for each pair of candidate entities (such as "silt clay" and "rock"):
[0088]
[0089] Where:
[0090] Word vector (generated by BERT or bag of words);
[0091] cos(·): cosine similarity;
[0092] δ dep (i,j): whether there is a legal dependency path (such as amod, nummod).
[0093] i,j are the indices or positions of the candidate entities in the text, for example, the 5th word is "silt clay" and the 7th word is "rock", then i = 5, j = 7;
[0094] Where, the cosine similarity measures the degree of semantic closeness, and the dependency path determines the grammatical validity.
[0095] When the score exceeds a certain threshold (such as 0.6), it is considered that there is an effective attribute relation, and the relation triple is output. Otherwise, the entity pair is filtered.
[0096] This method takes into account both semantic and syntactic rationality, ensuring the accuracy and interpretability of relation extraction.
[0097] Step S12, parse the table information, including table preprocessing, line frame detection and cell positioning, cell text extraction and structured output.
[0098] The table information parsing specifically includes the following contents:
[0099] For drilling data tables in Word / PDF, use table recognition modules (such as PaddleOCR + table header rules) for parsing.
[0100] Restore the input table image (such as a scanned drilling table) to a two-dimensional structure with clear "row-column cell boundaries" to support field recognition and data extraction. The specific processing flow is as follows:
[0101] 1. Table image recognition process description:
[0102] (1) Preprocessing, including:
[0103] Gray scale, binarization, and background noise removal for the image;
[0104] Apply Hough transform or OpenCV edge detection algorithm to obtain possible table borders.
[0105] (2) Line frame detection and table structure positioning, including:
[0106] Detect the intersection of horizontal and vertical lines and reconstruct the cell grid;
[0107] If the original image has no obvious line frame, use table structure models (such as TableNet or CascadeTabNet) for semantic segmentation of the image to predict each cell area.
[0108] (3) Text extraction (OCR), including:
[0109] Use OCR engine (such as PaddleOCR) to recognize text in each cell area;
[0110] Use table structure to map recognized text to corresponding field positions (such as "hole number", "lithology", "starting depth", etc.).
[0111] (4) Structured output, including:
[0112] Generate standard two-dimensional array or JSON representation, for example:
[0113]
[0114] Step S13, analyze the image information, including image preprocessing, feature extraction, feature fusion, and lithology classification.
[0115] Map recognition and fusion, specifically including the following:
[0116] 1. Image segmentation to identify rock layer area
[0117] For columnar section and profile section: use semantic segmentation network (such as U-Net) to extract layers; classify layer texture and color (combine gray scale features, texture filters) to determine lithology.
[0118] In geological survey report drawings, columnar section and profile section usually use different colors and patterns (textures) to distinguish different lithologies (such as clay, sand layer, gravel, etc.). For example: silty clay may be represented by light yellow + parallel fine lines, sand layer may be represented by gray + dotted texture, and bedrock may be represented by dark + blocky texture.
[0119] The traditional method relies on manual interpretation, while the present embodiment uses computer vision technology (such as semantic segmentation, texture analysis) to automatically identify these features, improving efficiency.
[0120] 2. Specific implementation steps
[0121] (1) Image preprocessing
[0122] Grayscale: Convert color images to grayscale to reduce color interference and highlight texture features.
[0123] Filtering and denoising: Use Gaussian filtering or median filtering to eliminate noise in scanned images.
[0124] (2) Texture feature extraction
[0125] Texture is a key feature for lithology discrimination. The present embodiment uses the following methods to quantify texture:
[0126] Gray Level Co-occurrence Matrix (GLCM): Statistics of spatial distribution of pixel gray values, extracting contrast, energy, entropy, correlation, etc.
[0127] For example: clay texture is uniform (low contrast, high energy), sand layer texture is rough (high contrast, low energy).
[0128] Texture filters (such as Gabor filter, LBP operator), including:
[0129] Gabor filter: Simulate human visual characteristics, capture texture in different directions (such as horizontal bedding, diagonal bedding).
[0130] LBP (Local Binary Pattern): Calculate local texture pattern, suitable for lithology classification (such as blocky vs. layered).
[0131] (3) Color feature extraction
[0132] Principal Component Analysis (PCA): Reduce dimensionality of RGB color, extract main color system (such as yellowish clay, grayish sand layer).
[0133] HSV color space: distinguish lithology (e.g. red clay vs. gray sandstone) by hue (H), saturation (S).
[0134] (4) Lithology classification
[0135] Feature fusion: concatenate texture (GLCM / Gabor) and color (PCA / HSV) features into a comprehensive feature vector.
[0136] Classification model: train a classifier using support vector machines (SVM) or convolutional neural networks (CNN) to match the lithology labels in the legend.
[0137] Example: input the feature vector of a certain layer area, output the most probable lithology category (e.g. "silty clay - 90%").
[0138] Step S14, the text, table and image analysis results are fused according to the preset fusion rule, and the pre-set format structured geological data is output.
[0139] In this embodiment, step S14 specifically includes: based on the confidence score voting mechanism, the multi-modal information including text, table and image is fused, and the unified format structured data is output, including project number, borehole number, borehole coordinates, and layer parameters.
[0140] Multi-source fusion and anomaly detection specifically includes the following contents:
[0141] Fusion of multi-modal information from text, table, and map, and establishment of fusion rules:
[0142]
[0143] Wherein:
[0144] r t : text recognition result;
[0145] r g : map recognition result;
[0146] r b : table recognition result;
[0147] r* : result obtained by comprehensive analysis of text, map, and table.
[0148] conf r : confidence score (derived from model output probability or rule score);
[0149] w r : weight parameter (trainable or expert set).
[0150] Fusion example: suppose the lithology of a certain stratum is judged, and the results of three data sources are as follows:
[0151]
[0152] Fusion calculation:
[0153] Total weight normalization: 0.36+0.21+0.24=0.810.36+0.21+0.24=0.81
[0154] Silty clay total score: 0.36(text) + 0.24(table) = 0.60
[0155] Clay total score: 0.21 (map)
[0156] Final result: silty clay (because 0.60>0.21)
[0157] Final output: unified format structured geological data as follows:
[0158]
[0159] In an operation example, the specific implementation process of step S1 geological data analysis and structured processing is as follows:
[0160] Obtain the geological survey report, including text description, drilling diagram, profile diagram, drilling table and other contents, and uniformly convert them into readable digital format (such as Word, PDF, picture);
[0161] Divide the text part into paragraphs and perform syntax normalization, and use the word segmentation tool supported by the geological dictionary to segment professional terms;
[0162] Call the BERT-BiLSTM-CRF model to identify geological entities in the text, including lithology, hydrological state, thickness, layer number, etc.;
[0163] Use the syntax dependency analysis tool (such as spaCy) to establish attribute-value dependency paths and filter effective entity relationships;
[0164] Combine PaddleOCR and other modules to reconstruct the drilling table and identify "hole number-stratigraphic position-lithology" fields;
[0165] Use the U-Net network to segment the rock layer for column chart and profile diagram, extract legend texture features and match lithology classification;
[0166] Fuse the information from the text, table and map, and use the confidence weighted voting mechanism to output the final result;
[0167] Output structured JSON data in a unified format, including borehole number, coordinates, geological parameters of each layer, etc., to provide a standard interface for subsequent calls.
[0168] Example: The input geological exploration report contains the description "ZK3 hole from the ground to 6.5m is silty clay, weakly permeable, sand at the bottom of the layer", the corresponding lithology code in the table is "FZNT", and the texture in the drawing matches the "silty clay" legend. The final output structured result is: {"hole number": "ZK3", "stratum": [{"start": 0.0, "end": 6.5, "lithology": "silty clay", "hydrological property": "weakly permeable"}]}
[0169] As shown in Figure 1 Step S2, combine engineering information and structured geological data to generate engineering-geological situation labels.
[0170] This step is to model the engineering context by receiving the structured geological data output by step S1 and combining the user-provided engineering background (project type, excavation depth, surrounding environment, etc.) to generate an "engineering context vector" as a precondition input for large language model reasoning.
[0171] The main functions include: explicitly constructing the scenario of "how the large model understands this project"; assisting in selecting appropriate design specification clauses and constraint conditions; forming an "engineering-geological combined situation" for subsequent reasoning.
[0172] The above steps specifically include:
[0173] Step S21, obtain engineering information including engineering type, excavation depth, site category, distance to surrounding structures or underground structures, and preprocess it to generate an engineering context embedding vector based on the processed engineering information.
[0174] Step S22, obtain different stratum information from structured geological data, including the thickness of each stratum, lithology type, hydrological property, and standard penetration value.
[0175] Step S23, dynamically combine engineering parameters and geological horizon features through rule matching or clustering to generate engineering-geological situation labels.
[0176] Specifically,
[0177] (1) Engineering type identification and context coding input content includes: engineering type, excavation depth d, site grade Cs, adjacent distance L, etc.
[0178] Each input parameter is obtained through the following methods:
[0179] The engineering type and excavation depth are input by the user or automatically parsed from the design file;
[0180] Site class inherits from the geotechnical data structuring result (step S1);
[0181] Adjacent distance is obtained by manual input or CAD / GIS tool measurement.
[0182] The system provides input verification functions to ensure that the parameters meet engineering common sense (such as excavation depth > 0).
[0183] Generate engineering scene context vector:
[0184]
[0185] Where the parameter meanings are:
[0186]
[0187] (2) Engineering-geological situation label generation
[0188] This part of the engineering-geological situation label generation is based on the output of the engineering context coding in (1), dynamically combines engineering parameters (such as excavation depth) with geological horizon characteristics (such as lithology, hydrology), and generates situation labels with engineering significance (such as 'deep foundation pit + soft soil'), providing structured input for subsequent risk reasoning.
[0189] Combination result example:
[0190] "Deep foundation pit + thick soft soil + high water level" → high-risk situation label
[0191] "Shallow foundation + gravel soil + dry area" → low-risk scene label
[0192] Output format example:
[0193]
[0194]
[0195] In a specific operation example, the specific implementation process of step S2 engineering context modeling is as follows:
[0196] Input engineering information by user or front-end system, including engineering type (such as foundation pit, slope), excavation depth, site category, distance from surrounding structures, etc.
[0197] Enumerate and normalize the input fields;
[0198] Use MLP or expert rule base to build engineering context embedding vector
[0199] Combine the geological horizon information extracted in step S1 to build stratum combinations, such as Each Including layer thickness, lithology type, hydrological properties, standard penetration value, etc.;
[0200] Through rule matching or clustering, "engineering-geological situation labels" are generated, such as "soft soil + rich water + deep foundation pit", and associated with the initial risk level judgment.
[0201] Example: Input: Project type = "underground station foundation pit"; Excavation depth = 18m; Site category = "Class III"; Distance to adjacent buildings = 3m; combined with the stratigraphic combination of "muddy silty clay + thickness > 10m + highly permeable", the final context vector label is "deep foundation pit + soft soil + rich water", which is used as input for subsequent model inference.
[0202] like Figure 1 As shown, in step S3, a geological risk knowledge graph is constructed, and knowledge graph matching is performed according to engineering-geological situation labels to obtain knowledge graph matching results.
[0203] This step provides a set of "engineering knowledge background" to support the identification of hidden risks before the large language model performs reasoning.
[0204] Main functions: Integrate regulatory provisions, expert rules, and historical cases; establish a knowledge graph between "geological conditions - risk types - countermeasures and suggestions"; support graph query, filtering, and prompt word enhancement.
[0205] The above steps specifically include:
[0206] Step S31: Integrate regulatory provisions, expert experience, and historical cases to construct a geological risk knowledge graph containing triples of geological scenarios, risk types, and control measures;
[0207] Step S32: Based on the engineering-geological situation label, similarity calculation and screening are performed with the nodes in the knowledge graph to obtain matching triple results.
[0208] Specifically:
[0209] ①Graph structure design
[0210] Triple form: (geological scenario, potential risk, control measures)
[0211] Example:
[0212] ("muddy clay + rich water + deep foundation pit", "bottom surge", "water-stop curtain + advanced reinforcement")
[0213] ("Strongly weathered granite + high stress", "rock fall", "anchor mesh shotcrete support")
[0214] The graph structure is represented as G = (V, E), where V is the entity set (geology, risk, countermeasure), and E is the relationship edge.
[0215] ②Graph call and semantic expansion
[0216] Graph structure: contains (geological scenario, risk type, control countermeasure) triple, pre-stored industry specifications and expert experience;
[0217] Call process: according to the engineering-geological label obtained in step S2, calculate the semantic similarity with the graph node, and return the matched triple; convert the matching result into natural language prompt words to guide the large model to generate standardized risk analysis.
[0218] Output example: {"scenario": "soft soil + rich water", "risk": "bottom sudden gushing", "suggested countermeasure": "water stop curtain"}.
[0219] In an operation example, the specific implementation process of step S3 knowledge graph construction and calling is as follows:
[0220] Construct a knowledge graph containing triples, with the basic structure of (geological scenario, risk type, countermeasure suggestion);
[0221] The graph node includes lithology, hydrological state, risk type, design countermeasure, etc. Semantic label, and the edge represents the cause and effect and control relationship;
[0222] Construct a vectorized query interface to filter similar risk cases from the graph according to the engineering context and geological combination output in S2;
[0223] Implementation method: the "similar risk case filtering" process in the graph is as follows:
[0224] (1) Context encoding generation:
[0225] The semantic labels of the current project such as "engineering type + geological combination + hydrological characteristics" (such as "deep foundation pit + silty clay + strong water permeability") are spliced into text description;
[0226] Input the pre-trained BERT (or RoBERTa) type encoder to get a 768-dimensional semantic vector.
[0227] (2) Graph node encoding:
[0228] The "scenario" field in each triple (such as "soft soil + rich water + excavation depth") is pre-encoded into a vector and stored in a vector index library; FAISS, Milvus, etc. Vector search engine for efficient recall.
[0229] (3) Similarity calculation and filtering:
[0230] Using cosine similarity calculation, set the matching threshold θ = 0.8, only keep the items with similarity higher than the threshold;
[0231] Return Top-K (such as 3) similar risk items, sorted by score.
[0232] (4) Failure handling mechanism:
[0233] If there is no matching item that meets the threshold (such as all less than 0.6), trigger the bottom prompt: "The current scene is a new type, no similar historical cases are found, please rely on the current geological parameters to infer the risk by the large model";
[0234] The prompt will be passed into the prompt of step S4, guiding the large language model to perform full-scene reasoning.
[0235] Output matching items and construct prompt word templates: such as "For {scenario}, is there {risk}? What countermeasures can be taken?";
[0236] Support manual maintenance of graph and historical cases, and improve the system's self-learning ability.
[0237] Example: input context label: "deep foundation pit + soft soil + water-rich"; matching graph item: ("soft soil + water-rich + excavation depth > 15m", "bottom sudden gushing", "set water curtain + deep mixing pile"). Used to construct the prompt: "Please judge whether there is a bottom sudden gushing risk and propose control suggestions".
[0238] As shown in Figure 3 Step S4, construct the prompt word based on the knowledge graph matching result, use the large language model to perform geological risk automatic identification and reasoning according to the constructed prompt word, and output the geological risk identification and reasoning result.
[0239] This step integrates geological information, engineering context and knowledge graph, and uses a large language model (LLM) to perform automatic identification and reasoning of geological risks; judge the risk type, risk level and reason; propose targeted control measures and give corresponding specification article references; ensure that the output has language interpretability and structured callability.
[0240] Explanation of each step (relationship with steps S1-S3)
[0241]
[0242] The above steps specifically include:
[0243] Step S41, create a prompt word template, which includes engineering information, stratum information and knowledge graph guide language, and the knowledge graph guide language is constructed according to the matching knowledge graph triple result and used to provide reference and guidance for large language model risk analysis.
[0244] Step S42, embedding the actual engineering information, stratum information and matched triple information into the prompt word template to obtain a prompt word;
[0245] Step S43, using a large language model to automatically identify and reason geological risks according to the obtained prompt word, and outputting a result including risk type, risk level, reasoning reason, control countermeasure suggestion and cited standard provisions.
[0246] The specific content is as follows:
[0247] ①Prompt word generation and input organization (Prompt Engineering)
[0248] Core principle: convert structured data into a combination of "natural language + structured description"; use a three-part prompt template of "scene description + risk question + output format guide"; automatically insert standard numbers and keywords to guide the model to output standardized content.
[0249] Among them, the knowledge graph plays the following roles in this step:
[0250] (1) Semantic query: the system retrieves risk triples with high similarity in the knowledge graph according to the "engineering-geology combined label" output in step S2 (such as "soft soil + water-rich + deep foundation pit" → "bottom gushing + water stop curtain");
[0251] (2) Prompt word enhancement: embed the matched "risk type" and "suggestion countermeasure" into the prompt to construct "implicit prior" reasoning guide language. For example, add "according to expert experience, this scene may have a bottom gushing risk, please judge its possibility combined with existing stratum and excavation conditions" in the original input of the model;
[0252] (3) Output comparison: after the model output, use the knowledge graph to do "reverse comparison" again to assist in identifying standard numbers and item content, ensuring that the recommended measures are consistent with engineering practice.
[0253] Convert the matched result triples (scene, risk type, control suggestion) in the graph into "guide language" segments that the large language model can understand, called knowledge graph guide language.
[0254] Example Prompt template:
[0255] "The project is {engineering type}, the foundation pit excavation depth is {H} meters. The stratum includes {stratum information}, among which {soft layer description}, the hydrological nature is {hydrological nature description}. The groundwater level is {W} meters from the ground surface.
[0256] Please identify the potential geological engineering risks, determine the risk type and level, and provide reasoning and control suggestions. Include the reference to the specification article number and content in the output.
[0257] Example (filled actual prompt):
[0258] "The project is an underground station foundation pit with an excavation depth of 18 meters. The stratum contains thick silt clay (10 meters thick), with strong hydrological properties and high permeability. The groundwater level is 2 meters above the ground surface.
[0259] Please determine the potential geological risk type and level, explain the reasons, and suggest control measures, referring to relevant specifications."
[0260] ② Risk reasoning method and scoring mechanism
[0261] Reasoning model principle:
[0262] Large language models extract the following logical chain from the prompt words through context semantic modeling:
[0263] Causal chain identification (e.g., "soft soil + rich water + deep foundation pit" → bottom sudden gushing);
[0264] Match the risk definition and applicable conditions against the knowledge base / specification clauses;
[0265] Empirical reasoning output: automatically generate corresponding levels and control suggestions based on geological background.
[0266] Risk level scoring mechanism:
[0267] Use a hybrid scoring mechanism based on key word rules and model scoring:
[0268] R score = α·R model + β·R rule + γ·R expert
[0269] Where:
[0270] R model : Risk level score determined by large model (e.g., high, medium, low → mapped to 3, 2, 1);
[0271] R rule : Pre-warning rule scoring based on "engineering type-soil-hydrological properties" conditions in the atlas;
[0272] R expert : (Optional) Expert intervention correction score;
[0273] α, β, γ: Configurable weights (e.g., α = 0.6, β = 0.3, γ = 0.1, indicating model dominance).
[0274] Wherein, the pre-warning rule scoring in the atlas is obtained by the following way:
[0275] Rule source: The atlas stores risk rule triples set by experts or extracted from historical cases, such as: ("deep foundation + silt soil + water-rich", "bottom sudden gushing", "medium-high"), ("shallow foundation + gravel soil + dry", "essentially no risk", "low");
[0276] Matching mechanism: The system performs semantic matching between the "engineering-geology combination tag" of the current scene and the atlas rules (BERT similarity / cosine similarity can be used), and sets a matching threshold (such as 0.8). Those exceeding the threshold are considered as hits;
[0277] Scoring calculation:
[0278] If the risk label item in the atlas is completely hit, the corresponding risk level scoring is given (such as medium-high = 2.5);
[0279] If partially hit (such as 2 / 3 condition matching), the score is weighted proportionally (such as assigned value 2.0);
[0280] If no rule is matched, a default value is set or not counted.
[0281] ③Output content and structured format
[0282] Output content composition: risk type (multiple); level (high / medium / low); reasoning reason (based on causal reasoning or norm basis); control suggestion (technical measures); norm reference (norm number + article number).
[0283] Output structure (JSON format):
[0284] {
[0285] "risk identification": [
[0287] {
[0288] "type": "bottom sudden gushing",
[0289] "level": "medium-high",
[0290] "reason": "soft soil layer thick + water-rich + excavation deep",
[0291] "control suggestion": "set water stop curtain, supplemented by deep mixing pile reinforcement",
[0292] "norm reference": "JGJ120-2012 Article 6.3.4"
[0293] },
[0294] {
[0295] "Type":"Pit wall instability",
[0296] "Grade":"Medium",
[0297] "Reason":"Large slope angle + low soil internal friction",
[0298] "Control Suggestions":"Set internal support system and monitor deformation in real time",
[0299] "Code Reference":"GB50497-2019 Article 7.2.1"
[0300] } ]
[0302] }
[0303] ④Support module:
[0304] Prompt word generation algorithm:
[0305] The prompt word generation module extracts keywords from structured data and assembles them into a Prompt:
[0306] Prompt = T proj + T geo + T hydro + T task
[0307] T proj : Engineering feature text (such as "subway foundation pit, excavation depth 18m");
[0308] T geo : Geological structure description (such as "10m thick silt clay");
[0309] T hydro : Hydrological conditions (such as "water level 2m above ground, strong water permeability");
[0310] Q task : Task instructions, such as "please determine the risk type and grade".
[0311] Supplementary notes:
[0312] Method for obtaining code references:
[0313] Code references can be obtained in the following two ways:
[0314] Output risk type by knowledge graph matching model, associate recommended articles;
[0315] Or through the large model retrieval prompt (such as: "for the control measures of bottom gushing, please quote the provisions in JGJ120-2012"), the model generates the corresponding reference clauses.
[0316] The output of this step in the system: provide structured risk information for visualization and report output; Access form filling, risk account automatic generation system; Support human-computer cooperation review and expert feedback mechanism.
[0317] In a specific operation example, the specific implementation process of step S4 large model reasoning is as follows:
[0318] Receive the output content of steps S1-S3, and generate Prompt in fixed format: contains "engineering information + geological combination + knowledge graph guide language";
[0319] Convert the matching result triplets (scene, risk type, control suggestion) in the graph into "guide language" segments that the large language model can understand, called knowledge graph guide language.
[0320] Construction method:
[0321] The system generates the following guide language according to the matching graph triplets:
[0322] "History shows that'soft soil + rich water + deep foundation pit' combination may cause bottom gushing risk, and often uses water stop curtain as control measures."
[0323] "Please analyze whether there is a bottom gushing risk according to the geological combination, and consider using water stop curtain and other technical means."
[0324] This type of statement is the middle paragraph of the Prompt, which plays a role in strengthening the model's attention.
[0325] Mechanism of action:
[0326] Guide the model to pay attention to typical risk paths (such as "soft soil → gushing");
[0327] Provide an empirical background to avoid irrelevant or missed model outputs;
[0328] Strengthen the rationality of quoting specification clauses (such as "gushing → JGJ120-2012 6.3.4").
[0329] Call a pre-set large language model (such as GPT-4), set temperature = 0.2 to ensure output stability;
[0330] Submit prompt to model API, parse the returned content, and extract risk type, level, cause analysis, control suggestion, specification reference, and other keywords;
[0331] Results are structured as JSON objects, with model confidence scores (e.g. logprob or explicit ratings);
[0332] If there are ambiguous or low-confidence items in the results, mark them as "need human review";
[0333] All results are written into the risk identification record table and input into the visualization module.
[0334] Example: Prompt: "The project is a subway station foundation pit with an excavation depth of 18m, adjacent to buildings. The stratum contains 10m thick silty clay, which is highly permeable and has high groundwater level. Please determine the potential geological risk type, level, and control measures." The large model returns: {"type": "bottom sudden inrush", "level": "medium-high", "reason": "soft soil layer thickness + water-rich + deep foundation pit", "control suggestion": "set up water stop curtain, supplemented by deep mixing pile reinforcement", "standard reference": "JGJ120-2012 Article 6.3.4"}.
[0335] The method of this embodiment further comprises:
[0336] Step S5, result output and visualization: the results output in step S4 are displayed in structured text and chart form, supporting interactive review and report generation, and facilitating engineers to use.
[0337] ① Structured result summary generation
[0338] Output "risk item + level + suggested measures + reason explanation" table
[0339] Support export of Word, PDF report format
[0340] Optional reference to related atlas or historical case links
[0341] ② Visualization risk map / cross-section map generation
[0342] Generate risk heat map combined with drilling data and risk items
[0343] Dangerous horizon cross-section map (automatically label inrush layer and sliding layer)
[0344] Support HTML / PNG output
[0345] Output example:
[0346]
[0347] In one specific operation example, the specific implementation process of step S5 result output and visualization is as follows:
[0348] Receive structured risk analysis results from step S4;
[0349] Automatically generate a report summary table, including "risk type-grade-suggested countermeasures-reference specification";
[0350] Build an interactive report in HTML format for users to quickly browse and edit;
[0351] Generate a two-dimensional profile map using borehole coordinates and risk horizon information to mark high-risk horizons and control recommendation areas;
[0352] Exportable to Word, PDF report or inserted into BIM / GIS system;
[0353] Record user modification opinions and update the database to form a feedback loop.
[0354] Example: User opens the risk analysis interface and sees:
[0355] Type: Bottom sudden inrush; Grade: Medium-high; Control suggestion: Waterproof curtain + mixing pile;
[0356] Click "profile" to view the high-risk layer 2 in ZK3 hole, highlighted on the map;
[0357] One-key export of "geological risk special analysis report".
[0358] Corresponding to the above disclosed geological condition engineering risk auxiliary analysis method based on a large language model, the embodiment of the present application also discloses a geological condition engineering risk auxiliary analysis system based on a large language model, as shown in Figure 3 Specifically includes:
[0359] A data analysis module is used to analyze and fuse multi-modal information of geological exploration data, and obtain structured geological data;
[0360] A label generation module is used to generate engineering-geological situation labels by combining engineering information and structured geological data;
[0361] A knowledge graph matching module is used to construct a geological risk knowledge graph, match the knowledge graph according to the engineering-geological situation label, and obtain a knowledge graph matching result;
[0362] A large model reasoning module is used to construct prompt words by combining the knowledge graph matching result, use a large language model to automatically identify and reason geological risks according to the constructed prompt words, and output a geological risk identification and reasoning result.
[0363] It should be noted that the detailed description of the application embodiment of the method for assisting in analyzing geological condition engineering risk based on a large language model can refer to the related description of the application embodiment of the method for assisting in analyzing geological condition engineering risk based on a large language model, which will not be repeated here.
[0364] In addition, the application embodiment further provides an electronic device, which comprises a processor and a memory; the memory is used to store one or more program instructions; and the processor is used to run the one or more program instructions to execute the steps of the method for assisting in analyzing geological condition engineering risk based on a large language model according to any one of the above.
[0365] It should be noted that the detailed description of the electronic device provided by the application embodiment can refer to the related description of the method for assisting in analyzing geological condition engineering risk based on a large language model provided by the application embodiment, which will not be repeated here.
[0366] In addition, the application embodiment further provides a computer readable storage medium, which stores a computer program; and the computer program is executed by a processor to implement the steps of the method for assisting in analyzing geological condition engineering risk based on a large language model according to any one of the above.
[0367] It should be noted that the detailed description of the computer readable storage medium provided by the application embodiment can refer to the related description of the method for assisting in analyzing geological condition engineering risk based on a large language model provided by the application embodiment, which will not be repeated here.
[0368] Those skilled in the art can understand that all or part of the functions of the various methods in the above embodiments can be realized by hardware or by a computer program. When all or part of the functions in the above embodiments are realized by a computer program, the program can be stored in a computer readable storage medium, which can include a read-only memory, a random access memory, a magnetic disk, an optical disk, a hard disk, etc. The above functions are realized by executing the program by a computer. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above functions are realized. In addition, when all or part of the functions in the above embodiments are realized by a computer program, the program can also be stored in a server, another computer, a disk, an optical disk, a flash disk or a mobile hard disk, etc. The program is downloaded or copied into the memory of the local device or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above embodiments are realized.
[0369] The above application of specific examples to illustrate the present invention, is only used to help understand the present invention, and does not limit the present invention. For the skilled in the art to which the present invention belongs, according to the idea of the present invention, several simple deductions, deformation or replacement can be made.
Claims
1. A method for auxiliary analysis of geological engineering risks based on a large language model, characterized in that: The method comprises: Step S1, performing multimodal information analysis and fusion on geological exploration data to obtain structured geological data; Step S2, combining engineering information with structured geological data to generate engineering-geological situation labels; Step S3: construct a geological risk knowledge graph, perform knowledge graph matching based on engineering-geological situation labels, and obtain knowledge graph matching results; Step S4: construct prompt words based on the knowledge graph matching results, use the large language model to automatically identify and reason about geological risks based on the constructed prompt words, and output the geological risk identification and reasoning results.
2. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 1, characterized in that: The step S1 specifically includes: Step S11, parsing the text information, including text preprocessing, geological entity recognition and entity relationship extraction; Step S12, parsing the table information, including table preprocessing, wireframe detection and cell positioning, cell text extraction and structured output; Step S13, analyzing the image information, including image preprocessing, feature extraction, feature fusion and lithology classification; Step S14: performing multimodal information fusion on the text, table and image analysis results according to preset fusion rules, and outputting structured geological data in a preset format.
3. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 2, characterized in that: The step S14 specifically includes: Based on the confidence scoring voting mechanism, multimodal information including text, tables and images is fused and structured data in a unified format is output, including project number, borehole number, borehole coordinates and stratigraphic parameters of each layer.
4. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 1, characterized in that: The step S2 specifically includes: Step S21: Acquire project information, including project type, excavation depth, site category, and distance to surrounding structures or underground structures, and perform preprocessing to generate a project context embedding vector based on the processed project information; Step S22, obtaining different stratum information from the structured geological data, including the thickness, lithology type, hydrological properties, and standard penetration value of each stratum; Step S23 : Dynamically combine engineering parameters with geological layer characteristics through rule matching or clustering to generate engineering-geological situation labels.
5. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 1, characterized in that: The step S3 specifically includes: Step S31: Integrate regulatory provisions, expert experience, and historical cases to construct a geological risk knowledge graph containing triples of geological scenarios, risk types, and control measures; Step S32: Based on the engineering-geological situation label, similarity calculation and screening are performed with the nodes in the knowledge graph to obtain matching triple results.
6. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 5, characterized in that: The step S32 specifically includes: If the similarity calculation result exceeds the preset threshold, it is considered a match, and the top-K similar risk entries are returned and sorted by score. If there is no match, the large language model is prompted to perform autonomous risk identification and reasoning.
7. The method for auxiliary analysis of geological engineering risks based on a large language model according to claim 1, characterized in that: The step S4 specifically includes: Step S41: creating a prompt word template, wherein the prompt word template includes engineering information, stratum information, and a knowledge graph guide. The knowledge graph guide is constructed based on the matched knowledge graph triples and is used to provide reference and guidance for risk analysis by the large language model. Step S42: embedding the actual engineering information, stratum information, and matched triplet information into a prompt word template to obtain a prompt word; Step S43: Automatically identify and reason about geological risks using a large language model based on the obtained prompt words, and output results including risk type, risk level, reasoning reasons, control countermeasures and referenced regulatory provisions.
8. A geological condition engineering risk auxiliary analysis system based on a large language model, characterized by: The system comprises: Data analysis module, used to perform multimodal information analysis and fusion of geological exploration data to obtain structured geological data; The label generation module is used to combine engineering information with structured geological data to generate engineering-geological situation labels; The knowledge graph matching module is used to construct a geological risk knowledge graph, perform knowledge graph matching based on engineering-geological situation labels, and obtain knowledge graph matching results; The large model reasoning module is used to construct prompt words based on the knowledge graph matching results, and use the large language model to automatically identify and reason about geological risks based on the constructed prompt words, and output the geological risk identification and reasoning results.
9. An electronic device, characterized in that: The device includes: a processor and a memory; The memory is used to store one or more program instructions; The processor is used to run one or more program instructions to execute the steps of the method for auxiliary analysis of geological condition engineering risks based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for auxiliary analysis of geological condition engineering risks based on a large language model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Identification ring
GB3020303S
Embedded hydraulic engineering risk intelligent regulation and control method based on knowledge graph
CN118297381A
Large language model knowledge question-answering method and system fused with multi-modal knowledge graph
CN118627628A
Water conservancy project risk response decision recommendation method based on cooperation of multi-modal knowledge graph and large model
CN120069378A
Large language model-based event processing method and apparatus, device and medium
WO2025086682A1
Cited By
Intelligent planning data processing method and system based on BIM and GIS
CN121561821A
Risk adaptation type intelligent compilation method for construction full-content type
CN121639121A
A risk-adaptive intelligent compiling method for construction full content type
CN121639121B
Tunnel geological anomalous body identification method, medium, equipment and product
CN121660109A
A method, medium, equipment and product for identifying geological anomalies in tunnels.
CN121660109B