Intelligent stratification method of geological model based on RAG technology

CN122614969APending Publication Date: 2026-08-21CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610538690.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

早期尝试采用机器学习算法,如支持向量机(SVM)、随机森林(Random Forest)等,对测井曲线进行分类以实现地层自动划分,但存在依赖人工提取特征(如自然伽马、声波时差的阈值分割)、泛化能力差(换一个矿区或地层类型时需重新调整参数甚至重建模型)、无法处理文本形式地质描述(如岩心记录中的“灰黑色致密块状玄武岩”)等缺陷

Benefits of technology

[0018]This application achieves the following effects: By acquiring multi-source geological data and vectorizing it using a text embedding model, it stores the data in a vector database. This allows for the construction of a unified geological knowledge base encompassing borehole lithology descriptions, stratigraphic characteristics, and mining area geological background information. This integrates private domain data such as geological reports and drilling data scattered across different systems into an efficient retrieval index, achieving the fusion and unified representation of multi-source heterogeneous data. Furthermore, by using a text embedding model to transform user-input geological stratification questions into query vectors and performing similarity searches in the vector database, it obtains multi-source geological knowledge fragments most semantically relevant to the question, overcoming the shortcomings of conventional keyword matching in existing technologies that easily miss deep-seated related knowledge. The retrieved knowledge fragments are then concatenated with geological domain prompt templates and input into a large language model. Leveraging the semantic reasoning and knowledge fusion capabilities of the large language model, it automatically generates structured stratification results containing stratigraphic names, lithological characteristics, and thickness ranges. The synergistic effect of these techniques enables this method to integrate multi-source heterogeneous geological data, reduce reliance on manual experience, improve the intelligence and accuracy of stratification analysis, and effectively solve the problems of low efficiency, strong subjectivity, and poor adaptability of traditional stratification methods and existing AI methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614969A_ABST
    Figure CN122614969A_ABST
Patent Text Reader

Abstract

The application belongs to the field of three-dimensional geological modeling and large language models, and specifically discloses an intelligent geological model layering method based on RAG technology, which comprises the following steps: obtaining multi-source geological data, vectorizing the multi-source geological data through a text embedding model, and storing the data in a vector database to build a geological knowledge base; converting a geological layering problem input by a user into a query vector through the text embedding model, performing similarity retrieval in the vector database, and obtaining multi-source geological knowledge segments related to the problem; splicing the retrieved multi-source geological knowledge segments with a preset geological field prompt template to form a structured prompt text, inputting the text into a large language model, and generating a structured layering result, wherein the structured layering result comprises a stratigraphic name, lithological characteristics and a thickness range. The application can integrate multi-source heterogeneous geological data and realize intelligent and high-precision geological layering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of three-dimensional geological modeling and large language modeling, and more specifically, relates to an intelligent layering method for geological models based on RAG technology. Background Technology

[0002] As modern geological exploration and resource development extend into deeper and more complex tectonic zones, geological stratification analysis, as a fundamental task, directly impacts the reliability of subsequent resource assessments, engineering designs, and safety evaluations. Traditional geological stratification relies primarily on manual interpretation of various geological data, determining stratigraphic sequences, lithological variations, and thickness distribution through comprehensive analysis. This model played a crucial role in early geological work, but its limitations have become increasingly apparent with the expansion of exploration scale and the diversification of data types.

[0003] On the one hand, the volume of geological data is exploding. Drilling data for a medium-sized deposit can reach tens of thousands of meters, and each report contains thousands of lithological descriptions; in oilfield exploration, well logging curves for a single block can reach hundreds of thousands of records; coupled with unstructured information such as remote sensing imagery and geophysical inversion data, the workload of manually processing this data is enormous. According to statistics, when a senior geological engineer analyzes a medium-sized mining area, data processing alone takes several weeks, and the lag in stratification results severely restricts the timeliness of exploration decisions.

[0004] On the other hand, artificial stratigraphy is highly subjective and reliant on experience. Descriptions of the same geological phenomenon may differ between different regions and engineers; for example, the distinction between "sandy mudstone" and "muddy sandstone" often varies from person to person. Misjudgments are prone to occur when tracing stratigraphic layers in complex structural areas (such as fault fracture zones and fold axes). Furthermore, the lack of unified quantitative standards during cross-regional stratigraphic correlation can lead to accumulated and amplified errors in stratigraphic alignment. In addition, young technicians require long-term experience to master regional geological patterns, resulting in a lengthy talent development cycle that is difficult to meet the demands of current high-intensity exploration.

[0005] In recent years, the application of artificial intelligence technology in the geological field has gradually deepened. Early attempts used machine learning algorithms, such as Support Vector Machines (SVM) and Random Forests, to classify well logging curves for automatic stratigraphic division. However, these methods suffer from drawbacks, including reliance on manually extracted features (such as threshold segmentation of natural gamma and sonic transit time), poor generalization ability (requiring parameter readjustment or even model reconstruction when changing mining areas or stratigraphic types), and inability to handle textual geological descriptions (such as "grayish-black dense massive basalt" in core records). Deep learning technology has alleviated these problems to some extent, but still faces bottlenecks such as difficulty in integrating multi-source heterogeneous data, lack of understanding of geological terminology (such as "integrated contact" and "angular unconformity") and professional logic, and reliance on large-scale labeled data. The emergence of Retrieval-Augmented Generation (RAG) technology has provided a new technical path for geological stratification, but current technologies suffer from severe fragmentation of geological knowledge, poor adaptability of retrieval strategies to stratification tasks, and a lack of a geological-specific prompting engineering framework.

[0006] Therefore, how to integrate multi-source heterogeneous geological data and achieve intelligent and high-precision geological stratification is an urgent problem to be solved. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the purpose of this application is to provide a smart geological model stratification method based on RAG technology, which can integrate multi-source heterogeneous geological data and achieve intelligent and high-precision geological stratification.

[0008] To achieve the above objectives, in a first aspect, this application provides a method for intelligent stratification of geological models based on RAG technology, comprising the following steps: S10, acquire multi-source geological data, vectorize the multi-source geological data through a text embedding model, and store it in a vector database to construct a geological knowledge base covering borehole lithology description, stratigraphic characteristics, and geological background information of the mining area; S20, the geological stratification question input by the user is transformed into a query vector through the text embedding model, and a similarity search is performed in the vector database to obtain multi-source geological knowledge fragments related to the question; S30: The retrieved multi-source geological knowledge fragments are spliced ​​with the preset geological domain prompt template to form a structured prompt text. This text is then input into a large language model to generate a structured layered result, which includes the layer name, lithological characteristics, and thickness range.

[0009] As a further preferred option, step S10, after acquiring multi-source geological data, also includes preprocessing: standardizing the format of the raw data, converting unstructured text into plain text through OCR and rule parsing, interpolating and completing numerical well logging data by depth or time, and marking missing fields; The multi-source geological data includes geological reports and drilling data, and is stored in private domain data; the private domain data includes drilling logs, well logging curve interpretation reports, historical stratification results tables, and digitized files of expert handwritten notes, in the format of JSON, WORD, CSV, TXT or PDF parsed text.

[0010] As a further preferred embodiment, in step S10, the text embedding model is a pre-trained Transformer-based text embedding model, or a variant model finely tuned for the geological field; the Transformer-based text embedding model is BERT, Sentence-BERT, or Text2Vec. The text embedding model generates high-dimensional vector representations for each geological description, each well logging interpretation record, and each borehole lithology segment, and embeds the numerical data into the same vector space through normalization and splicing operations.

[0011] As a further preferred embodiment, in step S10, the vector database supports an approximate nearest neighbor retrieval algorithm, which is Milvus, FAISS, or Weaviate; when storing vectorized data, the vector database attaches metadata tags, which include data source, data type, depth range, mining area, and collection time; the vector database adopts an HNSW or IVFFlat index structure.

[0012] As a further preferred embodiment, in step S20, the process of converting the user-input geological stratification question into a query vector includes: performing word segmentation and entity recognition on the user-input geological stratification question, extracting depth range, mining area name, and stratigraphic code, aligning the identified entities with a geological terminology dictionary, and then generating a query vector containing contextual semantics through the text embedding model.

[0013] As a further preferred embodiment, in step S20, the similarity retrieval uses cosine similarity or Euclidean distance as a metric. The most relevant knowledge fragments to the query vector are retrieved from the vector database, and the top K fragments are selected as candidate knowledge fragments after being sorted by similarity, where K is greater than or equal to 3. Then, the candidate knowledge fragments are filtered a second time to remove records that do not match the depth range or mineral section in the geological stratification question input by the user. The fragments are then sorted in descending order of similarity score to obtain the multi-source geological knowledge fragments.

[0014] As a further preferred embodiment, in step S30, the preset geological domain prompt template includes task instructions, input variable placeholders, output format constraints, and domain constraints; the task instructions require inferring geological strata based on retrieved geological knowledge fragments; the input variable placeholders include retrieved knowledge fragments and depth ranges; the output format constraints require outputting strata names, lithological characteristics, thickness ranges, and stratigraphic ages in tabular form; and the domain constraints are regional geological background constraints.

[0015] As a further preferred embodiment, in step S30, the stratification result also includes stratigraphic age and correlation basis; the stratification result output by the large language model is formatted and its rationality is verified, the rationality verification includes checking whether the thickness is non-negative and whether the stratigraphic age is consistent with the regional geological background.

[0016] As a further preferred embodiment, S40 is also included: converting the layering results into an input format recognizable by the 3D geological modeling software, the input format including the top and bottom depths of each layer, lithological codes, and color identifiers; importing the converted data into the 3D geological modeling software, automatically generating a 3D mesh based on the layering data, constructing a stratigraphic volume model, and applying spatial interpolation algorithms to fill the stratigraphic interfaces in areas without boreholes, making the model continuous throughout the entire area; visualizing the constructed 3D stratigraphic model, the visualization methods including generating contour lines, profiles, 3D mesh models, stratigraphic thickness heat maps, and borehole trajectory and stratigraphic layering distribution maps, and displaying the stratigraphic distribution through color coding, profile display, and depth annotation; the 3D geological modeling software is Surpac, GoCAD, or GOCAD; the visualization results are exported in DXF, STL, or OBJ format.

[0017] Secondly, this application provides a geological model intelligent stratification system based on RAG technology, which implements the steps of any one of the above methods, including: The vector database construction unit is used to acquire multi-source geological data, vectorize the multi-source geological data through a text embedding model, and store it in the vector database to build a geological knowledge base covering borehole lithology description, stratigraphic characteristics, and geological background information of mining areas. The retrieval unit is used to transform the geological stratification question input by the user into a query vector through the text embedding model, perform similarity retrieval in the vector database, and obtain multi-source geological knowledge fragments related to the question. The layered generation unit is used to splice the retrieved multi-source geological knowledge fragments with the preset geological domain prompt template to form a structured prompt text, which is then input into a large language model to generate a structured layered result. The structured layered result includes the layer name, lithological characteristics, and thickness range.

[0018] This application achieves the following effects: By acquiring multi-source geological data and vectorizing it using a text embedding model, it stores the data in a vector database. This allows for the construction of a unified geological knowledge base encompassing borehole lithology descriptions, stratigraphic characteristics, and mining area geological background information. This integrates private domain data such as geological reports and drilling data scattered across different systems into an efficient retrieval index, achieving the fusion and unified representation of multi-source heterogeneous data. Furthermore, by using a text embedding model to transform user-input geological stratification questions into query vectors and performing similarity searches in the vector database, it obtains multi-source geological knowledge fragments most semantically relevant to the question, overcoming the shortcomings of conventional keyword matching in existing technologies that easily miss deep-seated related knowledge. The retrieved knowledge fragments are then concatenated with geological domain prompt templates and input into a large language model. Leveraging the semantic reasoning and knowledge fusion capabilities of the large language model, it automatically generates structured stratification results containing stratigraphic names, lithological characteristics, and thickness ranges. The synergistic effect of these techniques enables this method to integrate multi-source heterogeneous geological data, reduce reliance on manual experience, improve the intelligence and accuracy of stratification analysis, and effectively solve the problems of low efficiency, strong subjectivity, and poor adaptability of traditional stratification methods and existing AI methods. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method provided in the embodiments of this application; Figure 2 This is a detailed flowchart illustrating the method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the meshed structure of the three-dimensional geological model provided in the embodiments of this application; Figure 4 This is a visualization example of geological stratification results provided in the embodiments of this application; Figure 5 This is a three-dimensional spatial mapping of borehole trajectory and stratigraphic distribution provided in the embodiments of this application; Figure 6 This is a thermal map of the surface elevation and stratum thickness of a three-dimensional geological model provided in the embodiments of this application; Figure 7 This is the visualization result of the three-dimensional stratigraphic model provided in the embodiments of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0021] This application's research reveals that the application of RAG to geological stratification is still in its early stages, and existing technologies suffer from the following problems: First, geological knowledge is severely fragmented, with borehole data, literature reports, and expert experience scattered across different systems, lacking a unified vectorized representation method, making it difficult to construct efficient retrieval indexes; second, retrieval strategies are poorly adapted to geological stratification tasks, and conventional keyword matching or semantic similarity retrieval easily misses deep-seated related knowledge (such as the control effect of regional tectonic evolution on stratigraphic distribution); third, there is a lack of a geological-specific Prompt engineering framework, and directly applying general Prompts to geological problems can easily lead to LLM outputs deviating from geological logic (such as confusing stratigraphic age with petrogenesis).

[0022] To address this, this application provides an intelligent stratification method for geological models based on RAG technology. By integrating private domain data such as geological reports and drilling data, and using an embedding model to transform them into vectors and store them in a vector database, a geological knowledge foundation covering borehole lithology, stratigraphic description, and mining area knowledge is constructed. On this basis, through user question-driven retrieval and semantic reasoning of a large language model, structured stratification results containing elements such as stratigraphic names, lithological characteristics, and thickness ranges are automatically generated, thereby significantly improving the intelligence level and accuracy of geological stratification analysis and reducing reliance on manual experience.

[0023] The specific plan is as follows: S1: Vectorized embedding of multi-source geological data and construction of knowledge base; S2: User-question-driven multi-source knowledge retrieval; S3: Generation of geological stratification results enhanced by large language models; S4: Stratigraphic model construction and visualization based on 3D geological modeling tools.

[0024] Specifically, S1 can be: multi-source geological data vectorization embedding and knowledge base construction, including: acquiring geological data from multiple sources, including geological reports and drilling data, and storing them in private domain data (such as structured or semi-structured data in JSON and EXCEL formats); semantically encoding and vectorizing various types of geological data through an embedding model to generate high-dimensional vectors; uniformly storing all vectorized data in a vector database to construct a "geological knowledge base" covering information such as borehole lithology descriptions, stratigraphic characteristics, and geological background of mining areas; the "embedding model" is a pre-trained Transformer-based text embedding model (such as BERT, Sentence-BERT, Text2Vec, etc.), or a variant model fine-tuned for the geological field, supporting high-quality semantic encoding of text such as geological terms, stratigraphic descriptions, and borehole records. "Private domain data" includes drilling logs, well logging curve interpretation reports, historical stratification result tables, and digitized expert handwritten notes accumulated within the enterprise, supporting formats such as JSON, WORD, CSV, TXT, and parsed PDF text.

[0025] The user-question-driven multi-source knowledge retrieval process in S2 first receives a geological stratification question input by the user, such as "What is the geological stratum in a certain depth range?". Second, the question is segmented and entity recognition (such as depth range, mining area name, stratigraphic code, etc.) is performed, and the recognized entities are aligned with a geological terminology dictionary. Then, the question is transformed into a query vector through an embedding model, and similarity matching is performed in the vector database. The top-K (K≥3) relevant knowledge fragments are selected as candidate inputs based on similarity ranking.

[0026] The geological stratification result generation process enhanced by the large language model in S3 first concatenates the retrieved multi-source geological knowledge fragments with a preset geological domain Prompt template to form a structured prompt text. The Prompt template includes task instructions, input variable placeholders, output format constraints, and domain constraints. Then, the prompt text is input into the large language model, which performs semantic reasoning and knowledge fusion to automatically generate a structured stratification result containing elements such as stratigraphic name, lithological characteristics, thickness range, stratigraphic age, and correlation basis.

[0027] The stratigraphic model construction and visualization process in S4, based on 3D geological modeling tools, first imports the structured stratification results into 3D geological modeling software (such as Surpac, GoCAD, GOCAD, etc.), automatically reads the stratigraphic name, top and bottom depth, and lithological properties, and generates contour lines, profiles, and 3D mesh models. Then, the stratigraphic distribution is visualized through color coding, profile display, depth annotation, etc., and can be exported to common formats such as DXF, STL, and OBJ for geological engineers to make decisions and verify.

[0028] In one embodiment, the technical solution for achieving the above objective can specifically be as follows: (Refer to...) Figure 1 This embodiment provides a geological stratification intelligent analysis method based on RAG technology, including the following process: S1: Vectorization and Knowledge Base Construction of Multi-Source Geological Data Furthermore, step S1 specifically includes: S11: Acquisition and Preprocessing of Multi-Source Geological Data Acquire geological data from multiple sources, including geological reports (PDF / Word format), drilling data (CSV / TXT / JSON), well logging curves (MAT / LAW), and enterprise private domain data (such as historical stratification result tables and digitized files of expert handwritten records). Standardize the format of the raw data, converting unstructured text into processable plain text through OCR and rule parsing. Perform interpolation to complete numerical well logging data by depth or time, and mark missing fields to ensure the integrity of subsequent vectorization processes.

[0029] S12: Embedding Model Selection and Vectorization A pre-trained Transformer-based text embedding model (such as Sentence-BERT or Text2Vec) was selected and fine-tuned based on geological corpora to accurately encode geological terms (such as "angular unconformity," "conformable contact," "sandstone and mudstone," etc.). High-dimensional vector representations were generated for each geological description, each well logging interpretation record, and each borehole lithology segment. Numerical data (such as depth, thickness, and porosity) were embedded into the same vector space through normalization and concatenation operations for unified storage with text vectors.

[0030] S13: Vector Database Construction All vectorized data is stored in a vector database that supports approximate nearest neighbor retrieval (such as Milvus, FAISS, Weaviate, etc.), an index structure is built (HNSW or IVFFlat), and metadata tags are attached, including data source, data type, depth range, mining area, collection time, etc., forming... Figure 2 The "geological knowledge base" shown provides efficient support for subsequent searches.

[0031] S2: User-Question-Driven Multi-Source Knowledge Retrieval Reference Figure 1 and Figure 2 Step S2 is as follows: S21: Problem Analysis and Query Generation The system receives geological stratification questions input by users, such as "What are the geological strata in a certain mining section at a depth of 500m to 800m?". It performs word segmentation and entity recognition on the question text, extracts key parameters (depth range, mining section name, target stratum features, etc.), and transforms the question into a query vector through an embedding model.

[0032] S22: Vector Similarity Retrieval In the vector database, using the Query Vector as input, the most relevant knowledge fragments in the Top-K (K≥3) are retrieved from the geological knowledge base using cosine similarity or Euclidean distance metrics. These fragments include lithological descriptions, well logging interpretations, regional geological backgrounds, and historical stratification results for the corresponding depth range.

[0033] S23: Filtering and sorting of search results The search results are filtered a second time to remove records that do not match the depth range or mining section of the question, and then sorted in descending order of similarity score to form the final "multi-source geological knowledge fragment set" to provide input for subsequent generation steps.

[0034] S3: Generation of Geological Layering Results Enhanced by Large Language Model Reference Figure 2 Step S3 is as follows: S31: Prompt Template Construction Based on the characteristics of geological stratification tasks, a domain-specific Prompt template is designed, including task instructions ("Please infer the geological strata of a certain depth range based on the following geological knowledge fragments"), input variable placeholders ({retrieved knowledge fragments}, {depth range}), output format constraints (outputting strata names, lithological characteristics, thickness ranges, and stratigraphic ages in tabular form), and regional constraints.

[0035] S32: Knowledge Integration and Reasoning The retrieved multi-source geological knowledge fragments are concatenated with a Prompt template to form structured prompt text, which is then input into a large language model (such as ChatGLM, Qwen, Llama3, etc.). Based on understanding the context, the model fuses the multi-source information and performs reasoning based on geological logic to generate preliminary hierarchical results.

[0036] S33: Result Structuring and Validation Format the LLM output to ensure that fields such as stratigraphic name, lithological characteristics, thickness range, and stratigraphic age are complete, and perform reasonableness checks on key data (such as non-negative thickness and consistency between stratigraphic age and regional geological background), and make manual or automatic corrections when necessary.

[0037] S4: 3D Geological Modeling and Visualization Reference Figures 3 to 7 Step S4 is as follows: S41: Data Transformation and Model Import The structured stratification results are converted into an input format that can be recognized by 3D geological modeling software (such as Surpac, GoCAD, GOCAD, etc.), including the top and bottom depths of each stratum, lithological codes, color codes, etc.

[0038] S42: Construction of 3D Stratigraphic Model In the modeling software, a 3D mesh is automatically generated based on the layered data to construct a stratigraphic volume model. Spatial interpolation algorithms are then applied to fill the stratigraphic interfaces in areas without boreholes, ensuring the model's continuity across the entire region. According to... Figure 4 The layered results are visualized and projected onto the borehole, generating... Figure 5 The diagram shows the borehole trajectory and stratigraphic distribution in three-dimensional space.

[0039] S43: Visualization and Results Output Visualize the three-dimensional stratigraphic model. Figure 6 The thermal map of the formation thickness shown Figure 7 The visualization results of the three-dimensional stratigraphic model shown are helpful for geological engineers to analyze and verify.

[0040] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A smart stratification method for geological models based on RAG technology, characterized in that, Includes the following steps: S10, acquire multi-source geological data, vectorize the multi-source geological data through a text embedding model, and store it in a vector database to construct a geological knowledge base covering borehole lithology description, stratigraphic characteristics, and geological background information of the mining area; S20, the geological stratification question input by the user is transformed into a query vector through the text embedding model, and a similarity search is performed in the vector database to obtain multi-source geological knowledge fragments related to the question; S30: The retrieved multi-source geological knowledge fragments are spliced ​​with the preset geological domain prompt template to form a structured prompt text. This text is then input into a large language model to generate a structured layered result, which includes the layer name, lithological characteristics, and thickness range.

2. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S10, after acquiring multi-source geological data, preprocessing is also included: standardizing the format of the raw data, converting unstructured text into plain text through OCR and rule parsing, interpolating and completing numerical well logging data by depth or time, and marking missing fields. The multi-source geological data includes geological reports and drilling data, and is stored in private domain data; the private domain data includes drilling logs, well logging curve interpretation reports, historical stratification results tables, and digitized files of expert handwritten notes, in the format of JSON, WORD, CSV, TXT or PDF parsed text.

3. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S10, the text embedding model is a pre-trained Transformer-based text embedding model, or a variant model finely tuned for the geological domain; the Transformer-based text embedding model is BERT, Sentence-BERT, or Text2Vec. The text embedding model generates high-dimensional vector representations for each geological description, each well logging interpretation record, and each borehole lithology segment, and embeds the numerical data into the same vector space through normalization and splicing operations.

4. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S10, the vector database supports an approximate nearest neighbor retrieval algorithm, which is Milvus, FAISS, or Weaviate; when storing vectorized data, the vector database attaches metadata tags, which include data source, data type, depth range, mining area, and collection time. The vector database uses an HNSW or IVFFlat index structure.

5. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S20, the process of converting the geological stratification question input by the user into a query vector includes: performing word segmentation and entity recognition on the geological stratification question input by the user, extracting depth range, mining area name, and stratigraphic code, aligning the identified entities with a geological terminology dictionary, and then generating a query vector containing contextual semantics through the text embedding model.

6. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S20, the similarity retrieval uses cosine similarity or Euclidean distance as a metric. The most relevant knowledge fragments to the query vector are retrieved from the vector database, and the top K fragments are selected as candidate knowledge fragments after being sorted by similarity, where K is greater than or equal to 3. Then, the candidate knowledge fragments are filtered a second time to remove records that do not match the depth range or mineral section in the geological stratification question input by the user. The fragments are then sorted in descending order of similarity score to obtain the multi-source geological knowledge fragments.

7. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S30, the preset geological domain prompt template includes task instructions, input variable placeholders, output format constraints, and domain constraints. The task instructions require inferring geological strata based on retrieved geological knowledge fragments. The input variable placeholders include retrieved knowledge fragments and depth ranges. The output format constraints require outputting the stratum name, lithological characteristics, thickness range, and stratigraphic age in tabular form. The domain constraints are regional geological background constraints.

8. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, In step S30, the stratification results also include stratigraphic age and correlation basis; the stratification results output by the large language model are formatted and validated for rationality, and the rationality validation includes checking whether the thickness is non-negative and whether the stratigraphic age is consistent with the regional geological background.

9. The intelligent stratification method for geological models based on RAG technology as described in claim 1, characterized in that, The method also includes S40: converting the layering results into an input format recognizable by the 3D geological modeling software, the input format including the top and bottom depths of each layer, lithological codes, and color identifiers; importing the converted data into the 3D geological modeling software, automatically generating a 3D mesh based on the layering data, constructing a stratigraphic volume model, and applying spatial interpolation algorithms to fill the stratigraphic interfaces in areas without boreholes, making the model continuous throughout the entire area; visualizing the constructed 3D stratigraphic model, the visualization methods include generating contour lines, profiles, 3D mesh models, stratigraphic thickness heat maps, and borehole trajectory and stratigraphic layer distribution maps, and displaying the stratigraphic distribution through color coding, profile display, and depth annotation; the 3D geological modeling software is Surpac, GoCAD, or GOCAD; the visualization results are exported in DXF, STL, or OBJ format.

10. A geological model intelligent stratification system based on RAG technology, characterized in that, The steps for implementing the method according to any one of claims 1 to 9 include: The vector database construction unit is used to acquire multi-source geological data, vectorize the multi-source geological data through a text embedding model, and store it in the vector database to build a geological knowledge base covering borehole lithology description, stratigraphic characteristics, and geological background information of mining areas. The retrieval unit is used to transform the geological stratification question input by the user into a query vector through the text embedding model, perform similarity retrieval in the vector database, and obtain multi-source geological knowledge fragments related to the question. The layered generation unit is used to splice the retrieved multi-source geological knowledge fragments with the preset geological domain prompt template to form a structured prompt text, which is then input into a large language model to generate a structured layered result. The structured layered result includes the layer name, lithological characteristics, and thickness range.