Geological text information automatic extraction method and system based on large model

By constructing a digital tagging system and combining it with a large model in the geological field, the problem of inaccurate and unreliable extraction of geological text information was solved, achieving efficient and accurate information extraction and meeting the professional needs of geological research.

CN121858864APending Publication Date: 2026-04-14武汉智博创享科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate geological expertise with large-scale models, resulting in inaccurate and unreliable extraction of geological text information, which fails to meet the professional needs of geological research.

Method used

A digital tagging system for the geological field is constructed, combining cosine similarity algorithm and Transformer self-attention mechanism. Geological text information is automatically extracted through a large model, including digital tag management, text standardization processing, and result verification and storage.

Benefits of technology

It significantly improves the accuracy and reliability of geological text information extraction, reduces manual intervention, increases processing efficiency, reduces costs, and meets the professional precision requirements of geological research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858864A_ABST
    Figure CN121858864A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of household kitchens, and discloses a geological text information automatic extraction method and system based on a large model, and the method comprises the following steps: S1, geological field digital label system construction, S2, digital label system management, S3, geological electronic text standardization processing, S4, geological electronic data automatic extraction, and S5, geological electronic extraction result verification and storage. According to the method, a digital tag system in the geological field is constructed, semantic association of text blocks and tags is screened in combination with a cosine similarity algorithm in S4, geological entities and relational semantics are captured by a Transform self-attention mechanism, a confidence score is generated by a Softmax algorithm, a large model accurately focuses on core information such as ore deposit types and geological structures, and irrelevant interference is avoided. According to the technical path, the problem that an existing model is insufficient in understanding of geological texts is solved, the extraction accuracy is remarkably improved, it is ensured that output original texts and triple results are reliable, and the core requirements of geological research and resource exploration for professional precision are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of home kitchen technology, specifically to a method and system for automatically extracting geological text information based on a large model. Background Technology

[0002] As geological research progresses, a massive amount of geological text data exists in unstructured or semi-structured forms, including electronic geological reports, digital exploration records, and electronic academic papers. These texts contain core professional information such as stratigraphic structure, rock characteristics, mineral resource distribution, and geological hazard conditions. They serve as important data support for geological research, resource exploration, and disaster prevention. In the traditional model, manual screening and extraction of key information are required, forming the basic scenario for geological text information processing.

[0003] In existing technologies, some geological texts are extracted using global knowledge extraction methods based on large models and RAG technology. These methods segment and vectorize the text, cluster it into a vector library, filter out relevant entity pairs, and then combine RAG retrieval with large model reasoning to generate entity relationships. There are also text information extraction methods based on geological constraints. These methods define a geological ontology model, train the model using labeled datasets, and then use the model to perform entity recognition and relationship prediction on unlabeled text.

[0004] The most critical shortcoming of existing technologies lies in their failure to effectively integrate geological expertise with large-scale modeling techniques to ensure extraction effectiveness. Methods based on RAG technology suffer from uncertainty in the logical relationships of entities inferred by the large model due to a lack of geological expertise. Methods based on geological constraints, on the other hand, are highly dependent on the quality of the geological entity relationship dataset. Incomplete geological ontology model definitions or biased labeled samples directly impact model training. Ultimately, neither method can meet the professional requirements of the geological field for accurate and reliable extraction of textual information. Therefore, we propose an automatic geological textual information extraction method and system based on large-scale models. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method and system for automatic extraction of geological text information based on large models, which solves the problem that existing technologies cannot meet the professional needs of the geological field for accurate and reliable extraction of text information.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an automatic extraction method for geological text information based on a large model, comprising the following steps: S1. Construction of a digital tagging system in the geological field: acquire electronic text data in the geological field through electronic data acquisition tools, sort out the core knowledge in the electronic text data, and construct a digital domain knowledge tagging system covering key information in the geological profession; S2, Digital Tag System Management: Based on the digital domain knowledge tag system built by S1, configure the digital management function of the tag system in the electronic data management system; S3. Geological electronic text standardization processing: Receive geological texts in different electronic formats, perform digital text extraction operations on the geological texts in different electronic formats to unify the electronic text format, and then, based on the digital characteristics of the geological electronic text associated with the tag system after digital management in S2, set preset parameters through the electronic text segmentation tool, and perform digital segmentation operations on the unified electronic text to obtain digital text blocks that meet the requirements of large model electronic data input. S4. Automatic extraction of geological electronic data: Input the digital text block obtained in S3 into the large model, combine it with the digital domain knowledge tag system after digital management in S2 to generate standardized digital extraction prompt words, and output the electronic extraction results of geological information through the large model. S5. Verification and storage of geological electronic extraction results: Perform electronic verification operation on the geological information electronic extraction results output by S4, and store the verified geological information electronic extraction results in the electronic structured database and the electronic map database respectively.

[0007] Preferably, the electronic text data in the geological field in S1 includes electronic geological reports, digital exploration records, and electronic academic papers. The digital field knowledge tagging system covers digital tags related to geological entities, geological relationships, and geological attributes. Among them, the digital tags related to geological entities include tags corresponding to deposit types and stratigraphic units, the digital tags related to geological relationships include tags corresponding to tectonic control and mineralization association, and the digital tags related to geological attributes include tags corresponding to geochemical characteristics and mineral composition.

[0008] Preferably, the digital management function in S2 includes adding, deleting, modifying, and querying electronic tags. Through this digital management function, the digital domain knowledge tag system is updated in the electronic data management system to match the knowledge and change requirements in newly added electronic text data in the geological field.

[0009] Preferably, the geological texts in different electronic formats in S3 include text document format electronic texts and PDF format electronic texts; a direct digital reading operation is performed on the text document format electronic texts to extract text, and an electronic OCR recognition operation is performed on the PDF format electronic texts containing image texts to convert the image texts into editable digital texts.

[0010] Preferably, the preset parameters in S3 are set according to the digital structure features of the geological electronic text. The digital structure features include the chapter digital identifiers and paragraph semantic digital associations of the geological electronic text. The preset parameters control the information integrity and electronic data length range of each digital text block after digital segmentation.

[0011] Preferably, the standardized digital extraction prompts in S4 are in digital format text, containing digital domain knowledge tag content and electronic extraction result output structure definition. The electronic extraction result output structure definition clearly defines the original text fragments in digital form and presents them in the form of confidence scores.

[0012] Preferably, the operation of storing the verified geological information electronic extraction results to the electronic structured database in S5 includes organizing the electronic extraction results according to a preset digital JSON format and writing them; the operation of storing the verified geological information electronic extraction results to the electronic map database includes generating a digital Cypher statement based on the triplet data in the electronic extraction results, and writing the data to the electronic map database through the digital Cypher statement.

[0013] The automatic geological text information extraction system based on large models includes the following modules: Domain Tag Management Module: Acquire geological knowledge data, organize key geological information, construct a domain knowledge tag system, and configure management functions for this tag system; Text standardization processing module: Receives geological texts in different formats, performs text extraction operations on the geological texts in different formats to unify the text format, sets the block parameters based on the geological text features associated with the domain knowledge tag system output by the domain tag management module, performs block operations on the unified text, and obtains text blocks that meet the input requirements of the large model. The geological information extraction and storage module inputs the text blocks output by the text standardization processing module into the large model, combines the domain knowledge tag system output by the domain tag management module to generate standardized extraction prompt words, outputs the geological information extraction results through the large model, performs a verification operation on the extraction results, and stores the verified extraction results in the structured database and the graph database respectively.

[0014] Preferably, the geological knowledge data acquired by the domain tag management module includes geological literature, exploration reports, and academic papers; the constructed domain knowledge tag system covers tags related to geological entities, geological relationships, and geological attributes, wherein geological entity-related tags include tags corresponding to deposit types and stratigraphic units, geological relationship-related tags include tags corresponding to tectonic control and mineralization association, and geological attribute-related tags include tags corresponding to geochemical characteristics and mineral composition; the configured management functions include tag addition, tag deletion, tag modification, and tag query.

[0015] Preferably, when the text standardization processing module performs text extraction, it performs direct reading of geological text in text document format and OCR recognition of geological text in PDF format that contains image text. The set block parameters are determined based on the chapter division and paragraph semantic association of the geological text. The standardized extraction prompts generated by the geological information extraction and storage module include domain knowledge tag content and extraction result output structure definition. The extraction result output structure definition clearly defines the presentation format of the original text fragment, entity 1-relation-entity 2 triple, and confidence score.

[0016] This invention provides a method and system for automatically extracting geological text information based on large models. It has the following beneficial effects: 1. This invention constructs a digital tagging system for the geological field and, within the S4 framework, combines a cosine similarity algorithm to filter semantic associations between text blocks and tags, a Transformer self-attention mechanism to capture the semantics of geological entities and relationships, and a Softmax algorithm to generate confidence scores. This allows the large model to accurately focus on core information such as deposit types and geological structures, avoiding irrelevant interference. This technical approach addresses the problem of insufficient understanding of geological text in existing models, significantly improves extraction accuracy, ensures the reliability of the output original text and triplet results, and meets the core requirements of professional precision in geological research and resource exploration.

[0017] 2. This invention utilizes S3 Chinese text standardization processing (OCR recognition of text in PDF images, combined with the Chonkie library and paragraph semantic association for segmentation), along with a fully automated algorithm, to replace the tedious traditional manual data annotation and page-by-page processing of different text formats. Compared to manual processing, the processing time for a single geological report is reduced from several days to hours, increasing efficiency by several to tens of times, significantly reducing repetitive work for geological experts and lowering project costs for reusing geological data.

[0018] 3. This invention utilizes the tag addition, deletion, modification, and query management functions in S2 and the standardized prompt word automatic generation algorithm in S4 to transform geological expert knowledge into reusable digital tags and guidance rules. Non-professionals do not need to master model operation or in-depth geological knowledge; they can complete the extraction simply by viewing confidence level markers and modifying results. Simultaneously, expert experience is solidified into knowledge assets, preventing experience loss. Newcomers can quickly learn through the knowledge graph in the graph database, providing structured data support for geological research and disaster prevention. Attached Figure Description

[0019] Figure 1 This is a step-by-step diagram of the method for automatically extracting geological text information based on a large model; Figure 2 This is a flowchart of a method for automatically extracting geological text information based on a large model. Detailed Implementation

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Example: Please see the appendix Figure 1 - Appendix Figure 2 This invention provides a method for automatically extracting geological text information based on a large model, comprising the following steps: S1. Construction of a digital tagging system in the geological field: acquire electronic text data in the geological field through electronic data acquisition tools, sort out the core knowledge in the electronic text data, and construct a digital domain knowledge tagging system covering key information in the geological profession; Before information extraction, a comprehensive approach combining data compilation and expert summarization was used to organize core knowledge within the geological field, thereby constructing a complete and targeted knowledge tagging system for the current domain. During data compilation, various textual materials, including geological literature, exploration reports, and academic papers, were systematically analyzed to extract frequently occurring key information points. Simultaneously, experienced experts in the geological field were invited to participate, combining their professional knowledge and practical experience to screen, classify, and supplement the extracted information points. Ultimately, a tagging system covering ore deposit types, mineralization processes, geological structures, and other relevant important information points (such as geochemical characteristics and mineral composition) was determined. S2, Digital Tag System Management: Based on the digital domain knowledge tag system built by S1, configure the digital management function of the tag system in the electronic data management system; The digital management function in S2 includes adding, deleting, modifying, and querying electronic tags. This function updates the digital domain knowledge tag system within the electronic data management system to match the knowledge and changing needs of newly added electronic text data in the geological field. By setting tags, the system uses preset domain knowledge tags as a basis for targeted processing of the input geological text information during information extraction. This allows the model to focus on relevant information points such as deposit type, mineralization process, and geological structure when extracting information, improving the accuracy and effectiveness of information extraction and effectively avoiding interference from irrelevant information and the omission of core information. S3. Geological electronic text standardization processing: Receive geological texts in different electronic formats, perform digital text extraction operations on the geological texts in different electronic formats to unify the electronic text format, and then, based on the digital characteristics of the geological electronic text associated with the tag system after digital management in S2, set preset parameters through the electronic text segmentation tool, and perform digital segmentation operations on the unified electronic text to obtain digital text blocks that meet the requirements of large model electronic data input. The geological texts in S3, in different electronic formats, include text document format electronic text and PDF format electronic text. For text documents, the system can directly read the text content to obtain text information. For PDF documents, the system automatically calls the integrated OCR recognition model to convert the text information in the image into editable text. After converting all input documents into text, the system uses the Chonkie library to segment the text. Different segmentation parameters (such as delimiters, paragraph semantics, etc.) are used to segment the text, ultimately dividing it into text paragraphs with complete information and context lengths that meet the input requirements of the large model. Through this text standardization process, unified processing of texts in different formats is achieved, reducing the complexity of data extraction. The standardized text maintains consistency in format and structure, providing a standardized data foundation for subsequent information extraction and reducing extraction errors caused by non-standard data. Scientific and reasonable text segmentation ensures that the context length of the processed text paragraphs meets the input limitations of the large model, enabling the large model to process the text more efficiently and accurately, thus improving the information extraction effect. The standardization process simplifies the data processing flow, reduces manual intervention, and improves the overall efficiency of information extraction. The preset parameters in S3 are set according to the digital structure features of the geological electronic text. The digital structure features include the chapter digital identifier and paragraph semantic digital association of the geological electronic text. The preset parameters control the information integrity and electronic data length range of each digital text block after digital segmentation. S4. Automatic extraction of geological electronic data: Input the digital text block obtained in S3 into the large model, combine it with the digital domain knowledge tag system after digital management in S2 to generate standardized digital extraction prompt words, and output the electronic extraction results of geological information through the large model. The standardized digital extraction prompts in S4 are in digital format text, containing digital domain knowledge tags and electronic extraction result output structure definitions. The electronic extraction result output structure definitions clearly define the original text fragments in digital form and present them in the form of confidence scores. During text extraction, the system uses the previously segmented text as context data for the large model, inputting it sequentially. Simultaneously, the system utilizes the pre-built domain knowledge tagging system, combining the tag content with built-in prompt templates to automatically generate standardized extraction prompts. These prompts specify the extraction content and return data structure (including the original text, triple content, and correlation degree), ensuring the model outputs standardized extraction results according to preset requirements, facilitating subsequent service parsing and processing. After obtaining the extraction results returned by the large model, the results are synchronized to the text display page, allowing users to view and modify the results on the page. This includes the following algorithms: Step 1: Semantic matching to filter related numeric tags The system of digital domain knowledge tags (such as "deposit type", "ore-controlling structure", and "metallogenic epoch") managed by S2 is invoked to perform semantic association judgment on the digital text blocks output by S3, filtering out tags that are semantically close to the text blocks (to avoid irrelevant tags interfering with the generation of prompt words). This judgment is implemented using the cosine similarity algorithm: first, each digital tag is converted into a vector. (The dimension is m, where m represents the total length of the vocabulary, determined by the geological vocabulary constructed during the text preprocessing stage), the value of the k-th dimension in the vector. The TF-IDF value of the k-th word in the tag text (the TF-IDF value is calculated through the key information point extraction logic in S1, reflecting the importance of the word in the tag), and the numeric text block is converted into a vector. The value of the k-th dimension in the vector Let be the TF-IDF value of the k-th word in the text block; then calculate the cosine similarity between the two vectors. The formula is:

[0022] in Representing vectors and The dot product (calculates the degree of overlap between two vectors; a larger value indicates greater semantic overlap). Representing vectors The modulus (calculated by the formula is) (reflecting the overall "strength" of the vector) Representing vectors The modulus (calculated by the formula is) Set a similarity threshold (e.g., 0.5). If a tag and a text block have a similarity threshold, the similarity threshold will be lowered. If the value is greater than or equal to the threshold, it is determined to be an associated tag and included in the subsequent prompt word generation; if If the threshold is reached, then the label will be excluded; Step 2: Generate standardized number extraction prompts Based on the associated tags selected in step 1, the system's built-in prompt word template is invoked (the template has a preset "extraction target + output structure" framework, such as "extract information related to {tag} from the text, and the output format includes: original text fragment, entity-relation-entity triple, degree of association"). The associated tags are filled into the {tag} position of the template, and the output structure definition of the electronic extraction results is clarified (such as the original text fragment must retain the information source context, the triple must conform to the "entity 1-relation-entity 2" format, and the degree of association must be presented as a value of 0-1), forming standardized digital extraction prompt words (the prompt words are UTF-8 encoded digital format text to ensure that large models can directly read and parse them). Step 3: Large-scale model extracts geological information based on self-attention mechanism The digital text blocks obtained in S3 (encoded in UTF-8 format) and the standardized digital extraction prompts generated in step 2 are input into a large model (such as a geological-specific large model based on Llama3 fine-tuning). The large model uses the self-attention mechanism of the Transformer architecture to capture key semantics related to the tags in the text blocks (such as the association between "a copper deposit", "porphyry type", and "belongs to" in "a copper deposit belongs to porphyry type"). First, word embedding processing is performed on the digital text blocks to obtain the word embedding matrix. (dimension is) , Indicates the number of words in a text block. The hidden layer dimension of the model is represented by 768, and then passed through three learnable linear transformation matrices. , , (All dimensions are) , The dimensions representing the query and key vectors are typically taken as follows: , For the number of attention heads, such as 12), Convert them into query vector matrices respectively Key vector matrix Value vector matrix The formulas are respectively , , ,in This represents the "query requirement" for each term (e.g., "porphyry type" requires querying the associated "mineral deposit" entity). The "feature identifier" representing each word (e.g., the feature identifier of "a copper deposit" is used to match queries for other words) The semantic content of each word is represented (e.g., the semantic content of "belongs to" is used to express the relationship between entities); then the attention score is calculated, first for... and transpose matrix Perform a dot product operation (to obtain) , dimension Each element represents the first The word and the first (Original semantic association score of each word), then divided by ( As a scaling factor, avoid Excessive weights can cause the dot product to overflow, leading to gradient vanishing in the subsequent Softmax function. The score is then normalized using the Softmax function (converting the score into 0-1 attention weights, ensuring the sum of weights in each row is 1; a larger weight indicates a stronger semantic contribution of the corresponding word to the current word). Finally, the normalized weights are summed with... Perform matrix multiplication to obtain the attention feature matrix, as shown in the formula:

[0023] Based on this attention feature matrix and combined with the requirements of standardized digit extraction prompts, the large model outputs preliminary electronic extraction results of geological information, including original text fragments and triplet content. Step 4: Calculate the confidence score using the Softmax algorithm. For the preliminary extraction results output in step 3 (such as each triplet), extract the raw prediction score from the output of the last fully connected layer of the large model. ( (This is a real number with no fixed range; a larger value indicates a higher initial confidence level of the model in the result.) The Softmax algorithm is then used to... Convert to 0-1 confidence score (i.e., the "degree of relevance" in the disclosure document), the formula is:

[0024] in is the natural constant (approximately 2.718, used to map the original fraction to a positive value). This represents the total number of results extracted within the same category (e.g., the number of all possible triples under the "Deposit Type" label, such as 3). This is the exponential sum of the raw scores of all extracted results within the same category (used for normalization to ensure that the sum of the confidence scores of all results within the same category is 1). For the first The final confidence score of each sampling result ( The closer the result is to 1, the higher its reliability; set a confidence threshold (e.g., 0.8). If the threshold is met, the extraction result is marked as "reliable"; otherwise... If the threshold is not met, it is marked as "to be verified"; By adopting the above-described automated text extraction and storage process, non-professionals can extract and modify geological text information without needing specialized knowledge or experience using complex tools, significantly reducing operational difficulty. From generating extraction prompts and modifying results to final storage, the entire process reduces manual intervention, lowers labor costs, and improves the overall efficiency of geological information processing. S5. Verification and storage of geological electronic extraction results: Perform electronic verification operation on the geological information electronic extraction results output by S4, and store the verified geological information electronic extraction results in the electronic structured database and the electronic map database respectively. The operation of storing the verified geological information electronic extraction results into the electronic structured database in S5 includes organizing the electronic extraction results according to a preset digital JSON format and writing them; the operation of storing the verified geological information electronic extraction results into the electronic map database includes generating a digital Cypher statement based on the triplet data in the electronic extraction results, and writing the data into the electronic map database through the digital Cypher statement.

[0025] The automatic geological text information extraction system based on large models includes the following modules: Domain Tag Management Module: Acquire geological knowledge data, organize key geological information, construct a domain knowledge tag system, and configure management functions for this tag system; Text standardization processing module: Receives geological texts in different formats, performs text extraction operations on the geological texts in different formats to unify the text format, sets the block parameters based on the geological text features associated with the domain knowledge tag system output by the domain tag management module, performs block operations on the unified text, and obtains text blocks that meet the input requirements of the large model. The geological information extraction and storage module inputs the text blocks output by the text standardization processing module into the large model, combines the domain knowledge tag system output by the domain tag management module to generate standardized extraction prompt words, outputs the geological information extraction results through the large model, performs a verification operation on the extraction results, and stores the verified extraction results in the structured database and the graph database respectively.

[0026] The domain tag management module acquires geological domain knowledge data including geological literature, exploration reports, and academic papers; the constructed domain knowledge tag system covers tags related to geological entities, geological relationships, and geological attributes, among which geological entity-related tags include tags corresponding to deposit types and stratigraphic units, geological relationship-related tags include tags corresponding to tectonic control and metallogenic association, and geological attribute-related tags include tags corresponding to geochemical characteristics and mineral composition; the configured management functions include tag addition, tag deletion, tag modification, and tag query.

[0027] When the text standardization processing module performs text extraction, it directly reads geological text in text document format and performs OCR recognition on geological text in PDF format that contains image-based text. The segmentation parameters are determined based on the chapter division and paragraph semantic relationships of the geological text. The standardized extraction prompts generated by the geological information extraction and storage module include domain knowledge tags and the extraction result output structure definition. The extraction result output structure definition clearly defines the original text fragments and presents them in the form of confidence scores.

[0028] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for automatically extracting geological text information based on a large model, characterized in that, Includes the following steps: S1. Construction of a digital tagging system in the geological field: acquire electronic text data in the geological field through electronic data acquisition tools, sort out the core knowledge in the electronic text data, and construct a digital domain knowledge tagging system covering key information in the geological profession; S2, Digital Tag System Management: Based on the digital domain knowledge tag system built by S1, configure the digital management function of the tag system in the electronic data management system; S3. Geological electronic text standardization processing: Receive geological texts in different electronic formats, perform digital text extraction operations on the geological texts in different electronic formats to unify the electronic text format, and then, based on the digital characteristics of the geological electronic text associated with the tag system after digital management in S2, set preset parameters through the electronic text segmentation tool, and perform digital segmentation operations on the unified electronic text to obtain digital text blocks that meet the requirements of large model electronic data input. S4. Automatic extraction of geological electronic data: Input the digital text block obtained in S3 into the large model, combine it with the digital domain knowledge tag system after digital management in S2 to generate standardized digital extraction prompt words, and output the electronic extraction results of geological information through the large model. S5. Verification and storage of geological electronic extraction results: Perform electronic verification operation on the geological information electronic extraction results output by S4, and store the verified geological information electronic extraction results in the electronic structured database and the electronic map database respectively.

2. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The geological electronic text data in S1 includes electronic geological reports, digital exploration records, and electronic academic papers. The digital domain knowledge tagging system covers digital tags related to geological entities, geological relationships, and geological attributes. Among them, the geological entity-related digital tags include tags corresponding to ore deposit types and stratigraphic units, the geological relationship-related digital tags include tags corresponding to tectonic control and mineralization association, and the geological attribute-related digital tags include tags corresponding to geochemical characteristics and mineral composition.

3. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The digital management function in S2 includes adding, deleting, modifying, and querying electronic tags. Through this digital management function, the digital domain knowledge tag system is updated in the electronic data management system to match the knowledge and change requirements in newly added electronic text data in the geological field.

4. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The geological texts in different electronic formats in S3 include text document format electronic texts and PDF format electronic texts; a direct digital reading operation is performed on the text document format electronic texts to extract the text, and an electronic OCR recognition operation is performed on the PDF format electronic texts containing image texts to convert the image texts into editable digital texts.

5. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The preset parameters in S3 are set according to the digital structure characteristics of the geological electronic text. The digital structure characteristics include the chapter digital identifiers and paragraph semantic digital associations of the geological electronic text. The preset parameters control the information integrity and electronic data length range of each digital text block after digital segmentation.

6. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The standardized digital extraction prompts in S4 are in digital format text, containing digital domain knowledge tags and an electronic extraction result output structure definition. The extraction result output structure definition clearly defines the presentation format of the original text fragment, entity 1-relation-entity 2 triples, and confidence score.

7. The method for automatically extracting geological text information based on a large model according to claim 1, characterized in that, The operation of storing the verified geological information electronic extraction results into the electronic structured database in S5 includes organizing the electronic extraction results according to a preset digital JSON format and writing them; the operation of storing the verified geological information electronic extraction results into the electronic map database includes generating a digital Cypher statement based on the triplet data in the electronic extraction results, and writing the data into the electronic map database through the digital Cypher statement.

8. The automatic geological text information extraction system based on a large model according to claim 1, characterized in that, Includes the following modules: Domain Tag Management Module: Acquire geological knowledge data, organize key geological information, construct a domain knowledge tag system, and configure management functions for this tag system; Text standardization processing module: Receives geological texts in different formats, performs text extraction operations on the geological texts in different formats to unify the text format, sets the block parameters based on the geological text features associated with the domain knowledge tag system output by the domain tag management module, performs block operations on the unified text, and obtains text blocks that meet the input requirements of the large model. The geological information extraction and storage module inputs the text blocks output by the text standardization processing module into the large model, combines the domain knowledge tag system output by the domain tag management module to generate standardized extraction prompt words, outputs the geological information extraction results through the large model, performs a verification operation on the extraction results, and stores the verified extraction results in the structured database and the graph database respectively.

9. The automatic geological text information extraction system based on a large model according to claim 8, characterized in that, The domain tag management module acquires geological domain knowledge data including geological literature, exploration reports, and academic papers; the constructed domain knowledge tag system covers tags related to geological entities, geological relationships, and geological attributes, among which geological entity-related tags include tags corresponding to deposit types and stratigraphic units, geological relationship-related tags include tags corresponding to tectonic control and metallogenic association, and geological attribute-related tags include tags corresponding to geochemical characteristics and mineral composition; the configured management functions include tag addition, tag deletion, tag modification, and tag query.

10. The automatic geological text information extraction system based on a large model according to claim 8, characterized in that, When the text standardization processing module performs text extraction, it directly reads geological text in text document format and performs OCR recognition on geological text in PDF format that contains image-based text. The segmentation parameters are determined based on the chapter division and paragraph semantic association of the geological text. The standardized extraction prompts generated by the geological information extraction and storage module include domain knowledge tags and the extraction result output structure definition. The extraction result output structure definition clearly defines the presentation format of the original text fragment, entity 1-relation-entity 2 triples, and confidence score.