A method and system for extracting knowledge from geological prospecting literature based on large models
A large model-based system automates the extraction of triadic groups from geological mining literature, improving efficiency and accuracy by reducing human error and adapting to diverse language environments, thus overcoming the inefficiencies of traditional human-based methods.
Patent Information
- Application Number
- CN202411915502.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The extraction of mineral literature in the prior art relies on manual reading and recognition, is inefficient and error-prone, especially in multilingual and complex language environments, it is difficult to effectively extract triple data, and requires professional assistance, and cannot handle the language structure and writing logic of ancient and rare or future documents.
The geological ore-prospecting literature knowledge extraction method is adopted based on large models. By establishing a triple-group learning model, the triple-group data is identified and extracted, the knowledge base is formed and the knowledge graph is drawn, the model is optimized using machine learning and iterative mechanisms, information is extracted across the language environment, and the dependence on external assistance is reduced.
It realizes efficient and accurate extraction of triple data from mineral literature, reduces the error rate of manual reading and recognition, improves information extraction efficiency, can handle multilingual and complex language literature, and reduces dependence on professionals.
Smart Images

Figure CN119760155B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a method and system for extracting knowledge from geological prospecting literature based on a large model. Background Art
[0002] With the progress of technology and the development of business, new concepts and terms will continuously emerge in newly generated literature. Traditional manual reading of literature to extract information may cause the inability to understand the literature due to problems such as knowledge reserves. Moreover, due to the large amount of knowledge in the field of prospecting and the different professional knowledge required for different perspectives, different literatures need to be assisted by experts in different sub - fields to be realized. Therefore, the existing work of extracting mineral literature is completed manually. At the same time, due to the large quantity and relatively long length of mineral - related literature, manual extraction requires reading the literature completely and identifying and extracting, which is very time - consuming and inefficient.
[0003] On the other hand, in the process of reading and recognition, triple data is usually introduced. Triple data is a basic data structure, consisting of three elements: subject (entity), predicate (relationship), and object (entity). This structure is used to describe the relationship between things, usually expressed as (subject, predicate, object). The subject and object are entities, which can be specific things or abstract concepts, and the predicate represents the relationship between the subject and the object. In the process of reading existing literature and identifying triples, some professional content cannot be understood by ordinary technical workers, which easily leads to difficulties in extracting some triples or not knowing how to extract them. It requires the assistance of professionals or experts in reading. If the support of professionals or experts cannot be obtained, the literature information cannot be extracted. Manual reading and triple information identification and extraction are limited by human energy and fatigue, and are more prone to errors than machines in the process of reading and recognition.
[0004] Mineral research and prospecting research have gone through multiple stages in the development process and will continue to move forward in the future. In this process, there are great differences in the language structure and writing logic of the literature written by researchers. Reading and extraction personnel can only recognize part of the language structure and writing logic, and there are difficulties in reading the language structure and writing logic used in ancient and rare literature and literature that has not yet emerged in the future.
[0005] Manual reading is very inefficient in the face of multi - language and complex - language literature environments, and may even be unable to complete the reading and extraction work. Summary of the Invention
[0006] In order to solve the above problems, the purpose of the present invention is to provide a method and system for extracting knowledge from geological prospecting literature based on a large model, which replaces manual work through the advantages of rapid reading and accurate recognition of the large model to realize the extraction of triples from mineral literature.
[0007] The present invention provides a method for extracting knowledge from geological prospecting literature based on a large model, and the method includes:
[0008] Establish a triple learning model, and train the triple learning model to identify triple data in geological prospecting literature and the forms in which the triple data exists in the geological prospecting literature;
[0009] Using the triple learning model as an engine, collect geological prospecting literature from various literature entrances, store the collected geological prospecting literature in a pre-set database, and perform classification and annotation work based on the geological prospecting literature;
[0010] Use the triple learning model to extract triple data from the geological prospecting literature stored in the database, and form leaf-shaped knowledge points based on the extracted triple data and store them in a knowledge base;
[0011] Organize the leaf-shaped knowledge points through business logic to form a tree-shaped or star-shaped knowledge structure;
[0012] Draw a knowledge graph based on the knowledge structure.
[0013] Optionally, the classification and annotation work based on the geological prospecting literature includes:
[0014] Annotate the naming names of geological entities;
[0015] Annotate the attribute information of geological entities;
[0016] Annotate the mutual relationships between geological entities and the semantic relationships between geological entities and attribute information;
[0017] Annotate the article viewpoints, numerical values and methods of the geological prospecting literature.
[0018] Optionally, the attribute information includes ore deposits, attribute characteristics, ore-controlling factors, and prospecting indicators;
[0019] The ore deposits include one or more of the following: environment / mining environment / geological environment / tectonic domain, geological phenomenon, method / geological method / geological technique / geological principle / step;
[0020] The attribute characteristics include one or more of the following: formation age / time / mining age, diagenetic age, crystallization age, grade / element grade, ore body shape, ore vein / ore body, ore section / ore zone, mining area / ore field / ore concentration area, region / location, rock type, ore deposit type, resource quantity / production;
[0021] The prospecting criteria include one or more of the following: geophysical anomaly / element anomaly, geochemical anomaly / element anomaly, remote sensing interpretation anomaly / element anomaly;
[0022] The ore - controlling factors include one or more of the following: geological body / rock mass / three major rock masses / host rock type, stratum, geological event / geological process, alteration type, mineralization stage / mineralization type, metallogenic process / mineralization process, structure.
[0023] Optionally, when extracting data using the triple - learning model, it includes: constructing a structured prompt using prompt engineering; the structured prompt includes: role, goal, {context, limitation}, {skill, tool}, {input rule, output rule}, output example.
[0024] The present invention also provides a geological prospecting literature knowledge extraction system based on a large - model. The system includes one or more processors and a non - transitory computer - readable storage medium storing program instructions. When the one or more processors execute the program instructions, the one or more processors are used to implement the geological prospecting literature knowledge extraction method according to any one of the above.
[0025] Optionally, the system includes: a mineral literature collection module, a literature information extraction research and development module, an information knowledge fine - processing module, and a literature knowledge product production module;
[0026] The mineral literature collection module is used to collect mineral literature through various channels;
[0027] The literature information extraction research and development module is used to generate a literature information extraction tool for information extraction based on the collected mineral literature;
[0028] The information knowledge fine - processing module is used to generate a prospecting prediction map;
[0029] The literature knowledge product production module is used to provide a prospecting special - topic service application.
[0030] Optionally, the mineral literature collection module collects mineral literature through commercial collection, self - construction, open access, and / or sharing and exchange.
[0031] The present invention also provides a computing device, which includes the geological prospecting literature knowledge extraction system according to any one of the above.
[0032] The present invention also provides a computer - readable storage medium, which is used to store program code, and the program code is used to execute the geological prospecting literature knowledge extraction method according to any one of the above.
[0033] According to the present invention, since the knowledge extraction method through the large model does not only rely on integrating entities, relationships, and attributes in structured data, semi-structured data, and unstructured data, but in addition to identifying triple data, it also identifies the forms in which triples exist, and these forms include text form, hypertext form, cross-text form, picture form, data form, cross-data form, etc. On the other hand, according to the present invention, not only are entity relationship attributes annotated, but also article viewpoints, numerical values, and methods are annotated, which can ensure the generation efficiency and accuracy of the knowledge graph.
[0034] In the present invention, an entity disambiguation mechanism is established. This mechanism will first learn a large number of template documents to establish a series of entity object clustering sets, and then form a semantic dictionary pointing to the object based on the context of the target document, aggregate this dictionary under the established object clustering, and when constructing triples, a secondary tracking strategy of first matching the entity object clustering and then matching the semantic dictionary is used to achieve entity disambiguation.
[0035] The large model-based geological prospecting literature knowledge extraction system of the present invention has accumulated a lot of professional knowledge through the learning of the large model, can accurately identify the information to be extracted in professional content, and reduces the dependence on external assistance. By using machine reading and machine extraction to replace manual work, the error rate caused by fatigue in the reading and extraction process is eliminated. Further, the large model is updated and upgraded through machine learning, so that the large model can always understand the language structure and writing logic of each document and accurately extract the required information. For this purpose, the present invention establishes a model iteration mechanism, which includes steps such as global new reading, prompt engineering, new structure recognition, and new prompt engineering iteration. Through global new reading, the language structure and prompt engineering in the new document are obtained, and similarity matching is performed with the existing model. When it is lower than the set threshold, the language structure extracted from reading the new document is set as a new structure and enters the structure library, and a secondary prompt engineering iteration model is established, where correcting the original prompt engineering is the first-level iteration, and updating the prompt engineering model library is the second-level iteration. Utilizing the advantages of the large model in a cross-language environment, without using translation, the large model directly reads multi-language and complex language documents to extract effective information. Through pre-training, the ability to identify special corpora is established, and through multi-sample drills, the disassembling and entity extraction of documents in a special corpus environment are formed. In the system, full-conventional corpus documents, special corpus environment documents, and mixed corpus documents can be configured, so that the large language model can call different rules for extraction.
[0036] In the present invention, in order to effectively detect errors in the knowledge graph data and thereby improve the data quality, the system uses a deep learning mechanism to continuously optimize the large language model for extracting triples, and establishes a conventional literature prompt engineering and a comparison literature prompt engineering. When the output result has a large difference from the expectation, the deep learning mechanism is started to disassemble the manually extracted triple samples step by step, and understand and learn step by step.
[0037] Iteratively learn the principle of entity extraction in long-distance context information, and use the existing samples and the new samples after manual error correction to train the model bidirectionally, so that the prompt engineering can intelligently judge the influence of long-distance context on triples and accurately perform grouping and matching.
[0038] Through the following detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more clear about the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0040] Figure 1 is a schematic flow chart of a method for extracting knowledge from geological prospecting literature based on a large model according to an embodiment of the present invention;
[0041] Figure 2 is a schematic diagram of a triple learning model according to an embodiment of the present invention;
[0042] Figure 3 is a schematic diagram of a system for extracting knowledge from geological prospecting literature based on a large model according to an embodiment of the present invention;
[0043] Figure 4 is a schematic diagram of a system for extracting knowledge from geological prospecting literature based on a large model according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the present invention and are not restrictive.
[0045] The embodiment of the present invention provides a method for extracting knowledge from geological prospecting literature based on a large model. As Figure 1 shown, the method for extracting knowledge from geological prospecting literature based on a large model according to the embodiment of the present invention may include the following steps S1 to S6.
[0046] S1. Establish a triple learning model and train the triple learning model to identify triple data in geological prospecting literature and the forms in which the triple data exists in the geological prospecting literature.
[0047] Triple data is a basic data structure consisting of three elements: a subject (entity), a predicate (relationship), and an object (entity). This structure is used to describe the relationships between things and is usually expressed as (subject, predicate, object). The subject and the object are entities, which can be specific things or abstract concepts, and the predicate represents the relationship between the subject and the object.
[0048] The triple learning model mentioned in this embodiment is a machine learning model. Using the triple learning model, triple data in geological prospecting literature and the forms in which the triple data exists in the geological prospecting literature can be identified. The triple data can include subject entities, relationships, and object entities, and the forms in which the triple data exists in the geological prospecting literature can include text form, hypertext form, cross - text form, picture form, data form, cross - data form, etc.
[0049] S2. Use the triple learning model as an engine to collect geological prospecting literature from various literature entrances; among them, the collection condition is that it contains the identified triple data related to prospecting.
[0050] S3. Store the collected geological prospecting literature into a pre - set database and perform classification and annotation work based on the geological prospecting literature.
[0051] S4. Use the triple learning model to extract triple data from the geological prospecting literature stored in the database and form leaf - shaped knowledge points based on the extracted triple data and store them in the knowledge base.
[0052] S5. Organize the leaf - shaped knowledge points to form a tree - shaped or star - shaped knowledge structure. The leaf - shaped structure is an important concept in mathematics, mainly used to describe the hierarchical structure on a manifold. For the formed leaf - shaped knowledge points, the leaf - shaped knowledge points can be organized into a tree - shaped or star - shaped knowledge structure through business logic. The tree - shaped structure usually represents a hierarchical relationship, while the star - shaped structure may represent the relationship between a central entity and multiple related entities. When constructing a tree - shaped knowledge structure, the root node can be determined first: select a core entity as the root node, usually the broadest and most abstract concept; then determine the branches: starting from the root node, add child nodes according to business logic, and each child node represents a more specific entity or concept; and then determine the leaf nodes: the bottom - most nodes, representing specific knowledge points or data.
[0053] When building a star structure, you can first determine the central node: select a core entity as the central node, which is usually the most important entity in the business logic; then determine the radial nodes: starting from the central node, connect all related entities, and each connection represents a knowledge point.
[0054] S6, drawing a knowledge graph based on the knowledge structure.
[0055] The method for extracting knowledge from geological prospecting literature based on a large model in an embodiment of the present invention replaces manual work by taking advantage of the large model's rapid reading and accurate recognition to achieve the extraction of triple data from mineral literature. Since the triple learning model has accumulated a lot of professional knowledge through learning, it can accurately identify the information that needs to be extracted from professional content, reducing the reliance on external assistance. By replacing manual work with machine reading and machine extraction, the error rate caused by fatigue in the reading and extraction process is eliminated. At the same time, the advantage of the large model's cross-language environment can be utilized. Without translation, the large model can directly read multilingual and complex language literature to extract effective information.
[0056] In an optional embodiment of the present invention, when drawing a knowledge graph for knowledge fusion, the ambiguity between entities, relationships, attributes and other referents and factual objects is eliminated, so that knowledge from different sources can be standardized and integrated. Therefore, an entity disambiguation mechanism can be established. First, a large number of template documents are learned to establish a series of entity object clustering sets, and then a semantic dictionary pointing to the object is formed based on the context of the target document. The dictionary is aggregated under the established object cluster. When constructing triples, a secondary tracking strategy of matching the entity object cluster first and then matching the semantic dictionary is used to achieve entity disambiguation.
[0057] Optionally, the knowledge graph is not generated in one go, but is a slowly accumulated process that requires continuous updating and iteration. Therefore, a model iteration mechanism can also be established, which includes steps such as global new reading, prompt word engineering, new structure identification, and new prompt word engineering iteration. The language structure and prompt word engineering in the new document are obtained through global new reading, and the similarity is matched with the existing model. When it is lower than the set threshold, the language structure extracted from reading the new document is set as the new structure and enters the structure library, and a secondary prompt word engineering iteration model is established, in which the correction of the original prompt word engineering is the first iteration, and the updating of the prompt word engineering model library is the second iteration.
[0058] Furthermore, a deep learning mechanism can be used to continuously optimize the triple extraction learning model, and a conventional literature prompt engineering and a comparison literature prompt engineering can be established. When the output result has a large difference from the expectation, the deep learning mechanism is activated to disassemble the steps of the manually extracted triple samples and understand and learn step by step. Iteratively learn the principles of entity extraction in long-distance context information, and use the existing samples and the new samples after manual error correction to train the model bidirectionally, so that the prompt engineering can intelligently judge the influence of long-distance context on triples and accurately perform grouping and matching, thereby effectively detecting errors in the knowledge graph data and improving the data quality.
[0059] In an alternative embodiment of the present invention, the triple information in the literature should be marked and standardized before extraction, that is, the classification and marking work based on the geological prospecting literature includes:
[0060] (1) Mark the naming names of geological entities;
[0061] (2) Mark the attribute information of geological entities;
[0062] (3) Mark the mutual relationships between geological entities and the semantic relationships between geological entities and attribute information;
[0063] (4) Mark the article viewpoints, numerical values, and methods of the geological prospecting literature.
[0064] Combined with Figure 2 it can be known that the attribute information includes ore deposits, attribute characteristics, ore-controlling factors, and prospecting criteria.
[0065] The ore deposits include one or more of the following: environment / mining environment / geological environment / tectonic domain, geological phenomenon, method / geological method / geological technique / geological principle / step.
[0066] The attribute characteristics include one or more of the following: formation age / time / mining age, diagenetic age, crystallization age, grade / element grade, ore body morphology, ore vein / ore body, ore section / ore zone, mining area / ore field / ore concentration area, area / location, rock type, ore deposit type, resource volume / production.
[0067] The prospecting criteria include one or more of the following: geophysical anomaly / element anomaly, geochemical anomaly / element anomaly, remote sensing interpretation anomaly / element anomaly.
[0068] The ore-controlling factors include one or more of the following: geological body / rock body / three major rock bodies / surrounding rock type, stratum, geological event / geological process, alteration type, mineralization stage / mineralization type, mineralization process, structure.
[0069] In an alternative embodiment of the present invention, when using a triple learning model for knowledge extraction, a structured prompt can be constructed using prompt engineering; the structured prompt includes: role, objective, {context, constraints}, {skills, tools}, {input rules, output rules}, and output examples. The construction example is as follows:
[0070] Role: You are a professional geological prospecting expert and gold depositologist, and I will input a literature to you.
[0071] Objective:
[0072] (1) Efficiently and accurately extract various entities closely related to gold prospecting from the paragraphs of the given literature, and clearly extract the specific relationships between the entities.
[0073] (2) Strictly draw the extracted results into an accurate and clear table according to the table header of "serial number, head entity type, head entity, relationship between entities, tail entity type, tail entity, relationship attribute, other attributes, original sentence, article number, remarks", and pay attention not to repeat the extraction of the same or similar triples from the same sentence.
[0074] Output logic:
[0075] (3) The extraction logic is: first separate each natural paragraph of the literature; then extract triples from each natural paragraph and output the results in batches.
[0076] (4) The extracted entities must comprehensively cover the key entities and their relationships mentioned in the document to construct a complete relationship network.
[0077] For the identified triple data, it can be filled into the table according to the table relationship. The table form is shown in the following table.
[0078] Table 1
[0079] Serial number Head entity type Head entity Relationship between entities Tail entity type Tail entity Relationship attribute Other attributes Original paragraph Article number
[0080] Among them, the head entity type can include 26 types such as region (administrative division, geological body, plateau, spray coating), mining area, ore belt, ore deposit, stratum, rock mass, etc.;
[0081] The relationship between entities can include 27 relationship types such as belong to, located in, developed in, experienced, related to, controlled by, shown by, predicted by, etc.
[0082] The tail entity type can include 26 entity types such as region (administrative division, geological body, plateau, basin), mining area, ore belt, ore deposit, ore vein, stratum, rock mass, etc.
[0083] The relationship attribute can include: fact, view, method.
[0084] Other attributes may include: people, time, location, and numerical values.
[0085] Finally, the reviewed content can be stored in a database, such as Table 2.
[0086] Table 2
[0087] Serial number Head entity type Head entity Relationship between entities Tail entity type Tail entity Relationship attribute Other attributes Original paragraph Article number 1 Region (geological body) Southern Qinling tectonic belt Formed in Geological event Indosinian - Yanshanian Qinling orogenic process Fact Location The Southern Qinling tectonic belt is a composite orogenic belt formed by the superposition of multiple - stage tectonic movements 42 2 Deposit Rich gold deposit Belongs to Rock mass Biotite monzogranite Fact Location The rich gold deposit in a certain county of a certain place is located in the Dongjiangkou rock mass of the Southern Qinling tectonic belt 42 3 Deposit Rich gold deposit Formed in Ore - forming age About 200Ma Fact Time The deposit was formed in the post - collision tectonic environment during the transition from collision to extension tectonic regime in the Indosinian - Yanshanian Qinling orogenic process 42 4 Deposit type Rich gold deposit Belongs to Ore - forming type Sulfide - quartz vein type Viewpoint The deposit types are sulfide - quartz vein type and fracture zone altered rock type 42 5 Geological phenomenon Occurrence state of gold Has Characteristics Included gold, fissure gold and intergranular gold Fact The occurrence states of gold are included gold, fissure gold and intergranular gold 42 6 Geology (geological body) Qinling tectonic belt Is Region (administrative division) A certain country Fact Location The Qinling tectonic belt is one of the main gold - metallogenic belts in a certain country 42 7 Deposit Rich gold deposit Belongs to Deposit scale Medium - sized Fact The deposit scale is medium - sized 42 8 Ore body Gold ore body Located in Structure Gentle - dipping fault Fact The ore body is mainly hosted in the gentle - dipping faults developed in the rock mass 42 9 Ore type Ore Includes Type Sulfide - quartz vein type and fracture zone altered rock type Viewpoint The ore types can be divided into sulfide - quartz vein type and fracture zone altered rock type 42 10 Geological body Dongjiangkou rock mass Formed in Geological event Post - collision tectonic environment Fact Location The Dongjiangkou rock mass was formed in a typical post - collision tectonic environment 42 11 Geological event Magmatic activity Formed in Geological event Indosinian period Fact Time In recent years, a certain mining development and trade company has discovered rich gold deposits in the Southern Qinling region 42 12 Deposit Rich gold deposit Formed in Geological body Dongjiangkou rock mass Fact Location The rich gold deposit is located in the Dongjiangkou rock mass 42 13 Geological body Dongjiangkou rock mass Includes Rock mass Gaoqiaojie, Dongjiangkou, Zhashui, Pingganchuan rock masses Fact Location The Dongjiangkou rock mass group includes four geologically unconnected rock masses to the east 42 14 Geological body Dongjiangkou rock mass Formed in Geological event Between 220 - 200Ma Fact Time The Dongjiangkou rock mass was formed after the subduction, collision, folding and extension of the Yangtze plate and the North China plate to form post - collision granite 42 15 Geological structure Fault Controls Geological phenomenon Spatial distribution characteristics of hypabyssal - ultra - hypabyssal intrusions Fact The fault controls the spatial distribution characteristics of hypabyssal - ultra - hypabyssal intrusions 42
[0088] An embodiment of the present invention also provides a large model-based geological prospecting literature knowledge extraction system for implementing the large model-based geological prospecting literature knowledge extraction method of the above embodiment.
[0089] As Figures 3 - 4 shown, the large model-based geological prospecting literature knowledge extraction system of the embodiment of the present invention includes: a mineral literature collection module, a literature information extraction research and development module, an information knowledge fine processing module, and a literature knowledge product production module.
[0090] Among them, the mineral literature collection module is used to collect mineral literature through various channels; the mineral literature collection module collects mineral literature through commercial collection, self-construction, open access, and / or sharing and exchange. The literature information extraction research and development module is used to generate literature information extraction tools for information extraction based on the collected mineral literature; the information knowledge fine processing module is used to generate prospecting prediction maps; the literature knowledge product production module is used to provide prospecting special topic service applications.
[0091] An embodiment of the present invention also provides a computing device, which includes the large model-based geological prospecting literature knowledge extraction system described in the above embodiment.
[0092] An embodiment of the present invention also provides a computer-readable storage medium, which is used to store program codes for executing the large model-based geological prospecting literature knowledge extraction method described in the above embodiment.
[0093] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A method for extracting knowledge from geological prospecting literature based on large models, characterized in that, The method includes: Establishing a triple learning model, training the triple learning model to identify triple data in geological prospecting literature and the forms in which the triple data exists in the geological prospecting literature; Using the triple learning model as an engine to collect geological prospecting literature from various literature entrances; Storing the collected geological prospecting literature in a pre-set database and performing classification and annotation work based on the geological prospecting literature; Using the triple learning model to extract triple data from the geological prospecting literature stored in the database and forming leaf-shaped knowledge points based on the extracted triple data and storing them in the knowledge base; Organizing the leaf-shaped knowledge points to form a tree-shaped or star-shaped knowledge structure; Drawing a knowledge graph based on the knowledge structure; The triple learning model has an entity disambiguation mechanism; the entity disambiguation mechanism first learns a large number of template literatures to establish a series of entity object clustering sets, then forms a semantic dictionary pointing to the object based on the context of the target literature, aggregates the dictionary under the established object clustering, and implements entity disambiguation by using a secondary tracking strategy that first matches the entity object clustering and then matches the semantic dictionary when constructing triples; The triple learning model has a model iteration mechanism; the model iteration mechanism includes steps of global new reading, prompt engineering, new structure recognition, and new prompt engineering iteration; through global new reading, the language structure and prompt engineering in the new literature are obtained, and similarity matching is performed with the existing model. When the similarity is lower than the set threshold, the language structure extracted from the new literature is set as a new structure and enters the structure library, and a secondary prompt engineering iteration model is established, where correcting the original prompt engineering is the first iteration and updating the prompt engineering model library is the second iteration; Using the deep learning mechanism to continuously optimize and extract the triple learning model, establishing a conventional literature prompt engineering and a comparison literature prompt engineering; when the output result has a large difference from the expectation, start the deep learning mechanism to disassemble the steps of the manually extracted triple samples and understand and learn step by step; iterate and learn the principle of entity extraction in long-distance context information, and use the existing samples and the new samples after manual correction to train the model bidirectionally, so that the prompt engineering can intelligently judge the influence of the long-distance context on the triples and accurately perform grouped matching.
2. The method according to claim 1, wherein The classification and annotation work based on the geological prospecting literature includes: Annotating the naming names of geological entities; Annotating the attribute information of geological entities; Annotating the mutual relationships between geological entities and the semantic relationships between geological entities and attribute information; Annotating the article viewpoints, numerical values, and methods of the geological prospecting literature.
3. The method according to claim 2, wherein The attribute information includes ore deposits, attribute characteristics, ore-controlling factors, and prospecting criteria; The ore deposits include one or more of the following: metallogenic environment, geological phenomenon, geological method; The attribute characteristics include one or more of the following: formation age, grade, ore body shape, ore vein, ore section, mining area, location, rock type, ore deposit type, resource quantity; The prospecting criteria include one or more of the following: geophysical anomaly, geochemical anomaly, remote sensing interpretation anomaly; The ore-controlling factors include one or more of the following: geological bodies, strata, geological processes, alteration types, mineralization stages, mineralization processes, and structures.
4. A geological prospecting literature knowledge extraction system based on a large model, characterized in that, The system includes one or more processors and a non-transitory computer-readable storage medium storing program instructions. When the one or more processors execute the program instructions, the one or more processors are configured to implement the method for extracting knowledge from geological prospecting literature based on a large model according to any one of claims 1-3.
5. The system according to claim 4, wherein The system includes: a mineral literature collection module, a literature information extraction R & D module, an information knowledge refinement module, and a literature knowledge product production module; The mineral literature collection module is configured to collect mineral literature through various channels; The literature information extraction R & D module is configured to generate a literature information extraction tool for information extraction based on the collected mineral literature; The information knowledge refinement module is configured to generate a prospecting prediction map; The literature knowledge product production module is configured to provide a prospecting special topic service application.
6. The system according to claim 5, wherein The mineral literature collection module collects mineral literature through commercial collection, self-construction, open access, and / or sharing and exchange.
7. A computing device, characterized in that, The computing device includes the system for extracting knowledge from geological prospecting literature based on a large model according to any one of claims 4-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store program code for executing the method for extracting knowledge from geological prospecting literature based on a large model according to any one of claims 1-3.
Citation Information
Patent Citations
Event extraction model training method, event extraction method and related equipment
CN115525776A
Power grid data search method and system based on core data identification and electronic equipment
CN116738979A
Knowledge graph construction method and device, electronic equipment and storage medium
CN117633245A
Intelligent algorithm knowledge extraction method and device based on pre-training language model
CN117874186A