Construction method and application of ore deposit prospecting prediction ontology architecture
By constructing an ontology architecture for mineral prospecting and prediction, combining large language models with keyword analysis, and creating a knowledge question-answering system and graphic teaching aids, the problem of fragmentation of geological prospecting teaching resources is solved, efficient teaching and learning effects are achieved, and the needs of mineral education in colleges and universities are met.
Patent Information
- Application Number
- CN202510739431.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are unable to effectively combine large language models to construct structured knowledge graphs suitable for the field of geological prospecting, resulting in fragmentation of teaching resources and inefficient reasoning of mineralization laws, making it difficult to meet the needs of mineral teaching in universities.
Construct an ontology architecture for mineral deposit exploration and prediction, determine the initial entity set through keyword co-occurrence analysis and word cloud statistics, combine large model prompt engineering for labeling and relationship extraction, create a knowledge question-answering system and knowledge graph teaching aids for systematic teaching deduction and student training.
It achieves the structured display of mineral prospecting and prediction knowledge and the logically rigorous knowledge support, improves teaching efficiency and students' learning effects, meets the iterative needs of AI4S teaching tasks, and assists exploration planning and data analysis.
Smart Images

Figure CN120670550A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prospecting and prediction, and in particular to a method for constructing a mineral deposit prospecting and prediction ontology architecture and its application. Background Art
[0002] With the sustainable development of the global economy and society and the improvement of people's living standards, the contradiction between the increasing demand for "industrial food" mineral resources and the decreasing number of "surface, shallow, and identified minerals" and the increasing difficulty of mineral exploration is becoming increasingly prominent. Mineral resources continue to develop towards hidden deposits and "deep" secondary exploration areas. Therefore, based on the background of mineral deposit exploration big data, the integration of artificial intelligence technology into the entire process of mineral exploration data processing, analysis, and application is of great significance for identifying mineral deposits and improving mineral exploration efficiency. This development situation places higher demands on geological and mineral exploration talents. The combination of teaching and artificial intelligence tools can cultivate more innovative talents.
[0003] At present, university geology courses, such as mineral resources teaching, face two technical barriers: first, traditional case teaching of mineral prospecting and prediction relies on isolated geological maps and static specimen libraries, which cannot present the temporal and spatial evolution laws of mineralization elements. As a result, it is difficult for undergraduate and graduate students in geology to establish a systematic understanding of "structure-magma-alteration" coordinated mineralization during their studies; second, the existing knowledge system is fragmented, and academic literature, exploration reports and teaching resources lack structured connections, resulting in low efficiency in reasoning about mineralization laws in case teaching.
[0004] Using large language models to assist in building a knowledge graph for prospecting and prediction can aid teaching exercises and reasoning. However, existing models lack accurate and professional background knowledge, making them unsuitable for teaching and research in the field of geological prospecting. Therefore, building a structured prospecting and prediction knowledge ontology, providing a systematic knowledge graph and knowledge question-and-answer system for mineral deposit teaching, will help students adapt more quickly to the current era of artificial intelligence and meet the iterative requirements of AI4S teaching tasks, which is of great significance. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for constructing a mineral deposit prospecting and prediction ontology architecture in response to the shortcomings of the existing technology. An extensible mineral deposit prospecting and prediction knowledge ontology architecture is constructed based on the mineral deposit prospecting and prediction knowledge field. By constructing structured entities and relationships, a knowledge question-and-answer system and a knowledge graph teaching aid are created for systematic teaching deduction, display and student training, and the teaching is highly practical.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for constructing a mineral deposit prospecting prediction ontology architecture comprises the following steps:
[0008] S1: Collect mineral prospecting data and pre-process the data;
[0009] S2: Analyze the preprocessed data to obtain the initial entity set A n , complete the determination of entity tags in the initial ontology architecture;
[0010] S3: Based on A n The representative data pre-processed in step S1 are labeled, and the final entity set and relationship labels are determined through the labeled data to form the mineral prospecting prediction ontology architecture.
[0011] Furthermore, in step S1, the collected prospecting data is organized into a PDF format and then preprocessed. The preprocessing is to clean redundant data, erroneous data, and duplicate data, and to normalize the data and convert it into a MarkDown format.
[0012] Furthermore, in step S2, the pre-processed data is analyzed by using CiteSpace software to obtain a co-occurrence network diagram of keywords, which is recorded as C b ={w b1 ,w b2 ,w b3 ,...}, where C b represents the associated word set of the b-th keyword, w b1 ,w b2 ,w b3 ,...represent the 1st, 2nd, 3rd... associated words of the bth keyword respectively;
[0013] After preprocessing, we remove stop words and perform word cloud statistics. The first 40 words in each file are recorded as P. m ={p m1 ,p m2 ,p m3 ,...,p m40}, where P m is the word frequency set of the mth file, p m1 ,p m2 ,p m3 ,...p m40 Represent the top 40 word frequencies of the mth file respectively;
[0014] C b and P m Request the Prompt model to summarize the research fields in the form of prompt requests, and determine 10-15 research fields. For each field, list 50 entities according to importance, and obtain the initial entity set A. n ={a n1 ,a n2,a n3 ,...a n50}, where A n is the initial entity set of the nth research field, a n1 ,a n2 ,a n3 ,...represent the 1st, 2nd, 3rd...entities in the nth research field respectively. A n is the first-level initial entity label, a n1 ,a n2 ,a n3 ,...a n50 It is the secondary initial entity label.
[0015] Furthermore, in step S3, the initial entity set A n ={a n1 ,a n2 ,a n3 ,...a n50 For entity labels, the relationship labels are set as measurement and comparison, composition attribute relationship, time order development relationship, process and mechanism, location and geographical relationship, conditional hypothesis and logical reasoning. Doccano is used to perform sequence labeling on representative data to obtain the labeled data and determine the final entity set A' n ={a n1 ,a n2 ,a n3 ,...,a nt}, A' n is the final entity set of the nth research field, t is the number of entities in the final entity set of the nth research field, a nt is the t-th entity in the n-th research field.
[0016] An application of a mineral deposit prospecting prediction ontology architecture is used to build a question-answering system teaching aid, which is constructed by the following steps:
[0017] S1: Extract entity relationships from multiple large models using prompts and few-sample prompts on the labeled data, and determine the large model with the strongest extraction capability.
[0018] S2: Through the final entity set A n‘ The best model determined by S1 was used to complete the construction of the mineral prospecting prediction question-answering system teaching aids through Graphrag.
[0019] Furthermore, in step S1, multiple models are tested through a few-sample prompting project. The data returned by LLM is json data. The entities and relationships are programmed in a BIO-like manner to perform weighted calculations of precision P, recall R, and F1. The model with the highest proportion of the highest values of P, R, and F1 is the best model.
[0020] Furthermore, in step S2, the best model determined in S1 is used to convert A' n Input it into {entity_types} of Graphrag Promptfine-tune, and use the Markdown document segmentation structure to split the pre-processed MarkDown data into text blocks and process it into .csv files, which specifically include publication time, document title, chapter title, and text blocks. Input it into Graphrag project data and perform multiple Auto Prompt Fine-tunes to complete the construction of the question-answering intelligent agent.
[0021] An application of a mineral deposit prospecting and prediction ontology architecture is used to construct a knowledge graph teaching aid, and the annotated data is imported into Neo4j to construct a knowledge graph as a teaching aid.
[0022] The beneficial effects of the present invention are:
[0023] 1. The present invention discloses a method for constructing an ontology architecture for mineral deposit prospecting and prediction. Based on data sources such as papers, monographs and web page data, the method utilizes keyword co-occurrence analysis and word cloud statistics technology, and combines large model prompt engineering to summarize and improve the knowledge structure of mineral deposit prospecting and prediction. On the basis of the knowledge field of mineral deposit prospecting and prediction, the present invention constructs an extensible knowledge ontology architecture for mineral deposit prospecting and prediction. By constructing structured entities and relationships, a knowledge question-and-answer system and a case teaching aid of knowledge graph are created, which are used for systematic teaching deduction, demonstration and student training, and have strong teaching practicality.
[0024] 2. In order to enhance the accuracy of the prospecting and prediction ontology architecture in case teaching, this application first determines the prospecting and prediction research field through keyword co-occurrence analysis, word cloud statistics technology and large model prompt engineering, creates an initial entity set, and then labels the representative data. During the labeling process, the initial ontology architecture is improved and optimized to obtain the final ontology architecture. This creation method not only covers more entity tags, but also makes the obtained ontology architecture more accurate, which can provide logically rigorous knowledge support for question-answering models and knowledge graph case teaching aids.
[0025] 3. The knowledge graph constructed based on the above precise ontology architecture is used for CRUD-like teaching. At the same time, existing NLP technology can be used to extract entity relationships, allowing students to quickly learn and get started with mineral prospecting knowledge graph research, meeting the teaching task iteration of AI4S.
[0026] On the basis of the final ontology architecture, Graphrag Prompt Fine-tune is used to complete the prompt engineering file and build a knowledge question-answering system, so that students can quickly query the prospecting and mineralization knowledge of specific mineral deposits, dialectically analyze the contradictions in detailed knowledge and form knowledge discoveries, and assist in completing exploration planning, data analysis, report writing and other tasks, which is highly practical for teaching and learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flow chart of the present invention;
[0028] Figure 2 This is the co-occurrence network diagram of the keywords "geochemistry" and "ore prospecting prediction";
[0029] Figure 3 This is the co-occurrence network diagram of the keyword "mineral prospecting prediction";
[0030] Figure 4 This is a partial ontology architecture diagram. DETAILED DESCRIPTION
[0031] The present invention will be further described below with reference to the accompanying drawings and examples.
[0032] Example 1
[0033] like Figure 1 As shown in the figure, a mineral deposit prospecting prediction ontology architecture is constructed by the following steps:
[0034] S1: Collect prospecting data of mineral deposits and pre-process the data;
[0035] The data sources include open source data, closed source data and monographs. Open source data can be crawled on the website using a crawler. Closed source data includes papers and geological survey reports downloaded from CNKI. Monographs are obtained by purchasing and then scanned and recognized by OCR.
[0036] The search subject for data acquisition was “ore deposit” and abstract “tin and polymetallics”, not journal name “newspaper” or “conference”, and the keywords “environmental pollution”, “environmental governance” and “mining and metallurgy technology” were excluded;
[0037] The collected prospecting data (valid literature) are organized into PDF format and then preprocessed. The preprocessing is to clean up redundant data, erroneous data, duplicate data, normalize the data and convert it into MarkDown format;
[0038] The data source of this embodiment is to collect journals and dissertations published from January 2000 to December 2023 from CNKI. The search conditions are: subject "specific mineral deposit" and abstract "tin" and not journal name "newspaper" and "conference". The keywords "environmental pollution", "environmental governance" and "mining and metallurgy process" are screened out. After preprocessing, 358 valid preprocessed data are finally obtained.
[0039] S2: Analyze the preprocessed data to obtain the initial entity set A n , complete the determination of entity tags in the initial ontology architecture;
[0040] The step S2 is specifically as follows: for the pre-processed data, the keyword co-occurrence analysis is performed using the CiteSpace software to obtain the keyword co-occurrence network diagram and record it as C b ={w b1 ,w b2 ,w b3 ,...}, where C b represents the associated word set of the b-th keyword, w b1 ,w b2 ,w b3 ,...represent the 1st, 2nd, 3rd... associated words of the bth keyword respectively; the co-occurrence network diagram with "geochemistry" and "mineral prospecting prediction" as keywords can be found in Figure 2 and Figure 3 ;
[0041] After preprocessing, we remove stop words and perform word cloud statistics. The first 40 words in each file are recorded as P. m ={p m1 ,p m2 ,p m3 ,...p m40}, where P m is the word frequency set of the mth file, p m1 ,p m2 ,p m3 ,...p m40 The frequencies of the first 40 words 1, 2, 3, etc. in the m-th document;
[0042] C b ={w b1 ,w b2 ,w b3 ,...} and P m Request the Prompt model to summarize the research fields in the form of prompt requests, and determine 10-15 research fields. For each field, please list 50 entities according to their importance, and obtain the initial entity set A. n ={a n1 ,an2 ,a n3 ,...a n50}, where A n is the initial entity set of the nth research field, a n1 ,a n2 ,a n3 ,...represent the 1st, 2nd, 3rd...entities in the nth research field, A n is the first-level initial entity label, a n1 ,a n2 ,a n3 ,...a n50 It is the secondary initial entity label.
[0043] The specific operation method of this embodiment is: The prompt request is: As a geological prospecting expert, please calculate the co-occurrence statistics of the following literature keywords C b ={w b1 ,w b2 ,w b3 ,...} and word cloud statistics P m =
[0044] p m1 ,p m2 ,p m3 ,...p m40}, categorize the main research areas for building knowledge graphs and give the main 15 research areas. Please list 50 entities in each area according to their importance;
[0045] The following are some of the results returned by Prompt:
[0046] A4 is "Rock-magma related", A4 = {"rock", "rock mass", "magma", "magmatic activity", "rock type", "rock structure", "rock microstructural characteristics" ....}
[0047] A5 is "Area and Location", A5 = {"Rock Formation", "Stratum", "Azimuth", "Direction", "Mineral Cluster", "Mineral Location"...};
[0048] S3: Based on A n Label the representative data after preprocessing in step S1, determine the final entity set and relationship labels through the labeled data, and form the prospecting prediction ontology architecture;
[0049] Take the initial entity set A n ={a n1 ,a n2 ,a n3 ,...a n50} is the entity label, and the relationship label is set to consider the six relationships of measurement and comparison (including: about, greater than, less than, weighted, average), composition attribute relationship (including and, and, or, minority, majority, etc.), time order development relationship (including earlier than, later than, continuous, before, after, etc.), process and mechanism (including transition, invasion, symbiosis, development, etc.), location and geographical relationship (including close to, along, far away, distance, surrounded, concentrated on, etc.), conditional hypothesis and logical reasoning (including explanation, belief, visible / identification, cause, control, prediction, speculation, etc.). Doccano is used to perform sequence annotation on representative data. The representative data refers to: for the monographs in the past five years, select the ones covering A n At most 1-2 monographs. If there are no suitable monographs, select master's and doctoral dissertations published in the past three years, covering A n The maximum number of papers is 50-100. This example selects master's and doctoral dissertations published in the past three years, covering A n A maximum of 100 papers.
[0050] The specific labeling method is: for data that cannot be labeled with entity tags, entity tags can be added. For data with entity and relationship labeling frequencies less than 5, they are merged with existing semantically similar entity or relationship labels (entity replacement, synonym replacement, sentence structure transformation, etc.). In this embodiment, the data with labeling frequencies less than 5 are merged to obtain the final entity set A' n ={a n1 ,a n2 ,a n3 ,...,a nt} and relation label, A' n is the final entity set of the nth research field, t is the number of entities in the final entity set of the nth research field, a nt It is the t-th entity in the n-th research field, forming the prospecting prediction ontology architecture. The entity set (entity label) and relationship label summary table of the prospecting prediction ontology architecture in this embodiment are shown in Table 1 and Table 2, and some ontology architecture diagrams are shown in Figure 4 .
[0051] Table 1 Summary of prospecting prediction entity labels
[0052]
[0053]
[0054] Table 2 Prospecting prediction relationship labels
[0055]
[0056]
[0057] Among them, the labeled data can be verified for labeling consistency using the Fleiss' Kappa formula, and for logical consistency using the Consistency formula. If both labeling consistency and logical consistency are greater than 0.8, the labeling is valid.
[0058] Example 2
[0059] An application of the mineral deposit prospecting prediction ontology architecture is used to build a case teaching aid for the question-answering system. It is constructed by the following steps:
[0060] S1: Extract entity relationships from multiple large models using prompts and few-sample prompts on the labeled data, and determine the large model with the strongest extraction capability.
[0061] The specific operation steps are: test the Kimi, Deepseek, GLM, and Gpt-4o models through the few-shot prompting project.
[0062] In this embodiment, the entity tags are the entity tags in Embodiment 1; and the relationship tags are all the relationship tags in Embodiment 1.
[0063] 1) Entity extraction prompt example:
[0064] As an expert in tin prospecting, please complete the following tasks.
[0065] Main task: Please extract entities based on corpus data.
[0066] The following are some examples of entity tags:
[0067] 1. Mineral deposit knowledge ontology: document titles, book chapter titles, such as "Chapter 3: Deep mineralization geological background of the Gejiu tin-copper polymetallic deposit";
[0068] 2. Subtitle: first-level, second-level, and third-level titles, such as "3.2 Extraction of Gravity-Induced Mineral Anomaly Information";
[0069] 3. Mineral deposit: Content ending with "mineral deposit", such as "Gejiu tin-copper polymetallic deposit", "Kafang deposit", "tin-copper deposit";
[0070] 4. Mineral deposit category: semantically clarify the mineral deposit category, such as "weathering deposit", "endogenous deposit", "hydrothermal deposit", "skarn deposit";
[0071] 5. Mineral deposit lineage, mineral deposit series: Entities ending with mineral deposit lineage, mineral deposit series, such as "rare, nonferrous and precious metal mineralization series";
[0072] 6. Region, mining area: mining area, or the content ending with "domain" or "mining area" such as "Gejiu mining area", "metamorphic mudstone source area", "syn-collision granite area";
[0073] 7. Mineralization stages and periods: basic knowledge of mineral deposits, such as the "skarn stage," "sulfide stage," and "high-temperature hydrothermal mineralization stage."
[0074] 8. Characteristics of mineralization conditions: mineralization characteristics, mineralization conditions;
[0075] 10. Ore body: such as "skarn ore body", "blind ore body", "deep ore body", etc.
[0076] 11. Ore body location: The specific location of the ore body, such as "middle of the ore body", "head of the ore body", "tail of the ore body (960m elevation)", etc.
[0077] 12. Ore body morphology: such as "layered ore body", "vein ore body", "impregnation ore body", etc.
[0078] 13. Ore body distribution: There are ore body hints in the previous text, such as "multi-layer distribution" and "right-leaning type";
[0079] Please mark the text location of the entity tag in the specific text and extract it into jsonal format as shown below:
[0080]
[0081] See the following few-sample tip legend:
[0082] Sample 1 Input:
[0083]
[0084] Sample 1 output:
[0085]
[0086] Sample 2 input:
[0087]
[0088] Sample 2 output:
[0089]
[0090] Notice:
[0091]
[0092] 2) Relationship extraction prompt example: Partial relationship label:
[0093] 1. Contains: Attribute (level) Function (explanation) "Deposit instance" contains "Ore type" etc., which means "includes", "contains", etc.;
[0094] 2. "Have": "Have" mainly marks labels such as "chart" and "formula".
[0095] 3. To: assign values, attributes, and characteristics;
[0096] 4. Use: what methods or means;
[0097] 5. Possessing: specific "characteristics", "anomalies", trends, data distribution, degree, sequence", etc.;
[0098] 6. Cause: The cause is the starting tag header, followed by the cause, which is connected to the specific entity;
[0099] 7. Formation: emphasizes the cause of formation, similar to genesis;
[0100] 8. Lead to: emphasizes the cause of formation, similar to cause;
[0101] 9. Summary: "Personnel" summary;
[0102] 10. Speculation or prediction: "guess" based on characteristics;
[0103] 11. Possible: The text is marked with “possible”;
[0104] 12. According to: According to the viewpoint, situation, experimental conditions, etc., replace the "combination", "observation", etc. in the text;
[0105] 13. Think: The author believes;
[0106] 14. Explain: show, elaborate;
[0107] 15. Impact: The article contains “impact-related statements”;
[0108] 16. The results show: emphasize the results show
[0109] 17. portrayed, reflected: "portrayed" and "reflected";
[0110] 18. Derived from: “derived from”, “source”, etc. in the text;
[0111] 19. For “targeted”, see if there is any annotation of geological bodies in the preceding and following text;
[0112] 20. Mainly: The text emphasizes "mainly" and "main type";
[0113] 21. Important: The text emphasizes "important";
[0114] 22. Secondary / Second: Second in importance;
[0115] 23. Part: “part” in the text;
[0116] 24. Account: numerical proportion "account";
[0117] 25. Majority: “majority”, “most”, “many” in the text;
[0118] 26. Minority: The words "minority" and "small amount" in the text emphasize the small amount of data, or indicate that the situation is rare. For example, "sometimes" can also be replaced by the label "minority".
[0119] Output format:
[0120]
[0121] See the following diagram for the implementation of the few-shot hint:
[0122] Sample 1 input
[0123]
[0124] Sample 1 output:
[0125]
[0126] Notice:
[0127]
[0128] In the above extraction prompt project, the data returned by LLM is json data. Then, through programming code, entities and relationships are weightedly calculated using a BIO-like method to calculate the precision rate P, recall rate R, and F1;
[0129] The P of a single entity or relationship class is calculated as follows:
[0130]
[0131] The class R of a single entity or relation is calculated as follows:
[0132]
[0133] In the above formula, TP (True Positive) is the number of correctly predicted positive examples, FP (False Positive) is the number of incorrectly predicted positive examples, and FN (False Negative) is the number of true positive examples that were not predicted;
[0134] The F1 of a single entity or relationship class is calculated as follows:
[0135]
[0136] The weighted F1 is calculated as follows:
[0137]
[0138] Among them F1 i is the F1 score of the i-th category (calculated by the precision and recall of the category; w i is the weight of the i-th category (the number of samples in this category), and n is the total number of categories; the weighted P, R and weighted F1 are calculated in the same way.
[0139] The results of the precision rate P, recall rate R, and F1 extracted by the large model of this embodiment are shown in Table 3. As can be seen from Table 3, the maximum values of P, R, and F1 in Deepseek account for the largest proportion, so it has the best effect in extracting multiple entity tags;
[0140] Table 3 Extraction effects of kimi, deepseek, GLM, and Gpt-40 models
[0141]
[0142] S2: Through the final entity set A n‘ The best model determined by S1 is Deepseek, and the construction of the mineral prospecting prediction question-answering system is completed through Graphrag.
[0143] Using the best model (deepseek) in S1, first A n‘ Input it into {entity_types} of Graphrag Prompt fine-tune, and use the Markdown document segmentation structure to split the pre-processed MarkDown data into text blocks and process it into .csv files, specifically including publication time, document title, chapter title, and text blocks. Input it into Graphrag project data and perform multiple Auto Prompt Fine-tunes to complete the construction of the question-answering intelligent agent.
[0144] The main functions of the constructed intelligent question-answering system include but are not limited to:
[0145] 1. Combine existing metallogenic knowledge to conduct reasoning and complete the query of metallogenic prospecting knowledge such as metallogenic geological bodies, metallogenic models, metallogenic causes, mineral combinations, and metallogenic characteristics;
[0146] 2. Complete dialectical analysis and knowledge discovery of contradictions in the mineralization causes, mineralization models, mineralization characteristics, etc. in the existing mineralization models.
[0147] 3. Through inquiries and inquiries in cross-disciplinary fields such as geophysics, geochemistry, and remote sensing geology, it can assist mineral exploration researchers in completing tasks such as exploration planning, data analysis, and report writing.
[0148] Example 3
[0149] An application of a mineral deposit prospecting and prediction ontology architecture is used to construct a knowledge graph teaching aid, and the annotated data is imported into Neo4j to construct a knowledge graph as a case teaching aid.
[0150] For teaching purposes, for example, by using Cypher to query a specific entity (e.g., "Malage Deposit" (a type of Gejiu deposit)) with a specific relationship (e.g., "ore formation"), one can quickly find the corresponding entity-relationship diagram. The knowledge graph constructed here is used for both classroom demonstrations and after-class assignments. These assignments involve using the provided partially annotated data, having students complete CRUD tasks such as querying the knowledge graph by importing it into Neo4j, creating nodes, and outputting the specific results as a graph to complete their knowledge of mineralization and prospecting for a specific mineral deposit.
[0151] In addition, the labeled data can be divided into training sets, validation sets, and test sets to conduct NLP entity relationship extraction experiments. RoBERTa-BiLSTM-CRF is used for entity extraction, and RoBERTa-BiLSTM is used for relationship extraction.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not limiting. Other modifications or equivalent substitutions made to the technical solution of the present invention by ordinary technicians in this field should be included in the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for constructing a mineral deposit prospecting prediction ontology architecture, characterized in that: The following steps are involved: S1: Collect mineral prospecting data and pre-process the data; S2: Analyze the preprocessed data to obtain the initial entity set A n , complete the determination of entity tags in the initial ontology architecture; S3: Based on A n The representative data pre-processed in step S1 are labeled, and the final entity set and relationship labels are determined through the labeled data to form the mineral prospecting prediction ontology architecture.
2. The method for constructing a mineral deposit prospecting prediction ontology architecture according to claim 1, characterized in that: In the step S1, the collected prospecting data is organized into a PDF format and then preprocessed. The preprocessing is to clean the redundant data, erroneous data, and duplicate data, normalize the data, and convert it into a MarkDown format.
3. The method for constructing a mineral deposit prospecting prediction ontology architecture according to claim 1, characterized in that: In step S2, the pre-processed data is analyzed by using CiteSpace software to obtain a keyword co-occurrence network diagram and record it as C b ={w b1 ,w b2 ,w b3 ,...}, where C b represents the associated word set of the b-th keyword, w b1 ,w b2 ,w b3 ,...represent the associated words of the bth keyword respectively; After preprocessing, we remove stop words and perform word cloud statistics. The first 40 words in each file are recorded as P. m ={p m1 ,p m2 ,p m3 ,...,p m40 }, where P m is the word frequency set of the mth file, p m1 ,p m2 ,p m3 ,...p m40 Represent the top 40 word frequencies of the mth file respectively; C b and P m Request the Prompt model to summarize the research fields in the form of prompt requests, and determine 10-15 research fields. For each field, list 50 entities according to importance, and obtain the initial entity set A. n ={a n1 ,a n2 ,a n3 ,...a n50 }, where A n is the initial entity set of the nth research field, a n1 ,a n2 ,a n3 ,...represent the entities in the nth research field, A n is the first-level initial entity label, a n1 ,a n2 ,a n3 ,...a n50 It is the secondary initial entity label.
4. The method for constructing a mineral deposit prospecting prediction ontology architecture according to claim 1, characterized in that: In step S3, the initial entity set A n ={a n1 ,a n2 ,a n3 ,...a n50 } is the entity label, and the relationship labels are set as measurement and comparison, composition attribute relationship, time order development relationship, process and mechanism, location and geographical relationship, conditional hypothesis and logical reasoning. Doccano is used to perform sequence annotation on representative data to obtain the annotated data and determine the final entity set A' n ={a n1 ,a n2 ,a n3 ,...,a nt }, where t is the number of entities in the final entity set of the nth research field.
5. An application of the mineral deposit prospecting prediction ontology architecture according to claim 1, characterized in that: The teaching aid for building a question-answering system is constructed by the following steps: S1: Extract entity relationships from multiple large models using prompts and few-sample prompts on the labeled data, and determine the large model with the strongest extraction capability. S2: Through the final entity set A n‘ The best model determined by S1 was used to complete the construction of the mineral prospecting prediction question-answering system teaching aids through Graphrag.
6. The application according to claim 5, characterized in that: In step S1, multiple models are tested through the few-sample prompting project. The data returned by LLM is json data. The entities and relationships are programmed in a BIO-like manner to perform weighted calculation of precision P, recall R, and F1. The model with the highest proportion of the highest values of P, R, and F1 is the best model.
7. The use according to claim 5, characterized in that: In step S2, the best model determined in step S1 is used to convert A' n Input it into {entity_types} of Graphrag Prompt fine-tune, and use the Markdown document segmentation structure to split the pre-processed MarkDown data into text blocks and process it into .csv files, specifically including publication time, document title, chapter title, and text blocks. Input it into Graphrag project data and perform multiple Auto Prompt Fine-tunes to complete the construction of the question-answering intelligent agent.
8. An application of the mineral deposit prospecting prediction ontology architecture according to claim 1, characterized in that: Used to construct a knowledge graph teaching aid, the annotated data is imported into Neo4j to construct a knowledge graph as a teaching aid.