Geological and mineral large language model construction method and geological and mineral industry question and answer method

By constructing a mixed geological corpus and fine-tuning and training the large language model using the deposit knowledge map, the existing model is solved the problem of insufficient professionalism in the geological and mineral field, and higher professionalism and reasoning ability are achieved, and more accurate answers are generated.

CN120179768APending Publication Date: 2025-06-20CHINA UNIV OF GEOSCIENCES (WUHAN) +2

Patent Information

Application Number
CN202510097741.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing general-purpose large-language model has problems such as insufficient professionalism and difficulty in understanding complex knowledge when dealing with professional problems in the geological and mineral industry, resulting in the lack of domain-specific professional knowledge and accuracy when providing answers.

Method used

By cleaning and pre-processing the literature data in the field of geological and minerals, a mixed geological corpus is constructed, and the basic large language model is fine-tuned and trained using the mineral deposit knowledge map to obtain the large language model of geological and minerals.

Benefits of technology

It significantly improves the model's intelligent Q&A ability in the field of geology and minerals, so that it can understand professional terms and complex relationships more comprehensively and accurately, thereby generating concise, professional and accurate answers, and improving the professionalism and reasoning ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179768A_ABST
    Figure CN120179768A_ABST
Patent Text Reader

Abstract

The invention discloses a geological mineral large language model construction method and a geological mineral industry question and answer method, and relates to the technical field of geological mineral large language models.The geological mineral large language model construction method mainly comprises the steps that data cleaning and preprocessing are conducted on geological mineral field literature data to obtain a geological mineral field special data set; obtaining a mixed geological corpus according to the special data set for the geological mineral product field and the mineral deposit knowledge graph; and performing fine tuning training on the basic large language model by using the mixed geological corpus to obtain a geological mineral large language model. By implementing the geological mineral product large language model construction method and the geological and mineral industry question and answer method provided by the invention, the professionality and accuracy of geological and mineral industry question and answer performed by the geological mineral product large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models in the geological and mining industries, and more specifically, to a method for constructing a large language model for geological minerals and a question-answering method for the geological and mining industries. Background Art

[0002] In the field of geological and mineral exploration, traditional research and exploration usually rely on the practical experience and professional knowledge of experts and require a large amount of time. Especially in the process of ore deposit exploration and resource assessment, the complexity of data analysis and the difficulty of knowledge acquisition increase the challenges of geological exploration. These tasks require the accumulation of professional knowledge and experience, but young geologists often face greater challenges due to lack of experience and knowledge. In recent years, the rapid development of large language model technology has provided new ideas for solving this problem. It has significant advantages in text understanding, knowledge learning, generating intelligent answers, etc., and can provide efficient and reliable intelligent support for geological and mineral exploration work. However, existing general large models still have some limitations when dealing with professional problems in the geological and mining industries. Since the training data relies on a large amount of open encyclopedia text data, while the core data in the field of geological and mineral exploration is often private or restricted data, these important professional data are not fully utilized, resulting in the inability to guarantee the professionalism and accuracy of the model. Therefore, the model usually faces problems such as lack of domain-specific professional knowledge, insufficient understanding, and information hallucination. To address these problems, it is necessary to propose a method for constructing a large language model that can build and provide systematic, structured, and accurate geological and mineral knowledge, effectively reducing the manual burden in traditional geological and mineral exploration work, and at the same time providing strong intelligent support for geologists in high-difficulty tasks such as field exploration.

[0003] Chinese Patent "CN118260388A A Question-Answering Method for the Geological and Mining Industries Based on a Large Language Model and a Knowledge Base" proposes a question-answering system that combines a large language model and a knowledge base. This method first constructs a text knowledge corpus for the geological and mining industries and uses a voice chat program to generate fine-tuning samples; performs unsupervised autoregressive training and fine-tuning on the ChatGLM-6B model through these fine-tuning samples; then creates a vector knowledge base through the text knowledge corpus and conducts question searching based on this. The model can comprehensively understand user questions and generate accurate answers. This method does not require manual data annotation, saves costs, and improves the accuracy and diversity of answers by combining retrieval-based and generative question answering. However, the construction of the system still relies on the text knowledge corpus and the vector knowledge base, which requires a large amount of time and computing resources, and the model may still have certain limitations in answering questions in professional fields.

[0004] The Chinese patent "CN118193708A A Mineral Knowledge Q&A Method and System Based on Large Language Model" proposes a large language model mineral Q&A system that combines retrieval-augmented generation. First, it uses web crawler technology to obtain mineral data from multiple platforms and clean it, constructs a mineral knowledge base, and obtains the mineral knowledge documents that ultimately assist the generation model in generating answers based on the multi-feature fusion and fine-ranking algorithm of XGBoost. Then, it uses the LoRA technology to fine-tune the large language model and designs prompt words to guide it to generate content in the mineral field. This method can provide flexible answers to mineralogy questions, especially performing well in scenarios of multi-sentence answers and multi-turn conversations. However, this method still faces problems such as insufficient expansion of professional vocabulary and information overload in multi-turn conversations, which may affect the efficiency and accuracy of the system. Summary of the Invention

[0005] The object of the present invention is to provide a method for constructing a large language model for geological minerals and a Q&A method for the geological and mining industries, which can improve the professionalism and accuracy of the large language model for geological minerals in answering questions in the geological and mining industries.

[0006] The present invention provides a method for constructing a large language model for geological minerals, including the following steps: S1: Clean and preprocess the literature materials in the field of geological minerals to obtain a special data set for the field of geological minerals; S2: Obtain a mixed geological corpus according to the special data set for the field of geological minerals and the deposit knowledge graph; S3: Use the mixed geological corpus to fine-tune and train the basic large language model to obtain a large language model for geological minerals.

[0007] Further, step S1 specifically includes: S11: Obtain the literature materials in the field of geological minerals, clean and preprocess the literature materials in the field of geological minerals to obtain the cleaned text data; S12: Obtain a special data set for the field of geological minerals according to the cleaned text data.

[0008] Further, the above-mentioned literature materials in the field of geological minerals include textbooks on ore depositology, geological exploration reports, monographs on mineral exploration, academic journal papers, and industry research reports.

[0009] Further, the above-mentioned cleaning and preprocessing of the literature materials in the field of geological minerals include removing low-quality data and graphics, correcting recognition errors, filling in missing data, removing duplicate information and noise data, and unifying the format.

[0010] Further, step S2 specifically includes: S21: Extract the head entity, tail entity, and relationship from the deposit knowledge graph, and construct a triple set according to the head entity, tail entity, and relationship; S22: Construct a Q&A template and check whether the triple set matches the Q&A template: If not, correct the triple set; if it matches, perform entity word substitution to obtain Q&A pairs. S23: Convert the format of the Q&A pairs to obtain Q&A data in training format. Based on the Q&A data in training format and the special dataset in the field of geology and mineral resources, obtain a mixed geological corpus.

[0011] Further, step S3 specifically includes: Select Baichuan-2 as the basic large language model, and perform fine-tuning training on the basic large language model according to the settings of a learning rate of 0.00002 and a training epoch number of 2 to obtain a large language model for geology and mineral resources.

[0012] Further, the above method for constructing a large language model for geology and mineral resources also includes: Set multiple questions in multiple sub-fields of the geology and mineral resources field, and conduct comprehensive reasoning evaluation on the large language model for geology and mineral resources; Compare the answers of the large language model for geology and mineral resources with the answers of the Baichuan-2 basic model; Compare the large language model for geology and mineral resources with general large language models to evaluate its reasoning speed, logical clarity, answer accuracy, and integrity when dealing with professional geological questions.

[0013] The present invention also provides a Q&A method for the geology and mining industry, including the following steps: Use the above method for constructing a large language model for geology and mineral resources to obtain a large language model for geology and mineral resources; Input the received question to be answered into the large language model for geology and mineral resources to obtain the answer to the question to be answered.

[0014] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above Q&A method for the geology and mining industry are implemented.

[0015] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above Q&A method for the geology and mining industry are implemented.

[0016] Implementing the method for constructing a large language model for geology and mineral resources and the Q&A method for the geology and mining industry provided by the present invention has the following beneficial effects: The present invention significantly enhances the intelligent question - answering ability by embedding the professional knowledge graph in the geological and mineral field into the large - language model. The present invention uses the structured knowledge graph to enhance the model, enabling it to more comprehensively and accurately understand the professional terms and complex relationships in the geological and mineral field, thereby generating concise, professional, and accurate answers. Compared with the traditional intelligent question - answering system in the geological and mineral field, the present invention solves the problems of insufficient professionalism of the original model and difficulty in understanding complex knowledge through the embedding of the knowledge graph, greatly improving the professionalism and reasoning ability of the model. In addition, on the basis of constructing a large - scale, hybrid special corpus in the geological and mineral field, the present invention successfully balances the contradiction between the limitation of computing resources and the data scale, enabling this method to still maintain a high training efficiency and reasoning performance under the existing hardware resource conditions, and reducing the computing cost. By further optimizing hyperparameters (such as learning rate, batch size, number of training epochs, etc.), it is ensured that the model can operate efficiently and stably, improving the overall performance. Generally speaking, the present invention not only enhances the professionalism of the model in the geological and mineral field through the embedding of the knowledge graph, but also enables the model to train and reason efficiently and economically by optimizing the use of computing resources, thereby providing high - quality intelligent question - answering services. This technical solution provides strong intelligent support for geological resource exploration and mineral development, provides a prior reference for the efficient utilization of exploration data, and promotes the intelligent process in the geological and mineral field.

[0017] Compared with the prior art, the technical solution proposed by the present invention significantly reduces the training and reasoning costs while ensuring the accuracy and efficiency of the model. The trained model performs well in a series of geological tasks such as mineral resource exploration, analysis of mineral types and distribution, metallogenic regularity, and background geological knowledge. It not only demonstrates a profound understanding of geological knowledge but also shows excellent ability to solve practical geological problems in actual mineral exploration scenarios. The embedding of the knowledge graph significantly improves the professionalism and conciseness of the model's answers. Compared with traditional general large models, the large - language model in the geological and mineral field can provide more professional and accurate answers, providing stronger support for geological survey and mineral exploration work. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings: Figure 1 is the flow chart of the method for constructing the large - language model in the geological and mineral field provided by the present invention; Figure 2 is the technical route flow chart of the method for constructing the large - language model in the geological and mineral field provided by the present invention; Figure 3 is the flow chart of constructing question - answer pairs for the knowledge graph provided by the present invention; Figure 4 is the model training performance under different parameter settings provided by the present invention; Figure 5 It is a structural block diagram of the computer device provided by the present invention. Specific Embodiments

[0019] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0020] Figure 1 A schematic diagram of the method for constructing a geological and mineral language model according to this embodiment is shown. In this embodiment, the method for constructing a geological and mineral language model includes the following steps: S1: Clean and preprocess the literature data in the geological and mineral field to obtain a special dataset for the geological and mineral field; In an exemplary embodiment, step S1 specifically includes: S11: Obtain the literature data in the geological and mineral field, clean and preprocess the literature data in the geological and mineral field to obtain the cleaned text data; S12: Obtain a special dataset for the geological and mineral field according to the cleaned text data; In an exemplary embodiment, the literature data in the geological and mineral field includes textbooks on ore deposit geology, geological exploration reports, monographs on mineral exploration, academic journal papers, and industry research reports; In an exemplary embodiment, cleaning and preprocessing the literature data in the geological and mineral field includes removing low-quality data and graphics, correcting recognition errors, filling in missing data, removing duplicate information and noise data, and unifying the format; As an exemplary embodiment, in step S, collect the public and private restricted literature data in the geological and mineral field, including textbooks on ore deposit geology, geological exploration reports, monographs on mineral exploration, academic journal papers, and industry research reports, etc.; clean the collected text data, including removing low-quality data and graphics, correcting recognition errors, filling in missing data, and at the same time processing the duplicate information, inconsistent format, and noise data existing in the original data, and finally obtain a high-quality geological pure text corpus as the basis for subsequent model training; based on the cleaned text data, construct a special dataset for the geological and mineral field, covering geological knowledge in multiple aspects such as ore deposit types, mineral resource exploration, and mineralization mechanisms; S2: Obtain a mixed geological corpus according to the special dataset for the geological and mineral field and the ore deposit knowledge graph; In an exemplary embodiment, step S2 specifically includes: S21: Extract the head entity, tail entity, and relationship from the ore deposit knowledge graph, and construct a triple set according to the head entity, tail entity, and relationship; S22: Construct a Q&A template, and check whether the triple set matches the Q&A template: if not, correct the triple set; if it matches, perform entity word replacement to obtain Q&A pairs; S23: Convert the format of the Q&A pairs to obtain Q&A data in training format, and obtain a mixed geological corpus based on the Q&A data in training format and the special dataset in the field of geology and mineral resources; As an exemplary embodiment, in step S2, using a deposit knowledge graph containing multiple entity nodes and multiple relationships, based on geological expert knowledge and the actual application requirements in the mineral resources field, design multiple Q&A templates, and generate a large number of Q&A pairs through template matching for training the model; as Figure 3 shown is the specific process of constructing Q&A pairs through the knowledge graph; first, extract the head entity, tail entity, and relationship from the knowledge graph to construct a triple set; then check whether these triple sets match the predefined templates; if they match, replace the entity words with the placeholders in the template, and then fill in the question and answer templates to generate Q&A pairs; if the triple set does not match the template, correct it and try to match again; this process will continue to iterate until Q&A pairs that can be used in the geological corpus are successfully generated; It should be noted that in these templates, the domain and scope types of the semantic relationships in the ontology determine the types of Q&A templates, while the categories of geological entities and their corresponding semantic relationships determine their corresponding Q&A templates; in this way, geological entities can replace the vocabulary in the template to generate new Q&A pairs; The Q&A pairs generated in batches using the above method are converted into a data format suitable for training through a Python script and combined with the text dataset to form a mixed geological corpus; S3: Use the mixed geological corpus to fine-tune and train the basic large language model to obtain a large language model for geology and mineral resources; In an exemplary embodiment, step S3 specifically includes: selecting Baichuan-2 as the basic large language model, and performing fine-tuning training on the basic large language model according to the settings of a learning rate of 0.00002 and a training epoch of 2 to obtain a large language model for geology and mineral resources; In an exemplary embodiment, the method for constructing a large language model for geology and mineral resources further includes: setting multiple questions in multiple sub-domains of the geology and mineral resources field, and comprehensively reasoning and evaluating the large language model for geology and mineral resources; comparing the answers of the large language model for geology and mineral resources with the answers of the Baichuan-2 basic model; comparing the large language model for geology and mineral resources with the general large language model to evaluate its reasoning speed, logical clarity, answer accuracy, and integrity when dealing with professional geological questions.

[0021] This embodiment provides a question-and-answer method for the geological and mining industry, including the following steps: Input the received question to be answered into the geological and mineral large language model to obtain the answer to the question to be answered.

[0022] In an exemplary embodiment, the method for constructing a geological and mineral large language model can be implemented in the following manner: As Figure 2 shown is the main flow chart of the technical solution of this embodiment. The method for constructing a geological and mineral large language model includes the following steps: (1) Collect public and private restricted literature materials in the field of geology and minerals, including textbooks on ore deposit geology, geological exploration reports, monographs on mineral exploration, academic journal papers, and industry research reports, etc.; (2) Clean the collected text data, including removing low-quality data and graphics, correcting recognition errors, filling in missing data, and at the same time processing duplicate information, inconsistent formats, and noise data existing in the original data, and finally obtain a high-quality geological pure text corpus as the basis for subsequent model training; (3) Based on the cleaned text data, construct a special dataset in the field of geology and minerals containing 3,196,400 words, covering various geological knowledge such as ore deposit types, mineral resource exploration, and mineralization mechanisms; (4) Use a knowledge graph of ore deposits containing 3,300 entity nodes and 6,688 relationships. Based on geological expert knowledge and the actual application requirements in the mineral field, design 60 question-and-answer templates, and generate a large number of question-and-answer pairs through template matching for training the model; (5) Use the above method to batch generate question-and-answer pairs containing 1,969,600 words, convert them into a data format suitable for training through a Python script, and combine them with the text dataset to form a mixed geological corpus of 5,166,000 words; (6) Select Baichuan-2 as the basic large language model. This model has strong multilingual processing capabilities, especially good compatibility with Chinese texts, and its model parameters are relatively small, and the training speed is relatively fast; (7) Use the constructed mixed corpus to fine-tune and train the basic model to improve its understanding and reasoning capabilities in the field of geology and minerals; (8) Adjust key training parameters such as learning rate, batch size, and number of training epochs to optimize the training efficiency and accuracy of the model, and initially judge the model performance through the loss function value; (9) Set at least 20 questions in multiple sub-fields in the field of geology and minerals, such as ore deposit types and distributions, mineralization laws, geological characteristics, etc., and conduct comprehensive reasoning evaluation on the trained model (i.e., the geological and mineral large language model, hereinafter referred to as GeoMinLM) GeoMinLM to test its performance on professional questions; (10) Compare the answers of the GeoMinLM model with those of the Baichuan-2 base model before adding the mixed corpus, and analyze the differences in the answer content and accuracy between the two; (11) Compare GeoMinLM with general large language models, and evaluate its inference speed, logical clarity, answer accuracy and integrity when dealing with professional geological problems; Figure 3 The specific process of constructing question-answer pairs through the knowledge graph: First, extract the head entity, tail entity and relationship from the knowledge graph to construct a set of triples; then check whether these triples match the predefined templates; if they match, replace the entity words with the placeholders in the templates, and then fill in the question and answer templates to generate question-answer pairs; if the triples do not match the templates, correct them and try to match again; this process will continue to iterate until question-answer pairs that can be used in the geological corpus are successfully generated; among these templates, the domain and range types of the semantic relationships in the ontology determine the types of question-answer templates, while the categories of geological entities and their corresponding semantic relationships determine their corresponding question-answer templates; in this way, geological entities can replace the words in the templates to generate new question-answer pairs; Figure 4 The performance of the GeoMinLM model during training under different parameter settings; in the figure, the horizontal axis represents the learning rate, and the vertical axis represents the value of the loss function. Different training epochs are distinguished by different curves, representing the cases of training epochs 1, 2, and 4 respectively; it can be seen from the data in the figure that the learning rate and the training epoch have a greater impact on the change of the loss function value; when the training epoch is 1, the loss function value is generally high; in the lower learning rate range (about 0.00002 and below), when the training epochs are 2 and 4, the change of the loss function value tends to be flat, and the loss function value is low; at the same time, by comparing the curves of different training epochs, it is found that the increase in the training epoch has a limited improvement on the loss function value. Especially when the training epoch is 4, the model performance has a slight improvement, but due to the high training cost, the performance improvement is small. Therefore, the training epoch is finally set to 2 and the learning rate is set to 0.00002; this step provides a guiding basis for parameter optimization, and the experimental results are of great significance for determining the best hyperparameter configuration of the GeoMinLM model in practical applications; Table 1 shows the actual output comparison table after the model is added to the mixed corpus; during the research process, this embodiment compared the performance of the Baichuan-2 model and the GeoMinLM model trained with a mixed corpus when answering questions related to geology and mineral resources; it can be seen from the table that GeoMinLM shows higher professionalism and accuracy when providing answers, and can give more detailed and systematic geological knowledge, especially in the fields of genetic analysis, geological characteristics, ore deposit classification, etc., demonstrating its in-depth understanding of the professional field; while the answers of Baichuan-2 are relatively brief, and the details and accuracy in the answers are not as good as GeoMinLM, thus showing that the latter has significant advantages in the application of the geology and mineral resources industry; Table 1 Actual Output Comparison Table after the Model is Added to the Mixed Corpus

[0023] The focus of the present invention lies in significantly enhancing the intelligent question-answering ability by embedding a professional knowledge graph in the geological and mineral field into a large language model. The key point of this method is to enhance the model using a structured knowledge graph, enabling it to more comprehensively and accurately understand professional terms and complex relationships in the geological and mineral field, thereby generating concise, professional, and accurate answers. Compared with traditional intelligent question-answering systems in the geological and mineral field, the present invention solves problems such as insufficient professionalism of the original model and difficulty in understanding complex knowledge through the embedding of the knowledge graph, greatly improving the professionalism and reasoning ability of the model. In addition, based on the construction of a large-scale, hybrid special corpus in the geological and mineral field, the present invention successfully balances the contradiction between the limitation of computing resources and the data scale, enabling this method to maintain a high training efficiency and reasoning performance under the existing hardware resource conditions and reducing the computing cost. By further optimizing hyperparameters (such as learning rate, batch size, number of training epochs, etc.), it is ensured that the model can operate efficiently and stably, improving the overall performance. Generally speaking, the present invention not only enhances the professionalism of the model in the geological and mineral field through the embedding of the knowledge graph, but also enables the model to be trained and reasoned efficiently and economically by optimizing the use of computing resources, thereby providing high-quality intelligent question-answering services. This technical solution provides strong intelligent support for geological resource exploration and mineral development, provides a prior reference for the efficient use of exploration data, and promotes the intelligent process in the geological and mineral field. Compared with the existing technology, the technical solution proposed by the present invention significantly reduces the training and reasoning costs while ensuring the accuracy and efficiency of the model. The trained model performs well in a series of geological tasks such as mineral resource exploration, analysis of mineral types and distributions, metallogenic laws, and background geological knowledge, not only demonstrating a profound understanding of geological knowledge but also showing excellent ability to solve practical geological problems in actual mineral exploration scenarios. The embedding of the knowledge graph significantly improves the professionalism and conciseness of the model's answers. Compared with traditional general large models, the large language model in the geological and mineral field can provide more professional and accurate answers, providing stronger support for geological survey and mineral exploration work.

[0024] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned question-answering method in the geological and mining industry. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memories.

[0025] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned Q&A method for the geological and mining industry are implemented.

[0026] As Figure 5 shown, the computer device 120 may include: at least one processor 121, such as a Central Processing Unit (CPU), at least one communication interface 123, a memory 124, and at least one communication bus 122. Among them, the communication bus 122 is used to realize the connection and communication between these components. Among them, the communication interface 123 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the communication interface 123 may further include a standard wired interface and a wireless interface. The memory 124 may be a high-speed random access memory (Random Access Memory, RAM), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 124 may further be at least one storage device located far from the aforementioned processor 121. Among them, an application program is stored in the memory 124, and the processor 121 calls the program code stored in the memory 124 to execute any of the above method steps. Among them, the communication bus 122 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 122 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5It is represented by only one line, but it does not mean that there is only one bus or one type of bus. Among them, the memory 124 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 124 may further include a combination of the above types of memories. Among them, the processor 121 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. Among them, the processor 121 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Optionally, the memory 124 is further configured to store program instructions. The processor 121 may call the program instructions to implement the method for answering questions in the mining industry as described in this embodiment.

[0027] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope of the present invention as protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A method for constructing a geological and mineral language model, characterized in that: The following steps are involved: S1: Data cleaning and preprocessing of literature in the field of geology and mineral resources to obtain a special data set in the field of geology and mineral resources; S2: obtaining a mixed geological corpus based on the geological and mineral field-specific dataset and the mineral deposit knowledge graph; S3: Use the mixed geological corpus to fine-tune the basic large language model to obtain a geological and mineral large language model.

2. The method for constructing a geological and mineral language model according to claim 1, characterized in that: Step S1 specifically includes: S11: Acquire literature data in the field of geology and mineral resources, perform data cleaning and preprocessing on the literature data in the field of geology and mineral resources, and obtain cleaned text data; S12: Obtain a dataset dedicated to the field of geology and mineral resources based on the cleaned text data.

3. The method for constructing a geological and mineral language model according to claim 1, characterized in that: The literature materials in the field of geology and mineral resources include ore deposit textbooks, geological exploration reports, mineral exploration monographs, academic journal papers and industry research reports.

4. The method for constructing a geological and mineral language model according to claim 1, characterized in that: The data cleaning and preprocessing of literature in the field of geology and mineral resources includes removing low-quality data and graphics, correcting recognition errors, filling in missing data, removing duplicate information and noise data, and unifying the format.

5. The method for constructing a geological and mineral large language model according to claim 1, characterized in that: Step S2 specifically includes: S21: extracting a head entity, a tail entity, and a relationship from the mineral deposit knowledge graph, and constructing a triple set according to the head entity, the tail entity, and the relationship; S22: construct a question-answer template, and check whether the triple set matches the question-answer template: if not, modify the triple set; if matching, perform entity word replacement to obtain a question-answer pair; S23: Format conversion is performed on the question-answer pairs to obtain training format question-answer data, and a mixed geological corpus is obtained based on the training format question-answer data and the geological and mineral field-specific data set.

6. The method for constructing a geological and mineral language model according to claim 1, characterized in that: Step S3 specifically includes: selecting Baichuan-2 as the basic large language model, fine-tuning the basic large language model according to the settings of a learning rate of 0.00002 and a training number of 2, to obtain a geological and mineral large language model.

7. The method for constructing a geological and mineral language model according to claim 1, characterized in that: The method for constructing a geological and mineral big language model also includes: setting multiple questions in multiple sub-fields of the geological and mineral field, and conducting a comprehensive reasoning evaluation of the geological and mineral big language model; comparing the answers of the geological and mineral big language model with the answers of the Baichuan-2 basic model; comparing the geological and mineral big language model with the general big language model to evaluate its reasoning speed, logical clarity, answer accuracy and completeness when dealing with professional geological problems.

8. A question-answering method for the mining industry, characterized in that: The following steps are involved: A geological and mineral large language model is obtained by using the geological and mineral large language model construction method as described in any one of claims 1 to 7; and the received question to be answered is input into the geological and mineral large language model to obtain the answer to the question to be answered.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the geological and mining industry question and answer method as described in claim 8 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the geological and mining industry question and answer method as described in claim 8 are implemented.

Citation Information

Patent Citations

  • Mineral knowledge question answering method and system based on large language model

    CN118193708A

  • Geographic and mineral industry question and answer method based on large language model and knowledge base

    CN118260388A

Cited By

  • Structured corpus generation method and device for geological map multi-modal large model training

    CN121353573A