Intelligent prospecting question-answering system based on LLM and RAG
Through an intelligent mineral exploration question and answer system based on LLM and RAG, a knowledge base in the field of geological deposits was built and vector measurement index query was conducted, which solved the illusion problem of the general large language model in the field of professional knowledge, and achieved high-accurate mineral exploration question and answer services.
Patent Information
- Application Number
- CN202411689535.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-06
AI Technical Summary
The existing general-purpose large language model has hallucinations in the field of expertise, and cannot effectively deal with input tasks in vertical fields, maintaining output context coherence and consistency with real-world facts.
An intelligent mineral exploration question and answer system based on LLM and RAG is adopted to extract document texts in the field of geological deposits, perform sentence word segmentation and keyword recognition, generate text embedding vectors containing keyword information annotations, store them in vector database, build a knowledge base in the field of geological deposits, and query relevant knowledge texts using vector metric indexes to provide accurate question and answer replies.
It achieves a good match between user problems and the knowledge base, improves the accuracy of retrieval, avoids the hallucination problems of the general large language model in the field of professional knowledge, and provides high-quality mineral search Q&A services.
Smart Images

Figure CN119938814A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mineral prospecting and large-scale models, and more specifically, to an intelligent mineral prospecting question-answering system based on LLM and RAG. Background Art
[0002] Existing general-purpose large-scale language models, such as ChatGPT, ChatGLM, Qwen, etc., are all rooted in the Transformer architecture and incorporate deep learning neural networks with self-attention mechanisms. These models have demonstrated excellent language generation capabilities through pre-training on massive general knowledge data. However, when facing vertical fields, due to the limitations of training data, the models have deviations or errors in processing input tasks, maintaining output contextual coherence, and maintaining consistency with real-world facts, the so-called "hallucinations". In the field of earth sciences, there are currently geoGPT for geospatial data collection, processing, and analysis; Pangu models for meteorological research, but in the field of mineral prospecting, there are currently no more mature models.
[0003] There are two common methods for constructing large language models in vertical fields. One is to use the data in the research field for secondary training based on the pre-trained model, which is also called fine-tuning (e.g. Figure 1 ), this method requires a long time cost, high hardware level, poor plasticity, and cannot be quickly and happily updated with the update of data; another method is to use external professional knowledge base, through retrieval enhancement and general large language model prompt word engineering construction (such as Figure 2 ), this construction method not only has low hardware requirements but is also accurate and practical, and can be quickly updated and iterated as data is updated. Retrieval enhancement refers to triggering knowledge base queries based on the content of the question, and then providing the query results to the general large model as knowledge enhancement to generate better answers. There are many retrieval enhancement technical solutions, and the common ones include recommending knowledge base data by building a special trigger mechanism, building a database query index, using a text similarity algorithm, or using machine learning, deep learning, and ensemble learning algorithms to re-rank query results.
[0004] The most important technologies for building vertical domain knowledge bases are data collection, data purification, and data storage. The authenticity, accuracy, and professionalism of the data are particularly important. Currently, there is no unified standard for this technology, especially in vertical disciplines. Some cases use graph databases such as neo4j to build, and some use relational databases and vector databases to build. Among them, the open source vector knowledge base milvus officially provides methods on how to create and manage vector databases, but does not provide methods on how to purify and build data templates from massive text data. In addition, the knowledge base created directly using raw data has no knowledge focus or does not include context, which leads to inaccurate queries in later stages, thus affecting the answers.
[0005] At the same time, after the professional knowledge is stored, the most important thing is to establish an accurate match between the user's questions and the database knowledge. This problem directly affects the accuracy of the answer and is the most important aspect to avoid the illusion of a universal large language model. The common methods at present are: recommending knowledge base data by building a special trigger mechanism; building a database query index; using a text similarity algorithm; using machine learning, deep learning, and integrated learning algorithms to reorder the query results. These solutions are all methods used outside the data itself. After such data is stored in the knowledge base, the knowledge text itself does not contain key information tags, which will cause inaccurate query results at the data level itself, and then affect the relevance of questions and retrieval results. Summary of the invention
[0006] One of the purposes of the present invention is to provide an intelligent mineral prospecting question and answer method based on LLM and RAG, which solves the hallucination problem of the existing general large language model in the field of professional knowledge; the second purpose of the present invention is to provide an intelligent mineral prospecting question and answer system based on LLM and RAG; the third purpose of the present invention is to provide a computer medium.
[0007] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0008] The first aspect of the present invention provides an intelligent prospecting question-answering method based on LLM and RAG, comprising the following steps:
[0009] Extracting textual content from literature in the field of geological deposits;
[0010] Performing sentence segmentation on the text content, and performing keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information;
[0011] According to the preset keyword weights, a text embedding vector containing keyword information annotations and corresponding weights is obtained;
[0012] Storing the text embedding vector in a vector database to obtain a geological mineral deposit domain knowledge base;
[0013] According to the question text input into the large language model, after performing sentence segmentation and keyword recognition on the question text, the recognized question text is vectorized;
[0014] According to the vector metric index, the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base is searched to obtain the prompt word;
[0015] The large language model outputs a response to the question based on the question text and the prompt word.
[0016] Furthermore, the text content of the literature in the field of geological deposits is extracted, including:
[0017] Extraction of long texts from literature in the field of geological deposits;
[0018] The extracted long text is split into paragraphs of specified size through the RecursiveCharacterTextSplitter function of the langchain library. Then, regular expressions are used to remove useless information including non-printing characters, page numbers, headers and footers. The redundant blanks are replaced with single spaces to obtain the text content of the literature in the field of geological deposits.
[0019] Furthermore, the text content is segmented into sentences, and the results of the sentence segmentation are subjected to keyword recognition to obtain text content annotated with keyword information, including:
[0020] The text content is segmented using a pre-trained sentence segmentation model, and the results of sentence segmentation are identified using a pre-trained keyword recognition model, including the text content annotated with keyword information.
[0021] Furthermore, the pre-trained sentence segmentation model includes:
[0022] A Bert model is used as a base model, and a first preset label sample is used to train the Bert model to obtain the pre-trained sentence segmentation model, wherein the first preset label sample is created according to the description level of the geological mineral deposit field, and the description level of the geological mineral deposit field includes the deposit type, geological structure, mineral rock combination, geophysical and geochemical anomalies, mineralization type and mineralization time, and the first preset label sample includes text and corresponding segmentation words.
[0023] Furthermore, the pre-trained keyword recognition model includes:
[0024] The Bert model is used as a base model, and the Bert model is trained using a second preset label sample to obtain the pre-trained keyword recognition model, wherein the second preset label sample is created based on common keyword corpus in the field of geological mineral deposits. The common keyword corpus in the field of geological mineral deposits includes 7 categories, namely mine names, place names, personal names, time names, stratum names, structure names and other nouns. The second preset label sample includes keywords and corresponding categories.
[0025] Furthermore, the preset keyword weights include
[0026] Different weights are assigned to the keywords according to their corresponding categories. The weight of keywords corresponding to the categories of mine names, time names, stratum names and structure names is 0.2, the weight of keywords corresponding to the categories of human names is 0.1, and the weight of keywords corresponding to the categories of place names and other nouns is 0.05.
[0027] Furthermore, according to the preset keyword weights, a text embedding vector containing keyword information annotations is obtained, including:
[0028] Different weights are assigned to the keywords according to their corresponding categories, and then the keywords are embedded using the Bert model to obtain a text embedding vector containing keyword information annotations.
[0029] Further, according to the question text input into the large language model, sentence segmentation and keyword recognition are performed on the question text, including:
[0030] The question text is segmented using a pre-trained sentence segmentation model, and the question text is segmented using a pre-trained keyword recognition model.
[0031] The second aspect of the present invention provides an intelligent mineral prospecting question-answering system based on LLM and RAG, comprising the following steps:
[0032] An extraction module, which extracts textual content of documents in the field of geological deposits;
[0033] A first keyword recognition module, which performs sentence segmentation on the text content and performs keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information;
[0034] A word embedding module, wherein the word embedding module obtains a text embedding vector including keyword information annotations and corresponding weights according to preset keyword weights;
[0035] A knowledge base module, wherein the knowledge base module stores the text embedding vector into a vector database to obtain a knowledge base in the field of geological deposits;
[0036] A second keyword recognition module, which performs sentence segmentation and keyword recognition on the question text according to the question text input into the large language model, and then vectorizes the recognized question text;
[0037] A prompt word module, wherein the prompt word module searches the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base according to the vector metric index to obtain a prompt word;
[0038] An output module uses a large language model to output a response to the question based on the question text and prompt words.
[0039] The third aspect of the present invention provides a computer medium having a computer program stored thereon, and when the computer program is executed by a processor, the intelligent prospecting question-and-answer method based on LLM and RAG is implemented.
[0040] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0041] The present invention is triggered from within the data itself, processes the context information and keywords in the data, and then stores them in a vector database to obtain a knowledge base in the field of geological deposits, thereby achieving a good match between user questions and the knowledge base, improving the accuracy of retrieval, and avoiding the hallucination problem of general large language models in the field of professional knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a diagram of the architecture of a fine-tuning vertical domain large language model in the prior art;
[0043] Figure 2 This is another architecture diagram of a fine-tuned vertical domain large language model in the prior art;
[0044] Figure 3 A schematic diagram of a flow chart of an intelligent prospecting question-answering method based on LLM and RAG is provided for an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of extracting text content of literature in the field of geological deposits provided by an embodiment of the present invention;
[0046] Figure 5 A model parameter diagram of a pre-trained sentence segmentation model provided in an embodiment of the present invention;
[0047] Figure 6 A model parameter diagram of a pre-trained keyword recognition model provided in an embodiment of the present invention;
[0048] Figure 7 A flow question-answering process architecture diagram of an intelligent mineral prospecting question-answering method based on LLM and RAG provided in an embodiment of the present invention;
[0049] Figure 8 A large language model answer evaluation parameter diagram provided by an embodiment of the present invention;
[0050] Fig. 9 A module schematic diagram of an intelligent mineral prospecting question-and-answer system based on LLM and RAG provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0052] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0053] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0054] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0055] Example 1
[0056] The embodiment of the present invention provides an intelligent prospecting question-answering method based on LLM and RAG, such as Figure 3 As shown, the following steps are included:
[0057] Extracting textual content from literature in the field of geological deposits;
[0058] Performing sentence segmentation on the text content, and performing keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information;
[0059] According to the preset keyword weights, a text embedding vector containing keyword information annotations and corresponding weights is obtained;
[0060] Storing the text embedding vector in a vector database to obtain a geological mineral deposit domain knowledge base;
[0061] According to the question text input into the large language model, after performing sentence segmentation and keyword recognition on the question text, the recognized question text is vectorized;
[0062] According to the vector metric index, the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base is searched to obtain the prompt word;
[0063] The large language model outputs a response to the question based on the question text and the prompt word.
[0064] In a further embodiment, the extracting text content of the literature in the field of geological deposits includes:
[0065] Extraction of long texts from literature in the field of geological deposits;
[0066] Through the RecursiveCharacterTextSplitter function of the langchain library, the extracted long text is split into paragraphs of specified size, and then regular expressions are used to remove useless information including non-printing characters, page numbers, headers and footers, and then the redundant spaces are replaced with single spaces to reduce the redundant space in the text, making the text more compact and easier to process. Through these processes, the text content of the literature in the field of geological deposits is obtained, such as Figure 4 As shown, it is obvious that the extraction process has extracted the main content of the document, which provides data preparation for the subsequent construction of the knowledge base.
[0067] In a further embodiment, the text content is segmented into sentences, and the results of the sentence segmentation are subjected to keyword recognition to obtain keywords, including:
[0068] Since the information in text sentences is often contained in keywords, how to find keywords is an important task. The text content is segmented using a pre-trained sentence segmentation model, and the results of sentence segmentation are identified using a pre-trained keyword recognition model to obtain keywords.
[0069] In a further embodiment, the pre-trained sentence segmentation model includes:
[0070] A Bert model is used as a base model, and a first preset label sample is used to train the Bert model to obtain the pre-trained sentence segmentation model, wherein the first preset label sample is created according to the description level of the geological mineral deposit field, and the description level of the geological mineral deposit field includes the deposit type, geological structure, mineral rock combination, geophysical and geochemical anomalies, mineralization type and mineralization time, and the first preset label sample includes text and corresponding segmentation words.
[0071] In this embodiment, the Bert model is used as the base model, and the first preset label sample is used to fine-tune the Bert model to obtain a sentence segmentation model. The field of geological deposits is fully considered, and the main focus is on the type of deposit, geological structure, mineral rock combination, geophysical and geochemical anomalies, mineralization type, mineralization time and other information. According to the commonly used description levels in the field of mineral deposits, a total of 900 label samples (see Table 1) are created for fine-tuning the sentence segmentation model. Because the bidirectional Transformer architecture of Bert has high value in context understanding and segmentation, the sentence segmentation model obtained by fine-tuning has an accuracy rate of more than 99% in sentence segmentation. Its model parameters are as follows: Figure 5 shown.
[0072] Table 1. Sentence segmentation model training data display table
[0073]
[0074] In a further embodiment, the pre-trained keyword recognition model includes:
[0075] The Bert model is used as a base model, and the Bert model is trained using a second preset label sample to obtain the pre-trained keyword recognition model, wherein the second preset label sample is created based on common keyword corpus in the field of geological mineral deposits. The common keyword corpus in the field of geological mineral deposits includes 7 categories, namely mine names, place names, personal names, time names, stratum names, structure names and other nouns. The second preset label sample includes keywords and corresponding categories.
[0076] In this embodiment, 1400 common keyword corpora in the geological field are used (see Table 2). Since the structure of the question-answering model is generally who / what place / what time / what was done (happened) / what happened, the keywords are divided into 7 categories in Table 2. Finally, a keyword recognition model is obtained based on Bert model fine-tuning, and its recognition accuracy is above 95%. Other training parameters such as Figure 6 shown.
[0077] Table 2. Keyword recognition model training data display table
[0078]
[0079]
[0080] In a further embodiment, the preset keyword weights include
[0081] Different weights are assigned to the keywords according to their corresponding categories. The weight of keywords corresponding to the categories of mine name, time name, stratum name and structure name is 0.2, the weight of keywords corresponding to the category of person name is 0.1, and the weight of keywords corresponding to the category of place name and other nouns is 0.05, see Table 3.
[0082] Table 3. Keyword categories and their weight information table
[0083]
[0084] In a further embodiment, according to the preset keyword weights, a text embedding vector containing keyword information annotations is obtained, including:
[0085] Different weights are assigned to the keywords according to their corresponding categories, and then the keywords are embedded using the Bert model to obtain a text embedding vector containing keyword information annotations.
[0086] In this embodiment, through the processing of this embodiment, the word embedding vector of the keyword not only contains context information, but also contains keyword information, and is then stored in a vector database to obtain a geological mineral deposit field knowledge base, which achieves a good match between user questions and the knowledge base, improves the accuracy of retrieval, and avoids the hallucination problem of general large language models in the field of professional knowledge.
[0087] In a further embodiment, according to the question text input into the large language model, sentence segmentation and keyword recognition are performed on the question text, including:
[0088] The question text is segmented using a pre-trained sentence segmentation model, and the question text is segmented using a pre-trained keyword recognition model.
[0089] In this embodiment, when a user asks a question, the question is also segmented and keyword recognized, and then the knowledge base is queried according to the vector query index. The specific process is as follows: Figure 7 shown.
[0090] In a further embodiment, the method of the embodiment can be developed into a human-machine interactive intelligent prospecting robot based on a microservice architecture.
[0091] In a specific embodiment, the built large model of the vertical field of prospecting was verified for question answering, and compared with the most representative large language models with real-time networking functions such as ChatGPT, Kimi, and Wenxin Yiyan. 300 questions in the field of prospecting in the Qinhang metallogenic belt were used respectively. According to the guidance of experts, reference answers to 300 questions were written. Then, questions were asked to the above models and intelligent prospecting robots respectively, and their answers were evaluated using the Bert-score indicator. The results are as follows: Figure 8 As shown, by comparison, it is found that the questions of the intelligent mineral prospecting answering method of this embodiment are significantly higher than those of other models in terms of Precision and F1, and in terms of Recall, the model of this embodiment is also in an advantageous position. From the perspective of the question answer itself, it effectively avoids the "hallucination" problem of answering professional questions, while the answer of the general large language model has both real data and "hallucination" data.
[0092] Example 2
[0093] The embodiment of the present invention provides an intelligent prospecting question-answering system based on LLM and RAG, and the system implements the question-answering method described in Example 1, such as Fig. 9 As shown, the following steps are included:
[0094] An extraction module, which extracts textual content of documents in the field of geological deposits;
[0095] A first keyword recognition module, which performs sentence segmentation on the text content and performs keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information;
[0096] A word embedding module, wherein the word embedding module obtains a text embedding vector including keyword information annotations and corresponding weights according to preset keyword weights;
[0097] A knowledge base module, wherein the knowledge base module stores the text embedding vector into a vector database to obtain a knowledge base in the field of geological deposits;
[0098] A second keyword recognition module, which performs sentence segmentation and keyword recognition on the question text according to the question text input into the large language model, and then vectorizes the recognized question text;
[0099] A prompt word module, wherein the prompt word module searches the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base according to the vector metric index to obtain a prompt word;
[0100] An output module uses a large language model to output a response to the question based on the question text and prompt words.
[0101] Example 3
[0102] This embodiment provides a computer medium, on which a computer program is stored. When the computer program is executed by a processor, the intelligent mineral prospecting question and answer method based on LLM and RAG described in Example 1 is implemented.
[0103] The same or similar reference numerals correspond to the same or similar components;
[0104] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;
[0105] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. An intelligent prospecting question-answering method based on LLM and RAG, characterized in that: The following steps are involved: Extracting textual content from literature in the field of geological deposits; Performing sentence segmentation on the text content, and performing keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information; According to the preset keyword weights, a text embedding vector containing keyword information annotations and corresponding weights is obtained; Storing the text embedding vector in a vector database to obtain a geological mineral deposit domain knowledge base; According to the question text input into the large language model, after performing sentence segmentation and keyword recognition on the question text, the recognized question text is vectorized; According to the vector metric index, the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base is searched to obtain the prompt word; The large language model outputs a response to the question based on the question text and the prompt word.
2. The intelligent prospecting question-answering method based on LLM and RAG according to claim 1, characterized in that: The text content of the literature in the field of extracting geological deposits includes: Extraction of long texts from literature in the field of geological deposits; The extracted long text is split into paragraphs of specified size through the RecursiveCharacterTextSplitter function of the langchain library. Then, regular expressions are used to remove useless information including non-printing characters, page numbers, headers and footers. The redundant blanks are replaced with single spaces to obtain the text content of the literature in the field of geological deposits.
3. The intelligent prospecting question-answering method based on LLM and RAG according to claim 1 is characterized in that: Sentence segmentation is performed on the text content, and keyword recognition is performed on the result after sentence segmentation to obtain text content marked with keyword information, including: The text content is segmented using a pre-trained sentence segmentation model, and the results of sentence segmentation are identified using a pre-trained keyword recognition model, including the text content annotated with keyword information.
4. The intelligent prospecting question-answering method based on LLM and RAG according to claim 3 is characterized in that: The pre-trained sentence segmentation model includes: A Bert model is used as a base model, and a first preset label sample is used to train the Bert model to obtain the pre-trained sentence segmentation model, wherein the first preset label sample is created according to the description level of the geological mineral deposit field, and the description level of the geological mineral deposit field includes the deposit type, geological structure, mineral rock combination, geophysical and geochemical anomalies, mineralization type and mineralization time, and the first preset label sample includes text and corresponding segmentation words.
5. The intelligent prospecting question-answering method based on LLM and RAG according to claim 4 is characterized in that: The pre-trained keyword recognition model includes: The Bert model is used as a base model, and the Bert model is trained using a second preset label sample to obtain the pre-trained keyword recognition model, wherein the second preset label sample is created based on common keyword corpus in the field of geological mineral deposits. The common keyword corpus in the field of geological mineral deposits includes 7 categories, namely mine names, place names, personal names, time names, stratum names, structure names and other nouns. The second preset label sample includes keywords and corresponding categories.
6. The intelligent prospecting question-answering method based on LLM and RAG according to claim 5, characterized in that: The preset keyword weights include: Different weights are assigned to the keywords according to their corresponding categories. The weight of keywords corresponding to the categories of mine names, time names, stratum names and structure names is 0.2, the weight of keywords corresponding to the categories of human names is 0.1, and the weight of keywords corresponding to the categories of place names and other nouns is 0.
05.
7. The intelligent prospecting question-answering method based on LLM and RAG according to claim 6, characterized in that: According to the preset keyword weights, a text embedding vector containing keyword information annotations is obtained, including: Different weights are assigned to the keywords according to their corresponding categories, and then the keywords are embedded using the Bert model to obtain a text embedding vector containing keyword information annotations.
8. The intelligent prospecting question-answering method based on LLM and RAG according to claim 3 is characterized in that: According to the question text input into the large language model, sentence segmentation and keyword recognition are performed on the question text, including: The question text is segmented using a pre-trained sentence segmentation model, and the question text is segmented using a pre-trained keyword recognition model.
9. An intelligent mineral prospecting question-answering system based on LLM and RAG, characterized in that: The following steps are involved: An extraction module, which extracts textual content of documents in the field of geological deposits; A first keyword recognition module, which performs sentence segmentation on the text content and performs keyword recognition on the result after the sentence segmentation to obtain text content marked with keyword information; A word embedding module, wherein the word embedding module obtains a text embedding vector including keyword information annotations and corresponding weights according to preset keyword weights; A knowledge base module, wherein the knowledge base module stores the text embedding vector into a vector database to obtain a knowledge base in the field of geological deposits; A second keyword recognition module, which performs sentence segmentation and keyword recognition on the question text according to the question text input into the large language model, and then vectorizes the recognized question text; A prompt word module, wherein the prompt word module searches the knowledge text with the highest correlation with the question text in the geological deposit field knowledge base according to the vector metric index to obtain a prompt word; An output module uses a large language model to output a response to the question based on the question text and prompt words.
10. A computer medium, characterized in that The computer medium stores a computer program, and when the computer program is executed by the processor, it implements the intelligent mineral prospecting question-and-answer method based on LLM and RAG as described in any one of claims 1 to 8.
Citation Information
Patent Citations
A text sentence vector representation method and system
CN109408797A
RAG knowledge question-answering method and device based on fusion vector and keyword retrieval
CN117951274A
Mineral knowledge question answering method and system based on large language model
CN118193708A
Cited By
Synthetic data set construction method and electronic equipment
CN120975247A