Integrated cue word driven mine large model question and answer data set generation method

By integrating prompt-driven methods, a mining industry question-and-answer dataset is generated, which solves the problems of lack of knowledge and insufficient reliability of datasets in the application of general large models in the mining industry. It achieves efficient and professional dataset construction and improves the quality and coverage of question-and-answer.

CN121920518APending Publication Date: 2026-04-24CHINA COAL RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA COAL RES INST
Filing Date
2025-12-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing general-purpose large models lack expertise in mining industry applications, rely on high-quality data for fine-tuning, have low construction efficiency, and suffer from insufficient question-answering quality and coverage, resulting in inadequate dataset reliability.

Method used

By integrating prompt-driven methods, we acquire mining industry-related corpora, generate question-and-answer pairs, and score them based on domain information and text quality. We then combine scenario evaluation to optimize the question-and-answer pairs, ensuring semantic quality and domain reliability, and constructing a high-quality training dataset.

Benefits of technology

It has enabled the automated condensation of mining industry knowledge and the construction of high-quality training datasets, improving generation efficiency and data quality, and ensuring the professionalism and reliability of question-and-answer content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920518A_ABST
    Figure CN121920518A_ABST
Patent Text Reader

Abstract

The invention provides an integrated cue word driven mine large model question and answer data set generation method, and relates to the technical field of data processing.The method comprises the steps that a first knowledge base containing corpora related to the mine industry is obtained; extracting a first question-answer pair from the first knowledge base, and determining field information to which the first question-answer pair belongs and a first quality score of text quality of the first question-answer pair; performing scene evaluation on the first question and answer pair based on the domain information to obtain a second quality score of scene safety; optimizing the first question and answer pair according to the first quality score and the second quality score to obtain a second question and answer pair; according to the second question and answer pair, the domain information, the first quality score and the second quality score, the training data set is determined, automatic condensation of the mine industry knowledge and construction of the high-quality training data set are achieved, and the generation efficiency and the data quality are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for generating a large-scale mining model question-answering dataset driven by integrated prompt words. Background Technology

[0002] With the continuous advancement of intelligent, information-based, and digital construction in the mining industry, the need for industry knowledge accumulation and application is becoming increasingly urgent. Artificial intelligence, especially Large Language Model (LLM) technology, has shown great potential in natural language processing, knowledge extraction, and intelligent question answering. However, existing general-purpose large models have the following shortcomings in specialized mining scenarios: (1) Lack of industry knowledge: Although the general model has strong language capabilities, it lacks specialized knowledge in areas such as mine safety, production scheduling, and equipment operation and maintenance.

[0003] (2) Fine-tuning depends on high-quality data: Fine-tuning of large industry models requires a large number of high-quality question-and-answer datasets, but the existing data in the mining industry are mostly raw corpora, documents or unstructured data, lacking systematic question-and-answer pairs.

[0004] (3) Low construction efficiency: Existing methods for constructing question-answer datasets often rely on multiple rounds of prompt word sequential calls. For example, first generate question-answer pairs, then perform sub-domain annotation, then perform quality assessment, and finally revise low-quality data. This "serial" method results in long inference time and significant computational and cost overhead.

[0005] (4) Question and answer quality and insufficient coverage: Question and answer tasks generated by single prompt words are prone to problems that deviate from the industry scenario, resulting in insufficient question and answer coverage, a high proportion of low-quality questions and answers, and affecting the reliability of the dataset. Summary of the Invention

[0006] This application aims to at least partially address one of the technical problems in the related art.

[0007] Therefore, the first objective of this application is to propose an integrated prompt-driven method for generating large-scale mining model question-answering datasets, so as to achieve the generation of high-quality datasets and improve the accuracy of model fine-tuning.

[0008] The second objective of this application is to propose an integrated prompt-driven device for generating large-scale mining model question-answering datasets.

[0009] The third objective of this application is to propose an electronic device.

[0010] The fourth objective of this application is to provide a computer-readable storage medium.

[0011] The fifth objective of this application is to provide a computer program product.

[0012] To achieve the above objectives, the first aspect of this application proposes a method for generating a large-scale mining model question-answering dataset driven by prompt words, comprising: Obtain the first knowledge base containing corpus related to the mining industry; Extract the first question-answer pair from the first knowledge base, and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair; Based on the domain information, a scenario assessment is performed on the first question-and-answer pair to obtain a second quality score for scenario safety. The first question-answer pair is optimized based on the first quality score and the second quality score to obtain the second question-answer pair; The training dataset is determined based on the second question-answer pair, the domain information, the first quality score, and the second quality score.

[0013] To achieve the above objectives, a second aspect of this application proposes an integrated prompt-driven large-scale mining model question-answering dataset generation device, comprising: The first acquisition module is used to acquire a first knowledge base containing corpus related to the mining industry. The second acquisition module is used to extract the first question-answer pair from the first knowledge base and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair. The third acquisition module is used to perform scenario evaluation on the first question-answer pair based on the domain information to obtain a second quality score for scenario safety. The optimization module is used to optimize the first question-answer pair based on the first quality score and the second quality score to obtain a second question-answer pair; The generation module is used to determine the training dataset based on the second question-answer pair, the domain information, the first quality score, and the second quality score.

[0014] To achieve the above objectives, a third aspect of this application provides an electronic device comprising: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method described in the first aspect embodiment.

[0015] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method described in the first aspect embodiment.

[0016] To achieve the above objectives, a fifth aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0017] This application provides a method for generating a large-scale mining model question-answering dataset driven by integrated prompt words. It generates a first question-answer pair using a first knowledge base of mining industry-related corpora, simultaneously generating a first quality score for the domain information and text quality of the first question-answer pair. Based on the domain information, a scenario evaluation is performed to obtain a second quality score related to scenario safety. While ensuring the semantic quality of the text, the reliability of the relevant knowledge in the first question-answer pair within the corresponding domain is guaranteed. The first question-answer pair is optimized based on the first and second quality scores to obtain an optimized second question-answer pair, thereby determining the training dataset. This method achieves automated condensation of mining industry knowledge and construction of a high-quality training dataset, effectively improving generation efficiency and data quality.

[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an integrated prompt-driven method for generating a large-scale mining model question-answering dataset, as provided in an embodiment of this application; Figure 2 A flowchart illustrating an integrated prompt-driven method for generating a large-scale mining model question-answering dataset, as provided in an embodiment of this application; Figure 3 A logical schematic diagram of an integrated prompt-driven method for generating a large-scale mining model question-answering dataset, provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an integrated prompt-driven large-scale mining model question-answering dataset generation device provided in an embodiment of this application. Detailed Implementation

[0020] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0021] The following describes, with reference to the accompanying drawings, an embodiment of this application of a method for generating an integrated prompt-driven large-scale mining model question-answering dataset.

[0022] Figure 1 This is a flowchart illustrating a method for generating a large-scale mining model question-answering dataset driven by prompt words, as provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps: S101, obtain the first knowledge base containing corpus related to the mining industry.

[0023] To obtain initial multi-source text data for the mining industry, in this embodiment, the sources of initial multi-source text data can be mining safety regulations, industry standards, equipment operation manuals, training materials, accident case databases, scientific research papers, and industry technical reports, etc., to ensure that the corpus can comprehensively cover the knowledge scope of mine safety production and management.

[0024] The initial multi-source text data collected often has problems such as redundancy, noise and inconsistent format, so it is necessary to clean and preprocess the initial multi-source text data.

[0025] The initial multi-source text data is denoised by removing web page tags, advertising content, and redundant symbols to obtain the first multi-source text data.

[0026] Calculate the similarity between any two text segments in the first multi-source text data. The corresponding similarity is obtained from the cosine similarity, specifically:

[0027] in, For text fragments The similarity between them; Represents a text fragment The length of (i.e., the Euclidean norm).

[0028] Based on similarity, duplicate texts in the first multi-source text data are removed to obtain the second multi-source text data. In this embodiment, two text segments with a similarity greater than or equal to a preset similarity threshold are identified as duplicate texts. The corpus identified as duplicate texts is then removed to include only one of the text segments, resulting in the deduplicated second multi-source text data.

[0029] Furthermore, the second multi-source text data is segmented and formatted. Using sentence segmentation and segmentation algorithms, long texts are divided into independent knowledge units to ensure the contextual integrity of the subsequent input large model corpus. In terms of format, this embodiment uniformly converts the text into the lightweight markup language Markdown format to ensure structure and parsability, thereby obtaining the first knowledge base.

[0030] In some embodiments, to reduce semantic ambiguity, before generating the first knowledge base, a mining-related terminology database can be used to unify and standardize synonyms and abbreviations appearing from different sources, thereby ensuring the consistency and professionalism of the corpus and improving the generation effect of the first knowledge base.

[0031] S102, extract the first question-answer pair from the first knowledge base, and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair.

[0032] Optionally, a pre-trained large model can be used to perform reasoning analysis on the first knowledge base, automatically generating questions and corresponding answers to obtain the first question-answer pair in the first knowledge base. The design of the questions in the first question-answer pair covers key knowledge points in the corpus, such as mine safety regulations, production scheduling processes, equipment operation and maintenance points, and emergency response measures, ensuring that the generated questions closely meet the actual needs of the mining industry. The answer part focuses on accuracy and practicality, avoiding redundant expressions, making it both concise and clear, and able to reflect the professional information needed by mining practitioners in actual operation.

[0033] In some embodiments, during the process of generating the first question-and-answer pair through a pre-trained large model, the large model can also acquire the domain based on the pre-generated first question-and-answer pair, thereby obtaining domain information. In this embodiment, key domains of the mining industry can be condensed based on expert experience, such as underground mining technology, open-pit mining technology, roadway surrounding rock stability, rockburst prevention, coal preparation theory and technology, etc. Domain matching is performed according to the actual situation of the content involved in the first question-and-answer pair to obtain domain information.

[0034] In some embodiments, during the process of generating the first question-answer pair using a pre-trained large model, the large model can also perform text quality assessment on aspects such as the completeness and accuracy of the content in the first question-answer pair to obtain a first quality score for the first question-answer pair.

[0035] S103, based on domain information, perform scenario evaluation on the first question-answer pair to obtain a second quality score for scenario security.

[0036] Mining industry Q&As often imply complex production scenarios and safety risks. For example, the same textual description may correspond to completely different safety risks in high-gas mines and low-gas mines, thus requiring a second quality score for the safety of the current Q&A in the specific scenario.

[0037] Optionally, the corresponding scenarios in the domain can be obtained based on the domain information, and the accuracy of risk identification and safety measures in the first question-and-answer pair can be inferred and analyzed through a large model. If the answer ignores the key risk scenarios in the domain, the second quality score will be lower.

[0038] S104, optimize the first question-and-answer pair based on the first quality score and the second quality score to obtain the second question-and-answer pair.

[0039] Optionally, the total quality score can be obtained by summing the first quality score and the second quality score. The higher the total quality score, the higher the quality of the first correct answer; conversely, the lower the total quality score, the lower the quality of the first correct answer.

[0040] Optionally, the first question-and-answer pair with a low overall quality score can be optimized. For example, if the overall quality score is lower than the quality threshold, it means that the quality of the current first question-and-answer pair is poor. In this case, the first question-and-answer pair can be repaired and optimized to obtain the second question-and-answer pair, thereby improving the quality of the question-and-answer pair.

[0041] In some embodiments, the first question-and-answer pair can be repaired and optimized based on the knowledge of mining professional corpus referenced in the generative large model, and the content of the first question-and-answer pair can be reprocessed to obtain a second question-and-answer pair with higher professionalism and reliability.

[0042] In some embodiments, if the total quality score is higher than the quality threshold, it indicates that the quality of the first question-and-answer pair is good and no optimization is needed. In other words, the second question-and-answer pair is the same as the first question-and-answer pair.

[0043] S105. Determine the training dataset based on the second question-answer pair, domain information, first quality score, and second quality score.

[0044] Optionally, the total quality score can be determined based on the first quality score and the second quality score. In this embodiment, each training sample includes at least the second question-answer pair, domain information, and the total quality score, thereby obtaining a training dataset based on all training samples, which effectively improves the generation efficiency and data generation quality of the training dataset.

[0045] In this embodiment, a first question-and-answer pair is generated based on a first knowledge base of mining industry-related corpus. Simultaneously, a first quality score is generated, including domain information and text quality. Based on the domain information, a scenario evaluation is performed to obtain a second quality score related to scenario safety. This ensures both the semantic quality of the text and the reliability of the relevant knowledge in the first question-and-answer pair within the corresponding domain. The first question-and-answer pair is then optimized based on the first and second quality scores to obtain an optimized second question-and-answer pair, thereby determining the training dataset. Domain labels are added synchronously during the question-and-answer generation process to achieve structured data management, providing support for fine-tuning the industry model. This achieves automated condensation of mining industry knowledge and the construction of a high-quality training dataset, effectively improving generation efficiency and data quality.

[0046] Figure 2 This is a flowchart illustrating a method for generating a large-scale mining model question-answering dataset driven by prompt words, as provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps: S201, obtain the first knowledge base containing corpus related to the mining industry.

[0047] In this application embodiment, the implementation method of step S201 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0048] S202, Generate an integrated prompt based on the prompt word template.

[0049] The integration prompt includes one or more target tasks, and may also include output format specifications, enabling large models to complete the entire data generation process in a single inference. The target tasks are used to indicate the generation of the first question-answer pair, the domain information to which the first question-answer pair belongs, and the first quality score of the first question-answer pair. This avoids the token redundancy and computational waste caused by multiple rounds of prompt calls in traditional methods, while ensuring the structured, consistent, and high-quality output results.

[0050] S203 uses a pre-trained large model to extract the first question-answer pair from the first knowledge base based on the integrated prompt and the first knowledge base, and obtains the domain information to which the first question-answer pair belongs and the first quality score of the first question-answer pair.

[0051] Optionally, question-answer pair set It can be represented as:

[0052] in, For the question, The answer is N, and the number of question-answer pairs generated is N.

[0053] When generating domain information based on the integrated prompt, key areas of the mining industry (such as underground mining technology, open-pit mining technology, roadway surrounding rock stability, rockburst prevention, coal preparation theory and technology, etc.) can be used as keywords to guide the generation of question-answer pairs focusing on content strongly related to these areas, thereby improving the relevance of the generated question-answer data to the mining industry. Domain segmentation not only makes the dataset more hierarchical and structured but also provides targeted support for fine-tuning the large model on subdivided tasks, thus improving the accuracy of industry applications. Determining the most suitable domain label for each first question-answer pair yields the domain information for that pair, which can be specifically represented as:

[0054] in, The first question and answer were correct. The domain tag to which it belongs; The probability distribution of the first question-answer pair belonging to each domain label is estimated by the large model. , This is a predefined set of domain labels. In this embodiment, it covers fields such as underground mining technology, open-pit mining technology, roadway surrounding rock stability, rockburst prevention, and coal preparation theory and technology.

[0055] In the quality assessment phase of the large model, this embodiment sets the quality index set as {r,c,a,l}, representing relevance, completeness, accuracy, and language clarity, respectively. The large model performs quality assessment based on the ensemble prompt, obtaining the score corresponding to each first question-answer pair under the quality index, thus obtaining the first quality score for each first question-answer pair. The first quality score is expressed as:

[0056] in, The first quality score is given to the first correct answer. , , and The scores are the relevance score, completeness score, accuracy score, and language clarity score for the first question-answer pair. , , and These are the weights for relevance, completeness, accuracy, and linguistic clarity, respectively. .

[0057] S204, Identify the key scenarios corresponding to the domain information.

[0058] Mining industry Q&A often implicitly contains complex production scenarios and safety risks. For example, the same textual description may correspond to completely different safety risks in high-gas mines and low-gas mines. Such implicit scenarios and risks cannot be identified solely by large models. The types of implicit scenarios and risks involved in mining industry Q&A are highly dependent on their respective domain labels. If the domain labels automatically identified by the Q&A are not included in the safety assessment, large models may be unable to identify hidden scenario differences, leading to misjudgments in safety. Therefore, this embodiment proposes a "scenario + risk" consistency index based on domain labels to measure the consistency of generated answers within their respective domain labels. The safety, correctness, and risk integrity of the system.

[0059] Optionally, based on the domain tags in the domain information, the key scenarios for the current domain tag can be determined from the scenario set. Represented as:

[0060] That is, when the domain label is underground mining-ventilation, the key scenarios include, but are not limited to, high-gas mines and low-gas mines. When the domain label is open-pit mining-blasting, the key scenarios include, but are not limited to, multi-hole blasting, pre-splitting blasting and slope hazards.

[0061] S205, Obtain the probability distribution of the first question-answer pair in the key scenario.

[0062] Alternatively, the probability distribution model can be expressed as: , indicating a question-and-answer pair X={q,a} In subdomains Key scenarios ( m The probability distribution of the first question-answer pair is determined based on the probability distribution model. The probability distribution of each key scenario under the corresponding domain label can be obtained by reasoning and analysis of the pre-trained model, which improves the accuracy of the results.

[0063] S206, Obtain a predefined vector knowledge base, and determine the credibility of the answer in the first question-and-answer pair based on the vector knowledge base and key scenarios.

[0064] In some embodiments, after obtaining the first knowledge base, the first knowledge base can be input into a general large model for reasoning, and combined with Retrieval-Augmented Generation (RAG) technology, key information and knowledge bases can be constructed using the core standard library of the mine to improve the professionalism and accuracy of subsequent question-answer generation; during the knowledge base construction process, the vector model (Bidirectional Encoder Representations from Transformers, BERT) is used to obtain the vector of each text segment for storage, resulting in a vector knowledge base.

[0065] The search is based on a vector knowledge base to determine the accuracy of risk identification and security measures in scenario m of the answer in the first question-answer pair, which is also the credibility of the answer in the first question-answer pair. Understandably, if the answer ignores the key risks associated with this area (such as gas, carbon monoxide, roof collapse, blasting vibration, slope slippage, etc.), then the corresponding scenario... This will decrease significantly, resulting in a lower data quality assessment.

[0066] S207. Determine the second quality score for the first question-answer pair based on probability distribution and credibility.

[0067] Optionally, the entropy of the first question-and-answer pair can be determined based on the probability distribution to penalize the uncertainty of the scenario distribution. The specific calculation is as follows:

[0068] in, The entropy corresponding to the first question and answer pair; Let be the probability distribution of the first question-answer pair in the key scenario m.

[0069] Furthermore, based on probability distribution, credibility, and entropy, a second quality score is determined, which is:

[0070] in, The second quality score for the first correct answer; To assess the credibility of the answer; Entropy; Let be the probability distribution of the first question-answer pair in key scenario m; The entropy penalty coefficient is... .

[0071] S208, optimize the first question-and-answer pair based on the first quality score and the second quality score to obtain the second question-and-answer pair.

[0072] In some embodiments, the first quality score and the second quality score can be combined to obtain a comprehensive quality score for the first question-answer pair. The comprehensive quality score is:

[0073] in, The overall quality score for the first correct answer; The first quality score is given to the first correct answer. The second quality score for the first correct answer; These are the weighting coefficients. , When =1, it relies entirely on the quality assessment of the large model and ignores the safety and consistency of the scenario assessment. When =0, it relies entirely on the security consistency of the scenario assessment and ignores the semantic quality of the question-answer pair.

[0074] In response to a comprehensive quality score lower than a quality threshold, the first question-answer pair is optimized based on a pre-trained optimization model to obtain a second question-answer pair; for example, the quality threshold is... Then in If the first question-and-answer pair is determined to be of low quality, it needs to be optimized.

[0075] Optionally, relevant content can be retrieved from the vector knowledge base using a pre-trained optimization model, and the first question-answer pair can be optimized based on the relevant content to obtain the second question-answer pair, which is represented as follows:

[0076] in, As a vector knowledge base, retrieve strongly related content pairs from the vector knowledge base for the first question and answer pair. ) to be optimized, To reconstruct the function, this includes the hidden reasoning process of the large model, supplementing missing mining professional points, eliminating vague or ambiguous expressions, and reprocessing the answers using a retrieval-enhanced corpus to make them more in line with the knowledge system and expression habits of the mining industry.

[0077] If the overall quality score is greater than or equal to the quality threshold, indicating that the current first question-answer pair is not a low-quality question-answer pair, then the first question-answer pair is directly determined as the second question-answer pair.

[0078] In this application embodiment, the implementation method of step S208 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0079] S209. Determine the training dataset based on the second question-answer pair, domain information, first quality score, and second quality score.

[0080] Understandably, a comprehensive quality score is obtained by weighted summation of the first and second quality scores. This comprehensive score is then used as a training sample, concatenated with the second question-answer pair, domain information, and the comprehensive quality score. Each training sample can then be represented as: The final generated training dataset is .

[0081] In this application embodiment, the implementation method of step S209 can be implemented in any of the embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0082] In this embodiment, a first question-and-answer pair is generated based on a first knowledge base of mining industry-related corpora. Simultaneously, a domain label and a first quality score (relevance, completeness, accuracy, and linguistic clarity) are generated for the first question-and-answer pair. The probability distribution of the first question-and-answer pair in each scenario is determined based on the corresponding scenario of the domain information. The reliability of the answer is determined based on vector knowledge base reasoning. A second quality score is obtained by combining the penalty term entropy. The first and second quality scores are weighted and summed to obtain a comprehensive quality score for evaluating the first question-and-answer pair. This optimizes the lower-quality first question-and-answer pairs, resulting in a second question-and-answer pair. This improves the semantic quality and scenario safety consistency of the question-and-answer pairs. A training dataset is determined based on the comprehensive quality score, the second question-and-answer pair, and the domain label. This achieves automated condensation of mining industry knowledge and the construction of a high-quality training dataset. The final output question-and-answer content is not only more readable but also significantly improved in terms of professionalism and reliability.

[0083] Figure 3 This is a logical diagram illustrating a method for generating a large-scale mining model question-answering dataset driven by integrated prompts, as provided in an embodiment of this application. First, mining industry corpus is acquired. This corpus is then cleaned and preprocessed to obtain a first knowledge base. A vector knowledge base (RAG knowledge base) is constructed based on this first knowledge base. A pre-trained large model performs multi-task processing based on integrated prompts to generate a first question-answer pair, domain information, and a first quality score. Further, the domain information is combined with scenario evaluation to obtain a second quality score. The first and second quality scores are then fused to obtain a comprehensive quality score. Low-quality data is identified based on the comprehensive quality score and reconstructed and optimized to obtain the final second question-answer pair. Finally, based on the second question-answer pair, the comprehensive quality score, and the domain information, a training dataset is obtained, realizing the automated condensation of mining industry knowledge and the construction of a high-quality training dataset.

[0084] To implement the above embodiments, this application also proposes an integrated prompt-driven large-scale mining model question-answering dataset generation device.

[0085] Figure 4This is a schematic diagram of an integrated prompt-driven large-scale mine model question-answering dataset generation device provided in an embodiment of this application. Figure 4 As shown, the integrated prompt-driven large-scale mining model question-answering dataset generation device 400 includes: The first acquisition module 401 is used to acquire a first knowledge base containing corpus related to the mining industry; The second acquisition module 402 is used to extract the first question-answer pair from the first knowledge base and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair. The third acquisition module 403 is used to perform scenario evaluation on the first question-answer pair based on domain information to obtain a second quality score for scenario security. Optimization module 404 is used to optimize the first question-answer pair based on the first quality score and the second quality score to obtain the second question-answer pair; The generation module 405 is used to determine the training dataset based on the second question-answer pair, domain information, first quality score, and second quality score.

[0086] Furthermore, in one possible implementation of this application embodiment, the second acquisition module 402 is used for: Based on the prompt word template, an integrated prompt is generated. The integrated prompt includes one or more target tasks, which are used to indicate the generation of the first question-answer pair, the domain information to which the first question-answer pair belongs, and the first quality score of the first question-answer pair. Using a pre-trained large model, based on the ensemble prompt and the first knowledge base, the first question-answer pair in the first knowledge base is extracted, and the domain information to which the first question-answer pair belongs and the first quality score of the first question-answer pair are obtained.

[0087] Furthermore, in one possible implementation of this application embodiment, the third acquisition module 403 is used for: Identify the key scenarios corresponding to the domain information; Obtain the probability distribution of the first question-answer pair in key scenarios; Obtain a predefined vector knowledge base, and determine the credibility of the answer in the first question-and-answer pair based on the vector knowledge base and key scenarios; Based on the probability distribution and credibility, determine the second quality score for the first question-answer pair.

[0088] Furthermore, in one possible implementation of this application embodiment, the third acquisition module 403 is used for: Determine the entropy of the first question-and-answer pair based on the probability distribution; The second quality score is determined based on probability distribution, credibility, and entropy.

[0089] Furthermore, in one possible implementation of this application embodiment, the optimization module 404 is used for: The first quality score and the second quality score are combined to obtain the comprehensive quality score of the first question-answer pair; In response to the overall quality score being less than the quality threshold, the first question-answer pair is optimized based on the pre-trained optimization model to obtain the second question-answer pair; In response to a comprehensive quality score greater than or equal to a quality threshold, the first question-answer pair is identified as the second question-answer pair.

[0090] Furthermore, in one possible implementation of this application embodiment, the optimization module 404 is used for: By using a pre-trained optimization model, relevant content is retrieved from the vector knowledge base, and the first question-answer pair is optimized based on the relevant content to obtain the second question-answer pair.

[0091] Furthermore, in one possible implementation of this application embodiment, the first acquisition module 401: Obtain initial multi-source text data from the mining industry; The initial multi-source text data is denoised to obtain the first multi-source text data. Calculate the similarity between every two text segments in the first multi-source text data; Duplicate text in the first multi-source text data is removed based on similarity to obtain the second multi-source text data. The second multi-source text data is segmented and formatted to obtain the first knowledge base.

[0092] It should be noted that the foregoing explanation of the embodiment of the integrated prompt word-driven question-answering dataset generation method for large mining models also applies to the integrated prompt word-driven question-answering dataset generation device for large mining models in this embodiment, and will not be repeated here.

[0093] In this embodiment, a first question-and-answer pair is generated based on a first knowledge base of mining industry-related corpora. Simultaneously, a domain label and a first quality score (relevance, completeness, accuracy, and linguistic clarity) are generated for the first question-and-answer pair. The probability distribution of the first question-and-answer pair in each scenario is determined based on the scenario corresponding to the domain information. The reliability of the answer is determined based on vector knowledge base reasoning. A second quality score is obtained by combining the penalty term entropy. The first and second quality scores are weighted and summed to obtain a comprehensive quality score for evaluating the first question-and-answer pair. This optimizes the lower-quality first question-and-answer pairs, resulting in a second question-and-answer pair. This improves the semantic quality and scenario safety consistency of the question-and-answer pairs. A training dataset is determined based on the comprehensive quality score, the second question-and-answer pair, and the domain label. This achieves automated condensation of mining industry knowledge and the construction of a high-quality training dataset. The final output question-and-answer content is not only more readable but also significantly improved in terms of professionalism and reliability.

[0094] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments. To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0095] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0096] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0097] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0098] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0099] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0100] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0101] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0102] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0103] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0104] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0106] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for generating a large-scale mining model question-answering dataset driven by integrated prompt words, characterized in that, The method includes: Obtain the first knowledge base containing corpus related to the mining industry; Extract the first question-answer pair from the first knowledge base, and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair; Based on the domain information, a scenario assessment is performed on the first question-and-answer pair to obtain a second quality score for scenario safety. The first question-answer pair is optimized based on the first quality score and the second quality score to obtain the second question-answer pair; The training dataset is determined based on the second question-answer pair, the domain information, the first quality score, and the second quality score.

2. The method according to claim 1, characterized in that, The step of extracting the first question-answer pair from the first knowledge base and determining the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair includes: Based on the prompt template, an integrated prompt is generated, which includes one or more target tasks. The target tasks are used to indicate the generation of a first question-answer pair, the domain information to which the first question-answer pair belongs, and the first quality score of the first question-answer pair. Using a pre-trained large model, based on the integrated prompt and the first knowledge base, the first question-answer pair is extracted from the first knowledge base, and the domain information to which the first question-answer pair belongs and the first quality score of the first question-answer pair are obtained.

3. The method according to claim 2, characterized in that, The step of evaluating the first question-answer pair based on the domain information to obtain a second quality score for scenario safety includes: Identify the key scenarios corresponding to the domain information; Obtain the probability distribution of the first question-answer pair in the key scenario; Obtain a predefined vector knowledge base, and determine the credibility of the answer in the first question-answer pair based on the vector knowledge base and the key scenario; Based on the probability distribution and the credibility, a second quality score is determined for the first question-answer pair.

4. The method according to claim 3, characterized in that, Determining the second quality score of the first question-answer pair based on the probability distribution and the credibility includes: The entropy of the first question-and-answer pair is determined based on the probability distribution. The second quality score is determined based on the probability distribution, the credibility, and the entropy.

5. The method according to claim 3 or 4, characterized in that, The step of optimizing the first question-answer pair based on the first quality score and the second quality score to obtain the second question-answer pair includes: The first quality score and the second quality score are combined to obtain the comprehensive quality score of the first question-answer pair; In response to the overall quality score being less than the quality threshold, the first question-answer pair is optimized based on the pre-trained optimization model to obtain the second question-answer pair; In response to the overall quality score being greater than or equal to the quality threshold, the first question-answer pair is determined to be the second question-answer pair.

6. The method according to claim 5, characterized in that, The pre-trained optimization model optimizes the first question-answer pair to obtain a second question-answer pair, including: The pre-trained optimization model retrieves relevant content from the vector knowledge base, and optimizes the first question-answer pair based on the relevant content to obtain the second question-answer pair.

7. The method according to claim 1, characterized in that, The acquisition of the first knowledge base containing mining industry-related corpus includes: Obtain initial multi-source text data from the mining industry; The initial multi-source text data is denoised to obtain the first multi-source text data; Calculate the similarity between every two text segments in the first multi-source text data; Based on the similarity, duplicate text is removed from the first multi-source text data to obtain the second multi-source text data; The second multi-source text data is segmented and formatted to obtain the first knowledge base.

8. A device for generating a large-scale mining model question-answering dataset with integrated prompt words driven by the device, characterized in that, include: The first acquisition module is used to acquire a first knowledge base containing corpus related to the mining industry. The second acquisition module is used to extract the first question-answer pair from the first knowledge base and determine the domain information to which the first question-answer pair belongs and the first quality score of the text quality of the first question-answer pair. The third acquisition module is used to perform scenario evaluation on the first question-answer pair based on the domain information to obtain a second quality score for scenario safety. The optimization module is used to optimize the first question-answer pair based on the first quality score and the second quality score to obtain a second question-answer pair; The generation module is used to determine the training dataset based on the second question-answer pair, the domain information, the first quality score, and the second quality score.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.