Method for generating fine-tuning question and answer data set based on retrieval enhancement

By generating a fine-tuned question-answering dataset based on a retrieval enhancement method and performing multiple rounds of optimization using a large language model and vector library, the problems of small number of samples and weak domain relevance in the fine-tuning dataset are solved, the quality and efficiency of the dataset are improved, and it is suitable for question-answering tasks in vertical fields.

CN120745855AActive Publication Date: 2025-10-03INST OF AUTOMATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510712006.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-03
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In existing technologies, fine-tuning datasets have a small number of samples and weak domain relevance, resulting in low efficiency and poor quality when large models are applied in vertical fields.

Method used

A fine-tuned question-answering dataset is generated through a retrieval-enhancement-based method. Multiple rounds of optimization are performed using a large language model and vector library to generate domain-specific question-answer pairs, including file data preprocessing, seed dataset construction, multiple prompts for questions and answers, and similarity retrieval to ensure the accuracy and relevance of the data.

Benefits of technology

It improves the quality and efficiency of fine-tuning datasets, ensures the relevance of data to the field, makes it applicable to vertical fields, and improves the question-answering performance of large models in specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745855A_ABST
    Figure CN120745855A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large model fine tuning, and provides a retrieval enhancement-based fine tuning question and answer data set generation method, which comprises the following steps of: inputting file data, a seed data set and a question cue word of a single field into a large language model to obtain an optimization question; the seed data set is a question-answer pair generated based on file data; inputting the optimization problem, the file data, the seed data set and the first answer prompt word into a large language model to obtain an initial answer; determining original text information related to the optimization problem and the initial answer in the file data; inputting the original text information, the optimization question, the initial answer and the second answer cue word into a large language model to obtain an optimized answer; and determining a fine tuning question and answer data set according to a question and answer pair composed of the optimization question and the optimization answer. The method is simple, fast and low in cost, effectively solves the problems that the number of fine tuning data set samples is small and the field relevance is weak, well balances the requirements for efficiency and quality, and is suitable for the vertical field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model fine-tuning, and in particular to a method for generating a fine-tuning question-answering dataset based on retrieval enhancement. Background Art

[0002] Large models have demonstrated powerful capabilities in fields such as natural language processing and computer vision, and their development prospects are broad. With technological advancements, large models are increasingly being applied to vertical fields (such as healthcare, finance, and law) to solve specific problems. However, the application of large models in these vertical fields often requires fine-tuning to adapt to the needs of specific tasks.

[0003] Fine-tuning involves further training a large-scale pre-trained model using domain-specific data to better adapt the model to a specific task. The key to fine-tuning is producing a large number of appropriate samples that accurately reflect the characteristics of the target domain and the task requirements. However, producing these samples is often labor-intensive and time-consuming, especially in vertical fields where data is scarce or difficult to annotate. Summary of the Invention

[0004] The present invention provides a method for generating a fine-tuned question-answering dataset based on retrieval enhancement, which is used to solve the problems of small number of samples and weak domain relevance of fine-tuned datasets in the prior art.

[0005] The present invention provides a method for generating a fine-tuned question-answering dataset based on retrieval enhancement, comprising: inputting file data, a seed dataset, and question prompt words into a large language model to obtain an optimization question output by the large language model; wherein the seed dataset is a question-answer pair generated based on the file data, and the file data is data of a single domain; inputting the optimization question, the file data, the seed dataset, and a first answer prompt word into the large language model to obtain an initial answer output by the large language model; determining original text information related to the optimization question and the initial answer in the file data; inputting the original text information, the optimization question, the initial answer, and the second answer prompt word into the large language model to obtain an optimized answer output by the large language model; and determining a fine-tuned question-answering dataset based on the question-answer pair consisting of the optimization question and the optimized answer.

[0006] According to a retrieval-enhanced fine-tuning question-answering dataset generation method provided by the present invention, file data, a seed dataset, and question prompt words are input into a large language model to obtain an optimization problem output by the large language model, including: inputting file data, a seed dataset, and a first question prompt word into the large language model to obtain an initial question output by the large language model; and inputting the initial question, file data, a seed dataset, and a second question prompt word into the large language model to obtain an optimization problem output by the large language model.

[0007] According to a retrieval-enhanced fine-tuning question-answering dataset generation method provided by the present invention, file data, a seed dataset, and question prompt words are input into a large language model. Before obtaining the optimization problem output by the large language model, the method also includes: preprocessing the file including preset domain information to obtain file data; constructing a vector library based on the file data; and determining the seed dataset based on the file with preset domain information.

[0008] According to a retrieval-enhanced fine-tuning question-answering dataset generation method provided by the present invention, original text information related to the optimization question and the initial answer is determined in the document data, including: performing similarity search on the vectors composed of the optimization question and the initial question and answer in the vector library, and determining a set of paragraphs related to the vectors in the document data as the original text information.

[0009] According to a retrieval-enhanced fine-tuning question-answer dataset generation method provided by the present invention, a similarity search is performed on the vector composed of the optimization question and the initial question and answer in a vector library, and a set of paragraphs related to the vector in the file data is determined as the original text information, including: splicing the optimization question and the initial question and answer into a query text set; vectorizing the query text set to form a vectorized query text set; and performing a similarity search in the vector library to find the set of paragraphs most relevant to the vectorized query text set as the original text information.

[0010] According to a method for generating a fine-tuned question-answering dataset based on retrieval enhancement provided by the present invention, a file including preset domain information is preprocessed to obtain file data, including: converting a PDF file including preset domain information into file data in TXT format through data processing.

[0011] According to a retrieval-enhanced fine-tuning question-answering dataset generation method provided by the present invention, a PDF file including preset domain information is converted into file data in TXT format through data processing, including: converting the PDF file including preset domain information into a text collection in TXT format; converting the text collection into a sentence list; based on the Transformer pre-training model, encoding the sentence list to obtain a vector representation, and using the vectorized text as file data; wherein the file data is stored in a vector library.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement as described above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating a fine-tuned question-answering dataset based on retrieval enhancement.

[0014] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for generating a fine-tuned question-answering dataset based on retrieval enhancement.

[0015] The present invention provides a method for generating a fine-tuning question-answering dataset based on retrieval enhancement, comprising: inputting file data, a seed dataset, and a question prompt word into a large language model to obtain an optimization problem output by the large language model; wherein the seed dataset is a question-answer pair generated based on the file data, and the file data is data of a single field; inputting the optimization problem, the file data, the seed dataset, and the first answer prompt word into the large language model to obtain an initial answer output by the large language model; determining original text information related to the optimization problem and the initial answer in the file data; inputting the original text information, the optimization problem, the initial answer, and the second answer prompt word into the large language model to obtain an optimized answer output by the large language model; determining a fine-tuning question-answering dataset based on the question-answer pair consisting of the optimization problem and the optimized answer. Through the above-mentioned method, the present invention is simple, fast, and low-cost, and effectively solves the problems of small number of samples and weak domain relevance of the fine-tuning dataset. When expanding the data, it can ensure the relevance of the data to the field and maintain data consistency to further improve the quality of the fine-tuning dataset, better balance the requirements of efficiency and quality, make it suitable for vertical fields, and help the precise application of large models in various fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is one of the flow charts of the method for generating a fine-tuned question-answering dataset based on retrieval enhancement provided by an embodiment of the present invention.

[0018] Figure 2 This is the second flow chart of the method for generating a fine-tuned question-answering dataset based on retrieval enhancement provided by an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of a specific flow of data processing provided by an embodiment of the present invention.

[0020] Figure 4It is a schematic diagram of a specific process of generating an initial question provided by an embodiment of the present invention.

[0021] Figure 5 It is a schematic diagram of a specific process of generating an optimization problem provided by an embodiment of the present invention.

[0022] Figure 6 It is a schematic diagram of a specific process of generating an initial answer provided by an embodiment of the present invention.

[0023] Figure 7 It is a schematic diagram of a specific process of generating optimized answers provided by an embodiment of the present invention.

[0024] Figure 8 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] When fine-tuning large-scale pre-trained models, they require domain-specific data as training samples. When preparing these samples, the data should be carefully inspected and cleaned to ensure that the dataset contains no useless text or other noise. Each data sample in the dataset should have clear practical meaning so that the model can better understand its meaning. Furthermore, the size of the dataset can also affect the effectiveness of fine-tuning. The data samples should be sufficient to ensure that the model can accurately learn the format and patterns of the specific task. It is generally recommended that the number of texts in a dataset should be between a few thousand and several hundred thousand. Constructing a high-quality fine-tuning dataset can enable the fine-tuned model to perform better on the specific task.

[0027] However, creating these samples usually requires a lot of manpower and time, especially in vertical fields where data is scarce or difficult to label. Therefore, high-quality teaching methods are needed to expand data.

[0028] Data augmentation is a key technology in large-scale model fine-tuning. Data augmentation methods can include expert knowledge building, human-machine hybrid building, and model generation. Expert knowledge building relies on the knowledge and experience of experts to manually construct a dataset, forming a dataset by manually designing instructions and corresponding outputs. For example, in the medical field, medical experts can be asked to construct a dataset for fine-tuning instructions for large language models, such as designing an instruction for "explaining the symptoms of heart disease" and its detailed output. This method is particularly suitable for fields requiring a high degree of expertise and precision. It provides high-quality and accurate data and can be customized to specific needs. However, it also has the disadvantages of high cost, long time consumption, and limited dataset size.

[0029] Human-machine hybrid construction combines human creativity with machine efficiency. Initial data is generated using a large model, which is then manually screened and optimized. For example, when constructing a dataset for fine-tuning instructions for a tax scenario, a large language model can be used to generate a batch of initial instructions and outputs, which are then screened and refined by tax experts. This augmentation approach rapidly generates large amounts of data while ensuring data quality, reducing labor costs and time. However, since automatically generated data may be biased, it requires specialized knowledge and technical support.

[0030] To increase the complexity of fine-tuning data, it's also possible to generate more complex fine-tuning data using evolutionary methods based on existing data from a large model. Model generation involves leveraging a pre-trained large model to automatically generate a dataset using specific prompts or instructions. This approach is suitable for scenarios requiring large amounts of data and high data diversity. For example, when building datasets for natural language processing tasks, pre-trained models like GPT can be used to generate sample data for tasks like conversation and text classification. This approach can quickly generate large amounts of data with high data diversity. However, automatically generated data may contain noise and bias, requiring careful model tuning to ensure data quality.

[0031] Data augmentation methods often rely on expert knowledge and experience, using manual annotation or simple rules to generate fine-tuned datasets. These methods suffer from issues such as small size, insufficient diversity, high construction costs, and time-consuming nature. Model-generated data may contain quality flaws, including duplication, logical errors, and a lack of diversity, requiring subsequent screening and correction. Data augmentation methods also vary widely in their adaptability across different fields, making it difficult to develop a universal strategy.

[0032] Based on this, the present invention provides a simple and fast method for generating fine-tuned question-answering datasets based on retrieval enhancement, which can make up for the shortcomings of traditional data expansion methods, make full use of the advantages of large-scale knowledge bases and large models, and make the generated question-answering data more domain-specific and adapt to complex vertical domain environments, while taking into account data generation quality and efficiency.

[0033] See also Figure 1 , Figure 1 This is one of the flow charts of the method for generating a fine-tuned question-answering dataset based on retrieval enhancement provided by an embodiment of the present invention. In this embodiment, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement may include steps S110 to S150, each of which is specifically as follows: S110: Input the file data, seed data set and question prompt words into the large language model to obtain the optimization problem output by the large language model; wherein the seed data set is a question-answer pair generated based on the file data, and the file data is data in a single field.

[0034] In this step, the file data, seed dataset, and question prompts are fed into the large language model, allowing the model to generate optimization questions based on this information. The file data, which is single-domain data, makes the generated questions more targeted and specialized, focusing on key points in a specific field. The seed dataset provides the model with an initial sample of question-answer pairs, helping to guide the model in generating questions that better meet requirements and quality standards. Question prompts further clarify the direction and format of the questions, making the generated questions more standardized and more relevant to practical applications.

[0035] In some embodiments, the seed dataset may also be a small amount of high-quality question-answer pairs manually collected from various sources.

[0036] S120: Input the optimization problem, file data, seed data set, and first answer prompt word into the large language model to obtain an initial answer output by the large language model.

[0037] Based on the optimization problem generated in the previous step, the large language model is fed again, and a first-answer prompt is added so that the model can output the corresponding initial answer. The first-answer prompt clarifies the elements, format, and angle of the answer, making the initial answer more directional and usable, while also improving the efficiency and accuracy of the model's answer generation.

[0038] S130: Determine original text information related to the optimization question and the initial answer in the document data.

[0039] In this step, the original text related to the optimization question and initial answer must be found within the document data. This step retrieves and locates the original document data, ensuring the generated question-answer pairs are closely related to and accurate to the original data, avoiding the generation of inaccurate or fabricated answers. It also provides a reliable basis and reference for subsequent optimization of the answers, ensuring their credibility and authenticity.

[0040] S140: Input the original text information, the optimization question, the initial answer, and the second answer prompt word into the large language model to obtain the optimized answer output by the large language model.

[0041] Based on the initial answer and source text obtained in the previous steps, the second answer prompt is combined with the input of the large language model to further optimize the initial answer and generate the final optimized answer. The source text provides accurate context and basis for answer optimization. The optimization question clarifies the direction and focus of the answer. The initial answer serves as a foundation, while the second answer prompt helps the model further process and refine the answer, improving its content completeness, logic, accuracy, and language expression.

[0042] S150: Determine a fine-tuned question-answering dataset based on the question-answering pair consisting of the optimized question and the optimized answer.

[0043] Finally, we construct a fine-tuning Q&A dataset by combining the optimized questions and answers into question-answer pairs. These optimized Q&A pairs are of higher quality and better suited to the needs and characteristics of specific domains. They provide higher-quality and more accurate data support for subsequent fine-tuning of large language models, thereby improving the model's Q&A performance and effectiveness in specific domains.

[0044] For example, when a fine-tuned question-answering dataset in the medical field needs to be obtained, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement may include the following steps: 1) A large amount of medical literature, medical records, and other document data, as well as medical question-answer pairs initially generated based on this data, are used as a seed dataset. Professional prompts for medical questions are added and input into a large language model to obtain optimized medical diagnosis-related questions.

[0045] 2) Input the above optimization problem, medical document data, seed dataset, and first answer prompt word into the model to generate an initial answer. The initial answer can include diagnostic ideas and basic examination indicators for some common diseases.

[0046] 3) Retrieve original information such as symptom descriptions, examination methods, diagnostic criteria, etc. related to the diagnosis of the disease from medical document data.

[0047] 4) Input the original information, optimization question, initial answer, and second answer prompt into the model, and output the optimized answer, detailing the diagnostic ideas, identification points, selection of examination methods, and specific application of diagnostic criteria for various possible diseases.

[0048] 5) Collect the question-answer pairs consisting of the above optimization questions and optimized answers to form a fine-tuned question-answering dataset in the medical field, which can be used to fine-tune the large language model to enable it to have more accurate and professional question-answering capabilities in medical diagnosis.

[0049] The seed dataset is preferably a carefully selected dataset of doctor-patient Q&A questions and answers that matches the tone of patients' questions and doctors' professional responses. The purpose of using a seed dataset is to ensure that the generated Q&A pairs conform to the tone, voice, and language habits of different roles in vertical domain Q&A scenarios. The use of vertical domain literature knowledge is to make the generated answers more professional.

[0050] For example, when a fine-tuned question-answering dataset in the legal field needs to be obtained, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement may include the following steps: 1) Using legal documents, case analysis reports, and other document data as the seed dataset, we generate legal question-and-answer pairs based on these documents. We also add question prompts (e.g., "Please ask common questions regarding the application of this legal provision") and input them into a large language model to generate optimized legal questions, such as "Under what circumstances can the [specific content] stipulated in this legal provision be applied?"

[0051] 2) Input the optimization question, file data, seed dataset, and first answer prompt (e.g., "Please provide the basic conditions for the application of this clause based on legal provisions and relevant cases") into the model to obtain an initial answer. The initial answer may include a simple enumeration of the applicable circumstances of the clause.

[0052] 3) Find original information such as specific case analysis, legal interpretation, etc. related to the application of the legal clause in the document data.

[0053] 4) The model is fed with the original text, the optimization question, the initial answer, and the second answer prompt (e.g., "Please analyze in depth the various complex situations and exceptions to which this clause applies") to generate an optimized answer, detailing the various practical scenarios in which the clause applies, the handling of special circumstances, and its connection to related clauses.

[0054] 5) Collect the question-answer pairs consisting of the above optimization questions and optimized answers to form a fine-tuned question-answering dataset in the legal field, which will help improve the accuracy and professionalism of large language models in legal consultation and case analysis.

[0055] As described above, the embodiment of the present invention has achieved significant improvements in the accuracy, completeness, logic, and professionalism of the generated question and answer pairs through multiple rounds of optimization processes, including question optimization, generation and optimization of initial answers, and retrieval and reference based on original text information. This can better meet the actual needs of specific fields and provide a high-quality data foundation for model fine-tuning. The file data and seed dataset used are from a single domain, which can make the generated fine-tuned question and answer dataset focus on a single domain, helping the large language model to better adapt to the question and answer tasks in a specific domain after fine-tuning, and improve the professionalism and effectiveness of the model in that domain. For example, it can demonstrate more accurate and industry-standard question and answer capabilities in professional fields such as medicine, law, and finance. In addition, this embodiment also improves the reliability and credibility of the answers by retrieving original text information related to questions and answers from the file data and optimizing the initial answers based on this, thereby increasing the user's trust in the question and answer results of the fine-tuned model.

[0056] As described above, the embodiments of the present invention provide a method for generating a fine-tuned question-answering dataset based on retrieval enhancement, which can be simple, fast, and low-cost, and effectively solve the problems of small number of samples and weak domain relevance in the fine-tuning dataset. When expanding the data, it can ensure the relevance of data to the domain and maintain data consistency to further improve the quality of the fine-tuning dataset, and better balance the requirements of efficiency and quality, making it suitable for vertical fields and facilitating the precise application of large models in various fields.

[0057] Based on any of the above embodiments, the steps of inputting the file data, the seed data set, and the question prompt words into the large language model to obtain the optimization problem output by the large language model may specifically include: Input the file data, seed data set and the first question prompt word into the large language model to obtain the initial question output by the large language model; input the initial question, file data, seed data set and the second question prompt word into the large language model to obtain the optimization problem output by the large language model.

[0058] In this embodiment, the question prompt words include a first question prompt word and a second question prompt word. The first question prompt word is used to prompt the generation of an initial question, and the second question prompt word is used to prompt the generation of an optimization question.

[0059] Among them, the first question prompt word is used to guide the model's focus, clarify the question type and direction, and prevent the initial question from being too broad or off-topic; the second question prompt word contains more specific requirements, such as making the question more targeted, more in line with certain specific questioning rules, and more deeply exploring the key points in the file data. It is used to prompt the large language model to improve the initial question after comprehensively considering these factors, so as to improve its quality and better meet the needs of subsequent generation of high-quality question and answer pairs.

[0060] Therefore, in this embodiment, by utilizing the large language model twice and combining different prompts, we can gradually optimize the question, making it more precise and improving the quality of the question. Furthermore, the step-by-step approach of generating the initial question and then optimizing it helps the large language model construct questions more specifically, avoiding the complexity and difficulty of generating high-quality questions all at once and improving the efficiency of question generation.

[0061] Based on any of the above embodiments, the steps before inputting the file data, the seed data set, and the question prompt words into the large language model and obtaining the optimization problem output by the large language model may further include: Preprocess the file including the preset domain information to obtain file data; construct a vector library based on the file data; and determine a seed data set based on the file including the preset domain information.

[0062] Optionally, preprocessing may include formatting and standardizing the original file containing preset domain information, such as unifying text encoding, removing useless formatting symbols, and organizing paragraph structures, so as to make the file content neater and more standardized, and facilitate efficient reading and understanding by large language models.

[0063] Preprocessing can include content cleaning of original files containing preset domain information, such as removing noise information in the file, such as advertising content, irrelevant repeated information, incorrectly formatted content, etc., to avoid these invalid or interfering content from affecting the model's recognition and processing of key information, and improve the quality and availability of file data.

[0064] Furthermore, during preprocessing, the document content can be segmented into appropriate segments or chapters, facilitating more detailed analysis and processing by the model. For example, long academic documents can be segmented by chapter or topic, allowing the model to extract and utilize information more specifically.

[0065] This embodiment also constructs a vector library. This library converts text content in file data into vector form and stores it. This allows for rapid calculation of text similarities and quick location of the closest original text when searching for information relevant to a question within the file data. The vector library more precisely captures semantic similarity, improving search accuracy and efficiency and providing a more accurate basis for optimizing answers.

[0066] In addition, the vector library can better retain the contextual information of the text, so that during the retrieval process, it can not only find matching content based on a single keyword, but also take into account the semantic relationship between the contexts, thereby more accurately understanding the meaning and role of information fragments in the file data in the entire text, and further improving the model's utilization of file data.

[0067] In this embodiment, a seed dataset is also determined. The seed dataset is closely related to the target field to ensure that its question-answer pairs are generated for the actual problems and content in the field, providing a basic sample that is more relevant to the field for the optimization of subsequent questions and the generation of answers, avoiding interference from irrelevant fields, and improving the relevance and effectiveness of the entire question-answer dataset generation process.

[0068] Based on any of the above embodiments, the step of determining original text information related to the optimization question and the initial answer in the document data may specifically include: In the vector library, similarity retrieval is performed on the vectors composed of the optimization question and the initial question and answer, and a set of paragraphs related to the vectors in the file data is determined as the original text information.

[0069] In this embodiment, leveraging the efficient retrieval capabilities of the vector library and the semantics captured by vector representations, after converting the optimization question and initial question and answer into vectors, the similarity between these vectors and the text paragraph vectors stored in the vector library can be calculated to quickly find the set of paragraphs that are semantically closest to the question and answer. This vector similarity-based retrieval approach provides a deep understanding of the semantic characteristics of the text, going beyond superficial keyword matching. This allows for more accurate location of original textual information within the file data that is relevant to the current optimization question and initial answer, ensuring that the identified original textual information is closely related to the question and answer. This provides high-quality reference content for the subsequent generation of optimized answers, helping to improve the accuracy and rationality of the optimized answers.

[0070] Based on any of the above embodiments, the steps of performing similarity search on the vectors formed by the optimization question and the initial question and answer in the vector library and determining a set of paragraphs in the document data related to the vectors as the original text information may specifically include: The optimization question and the initial question and answer are spliced ​​into a query text set; the query text set is vectorized to form a vectorized query text set; a similarity search is performed in the vector library to find the paragraph set most relevant to the vectorized query text set as the original text information.

[0071] In this embodiment, the optimization question and the initial question and answer are combined into a query text set, forming a comprehensive text set. The optimization question represents the core point to be solved, while the initial question and answer provide the initial content and direction of the answer. By combining them together, the semantic characteristics of the question and answer can be more comprehensively captured.

[0072] The vector library stores vector representations of each text paragraph in the file data. Vectorization converts text into mathematical vectors, enabling quantitative analysis and comparison of text in a mathematical space. By vectorizing a query text set, the semantic features of the text are converted into numerical values, enabling mathematical methods (such as cosine similarity) to measure the similarity between texts.

[0073] By calculating the similarity between the vectorized query text set and the paragraph vectors in the vector library, we can quickly find the set of paragraphs that are semantically closest to the query text. This similarity search method not only considers keyword matching but also delves into semantic similarity, enabling more accurate location of original text information related to the question and answer, providing a reliable reference for optimizing answers.

[0074] Based on any of the above embodiments, the step of preprocessing the file including the preset domain information to obtain the file data may specifically include: Through data processing, the PDF file including the preset field information is converted into file data in TXT format.

[0075] This embodiment involves file format conversion and text content extraction. PDF files often contain complex layout information, embedded images, fonts, and other information, which is redundant for text analysis and processing. Therefore, this embodiment converts PDF files to TXT format through data processing, removing this irrelevant information and retaining only the text content, simplifying the subsequent text processing process.

[0076] Based on any of the above embodiments, the step of converting the PDF file including the preset domain information into file data in TXT format through data processing may specifically include: Convert a PDF file containing preset domain information into a text collection in TXT format; convert the text collection into a sentence list; based on the Transformer pre-trained model, encode the sentence list to obtain a vector representation, and use the vectorized text as file data; the file data is stored in a vector library.

[0077] In this embodiment, the text set is further converted into a sentence list. This step can use the text segmentation method in natural language processing technology to segment the continuous text into independent sentences.

[0078] Optionally, punctuation marks such as periods, question marks, and exclamation points can be used as sentence boundary markers. In some embodiments, special cases such as abbreviations and periods in numbers can also be considered to ensure accurate segmentation. The sentence list breaks down the text content into smaller semantic units, facilitating subsequent processing and analysis.

[0079] The Transformer pre-trained model, through pre-training on a large-scale corpus, learns the deep semantic features and contextual relationships of language. When a list of sentences is fed into the Transformer pre-trained model, it outputs a vector representation of each sentence. These vectors capture the semantic information of the sentence and map it into a high-dimensional vector space, so that semantically similar sentences have similar positions in the vector space. This provides rich semantic information for subsequent text processing, facilitating operations such as similarity comparison and classification.

[0080] To facilitate subsequent efficient retrieval and query, this embodiment will also facilitate subsequent efficient retrieval and query. The vector library is a data structure or system specifically used to store and retrieve vector data. It can quickly calculate the similarity between vectors and find the vector set most similar to the target vector.

[0081] Optionally, the vector library may be a Faiss, Annoy or other vector library, which is optimized through indexing and search algorithms to enable efficient storage and retrieval of large-scale vector data.

[0082] The method for generating a fine-tuned question-answering dataset based on retrieval enhancement described in an embodiment of the present invention is mainly a data expansion algorithm based on a large model, which is introduced in detail below.

[0083] See also Figure 2 , Figure 2 This is the second flow chart of the method for generating a fine-tuned question-answering dataset based on retrieval enhancement provided by an embodiment of the present invention.

[0084] This embodiment converts PDF text into a txt file and its vector library through data processing, and simultaneously constructs a small amount of high-quality seed data sets. The processed data, seed data set, and first prompt word are input into a large language model to obtain an initial question. The processed data, seed data set, initial question, and second prompt word are input into the large language model to obtain an optimized question. The processed data, seed data set, optimized question, and third prompt word are input into the large language model to obtain an initial answer. Finally, based on the optimized question and initial answer, the relevant paragraph position is searched in the vector library. The relevant paragraph, seed data set, optimized question, initial answer, and fourth prompt word are input into the large language model to obtain an optimized answer, thereby generating new question-answer pairs and completing high-quality data expansion.

[0085] This embodiment effectively utilizes retrieval technology and generative models, significantly improving the efficiency and domain relevance of generated data. Retrieval technology can quickly locate relevant document fragments, preventing redundant information in the article from affecting the generated question-answer pairs. This saves system costs while still ensuring the accuracy of the generated data. Through a multi-step generation and optimization process, a large language model is used to quickly generate high-quality, diverse question-answer data pairs, improving the quality and accuracy of the generated content and ensuring that the newly generated question-answer pairs are accurate, reliable, and domain-appropriate.

[0086] This embodiment may specifically include the following steps: Step A: Convert a large number of input domain documents (PDF format) into txt format text ,in, The text in txt is then converted into a vector using Embedding technology and stored in the vector library V for the model to calculate and process.

[0087] At the same time, a small number of appropriate questions and answers are manually created from the above PDF documents to form question-answer pairs as a high-quality seed dataset. .in, represents the kth manual question-answer pair.

[0088] Step B: Input the txt format text D, seed dataset S, and prompt word prompt1 into the large language model to obtain the initial question set .

[0089] in, represents the i-th initial question generated by the model.

[0090] Step C: Since the initial questions have problems such as the question-answer pair performance not being consistent with the domain knowledge and the expert tone, the initial question set obtained in step B needs to be Optimize. txt format text D, seed dataset S, initial question , prompt word prompt2 inputs the large language model to obtain the optimization problem set .in, represents the i-th optimization problem generated by the model.

[0091] Step D: Convert the txt format text D, seed dataset S, and optimization problem , prompt word prompt3 input large language model to obtain the initial answer set .

[0092] in, represents the i-th initial answer generated by the model.

[0093] Step E: Since the initial answer may be affected by redundant information in the article, resulting in inaccurate answers, the initial answer obtained in step D needs to be adjusted. First, optimize the optimization problem in the vector library V and the initial answer The vector is similarity searched to find the paragraph set P related to the vector. Then, the paragraph set P, seed dataset S, and optimization problem , initial answer , prompt word prompt4 inputs the large language model to obtain the optimized answer set .

[0094] in, Represents the i-th optimized answer generated by the model.

[0095] Optimization Problem and optimize answers The composed question-answer pairs can be constructed into a fine-tuning question-answer dataset, and finally the generated new question-answer pairs are output.

[0096] The above is the overall process of generating a retrieval-enhanced fine-tuned question-answering dataset. It mainly includes five sub-modules: data processing, initial question generation, optimized question generation, initial answer generation, and optimized answer generation. The following describes each of these sub-modules.

[0097] See also Figure 3 , Figure 3 It is a schematic diagram of a specific flow of data processing provided by an embodiment of the present invention.

[0098] Step A1: Input a large number of domain documents (PDF format), convert the PDF documents into plain text (txt format), and obtain a text collection ,in, Represents the nth document.

[0099] Step A2: Convert the text collection D into a list of sentences ,in Represents the sentence list of the i-th document, Represents the mth sentence in the i-th document.

[0100] Step A3: Use the transformers library to load a pre-trained model (such as BERT) and the corresponding tokenizer, and encode the sentence list Sentence to obtain a vector representation.

[0101] Step A4: Select the FAISS vector library, which can be used for efficient similarity search and clustering of dense vectors. Create an index, add vectors to the index, and store the vectorized text in the vector library V.

[0102] Step A5: Manually create a small number of question-answer pairs from domain documents (PDF format) as a high-quality seed dataset .in, represents the kth manual question-answer pair.

[0103] See also Figure 4 , Figure 4 It is a schematic diagram of a specific process of generating an initial question provided by an embodiment of the present invention.

[0104] Step B1: Input text D in txt format and seed dataset S.

[0105] Step B2: Take the seed dataset S as an example and construct the prompt word prompt1.

[0106] Step B3: Input the txt format text D in step B1, the seed dataset S, and the prompt word prompt1 in step B2 into the large language model.

[0107] Step B4: Generate an initial set of questions , .

[0108] in, represents the i-th initial question generated by the model.

[0109] See also Figure 5 , Figure 5 It is a schematic diagram of a specific process of generating an optimization problem provided by an embodiment of the present invention.

[0110] Step C1: Input text format D, seed dataset S, and initial question .

[0111] Step C2: Based on the seed dataset S and the initial question , construct the prompt word prompt2 to make the generated questions more in line with the vertical field.

[0112] Step C3: Convert the txt format text D, seed dataset S, and initial question in step C1 to , the prompt word prompt2 in step C2 is input into the large language model.

[0113] Step C4: Generate a set of optimization problems , .

[0114] in, represents the i-th optimization problem generated by the model.

[0115] See also Figure 6 , Figure 6 It is a schematic diagram of a specific process of generating an initial answer provided by an embodiment of the present invention.

[0116] Step D1: Input text format D, seed dataset S, and optimization problem .

[0117] Step D2: Based on the seed dataset S and the optimization problem , construct the prompt word prompt3 so that the large model can generate the corresponding answer.

[0118] Step D3: Convert the txt format text D, seed dataset S, and optimization problem in step D1 to , the prompt word prompt3 in step D2 is input into the large language model.

[0119] Step D4: Generate an initial answer set , .

[0120] in, represents the i-th initial answer generated by the model.

[0121] See also Figure 7 , Figure 7 It is a schematic diagram of a specific process of generating optimized answers provided by an embodiment of the present invention.

[0122] Step E1: Concatenate the optimization question and initial answer pairs into a query text set , .

[0123] in, Represents the query text composed of the i-th initial answer pair.

[0124] Step E2: Vectorize the query text , forming a vectorized query text set .

[0125] Step E3: Perform similarity search in the vector library V to find the vectorized query text set The most relevant paragraph set P locates the context scope in the relevant paragraphs, reducing the impact of redundant information in the article on the accuracy of the answer.

[0126] Step E4: Input relevant paragraph P, seed dataset S, and optimization problem , initial answer .

[0127] Step E5: Based on relevant paragraphs P, seed dataset S, and optimization problem , initial answer , construct the prompt word prompt4 so that the large model can generate the optimized answer.

[0128] Step E6: Combine the paragraph P, seed dataset S, and optimization problem in step E4 , initial answer , the prompt word prompt4 in step E5 is input into the large language model.

[0129] Step E7: Generate optimized answer set , .

[0130] in, Represents the i-th optimized answer generated by the model. Optimization problem and optimize answers The composed question-answer pairs can be constructed into a fine-tuning question-answering dataset.

[0131] As described above, this embodiment generates a fine-tuned question-answering dataset based on a large model, and uses a pre-trained large model to automatically generate a dataset through specific prompts or instructions to achieve data expansion. Compared with related data expansion methods, it fully utilizes the advantages of large models, greatly reduces the time cost of manual participation, and improves the efficiency, scale and diversity of data construction. This embodiment uses RAG technology to first calculate the similarity between an external knowledge base (such as a document library, database) and the question through vector retrieval, screens out highly correlated content, and then integrates the retrieval results into the generation module to assist the model in outputting more accurate answers and avoid the influence of redundant information in the article. Compared with directly using all text content, the quality of question and answer pairs is improved, especially the content related to professional domain knowledge, so that the generated question and answer data can be more domain-specific and adapt to complex vertical domain environments.

[0132] The following describes the device for generating a fine-tuned question and answer dataset based on retrieval enhancement provided by the present invention. The device for generating a fine-tuned question and answer dataset based on retrieval enhancement described below and the method for generating a fine-tuned question and answer dataset based on retrieval enhancement described above can refer to each other.

[0133] The device for generating a fine-tuned question-answering dataset based on retrieval enhancement includes a question generation module, an initial answer generation module, an original text information determination module, an optimized answer generation module and a dataset generation module.

[0134] The question generation module is used to input file data, a seed dataset, and question prompt words into the large language model to obtain the optimization problem output by the large language model; the seed dataset is generated based on the question-answer pairs of the file data, and the file data is data from a single domain; The initial answer generation module is used to input the optimization problem, file data, seed dataset, and first answer prompt word into the large language model to obtain the initial answer output by the large language model; An original information determination module, used to determine original information related to the optimization question and the initial answer in the file data; The optimized answer generation module is used to input the original information, optimization question, initial answer and second answer prompt word into the large language model to obtain the optimized answer output by the large language model; The dataset generation module is used to determine the fine-tuning question-answering dataset based on the question-answer pairs consisting of the optimization question and the optimization answer.

[0135] Based on any of the above embodiments, the question generation module can be specifically used to: input file data, seed data set and first question prompt words into the large language model to obtain the initial question output by the large language model; input the initial question, file data, seed data set and second question prompt words into the large language model to obtain the optimization problem output by the large language model.

[0136] Based on any of the above embodiments, the retrieval-enhanced fine-tuned question-answering dataset generation device also includes a data preprocessing module, which can be specifically used to: preprocess files including preset domain information to obtain file data; construct a vector library based on the file data; and determine a seed dataset based on the files with preset domain information.

[0137] Based on any of the above embodiments, the original information determination module can be specifically used to: perform similarity search on the vectors composed of the optimization question and the initial question and answer in the vector library, and determine a set of paragraphs related to the vector in the file data as the original information.

[0138] Based on any of the above embodiments, the original information determination module can also be specifically used to: splice the optimization question and the initial question and answer into a query text set; vectorize the query text set to form a vectorized query text set; perform similarity search in the vector library to find the paragraph set that is most relevant to the vectorized query text set as the original text information.

[0139] Based on any of the above embodiments, the data preprocessing module may be further configured to convert a PDF file including preset domain information into file data in TXT format through data processing.

[0140] Based on any of the above embodiments, the data preprocessing module can also be specifically used to: convert a PDF file including preset domain information into a text collection in TXT format; convert the text collection into a sentence list; based on the Transformer pre-training model, encode the sentence list to obtain a vector representation, and use the vectorized text as file data; wherein the file data is stored in a vector library.

[0141] On the other hand, an embodiment of the present invention further provides an electronic device, see Figure 8 , Figure 8 FIG is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention, such as Figure 8 As shown, the electronic device may include a memory 820, a processor 810, and a computer program stored in the memory 820 and executable on the processor 810. When the processor 810 executes the program, a method for generating a fine-tuned question-answering dataset based on retrieval enhancement may be implemented. The method may include: Input the file data, seed dataset and question prompt words into the large language model to obtain the optimization problem output by the large language model; the seed dataset is a question-answer pair generated based on the file data, and the file data is data from a single field; input the optimization problem, file data, seed dataset and first answer prompt words into the large language model to obtain the initial answer output by the large language model; determine the original text information related to the optimization problem and the initial answer in the file data; input the original text information, optimization problem, initial answer and second answer prompt words into the large language model to obtain the optimized answer output by the large language model; determine the fine-tuning question and answer dataset based on the question-answer pair consisting of the optimization problem and the optimized answer.

[0142] Optionally, the electronic device may further include a communication bus 830 and a communication interface (Communications Interface) 840, wherein the processor 810, the communication interface 840, and the memory 820 communicate with each other via the communication bus 830. The processor 810 may call the computer program in the memory 820 to execute the methods for generating a retrieval-enhanced fine-tuned question-answering dataset provided by the above methods.

[0143] Furthermore, the logic instructions in the aforementioned memory 820 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0144] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the retrieval-enhanced fine-tuning question and answer dataset generation method provided by the above methods. Its steps and principles have been introduced in detail in the above methods and will not be repeated here.

[0145] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the retrieval-enhanced fine-tuning question-answering dataset generation method provided by the above methods. Its steps and principles have been introduced in detail in the above methods and will not be repeated here.

[0146] The non-transitory computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid-state drives (SSDs)), etc.

[0147] As described above, the embodiment of the present invention provides a method for generating a fine-tuned question-answering dataset based on retrieval enhancement combined with a large model. It uses a pre-trained large model to automatically generate a dataset through specific prompts or instructions to achieve data expansion, and combines the advantages of retrieval enhancement technology and large models to create a high-quality question-answering dataset. At the same time, external knowledge sources are used to enhance the generation ability of the language model, improve the quality and accuracy of the generated content, and make the newly generated question and answer pairs consistent with the domain knowledge and the tone of domain experts. In addition, an answer optimization method based on a vector library is also used, including specific implementation methods such as vectorization processing, similarity measurement, and paragraph retrieval, so as to accurately match relevant paragraphs and eliminate the interference of irrelevant paragraphs on the answer to the question.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0149] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating a fine-tuned question answering dataset based on retrieval enhancement, characterized in that: include: Input the file data, the seed data set, and the question prompt words into the large language model to obtain an optimization problem output by the large language model; The seed dataset is a question-answer pair generated based on the file data, and the file data is data of a single field; Inputting the optimization problem, the file data, the seed data set, and the first answer prompt word into a large language model to obtain an initial answer output by the large language model; Determining textual information related to the optimization problem and the initial answer in the document data; Inputting the original text information, the optimization question, the initial answer, and the second answer prompt word into a large language model to obtain an optimized answer output by the large language model; A fine-tuned question-answering dataset is determined based on a question-answer pair consisting of the optimization question and the optimized answer.

2. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 1, characterized in that: The step of inputting the file data, the seed data set, and the question prompt words into the large language model to obtain the optimization problem of the large language model output includes: Inputting the file data, the seed data set, and the first question prompt word into a large language model to obtain an initial question output by the large language model; The initial question, the file data, the seed data set, and the second question prompt word are input into a large language model to obtain an optimization problem output by the large language model.

3. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 1, characterized in that: Before inputting the file data, the seed data set, and the question prompt words into the large language model to obtain the optimization problem output by the large language model, the process further includes: Preprocessing the file including the preset domain information to obtain the file data; Constructing a vector library according to the file data; The seed data set is determined according to the file of the preset domain information.

4. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 3, characterized in that: The determining of original text information related to the optimization problem and the initial answer in the file data includes: In the vector library, similarity search is performed on the vectors formed by the optimization problem and the initial question and answer, and a set of paragraphs related to the vectors in the document data is determined as the original text information.

5. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 4, characterized in that: The method of performing similarity search on the vectors formed by the optimization problem and the initial question and answer in the vector library and determining a set of paragraphs in the document data related to the vectors as the original text information includes: Splicing the optimization question and the initial question and answer into a query text set; Vectorizing the query text set to form a vectorized query text set; A similarity search is performed in the vector library to find a paragraph set that is most relevant to the vectorized query text set as the original text information.

6. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 3, characterized in that: The preprocessing of the file including the preset domain information to obtain the file data includes: Through data processing, the PDF file including the preset field information is converted into file data in TXT format.

7. The method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to claim 6, characterized in that: The data processing includes converting the PDF file including the preset field information into file data in TXT format, including: Convert the PDF file containing the preset field information into a text collection in TXT format; Converting the text collection into a list of sentences; Based on the Transformer pre-training model, the sentence list is encoded to obtain a vector representation, and the vectorized text is used as the file data; wherein the file data is stored in the vector library.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement is implemented as described in any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a fine-tuned question-answering dataset based on retrieval enhancement according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Question and answer processing method and device, computer equipment, readable storage medium and program product

    CN119066172A

  • Policy problem reply method, apparatus and device, and computer program product

    CN119202167A

  • Retrieval enhancement generation improvement method based on Embedded-FineTuning

    CN119621896A

  • Mobile phone retail store knowledge answer generation method based on LIama3 and retrieval enhancement

    CN119988533A

  • Automatic database enrichment and curation using large language models

    US20250045256A1

Cited By

  • Construction and evaluation method of manufacturing industry multi-mode question and answer data set

    CN121835908A