Automatic long text fine tuning instruction set construction method based on large language model

The high-quality fine-tuning instruction set is generated through recursive character segmentation method and self-guided learning method, which solves the problems of low fine-tuning data quality and insufficient instruction length in the existing technology, and realizes the efficient long text processing capability of large language models.

CN120181045APending Publication Date: 2025-06-20INST OF AUTOMATION CHINESE ACAD OF SCI +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510119646.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The quality of fine-tuning data generated in the prior art is uneven, making it difficult to ensure the performance of the model after fine-tuning, and the generated instruction length is short, which cannot meet the ability of large language models to obtain long context lengths.

Method used

Recursive character segmentation method is used to segment the input text to generate a paragraph set; then, based on the large language model, the preset question type set and prompt template are used to generate a question set and an answer set through a self-guided learning method; finally, based on the generated question set and answer set, an instruction set is generated, and the instruction set is optimized through multi-dimensional evaluation.

Benefits of technology

It ensures high quality of fine-tuning data, improves the model's long text processing capabilities, meets the ability of large language models to obtain long context lengths, and reduces the needs and costs of manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181045A_ABST
    Figure CN120181045A_ABST
Patent Text Reader

Abstract

The invention provides an automatic long text fine-tuning instruction set construction method based on a large language model, which comprises the following steps: acquiring an input text, and segmenting the input text by adopting a recursive character segmentation method to generate a paragraph set; aiming at the generated paragraph set, the large language model generates a question set and an answer set by adopting a self-guidance learning method through a preset question type set and a prompt template according to the task type; and generating an instruction set based on the generated question set and answer set, evaluating the quality of the instruction set in multiple dimensions, and optimizing the instruction set according to an evaluation result to obtain an optimized instruction set. According to the method, the high-quality long text fine tuning instruction set can be automatically generated, so that the performance of a large language model is improved, and meanwhile, the challenge of long context processing is solved. By automatically constructing the long text fine-tuning instruction set, the requirement of manual annotation is reduced, the cost is reduced, and meanwhile, the efficiency of the fine-tuning process and the long text processing capacity of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for constructing an automated long-text fine-tuning instruction set based on a large language model. Background Art

[0002] With the development of artificial intelligence technology, large language models (LLMs) have made remarkable progress in the field of natural language processing. These models can understand and generate natural language text through pre-training on large-scale datasets. However, despite the powerful language capabilities of pre-trained models, their performance in processing long contexts still has room for improvement. The ability to process long contexts is another challenge for large language models. There are many practical applications in daily production and life, such as reading scientific papers and analyzing book documents. Currently, the long-text understanding ability of closed-source large models is still weak, and in specific long-context practical application scenarios, the performance of open-source models is still significantly lower than that of closed-source large models.

[0003] To enable the model to acquire the ability to understand and process long-text data, researchers have proposed some long-context modeling techniques. Among them, using long-text pre-training data and long-text fine-tuning instruction sets in the pre-training and fine-tuning stages for training and fine-tuning is a key technology to improve the long-context ability of large language models.

[0004] Fine-tuning refers to further training on specific question-and-answer annotation data based on a pre-trained model, enabling the model to adapt to specific tasks and domains. Traditional fine-tuning methods rely on a large amount of manually annotated instruction data, which is not only time-consuming and laborious but also costly. The manually annotated instruction data may have problems such as bias and inaccuracy, affecting the performance of the model.

[0005] To improve the efficiency and effectiveness of fine-tuning, researchers have proposed some automated fine-tuning data generation methods. For example, some use data augmentation techniques, such as synonym replacement and sentence restructuring, to expand the training set. Other studies have explored techniques for generating a large amount of pseudo-parallel data guided by a small amount of annotated data. However, the quality of the fine-tuning data generated by these methods varies, making it difficult to guarantee the performance of the fine-tuned model, and the generated instructions are often short, unable to meet the needs of large language models for acquiring long-context lengths. Summary of the Invention

[0006] The present invention provides a method for constructing an automated long-text fine-tuning instruction set based on a large language model to solve the defects in the prior art that the quality of the generated fine-tuning data varies, it is difficult to guarantee the performance of the fine-tuned model, and the generated instructions are often short and cannot meet the needs of large language models for acquiring long-context lengths. The technical solutions proposed by the present invention are as follows: In a first aspect, the present invention provides a method for constructing an automated long text fine-tuning instruction set based on a large language model, including: Obtain the input text, and use a recursive character segmentation method to segment the input text to generate a paragraph set; For the generated paragraph set, the large language model generates a question set and an answer set by means of self-guided learning according to the task type through a preset question type set and prompt templates; Generate an instruction set based on the generated question set and answer set, evaluate the quality of the instruction set from multiple dimensions, and optimize the instruction set according to the evaluation results to obtain an optimized instruction set.

[0007] Optionally, the obtaining the input text and using a recursive character segmentation method to segment the input text to generate a paragraph set includes: Obtain the input text and a predefined separator set, and the separator set includes separators with different priorities; Obtain the text length of the input text. If the text length of the input text is greater than the target segmentation length, then select a separator from the predefined separator set according to the priority to segment the input text to obtain a cut text block; wherein, the selected separator is the first separator that makes the length of at least one cut text block not exceed the target segmentation length among all possible cutting methods; If the text length of the cut text block is greater than the target segmentation length, then reselect a separator to segment the cut text block until a paragraph set that meets the first condition is obtained.

[0008] Optionally, the question set includes a single paragraph retrieval task question set, and the answer set includes a single paragraph retrieval task answer set; for the generated paragraph set, the large language model generates a question set and an answer set by means of self-guided learning according to the task type through a preset question type set and prompt templates, including: If the task type is a single paragraph retrieval task, then perform the following operations on each text paragraph in the paragraph set to obtain a corresponding single paragraph retrieval task question set and a single paragraph retrieval task answer set: For each question type in the question type set, the large language model uses a pre-constructed question generation function to generate a corresponding single paragraph question based on the text paragraph and the question type to obtain a single paragraph retrieval task question set; wherein, the single paragraph question is generated based on a prompt template for generating the single paragraph question; For each single paragraph question in the single paragraph retrieval task question set, the large language model uses an answer generation function to generate a single paragraph answer based on the text paragraph and the single paragraph question to obtain a single paragraph retrieval task answer set.

[0009] Optionally, the question set includes a multi-paragraph retrieval task question set, and the answer set includes a multi-paragraph retrieval task answer set; for the generated paragraph set, the large language model uses a self-guided learning method to generate a question set and an answer set according to the task type through a preset question type set and prompt templates, including: If the task type is a multi-paragraph retrieval task, the following operations are performed on each multi-paragraph combination in the paragraph set to obtain the corresponding multi-paragraph retrieval task question set and multi-paragraph retrieval task answer set: For each question type in the question type set, the large language model uses a pre-constructed question generation function to generate corresponding multi-paragraph questions based on the multi-paragraph combination and the question type, obtaining a multi-paragraph retrieval task question set; among them, the multi-paragraph questions are generated based on the prompt templates for generating multi-paragraph questions. For each multi-paragraph question in the question set, the large language model uses a pre-constructed answer generation function to generate a multi-paragraph answer based on the multi-paragraph combination and the multi-paragraph question, obtaining a multi-paragraph retrieval task answer set.

[0010] Optionally, the answer set includes full-text-based answers; for the generated paragraph set, the large language model uses a self-guided learning method to generate a question set and an answer set according to the task type through a preset question type set and prompt templates, including: If the task type is a full-text understanding task, a task seed is determined, and the large language model performs text recursive full-text Q&A on the paragraph set according to the determined task seed and the preset prompt templates, obtaining questions and full-text-based answers; where each task seed corresponds to a question generation prompt template.

[0011] Optionally, the preset prompt templates include a question generation prompt template, a first answer prompt template, and a second answer prompt template; if the task type is a full-text understanding task, a task seed is determined, and the large language model performs text recursive full-text Q&A on the paragraph set according to the determined task seed and the preset prompt templates, obtaining questions and full-text-based answers, including: If the task type is a full-text understanding task, a task seed is randomly selected from a pre-constructed task seed pool. A text paragraph is extracted from the paragraph set, and the large language model uses the question generation prompt template corresponding to the selected task seed to generate a question for the extracted text paragraph. The large language model uses the pre-constructed first answer prompt template to carry the question to ask the first text paragraph in the paragraph set, obtaining an answer to the first text paragraph. The large language model uses a pre - constructed second answer prompt template to carry the question and the answer to the previous text paragraph to question the current text paragraph until all text paragraphs are traversed, and an answer based on the full text is obtained.

[0012] In a second aspect, the present invention also provides an apparatus for constructing an automated long - text fine - tuning instruction set based on a large language model, including the following modules: A text paragraph cutting module, configured to obtain the input text, and use a recursive character splitting method to split the input text to generate a paragraph set; An instruction question and answer generation module, configured to, for the generated paragraph set, the large language model generates a question set and an answer set by using a self - guided learning method according to the task type through a preset question type set and a prompt template; An instruction set evaluation module, configured to generate an instruction set based on the generated question set and answer set, evaluate the quality of the instruction set in multiple dimensions, and optimize the instruction set according to the evaluation result to obtain an optimized instruction set.

[0013] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the method for constructing an automated long - text fine - tuning instruction set based on a large language model as described in the first aspect above.

[0014] In a fourth aspect, the present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing an automated long - text fine - tuning instruction set based on a large language model as described in the first aspect above.

[0015] In a fifth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for constructing an automated long - text fine - tuning instruction set based on a large language model as described in the first aspect above.

[0016] Based on the above technical solutions, the beneficial effects of the present invention compared with the prior art are: The method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention utilizes the powerful generation capability of the large language model, combines a preset question type set and a prompt template Prompt, and generates high-quality questions and answers for the paragraph set obtained by cutting. Since the large language model has been trained on a large amount of text data, it can generate questions and answers that are closely related to the text content and logically reasonable, thereby ensuring the high quality of the fine-tuning data. After generating a preliminary instruction set, the method strictly screens and optimizes the instruction set through multi-dimensional evaluation. This step ensures that only high-quality instructions are retained for subsequent fine-tuning tasks, thereby further improving the quality of the fine-tuning data. The method uses a recursive character segmentation method to segment a long text and generate a paragraph set containing multiple text paragraphs. These text paragraphs have a longer context length than traditional short instructions, and can meet the ability requirements of the large language model to obtain a long context length. On the basis of generating question-answer pairs for the paragraph set, the method further integrates these question-answer pairs into an instruction set. Since each text paragraph contains relatively long context information, the generated instruction set also has a longer length and richer context information accordingly. This method fully utilizes the advantages of large language models in processing long texts and complex contexts. The paragraph set generated by the recursive character segmentation method, as well as the question-answer pairs and instruction sets generated for the paragraph set, fully consider the processing capability requirements of large language models for long context lengths.

[0017] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 It is a flow chart of the method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention.

[0021] Figure 2It is a schematic diagram of the overall solution of the method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention.

[0022] Figure 3 It is a schematic diagram of the problem generation solution for the full text understanding task provided by the present invention.

[0023] Figure 4 It is a schematic diagram of the answer generation solution for the full text understanding task provided by the present invention.

[0024] Figure 5 It is a schematic diagram of the structure of the device for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention.

[0025] Figure 6 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0027] The following combines Figures 1-4 to describe the method for constructing an automated long text fine-tuning instruction set based on a large language model of the present invention.

[0028] The method for constructing an automated long text fine-tuning instruction set based on a large language model processes long texts using a large language model through automated means to generate a high-quality fine-tuning instruction set, so as to improve the performance of the large language model and at the same time solve the challenges of long context processing. By automatically constructing a long text fine-tuning instruction set, the method of the present invention reduces the need for manual annotation, reduces costs, and at the same time improves the efficiency of the fine-tuning process and the long text processing ability of the model. Referring to Figure 1 as shown, the method includes the following: Step S110: Obtain the input text, and segment the input text using the recursive character segmentation method to generate a paragraph set; wherein, the text block set includes a plurality of segmented text blocks.

[0029] First, receive a long text as input from a user or a data source. Then, use the recursive character segmentation method (or other suitable text segmentation algorithms) to segment the input text. This method recursively divides the text into smaller units until a predetermined text paragraph size is reached or other segmentation conditions are met. After segmentation, a paragraph set containing multiple text paragraphs is generated, and these text paragraphs will be used as the input for subsequent processing. Through the recursive character segmentation method, long texts can be effectively divided into more manageable text paragraphs, facilitating the generation of question-answer pairs and instruction sets in the subsequent steps. The segmented paragraph set has higher flexibility and scalability and can be adjusted and optimized according to actual needs.

[0030] Step S120: For the generated paragraph set, the large language model uses the self-guided learning method to generate a question set and an answer set according to the task type through a preset question type set and prompt templates.

[0031] For the generated paragraph set, the large language model processes it according to the preset task types (such as information extraction, text summarization, sentiment analysis, etc.) and question type sets (such as "What is the main idea of this article?", "What is the author's opinion?", etc.). Use the prompt template Prompt to guide the large language model to generate questions related to the content of the text paragraphs. The design of Prompt should be concise and clear, capable of stimulating the model's ability to generate high-quality questions. The large language model uses the self-guided learning method to find answers in the text paragraphs based on the generated questions and generates corresponding question-answer pairs.

[0032] Through the preset question type set and the prompt template Prompt, the large language model can be guided to generate questions closely related to the task type, improving the pertinence and effectiveness of the questions. The self-guided learning method enables the large language model to find answers in the text paragraphs based on the generated questions without external supervision, thus generating high-quality question-answer pairs.

[0033] Step S130: Generate an instruction set based on the generated question set and answer set, evaluate the quality of the instruction set from multiple dimensions, and optimize the instruction set according to the evaluation results to obtain an optimized instruction set.

[0034] Using the question set and answer set generated by S120, construct a preliminary instruction set. Each instruction contains a question, the corresponding answer, and any text passages that may be required (for multi-paragraph or full-text understanding tasks). The design of the instruction set should ensure that it covers the main content and key information points of the text. At the same time, the diversity and complexity of the questions should be appropriate to adapt to different levels of understanding tasks. Specifically, from the question set and answer set generated by S120, remove duplicates, correct formatting errors, and handle any possible text encoding issues. Find the corresponding answer for each question and verify the matching relationship between them. This can be done through string matching or more complex algorithms to identify the semantic relationship between the question and the answer. For instructions that require text passage support, use text matching algorithms or NLP techniques to extract relevant passages from the original text. Format the generated instruction set into structured data suitable for subsequent processing. Such as creating a JSON or XML file, where each instruction contains the question, the answer, and the relevant text passages.

[0035] Evaluate the instruction set through multiple dimensions including question uniqueness (U), question relevance (R), answer accuracy (A), answer clarity (C), and evidence integrity (E), with each dimension having a full score of 5 points. Question uniqueness (U) refers to evaluating whether the question is novel, avoiding repetitive questions, and ensuring the diversity and richness of the instruction set. Question relevance (R) refers to checking the relevance of the question to the text content, whether the question is highly relevant to the text content, and ensuring that the question can accurately reflect the key information of the text. Answer accuracy (A) refers to verifying whether the answer is based on the text content and is accurate, ensuring the validity of the instruction set. Answer clarity (C) refers to evaluating whether the answer is concise and easy to understand, avoiding long or ambiguous descriptions. Evidence integrity (E) refers to for instructions that require evidence, checking whether the evidence fully supports the answer, ensuring the reliability of the instruction set.

[0036] The constructed Chinese Prompt is as follows: [INST] You are a QA judge. I will provide you with a piece of original text and questions, answers, and evidence generated by AI based on this original text. This original text is excerpted from a long text. Eventually, the long text and the questions will together form an instruction set for model fine-tuning. Please answer the relevant questions based on the original text content, and based on the answers you give, judge: 1. Answer uniqueness; When the generated questions come from a certain paragraph, but when the full text is given, is the answer to the question the same as when only a certain paragraph is given? 2. Question relevance; Are the generated questions relevant to the original text content? 3. Answer accuracy; Is the generated answer consistent with the original text content? 4. Clarity of answer; Is the generated answer clear and easy to understand? Is it the minimum description that can answer the question? 5. Completeness of answer; Does the generated answer cover all the content? Please give your judgment result according to the judgment basis. The judgment result is based on a full score of 5 points, and the score for each item is an integer. You can also give any evaluation. Your evaluation will be used to train the model. Please generate it in the following format: { "Uniqueness of answer": n, "Relevance of question": n, "Accuracy of answer": n, "Clarity of answer": n, "Completeness of answer": n, "Overall score": n, "QA evaluation": "Here is your evaluation" }[ / INST] The QA judge is responsible for judging based on the given original text, question, and answer. The QA judge needs to understand the goal of the task, that is, to evaluate the quality of the QA pairs generated by AI, and give scores and evaluations according to multiple dimensions such as question uniqueness (U), question relevance (R), answer accuracy (A), answer clarity (C), and evidence completeness (E).

[0037] The QA judge receives a piece of original text, a question, and an answer as input for judgment. The QA judge first reads the original text to understand its content and structure. Analyze each question one by one to judge its relevance to the original text. For each question, evaluate the accuracy, clarity, and completeness of the answer. If evidence is provided, the degree of support of the evidence also needs to be checked. According to the judgment criteria, give a score from 1 to 5 for each aspect. Fill in the judgment result in the specified format, including the scores for each judgment criterion, the overall score, and the QA evaluation. The QA judge can provide a more detailed evaluation, pointing out the advantages and disadvantages of the QA pair, as well as possible improvement suggestions. Submit the judgment result to the system or relevant personnel for subsequent analysis and improvement. The QA judge will receive feedback on their judgment result, which helps to improve the judgment process and the quality of judgment. Based on the feedback, the QA judge can adjust their judgment criteria and methods to better meet the task requirements.

[0038] The evaluation can be carried out through manual review, automated testing, or a combination of both. Manual review can ensure the accuracy and depth of the evaluation, while automated testing can improve the evaluation efficiency and consistency. Automated testing can: Use NLP techniques to compare the answer with the actual information in the text to ensure the accuracy of the answer. Write rules or use machine learning models to detect whether the answer is too verbose or ambiguous. For instructions that require evidence, use NLP techniques to verify whether the evidence sufficiently supports the answer.

[0039] Based on the evaluation results, optimize the instruction set. The optimization measures include modifying the expression of the questions to improve clarity and accuracy, adjusting the order or priority of the instructions to adapt to different training needs, adding or deleting certain instructions to improve the coverage and difficulty distribution of the instruction set. The optimized instruction set will be more in line with the actual needs and application scenarios, improving the efficiency and effect of the fine-tuning task.

[0040] Optimize the question expression, such as rephrasing the question using more concise and clear language. Adjust the structure of the question to make it easier to understand and answer. Instruction set adjustment includes difficulty analysis, relevance sorting, adding / deleting instructions. Difficulty analysis refers to analyzing the difficulty of the questions to ensure that the instruction set can cover tasks of different difficulty levels. Relevance sorting refers to sorting the instructions according to the relevance of the questions to the text content to ensure that the most important instructions are processed first. Adding / deleting instructions refers to performing redundancy checks, identifying and deleting duplicate instructions. Conduct missingness analysis, analyze the coverage of the instruction set to ensure that no key information or tasks are omitted. Conduct diversity improvement, add new instructions to cover a wider range of topics and scenarios. Re-evaluate the optimized instruction set to verify the improvement effect. Based on the evaluation results and expert feedback, conduct further optimization.

[0041] Through the multi-dimensional evaluation and optimization process, it can be ensured that the generated instruction set has a high degree of clarity, accuracy, and executability. This helps to improve the efficiency and effect of the fine-tuning task, enabling the model to better understand and process text content. The diversity and richness of the instruction set help the model to exhibit better generalization ability when dealing with different types and difficulties of text. By training the model to generate accurate answers under various instructions, the model's understanding and processing ability for unseen text can be improved. The optimized instruction set is more in line with the actual needs and application scenarios, and can provide a better experience and effect for users. For example, in a question-answering system or a reading comprehension application, a high-quality instruction set can ensure that the system can accurately answer users' questions, improving user satisfaction and trust. As an important input for model training, the quality and effect of the instruction set directly affect the training effect and performance of the model. By continuously optimizing the instruction set, the training iteration and performance improvement of the model can be promoted, driving the development and application of natural language processing technology.

[0042] The method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention utilizes the powerful generation capability of the large language model, combines a preset question type set and a prompt template Prompt, and generates high-quality questions and answers for the paragraph set obtained by cutting. Since the large language model has been trained on a large amount of text data, it can generate questions and answers that are closely related to the text content and logically reasonable, thereby ensuring the high quality of the fine-tuning data. After generating a preliminary instruction set, the method strictly screens and optimizes the instruction set through multi-dimensional evaluation. This step ensures that only high-quality instructions are retained for subsequent fine-tuning tasks, thereby further improving the quality of the fine-tuning data. The method uses a recursive character segmentation method to segment a long text and generate a paragraph set containing multiple text paragraphs. These text paragraphs have a longer context length than traditional short instructions, and can meet the ability requirements of the large language model to obtain a long context length. On the basis of generating question-answer pairs for the paragraph set, the method further integrates these question-answer pairs into an instruction set. Since each text paragraph contains relatively long context information, the generated instruction set also has a longer length and richer context information accordingly. This method fully utilizes the advantages of large language models in processing long texts and complex contexts. The paragraph set generated by the recursive character segmentation method, as well as the question-answer pairs and instruction sets generated for the paragraph set, fully consider the processing capability requirements of the large language model for long context lengths. In the process of generating the instruction set, the method also ensures that the instruction set is not only of high quality, but also has good adaptability and scalability through a multi-dimensional evaluation and optimization process. This means that the generated instruction set can be flexibly applied to various natural language processing tasks to meet the needs of different scenarios.

[0043] In an optional embodiment, considering that the current large model has limited long-context understanding and processing capabilities, long texts often exceed the context window of the model. Directly using overly long text corpora to generate question-answer pairs may cause inaccurate generation instructions. Therefore, the goal of the text paragraph cutting module is to generate text paragraphs (chunks) of moderate length based on semantic integrity, so as to improve the accuracy of subsequent question-answer generation and processing. The above step S110 obtains input text, uses a recursive character segmentation method to segment the input text, and generates a paragraph set, including: S1101. Obtain input text and a predefined delimiter set, where the delimiter set includes delimiters of different priorities.

[0044] The above input text is a long text corpus to be processed, such as data in long books, papers, code, etc. The delimiter set is a set containing delimiters with different priorities. These delimiters can be punctuation marks (such as full stops, commas, semicolons), spaces, line breaks, etc., depending on the application scenario. The priority of the delimiters determines their priority order when splitting the text.

[0045] S1102. Obtain the text length of the input text. If the text length of the input text is greater than the target segmentation length, then select a delimiter from the predefined delimiter set according to the priority to split the input text, obtaining the cut text blocks; wherein, the selected delimiter is the first delimiter among all possible cutting methods that makes the length of at least one of the cut text blocks not exceed the target segmentation length. If the text length of the cut text block is greater than the target segmentation length, then reselect a delimiter to split the cut text block until a paragraph set that meets the first condition is obtained.

[0046] The above target segmentation length L is a preset value used to determine the ideal length of the split text blocks (paragraphs). Obtain the text length of the input text, and its length may exceed the target segmentation length . If the text length is less than or equal to the target segmentation length, then no further splitting is required and it can be directly used as an element of the text block set. If the text length is greater than the target segmentation length, then select a delimiter from the delimiter set according to the priority to split the input text.

[0047] The selected delimiter should meet the following conditions: among all possible cutting methods, it is the first delimiter that makes the length of at least one of the cut text blocks not exceed the target segmentation length. This means selecting the delimiter that causes the part with the smallest text block length exceeding the target segmentation length to be split, but as close to the start of the text as possible to ensure the flexibility of subsequent splitting.

[0048] For each cut text block, repeat the above steps (starting from judging the text length) until the text lengths of all the cut paragraphs are close to the target segmentation length. If a cut text block still exceeds the target segmentation length, then reselect a delimiter (considering the available delimiters and their priorities within the current text block) to split it until the first condition is met.

[0049] The recursive segmentation of the present invention can flexibly divide long texts into multiple shorter and more manageable text paragraphs, suitable for various text processing requirements. Appropriate text segmentation can improve the readability of the text. Especially when reading long texts, dividing the text into paragraphs can make the content clearer and easier to understand. The present invention uses multiple delimiters and sets priorities, enabling the segmentation strategy to be adjusted according to specific application scenarios, improving the generality and adaptability of the method. By processing each text block recursively, it ensures that all text blocks meet the length requirements, avoiding the problem of text blocks being too long or too short due to improper single segmentation.

[0050] Specifically, the recursive character segmentation method with complete semantics is selected to achieve a balance between semantic integrity and implementation cost. This method ensures that the semantic integrity is not damaged as much as possible by gradually trying to cut the paragraph with different delimiters. The complete text paragraph cutting process is as follows: First, define the delimiter set , and these delimiters are sorted according to priority. For example, paragraph delimiters (such as consecutive line breaks \n\n), sentence delimiters (such as.,!,?), space delimiters (such as a single space), etc.

[0051] Second, for the input text T, if its text length (the target segmentation length), then start to cut one by one according to the priority of the delimiter set . The recursive cutting process is defined as follows: Let the target segmentation length be L, and the delimiter set be: In the formula, represents the number of delimiters in the delimiter set .

[0052] Select the delimiter with the highest priority from the delimiter set for trial. Use this delimiter to cut the text , and obtain the set of cut text blocks . Among all possible cutting methods, select the first cutting method that makes at least one of the cut text blocks not exceed the target segmentation length . represents the index of the delimiter with the minimum priority that meets the conditions, represents the cutting operation.

[0053] For each text block in the set of cut text blocks , if its text length is still greater than the target segmentation length , then the above steps are recursively executed for this text block . That is, further cutting is performed using the delimiter set . The recursive process continues until the lengths of all the cut text paragraphs are close to the target segmentation length . Finally, the generated paragraph set satisfies the first condition: the length of each text paragraph is close to the target segmentation length , and is within the range . represents the number of text paragraphs in the paragraph set .

[0054] To determine whether it is close to , the specific method can be to determine whether the absolute value of the difference between and the target segmentation length is less than a preset threshold.

[0055] During the cutting process, the semantic integrity of the text should be maintained as much as possible. For example, avoid cutting in the middle of a sentence or phrase. In practical applications, the priority of the delimiter can be adjusted according to the characteristics and requirements of the text. For example, when processing dialogue text, the priority of the line break character can be increased to better preserve the integrity of the dialogue. Through the recursive character segmentation method with complete semantics, the long text can be segmented into multiple text paragraphs that are easy to process and read while maintaining the semantic integrity of the text.

[0056] In an optional embodiment, as shown in Figure 2 , the task types of the above step S120 include single-paragraph retrieval tasks, multi-paragraph retrieval tasks, and full-text understanding tasks. For single-paragraph and multi-paragraph retrieval tasks, the present invention uses the self-instruct method to generate questions for each of the cut text paragraphs. Through the preset question types and Prompt templates, the large language model can be guided to generate structured and hierarchical high-quality single-paragraph, multi-paragraph retrieval, and understanding questions.

[0057] For the generated paragraph set described in the above step S120, the large language model generates a question set and an answer set according to the task type through the preset question type set and prompt template, including: If the task type is a single paragraph retrieval task, the following operations are performed for each text paragraph in the paragraph set to obtain the corresponding single paragraph retrieval task question set and single paragraph retrieval task answer set: S210, for each question type in the question type set, the large language model generates a corresponding single paragraph question based on the text paragraph and the question type using a pre-built question generation function, thereby obtaining a single paragraph retrieval task question set; wherein the single paragraph question is generated based on a prompt template for generating a single paragraph question; First, initialize the text paragraph Given. Question type set Include Different question types. Generate question function and generate the answer function Pre-built and trained.

[0058] For the question type set Each question type in (in =1,2,…, ): Use the Generate Question function , enter the text paragraph and question type , generating the corresponding question, here called a single paragraph question . Generate the problem function The Chinese prompt for generating single-paragraph questions provides clear instructions for writing a question related to the paragraph content based on the text fragment and providing the correct answer and original evidence. Aggregate into a single paragraph retrieval task problem set .

[0059] S220. For each single paragraph question in the single paragraph retrieval task question set, the large language model generates a single paragraph answer based on the text paragraph and the single paragraph question using a function for generating answers, thereby obtaining a single paragraph retrieval task answer set.

[0060] For the single paragraph retrieval task question set Each single paragraph question in : Use the function to generate the answer , enter the text paragraph and single paragraph questions , generate the corresponding answer, here called a single paragraph answer . The single-paragraph answer here It must be based on the text content, and corresponding evidence should be provided from the original text to support the correctness of the answer. All single-paragraph answers generated will be aggregated into a single-paragraph retrieval task answer set .

[0061] For each text paragraph , a corresponding single-paragraph retrieval task question set and a single-paragraph retrieval task answer set will be generated. Denote any single-paragraph question in the single-paragraph retrieval task question set .

[0062] Through the automated generation of questions and answers, the present invention can greatly accelerate the construction speed of the dataset, especially for large-scale text collections. Different styles and difficulties of questions and answers can be generated by adjusting the question type set and the parameters of the functions for generating questions or answers to adapt to different application scenarios. Through the carefully designed Prompt and well-trained generation functions, high-quality questions and answers closely related to the text content can be generated. Compared with manual annotation, using large language models to generate questions and answers can significantly reduce the cost of data annotation. The generated single-paragraph retrieval task dataset can be used to train and evaluate reading comprehension models, promoting research and development in related fields.

[0063] Among them, the Chinese Prompt for generating single-paragraph questions is constructed as follows: [INST] You will be given a fragment intercepted from a long text. Please write a question related to the paragraph content and provide the correct answer. The answer must be based on the text content. At the same time, please provide corresponding evidence from the original text to prove that your answer is correct. This question will be used as a reading comprehension test for the whole document. Please wrap the question and answer with XML tags ( <question>and< / question> , <answer>and < / answer>, <evidence>and< / evidence> ).[ / INST] In an optional embodiment, the question set includes a multi-paragraph retrieval task question set, and the answer set includes a multi-paragraph retrieval task answer set. The multi-paragraph retrieval task refers to, in a given set of paragraphs, for each multi-paragraph combination, according to a preset set of question types and prompt templates, generating a corresponding question set and answer set. For the generated set of paragraphs described in step S120 above, the large language model generates a question set and an answer set by means of self-guided learning according to the task type through the preset set of question types and prompt templates, including: If the task type is a multi-paragraph retrieval task, then the following operations are performed for each multi-paragraph combination in the set of paragraphs to obtain a corresponding multi-paragraph retrieval task question set and multi-paragraph retrieval task answer set: S310. For each question type in the set of question types, the large language model uses a pre-constructed question generation function to generate a corresponding multi-paragraph question based on the multi-paragraph combination and the question type, obtaining a multi-paragraph retrieval task question set; wherein, the multi-paragraph question is generated based on a prompt template Prompt for generating multi-paragraph questions.

[0064] The preset set of question types includes various possible question types, such as detail understanding, main idea generalization, inference and judgment, etc. These types of questions can comprehensively cover the requirements of the multi-paragraph retrieval task.

[0065] The large language model uses a pre-constructed question generation function, which can generate a corresponding multi-paragraph question according to the multi-paragraph combination and the question type. When generating a multi-paragraph question, it is based on a prompt template Prompt for generating multi-paragraph questions. This template provides the basic structure and guiding information of the question, helping the large language model generate questions that meet the requirements. For each question type in the set of question types, the large language model will use the question generation function and the prompt template Prompt for generating multi-paragraph questions to generate corresponding questions based on the multi-paragraph combination. In this way, a multi-paragraph retrieval task question set containing various question types can be obtained.

[0066] S320. For each multi-paragraph question in the question set, the large language model uses a pre-constructed answer generation function to generate a multi-paragraph answer based on the multi-paragraph combination and the multi-paragraph question, obtaining a multi-paragraph retrieval task answer set.

[0067] The large language model uses a pre-constructed answer generation function , the function can generate corresponding answers according to multi-paragraph combinations and multi-paragraph questions. For each multi-paragraph question in the question set, the large language model will use the answer generation function to generate corresponding multi-paragraph answers based on the multi-paragraph combination. . In this way, a set of multi-paragraph retrieval task answers corresponding to the question set can be obtained. .

[0068] Let the multi-paragraph combination be , where represents a combination of several paragraphs selected from the paragraph set . are all paragraphs in the multi-paragraph combination , is the number of paragraphs in the multi-paragraph combination . The total number of optional combinations represents how many ways there are to select different paragraph combinations from . That is: In the formula, represents the multi-paragraph retrieval task question set, represents the question generation function, represents the question type, represents the question type set, represents the multi-paragraph retrieval task answer set, represents any multi-paragraph answer in the multi-paragraph retrieval task answer set, represents the multi-paragraph retrieval task question set in any multi-paragraph question.

[0069] For each multi-paragraph combination , the following tasks need to be performed to generate comprehensive questions and extract answers: First, construct a Chinese Prompt for generating multi-paragraph questions. The construction of the Chinese Prompt aims to guide the large language model (LLM) to integrate multi-paragraph information, generate a comprehensive question (i.e., the above multi-paragraph question), and require the answer to be unique and based on the text fragment. The Prompt example is as follows: [INST] You will be given some fragments intercepted from a long text, which are from different parts of the article. After these text fragments, there are questions, answers, and original text evidence generated based on these paragraphs. Please write a comprehensive question that integrates the information from each paragraph. Ensure that the newly generated question does not contain independent sub-questions and that the answer is as unique as possible. The generated answer must be based on these text fragments. At the same time, provide corresponding evidence from the original text to prove that your answer is correct. This question will be used as a reading comprehension test for the entire document. Please wrap the question and answer with XML tags <question>and< / question> , <answer>and< / answer> , <evidence>and< / evidence> ).[ / INST] Secondly, select a specific multi-paragraph combination from the set of paragraphs . .

[0070] Then, based on the selected multi-paragraph combination and the Chinese Prompt, the LLM generates a comprehensive question, namely the above multi-paragraph question. This question needs to integrate all the information in the paragraphs, ensuring that the question does not contain independent sub-questions and that the answer is as unique as possible.

[0071] Finally, the LLM extracts the answer, namely the above multi-paragraph answer, based on the generated multi-paragraph question and the multi-paragraph combination , and finds the evidence in the original paragraphs to support the answer. The answer and evidence are respectively <answer>And <evidence>Label packaging.

[0072] For each multi-paragraph combination , the LLM outputs the following results: multi-paragraph question , multi-paragraph answer and original text evidence. The multi-paragraph question uses <question>Label packaging. Multi-paragraph answer With <answer>Label packaging. The original evidence is <evidence>Label packaging.

[0073] Through the self-guided learning method of the large language model, the present invention can generate corresponding question and answer sets for each multi-paragraph combination. This process has a high degree of automation and can significantly improve the efficiency of multi-paragraph retrieval tasks. The large language model has powerful language understanding and reasoning abilities and can generate accurate questions and answers based on multi-paragraph combinations. This helps to more accurately locate the required information in multi-paragraph retrieval tasks and improve the accuracy of retrieval. Through the preset question type set and prompt templates, various types of questions can be generated, thus covering different requirements in multi-paragraph retrieval tasks. This helps to generate a more comprehensive and diverse answer set to meet the diverse needs of users. The large language model has learned a large amount of language knowledge and reasoning rules during the training process, which enables it to adapt to different tasks and scenarios. In multi-paragraph retrieval tasks, the large language model can generate question and answer sets that meet the requirements based on the given paragraph set and question types, demonstrating strong generalization ability.

[0074] In an optional embodiment, for the generated paragraph set in step S120 above, the large language model generates a question set and an answer set by means of a self-guided learning method according to the task type through the preset question type set and prompt templates, including: S410. If the task type is a full-text understanding task, determine the task seed, and the large language model performs text recursive full-text Q&A on the paragraph set according to the determined task seed and the preset prompt template to obtain questions and full-text-based answers; where each task seed corresponds to a question generation prompt template.

[0075] Specifically, first, select one from the preset task seed types, and these types include full-text information extraction and statistics, theme induction and analysis, sentiment analysis and attitude evaluation, speculation and prediction, etc. Each task seed corresponds to a specific question generation prompt template, which guides the LLM on how to generate questions according to the selected task seed. Secondly, randomly select one or more consecutive paragraphs from the given paragraph set as the starting point for analysis. This step ensures the diversity and randomness of the questions. According to the selected segment and task seed, the LLM uses the corresponding question generation prompt template to construct a targeted question. When constructing the question, the LLM will follow the guidance in the prompt template to ensure that the question comprehensively considers the entire text as the context and avoids using piecemeal descriptive attributives. For the generated questions, the LLM will generate an answer based on the full-text content. This answer needs to cover the relevant information in the full text, rather than just the extracted segment. The LLM's answer should be accurate, comprehensive, and closely related to the question.

[0076] The above full-text information extraction and statistics can be: "In this article, how many times has XXX appeared in total?", "In this article, how many main character roles are there in total and who are they respectively?", "In this article, what hardships has XX mainly experienced?", and the Chinese Prompt1 is constructed as follows: [INST]Task seed: Full-text information extraction and statistics task Task generation: I will give you a fragment from a long text. Based on this fragment, generate a question related to information extraction and statistics. The following are some example questions: - "How many times has XXX appeared in the full text?" - "Who are the main characters mentioned in this fragment and what is the background in which they appear?" - "What key events has XX experienced in this fragment?" According to the content of the fragment, you can freely expand or modify this question. Your question should be as comprehensive as possible, considering the entire text as the context. It must be in the same language as the original text. Note that your question should be simple and direct and should be oriented towards the full text, without segmented descriptions. Avoid using attributive descriptions such as "in this fragment" or "in this text". [ / INST] The above topic induction and analysis can be: "What are the core topics discussed in this article?", "What are the main viewpoints of this article?" The Chinese Prompt2 is constructed as follows: [INST]Task seed: Topic induction and analysis Task generation: I will provide you with a fragment from a long text. Based on this fragment, generate a question related to summarizing or analyzing the main viewpoints. The following are some example questions: - "What core topics does this fragment discuss?" - "What main argument or viewpoint does this fragment present?" According to the content of the fragment, you can freely expand or modify this question. Your question should be as comprehensive as possible, considering the entire text as the context. It must be in the same language as the original text. Note that your question should be simple and direct and should be oriented towards the full text, without segmented descriptions. Avoid using attributive descriptions such as "in this fragment" or "in this text". [ / INST] The above sentiment analysis and attitude evaluation can be: "What is the overall sentiment tone of the article?", "What is the author's attitude towards something?" The Chinese Prompt3 is constructed as follows: [INST]Task seed: Sentiment analysis and attitude evaluation Task generation: I will provide you with a fragment of a long text. Based on this fragment, generate a relevant question involving the evaluation of tone, emotion, or attitude. The following are some examples of questions: - "What is the overall emotional tone of this fragment?" - "How does the author feel about a particular event or theme in this fragment?" Based on the content of the fragment, you can freely expand or modify this question. Your question should be as comprehensive as possible, considering the entire text as the context. It must be in the same language as the original text. Note that your question should be simple and direct and should be directed at the full text, without segmental descriptions. Avoid using attributive descriptions such as "in this fragment" or "in this text".[ / INST] The above speculation and prediction can be: "If the trend in the article continues, what might the result be?" Chinese Prompt4 is constructed as follows: [INST]Task seed: Reasoning and prediction Task generation: I will provide you with a fragment of a long text. Based on this fragment, generate a relevant question involving reasoning or prediction. The following are some examples of questions: - "If the trend described in this fragment continues, what might the result be?" Based on the content of the fragment, you can freely expand or modify this question. Your question should be as comprehensive as possible, considering the entire text as the context. It must be in the same language as the original text. Note that your question should be simple and direct and should be directed at the full text, without segmental descriptions. Avoid using attributive descriptions such as "in this fragment" or "in this text".[ / INST] Through the text recursive full-text Q&A task, the LLM can understand the full-text content more deeply, grasp the key information such as the main idea, emotion, and trend of the article. This helps the LLM perform better in dealing with similar full-text understanding tasks. Using the preset task seeds and prompt templates, the LLM can generate diverse and high-quality questions. These questions not only cover the key information of the full text but also can guide the LLM to generate more in-depth and comprehensive answers. The text recursive full-text Q&A task enables the Q&A system to handle more complex and diverse full-text understanding tasks. By processing different types of task seeds and randomly selected fragments, the LLM can learn more language knowledge and reasoning rules. This helps the LLM show stronger generalization ability when dealing with unseen texts and tasks.

[0077] Taking full-text information extraction and statistics as an example, when the task seed is determined to be "full-text information extraction and statistics", the LLM will use the corresponding prompt template to generate questions. For example, it may generate a question: "In this article, how many times does the main character XXX appear in total?" Then, the LLM will answer this question based on the full text content and count the number of times the main character XXX appears in the full text. This process not only demonstrates the LLM's full-text understanding ability but also reflects its question generation and answer generation capabilities.

[0078] Similarly, for other types of task seeds, such as topic induction and analysis, sentiment analysis and attitude assessment, speculation and prediction, etc., the LLM will use the corresponding prompt template to generate questions and generate answers based on the full text content. This enables the LLM to handle various types of full-text understanding tasks and demonstrates powerful language understanding and reasoning capabilities.

[0079] In an optional embodiment, the answer set includes full-text-based answers; the preset prompt templates include a question generation prompt template, a first answer prompt template, and a second answer prompt template; in step S410, if the task type is a full-text understanding task, then determine the task seed, and the large language model performs text recursive full-text Q&A on the paragraph set according to the determined task seed and the preset prompt template to obtain questions and full-text-based answers. Refer to Figure 3 as shown, including: S4101. If the task type is a full-text understanding task, randomly select a task seed from a pre-constructed task seed pool.

[0080] When the task type is determined to be a full-text understanding task, randomly select a task seed from a pre-constructed task seed pool (such as full-text information extraction and statistics, topic induction and analysis, sentiment analysis and attitude assessment, speculation and prediction, etc.). The task seed contains some specific instructions or topics for guiding the generation of questions and the extraction of answers.

[0081] S4102. Extract a text paragraph from the paragraph set, and the large language model uses the question generation prompt template corresponding to the selected task seed to generate a question for the extracted text paragraph.

[0082] Extract a text paragraph from the paragraph set. The large language model uses the question generation prompt template corresponding to the selected task seed to generate a question for the extracted text paragraph. This question should be able to summarize or target the core content of the paragraph.

[0083] S4103. The large language model uses the pre-constructed first answer prompt template to carry the question and ask the first text paragraph in the paragraph set to obtain an answer to the first text paragraph.

[0084] Use the pre - constructed first answer prompt template Prompt5, carry the just - generated question, and ask the first text paragraph in the paragraph set. The large - language model extracts an answer from the first text paragraph according to the prompt template and the question. Prompt5 is used to generate an answer for the first text paragraph, and it guides the large - language model to generate an accurate and clear answer based on the provided document fragment. The constructed Chinese Prompt5 is as follows: [INST]You are an assistant who answers questions based on the provided document content. I will give you a fragment extracted from a long document. Based on the content of this fragment, please provide an accurate and clear answer to the question. Ensure that your answer is as precise as possible and strictly follows the information in the text. Please answer the question as concisely as possible.[ / INST] S4104. The large - language model uses the pre - constructed second answer prompt template to carry the question and the answer of the previous text paragraph to ask the current text paragraph until all text paragraphs are traversed to obtain an answer based on the full text; among them, when asking the second text paragraph, the answer of the previous text paragraph is the answer of the first text paragraph.

[0085] Starting from the second text paragraph, iterate through the following steps: Use the pre - constructed second answer prompt template, carry the current question and the answer of the previous text paragraph, and ask the current text paragraph. The large - language model extracts an answer from the current text paragraph according to the second answer prompt template Prompt6, the question, and the answer of the previous text paragraph. Integrate the answer of the current text paragraph with the answer of the previous text paragraph to form a comprehensive answer based on all the processed paragraphs so far. The iteration process continues until all text paragraphs are traversed. Prompt6 is used for each round of questioning in the iteration process, and it guides the large - language model to generate a comprehensive and concise answer based on the current paragraph, the current question, and the answer of the previous text paragraph. The constructed Chinese Prompt6 is as follows: [INST]You are an assistant who answers questions based on segmented documents. I will provide you with a paragraph from a long document enclosed by <text_chunk> and < / text_chunk>, and a <question>and< / question> The problem of clamping, and the answer to this problem in the previous paragraph clamped by <last_answer> and < / last_answer>. Please answer the question based on the current paragraph, but note that your final answer should combine the answer to the current paragraph and the answer to the previous paragraph (i.e., the content clamped by <last_answer> and < / last_answer>). Your answer should accurately solve the problem based on the combined text of this paragraph and the previous paragraph, taking into account the content of both paragraphs. Please answer the question as concisely as possible. Note that your answer should be directly addressed to the full text, without segmented descriptions, and should be simple and direct, so please avoid attributive descriptions such as "this paragraph", "the previous paragraph", "combining the previous content".[ / INST] After the iterative loop, what is finally obtained is an answer based on the full text. This answer integrates the information of all text paragraphs and accurately answers the initially generated question.

[0086] Through the method of integrating paragraph-by-paragraph questions and answers, the large language model can understand the text content more deeply, capture the logical relationships and information connections between paragraphs. Each round of answer is based on the current paragraph and the previous round's answer, which helps to ensure the accuracy and consistency of the answer, and avoid information omission or misunderstanding. Through the pre-constructed prompt template and iterative process, questions can be automatically generated and answers can be extracted, greatly improving the efficiency of the Q&A task. The final comprehensive answer integrates the information of the full text, providing a comprehensive and coherent perspective, which helps to better understand and grasp the text content.

[0087] Specifically, in the following order, questions are sequentially asked about the paragraph set: a. Prompt5 carries the generated question to ask the first text paragraph, and obtains the answer to the first text paragraph.

[0088] b. Prompt6 carries the question and the answer to the first text paragraph to ask the second text paragraph, obtains the answer to the second text paragraph and integrates the answer to the previous text paragraph.

[0089] c. Prompt6 carries the question and the answer to the second text paragraph to ask the third text paragraph, obtains the answer to the third text paragraph and integrates the answer to the previous text paragraph.

[0090] d. After the iterative loop, Prompt6 carries the question and the answer to the (n - 1)th text paragraph to ask the nth text paragraph, obtains the answer to the nth text paragraph, and integrates the answer to the previous text paragraph. What is finally obtained is the answer based on the full text.

[0091] The following is an example step. The selected input text is the first five chapters of *Romance of the Three Kingdoms*, which are divided into six text paragraphs (chunks): Step 1: Select the seeds for the full-text information extraction and statistics task. In the example, Task 1 is selected.

[0092] Step 2: Extract a text paragraph (chunk) and generate questions based on the task seed Prompt1. The questions generated in the example are as follows: Q: What key events did Ding Yuan experience in this text? Step 3: Refer to Figure 4 As shown, use Prompt5 to carry Q and ask chunk0, getting the answer A0: A0: Ding Yuan is not mentioned in this text, so it is impossible to answer what key events he experienced.

[0093] Use Prompt6 to carry Q and A0 and ask chunk1, getting the answer A1: A1: Ding Yuan is not mentioned in this text, so it is impossible to answer what key events he experienced.

[0094] Use Prompt6 to carry Q and A1 and ask chunk2, getting the answer A2: A2: Ding Yuan is not mentioned in this text, so it is impossible to answer what key events he experienced.

[0095] Use Prompt6 to carry Q and A2 and ask chunk3, getting the answer A3: A3: Ding Yuan experienced one key event in this text: At the banquet in Wenming Garden summoned by Dong Zhuo, he publicly opposed Dong Zhuo's proposal to depose and replace the emperor, saying that Dong Zhuo's actions were acts of usurpation. Dong Zhuo was so enraged that he drew his sword and wanted to kill Ding Yuan, but was temporarily stopped by Li Ru and other ministers.

[0096] Use Prompt6 to carry Q and A3 and ask chunk4, getting the answer A4. Note that the answer A4 incorporates the content of A3, even though it does not exist in chunk4: A4: Ding Yuan experienced two key events in this text: First, at the banquet summoned by Dong Zhuo, he publicly opposed Dong Zhuo's proposal to depose and replace the emperor, calling it an act of usurpation, which enraged Dong Zhuo. Dong Zhuo drew his sword and wanted to kill Ding Yuan, but was temporarily stopped by Li Ru and other ministers. Second, Ding Yuan's adopted son, Lü Bu, was instigated by Li Su to rebel against Ding Yuan and finally killed Ding Yuan with a knife at night and presented his head to Dong Zhuo.

[0097] Prompt6 is used to ask questions about chunk5 with Q and A4, and the answer A5 is obtained. Since there is no longer any content related to Ding Yuan in chunk5, A5 remains the same as A4.

[0098] A5: Ding Yuan experienced two key events in this text: First, he publicly opposed Dong Zhuo's proposal to depose and replace the emperor at the banquet summoned by Dong Zhuo, calling it an act of usurpation, which angered Dong Zhuo. Dong Zhuo drew his sword and wanted to kill Ding Yuan, but was temporarily stopped by Li Ru and other ministers. Second, Lu Bu, Ding Yuan's adopted son, rebelled against Ding Yuan instigated by Li Su, and finally killed Ding Yuan with a knife at night and presented his head to Dong Zhuo.

[0099] By parsing complex text segments into questions, it provides a verification scenario for the model's reasoning and language understanding capabilities, and is the transformation bridge from text to task in the whole process.

[0100] The following describes the device for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention. The device for constructing an automated long text fine-tuning instruction set based on a large language model described below can be mutually referred to corresponding to the method for constructing an automated long text fine-tuning instruction set based on a large language model described above.

[0101] The device for constructing an automated long text fine-tuning instruction set based on a large language model provided by the present invention, referring to Figure 5 as shown, includes: A text paragraph cutting module 510, configured to obtain the input text, and use a recursive character segmentation method to segment the input text to generate a paragraph set; An instruction question and answer generation module 520, configured to, for the generated paragraph set, the large language model generates a question set and an answer set according to the task type through a preset question type set and a prompt template by using a self-guided learning method; An instruction set evaluation module 530, configured to generate an instruction set based on the generated question set and answer set, evaluate the quality of the instruction set in multiple dimensions, and optimize the instruction set according to the evaluation result to obtain an optimized instruction set.

[0102] Figure 6 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the method for constructing an automated long text fine-tuning instruction set based on a large language model.

[0103] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0104] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the above-mentioned various methods.

[0105] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the method for constructing an automated long text fine-tuning instruction set based on a large language model provided by the above-mentioned various methods.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0107] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / evidence> < / answer> < / question> < / evidence> < / answer> < / answer>

Claims

1. A method for constructing an automated long text fine-tuning instruction set based on a large language model, characterized in that: include: Obtaining input text, segmenting the input text using a recursive character segmentation method, and generating a paragraph set; For the generated paragraph set, the large language model uses a self-guided learning method to generate a question set and an answer set based on the task type through a preset question type set and prompt template; An instruction set is generated based on the generated question set and answer set, the quality of the instruction set is evaluated in multiple dimensions, and the instruction set is optimized according to the evaluation results to obtain an optimized instruction set.

2. The method for constructing an automated long text fine-tuning instruction set based on a large language model according to claim 1, characterized in that: The step of obtaining an input text and segmenting the input text using a recursive character segmentation method to generate a paragraph set includes: Get the input text and a predefined delimiter set, which includes delimiters of different priorities; The text length of the input text is obtained. If the text length of the input text is greater than the target segment length, a delimiter is selected from a predefined delimiter set according to a priority to segment the input text to obtain segmented text blocks; wherein the selected delimiter is the first delimiter in all possible segmentation methods that makes at least one segmented text block not longer than the target segment length; If the text length of the segmented text block is greater than the target segment length, a delimiter is reselected to segment the segmented text block until a paragraph set that meets the first condition is obtained.

3. The method for constructing an automated long text fine-tuning instruction set based on a large language model according to claim 1, characterized in that: The question set includes a single paragraph retrieval task question set, and the answer set includes a single paragraph retrieval task answer set; for the generated paragraph set, the large language model generates the question set and the answer set using a self-guided learning method according to the task type through a preset question type set and a prompt template, including: If the task type is a single paragraph retrieval task, the following operations are performed for each text paragraph in the paragraph set to obtain the corresponding single paragraph retrieval task question set and single paragraph retrieval task answer set: For each question type in the question type set, the large language model uses a pre-built question generation function to generate a corresponding single-paragraph question based on the text paragraph and the question type, thereby obtaining a single-paragraph retrieval task question set; wherein the single-paragraph question is generated based on a prompt template for generating a single-paragraph question; For each single-paragraph question in the single-paragraph retrieval task question set, the large language model uses the answer generation function to generate a single-paragraph answer based on the text paragraph and the single-paragraph question, thereby obtaining the single-paragraph retrieval task answer set.

4. The method for constructing an automated long text fine-tuning instruction set based on a large language model according to claim 1, characterized in that: The question set includes a multi-paragraph retrieval task question set, and the answer set includes a multi-paragraph retrieval task answer set; for the generated paragraph set, the large language model generates the question set and the answer set by a self-guided learning method according to the task type through a preset question type set and a prompt template, including: If the task type is a multi-paragraph retrieval task, the following operations are performed for each multi-paragraph combination in the paragraph set to obtain the corresponding multi-paragraph retrieval task question set and multi-paragraph retrieval task answer set: For each question type in the question type set, the large language model uses a pre-built question generation function to generate a corresponding multi-paragraph question based on a multi-paragraph combination and the question type, thereby obtaining a multi-paragraph retrieval task question set; wherein the multi-paragraph question is generated based on a prompt template for generating multi-paragraph questions; For each multi-paragraph question in the question set, the large language model uses a pre-built answer generation function to generate a multi-paragraph answer based on the multi-paragraph combination and the multi-paragraph question, and obtains a multi-paragraph retrieval task answer set.

5. The method for constructing an automated long text fine-tuning instruction set based on a large language model according to claim 1, characterized in that: The answer set includes answers based on the full text; for the generated paragraph set, the large language model generates a question set and an answer set using a self-guided learning method based on the task type through a preset question type set and prompt template, including: If the task type is a full-text comprehension task, the task seed is determined, and the large language model performs recursive full-text question and answering on the paragraph set according to the determined task seed and the preset prompt template to obtain questions and answers based on the full text; wherein each task seed generates a prompt template for a question.

6. The method for constructing an automated long text fine-tuning instruction set based on a large language model according to claim 5, characterized in that: The preset prompt template includes a question generation prompt template, a first answer prompt template, and a second answer prompt template; if the task type is a full-text comprehension task, a task seed is determined, and the large language model performs text recursive full-text question and answer on the paragraph set according to the determined task seed and the preset prompt template to obtain a question and an answer based on the full text, including: If the task type is a full-text comprehension task, a task seed is randomly selected from the pre-built task seed pool; A text paragraph is extracted from the paragraph set, and the large language model generates questions for the extracted text paragraph using the question generation prompt template corresponding to the selected task seed; The large language model uses a pre-built first answer prompt template to carry a question to ask a first text paragraph in the paragraph set to obtain an answer to the first text paragraph; The large language model uses a pre-built second answer prompt template to carry the question and the answer of the previous text paragraph to ask questions about the current text paragraph until all text paragraphs are traversed to obtain an answer based on the entire text.

7. A device for constructing an automatic long text fine-tuning instruction set based on a large language model, characterized in that: include: A text paragraph segmentation module is used to obtain input text, segment the input text using a recursive character segmentation method, and generate a paragraph set; The instruction question and answer generation module is used to generate a set of questions and a set of answers for the generated paragraph set. The large language model uses a preset set of question types and prompt templates according to the task type and adopts a self-guided learning method; The instruction set evaluation module is used to generate an instruction set based on the generated question set and answer set, evaluate the quality of the instruction set in multiple dimensions, and optimize the instruction set according to the evaluation results to obtain an optimized instruction set.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for constructing an automated long text fine-tuning instruction set based on a large language model as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing an automated long text fine-tuning instruction set based on a large language model as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing an automated long text fine-tuning instruction set based on a large language model as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Data generation method and related product

    CN121478919A