Instruction fine tuning data set construction method and system for large model long text generation

By hierarchically decomposing and retrieving semi-structured knowledge bases, the problems of insufficient structure and citation annotation in long text generation datasets are solved, enabling the construction of high-quality long text generation datasets and improving the logical consistency and citation accuracy of generation tasks.

CN120910193AActive Publication Date: 2025-11-07MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510963977.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-07
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing long text generation datasets lack a clear hierarchical structure and fine-grained annotations, resulting in defects in logical reasoning, content consistency, and structural completeness of the generated texts. Furthermore, insufficient annotation of citation information affects the credibility and academic rigor of the generated content.

Method used

By hierarchically decomposing semi-structured or structured open-source knowledge base documents, extracting structured content, constructing search queries to retrieve the most similar document entries in the entire network or a specified knowledge base, generating fine-grained question-answer pairs, and performing quality filtering and confidence assessment to form a structured question-answer dataset.

Benefits of technology

It enhances the structural control and content consistency of long text generation tasks, making it particularly suitable for academic writing generation tasks with high requirements for citation consistency, and improving the quality and diversity of datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910193A_ABST
    Figure CN120910193A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data set construction, in particular to an instruction fine-tuning data set construction method and system for large-model long text generation, and the method comprises the steps: carrying out hierarchical decomposition on a document of a semi-structured or structured open source knowledge base, extracting structured contents according to three levels of themes, outlines and paragraphs, and removing noise; forming a structured unit; constructing a retrieval formula for each structured unit, retrieving in a whole network, a specified knowledge base and / or a retrieval system to obtain a plurality of literature entries, selecting the literature entry with the highest similarity from the literature entries, and constructing a corresponding reference literature abstract; on the basis of each structured unit and the corresponding reference abstract, generating a fine-grained question-answer pair; and respectively carrying out quality filtering and confidence evaluation on all the fine-grained question and answer pairs to form a structured question and answer data set. Through a multi-agent cooperation mechanism and a hierarchical task decomposition strategy, the quality and efficiency of generated data are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data set construction, and particularly relates to a method and system for constructing instruction fine-tuning data set for long text generation of large model. BACKGROUND

[0002] With the rapid development of deep learning and natural language processing technology, long text generation (LFAG) tasks based on large pre-training language models (such as GPT, BERT, LLaMA, etc.) have achieved remarkable results in intelligent writing, content generation, and automatic question answering applications. Long text generation is a complex task that requires the generated content to not only have logical consistency and contextual coherence, but also cover rich knowledge content and structured information.

[0003] Long text generation data sets are used to further train existing text generation models to enhance their long text generation capabilities. Existing long text generation data sets mostly lack explicit hierarchical structure and fine-grained annotation, resulting in generated text often having defects in logical reasoning, content consistency, and structural completeness. For example, traditional text generation methods often rely on a single information source and lack multi-dimensional evaluation and verification of generated content. This is particularly true in long article generation, where the generated content often lacks coherence and structure, making it difficult to meet the needs of high-quality generation tasks.

[0004] Secondly, the reference information involved in the long text generation task often lacks accurate annotation. In order to ensure the accuracy and reliability of the generated content, the annotation of the reference information is crucial. However, in existing data sets, the accuracy and completeness of the references are often not given enough attention, which directly affects the credibility and academic nature of the generated content. In addition, in existing data sets, the generation task often lacks the support of multi-stage training, and cannot effectively be divided into multiple stages for processing and optimization. The model in the generation process often only relies on a fixed input, lacks adaptability and flexibility for different stage requirements, and the generated content is prone to information omission or logical confusion.

[0005] Currently, although multiple generative models have made some progress in long text generation, existing training data mostly relies on traditional text datasets (such as Wikipedia, news reports, etc.). These datasets often lack detailed annotations and hierarchical support for long text structures, and often only provide simple paragraph and content structures, making it difficult to cover the complex structures and diverse content required for long article generation. Some existing text generation methods have considered the overall structure of the article, but still have problems such as over-reliance on a single data source and insufficient annotations during data processing, resulting in limitations in the diversity and accuracy of the generated articles.

[0006] In recent years, semi-structured data (such as Wikipedia, technical documents, XML documents, etc.) has become an important data source for generation tasks. Semi-structured data is a type of data that is between structured data and unstructured data. It has a certain organizational structure, but does not fully comply with the traditional relational database model. Common semi-structured data includes XML, JSON files, HTML documents, and log files. Semi-structured data often contains rich hierarchical structure information such as titles, paragraphs, chapters, etc. By parsing these structured tags, the abstract, keywords, and other structured information of the article can be effectively extracted, thereby improving the ability of text generation, abstract generation, or question answering tasks. The advantage of semi-structured data is that its structured hierarchical information can provide rich contextual clues for generation tasks, helping the generation model understand the relationship between content and improving the accuracy and efficiency of dataset construction. However, processing semi-structured data is difficult, and data parsing, information extraction, and cleaning work are very tedious, requiring a large amount of resources for deep mining and structured processing.

[0007] With the explosive growth of data, the size of semi-structured data has far exceeded that of structured data, becoming a valuable resource for long text generation tasks. However, due to its high data complexity and strict processing requirements, the effective development and utilization of semi-structured data still face great challenges. Currently, many datasets still fail to fully exploit and apply these hierarchical information, resulting in a large amount of potential data value not being fully utilized. How to efficiently extract key information from semi-structured data and convert it into high-quality datasets that support long text generation tasks is a technical problem that needs to be solved. SUMMARY

[0008] In view of the problems of insufficient structured annotation, missing citation information and insufficient task decomposition in the existing long text generation dataset construction method. The existing datasets rely on a single information source, and lack hierarchical and fine-grained annotation for long article generation tasks, resulting in defects in the logical consistency, information integrity and citation accuracy of the generated content. The purpose of the present application is to overcome the above technical defects, and a kind of instruction fine-tuning dataset construction method for large model long text generation is proposed.

[0009] Therefore, the present application proposes a kind of instruction fine-tuning dataset construction method for large model long text generation, comprising:

[0010] Step 1: Hierarchical decomposition is carried out on the documents of semi-structured or structured open source knowledge base, structured content is extracted according to three levels of theme, outline and paragraph, noise is removed, and structured unit is formed;

[0011] Step 2: For each structured unit, a search formula is constructed, and a plurality of literature entries are searched in the whole network, specified knowledge base and / or search system, the literature entry with the highest similarity is selected, and the corresponding reference abstract is constructed;

[0012] Step 3: Based on each structured unit and the corresponding reference abstract, fine-grained question and answer pairs are generated;

[0013] Step 4: All fine-grained question and answer pairs are respectively filtered and confidence evaluated to form a structured question and answer dataset.

[0014] Preferably, the step 1 comprises:

[0015] Automatic script is used to extract documents, and the extraction rule Hierarchy(D i ) is:

[0016] Hierarchy(D i )={Topic,{Outline ij},{Paragraph ijk}}

[0017] Wherein, D i Indicates the i th document, Topic indicates theme, Outline ij Indicates the j th outline title of the i th document, and Paragraph ijk Indicates the k th paragraph content under the j th outline in the i th document;

[0018] Delete blank paragraph, format error paragraph;

[0019] Merge small paragraphs with continuous theme consistent;

[0020] Standardize hierarchical relationships to ensure a reasonable number of outlines and paragraphs under the same theme;

[0021] D i Decompose into triples: U ijk =(Topic) i Outline ij Paragraph ijk )

[0022] Repeat the above steps until you obtain triples for all documents, which then form a set U of structured units:

[0023] U={U ijk}

[0024] Preferably, the automated script includes: regular expression matching, an HTML / XML parser, and a structure extraction algorithm based on heading level changes.

[0025] Preferably, step 2 includes:

[0026] For each structured unit U in set U ijk Topic i Sub-topic Outline ij Paragraph content ijk Keyword extraction, stop word removal, and phrase recombination were performed to construct the retrieval query Q. ijk :

[0027] Q ijk =f(Topic i Outline ij Paragraph ijk )

[0028] By using keyword enhancement, domain limitation, and phrase recognition, Q... ijk Enhance;

[0029] For the enhanced search query Q′ ijk Perform a search operation on the entire internet, a specified knowledge base, and / or a search system to obtain a search result set R. ijk :

[0030] R ijk ={r1,r2,r3,...r m}

[0031] Where, r m This represents the m-th retrieved document entry, including its title, abstract, and a portion of the text.

[0032] U is calculated according to the following formula. ijk With Rijk Similarity of each document entry r m to U ijk , m from which the document entry with the highest similarity is selected;

[0033] Sim(U ijk ,r m ) = cosine_similarity(U ijk ,r m )

[0034] where cosine_similarity denotes cosine similarity;

[0035] For the selected document entry, record the document title Title, authors Authors, publication year Year, source Source and excerpt Excerpt, to form a structured unit U ijk and its corresponding citation summary Citation ijk :

[0036] Citation ijk = (Title, Authors, Year, Source, Excerpt).

[0037] Preferably, the step 3 comprises:

[0038] For each structured content unit U ijk and its corresponding citation summary Citation ijk , the system predefines a set of multi-dimensional question templates T = {t1, t2,..., t l}, where each template t l represents a type of question;

[0039] Combine the template set with U ijk semantics to automatically generate a set of candidate questions Q ijk = {t1(U ijk ), t2(U ijk ),...}.

[0040] For each candidate question q p ijk ∈ Q ijk , based on the structured content U ijk and the reference summary Citation ijk , call the imperative question answering generation module GenAns to generate an answer a p ijk :

[0041] a pijk = GenAns(q p ijk , U ijk , Citation ijk )

[0042] formatting all the generated results, output a structured question and answer set QA ijk = {(q 1 , a ijk , Citation 1 ) ijk , (q 2 , a ijk , Citation 2 ) ijk , ...}

[0043] Each question and answer pair is accompanied by a reference abstract Citation ijk , forming a fine-grained question and answer pair.

[0044] Preferably, the step 4 comprises:

[0045] Each fine-grained question and answer pair generated in step 3 is subjected to quality filtering including content integrity, answer accuracy and language standardization;

[0046] Conduct confidence evaluation in combination with the output probability or confidence scoring mechanism of the language model. If the confidence is lower than the threshold θ, or there are logical contradictions, language barriers and / or answers without basis, the corresponding fine-grained question and answer is marked as a low-quality sample and removed;

[0047] Using semantic similarity calculation and content hash matching, identify duplicate or similar expression question and answer items, merge or retain high-confidence versions of duplicate content, and remove the rest;

[0048] Uniform structure verification is performed on the retained fine-grained question and answer to ensure that the fields are complete, the reference annotation is legal, and the content character set and encoding format comply with the JSON / CSV specification; At the same time, check for format abnormalities, empty fields and illegal symbols, and store all verified data as a standardized sample set:

[0049] D QA = {(q i , a i , Citation i )}.

[0050] On the other hand, the present application provides a system for constructing instruction fine-tuning data set for large model long text generation, comprising:

[0051] A hierarchical decomposition module is configured to hierarchically decompose the documents of the semi-structured or structured open source knowledge base, extract structured content according to three levels of topics, outlines and paragraphs, remove noise, and form structured units.

[0052] A retrieval module is configured to construct a retrieval formula for each structured unit, retrieve a plurality of literature items in the whole network, a specified knowledge base and / or a retrieval system, select the literature item with the highest similarity, and construct a corresponding reference abstract.

[0053] A question and answer pair generation module is configured to generate fine-grained question and answer pairs based on each structured unit and the corresponding reference abstract.

[0054] A data set formation module is configured to perform quality filtering and confidence evaluation on all fine-grained question and answer pairs, and form a structured question and answer data set.

[0055] Compared with the prior art, the advantages of the present application are that:

[0056] The present application provides a method for constructing an instruction fine-tuning data set for large model long text generation: the method uses a multi-agent collaboration mechanism and a hierarchical task decomposition strategy, uses a structured decomposition and task decoupling method to efficiently convert semi-structured data sources into structured information, and then generates a data set supporting the long text generation task. The method designs a multi-module collaborative work process with clear division of labor, realizes the whole process management from raw data analysis, outline construction, content generation to reference extraction, and improves the efficiency and accuracy of data construction.

[0057] The present application designs a multi-stage annotation method supporting reference tracking and structural consistency: the method introduces layer-by-layer annotation of structured elements such as outlines, paragraph contents and references during data construction, to ensure that the generated data has clear structure and accurate reference information. The method effectively improves the structure control ability and content consistency in the long text generation task, and is particularly suitable for academic writing type generation tasks with high reference consistency requirements. BRIEF DESCRIPTION OF DRAWINGS

[0058] Figure 1 is a schematic diagram of a hierarchical decoupling and fine-grained annotation long text generation data set construction method;

[0059] Figure 2 is a flowchart of a hierarchical decoupling and fine-grained annotation long text generation data set construction method. DETAILED DESCRIPTION

[0060] The application selects semi-structured data represented by Wikipedia as the data source, and fully utilizes the advanced capabilities of large language models in text processing and knowledge extraction. Through multiple steps such as data mining, citation retrieval, question and answer annotation, and data cleaning, the application can efficiently extract useful information from semi-structured data and generate a dataset suitable for long text generation tasks. Compared with traditional dataset construction methods, the application significantly enhances the model's ability to process complex information sources, improves the processing effect of structure and citation in long text generation tasks, and effectively improves the quality and diversity of the dataset.

[0061] In addition, the application introduces a hierarchical task decomposition and phased training strategy, so that the long text generation dataset construction task can be decomposed into multiple sub-tasks for processing, greatly improving the training efficiency and effect of the generation model. With the support of fine-grained annotation and multi-stage training, the application provides a more flexible and efficient way to improve the performance of long text generation models in multiple application scenarios, thereby providing a new technical path for the development of long text generation.

[0062] The technical solutions of the application will be described in detail below in conjunction with the drawings and embodiments.

[0063] Embodiment 1

[0064] The application proposes a method for constructing an instruction fine-tuning dataset for large model long text generation, which includes a content decoupling module, a literature retrieval and summary module, a fine-grained question and answer annotation module, and a data cleaning and quality control module. The method includes the following steps:

[0065] Step 1) Hierarchical decomposition of the original semi-structured Baidu Encyclopedia open source data, extracting structured content according to three levels of theme, outline and paragraph, forming a multi-level content framework, and ensuring the integrity and hierarchy of the literature knowledge system.

[0066] Step 2) For each content unit, search for relevant authoritative references through the retrieval system, select the literature fragment with the highest similarity from the retrieval results as the reference summary of the content unit, and record the citation information to support subsequent generation.

[0067] Step 3) Based on the hierarchical content and corresponding reference information, generate multi-dimensional fine-grained question and answer pairs for each knowledge unit, covering multiple knowledge point angles such as definition, principle, application, and limitation, to ensure the breadth and depth of the question and answer content.

[0068] Step 4) Quality audit of all generated data, including content consistency check, reference literature matching degree evaluation and question and answer quality score, to eliminate low-quality samples and ensure the accuracy and high-quality standard of the final dataset.

[0069] In the technical solution, the step 1) specifically comprises:

[0070] Step 101) Obtain semi-structured knowledge document dataset: This is the basis for constructing long text generation dataset. In this step, various semi-structured or structured knowledge document sets (denoted as set K) need to be collected and organized, such as data from open source knowledge bases such as Wikipedia.

[0071] Specifically, let the original data set be:

[0072] K={D1,D2,D3,...D i}

[0073] Where D i represents the i-th original document. Each document may contain multiple levels of content structure, such as document topic, title, sub-title, paragraph, etc. Through unified processing and pre-screening, it is ensured that the original document has basic hierarchical potential, preparing for subsequent structured decomposition.

[0074] Step 102) Use script to process knowledge document dataset with high degree of structure: For the data after preliminary screening, use automated scripts for preliminary hierarchical analysis. In the analysis process, according to the structured marks in the text (such as HTML tags, markdown marks, explicit outline numbering, etc.), the hierarchical information is extracted: 1. First level: Topic 2. Second level: Outline 3. Third level: Paragraph content. The extraction rule can be expressed as:

[0075] Hierarchy(D i )={Topic,{Outline ij},{Paragraph ijk}}

[0076] Where Outline ij represents the j-th outline title, and Paragraph ijk represents the k-th paragraph content under the j-th outline. The processing script includes but is not limited to: regular matching, HTML / XML parser, structure extraction algorithm based on title level change, etc.

[0077] Step 103) Clean up noise content and abnormal structure: Due to the possible existence of format disorder, paragraph disorder, empty paragraph or error annotation in open source knowledge base, data cleaning is needed in this step, including: 1. Delete blank paragraphs, format error paragraphs; 2. Merge small paragraphs with continuous topics; 3. Standardize hierarchical relationship, ensure reasonable number of outlines and paragraphs under the same topic.

[0078] The noise filtering function is defined as:

[0079] Clean(D i )=D′ i

[0080] where D′ i is the cleaned labeled document.

[0081] Step 104) After cleaning, each document is formally decomposed into a structured unit, which is specifically defined as a triple:

[0082] U ijk =(Topic i ,Outline ij ,Parargraph ijk )

[0083] set:

[0084] U={U ijk}

[0085] will be the basis for subsequent retrieval, question and answer generation, labeling and other steps.

[0086] In the above technical solution, the step 2) specifically includes:

[0087] Step 201) Constructing a search query: After obtaining the set U from step one, for each content unit U ijk , a search query Query needs to be automatically generated to retrieve relevant documents in the external knowledge base. The search query Q ijk is constructed by the following method:

[0088] Q ijk =f(Topic i ,Outline ij ,Paragraph ijk )

[0089] Where the function f(·) represents keyword extraction, stop word removal, phrase reorganization, etc. on the topic, subtopic and paragraph content to generate a concise and representative query. In order to improve the retrieval effect, this step can use: 1. Keyword expansion: expand the core word with synonyms. 2. Domain restriction: If the system supports, domain restriction words (such as physics, finance, medicine, etc.) can be added. 3. Phrase mining: Keep key expressions such as proper nouns and technical terms.

[0090] The formulaic representation of the search query enhancement is:

[0091] Q′ ijk =Enhance(Q) ijk )

[0092] Step 202) Perform a literature search. Use search query Q′ ijk Perform a search operation on the entire internet or a specified knowledge base or retrieval system (such as a self-built academic database, open-source literature repository, or API interface such as Google Serper API or Bing API), and return a set of search results:

[0093] R ijk ={r1,r2,r3,...r m}

[0094] Where, r m This represents the m-th retrieved document entry, typically including basic information such as title, abstract, and excerpt from the text. To improve search accuracy, the system supports: 1. multi-round searches; 2. exclusion of irrelevant documents.

[0095] Step 203) Filtering highly relevant document fragments based on similarity matching: For the search result set R ijk It is necessary to evaluate the relationship between each document and the original content unit U. ijk The semantic similarity is used to select the segment with the highest similarity as the reference summary. The similarity function is:

[0096] Sim(U ijk ,r t =cosine_similarity(U ijk ,r m )

[0097] Here, cosine_similarity represents the cosine similarity.

[0098] Step 204) Record references and corresponding citation information: For each selected highly relevant fragment r * The following citation information needs to be standardized and recorded: 1. Title; 2. Authors; 3. Publication Year; 4. Source (Journal / Conference / Web); 5. Excerpt. The citation metadata record is defined as follows:

[0099] Citation ijk =(Title,Authors,Year,Source,Excerpt)

[0100] Finally, for each unit U ijk Link a corresponding reference abstract (Citation)ijk , as an important support for subsequent question generation and fine-grained annotation.

[0101] In the technical solution, the step 3) specifically includes:

[0102] Step 301) constructing fine-grained question templates and generating candidate questions: for each structured content unit U ijk and its corresponding reference abstract r * , the system pre-sets a multi-dimensional question template set:

[0103] T = {t1, t2,.....t l}

[0104] Where each template t l represents a type of question, such as definition, principle, application or advantages and disadvantages. Combining the template with the content unit semantics, a set of candidate questions is automatically generated:

[0105] Q ijk = {t1(U ijk ), t2(U ijk ),…}

[0106] For example, for "convolutional neural network", questions such as "what is convolutional neural network" and "how does it work" can be generated, covering key knowledge dimensions.

[0107] Step 302) generating answer text based on content and reference abstract: for each candidate question q p ijk ∈ Q ijk , the system generates answers based on structured content U ijk and reference abstract Citation ijk , calling the instructional question generation module:

[0108] a p ijk = GenAns(q p ijk , U ijk , Citation ijk )

[0109] The generation process integrates the extraction of question-related information through a language model, ensuring that the answers are accurate, the content is derived from reference literature, the language expression is natural and standard, and it has traceability and context consistency.

[0110] Step 303) standardizing output fine-grained question and answer pairs: the system formats all generated results, outputting a structured question and answer set:

[0111] QA ijk = {(q1 ijk a 1 ijk ),...} 2 ijk a 2 ijk ),...}

[0112] Each question-answer pair is accompanied by a reference abstract Citation ijk , forming a complete triple (q, a, Citation). Finally, it is exported in JSON or CSV format, constituting a high-quality fine-grained question-answer dataset that is structurally unified, source-specific, and directly usable for fine-tuning instructions.

[0113] In the above technical solution, the step 4) specifically comprises:

[0114] Step 401) Low-quality question-answer pair screening and confidence evaluation: The system first conducts quality detection on each question-answer pair generated in step 3), including automatic evaluation of content integrity, answer accuracy, language standardization, etc., and combines the output probability or confidence scoring mechanism of the language model to determine its credibility. If the confidence of a question-answer pair is lower than the threshold θ, or there are logical contradictions, language barriers, or no basis for answering, etc. abnormal situations, it will be marked as a low-quality sample and removed:

[0115] (q, a, Citation) → discard, if Conf(a) < θ

[0116] This process ensures that only question-answer samples with clear semantics, standard structure, and reference support are retained for subsequent steps.

[0117] Step 402) Duplicate question-answer and redundant content deduplication: In the entire question-answer data, the system uses semantic similarity calculation and content hash matching to identify duplicate or similar question-answer items, and merges or retains the high-confidence version of the duplicate content, and removes the rest. The semantic repetition judgment function can be represented as:

[0118]

[0119] where δ is the similarity threshold. This process can significantly improve the uniqueness and information density of question-answer samples and avoid training data bias.

[0120] Step 403) Final data verification and format standardization: After screening and deduplication, the system verifies the retained question-answer samples for uniform structure, ensuring that the fields are complete, the reference annotations are legal, the content character set and encoding format comply with JSON / CSV specifications; At the same time, check if there are format abnormalities, empty fields or illegal symbols, etc. All verified data is stored as a standardized sample set:

[0121] D QA = {(q i ,a i ,Citation i )}

[0122] This step ensures that the final output data has a stable structure, strong training readability, and can be directly input as high-quality instruction fine-tuning data.

[0123] Example 2

[0124] As Figure 1 shown, the decomposition decoupling and fine-grained annotation long text generation dataset construction method of the present application mainly includes the following steps:

[0125] Step 1: Prepare data source

[0126] The original dataset comes from open source knowledge bases such as Wikipedia. Data miners will extract the required content from these documents, which includes:

[0127] 1. Topic

[0128] 2. Outline

[0129] 3. Paragraph content

[0130] This step first ensures that the document has hierarchical potential and prepares for subsequent structured decomposition. Redundant parts of the text should be deleted to ensure that the hierarchical structure is clear and operable. Use automated scripts (such as HTML tags, Markdown markers, etc.) to extract layers to ensure that the document structure is standardized and easy to process later.

[0131] Table 1: Term index example

[0132] Index field Field explanation Title Article title Topic Overall topic of the document, as the highest level of title Outline Structured titles of various sections Paragraph Detailed paragraph content under each section Reference Detailed information of the cited literature, including title, url, snippet, etc.

[0133] Step 2: Reference retrieval and abstract extraction

[0134] In each content unit, search for relevant authoritative references through the retrieval system. The system searches according to the content and selects the literature abstracts related to the theme to generate content abstract pairs and record the reference information. In this process, multiple retrieval tools (such as Google Serper API, Bing API, etc.) will participate in multiple rounds of screening, and the most relevant literature will be selected through similarity algorithms to ensure the authority and accuracy of the reference information.

[0135] Table 2: Reference retrieval and abstract extraction example

[0136]

[0137] Step 3: Fine-grained question and answer annotation

[0138] Based on the information in steps 1 and 2, the question and answer annotator designs question and answer pairs according to the document content. Each question and answer pair not only covers multiple knowledge points such as definition, principle, and application, but also ensures the breadth and depth of the content. Question and answer generation is not limited to simple questions and answers, but also includes in-depth analysis of each paragraph or knowledge point. The answer to each question should be detailed to ensure coverage of various aspects from definition to application.

[0139] Table 3: Question and answer pair data sample

[0140]

[0141] Step 4: Data cleaning and quality control

[0142] In this step, through multi-angle review, all generated question and answer data are strictly quality screened. Including: 1. Content consistency check 2. Reference literature matching degree evaluation; question and answer quality score.

[0143] Data cleaning is not only to remove low-quality samples, but also to ensure that the generated content meets academic standards and can efficiently support the generation of training data. Quality review should cover content consistency, literature citation accuracy, and whether it meets the expected depth and breadth.

[0144] Figure 2 is a flowchart of the method of the present invention.

[0145] Example 3

[0146] Example 3 of the present invention provides a system for constructing instruction fine-tuning data sets for large model long text generation, based on the method of example 1, including:

[0147] Hierarchical decomposition module, for hierarchical decomposition of semi-structured or structured open source knowledge base documents, extracting structured content according to three levels of theme, outline and paragraph, removing noise, and forming structured units;

[0148] Retrieval module, for constructing a retrieval formula for each structured unit, retrieving a number of literature items in the entire network, specified knowledge base and / or retrieval system, and selecting the literature item with the highest similarity to construct the corresponding reference literature abstract;

[0149] Question and answer pair generation module, for generating fine-grained question and answer pairs based on each structured unit and the corresponding reference literature abstract;

[0150] Dataset composition module, for quality filtering and confidence evaluation of all fine-grained question and answer pairs, to form a structured question and answer dataset.

[0151] It is worth noting that in the above embodiment of the system, each module included is only divided according to functional logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional module is only for easy mutual differentiation, and is not used to limit the protection scope of the present application.

[0152] Innovation points:

[0153] The present application proposes a long text generation instruction fine-tuning data set construction method combining structured task decomposition and multi-functional module cooperation, which has multi-level structure modeling capability and reference information tracking capability. The method divides the outline construction, paragraph generation, reference retrieval and quality control sub-tasks, so as to ensure that the generated data has structural integrity and content accuracy. The possible alternative solution is a direct extraction or rewriting method based on a large model, which lacks task controllability and information tracing mechanism, and is easy to cause structural confusion and factual errors.

[0154] The present application designs a question and answer generation mechanism combining outline information driving and fine-grained reference embedding, which supports generating training samples with high readability and verifiability from multiple dimensions such as structure level, content points and reference relationship. Compared with the question and answer construction method based on the original content only, the mechanism proposed in the present application can significantly enhance the organization and knowledge tracing ability of the generated content, and improve the modeling ability of the model for complex structure content.

[0155] The present application establishes a multi-stage cooperative process composed of outline generation, reference retrieval, paragraph question and answer generation and quality filtering, integrates language model generation capability and regularized processing mechanism, and can continuously and stably output high-quality data set with clear structure and clear reference. The possible alternative solution is to use an end-to-end instructional generation method, but such method often lacks quality control and structure consistency guarantee, and is difficult to meet the strict requirements of long text generation task in training data.

[0156] Finally, it should be explained that the above embodiments are only used to illustrate the technical solutions of the present application and are not limited. Although the present application has been described in detail with reference to the embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A method for constructing an instruction fine-tuning dataset for large model long text generation, comprising: Step 1: Hierarchical decomposition of the documents of semi-structured or structured open source knowledge bases, extraction of structured content according to three levels of theme, outline and paragraph, removal of noise, and formation of structured units; Step 2: Constructing a search formula for each structured unit, searching for several literature entries in the entire network, specified knowledge base and / or search system, selecting the literature entry with the highest similarity, and constructing the corresponding reference literature abstract; Step 3: Generating fine-grained question and answer pairs based on each structured unit and the corresponding reference literature abstract; Step 4: Quality filtering and confidence evaluation of all fine-grained question and answer pairs to form a structured question and answer dataset.

2. The instruction fine-tuning data set construction method for large model long text generation according to claim 1, characterized in that, The step 1 includes: The document is extracted by using an automated script, the extraction rule Hierarchy(D i ) being: Hierarchy(D i ) = { Topic, { Outline ij}, { Paragraph ijk}} wherein D i represents the ith document, Topic represents a topic, Outline ij represents the jth outline title of the ith document, Paragraph ijk represents the kth paragraph content in the ith document belonging to the jth outline Delete blank paragraphs and format error paragraphs; Merge small paragraphs with continuous and consistent themes; Standardize hierarchical relationships to ensure reasonable number of outlines and paragraphs under the same theme; D i Decomposed into triples: U ijk = (Topic i , Outline ij , Paragraph ijk ) Repeat the above steps until the triplets of all documents are obtained, and then form a set U of structured units: U = {U ijk}.

3. The instruction fine-tuning dataset construction method for large model long text generation according to claim 2, characterized in that, The automation script includes: regular matching, HTML / XML parser and structure extraction algorithm based on title level change.

4. The instruction fine-tuning dataset construction method for large model long text generation according to claim 2, characterized in that, The step 2 includes: For each structured unit U in the set U ijk , the subject Topic i , the sub- subject Outline ij , and the paragraph content Paragraph ijk , keyword extraction, stop word removal, and phrase reorganization are performed to construct a search expression Q ijk : Q ijk = f(Topic i , Outline ij , Paragrapph ijk ) Q ijk is enhanced by key word enhancement, domain restriction and phrase recognition; Q' = Q + Q'Q ijk performing a search operation in the network-wide, designated repository and / or search system to obtain a set of search results R ijk : R ijk ={r1,r2,r3,...r m} wherein r m represents the mthretrieved document entry, including title, abstract and text segment; U is calculated according to the formula ijk with R ijk each document entry r t The similarity Sim(U ijk , r m ) is calculated for each document entry r ; the document entry with the highest similarity is selected. Sim(U ijk ,r m ) = cosine_similarity(U ijk ,r m ) Wherein, cosine_similarity represents the cosine similarity; For the selected document entry, record the document title Title, the authors Authors, the year of publication Year, the source Source and the text of the excerpt Excerpt, constituting the structured unit U ijk of the reference Citation ijk : Citation ijk = (Title, Authors, Year, Source, Excerpt).

5. The instruction fine-tuning dataset construction method for large model long text generation according to claim 1, characterized in that, The step 3 includes: For each structured content unit U ijk and its corresponding reference abstracts Citation ijk The system pre-sets a multi-dimensional problem template set T = {t1, t2, ..., t} l }, where each template t l Indicates a type of question; The template set is combined with U ijk semantics to automatically generate a candidate question set Q ijk = {t1(U ijk ), t2(U ijk ),...}; For each candidate question q p ijk ∈ Q ijk , based on structured content U ijk and reference abstract Citation ijk , invoke the imperative QA generation module GenAns to generate answer a p ijk : a p ijk = GenAns(q p ijk ,U ijk ,Citation ijk ) Format processing of all generated results, output structured question and answer set QA ijk = {(q 1 ijk ,a 1 ijl )(q 2 ijk ,a 2 ijk ),...} Citation ijk , forming fine-grained question-answer pairs.

6. The instruction fine-tuning dataset construction method for large model long text generation according to claim 1, characterized in that, The step 4 includes: Quality filtering of each fine-grained question and answer pair generated in step 3, including content integrity, answer accuracy and language standardization; Conduct confidence evaluation combined with the output probability or confidence scoring mechanism of the language model. If the confidence is lower than the threshold θ, or there are logical contradictions, language barriers and / or no basis for answering, the corresponding fine-grained question and answer is marked as a low-quality sample and removed; Use semantic similarity calculation and content hash matching to identify duplicate or similar question and answer items. Merge or retain the high-confidence version of the duplicate content, and remove the rest; Uniform structure verification of the retained fine-grained question and answer to ensure that the fields are complete, the reference annotations are legal, and the content character set and encoding format comply with the JSON / CSV specification; At the same time, check for format abnormalities, empty fields and illegal symbols. All data that passes the verification is stored as a standardized sample set: D QA = {(q i ,a i ,Citation i )}.

7. A system for constructing an instruction fine-tuning dataset for large model long text generation, characterized in that, It includes: A hierarchical decomposition module for hierarchical decomposition of the documents of semi-structured or structured open source knowledge bases, extraction of structured content according to three levels of theme, outline and paragraph, removal of noise, and formation of structured units; A search module for constructing a search formula for each structured unit, searching for several literature entries in the entire network, specified knowledge base and / or search system, selecting the literature entry with the highest similarity, and constructing the corresponding reference literature abstract; A question and answer pair generation module for generating fine-grained question and answer pairs based on each structured unit and the corresponding reference literature abstract; And A dataset forming module for quality filtering and confidence evaluation of all fine-grained question and answer pairs to form a structured question and answer dataset.

Citation Information

Patent Citations

  • Precise reference article generation method and system based on large model

    CN119357386A

  • Method and system for generating bionic hierarchical memory fusion document of large electric semantic model

    CN119474347A

  • Article generation method and device, equipment and storage medium

    CN119476496A

  • High-quality data set construction method and system for large model in vertical field

    CN119647595A

  • Tuning a generative artificial intelligence model

    US11875240B1