Large language model fine tuning and evaluation method based on retrieval enhancement generation
Through refined data processing and scientific evaluation, the problems of bias, illusion, and obsolescence in large language models have been solved, improving the accuracy and reliability of the model in complex question-answering scenarios and realizing the enhancement of the model's diversified capabilities.
Patent Information
- Application Number
- CN202511456625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing large language models suffer from bias, illusion, and obsolescence in practical applications, and existing optimization techniques based on retrieval enhancement cannot fundamentally solve these problems.
By employing refined data processing, diverse fine-tuning data construction, and a scientific evaluation system—including pre-defined character-level segmentation rules, knowledge graph construction, generation of noisy data, complex data, rejected data, and multi-hop data, as well as evaluation of answer relevance and similarity metrics—the accuracy and robustness of the model can be improved.
It improves the accuracy and reliability of the model in complex question-answering scenarios, enhances the model's ability to resist interference from interfering information, improves logical reasoning ability, and achieves a balance between specialized and general capabilities.
Smart Images

Figure CN120910281A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, in particular to a retrieval-augmented generation based large language model fine-tuning and evaluation method. BACKGROUND
[0002] As a language model based on artificial neural network, a large language model usually contains tens of billions or even more parameters, and is trained on a large amount of unlabeled text data through self-supervised learning or semi-supervised learning. It has shown great power in many fields such as natural language processing. However, in practical application, users can directly perceive three major problems of this kind of model: first, bias, the model output may imply a tendency that does not conform to objective fairness; second, illusion, generating false information that does not conform to the facts; third, obsolescence, due to the time limitation of training data, it is difficult to cope with emerging knowledge or dynamic information needs.
[0003] To solve the above problems, retrieval-augmented generation (RAG) technology has emerged. At present, there is no retrieval-augmented generation based large language model fine-tuning technology in the technology combining retrieval-augmented generation and large language model. Some researches optimize the question and answer effect of large language model by controlling the number of recalled documents, which can alleviate the bias, illusion and obsolescence of large language model to a certain extent, but cannot fundamentally solve the above core problems. Therefore, the existing retrieval-augmented generation based large language model question and answer optimization technology still has significant bottlenecks. SUMMARY
[0004] The present disclosure provides a retrieval-augmented generation based large language model fine-tuning and evaluation method, which improves the accuracy, robustness, reliability and reasoning ability of retrieval-augmented generation model through fine data processing, diversified fine-tuning data construction, efficient model fine-tuning and scientific evaluation system, so that it can better cope with complex and diverse question and answer scenarios in practical application.
[0005] According to a first aspect of the present disclosure, a retrieval-augmented generation based large language model fine-tuning and evaluation method is provided, comprising: dividing the input original text document content into multiple independent paragraphs according to a preset character level division rule; extracting entities and relationships between entities from the multiple independent paragraphs to construct a knowledge graph; generating noise data, complex data, rejection data and multi-hop data through a noise sampling module, a complex attribution module, a rejection analysis module and a thinking reasoning module based on the knowledge graph; The noise data, complex data, rejection data, multi-hop data and general data are format-converted to obtain fine-tuning data, wherein the general data is open source model pre-training question and answer data; The large language model is fine-tuned according to the fine-tuning data to obtain a fine-tuned retrieval enhancement generation model; The fine-tuned retrieval enhancement generation model is evaluated according to the answer relevance index and the answer similarity index.
[0006] As a preferred embodiment, the input original text document content is divided into multiple independent paragraphs according to a preset character-level division rule, which includes: According to the preset character-level division rule, the original text is divided based on a preset number of words and common punctuation marks; The independent paragraphs formed after division are subjected to semantic integrity detection; If the end of an independent paragraph is an incomplete sentence, the independent paragraph is merged into the next independent paragraph for re-division; If the last independent paragraph is an incomplete sentence, it is merged into the previous independent paragraph.
[0007] As a preferred embodiment, the entities and inter-entity relationships are extracted from the independent paragraphs to construct a knowledge graph, which includes: A large language model is used to analyze the entities and inter-entity relationships in the paragraphs based on a preset prompt word template; The extracted entities and inter-entity relationships are mapped to nodes and edges in the knowledge graph to construct the knowledge graph.
[0008] As a preferred embodiment, the preset prompt word template includes a target instruction and a step-by-step analysis rule, wherein the step-by-step analysis rule includes: All entities are identified and entity information is extracted, including entity name, entity type and entity description; Related entity pairs are identified from the identified entities, and inter-entity relationship information is extracted, including source entity name, target entity name, relationship description and relationship strength.
[0009] As a preferred embodiment, the extracted entities and inter-entity relationships are mapped to nodes and edges in the knowledge graph to construct the knowledge graph, which includes: The entity name is used as a unique node identifier, the entity type is used as a node classification tag, and the entity description is used as a node attribute; The source entity name and the target entity name correspond to the starting node and the ending node of the edge respectively, the relationship description is used as the type of the edge, and the relationship strength is used as the weight of the edge; The mapped nodes and edges are stored in a triple format, while the node attributes and edge weights are associated, forming an extended triple structure, to build a knowledge graph.
[0010] As a preferred embodiment, based on the knowledge graph, noise data, complex data, rejection data and multi-hop data are generated by a noise sampling module, a complex attribution module, a rejection analysis module and a thinking reasoning module, respectively, including: Based on the entities and the relationships between the entities in the knowledge graph, a first simple question and answer pair is constructed, and the paragraph information related to the first simple question and answer pair is selected from the independent paragraph set and added to the first simple question and answer pair as noise information according to a preset proportion, to form noise data; M entities with similar attributes or relationships are selected from the knowledge graph to form an entity cluster, and a complex question and answer pair is constructed based on the entity cluster, to form complex data; Based on the entities in the knowledge graph and the absolute paragraph containing the explicit answer, a second simple question and answer pair is constructed, and other paragraphs except the absolute paragraph are selected from the independent paragraph set as related text of the second question and answer pair, to form rejection data; Bridge entities are selected in the knowledge graph, and a candidate paragraph pair containing three entity relationships is generated based on the bridge entities and associated entities to construct a multi-hop question and answer pair, to form multi-hop data, wherein the bridge entity is associated with at least two other entities, and there is no direct relationship between the associated entities.
[0011] As a preferred embodiment, the noise data, complex data, rejection data, multi-hop data and general data are format-converted to obtain fine-tuning data, including: A triple data structure containing a task description field, a context information field and a target output field is called; The noise data, complex data, rejection data, multi-hop data and general data are mapped to the triple data structure according to a preset mapping structure; The mapped triple data is normalized to obtain fine-tuning data that meets the input requirements of the model fine-tuning framework.
[0012] As a preferred embodiment, the noise data, complex data, rejection data, multi-hop data and general data are mapped to the triple data structure according to a preset mapping structure, including: The original question of the noise data, complex data and multi-hop data is mapped to the task description field, the associated background text is mapped to the context information field, and the standard answer derived is mapped to the target output field; The original question of the rejection data is mapped to the task description field, the interference text without the correct answer is mapped to the context information field, and the preset knowledge rejection identifier is mapped to the target output field. The original task description of the general data is mapped to the task description field, the original input information is mapped to the context information field, and the original output result is mapped to the target output field.
[0013] In a preferred embodiment, the evaluation of the fine-tuned retrieval enhancement generation model based on answer relevance and answer similarity indicators includes: Based on the generated answer, at least one related question is derived in reverse, and the related question is combined with the original input question to form a question set; Calculate the semantic relevance score between each question in the question set, as an indicator of answer relevance; The generated answers are compared with the standard answers for semantic equivalence to obtain an answer similarity index.
[0014] As a preferred embodiment, the formula for calculating the answer relevance index is: ; in, It is the generated first Each question corresponds to a vector. It is the vector corresponding to the user's question. It represents the number of problems generated.
[0015] According to a second aspect of this disclosure, a device for fine-tuning and evaluating a large language model based on retrieval enhancement generation is provided, comprising: The text segmentation module is used to segment the input raw text document content into multiple independent paragraphs according to preset character-level segmentation rules; The knowledge graph construction module is used to extract entities and relationships between entities from the multiple independent paragraphs and construct a knowledge graph. The feature data generation module is used to generate noisy data, complex data, rejection data, and multi-hop data based on the knowledge graph through the noise sampling module, complex attribution module, rejection analysis module, and reasoning module, respectively. The data conversion module is used to convert the noisy data, complex data, rejection data, multi-hop data and general data into formats to obtain fine-tuning data, wherein the general data is the question-and-answer data used for pre-training of the open-source model; The large language model fine-tuning module is used to fine-tune the large language model based on the fine-tuning data to obtain the fine-tuned retrieval enhancement generation model; The large language model evaluation module is used to evaluate the fine-tuned retrieval enhancement generation model based on answer relevance and answer similarity metrics.
[0016] According to a third aspect of the present disclosure, an electronic device is provided, including at least one processor, and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of the preceding aspects.
[0017] Compared with the prior art, the present disclosure achieves the following beneficial effects: (1) The original text is segmented by pre-set character level segmentation rules, and the paragraphs are adjusted in combination with semantic integrity detection to ensure the semantic integrity of the paragraphs, thereby laying a high-quality data foundation for subsequent entity and relationship extraction. This segmentation method avoids semantic loss caused by paragraph fragmentation, making the extraction of entities and relationships between entities from the paragraphs more accurate, and thus the knowledge graph constructed is more in line with the real semantics of the text, providing reliable knowledge support for model training. (2) Based on the knowledge graph, noise data, complex data, refusal data and multi-hop data are generated, and fine-tuning data is formed in combination with general data. The diversified data types can specifically improve the model's ability in different scenarios: noise data can enhance the model's anti-interference ability to irrelevant information, so that it can still accurately focus on the core problem when there is irrelevant information; complex data helps the model to handle problems involving multiple similar entities or complex relationships, improving its ability to cope with complex tasks; the introduction of refusal data enables the model to learn to identify questions that cannot be answered, reducing the situation of fabricating answers and enhancing the reliability of the model's output; multi-hop data can exercise the model's logical reasoning ability, enabling it to handle problems that require multiple steps of deduction. The integration of general data ensures that the model has the ability to handle specific scenarios while not losing the general question and answer ability, achieving a balance between special capabilities and general capabilities. (3) The various data are mapped to a triple data structure and standardized through unified format conversion, so that data of different types and from different sources can adapt to the model fine-tuning framework, improving the efficiency and convenience of the fine-tuning process and enhancing the fine-tuning effect.
[0018] (4) The fine-tuned model is evaluated by combining the answer relevance index and the answer similarity index, which is more comprehensive and scientific than a single index. The answer relevance index ensures that the model-generated answers are closely related to the questions, avoiding irrelevant answers; the answer similarity index measures the degree of agreement between the generated answers and the standard answers, and the combination of the two can more accurately reflect the quality of the model-generated answers, providing a reliable basis for the evaluation of model performance.
[0019] It is to be understood that the description in the summary section is not intended to identify key or essential features of embodiments of the disclosure or to limit the scope of the disclosure. Other features, aspects, and advantages of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings. The drawings are intended to illustrate but not to limit the present disclosure. In the drawings: Figure 1 A flowchart of a large language model fine-tuning and evaluation method based on retrieval enhancement generation according to an embodiment of the present disclosure is shown; Figure 2 A flowchart of a large language model fine-tuning method based on retrieval enhancement generation according to an embodiment of the present disclosure is shown; Figure 3 A block diagram of a large language model fine-tuning and evaluation device based on retrieval enhancement generation according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0022] In addition, the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.
[0023] Embodiment One In combination Figure 1 And Figure 2 A large language model fine-tuning and evaluation method based on retrieval enhancement generation according to an embodiment of the present disclosure is explained and described.
[0024] As Figure 1 A flowchart of a large language model fine-tuning and evaluation method based on retrieval enhancement generation according to an embodiment of the present disclosure is shown, and the method 100 includes: S110: Cut the input original text document content into multiple independent paragraphs according to the preset character level segmentation rule.
[0025] In some embodiments, according to the preset character level segmentation rule, the original text is segmented based on the preset number of words and common punctuation marks; Wherein, the preset character level segmentation rule contains segmentation base and dynamic adjustment mechanism: first, based on the preset number of words (such as 500-800 words) and common punctuation marks (period, question mark, exclamation mark, semicolon), the original text is preliminarily segmented.
[0026] Secondly, in order to ensure the semantic integrity of the paragraph and lay a high-quality data foundation for subsequent entity and relationship extraction, the segmented paragraph is further detected by a semantic integrity detection model (a pre-trained text understanding model can be used) to identify whether the end of the paragraph is a complete sentence (the judgment standard is whether the end punctuation is a sentence end punctuation and the sentence component is complete).
[0027] If there is an independent paragraph ending with an incomplete sentence (such as ending with a comma, a colon, or a sentence lacking a predicate), the independent paragraph is merged into the next independent paragraph for re-segmentation; if the last independent paragraph is an incomplete sentence, it is merged into the previous independent paragraph.
[0028] Through this process of "preliminary segmentation-semantic verification-dynamic merging", this embodiment ensures that all independent paragraphs have complete semantics and provide a coherent context basis for subsequent entity extraction.
[0029] S120: Extract entities and relationships between entities from the multiple independent paragraphs to construct a knowledge graph.
[0030] In some embodiments, the extraction of entities and relationships between entities from multiple independent paragraphs to construct a knowledge graph adopts a two-level processing method of "large language model analysis + structured mapping": First, use a large language model (such as GPT-4, LLaMA series) to analyze the entities and relationships between entities in the paragraph based on a preset prompt template, wherein the preset prompt template includes explicit task objectives, a list of entity types (such as people, organizations, events, etc.), examples of relationship types (such as membership, causality, time sequence, etc.), and output format constraints.
[0031] Among them, the analysis of entities and relationships between entities in the paragraph based on the preset prompt template includes: First, identify all entities and extract entity information (need to remove duplicates, same name different entity mark distinction), including entity name, entity type, and entity description, etc. Secondly, relevant entity pairs are identified again from the identified entities, and the relationship information between entities is extracted, including source entity name (associated unique entity identifier), target entity name (associated unique entity identifier), relationship description, and relationship strength, etc.
[0032] It should be noted that the identification of relevant entity pairs can be based on co-occurrence frequency and semantic correlation degree screening. Entity pairs with a correlation degree below a threshold value are marked as "weak correlation", and entity pairs with a correlation degree above a threshold value are marked as "strong correlation".
[0033] Exemplarily, the prompt word template is as follows: Objective: Given a text document that may be related to the activity and a set of entity types, identify all entities of these types from the text and all relationships between the identified entities.
[0034] Steps: (1) Identify all entities, and for each identified entity, extract the following information: entity_name: entity name, first letter capitalized; entity_type: one of the following types: [{entity_types}]; entity_description: format each entity as ("entity"{tuple_delimiter}<entity_name>{tuple_delimiter}<entity_type>{tuple_delimiter}<entity_description>); (2) From the entities identified in step (1), identify all (source entity, target entity) pairs that are apparently related to each other. For each pair of related entities, extract the following information: source_entity: the name of the source entity, as identified in step (1); target_entity: the name of the target entity; relationship_description: explains why the source entity and the target entity are considered to be related to each other; relationship_strength: a numerical score representing the strength of the relationship between the source entity and the target entity. Format each relationship as: ("relationship"{tuple_delimiter}<source_entity>{tuple_delimiter}<target_entity>{tuple_delimiter}<relationship_description>{tuple_delimiter}<relationship_strength>)).
[0035] (3) English returns a single list containing all entities and relationships identified in steps (1) and (2), using **{record_delimiter}** as a list separator.
[0036] (4) After completion, output {completion_delimiter}.
[0037] Further, the extracted entities and relationships between entities are mapped to nodes and edges in the knowledge graph, and the knowledge graph is constructed, including: (a) Node mapping rules: The identified entity name is used as the unique identifier of the node. If there is a case of duplicate entity names but different entity types or entity descriptions, a unique identifier (such as entity type abbreviation + serial number, e.g. "Beijing_City_001", "Beijing_Company_002") is added after the entity name to distinguish them. The entity type is used as the classification label of the node, and a preset type hierarchy system is used to standardize the entity type to ensure that entities of the same category correspond to a unified label. The entity description is stored as a node attribute in the form of key-value pairs (e.g. {"date of birth": "1955", "nationality": "American"}), and the description content is de-duplicated, noise-reduced, and key information is retained, supporting the associated storage of multi-language descriptions.
[0038] (b) Edge mapping rules: The source entity name and target entity name identified are respectively mapped to the starting node and terminating node of the edge. If the entity name is duplicated, its unique identifier is associated to ensure the accuracy of the edge direction. The relationship description is used as the type of the edge, and a mapping table between the relationship description and the standard relationship type is established to normalize non-standard relationship descriptions. The relationship strength is used as the weight of the edge, using a numerical range of 0-1 to represent, calculated through quantitative indicators such as entity co-occurrence frequency and relationship mention times, and the weight value is kept to two decimal places. The higher the weight value, the closer the relationship.
[0039] (c) Knowledge graph storage and optimization: The mapped nodes and edges are stored in a triple (head entity, relation type, tail entity) format, while associating node attributes and edge weight information, forming an extended triple structure; Redundancy detection is performed on the constructed knowledge graph, and duplicate nodes, edges and triple data are deleted; Consistency checking is performed to ensure that the mapping logic of entities and relations is consistent. If there is a conflict relation, mark the conflict information and keep the original data source for manual verification; When new entities or relations are added, they are extended according to the mapping rules above, and the triple storage content is updated.
[0040] S130: Based on the knowledge graph, noise data, complex data, rejection data and multi-hop data are generated by a noise sampling module, a complex attribution module, a rejection analysis module and a thought reasoning module respectively.
[0041] In some embodiments, based on the knowledge graph, noise data, complex data, rejection data and multi-hop data are generated by a data feature module, wherein the data feature module includes four modules of noise sampling, complex attribution, rejection analysis and thought reasoning.
[0042] Since the entity and the relationship between entities have been obtained when constructing the knowledge graph, the noise sampling module can construct a first simple question and answer pair (such as "What is the relationship between entity A and entity B?") based on the entity and the relationship between entities in the knowledge graph, and then select relevant paragraphs with low relevance as noise information from the independent paragraph set through semantic similarity calculation (such as cosine similarity greater than or equal to 0.3), and combine them with the question and answer pair at a predetermined proportion (such as noise ratio 30%-50%) to form noise data containing interference information. For example, the question "What theory did Einstein propose?" + answer "relativity", add "Einstein's late research on unified field theory", "Newton discovered universal gravitation" and other paragraphs with low relevance as noise.
[0043] The complex attribution module needs to manually arrange M (such as 20-30) similar entities from the knowledge graph, which needs to cover at least 5 or more entity categories. For each entity cluster, generate complex question and answer pairs such as "What are the main differences between Apple and Microsoft in product ecology?" Use the associated paragraphs in the entity cluster to construct the context and enhance the model's ability to analyze complex problems.
[0044] The rejection analysis module, similar to the noise sampling module, can construct a second simple question and answer pair through known entities and absolute paragraphs containing explicit answers. In addition to the absolute paragraph, the remaining noise paragraphs have no positive effect on the actual question and answer. This module only uses noise paragraphs as related text, and this module can ensure that the model only answers known information and refuses to answer unknown questions.
[0045] The thought reasoning module is used to associate entity c with entities a and b in the knowledge graph. The entity c can be regarded as a bridging entity of entities a and b. The bridging entity is screened to generate a candidate paragraph pair. For example, “When was the singer of “Lan Hua Cao” born?” The singer of “Lan Hua Cao” needs to be determined first, and then the birth date is searched. This module often contains the relationship among three entities, thereby providing a potential clue for thought reasoning.
[0046] In summary, the data feature module can generate noise data, complex data, rejection data, and multi-hop data to obtain question and answer data that meets the requirements of knowledge retrieval enhancement.
[0047] Exemplarily, the question of the knowledge retrieval enhancement question and answer data needs to meet the following format: { --- Role --- You are a helpful assistant who answers questions about the data in the provided information.
[0048] --- Target --- Generate a reply that meets the target length and format to answer the user's question and summarize all relevant information in the input data to make it suitable for the length and format of the reply.
[0049] You should use the data provided in the following data as the main context for generating the reply.
[0050] If you do not know the answer or there is not enough information in the input data table to provide the answer, please state it directly. Do not fabricate any content.
[0051] The points supported by the data should list their data references as follows: {Reference Example} ---Target Answer Length and Format--- {Answer Example} ---Reference Information--- {Reference Content} } S140: Format conversion is performed on the noise data, complex data, rejection data, multi-hop data, and general data to obtain fine-tuning data, wherein the general data is open source model pre-training question and answer data.
[0052] In some embodiments, to ensure the general question and answer ability of the model, general data, which is open source model pre-training question and answer data, is still needed in addition to noise data, complex data, rejection data, and multi-hop data.
[0053] In addition, since there is a great difference between the question and answer data format and the format required for model fine-tuning, the data preprocessing module is needed to adapt and convert the format. The specific steps of format conversion include: First, a triple data structure containing "instruction" (task description field), "input" (context information field), and "output" (target output field) is called to adapt the input format of mainstream fine-tuning frameworks such as LLaMA Factory and Hugging Face Transformers. According to the preset mapping structure, the noise data, complex data, rejection data, multi-hop data, and general data are mapped to the triple data structure.
[0054] Specifically, the mapping rules are refined as follows: Noise data, complex data, and multi-hop data: the original question is mapped to the "instruction" field, the associated background text is mapped to the "input" field, and the derived standard answer is mapped to the "output" field. The original question of the rejection data is mapped to the "instruction" field, the interference text without the correct answer is mapped to the "input" field, and the preset knowledge rejection identifier is mapped to the "output" field. The original task description of the general data is mapped to the "instruction" field, the original input information is mapped to the "input" field, and the original output result is mapped to the "output" field.
[0055] Finally, the mapped triple data is standardized to obtain fine-tuning data that meets the input requirements of the model fine-tuning framework.
[0056] Exemplary fine-tuning data: { "instruction": "Identify and explain two scientific theories in the given list: cell theory and heliocentrism.", "input": "", "output": "Cell theory is a theory in biology that states that all living organisms are composed of tiny basic units called cells. This is a fundamental theory of biology that considers cells as the basic structural and functional units of all living organisms, and all living organisms are composed of one or more cells. Cells can only produce new cells through cell division. This theory was first proposed by Schrödinger, Schwann and Shelley in 1839. Heliocentrism refers to the theory that the sun is the center of the solar system, that is, the theory that planets revolve around the sun. This theory breaks the traditional geocentric view that the Earth is not the center of the universe. The proposer of the heliocentric theory is Nicolaus Copernicus, who published his work "On the Revolutions of the Heavenly Bodies" in the early 16th century, describing a model of the solar system in which planets revolve around the sun. It has made a great contribution to the development of astronomy." S150: Fine-tune the large language model according to the fine-tuning data to obtain a fine-tuned retrieval enhancement generation model.
[0057] In some embodiments, the open source model is fine-tuned using the fine-tuning data obtained in step S140. This embodiment has no special requirements for open source models, and domestic and foreign open source large language models can be used for fine-tuning, such as Qwen, Deepseek, Llama, etc. Open source large models.
[0058] Exemplarily, taking the LoRA fine-tuning of the llama-factory framework as an example, the optimized fine-tuning parameter settings are as follows: llamafactory-cli train \ --stage sft \ --do_train \ --model_name_or_path local_model\ --dataset local_data\ --dataset_dir. / data \ --template local_model_template \ --finetuning_type lora \ --output_dir. / saves / local_model \ --overwrite_cache \ --overwrite_output_dir \ --cutoff_len 1024 \ --preprocessing_num_workers 16 \ --per_device_train_batch_size 2 \ --per_device_eval_batch_size 1 \ --gradient_accumulation_steps 8 \ --lr_scheduler_type cosine \ --logging_steps 50 \ --warmup_steps 20 \ --save_steps 100 \ --eval_steps 50 \ --evaluation_strategy steps \ --load_best_model_at_end \ --learning_rate 5e-5 \ --num_train_epochs 5.0 \ --max_samples 1000 \ --val_size 0.1 \ --plot_loss \ --fp16 It is worth mentioning that the data balancing strategy is added during the fine-tuning process to ensure that the proportion of noise data, complex data, rejected data, multi-hop data and general data is 2:2:1:2:3, avoiding the model bias caused by the high proportion of a certain type of data.
[0059] S160: Evaluate the fine-tuned retrieval and enhancement generation model according to the answer relevance index and the answer similarity index.
[0060] In some embodiments, the evaluation index of answer relevance focuses on evaluating the close correlation between the generated answer and the input question. This embodiment innovatively uses an indirect but highly effective method, which requires the large language model to generate a related question that may be asked based on the generated answer, forms a question set with the original input question, and then measures the average cosine similarity between the generated related question and the original input question using text similarity calculation technology, thereby determining the answer relevance of the generated answer.
[0061] wherein the specific prompt is "generate an associated question for the given answer and determine if the answer is explicit, if the answer is not explicit, give the noncommittal a value of 1, if the answer is explicit, give the noncommittal a value of 0 (non-explicit answers are evasive, ambiguous, or equivocal responses). For example: I don't know or I am not sure are non-explicit answers.
[0062] Specifically, the answer relevance calculation formula is: ; wherein, is the generated first question corresponding vector, is the user question corresponding vector, is the number of generated questions (generally default is 3).
[0063] The answer similarity index is mainly dedicated to evaluating the semantic similarity between the generated answer and the input question. In the actual evaluation process, the evaluation is strictly based on the ideal dialogue and the generated dialogue, and the value range is also limited between 0 and 1. The higher the score, the better the consistency of the semantic connotation between the generated answer and the standard answer.
[0064] According to the above embodiments of the present disclosure, the following technical effects are achieved: (1) The original text is segmented by predefining the character-level segmentation rule, and the paragraph is adjusted in combination with semantic integrity detection, so as to ensure the semantic integrity of the paragraph, lay a high-quality data foundation for subsequent entity and relationship extraction. This segmentation method avoids the loss of semantics caused by paragraph fragmentation, makes the extraction of entities and relationships between entities from the paragraph more accurate, and thus the knowledge graph constructed is more in line with the real semantics of the text, providing reliable knowledge support for model training. (2) Based on the knowledge graph, noise data, complex data, refusal data and multi-hop data are generated, and fine-tuning data is formed in combination with general data. The diversified data types can specifically improve the ability of the model in different scenarios: noise data can enhance the anti-interference ability of the model to interference information, so that it can still accurately focus on the core problem when there is irrelevant information; complex data helps the model to handle problems involving multiple similar entities or complex relationships, and improves the ability to cope with complex tasks; the introduction of refusal data enables the model to learn to identify questions that cannot be answered, reduces the situation of fabricating answers, and enhances the reliability of the model output; multi-hop data can train the logical reasoning ability of the model, so that it can handle problems that require multiple steps of deduction. The integration of general data ensures that the model has the ability to handle specific scenarios while not losing the general question and answer ability, achieving a balance between special capabilities and general capabilities. (3) The unified format conversion maps multiple data to a triple data structure and normalizes it, so that data of different types and from different sources can adapt to the model fine-tuning framework, improving the efficiency and convenience of the fine-tuning process and enhancing the fine-tuning effect.
[0065] (4) The fine-tuned model is evaluated by combining the answer relevance index and the answer similarity index, which is more comprehensive and scientific than a single index. The answer relevance index ensures that the generated answer is closely related to the question and avoids irrelevant answers. The answer similarity index measures the degree of agreement between the generated answer and the standard answer. The combination of the two can more accurately reflect the quality of the generated answer, providing a reliable basis for evaluating the performance of the model.
[0066] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the disclosure is not limited by the order of the described actions, because according to the disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the disclosure.
[0067] The above is an introduction to the method embodiment. The following describes the scheme of the disclosure through a device embodiment.
[0068] Embodiment Two Figure 3 A block diagram of a large language model fine-tuning and evaluation device based on retrieval enhancement generation according to an embodiment of the disclosure is shown. As shown in Figure 3 The device 300 includes: A text segmentation module 310 is configured to segment the content of an input original text document into multiple independent paragraphs according to a preset character level segmentation rule. A knowledge graph construction module 320 is configured to extract entities and relationships between entities from the multiple independent paragraphs and construct a knowledge graph. A feature data generation module 330 is configured to generate noise data, complex data, rejection data, and multi-hop data based on the knowledge graph through a noise sampling module, a complex attribution module, a rejection analysis module, and a thinking reasoning module, respectively. A data conversion module 340 is configured to convert the noise data, complex data, rejection data, multi-hop data, and general data into a format to obtain fine-tuning data, wherein the general data is open source model pre-training question and answer data. A large language model fine-tuning module 350 is configured to fine-tune a large language model based on the fine-tuning data to obtain a fine-tuned retrieval enhancement generation model. The large language model evaluation module 360 is configured to evaluate the fine-tuned retrieval enhancement generation model according to the answer relevance index and the answer similarity index.
[0069] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0070] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0071] Figure 4 A schematic block diagram of an electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0072] The electronic device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM 402 or a computer program loaded into a RAM 403 from a storage unit 408. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O interface 405 is also connected to the bus 404.
[0073] Various components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc., an output unit 407, such as various types of displays, a speaker, etc., a storage unit 408, such as a magnetic disk, an optical disk, etc., and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0074] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the method 100. For example, in some embodiments, the method 100 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the method 100 described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the method 100 by any other appropriate means, such as by means of firmware.
[0075] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0076] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0077] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0078] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0079] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0080] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0081] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which are not limited herein.
[0082] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for large language model fine-tuning and evaluation based on retrieval-augmented generation, characterized in that, The method comprises the following steps: According to the preset character level segmentation rule, the content of the input original text document is segmented into multiple independent paragraphs; Extracting entities and inter-entity relationships from the multiple independent paragraphs to construct a knowledge graph; Based on the knowledge graph, noise sampling module, complex attribution module, rejection analysis module and thought reasoning module are used to generate noise data, complex data, rejection data and multi-hop data respectively; Format conversion is performed on the noise data, complex data, rejection data, multi-hop data and general data to obtain fine-tuning data, wherein the general data is open source model pre-training question and answer data; According to the fine-tuning data, the large language model is fine-tuned to obtain a fine-tuned retrieval enhancement generation model; According to the answer correlation index and the answer similarity index, the fine-tuned retrieval enhancement generation model is evaluated.
2. The method of claim 1, wherein, According to the preset character level segmentation rule, the content of the input original text document is segmented into multiple independent paragraphs; According to the preset character level segmentation rule, the content of the input original text document is segmented into multiple independent paragraphs; If the last independent paragraph is an incomplete sentence, it is merged into the previous independent paragraph. The method comprises the following steps: Using a large language model, the entities and inter-entity relationships in the paragraph are analyzed based on a preset prompt word template; 3. The method of claim 2, wherein, The extracted entities and inter-entity relationships are mapped into nodes and edges in the knowledge graph to construct the knowledge graph. The preset prompt word template includes target instructions and step-by-step analysis rules, wherein the step-by-step analysis rules include: Identify all entities and extract entity information, including entity name, entity type and entity description; 4. The method of claim 3, wherein, Identify related entity pairs from the identified entities and extract inter-entity relationship information, including source entity name, target entity name, relationship description and relationship strength. The extracted entities and inter-entity relationships are mapped into nodes and edges in the knowledge graph to construct the knowledge graph, which comprises: The entity name is used as the unique identifier of the node, the entity type is used as the classification tag of the node, and the entity description is used as the attribute of the node; 5. The method of claim 4, wherein, The source entity name and the target entity name correspond to the starting node and the ending node of the edge respectively, the relationship description is used as the type of the edge, and the relationship strength is used as the weight of the edge; The mapped nodes and edges are stored in a triple format, and the node attributes and edge weights are associated to form an extended triple structure to construct the knowledge graph. Based on the knowledge graph, noise sampling module, complex attribution module, rejection analysis module and thought reasoning module are used to generate noise data, complex data, rejection data and multi-hop data respectively, which comprises: Based on the entities and inter-entity relationships in the knowledge graph, a first simple question and answer pair is constructed, and the paragraph information related to the first simple question and answer pair is selected from the independent paragraph set and added to the first simple question and answer pair as noise information according to a preset proportion to form noise data; 6. The method of claim 3, wherein, Selecting M entities with similar attributes or relationships from the knowledge graph to form an entity cluster, constructing a complex question-answer pair based on the entity cluster, and forming complex data; Based on the entities in the knowledge graph and the absolute paragraph containing the explicit answer, a second simple question-answer pair is constructed, other paragraphs except the absolute paragraph are selected from the independent paragraph set as the relevant text of the second question-answer pair, and the rejection data is formed; Filtering bridge entities in the knowledge graph, generating candidate paragraph pairs containing three entity relationships based on the bridge entities and associated entities to construct multi-hop question-answer pairs, and forming multi-hop data, wherein the bridge entity is associated with at least two other entities, and there is no direct relationship between the associated entities.
7. The method of claim 6, wherein, Format conversion is performed on the noise data, complex data, rejection data, multi-hop data and general data to obtain fine-tuning data, including: Calling a triple data structure containing a task description field, a context information field and a target output field; According to the preset mapping structure, the noise data, complex data, rejection data, multi-hop data and general data are mapped to the triple data structure; The mapped triple data is normalized to obtain fine-tuning data that meets the input requirements of the model fine-tuning framework.
8. The method of claim 7, wherein, The noise data, complex data, rejection data, multi-hop data and general data are mapped to the triple data structure according to the preset mapping structure, including: Map the original question of the noise data, complex data and multi-hop data to the task description field, the associated background text to the context information field, and the derived standard answer to the target output field; Map the original question of the rejection data to the task description field, the interference text without correct answer to the context information field, and the preset knowledge rejection identifier to the target output field; Map the original task description of the general data to the task description field, the original input information to the context information field, and the original output result to the target output field.
9. The method of claim 1, wherein, The fine-tuned retrieval enhancement generation model is evaluated according to the answer relevance index and the answer similarity index, including: Based on the generated answer, at least one associated question is deduced in reverse, and the associated question and the input original question form a question set; Calculate the semantic association score between each question in the question set as the answer relevance index; Compare the generated answer with the standard answer for semantic equivalence to obtain the answer similarity index.
10. The method of claim 9, wherein, The answer relevance index calculation formula is: ; wherein, is the generated first question corresponding vector, is the user question corresponding vector, is the number of generated questions.
Citation Information
Patent Citations
Automatic evaluation method and system for retrieval enhancement generation system
CN119166785A
Criminal period prediction method based on sentencing rule knowledge graph driving
CN119416987A
Large language model generation method based on retrieval enhancement
CN119782470A
RAG-based construction technology multi-mode intelligent knowledge base question and answer processing method, medium and equipment
CN119938817A
Large model reasoning method and system based on tree diagram and knowledge graph retrieval enhancement
CN119961377A
Cited By
Entity pair guided scientific and technical literature document level relation extraction method and system
CN121501984A