Fine adjustment method and device for large model in vertical field and medium

By extracting the initial triplet from the knowledge graph and using Q&A prompt word engineering to generate Q&A pairs, screening high-quality Q&A pairs for fine-tuning of large models, the shortcomings in semantic accuracy and information integrity of large models in vertical fields are solved, and more efficient Q&A accuracy and professionalism are achieved.

CN120069098AActive Publication Date: 2025-05-30ZHEJIANG LAB

Patent Information

Application Number
CN202510549483.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing large language models are difficult to achieve the expected effects of users in the vertical field, mainly because the fixed generation rules lose complex attributes and relationship information in the knowledge graph, resulting in the generated question-and-answer incomplete semantics or insufficient accuracy.

Method used

By extracting the initial triplet from the knowledge graph of the specified domain, based on the pre-constructed Q&A prompt word engineering and initial triplet, the specified big model is called to generate Q&A pairs, and fine-tuning training is performed on the vertical domain big model by filtering high-quality Q&A pairs through quality scores.

Benefits of technology

It improves the accuracy of question-and-answer models in designated fields, avoids the loss of information in the knowledge graph, and meets the professional needs of specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069098A_ABST
    Figure CN120069098A_ABST
Patent Text Reader

Abstract

The invention discloses a fine tuning method and device for a large model in a vertical field and a medium, and the method comprises the steps: extracting an initial triple from a knowledge graph in a specified field; based on a pre-constructed question and answer prompt word project and the initial triple, calling a specified large model to generate question and answer pairs about the initial triple; determining a quality score for representing the quality of the question-answer pair; the higher the quality score is, the higher the quality of the question-answer pair is; and performing fine tuning training on the vertical field large model through the target question and answer pair with the mass score greater than the threshold. Therefore, based on the question and answer prompt words and the initial triple of the knowledge graph, the specified large model can quickly generate high-quality question and answer pairs, and effective information in the triple is prevented from being lost. In addition, based on screening of question and answer pair quality scores, fine tuning training is carried out on the vertical field large model through high-quality target question and answer pairs, and the question and answer accuracy of the large model in the vertical field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular, to a fine-tuning method, device, and medium for large models in vertical domains. Background Art

[0002] With the rapid development of deep learning technology, large language models are widely used in all aspects of people's lives, such as intelligent driving, sentiment analysis, and creative content generation. However, current large language models are usually general models and cannot meet the needs of vertical domains. Vertical domains usually require in-depth professional background knowledge to accurately understand and process relevant information. For example, in vertical domains such as geology, medicine, and law, there are a large number of professional terms and complex concept systems, and simply relying on general large language models often fails to achieve the expected results of users.

[0003] In order to meet the actual business needs of large models in different domains, based on fixed generation rules, triples in the knowledge graph of the vertical domain can be converted into natural language question-answer pairs, and the large language model can be fine-tuned through the question-answer pairs to obtain a large model applied to the vertical domain. However, in the process of converting triples into natural language question-answer pairs, the fixed generation rules will lose complex attribute and relationship information in the knowledge graph, resulting in incomplete semantics or insufficient accuracy of the generated question-answer pairs, so that the accuracy of the large model in the vertical domain is still difficult to achieve the expected results of users.

[0004] Therefore, how to improve the accuracy and interaction ability of large models in vertical domains and meet the business needs of vertical domains is an urgent problem for those skilled in the art. Summary of the Invention

[0005] In view of this, one aspect of this application provides a fine-tuning method for a large model in a vertical domain, and the method includes: Extract initial triples from the knowledge graph of the specified domain; Based on the pre-constructed question-answering prompt engineering and the initial triples, call a specified large model to generate question-answer pairs about the initial triples; Determine a quality score for characterizing the quality of the question-answer pairs; the larger the quality score, the higher the quality of the question-answer pairs; Fine-tune and train the large model in the vertical domain through the target question-answer pairs whose quality scores are greater than the threshold.

[0006] Optionally, the determining the quality score for characterizing the quality of the question-answer pairs includes: Obtain the pre-constructed inference prompt engineering; Mask the initial triples to obtain masked triples; the initial triples include a subject field, a relationship field, and a predicate field; Based on the above-mentioned inference prompt engineering, the masked triples are inferred by the specified large model to obtain the completed triples; Determine the quality score according to the initial triples and the completed triples.

[0007] Optionally, masking the initial triples to obtain masked triples includes: Masking any one or any two of the subject field, the relation field, and the predicate field to obtain the masked triples.

[0008] Optionally, the masked triples include a first masked triple obtained by masking the relation field and a second masked triple obtained by masking the predicate field; Correspondingly, the completed triples include a first completed triple corresponding to the first masked triple and a second completed triple corresponding to the second masked triple.

[0009] Optionally, determining the quality score according to the initial triples and the completed triples includes: Obtain a pre-constructed similarity scoring prompt engineering; Based on the similarity scoring prompt engineering, call a target large language model to perform similarity scoring on the initial triples and the first completed triple and the second completed triple respectively to obtain a first similarity score and a second similarity score; Take the average of the first similarity score and the second similarity score as the quality score; where the higher the average, the higher the quality of the Q&A pair corresponding to the initial triple.

[0010] Optionally, calling a specified large model to generate Q&A pairs for the initial triples includes: For each of the initial triples, generate a preset number of groups of initial Q&A pairs; the initial Q&A pairs include at least one of open-ended Q&A pairs and closed-ended Q&A pairs; Merge the initial Q&A pairs to obtain the Q&A pairs of the initial triples.

[0011] Optionally, extracting the initial triples from the knowledge graph of a specified domain includes: Extract entity nodes and the relationship information between the entity nodes from the knowledge graph; Perform data cleaning on the entity nodes and the relationship information; Construct the entity nodes and the relationship information after data cleaning into the initial triples in a preset format.

[0012] Another aspect of the present application provides a fine-tuning device for a vertical domain large model, the device comprising: A triple extraction module, configured to extract initial triples from a knowledge graph of a specified domain; A question-and-answer pair generation module, configured to call a specified large model to generate question-and-answer pairs about the initial triples based on a pre-constructed question-and-answer prompt engineering and the initial triples; A quality determination module, configured to determine a quality score for characterizing the quality of the question-and-answer pairs; the larger the quality score, the higher the quality of the question-and-answer pairs; A fine-tuning training module, configured to fine-tune and train a vertical domain large model with target question-and-answer pairs whose quality scores are greater than a threshold.

[0013] Another aspect of the present application provides a fine-tuning device for a vertical domain large model, comprising a memory and a processor, wherein a computer program executable on the processor is stored on the memory, and when the processor executes the program, the steps of the fine-tuning method for the vertical domain large model are implemented.

[0014] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the fine-tuning method for the vertical domain large model are implemented.

[0015] The beneficial effects of the fine-tuning method, device and medium for a vertical domain large model provided by the present application are as follows: Based on the initial triples of the knowledge graph in the specified domain and the question-and-answer prompts, the specified large model can quickly learn and reason about the knowledge in the initial triples, generate high-quality question-and-answer pairs, and avoid losing the effective information in the triples. In addition, through the screening based on the quality scores of the question-and-answer pairs, the vertical domain large model is fine-tuned and trained with high-quality target question-and-answer pairs, thereby improving the question-and-answer accuracy of the large model in the vertical domain, and the large model fine-tuning training can be completed according to the knowledge graphs of different domains, so as to meet the professional needs of specific domains. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application; Figure 2 It is a schematic principle diagram of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application; Figure 3 It is a schematic principle diagram of a triple mask provided by an embodiment of the present application; Figure 4 It is a schematic principle diagram of a fine-tuning method for a vertical domain large model provided by another embodiment of the present application; Figure 5Schematic structural diagram of a fine-tuning device for a vertical domain large model provided by an embodiment of the present application; Figure 6 Schematic structural diagram of a fine-tuning device for a vertical domain large model provided by another embodiment of the present application.

[0017] The reference numerals are as follows: 60 is a memory, 61 is a processor, 62 is a display screen, 63 is an input / output interface, 64 is a communication interface, 65 is a power supply, 66 is a communication bus, 601 is a computer program, 602 is an operating system, and 603 is data. Detailed implementation manners

[0018] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0019] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0020] Figure 1 Schematic flow diagram of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application, as Figure 1 shown, the method includes: S10: Extract initial triples from the knowledge graph of a specified domain; Figure 2 Schematic principle diagram of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application, as Figure 2 shown, in a specific embodiment, obtain the knowledge graph of a specified domain, and the specified domain may include, but is not limited to, geology, medicine, and law. The knowledge graph may be pre-constructed or an open-source knowledge graph obtained. That is, the present application does not limit the acquisition methods of the specified domain and the knowledge graph.

[0021] After obtaining the knowledge graph, in order to obtain the professional knowledge in the specified field for subsequent fine-tuning of the large model, that is, to let the large model learn the professional knowledge in the specified field so that the large model has the ability to answer questions in the specified field. In an alternative embodiment, initial triples are extracted from the knowledge graph.

[0022] Specifically, the entity nodes in the knowledge graph and the relationship information between the entity nodes can be first identified, and the entity nodes and the relationship information are combined to obtain the initial triples. It should be noted that in an alternative embodiment, in order to ensure the quality of the initial triples, the extracted initial triples can be filtered and screened to retain high-quality initial triples. For ease of understanding, the initial triples are illustrated below with examples. For example, the initial triple is (Apple, belongs to, Fruit).

[0023] S11: Based on the pre-constructed Q&A prompt engineering and the initial triples, call the specified large model to generate Q&A pairs about the initial triples; Furthermore, in order to enable the specified large model to quickly learn the professional knowledge in the initial triples and generate reliable question-answer pairs (i.e., Q&A pairs), in an alternative embodiment, a Q&A prompt (Prompt) engineering is pre-constructed. Under the guidance of the Q&A Prompt engineering, as Figure 2 shown, the specified large model learns and reasons about the initial triples, thereby generating Q&A pairs about the initial triples.

[0024] For example, if the initial triple is (Apple, belongs to, Fruit), the Q&A pairs that the specified large model can generate include: Q (Question): "What category does an apple belong to?", A (Answer): "An apple belongs to the fruit category."

[0025] It should be noted that the specified large model includes but is not limited to ChatGPT, LLaMA2, and LLaMA3, etc. This application does not limit the specified large model, and it can be selected according to actual business needs. For example, in a vertical field, the large model corresponding to the vertical field can be selected as the specified model, and the large model can be fine-tuned and trained through the fine-tuning method of this application, thereby improving the performance of the model in the vertical field.

[0026] It is understandable that Q&A Prompt engineering is used to guide a specified large model to generate natural language Q&A pairs based on the provided initial triples. Therefore, Q&A Prompt engineering is crucial for the quality of Q&A pairs. In a specific embodiment, Q&A Prompt engineering is concatenated with the initial triples to generate an input text, and the input text is input into the specified large model for inference to generate Q&A pairs. Q&A Prompt engineering requires the specified large model to make full use of the information provided by the initial triples to generate Q&A pairs with true content and natural language, and output them in a specified format (for example, Json format). For ease of understanding, Q&A Prompt engineering will be illustrated with examples below.

[0027] Q&A Prompt engineering: <Please generate Q&A pairs based on the initial triples I provided and meet the following requirements: 1. Generate Q&A pairs using the provided initial triples.

[0028] 2. Add the content you know to make the Q&A pairs as rich as possible. At the same time, please ensure that you utilize all the information provided in the initial triples.

[0029] 3. Ensure that the generated content is true.

[0030] 4. Conform to the laws of natural language in the Q&A pairs.

[0031] 5. Please output in Json format: { "Question": "……"; "Answer": "……"; } The following are the initial triples I provided for you: (Apple, belongs to, fruit). > S12: Determine the quality score used to characterize the quality of Q&A pairs; the larger the quality score, the higher the quality of the Q&A pairs; S13: Fine-tune and train the vertical domain large model with the target Q&A pairs whose quality scores are greater than the threshold.

[0032] Furthermore, select high-quality target Q&A pairs from all the Q&A pairs generated by the specified large model and fine-tune and train the vertical domain large model, that is, use the target Q&A pairs as the training data set to fine-tune and train the vertical domain large model, so as to obtain a large model with high precision in the specified domain.

[0033] In an alternative embodiment, the calculation of the quality score can be based on multiple dimensions. For example, the accuracy, completeness, logic of the answer, and whether it conforms to the professional knowledge within the field. In an alternative embodiment, if the generated answer completely conforms to the facts within the specified field and is clearly and completely expressed, then the quality score is relatively high. On the contrary, if the answer contains errors or is ambiguously expressed, the quality score will be relatively low. The larger the quality score, the higher the quality of the generated question-answer pair, and the more it can meet the requirements for accuracy and professionalism in practical applications.

[0034] Therefore, as Figure 2 shown, by screening out the target question-answer pairs with a quality score greater than the threshold, fine-tuning training is performed on the large model for the vertical domain. Among them, the purpose of fine-tuning training is to enable the large model to better adapt to the knowledge and language style of the vertical domain and improve its question-answer performance in this domain. After fine-tuning, the large model can more accurately understand and generate question-answer content related to the vertical domain, thereby improving the quality and practicality of the entire question-answer system.

[0035] It should be noted that the large model for the vertical domain can be the specified large model in step S11 or other models, and this application does not make any limitations in this regard. In addition, the vertical domain includes but is not limited to the geological field, medical field, legal field, sports field, etc.

[0036] In an alternative embodiment, in order to further improve the quality of the target question-answer pairs, after selecting high-quality question-answer pairs according to the instruction score, the target question-answer pairs can be further verified and corrected manually. For example, correcting errors and optimizing expressions to ensure the quality of the target question-answer pairs and further improving the accuracy of the large model for the vertical domain in the specified field.

[0037] In another alternative embodiment, a high-quality test set data in the specified field can be pre-constructed, which can be a test set obtained through manual annotation. In a specific embodiment, after training the large model for the vertical domain with the above-mentioned automatically generated target question-answer pairs, the performance of the model can be evaluated and fine-tuned through the test set data for indicators such as factuality, fluency, and interpretability, further improving the performance of the large model for the vertical domain.

[0038] Thus, the fine-tuning method for the large model in the vertical domain provided by the embodiments of this application is based on the initial triples of the knowledge graph in the specified field. Based on the question-answer prompt words, the specified large model can quickly learn and reason about the knowledge in the initial triples, generate high-quality question-answer pairs, and avoid losing the effective information in the triples. In addition, through the screening based on the quality score of the question-answer pairs, fine-tuning training is performed on the large model for the vertical domain with high-quality target question-answer pairs, thereby improving the question-answer accuracy of the large model in the vertical domain. Fine-tuning training of the large model can be completed according to the knowledge graphs of different fields, so as to meet the professional needs of specific fields.

[0039] In an alternative embodiment, determining a quality score for characterizing the quality of a question-and-answer pair includes: Obtain a pre-constructed inference prompt word project; Mask the initial triple to obtain a masked triple; the initial triple includes a subject field, a relation field, and a predicate field; Based on the inference prompt word project, use a specified large model to infer the masked triple to obtain a completed triple; Determine the quality score according to the initial triple and the completed triple.

[0040] Figure 3 FIG. is a schematic diagram of the principle of triple masking provided by an embodiment of the present application. As Figure 3 shown, in a specific embodiment, the initial triple extracted from the knowledge graph is masked to obtain a masked triple. Among them, masking the initial triple means hiding or replacing some field information in the initial triple by means of masking.

[0041] The initial triple includes a subject field, a relation field, and a predicate field. Therefore, masking the initial triple means hiding any one or more of the subject field, the relation field, and the predicate field. It should be noted that, in order to subsequently use a specified large model to infer the masked triple and thus complete the masked field, it can be understood that only one or two of the subject field, the relation field, and the predicate field can be masked to ensure that the specified large model can infer based on at least one piece of information in the masked triple. For ease of understanding, an example will be given below by masking one of the fields.

[0042] For example, the initial triple is (water, molecular formula, H 2 O), where water is the subject field, molecular formula is the relation field, and H 2 O is the predicate field. In an alternative embodiment, the predicate field can be masked to obtain a masked triple of (water, molecular formula, *), where "*" means hiding the predicate field.

[0043] In an alternative embodiment, in order to quickly infer and complete the masked triple, that is, restore the masked (i.e., hidden) field in the masked triple, a pre-constructed inference Prompt project can be built. The inference Prompt project is used to guide a specified large model to complete the masked part in the masked triple. Specifically, as Figure 3 described, based on the inference Prompt project, use the masked triple as the input of the specified large model, and use the specified large model to infer the masked triple, so as to obtain a completed triple with the masked field completed.

[0044] It is understandable that in the above embodiments, the specified large model obtains question-answer pairs through inferential learning of the initial triples. Therefore, the specified large model has learned the knowledge in the initial triples at this time. To detect whether the generated question-answer pairs are high-quality data, the specified large model can be used to complete the masked triples. If the restored completed triples are exactly the same as the initial triples, it indicates that the specified large model has learned the knowledge in the initial triples, and the corresponding generated question-answer pairs are of high quality. On the contrary, if there is a large semantic difference between the completed triples and the initial triples, it indicates that the quality of the generated question-answer pairs is low. Therefore, in an alternative embodiment, as Figure 3 shown, the quality score of the corresponding question-answer pairs can be calculated based on the initial triples and the completed triples. For ease of understanding, an example of the inference Prompt engineering will be given below.

[0045] For example, the inference Prompt engineering is as follows: <Please fill in the missing part (* part) of the masked triple according to the question-answer pairs I provide: "Question": "What is the molecular formula of water?" "Answer": "The molecular formula of water is H 2 O".

[0046] The masked triple is: (water, molecular formula, *)> Thus, the fine-tuning method of the vertical domain large model provided by the embodiments of the present application realizes the rapid completion of the masked triples based on the inference Prompt engineering, combines the completed triples and the initial triples, and evaluates whether the quality of the obtained question-answer pairs meets the expectations, so as to efficiently extract high-quality question-answer pairs and provide a high-quality training data set basis for subsequent vertical domain large model training.

[0047] Based on the above embodiments where the initial triples include a subject field, a relationship field, and a predicate field, as an alternative embodiment, the initial triples are masked to obtain masked triples, including: Masking any one or any two of the subject field, the relationship field, and the predicate field to obtain masked triples.

[0048] It is understandable that the obtained masked triples will be inferred later to restore the masked fields. Therefore, when masking the initial triples, at least one field should be retained for subsequent inference. That is, when masking, any one or any two of the subject field, the relationship field, and the predicate field are masked.

[0049] However, it can also be understood that in a specific embodiment, the subject field is very important for the inference of the mask field. If the subject field is masked, the result of subsequent inference may deviate significantly from the corresponding initial triple, which is not conducive to the evaluation of the quality of the question-answer pair.

[0050] Meanwhile, if multiple fields among the subject field, the relation field, and the predicate field are masked, that is, two fields are masked and only one field remains, it will result in a large number of completed triples, which is also not conducive to the evaluation of the quality of the question-answer pair and will greatly affect the quality evaluation efficiency.

[0051] Therefore, in an alternative embodiment, masking the initial triple to obtain a masked triple includes: masking the relation field to obtain a first masked triple, and masking the predicate field to obtain a second masked triple. That is, only one field in the initial triple is masked. Thus, the masked triple includes the first masked triple and the second masked triple. The corresponding completed triples include the first completed triple corresponding to the first masked triple and the second completed triple corresponding to the second masked triple.

[0052] It should be noted that in a specific embodiment, to ensure the accuracy of the quality score evaluation, the relation field and the predicate field can be masked separately, and the quality of the corresponding question-answer pair can be evaluated based on the generated first masked triple and second masked triple. Of course, to ensure the efficiency of the quality evaluation, in an alternative embodiment, only one of the relation field and the predicate field can be masked, and this application does not make a limitation on this.

[0053] Based on the above embodiments, as an alternative embodiment, determining the quality score according to the initial triple and the completed triple includes: Obtaining a pre-constructed similarity scoring prompt engineering; Based on the similarity scoring prompt engineering, calling a target large language model to perform similarity scoring on the initial triple and the first completed triple, and the second completed triple respectively, to obtain a first similarity score and a second similarity score; Taking the mean of the first similarity score and the second similarity score as the quality score; where the higher the mean, the higher the quality of the question-answer pair corresponding to the initial triple.

[0054] In an alternative embodiment, to quickly determine the quality score of the question-answer pair, a similarity scoring Prompt engineering regarding the similarity between the initial triple and the completed triple can be pre-constructed. It can be understood that the similarity scoring Prompt engineering is used to guide the target large language model to score the similarity between the initial triple and the completed triple.

[0055] Specifically, based on the similarity scoring Prompt project, the target large language model is called to score the initial triple and the first completed triple to obtain a first similarity score, and to score the initial triple and the second completed triple to obtain a second similarity score.

[0056] Furthermore, since the obtained similarity scores include multiple ones, when evaluating the quality of the corresponding question-answer pair, the average of the first similarity and the second similarity can be used as the quality score of the corresponding question-answer pair. It is worth noting that when the average is higher, the higher the semantic similarity between the initial triple and the completed triple is, the higher the quality of the corresponding generated question-answer pair is. For ease of understanding, the following example of the similarity scoring Prompt project is given by taking the completed triple as one, that is, only the relationship field or predicate field is masked.

[0057] The similarity score prompt project is: <I will provide two sets of triplets. The triplets in these two sets correspond to each other in order, but their objects or topics may be different. Please score their similarity according to the following rules.

[0058] 1. Scoring Rules If two triples are exactly the same, a score of 100 is assigned.

[0059] If two triples are not identical, score them based on how close they are in meaning and how they are expressed, assigning higher scores to those that are closely related in meaning and lower scores to those that are more different, on a scale of 1 to 99.

[0060] If the two meanings are completely different, assign 0 points.

[0061] 2. Make sure the scores in the output are the same and in the same order as the three scores in the set I provided, and then directly output the similarity scores.

[0062] The output format is: [score 1, score 2, ...] Triples include: (apple, belongs to, fruit), (apple, is, fruit)> Therefore, based on the similarity scoring Prompt project, the similarity score between the initial triple and the completed triple can be quickly output. Of course, in another optional embodiment, the similarity between the initial triple and the completed triple can also be calculated by a semantic similarity algorithm, and the quality score can be determined based on the calculated similarity value. This application does not limit the quality scores corresponding to different similarity values, but the corresponding quality score values ​​must comply with the rule that the larger the similarity value, the larger the quality score, which indicates that the quality of the question and answer pair is higher.

[0063] In an alternative embodiment, to improve the quality of generating question-answer pairs, a specified large model is called to generate question-answer pairs for the initial triples, including: For each initial triple, a preset number of groups of initial question-answer pairs are generated; the initial question-answer pairs include at least one of open-ended question-answer pairs and closed-ended question-answer pairs; The initial question-answer pairs are merged to obtain the question-answer pairs of the initial triples.

[0064] Figure 4 It is a schematic diagram of the principle of a fine-tuning method for a vertical domain large model provided by another embodiment of the present application. In a specific embodiment, multiple rounds of question-answer pair generation can be performed for each initial triple to ensure the diversity of the generated question-answer pairs. For example, as Figure 4 shown, the initial triple A is spliced based on the question-answer Prompt engineering to generate input text 1 and input text 2, and the initial question-answer pairs 1 and 2 are obtained through inference by the specified large model. Further, after merging the initial question-answer pairs 1 and 2, the question-answer pairs corresponding to the initial triple are obtained.

[0065] It should be noted that in a specific embodiment, the same initial triple may generate multiple different question-answer pairs. When merging the initial question-answer pairs, the semantic similarity between two question-answer pairs can be determined first, and further, the question-answer pairs with semantic similarity reaching the similarity threshold are merged, that is, the question-answer pairs with high similarity are merged.

[0066] In an alternative embodiment, the initial triples are extracted from the knowledge graph of the specified domain, including: Entity nodes and the relationship information between entity nodes are extracted from the knowledge graph; Data cleaning is performed on the entity nodes and relationship information; The entity nodes and relationship information after data cleaning are constructed into initial triples in a preset format.

[0067] In a specific embodiment, the knowledge graph of the specified domain can be obtained from various open-source data sources (such as text, databases, web pages, etc.), and entity nodes and the relationship information between entity nodes are extracted from the knowledge graph. Among them, the knowledge graph G=(V, R, E), where V is the set of entity nodes, R is the set of relationship information, and E is the set of triples. Specifically, according to the actual business requirements, entity nodes that meet the specified domain are screened out from the knowledge graph. For example, in the scientific field, key entity nodes such as "molecule", "reaction", and "experiment" are selected. Further, among the entity nodes, relationship information that is highly relevant to the actual business requirements is screened out. For example, in the scientific field, core relationship information such as "reaction conditions", "experimental methods", and "molecular structure" is screened out.

[0068] Furthermore, in order to improve the quality of subsequent initial triples, data cleaning is performed on entity nodes and relationship information, and the data cleaning includes but is not limited to removing noisy data. The cleaned entity nodes and relationship information are combined into initial triples in a preset format, thereby completing the construction of the initial triples. Among them, the preset format can be the standard format in the triples, that is, the standard format of (subject, relationship, predicate). Of course, it can also be constructed based on other orders, formats, etc., and this application does not make any limitations in this regard.

[0069] In order to make the technical solution of this application clearer to those skilled in the art, the following will take the subject field as Juan Antonio Pizzi as an example for illustration.

[0070] The initial triples include: (Juan Antonio Pizzi, place of birth, Santa Fe), (Santa Fe, country, Argentina), (Juan Antonio Pizzi, sports country, Spain).

[0071] The masked triples obtained by masking the initial triples include: (Juan Antonio Pizzi, place of birth, *), (Santa Fe, country, *), (Juan Antonio Pizzi, sports country, *).

[0072] Based on the Q&A Prompt engineering, the large model is specified to generate Q&A pairs for the initial triples as: "Question": "Where was Juan Antonio Pizzi born and in which countries during his sports career", "Answer": "Juan Antonio Pizzi was born in Santa Fe, Argentina, and during his sports career, he has been in Spain." Furthermore, based on the inference Prompt engineering, the masked triples are inferred to obtain the completed triples including: (Juan Antonio Pizzi, place of birth, Santa Fe Province), (Argentina, republic, Santa Fe), (Juan Antonio Pizzi, Spanish national team, sports national team).

[0073] Furthermore, based on the similarity scoring Prompt engineering, the similarity score between the initial triples and the completed triples is calculated, and then the quality score representing the quality of the Q&A pairs is determined. In a specific embodiment, the data with a quality score greater than a threshold, for example, greater than 80 points, can be used as the training data set to fine-tune the large model in the vertical domain, thereby improving the Q&A accuracy of the large model in the specified domain.

[0074] In addition, to enable those skilled in the art to more clearly understand the fine-tuning results of the vertical domain large model fine-tuning method provided by this application, in a specific embodiment, taking the knowledge graph AristoV4 as an example, the fine-tuning results obtained by the method provided by this application are introduced. Among them, AristoV4 is a knowledge graph in the scientific field, containing high-precision triples for the basic science field. AristoV4 contains more than 40,000 entities and 280,000 triples of knowledge, which can be used to answer science-related questions. Table 1 is a schematic table showing the number and score of generated question-and-answer pairs provided by the embodiments of this application. Through the large model fine-tuning method provided by this application, question-and-answer pairs 5 times the size were generated on AristoV4, and the specific quantity and quality are shown in Table 1.

[0075] Table 1 is a schematic table showing the number and score of generated question-and-answer pairs

[0076] Table 2 is a schematic table showing the results of different target-generated pairs and the accuracy of the large model provided by the embodiments of this application. Further, through the target question-and-answer pairs of different qualities in the above table, the vertical domain large model is fine-tuned to obtain the comparison results shown in Table 2. Among them, the vertical domain large model is illustrated by taking Llama2 and Llama3 as examples.

[0077] Table 2 is a schematic table showing the results of different target-generated pairs and the accuracy of the large model

[0078] Table 3 is a schematic table showing the results of different training data and model accuracy provided in the embodiments of this application. Through the above-generated question-and-answer pairs, the large models Llama2 and Llama3 are fine-tuned and trained. At the same time, as shown in Table 3, a non-fine-tuning method, a triple fine-tuning method, a retrieval answer fine-tuning method, and a long text fine-tuning method are also carried out. Among them, the non-fine-tuning method refers to a method of not fine-tuning the model, the triple fine-tuning method refers to a method of directly fine-tuning the model through triples, the retrieval answer fine-tuning refers to a method of fine-tuning the model through retrieval results and retrieval questions, and the long text fine-tuning refers to a method of converting triples into long texts to fine-tune the model.

[0079] Table 3 is a schematic table showing the results of different training data and model accuracy

[0080] As can be seen from Table 1, as the data scale (i.e., the number of initial triples) increases, the number of Q&A pairs with different quality scores of high quality increases significantly, but the proportion of high-quality data remains relatively stable (i.e., the proportion of 0 to 60 points is about 73%, the proportion of 60 to 80 points and above is about 48%, and the proportion of 80 to 100 points is about 27%). It can be seen from this that the technical solution provided by this application can maintain a high quality of Q&A pairs under different data scales. At 5 times the scale, there are more than 130,000 high-quality Q&A pairs with a score of 60 or above, providing sufficient high-quality data support for model fine-tuning. That is, the method provided by this application can generate high-quality Q&A pairs under different data scales, and the proportion of high-quality data remains stable, providing reliable data support for model fine-tuning.

[0081] As can be seen from Table 2, even when using Q&A pairs with scores of 0 to 60 for fine-tuning training, the model performance has a significant improvement (among them, the large model Llama2 is improved from 7.56 to 7.62, and the large model Llama3 is improved from 8.42 to 8.67). It can be seen from this that the method provided by this application can effectively screen out high-quality data for model fine-tuning. According to Table 2, the model performance of fine-tuning with Q&A pairs with scores of 80 to 100 is the highest. It can be seen that the method provided by this application can efficiently screen out high-quality Q&A pairs and achieve high-precision fine-tuning training of the model.

[0082] As can be seen from Table 3, on the large models Llama2 and Llama3, the fine-tuning results of Q&A pairs are higher than those of other fine-tuning methods. That is, the method provided by this application has significant advantages in improving model performance. Especially on the large model Llama3, the Q&A pair fine-tuning method provided by this application is significantly higher than other fine-tuning methods.

[0083] Therefore, the fine-tuning method of the vertical domain large model provided by this application realizes the efficient conversion of knowledge graph triples to natural language Q&A pairs through multi-level Prompt engineering, reducing information loss. And based on the triple completion task and similarity scoring mechanism, it quantifies the knowledge integrity and semantic accuracy of Q&A pairs to ensure data quality, generates Q&A pairs through an automated process and conducts quality evaluation, significantly reducing labor costs and time costs. At the same time, by generating Q&A pairs for a specified domain, it improves the adaptability and practicality of the large model in fields with high factuality requirements such as medical, geological, and legal.

[0084] In the above embodiments, the fine-tuning method of the vertical domain large model is described in detail. This application also provides an embodiment corresponding to a fine-tuning device for the vertical domain large model.

[0085] Figure 5 It is a schematic structural diagram of a fine-tuning device for a vertical domain large model provided by an embodiment of this application, as Figure 5As shown, the device includes: A triple extraction module 50 for extracting initial triples from a knowledge graph in a specified domain; A question-and-answer pair generation module 51 for calling a specified large model to generate question-and-answer pairs about the initial triples based on a pre-constructed question-and-answer prompt engineering and the initial triples; A quality determination module 52 for determining a quality score for characterizing the quality of the question-and-answer pairs; the larger the quality score, the higher the quality of the question-and-answer pairs; A fine-tuning training module 53 for fine-tuning and training a vertical domain large model with target question-and-answer pairs whose quality scores are greater than a threshold.

[0086] In addition, the fine-tuning device for the vertical domain large model provided by the embodiments of the present application further includes: A prompt engineering acquisition module for acquiring a pre-constructed inference prompt engineering; A masking module for masking the initial triples to obtain masked triples; the initial triples include a subject field, a relationship field, and a predicate field; A masked inference module for inferring the masked triples through a specified large model based on the inference prompt engineering to obtain completed triples; A quality score determination module for determining a quality score according to the initial triples and the completed triples.

[0087] A first masking sub-module for masking any one or any two of the subject field, the relationship field, and the predicate field to obtain masked triples.

[0088] The prompt engineering acquisition module is further used to acquire a pre-constructed similarity scoring prompt engineering; A similarity scoring model for calling a target large language model based on the similarity scoring prompt engineering to perform similarity scoring on the initial triples and a first completed triple and a second completed triple respectively to obtain a first similarity score and a second similarity score; An average value determination module for using the average value of the first similarity score and the second similarity score as the quality score; where the higher the average value, the higher the quality of the question-and-answer pairs corresponding to the initial triples.

[0089] An initial question-and-answer pair generation module for generating a preset number of groups of initial question-and-answer pairs for each initial triple; the initial question-and-answer pairs include at least one of open-ended question-and-answer pairs and closed-ended question-and-answer pairs; A merging module for merging the initial question-and-answer pairs to obtain question-and-answer pairs of the initial triples.

[0090] An extraction module for extracting entity nodes and relationship information between entity nodes from the knowledge graph; A data cleaning module for cleaning data of entity nodes and relationship information; A triple construction module for constructing the entity nodes and relationship information after data cleaning into initial triples in a preset format.

[0091] Figure 6 The structural schematic diagram of a fine-tuning device for a vertical domain large model provided by another embodiment of this application is as Figure 6 shown. The fine-tuning device for the vertical domain large model includes: a memory 60 for storing computer programs; A processor 61 for implementing the steps of the fine-tuning method of the vertical domain large model mentioned in the above embodiment when executing the computer program.

[0092] The fine-tuning device for the vertical domain large model provided in this embodiment may include but is not limited to a laptop computer or a desktop computer, etc.

[0093] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 may be implemented in at least one hardware form of a digital signal processor (DSP for short), a field programmable gate array (FPGA for short), and a programmable logic array (PLA for short). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU for short); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU for short), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may further include an artificial intelligence (AI for short) processor, and the AI processor is used to process computational operations related to machine learning.

[0094] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may further include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601. After the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the fine-tuning method of the vertical domain large model disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may further include an operating system 602 and data 603, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include, but is not limited to, the relevant data involved in the fine-tuning method of the vertical domain large model.

[0095] In some embodiments, the fine-tuning device of the vertical domain large model may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.

[0096] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the fine-tuning device of the vertical domain large model, and may include more or fewer components than those shown in the figure.

[0097] The fine-tuning device of the vertical domain large model provided by the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the fine-tuning method of the vertical domain large model in the above embodiment.

[0098] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

Claims

1. A method for fine-tuning a large vertical model, characterized in that: The method comprises: Extract initial triples from the knowledge graph of the specified field; Based on the pre-built question-answer prompt word project and the initial triples, calling a designated large model to generate question-answer pairs about the initial triples; Determine a quality score for characterizing the quality of the question-answer pair; the larger the quality score, the higher the quality of the question-answer pair; The vertical domain large model is fine-tuned and trained using the target question-answer pairs whose quality scores are greater than a threshold.

2. The method for fine-tuning a vertical domain large model as claimed in claim 1, characterized in that: The determining of a quality score for characterizing the quality of the question-answer pair comprises: Get pre-built inference cue word projects; Masking the initial triple to obtain a masked triple; the initial triple includes a subject field, a relation field and a predicate field; Based on the inference prompt word project, the mask triplet is inferred through the specified large model to obtain a completed triplet; The quality score is determined according to the initial triple and the completed triple.

3. The fine-tuning method for a vertical domain large model as claimed in claim 2, characterized in that: The step of masking the initial triplet to obtain a masked triplet includes: Mask any one or any two of the subject field, the relation field and the predicate field to obtain the masked triplet.

4. The method for fine-tuning a large vertical domain model as claimed in claim 2, characterized in that: The mask triples include a first mask triple for masking the relation field and a second mask triple for masking the predicate field; Correspondingly, the completion triplet includes a first completion triplet corresponding to the first mask triplet, and a second completion triplet corresponding to the second mask triplet.

5. The method for fine-tuning a large vertical domain model as claimed in claim 4, characterized in that: The determining the quality score according to the initial triplet and the completed triplet includes: Get a pre-built similarity scoring prompt word project; Based on the similarity scoring prompt word project, calling the target large language model, performing similarity scoring on the initial triple, the first completion triple, and the second completion triple respectively, to obtain a first similarity score and a second similarity score; The average of the first similarity score and the second similarity score is used as the quality score; wherein, when the average is higher, the quality of the question-answer pair corresponding to the initial triplet is higher.

6. The method for fine-tuning a large vertical domain model as claimed in claim 1, characterized in that: Calling a specified large model to generate a question-answer pair about the initial triples includes: For each of the initial triples, a preset number of initial question-answer pairs are generated; the initial question-answer pairs include at least one of an open-ended question-answer pair and a closed-ended question-answer pair; The initial question-answer pairs are merged to obtain the question-answer pairs of the initial triples.

7. The fine-tuning method for a vertical domain large model as claimed in claim 1, characterized in that: The step of extracting initial triples from the knowledge graph of the specified field includes: Extracting entity nodes and relationship information between the entity nodes from the knowledge graph; Performing data cleaning on the entity nodes and the relationship information; The entity nodes and the relationship information after data cleaning are constructed into the initial triples in a preset format.

8. A fine-tuning device for a large vertical model, characterized in that: The device comprises: The triple extraction module is used to extract initial triples from the knowledge graph of a specified field; A question-answer pair generation module, configured to call a designated large model to generate a question-answer pair about the initial triplet based on a pre-built question-answer prompt word project and the initial triplet; A quality determination module, used to determine a quality score for characterizing the quality of the question-answer pair; the larger the quality score, the higher the quality of the question-answer pair; The fine-tuning training module is used to fine-tune the vertical domain large model through the target question-answer pairs whose quality scores are greater than the threshold.

9. A fine-tuning device for a large vertical model, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method for fine-tuning the vertical field large model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for fine-tuning a vertical field large model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Question and answer pair generation method and system based on large language model

    CN118332086A

  • Question answering method, system and equipment based on domain-specific knowledge graph and medium

    CN118897886A

  • Large model training method and device, electronic equipment, storage medium and program product

    CN119293500A

  • Model fine tuning method and device, electronic equipment and storage medium

    CN119337832A

  • Knowledge graph open domain construction and RAG question and answer method and device and storage medium

    CN119416882A

Cited By

  • RAG system evaluation data set automatic synthesis method and device based on reinforcement learning

    CN120596663A

  • Reinforcement learning based rag system evaluation dataset automatic synthesis method and apparatus

    CN120596663B

  • Training method, device and equipment for question and answer model in designated domain and medium

    CN120822599A

  • Large language model structured multi-mode response method and device oriented to vertical field

    CN121168669A

  • Large language model structured multi-modal response method and device for vertical field

    CN121168669B