A Fine-tuning Method, Device and Medium for Large Models in Vertical Domains
By extracting the initial triplet from the knowledge graph and generating high-quality Q&A pairs, and fine-tuning the vertical field big model, the problem of insufficient accuracy of the vertical field big model is solved, and efficient Q&A and professional interaction in the designated field is achieved.
Patent Information
- Application Number
- CN202510549483.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing large language models cannot meet the professional needs of vertical fields, especially in areas such as geology, medical care and law, and cannot accurately understand and process complex professional terms and concepts, resulting in insufficient accuracy and interaction capabilities of large models in vertical fields.
By extracting initial triplets from the knowledge graph of the specified field, using pre-constructed Q&A prompt word engineering and designated big models to generate Q&A pairs, determine the quality scores of the Q&A pairs, and fine-tuning training of vertical field big models through high-quality Q&A pairs, ensure that the quality of the Q&A pair meets the professional requirements of the specified field.
It improves the accuracy and interaction ability of the vertical field big models in the professional field, meets the professional needs of specific fields, generates high-quality question-and-answer pairs, avoids the loss of information in the knowledge graph, and improves the adaptability and practicality of the model in the vertical field.
Smart Images

Figure CN120069098B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a fine-tuning method, device, and medium for large models in vertical domains. Background Art
[0002] With the rapid development of deep learning technology, large language models are widely used in all aspects of people's lives, such as intelligent driving, sentiment analysis, and creative content generation. However, current large language models are usually general models and cannot meet the needs of vertical domains. Vertical domains usually require in-depth professional background knowledge to accurately understand and process relevant information. For example, in vertical domains such as geology, medicine, and law, there are a large number of professional terms and complex concept systems, and simply relying on general large language models often fails to achieve the expected effects of users.
[0003] To meet the actual business needs of large models in different domains, based on fixed generation rules, triples in the knowledge graph of a vertical domain can be transformed into natural language question-answer pairs, and the large language model can be fine-tuned through the question-answer pairs to obtain a large model applied to the vertical domain. However, in the process of transforming triples into natural language question-answer pairs, the fixed generation rules will lose complex attribute and relationship information in the knowledge graph, resulting in incomplete semantics or insufficient accuracy of the generated question-answer pairs, thus making it still difficult for the accuracy of the large model in the vertical domain to reach the expected effects of users.
[0004] Therefore, how to improve the accuracy and interaction ability of large models in vertical domains and meet the business needs of vertical domains is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, one aspect of this application provides a fine-tuning method for large models in vertical domains, and the method includes:
[0006] Extract initial triples from the knowledge graph of a specified domain;
[0007] Based on a pre-constructed question-answering prompt engineering and the initial triples, call a specified large model to generate question-answer pairs about the initial triples;
[0008] Determine a quality score for characterizing the quality of the question-answer pairs; the larger the quality score, the higher the quality of the question-answer pairs;
[0009] Fine-tune and train the large model in the vertical domain through the target question-answer pairs with the quality score greater than a threshold.
[0010] Optionally, the determining a quality score for characterizing the quality of the question-answer pairs includes:
[0011] Obtain a pre-constructed inference prompt engineering;
[0012] Mask the initial triple to obtain a masked triple; the initial triple includes a subject field, a relation field, and a predicate field;
[0013] Based on the inference prompt engineering, use the specified large model to perform inference on the masked triple to obtain a completed triple;
[0014] Determine the quality score according to the initial triple and the completed triple.
[0015] Optionally, the masking of the initial triple to obtain a masked triple includes:
[0016] Mask any one or any two of the subject field, the relation field, and the predicate field to obtain the masked triple.
[0017] Optionally, the masked triple includes a first masked triple obtained by masking the relation field and a second masked triple obtained by masking the predicate field;
[0018] Correspondingly, the completed triple includes a first completed triple corresponding to the first masked triple and a second completed triple corresponding to the second masked triple.
[0019] Optionally, the determining of the quality score according to the initial triple and the completed triple includes:
[0020] Obtain a pre-constructed similarity scoring prompt engineering;
[0021] Based on the similarity scoring prompt engineering, call a target large language model to perform similarity scoring on the initial triple and the first completed triple, and the second completed triple respectively to obtain a first similarity score and a second similarity score;
[0022] Take the average of the first similarity score and the second similarity score as the quality score; where the higher the average, the higher the quality of the question-answer pair corresponding to the initial triple.
[0023] Optionally, calling a specified large model to generate a question-answer pair about the initial triple includes:
[0024] For each initial triple, generate a preset number of groups of initial question-answer pairs; the initial question-answer pairs include at least one of an open-ended question-answer pair and a closed-ended question-answer pair;
[0025] Merge the initial question-answer pairs to obtain the question-answer pair of the initial triple.
[0026] Optionally, extracting the initial triples from the knowledge graph in the specified domain includes:
[0027] Extracting entity nodes and relationship information between the entity nodes from the knowledge graph;
[0028] Performing data cleaning on the entity nodes and the relationship information;
[0029] Constructing the entity nodes and the relationship information after data cleaning into the initial triples in a preset format.
[0030] Another aspect of the present application provides a fine-tuning device for a vertical domain large model, and the device includes:
[0031] A triple extraction module, configured to extract initial triples from the knowledge graph in the specified domain;
[0032] A question-and-answer pair generation module, configured to call a specified large model to generate question-and-answer pairs about the initial triples based on a pre-constructed question-and-answer prompt engineering and the initial triples;
[0033] A quality determination module, configured to determine a quality score for characterizing the quality of the question-and-answer pairs; the larger the quality score, the higher the quality of the question-and-answer pairs;
[0034] A fine-tuning training module, configured to fine-tune and train the vertical domain large model through target question-and-answer pairs with a quality score greater than a threshold.
[0035] Another aspect of the present application provides a fine-tuning device for a vertical domain large model, including a memory and a processor. A computer program that can run on the processor is stored on the memory, and when the processor executes the program, the steps of the fine-tuning method for the vertical domain large model are implemented.
[0036] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the fine-tuning method for the vertical domain large model are implemented.
[0037] The beneficial effects produced by the fine-tuning method, device and medium for the vertical domain large model provided by the present application are as follows: Based on the initial triples in the knowledge graph of the specified domain and the question-and-answer prompts, the specified large model can quickly learn and reason about the knowledge in the initial triples, generate high-quality question-and-answer pairs, and avoid losing effective information in the triples. In addition, through the screening based on the quality score of the question-and-answer pairs, the vertical domain large model is fine-tuned and trained through high-quality target question-and-answer pairs, so as to improve the question-and-answer accuracy of the large model in the vertical domain, and the large model fine-tuning training can be completed according to the knowledge graphs of different domains, so as to meet the professional needs of specific domains. Description of the Drawings
[0038] Figure 1 It is a schematic flowchart of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the principle of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of the principle of a triple mask provided by an embodiment of the present application;
[0041] Figure 4 It is a schematic diagram of the principle of a fine-tuning method for a vertical domain large model provided by another embodiment of the present application;
[0042] Figure 5 It is a schematic structural diagram of a fine-tuning device for a vertical domain large model provided by an embodiment of the present application;
[0043] Figure 6 It is a schematic structural diagram of a fine-tuning device for a vertical domain large model provided by another embodiment of the present application.
[0044] The reference numerals are as follows: 60 is a memory, 61 is a processor, 62 is a display screen, 63 is an input / output interface, 64 is a communication interface, 65 is a power supply, 66 is a communication bus, 601 is a computer program, 602 is an operating system, and 603 is data. Detailed implementation manners
[0045] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms of "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0046] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0047] Figure 1 It is a schematic flowchart of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application, as Figure 1As shown, the method includes:
[0048] S10: Extract initial triples from the knowledge graph of the specified domain;
[0049] Figure 2 The figure is a schematic diagram of the principle of a fine-tuning method for a vertical domain large model provided by an embodiment of the present application. As Figure 2 shown, in a specific embodiment, obtain the knowledge graph of the specified domain. The specified domain may include, but is not limited to, geology, medicine, and law. The knowledge graph can be pre-built or an open-source knowledge graph obtained. That is, the present application does not limit the acquisition methods of the specified domain and the knowledge graph.
[0050] After obtaining the knowledge graph, in order to obtain the professional knowledge in the specified domain for subsequent fine-tuning of the large model, that is, to let the large model learn the professional knowledge in the specified domain so that the large model has the ability to answer questions in the specified domain knowledge. In an alternative embodiment, extract initial triples from the knowledge graph.
[0051] Specifically, first identify the entity nodes in the knowledge graph and the relationship information between the entity nodes, and combine the entity nodes and the relationship information to obtain the initial triples. It should be noted that in an alternative embodiment, in order to ensure the quality of the initial triples, the extracted initial triples can be filtered and screened to retain high-quality initial triples. For ease of understanding, the initial triples are illustrated below. For example, the initial triple is (apple, belongs to, fruit).
[0052] S11: Based on the pre-built question-and-answer prompt engineering and the initial triples, call the specified large model to generate question-and-answer pairs about the initial triples;
[0053] Furthermore, in order to enable the specified large model to quickly learn the professional knowledge in the initial triples and generate reliable question-and-answer pairs (i.e., question-and-answer pairs), in an alternative embodiment, a question-and-answer prompt (Prompt) engineering is pre-built. Under the guidance of the question-and-answer Prompt engineering, as Figure 2 shown, the specified large model learns and reasons about the initial triples to generate question-and-answer pairs about the initial triples.
[0054] For example, if the initial triple is (apple, belongs to, fruit), the question-and-answer pairs that the specified large model can generate include: Q (question): "What category does an apple belong to", A (answer): "An apple belongs to fruit".
[0055] It should be noted that the designated large models include but are not limited to ChatGPT, LLaMA2, LLaMA3, etc. This application does not limit the designated large models, and they can be selected according to actual business needs. For example, in a vertical field, a large model corresponding to the vertical field can be selected as the designated model, and the large model can be fine-tuned and trained through the fine-tuning method of this application, so as to improve the performance of the model in the vertical field.
[0056] It can be understood that the Q&A Prompt engineering is used to guide the designated large model to generate natural language Q&A pairs based on the provided initial triples. Therefore, the Q&A Prompt engineering is crucial for the quality of the Q&A pairs. In a specific embodiment, the Q&A Prompt engineering is concatenated with the initial triples to generate the input text, and the input text is input into the designated large model for inference to generate Q&A pairs. The Q&A Prompt engineering requires the designated large model to make full use of the provided initial triple information to generate Q&A pairs with true content and natural language, and output them in a specified format (for example, Json format). For the convenience of understanding, the Q&A Prompt engineering will be illustrated with examples below.
[0057] Q&A Prompt engineering:
[0058] <Please generate Q&A pairs according to the initial triples I provided and meet the following requirements:
[0059] 1. Generate Q&A pairs using the provided initial triples.
[0060] 2. Add the content you know to make the Q&A pairs as rich as possible. At the same time, please ensure that you utilize all the information provided in the initial triples.
[0061] 3. Ensure that the generated content is true.
[0062] 4. Conform to the laws of natural language in the Q&A pairs.
[0063] 5. Please output in Json format:
[0064] {
[0065] “Question”: “……”;
[0066] “Answer”: “……”;
[0067] }
[0068] The following are the initial triples I provided for you: (Apple, belongs to, fruit). >
[0069] S12: Determine the quality score used to characterize the quality of the Q&A pair; the larger the quality score, the higher the quality of the Q&A pair;
[0070] S13: Fine-tune and train the large model for the vertical domain using target Q&A pairs with a mass fraction greater than the threshold.
[0071] Furthermore, select high-quality target Q&A pairs from all the Q&A pairs generated by the specified large model, and use them to fine-tune and train the large model for the vertical domain. That is, use the target Q&A pairs as the training data set to fine-tune and train the large model for the vertical domain, so as to obtain a large model with high accuracy in the specified domain.
[0072] In an alternative embodiment, the calculation of the mass fraction can be based on multiple dimensions. For example, the accuracy, completeness, logic of the answer, and whether it conforms to the professional knowledge within the domain. In an alternative embodiment, if the generated answer completely conforms to the facts within the specified domain and is clearly and completely expressed, then the mass fraction is relatively high. On the contrary, if the answer contains errors or is vaguely expressed, the mass fraction will be relatively low. The larger the mass fraction, the higher the quality of the generated Q&A pair, and the more it can meet the requirements for accuracy and professionalism in practical applications.
[0073] Therefore, as Figure 2 shown, fine-tune and train the large model for the vertical domain by screening out target Q&A pairs with a mass fraction greater than the threshold. Among them, the purpose of fine-tuning and training is to make the large model better adapt to the knowledge and language style of the vertical domain and improve its Q&A performance in this domain. After fine-tuning, the large model can more accurately understand and generate Q&A content related to the vertical domain, thereby improving the quality and practicality of the entire Q&A system.
[0074] It should be noted that the large model for the vertical domain can be the specified large model in step S11 or other models, and this application does not make any limitations in this regard. In addition, the vertical domain includes but is not limited to the geological domain, medical domain, legal domain, sports domain, etc.
[0075] In an alternative embodiment, in order to further improve the quality of the target Q&A pairs, after selecting high-quality Q&A pairs according to the instruction score, the target Q&A pairs can be further verified and corrected manually. For example, correct errors and optimize expressions to ensure the quality of the target Q&A pairs and further improve the accuracy of the large model for the vertical domain in the specified domain.
[0076] In another alternative embodiment, a high-quality test set data for the specified domain can be pre-constructed, which can be a test set obtained through manual annotation. In a specific embodiment, after training the large model for the vertical domain with the above automatically generated target Q&A pairs, the factuality, fluency, interpretability and other indicators of the model can be evaluated and fine-tuned through the test set data, so as to further improve the performance of the large model for the vertical domain.
[0077] Thus, the fine-tuning method for the vertical domain large model provided by the embodiments of the present application, based on the initial triples of the specified domain knowledge graph and the Q&A prompts, enables the large model to quickly learn and reason about the knowledge in the initial triples, generate high-quality Q&A pairs, and avoid losing valid information in the triples. In addition, through the screening based on the quality scores of the Q&A pairs, the vertical domain large model is fine-tuned and trained with high-quality target Q&A pairs, thereby improving the Q&A accuracy of the large model in the vertical domain. The large model fine-tuning training can be completed according to the knowledge graphs of different domains, thus meeting the professional needs of specific domains.
[0078] In an alternative embodiment, determining the quality score for characterizing the quality of the Q&A pair includes:
[0079] Obtain a pre-constructed inference prompt engineering;
[0080] Mask the initial triples to obtain masked triples; the initial triples include a subject field, a relation field, and a predicate field;
[0081] Based on the inference prompt engineering, use the specified large model to reason about the masked triples to obtain completed triples;
[0082] Determine the quality score according to the initial triples and the completed triples.
[0083] Figure 3 FIG. is a schematic diagram of the principle of triple masking provided by the embodiments of the present application. As Figure 3 shown, in a specific embodiment, the initial triples extracted from the knowledge graph are masked to obtain masked triples. Among them, masking the initial triples means hiding or replacing some field information in the initial triples by means of masking.
[0084] The initial triples include a subject field, a relation field, and a predicate field. Therefore, masking the initial triples means hiding any one or more of the subject field, the relation field, and the predicate field. It should be noted that, in order to enable the subsequent specified large model to reason about the masked triples and thus complete the masked fields, it can be understood that only one or two of the subject field, the relation field, and the predicate field can be masked to ensure that the specified large model can reason based on at least one piece of information in the masked triples. For ease of understanding, the following will take masking one of the fields as an example for illustration.
[0085] For example, the initial triple is (water, molecular formula, H2O), where water is the subject field, molecular formula is the relationship field, and H2O is the predicate field. In an alternative embodiment, the predicate field can be masked to obtain a masked triple (water, molecular formula, *), where "*" indicates that the predicate field is hidden.
[0086] In an alternative embodiment, to quickly reason about and complete the masked triple, that is, to restore the masked (i.e., hidden) field in the masked triple, an inference Prompt project can be pre-constructed. The inference Prompt project is used to guide a specified large model to complete the masked part of the masked triple. Specifically, as Figure 3 described, based on the inference Prompt project, the masked triple is used as the input to the specified large model, and the specified large model reasons about the masked triple to obtain a completed triple with the masked field completed.
[0087] It can be understood that in the above embodiment, the specified large model learns question-answer pairs by reasoning about the initial triple. Therefore, the specified large model has learned the knowledge in the initial triple at this time. To detect whether the generated question-answer pairs are high-quality data, the specified large model can be used to complete the masked triple. If the restored completed triple is exactly the same as the initial triple, it indicates that the specified large model has learned the knowledge in the initial triple and the corresponding generated question-answer pairs have high quality. On the contrary, if the semantic difference between the completed triple and the initial triple is large, it indicates that the quality of the generated question-answer pairs is low. Therefore, in an alternative embodiment, as Figure 3 shown, the quality score of the corresponding question-answer pairs can be calculated based on the initial triple and the completed triple. For ease of understanding, the inference Prompt project will be illustrated by examples below.
[0088] For example, the inference Prompt project is:
[0089] <Please fill in the missing part (* part) of the masked triple according to the question-answer pairs I provide:
[0090] "Question": "What is the molecular formula of water?"
[0091] "Answer": "The molecular formula of water is H2O".
[0092] The masked triple is: (water, molecular formula, *)>
[0093] Thus, the fine-tuning method for the vertical domain large model provided by the embodiments of the present application realizes the rapid completion of masked triples based on inference Prompt engineering, combines the completed triples with the initial triples, and evaluates whether the quality of the obtained question-answer pairs meets the expectations, so as to efficiently extract high-quality question-answer pairs and provide a high-quality training data set basis for the subsequent training of the vertical domain large model.
[0094] Based on the above-mentioned initial triples in the embodiments including a subject field, a relation field, and a predicate field, as an optional embodiment, the initial triples are masked to obtain masked triples, including:
[0095] Mask any one or any two of the subject field, the relation field, and the predicate field to obtain masked triples.
[0096] It can be understood that the obtained masked triples will be inferred later to restore the masked fields. Therefore, when masking the initial triples, at least one field should be retained for subsequent inference. That is, when masking, any one or any two of the subject field, the relation field, and the predicate field are masked.
[0097] However, it can also be understood that in a specific embodiment, the subject field is very important for the inference of the masked field. If the subject field is masked, the subsequent inference results may be quite different from the corresponding initial triples, which is not conducive to the evaluation of the quality of the question-answer pairs.
[0098] At the same time, if multiple of the subject field, the relation field, and the predicate field are masked, that is, two fields are masked and only one field remains, it will result in a large number of completed triples, which is also not conducive to the evaluation of the quality of the question-answer pairs and will greatly affect the quality evaluation efficiency.
[0099] Therefore, in an optional embodiment, the initial triples are masked to obtain masked triples, including: masking the relation field to obtain the first masked triples, and masking the predicate field to obtain the second masked triples. That is, only one field in the initial triples is masked. Thus, the masked triples include the first masked triples and the second masked triples. The corresponding completed triples include the first completed triples corresponding to the first masked triples and the second completed triples corresponding to the second masked triples.
[0100] It should be noted that in specific embodiments, to ensure the accuracy of quality score evaluation, the relationship field and the predicate field can be masked respectively, and the quality of the corresponding question-answer pairs can be evaluated based on the generated first masked triple and second masked triple. Of course, to ensure the efficiency of quality evaluation, in an alternative embodiment, only one of the relationship field and the predicate field can be masked, and the present application does not make any limitations in this regard.
[0101] Based on the above embodiments, as an alternative embodiment, determining the quality score according to the initial triple and the completed triple includes:
[0102] Obtain a pre-constructed similarity scoring prompt engineering;
[0103] Based on the similarity scoring prompt engineering, call the target large language model to perform similarity scoring on the initial triple and the first completed triple and the second completed triple respectively, to obtain a first similarity score and a second similarity score;
[0104] Take the mean of the first similarity score and the second similarity score as the quality score; where, the higher the mean, the higher the quality of the question-answer pair corresponding to the initial triple.
[0105] In an alternative embodiment, to quickly determine the quality score of the question-answer pair, a similarity scoring Prompt engineering regarding the similarity between the initial triple and the completed triple can be pre-constructed. It can be understood that the similarity scoring Prompt engineering is used to guide the target large language model to score the similarity between the initial triple and the completed triple.
[0106] Specifically, based on the similarity scoring Prompt engineering, call the target large language model to score the initial triple and the first completed triple to obtain a first similarity score, and perform similarity scoring on the initial triple and the second completed triple to obtain a second similarity score.
[0107] Furthermore, since there are multiple obtained similarity scores, when evaluating the quality of the corresponding question-answer pair, the mean of the first similarity and the second similarity can be used as the quality score of the corresponding question-answer pair. It should be noted that, the higher the mean, the higher the semantic similarity between the initial triple and the completed triple, and the higher the quality of the corresponding generated question-answer pair. For ease of understanding, the following takes the completed triple as one, that is, only masking the relationship field or the predicate field as an example to illustrate the similarity scoring Prompt engineering.
[0108] The similarity scoring Prompt engineering is:
[0109] I will provide two sets of triples. The triples in these two sets correspond to each other in order, but their objects or subjects may be different. Please score their similarity according to the following rules.
[0110] 1. Scoring rules
[0111] If two triples are exactly the same, assign a score of 100 points.
[0112] If two triples are not exactly the same, score them according to the closeness of their meanings and expressions. Assign higher scores to those with closely related meanings and lower scores to those with greater differences. The score range is from 1 to 99 points.
[0113] If their meanings are completely different, assign 0 points.
[0114] 2. Ensure that the scores in the output are the same as the scores of the triples in the set I provided and in the same order, and then directly output the similarity scores.
[0115] The output format is: [score 1, score 2, ……]
[0116] The triples include: (apple, belongs to, fruit), (apple, is, fruit)
[0117] Thus, based on this similarity scoring Prompt engineering, the similarity scores between the initial triples and the completed triples can be quickly output. Of course, in an alternative embodiment, the semantic similarity algorithm can also be used to calculate the similarity between the initial triples and the completed triples, and based on the calculated similarity value, determine the quality score. This application does not limit the quality scores corresponding to different similarity values, but when assigning the corresponding quality score values, it is necessary to conform to the rule that the greater the similarity value, the greater the quality score, indicating the higher the quality of the question-and-answer pair.
[0118] In an alternative embodiment, to improve the generation quality of the question-and-answer pair, a specified large model is called to generate question-and-answer pairs for the initial triples, including:
[0119] For each initial triple, generate a preset number of groups of initial question-and-answer pairs; the initial question-and-answer pairs include at least one of open-ended question-and-answer pairs and closed-ended question-and-answer pairs;
[0120] Merge the initial question-and-answer pairs to obtain the question-and-answer pairs of the initial triples.
[0121] Figure 4 This is a schematic diagram of the principle of a fine-tuning method for a vertical domain large model provided by another embodiment of this application. In a specific embodiment, multiple rounds of question-and-answer pair generation can be performed for each initial triple to ensure the diversity of the generated question-and-answer pairs. For example, asFigure 4 As shown, the initial triple A is concatenated based on the Q&A Prompt engineering to generate input text 1 and input text 2, and the initial Q&A pairs 1 and 2 are obtained through inference by specifying a large model. Further, after merging the initial Q&A pairs 1 and 2, the Q&A pair corresponding to the initial triple is obtained.
[0122] It should be noted that in a specific embodiment, the same initial triple may generate multiple different Q&A pairs. When merging the initial Q&A pairs, the semantic similarity between two Q&A pairs can be determined first, and further, the Q&A pairs with the semantic similarity reaching the similarity threshold are merged, that is, the Q&A pairs with high similarity are merged.
[0123] In an optional embodiment, the initial triples are extracted from the knowledge graph of a specified domain, including:
[0124] Extract entity nodes and the relationship information between entity nodes from the knowledge graph;
[0125] Perform data cleaning on the entity nodes and relationship information;
[0126] Construct the entity nodes and relationship information after data cleaning into initial triples in a preset format.
[0127] In a specific embodiment, the knowledge graph of a specified domain can be obtained from various open-source data sources (such as texts, databases, web pages, etc.), and entity nodes and the relationship information between entity nodes are extracted from the knowledge graph. Among them, the knowledge graph G=(V, R, E), where V is the set of entity nodes, R is the set of relationship information, and E is the set of triples. Specifically, according to the actual business requirements, the entity nodes that meet the specified domain are screened out from the knowledge graph. For example, in the scientific field, key entity nodes such as "molecule", "reaction", and "experiment" are selected. Further, among the entity nodes, the relationship information highly relevant to the actual business requirements is screened out. For example, in the scientific field, core relationship information such as "reaction conditions", "experimental methods", and "molecular structure" is screened out.
[0128] Furthermore, in order to improve the quality of subsequent initial triples, data cleaning is performed on the entity nodes and relationship information. The data cleaning includes but is not limited to removing noise data. The entity nodes and relationship information after cleaning are combined into initial triples in a preset format, thereby completing the construction of the initial triples. Among them, the preset format can be the standard format in the triples, that is, the standard format of (subject, relationship, predicate). Of course, it can also be constructed based on other orders, formats, etc., and this application does not make any limitations in this regard.
[0129] To make the technical solution of this application clearer to those skilled in the art, the following will take the main field Juan Antonio Pizzi as an example for illustration.
[0130] The initial triples include: (Juan Antonio Pizzi, place of birth, Santa Fe), (Santa Fe, country, Argentina), (Juan Antonio Pizzi, sports country, Spain).
[0131] The masked triples obtained by masking the initial triples include: (Juan Antonio Pizzi, place of birth, *), (Santa Fe, country, *), (Juan Antonio Pizzi, sports country, *).
[0132] Based on the Q&A Prompt engineering, the large model is specified to generate Q&A pairs for the initial triples as: "Question": "Where was Juan Antonio Pizzi born and in which countries during his sports career", "Answer": "Juan Antonio Pizzi was born in Santa Fe, Argentina, and during his sports career, he has been in Spain."
[0133] Furthermore, based on the reasoning Prompt engineering, reasoning is performed on the masked triples, and the completed triples obtained include: (Juan Antonio Pizzi, place of birth, Province of Santa Fe), (Argentina, Republic, Santa Fe), (Juan Antonio Pizzi, Spanish national team, sports national team).
[0134] Furthermore, based on the similarity scoring Prompt engineering, the similarity score between the initial triples and the completed triples is calculated, and then the quality score representing the quality of the Q&A pairs is determined. In a specific embodiment, data with a quality score greater than a threshold, for example, greater than 80 points, can be used as the training data set to fine-tune the vertical domain large model, thereby improving the Q&A accuracy of the large model in the specified domain.
[0135] In addition, to make the fine-tuning results of the vertical domain large model fine-tuning method provided by this application clearer to those skilled in the art, in a specific embodiment, taking the knowledge graph AristoV4 graph as an example, the fine-tuning results by the method provided by this application are introduced. Among them, AristoV4 is a knowledge graph in the scientific field, containing high-precision triples for the basic science field. AristoV4 contains more than 40,000 entities and 280,000 triples of knowledge, which can be used to answer scientific-related questions. Table 1 is a schematic table of the number and score of generated Q&A pairs provided by the embodiment of this application. Through the fine-tuning method of the large model provided by this application, 5 times the size of Q&A pairs are generated on AristoV4, and the specific quantity and quality are shown in Table 1.
[0136] Table 1 is a schematic table of the number and score of generated Q&A pairs
[0137]
[0138] Table 2 is a result schematic table of different target generation pairs and the accuracy of the large model provided by the embodiments of the present application. Further, through the target Q&A pairs of different qualities in the above table, the vertical domain large model is fine-tuned to obtain the comparison results shown in Table 2. Among them, the vertical domain large models are illustrated by taking Llama2 and Llama3 as examples.
[0139] Table 2 is a result schematic table of different target generation pairs and the accuracy of the large model
[0140]
[0141] Table 3 is a result schematic table of different training data and the accuracy of the model provided by the embodiments of the present application. Through the Q&A pairs generated above, the large models Llama2 and Llama3 are fine-tuned. At the same time, as shown in Table 3, there are also non-fine-tuning methods, triple fine-tuning methods, retrieval answer fine-tuning methods, and long text fine-tuning methods. Among them, the non-fine-tuning method refers to the method of not fine-tuning the model, the triple fine-tuning method refers to the method of directly fine-tuning the model through triples, the retrieval answer fine-tuning refers to the method of fine-tuning the model through retrieval results and retrieval questions, and the long text fine-tuning refers to the method of converting triples into long texts to fine-tune the model.
[0142] Table 3 is a result schematic table of different training data and the accuracy of the model
[0143]
[0144] As can be seen from Table 1, as the data scale (i.e., the number of initial triples) increases, the number of Q&A pairs with different quality scores of high quality increases significantly, but the proportion of high-quality data remains relatively stable (i.e., the proportion of 0 to 60 points is about 73%, the proportion of 60 to 80 points and above is about 48%, and the proportion of 80 to 100 points is about 27%). It can be seen from this that the technical solution provided by the present application can maintain a high quality of Q&A pairs under different data scales. At 5 times the scale, there are more than 130,000 high-quality Q&A pairs above 60 points, providing sufficient high-quality data support for model fine-tuning. That is, the method provided by the present application can generate high-quality Q&A pairs under different data scales, and the proportion of high-quality data remains stable, providing reliable data support for model fine-tuning.
[0145] As can be seen from Table 2, even when fine-tuning and training with Q&A pairs ranging from 0 to 60 points, the model performance has been significantly improved (among them, the large model Llama2 has increased from 7.56 to 7.62, and the large model Llama3 has increased from 8.42 to 8.67). From this, it can be known that the method provided by this application can effectively screen out high-quality data for model fine-tuning. According to Table 2, the model performance fine-tuned with Q&A pairs ranging from 80 to 100 points is the highest. It can be seen that the method provided by this application can efficiently screen out high-quality Q&A pairs and achieve high-precision fine-tuning training of the model.
[0146] As can be seen from Table 3, on the large models Llama2 and Llama3, the fine-tuning results of the Q&A pairs are higher than those of other fine-tuning methods. That is, the method provided by this application has significant advantages in improving model performance. Especially on the large model Llama3, the Q&A pair fine-tuning method provided by this application is significantly higher than other fine-tuning methods.
[0147] Therefore, the fine-tuning method for the vertical domain large model provided by this application realizes the efficient conversion of knowledge graph triples to natural language Q&A pairs through multi-level Prompt engineering, reducing information loss. And based on the triple completion task and similarity scoring mechanism, it quantifies the knowledge integrity and semantic accuracy of the Q&A pairs, ensures data quality, generates Q&A pairs through an automated process and conducts quality evaluation, significantly reducing labor costs and time costs. At the same time, by generating Q&A pairs for the specified domain, it improves the adaptability and practicality of the large model in fields with high factual requirements such as medical, geological, and legal.
[0148] In the above embodiments, the fine-tuning method for the vertical domain large model is described in detail. This application also provides an embodiment corresponding to a fine-tuning device for the vertical domain large model.
[0149] Figure 5 It is a schematic structural diagram of a fine-tuning device for a vertical domain large model provided by an embodiment of this application, as Figure 5 shown. The device includes:
[0150] A triple extraction module 50, configured to extract initial triples from the knowledge graph of the specified domain;
[0151] A Q&A pair generation module 51, configured to call a specified large model to generate Q&A pairs about the initial triples based on the pre-constructed Q&A prompt word engineering and the initial triples;
[0152] A quality determination module 52, configured to determine a quality score for characterizing the quality of the Q&A pairs; the larger the quality score, the higher the quality of the Q&A pairs;
[0153] The fine-tuning training module 53 is used to fine-tune and train the vertical domain large model with target question-answer pairs whose quality scores are greater than the threshold.
[0154] In addition, the fine-tuning device for the vertical domain large model provided by the embodiments of the present application further includes:
[0155] The prompt engineering acquisition module is used to acquire the pre-constructed inference prompt engineering;
[0156] The masking module is used to mask the initial triple to obtain a masked triple; the initial triple includes a subject field, a relationship field, and a predicate field;
[0157] The masked inference module is used to infer the masked triple through a specified large model based on the inference prompt engineering to obtain a completed triple;
[0158] The quality score determination module is used to determine the quality score according to the initial triple and the completed triple.
[0159] The first masking sub-module is used to mask any one or any two of the subject field, the relationship field, and the predicate field to obtain a masked triple.
[0160] The prompt engineering acquisition module is further used to acquire the pre-constructed similarity scoring prompt engineering;
[0161] The similarity scoring model is used to perform similarity scoring on the initial triple and the first completed triple and the second completed triple respectively by calling the target large language model based on the similarity scoring prompt engineering to obtain a first similarity score and a second similarity score;
[0162] The mean determination module is used to take the mean of the first similarity score and the second similarity score as the quality score; where the higher the mean, the higher the quality of the question-answer pair corresponding to the initial triple.
[0163] The initial question-answer pair generation module is used to generate a preset number of groups of initial question-answer pairs for each initial triple; the initial question-answer pair includes at least one of an open-ended question-answer pair and a closed-ended question-answer pair;
[0164] The merging module is used to merge the initial question-answer pairs to obtain the question-answer pairs of the initial triple.
[0165] The extraction module is used to extract entity nodes and the relationship information between entity nodes from the knowledge graph;
[0166] The data cleaning module is used to clean the entity nodes and the relationship information;
[0167] A triple construction module for constructing the entity nodes and relationship information after data cleaning into initial triples in a preset format.
[0168] Figure 6 The following is a schematic structural diagram of a fine-tuning device for a vertical domain large model provided by another embodiment of the present application. As Figure 6 shown, the fine-tuning device for the vertical domain large model includes: a memory 60 for storing computer programs;
[0169] a processor 61 for implementing the steps of the fine-tuning method for the vertical domain large model as mentioned in the above embodiment when executing the computer program.
[0170] The fine-tuning device for the vertical domain large model provided in this embodiment may include, but is not limited to, a laptop computer or a desktop computer, etc.
[0171] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0172] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices and flash storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601. After being loaded and executed by the processor 61, the computer program can implement the relevant steps of the fine-tuning method of the vertical domain large model disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include, but is not limited to, relevant data involved in the fine-tuning method of the vertical domain large model.
[0173] In some embodiments, the fine-tuning device of the vertical domain large model may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.
[0174] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the fine-tuning device of the vertical domain large model, and may include more or fewer components than shown in the figure.
[0175] The fine-tuning device of the vertical domain large model provided by the embodiments of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the fine-tuning method of the vertical domain large model in the above embodiments.
[0176] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
Claims
1. A fine-tuning method for large models in vertical domains, characterized in that, The method includes: extracting initial triples from the knowledge graph of a specified domain; invoking a specified large model to generate question-and-answer pairs about the initial triples based on a pre-constructed question-and-answer prompt engineering and the initial triples; determining a quality score for characterizing the quality of the question-and-answer pairs; the larger the quality score, the higher the quality of the question-and-answer pairs; fine-tuning and training a vertical domain large model with the target question-and-answer pairs whose quality scores are greater than a threshold; The determining a quality score for characterizing the quality of the question-and-answer pairs includes: obtaining a pre-constructed inference prompt engineering; masking the initial triples to obtain masked triples; the initial triples include a subject field, a relation field, and a predicate field; performing inference on the masked triples through the specified large model based on the inference prompt engineering to obtain completed triples; determining the quality score according to the initial triples and the completed triples; The masked triples include a first masked triple obtained by masking the relation field and a second masked triple obtained by masking the predicate field; Correspondingly, the completed triples include a first completed triple corresponding to the first masked triple and a second completed triple corresponding to the second masked triple; The determining the quality score according to the initial triples and the completed triples includes: obtaining a pre-constructed similarity scoring prompt engineering; invoking a target large language model based on the similarity scoring prompt engineering to perform similarity scoring on the initial triples and the first completed triple and the second completed triple respectively to obtain a first similarity score and a second similarity score; taking the mean of the first similarity score and the second similarity score as the quality score; wherein, the higher the mean, the higher the quality of the question-and-answer pairs corresponding to the initial triples.
2. The fine-tuning method of the large model in the vertical domain according to claim 1, wherein Invoking a specified large model to generate question-and-answer pairs about the initial triples includes: generating a preset number of groups of initial question-and-answer pairs for each of the initial triples; the initial question-and-answer pairs include at least one of open-ended question-and-answer pairs and closed-ended question-and-answer pairs; merging the initial question-and-answer pairs to obtain the question-and-answer pairs of the initial triples.
3. The fine-tuning method of the large model in the vertical domain according to claim 1, wherein, The extracting initial triples from the knowledge graph of a specified domain includes: extracting entity nodes and the relationship information between the entity nodes from the knowledge graph; performing data cleaning on the entity nodes and the relationship information; constructing the entity nodes and the relationship information after data cleaning into the initial triples in a preset format.
4. A fine-tuning device for a large model in a vertical domain, characterized in that, The device includes: a triple extraction module for extracting initial triples from the knowledge graph of a specified domain; a question-and-answer pair generation module for invoking a specified large model to generate question-and-answer pairs about the initial triples based on a pre-constructed question-and-answer prompt engineering and the initial triples; a quality determination module for determining a quality score for characterizing the quality of the question-and-answer pairs; the larger the quality score, the higher the quality of the question-and-answer pairs; A fine-tuning training module for fine-tuning and training a vertical domain large model with the target Q&A pairs whose quality scores are greater than a threshold; A prompt engineering acquisition module for acquiring a pre-constructed inference prompt engineering; A masking module for masking the initial triple to obtain a masked triple; the initial triple includes a subject field, a relationship field, and a predicate field; the masked triple includes a first masked triple obtained by masking the relationship field and a second masked triple obtained by masking the predicate field; A masked inference module for inferring the masked triple through the specified large model based on the inference prompt engineering to obtain a completed triple; correspondingly, the completed triple includes a first completed triple corresponding to the first masked triple and a second completed triple corresponding to the second masked triple; A quality score determination module for determining the quality score according to the initial triple and the completed triple; The prompt engineering acquisition module is further configured to acquire a pre-constructed similarity scoring prompt engineering; A similarity scoring model for performing similarity scoring on the initial triple and the first completed triple and the second completed triple respectively by invoking a target large language model based on the similarity scoring prompt engineering to obtain a first similarity score and a second similarity score; An average value determination module for using the average value of the first similarity score and the second similarity score as the quality score; wherein, the higher the average value is, the higher the quality of the Q&A pair corresponding to the initial triple is.
5. A fine-tuning device for a large model in a vertical domain, comprising a memory and a processor, wherein a computer program that can run on the processor is stored on the memory, and is characterized in that, When the processor executes the program, it implements the steps of the fine-tuning method for the vertical domain large model according to any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the fine-tuning method for the vertical domain large model according to any one of claims 1 to 3.
Citation Information
Patent Citations
Model fine tuning method and device, electronic equipment and storage medium
CN119337832A