Large model fine tuning method and device in target field

By fine-tuning the target domain on the big model, using the input text and expected responses in the training dataset, we learn the empirical thinking paths in the vertical field, and form a thinking tree, solving the problem of the big model performing poorly in the vertical field, achieving higher quality answers.

CN120029673APending Publication Date: 2025-05-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510389327.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The performance of existing large models in vertical fields is not ideal enough, and they cannot effectively utilize the implicit background knowledge in vertical fields, resulting in low quality of answers.

Method used

The big model fine-tuning method under the target field is adopted. By obtaining a training data set composed of multiple samples, including input text and expected responses, the training data set is used to fine-tune the instruction of the big model, learning the empirical thinking paths under the target field, forming a thinking tree, and improving the model's answering ability in the vertical field.

Benefits of technology

Through fine-tuning, the big model can better understand and perform vertical domain tasks, improve the accuracy and relevance of the answers, and further improve the performance of the big model in vertical domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029673A_ABST
    Figure CN120029673A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a large model fine tuning method and device in a target field. The method comprises the following steps: acquiring a training data set formed by a plurality of samples; any sample comprises an input text and an expected response, and the input text comprises a request task of the target domain and an instruction for executing the request task; the expected response comprises a label thinking path and label answers obtained based on the label thinking path and the retrieved documents; the label thinking path comprises concept classification of the request task and associated information determined based on the concept classification; and performing instruction fine tuning on the large model by using the training data set to obtain the large model in the target field. And the performance of the large model in the vertical field can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the computer field, and in particular, to a large model fine-tuning method and apparatus in a target field. Background Art

[0002] Vertical fields, also known as vertical markets or vertical industries, refer to the focus on a specific industry or professional field. Different from general fields, vertical fields are more focused on meeting the unique needs of a specific industry and providing more specialized and customized solutions.

[0003] Usually, you can pre-train a large model in a general domain first, and then fine-tune the large model. Fine-tuning is the process of further training a pre-trained model for a specific task or domain, with the goal of making the pre-trained model better adaptable to specific application scenarios. Among them, the user's private data may be used, and it is necessary to protect the private data from being leaked. After fine-tuning in the general domain, the large model can obtain good results in the vertical domain, but if you want to further improve the performance of the large model in the vertical domain, you also need some implicit background in the vertical domain, because some knowledge in the vertical domain cannot be recalled.

[0004] Therefore, it is necessary to provide a method for fine-tuning the large model in the target field, so that the large model can provide better answers and further improve the performance of the large model in the vertical field. Summary of the invention

[0005] One or more embodiments of this specification describe a method and device for fine-tuning a large model in a target domain, which can further improve the performance of the large model in a vertical domain.

[0006] In a first aspect, a large model fine-tuning method in a target domain is provided, the method comprising:

[0007] A training data set consisting of multiple samples is obtained; any sample includes an input text and an expected response, wherein the input text includes a request task in the target domain and an instruction for executing the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification;

[0008] The training data set is used to fine-tune the instructions of the large model to obtain a large model in the target domain.

[0009] In a possible implementation manner, the input text also includes a number of retrieved documents.

[0010] In a possible implementation, the tag thinking path includes multiple levels of associated information, and the multiple levels of associated information form a thinking tree.

[0011] Furthermore, the plurality of associated information is determined in the following manner:

[0012] Determining first associated information based on the concept classification and the requested task;

[0013] According to the first associated information, the first precautions are determined as the second associated information, and the second precautions are determined as the third associated information.

[0014] In a possible implementation, the label answer is determined in the following manner:

[0015] Based on the association information, searching for target document information from the retrieved documents;

[0016] The target document information is input into a first neural network model for inference to obtain the label answer.

[0017] In a possible implementation, obtaining a training data set consisting of a plurality of samples includes:

[0018] Acquire a first number of manually annotated samples to form a first data subset;

[0019] Taking the samples in the first data subset as examples, a second number of samples are generated using a second neural network model to form a second data subset; the first data subset and the second data subset constitute the training data set.

[0020] Furthermore, the second neural network model has more parameters than the large model.

[0021] In a possible implementation manner, the input text includes a first type of special symbols, which are used to instruct the large model to think or answer.

[0022] In a possible implementation manner, the expected response includes a second type of special symbols, which are used to mark a label thinking path or a label answer.

[0023] In a possible implementation, the step of fine-tuning the instructions of the large model using the training data set includes:

[0024] Inputting the input text into the large model to obtain an output response corresponding to the input text;

[0025] The model parameters of the large model are adjusted with the minimization of the difference between the output response and the expected response as the training objective.

[0026] In a second aspect, a large model fine-tuning device in a target domain is provided, the device comprising:

[0027] An acquisition unit is used to acquire a training data set consisting of multiple samples; any sample includes an input text and an expected response, the input text includes a request task in the target field and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification;

[0028] A fine-tuning unit is used to fine-tune the instructions of the large model using the training data set obtained by the acquisition unit to obtain the large model in the target domain.

[0029] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0030] According to a fourth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0031] Through the method and device provided in the embodiment of this specification, a training data set consisting of multiple samples is first obtained; any sample includes an input text and an expected response, the input text includes a request task in the target field and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification; then the instruction fine-tuning of the large model is performed using the training data set to obtain a large model in the target field. As can be seen from the above, the embodiment of this specification adopts the Retrieval-Augmented Generation (RAG) technology and instruction fine-tuning, and uses human prior knowledge as a label thinking path, so that the large model learns to adopt different thinking paths under different concept classifications, thereby providing better answers, which can further improve the performance of the large model in vertical fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0033] Figure 1 A schematic diagram of an implementation scenario of an embodiment disclosed in this specification;

[0034] Figure 2 A flow chart of a large model fine-tuning method in a target domain according to an embodiment is shown;

[0035] Figure 3 A schematic diagram of a thinking path in a vertical field according to an embodiment is shown;

[0036] Figure 4 A schematic block diagram of a large model fine-tuning apparatus in a target domain according to an embodiment is shown. DETAILED DESCRIPTION

[0037] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0038] Figure 1 This is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification. This implementation scenario involves fine-tuning a large model in a target field. It can be understood that the target field is a vertical field. Before fine-tuning a large model, it is usually necessary to pre-train the large model. Figure 1 , pretraining and fine-tuning are two key steps in the development of large models. Pretraining refers to training a model on a large-scale dataset so that it can learn common features and patterns. These features usually have strong generalization capabilities and can serve as the basis for subsequent tasks. Fine-tuning is to further optimize the model performance by conducting a small amount of training on a dataset for a specific task based on the pretrained model. The goal of the fine-tuning stage is to make the model better adapt to the needs of a specific task. The model parameters are adjusted through pretraining, so that the large model is adjusted from the initial parameters to the pretrained parameters, and the model parameters are further adjusted through fine-tuning, so that the large model is adjusted from the pretrained parameters to the fine-tuned parameters.

[0039] The embodiment of this specification adopts the instruction fine-tuning method to achieve fine-tuning of the large model to obtain the large model in the target domain.

[0040] Among them, instruction fine-tuning improves the performance of the pre-trained large model through specific instructions. This method enables the large model to better understand and perform various tasks, not just generate coherent text. The main goal of instruction fine-tuning is to enable the large model to learn to follow the specific instructions given by humans, so as to show better performance on a variety of downstream tasks.

[0041] In addition, it should be noted that the big model can include but is not limited to any of the following: natural language processing (NLP) big model: focusing on processing and understanding natural language text data; multimodal big model: able to simultaneously process data in multiple modalities, such as text, images, audio, etc.; scientific computing big model: focusing on solving computational problems in the scientific field.

[0042] Figure 2 A flowchart of a large model fine-tuning method in a target domain according to an embodiment is shown. The method can be based on Figure 1 The implementation scenario shown in Figure 1 is as follows. Figure 2 As shown, the method for fine-tuning a large model in the target domain in this embodiment includes the following steps: Step 21, obtaining a training data set consisting of multiple samples; any sample includes an input text and an expected response, and the input text includes a request task in the target domain and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a conceptual classification of the request task and associated information determined based on the conceptual classification; Step 22, using the training data set to fine-tune the large model to obtain a large model in the target domain. The specific execution method of each of the above steps is described below.

[0043] First, in step 21, a training data set consisting of multiple samples is obtained; any sample includes an input text and an expected response, the input text includes a request task in the target field and an instruction to perform the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a conceptual classification of the request task and associated information determined based on the conceptual classification. It can be understood that the above-mentioned request task can be a question queried by a user, and the instruction is usually expressed in natural language to describe the task that the model needs to complete.

[0044] In the embodiments of this specification, the instruction may be a specific task description or a more abstract prompt. The instruction may also include contextual information or knowledge in a specific field. For example, in the medical field, the instruction may involve medical terms or specific medical scenarios.

[0045] In one example, the input text also includes a number of retrieved documents.

[0046] In this example, retrieval-augmented generation (RAG) technology is used, which is a technology that combines information retrieval and generation models and is often used to enhance the performance of natural language processing tasks. The core concept of RAG is to improve the accuracy and relevance of generated text by introducing information from external knowledge bases or document sets when generating text.

[0047] In one example, the tag thinking path includes multiple levels of associated information, and the multiple levels of associated information form a thinking tree.

[0048] In this example, the large model can be deduced in the form of a thinking tree. The thinking tree (ToT) is to think about the path through labels, so that the fine-tuned large model can think in the form of a tree when reasoning, and go down in the nodes of the tree until it reaches the leaf node for output.

[0049] Furthermore, the plurality of associated information is determined in the following manner:

[0050] Determining first associated information based on the concept classification and the requested task;

[0051] According to the first associated information, the first precautions are determined as the second associated information, and the second precautions are determined as the third associated information.

[0052] In this example, the first associated information, the second associated information and the third associated information are all nodes of the thinking tree. The first associated information and the second associated information form a thinking path, and the first associated information and the third associated information form another thinking path. The two thinking paths together constitute the thinking tree.

[0053] In one example, the label answer is determined as follows:

[0054] Based on the association information, searching for target document information from the retrieved documents;

[0055] The target document information is input into a first neural network model for inference to obtain the label answer.

[0056] In this example, a neural network can be used to automatically generate label answers, avoiding the inefficiency of manual labeling.

[0057] In one example, obtaining a training data set consisting of multiple samples includes:

[0058] Acquire a first number of manually annotated samples to form a first data subset;

[0059] Taking the samples in the first data subset as examples, a second number of samples are generated using a second neural network model to form a second data subset; the first data subset and the second data subset constitute the training data set.

[0060] In this example, the thinking path in the vertical field requires the help of human prior knowledge. It is a kind of path learning (path-learning). This kind of path reasoning with a prior order can greatly improve the correctness of model reasoning. Some seed data are set up based on human prior concept classification: examples involving concept classification and thinking paths under the concept intention. These seed data are used as few-shot learning examples to spread the intent-thought of all training data, and then train them, so that the model can learn the human experience thinking path under different intentions. In addition to the original form of chain of thought (CoT), it also has the thinking method of tree of thought (ToT).

[0061] Chain of Thought (CoT) refers to the large model thinking in a chain-like manner when reasoning.

[0062] The above-mentioned thinking path is composed of multiple nodes, and each node can be, but is not limited to, related information of a certain concept classification, such as precautions under the concept intention.

[0063] Furthermore, the second neural network model has more parameters than the large model.

[0064] In the embodiments of this specification, special tokens are used. They are some special symbols or marks introduced by the large model tokenizer when processing text data. These marks are used to implement specific functions and help the model better understand and process input data. They are usually decoded into a single identification ID.

[0065] In one example, the input text includes a first type of special symbols, which are used to instruct the large model to think or answer.

[0066] In this example, special symbols are used to let the model perform different tasks. For example, <|Reason|> represents thinking and <|ANSWER|> represents answering.

[0067] In one example, the expected response includes a second type of special symbols, which are used to mark a label thinking path or a label answer.

[0068] In this example, the special symbols are used to make the model output a labeled response, so that the user can easily identify the different parts of the response through the labels, for example, whether it belongs to a thinking path or an answer.

[0069] Then, in step 22, the instruction fine-tuning of the large model is performed using the training data set to obtain a large model in the target domain. It can be understood that instruction fine-tuning is a supervised training, and the performance of the large model can be improved by adjusting the model parameters.

[0070] In one example, fine-tuning the instruction of the large model using the training data set includes:

[0071] Inputting the input text into the large model to obtain an output response corresponding to the input text;

[0072] The model parameters of the large model are adjusted with the minimization of the difference between the output response and the expected response as the training objective.

[0073] In this example, instruction fine-tuning can enable the large model to learn the empirical thinking paths under different intentions, thereby improving the performance of the large model.

[0074] Figure 3 A schematic diagram of a thinking path in a vertical field according to an embodiment is shown. Figure 3 , the thinking path in this vertical field can be divided into three steps: concept classification, concept recall, and reasoning based on the associated information and document combination of concept recall. Among them, concept classification is to identify the intent by combining user questions and recalled documents, so that the subsequent model can take different paths according to different intents; concept recall enables the model to know what to focus on under different concepts and pay attention to some precautions to avoid going into the wrong reasoning path when using documents for reasoning later; reasoning based on the associated information and document combination of concept recall, that is, combining different intents to find the required content in the document for reasoning and answering.

[0075] The following is an example of an expected response:

[0076] [1] This question involves the type of ##service guide## for ##medical treatment in other places##. Through the interpretation of the concept of intent, we can know that this type of question usually deals with the situation where the insured person seeks medical treatment in a place other than the place where he / she is insured, especially cross-provincial medical treatment. According to the specific content of the question, this inquiry is about the registration of medical treatment in other places from Shaoxing to Hangzhou.

[0077] [2][2.1] First, regarding the identification of the place of insurance participation and the place of medical treatment, the question does not clearly state the place of insurance participation, but based on the question "from Shaoxing to Hangzhou", we can infer that the place of insurance participation should be Shaoxing and the place of medical treatment should be Hangzhou.

[0078] [2][2.2] Confirm the relevant policies of the place of insurance participation and the place of medical treatment [3] From the reference materials, we can see that ##begin_quote##No registration is required for medical treatment in other places within Zhejiang Province, but registration is required for medical treatment outside the province##end_quote## <cite>1< / cite> Therefore, it is not necessary to register for medical treatment in other places from Shaoxing to Hangzhou.

[0079] [3] Next, we need to find the specific steps for how to apply for a non-local medical treatment record. From the reference materials, we can see that ##begin_quote##… Click [Health Care] on the homepage or search for “Health Care” to enter… the [Health Care] applet, then click [All Services], click [Medical Insurance] – [Medical Insurance Application] – [Non-local Medical Treatment Record], and then follow the instructions to complete the operation##end_quote## <cite>2< / cite> .

[0080] <|ANSWER|>The registration process for medical treatment in other places from Shaoxing to Hangzhou is as follows:

[0081] 1. **Registration requirements**: - Shaoxing to Hangzhou is considered out-of-province medical treatment. According to Zhejiang Province policy, out-of-province medical treatment does not require registration.

[0082] 2. **Registration steps**: - Since there is no need to register for medical treatment in other places within the province, there is no need to register for medical treatment in other places from Shaoxing to Hangzhou.

[0083] Therefore, people from Shaoxing to Hangzhou can directly enjoy medical services without filing a record. "

[0084] As can be seen from the above, the expected response consists of two parts, namely the label thinking path and the label answer. Among them, [1], [2][2.1], [2][2.2], and [3] in the above examples are special symbols, which are used to identify nodes at different levels in the label thinking path. Specifically, [1] represents concept classification, [2] represents concept recall, [2.1] represents a node under concept recall, [2.2] represents another node under concept recall, and [3] represents reasoning based on the associated information and document combination of concept recall. <|ANSWER|> is used to identify the label answer. It can be understood that the expected response can also contain some other special symbols, such as ## in the above example, which also have corresponding meanings and will not be repeated here.

[0085] In the embodiments of this specification, the big model can learn the experiential thinking path in the target field, which includes how to use the content in the document. For example, in the medical insurance scenario, when asking about prenatal check-up subsidies, it actually implies a premise that it is subordinate to maternity subsidies. The big model needs to retrieve content related to maternity subsidies during RAG, generate fine-tuning through RAG in vertical fields, and use human prior knowledge as the thinking path, so that the big model knows the different focus points under different concept categories, so as to give better answers.

[0086] Through the method provided in the embodiment of this specification, a training data set consisting of multiple samples is first obtained; any sample includes an input text and an expected response, the input text includes a request task in the target field and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a conceptual classification of the request task and associated information determined based on the conceptual classification; and then the training data set is used to fine-tune the instructions of the large model to obtain a large model in the target field. As can be seen from the above, the embodiment of this specification adopts RAG technology and instruction fine-tuning, and uses human prior knowledge as a label thinking path, so that the large model learns to adopt different thinking paths under different concept classifications, so as to give better answers, which can further improve the performance of the large model in vertical fields.

[0087] According to another aspect of the embodiment, a large model fine-tuning device in a target domain is also provided, and the device is used to execute the method provided in the embodiment of this specification. Figure 4 FIG. 4 is a schematic block diagram of a large model fine-tuning device in a target domain according to an embodiment. Figure 4 As shown, the device 400 includes:

[0088] The acquisition unit 41 is used to acquire a training data set consisting of multiple samples; any sample includes an input text and an expected response, the input text includes a request task in the target field and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification;

[0089] The fine-tuning unit 42 is used to fine-tune the instructions of the large model using the training data set obtained by the obtaining unit 41 to obtain the large model in the target domain.

[0090] Optionally, as an embodiment, the input text also includes a number of retrieved documents.

[0091] Optionally, as an embodiment, the tag thinking path includes multiple levels of associated information, and the multiple levels of associated information form a thinking tree.

[0092] Furthermore, the plurality of associated information is determined in the following manner:

[0093] Determining first associated information based on the concept classification and the requested task;

[0094] According to the first associated information, the first precautions are determined as the second associated information, and the second precautions are determined as the third associated information.

[0095] Optionally, as an embodiment, the label answer is determined in the following manner:

[0096] Based on the association information, searching for target document information from the retrieved documents;

[0097] The target document information is input into a first neural network model for inference to obtain the label answer.

[0098] Optionally, as an embodiment, the acquiring unit 41 includes:

[0099] An acquisition subunit, used for acquiring a first number of manually labeled samples to form a first data subset;

[0100] A generating subunit is used to use the samples in the first data subset acquired by the acquiring subunit as examples, and to generate a second number of samples to form a second data subset using a second neural network model; the first data subset and the second data subset constitute the training data set.

[0101] Furthermore, the second neural network model has more parameters than the large model.

[0102] Optionally, as an embodiment, the input text includes a first type of special symbols, which are used to instruct the large model to think or answer.

[0103] Optionally, as an embodiment, the expected response includes a second type of special symbols, which are used to mark label thinking paths or label answers.

[0104] Optionally, as an embodiment, the fine-tuning unit 42 includes:

[0105] A prediction subunit, used for inputting the input text into the large model to obtain an output response corresponding to the input text;

[0106] The adjustment subunit is used to adjust the model parameters of the large model by taking minimizing the difference between the output response obtained by the prediction subunit and the expected response as the training goal.

[0107] Through the device provided in the embodiment of this specification, first, the acquisition unit 41 acquires a training data set consisting of multiple samples; any sample includes an input text and an expected response, and the input text includes a request task in the target field and an instruction to execute the request task; the expected response includes a label thinking path, a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a conceptual classification of the request task and associated information determined based on the conceptual classification; then the fine-tuning unit 42 uses the training data set to perform instruction fine-tuning on the large model to obtain a large model in the target field. As can be seen from the above, the embodiment of this specification adopts RAG technology and instruction fine-tuning, and uses human prior knowledge as a label thinking path, so that the large model learns to adopt different thinking paths under different concept classifications, so as to give better answers, which can further improve the performance of the large model in vertical fields.

[0108] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 2 The method described.

[0109] According to another embodiment of the present invention, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the Figure 2 The method described.

[0110] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0111] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A large model fine-tuning method in a target domain, the method comprising: Obtain a training data set consisting of multiple samples; Any sample includes an input text and an expected response, wherein the input text includes a request task in the target domain and an instruction for executing the request task; the expected response includes a label thinking path and a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification; The training data set is used to fine-tune the instructions of the large model to obtain a large model in the target domain.

2. The method of claim 1, wherein: The input text also includes a number of retrieved documents.

3. The method of claim 1, wherein: The tag thinking path includes multiple levels of associated information, and the multiple levels of associated information form a thinking tree.

4. The method of claim 3, wherein: The plurality of associated information is determined in the following manner: Determining first associated information based on the concept classification and the requested task; According to the first associated information, the first precautions are determined as the second associated information, and the second precautions are determined as the third associated information.

5. The method of claim 1, wherein: The label answer is determined as follows: Based on the association information, searching for target document information from the retrieved documents; The target document information is input into a first neural network model for inference to obtain the label answer.

6. The method of claim 1, wherein: The step of obtaining a training data set consisting of a plurality of samples includes: Acquire a first number of manually annotated samples to form a first data subset; Taking the samples in the first data subset as examples, a second number of samples are generated using a second neural network model to form a second data subset; the first data subset and the second data subset constitute the training data set.

7. The method of claim 6, wherein: The second neural network model has more parameters than the large model.

8. The method of claim 1, wherein: The input text includes a first type of special symbols, which are used to instruct the large model to think or answer.

9. The method of claim 1, wherein: The expected response includes a second type of special symbols, which are used to mark label thinking paths or label answers.

10. The method of claim 1, wherein: The step of fine-tuning the instructions of the large model using the training data set includes: Inputting the input text into the large model to obtain an output response corresponding to the input text; The model parameters of the large model are adjusted with the minimization of the difference between the output response and the expected response as the training objective.

11. A large model fine-tuning device in a target domain, the device comprising: An acquisition unit, used for acquiring a training data set consisting of multiple samples; Any sample includes an input text and an expected response, wherein the input text includes a request task in the target domain and an instruction for executing the request task; the expected response includes a label thinking path and a label answer obtained based on the label thinking path and a number of retrieved documents; the label thinking path includes a concept classification of the request task and associated information determined based on the concept classification; A fine-tuning unit is used to fine-tune the instructions of the large model using the training data set obtained by the acquisition unit to obtain the large model in the target domain.

12. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.

13. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Model training method, text processing method and related device

    CN121188470A