Model supervision fine tuning control method and device and medium

Through the splicing of the dialogue dataset and the adjustment of attention modules, the problem of inconsistent sequence lengths in large language model training is solved, efficient and accurate model fine-tuning training is achieved, and the model's performance ability to specific tasks is improved.

CN120597966APending Publication Date: 2025-09-05ZHEJIANG LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510697732.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the iterative fine-tuning training of large language models, due to the inconsistent length of the dialogue sample sequence, the computational resources are wasted and the training efficiency is low, and the filling strategy introduces meaningless filling marks, affecting the model performance.

Method used

By splicing the dialogue dataset, adjusting the attention module, isolating information between different initial dialogue samples, compute in parallel, reducing the number of invalid tokens, and optimizing the training process.

Benefits of technology

It improves the efficiency and accuracy of model training, reduces meaningless padding, and improves the model's expressiveness in responding to specific tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597966A_ABST
    Figure CN120597966A_ABST
Patent Text Reader

Abstract

The invention discloses a model supervision fine tuning control method and device and a medium. The method comprises the following steps: acquiring a dialogue data set and a pre-constructed initial large model; splicing the initial dialogue samples in the dialogue data set to obtain a plurality of spliced samples; according to the spliced samples, an attention module in the initial large model is adjusted, so that when the attention module calculates the spliced samples, information between different initial dialogue samples in the spliced samples is mutually isolated; and carrying out iterative fine tuning training on the adjusted initial large model through sample splicing to obtain a target large model. Therefore, by splicing the dialogue samples, introduction of a large number of meaningless filling marks through a sample filling strategy is avoided, and the number of invalid tokens of the samples is reduced. Besides, the attention module is adjusted according to the spliced samples, so that different initial dialogue samples in the spliced samples are isolated, the context attention is prevented from deviating, and the model training efficiency is improved while the model precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a control method, device, and medium for model supervised fine-tuning. Background Art

[0002] With the rapid development of artificial intelligence (AI), large language models are widely used in question-answering systems across various fields. To improve the performance of large language models across different tasks, supervised learning can be used to optimize pre-trained large models, thereby enhancing their performance. Specifically, dialogue datasets are used as input for iterative training of the large model, improving its ability to generate corresponding answers to specific commands.

[0003] During iterative fine-tuning training, the sequence lengths of different dialogue samples in a dialogue dataset vary significantly, hindering the learning of large models and ultimately causing them to fall short of expectations. Currently, to address this technical issue, a padding strategy is commonly employed. This involves padding shorter samples with tokens to ensure that all input dialogue samples have the same sequence length. However, while this approach can achieve uniform sequence length, shorter sequences introduce a large number of meaningless padding tokens, resulting in wasted computational and storage resources and reduced model training efficiency.

[0004] Therefore, how to improve the efficiency of model training without sacrificing the accuracy of the large model and avoiding the waste of computing and storage resources, thereby improving the expressiveness of the large model in responding to specific tasks, is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the Invention

[0005] In view of this, one aspect of the present application provides a control method for model-supervised fine-tuning, the method comprising:

[0006] Obtain a conversation dataset and a pre-built initial large model;

[0007] splicing the initial conversation samples in the conversation dataset to obtain a plurality of spliced ​​samples;

[0008] Adjusting the attention module in the initial large model according to the spliced ​​samples so that when the attention module calculates the spliced ​​samples, information between different initial dialogue samples in each spliced ​​sample is isolated from each other;

[0009] The adjusted initial large model is iteratively fine-tuned and trained using the spliced ​​samples to obtain a target large model.

[0010] Optionally, concatenating the initial conversation samples in the conversation dataset includes:

[0011] Segmenting the character string in the initial conversation sample into word units;

[0012] The initial dialogue samples are spliced ​​according to a splicing rule in which the initial dialogue samples are used as splicing units and the number of word units included in the splicing samples is less than a threshold; wherein adjacent initial dialogue samples are separated by separators.

[0013] Optionally, the splicing rule further includes: performing splicing based on a sorting result, where the sorting result is a sorting of the initial dialogue samples according to the number of included word units.

[0014] Optionally, the splicing rule further includes: performing splicing based on the original order of the initial conversation samples.

[0015] Optionally, adjusting the attention module in the initial large model according to the spliced ​​sample includes:

[0016] The first word unit after each delimiter is used as the starting position, and a position ID is sequentially configured for each word unit in the spliced ​​sample starting from the starting position; different starting positions correspond to the same position ID;

[0017] Determining the cumulative length of the position of the spliced ​​sample according to the position ID;

[0018] A block diagonal mask in the attention module is generated according to the spliced ​​samples, the separator and the position cumulative length.

[0019] Optionally, before iteratively fine-tuning the adjusted initial large model using the spliced ​​samples, the method further includes:

[0020] Determining the number of target dialogue samples included in the target splicing sample and the number of word units included in the target dialogue sample; wherein the target splicing sample is the sample currently input into the initial large model;

[0021] Collecting the loss function of the initial large model after each iterative fine-tuning training to obtain a loss function sequence;

[0022] The target loss function in the current iterative fine-tuning training of the initial large model is determined according to the number of samples, the number of word units and the loss function sequence.

[0023] Optionally, before iteratively fine-tuning the adjusted initial large model using the spliced ​​samples, the method further includes:

[0024] Obtaining a first number of the initial conversation samples and a second number of the spliced ​​samples;

[0025] The global batch size of the initial large model is adjusted according to the first number and the second number; wherein, when the second number is larger, the global batch size is larger.

[0026] Another aspect of the present application provides a control device for model-supervised fine-tuning, the device comprising:

[0027] The acquisition module is used to obtain the dialogue dataset and the pre-built initial large model;

[0028] a splicing module, configured to splice the initial conversation samples in the conversation dataset to obtain a plurality of spliced ​​samples;

[0029] an adjustment module, configured to adjust an attention module in the initial large model according to the spliced ​​samples, so that when the attention module calculates the spliced ​​samples, information between different initial dialogue samples in each of the spliced ​​samples is isolated from each other;

[0030] The training module is used to iteratively fine-tune the adjusted initial large model through the spliced ​​samples to obtain a target large model.

[0031] Another aspect of the present application provides a control device for model-supervised fine-tuning, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps of the control method for model-supervised fine-tuning are implemented.

[0032] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the control method of model-supervised fine-tuning when the program is executed by a processor.

[0033] This application provides a control method, device, and medium for supervised model fine-tuning, which have the following beneficial effects: by splicing conversation samples, the introduction of a large number of meaningless filler tags through sample padding strategies can be avoided. In other words, the number of invalid tokens in the samples can be reduced, thereby improving the efficiency of model fine-tuning training. In addition, the attention module in the large model is adjusted based on the spliced ​​samples to isolate the different initial conversation samples in the spliced ​​samples, avoiding deviations in contextual attention, thereby ensuring model accuracy while improving model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flow chart of a control method for model supervision fine-tuning provided in an embodiment of the present application;

[0035] Figure 2 A data diagram of a conversation dataset provided in an embodiment of the present application;

[0036] Figure 3 A flowchart of a control method for model supervision fine-tuning provided in another embodiment of the present application;

[0037] Figure 4 A schematic diagram of the structure of a control device for model supervision fine-tuning provided in an embodiment of the present application;

[0038] Figure 5 A schematic structural diagram of a control device for model supervision fine-tuning provided in another embodiment of the present application.

[0039] The accompanying drawings are marked as follows: 40 is an acquisition module, 41 is a splicing module, 42 is an adjustment module, 43 is a training module, 50 is a memory, 51 is a processor, 52 is a display screen, 53 is an input and output interface, 54 is a communication interface, 55 is a power supply, 56 is a communication bus, 501 is a computer program, 502 is an operating system, and 503 is data. DETAILED DESCRIPTION

[0040] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0041] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0042] Figure 1 A flow chart of a control method for model supervision fine-tuning provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:

[0043] S10: Obtain the conversation dataset and the pre-built initial large model;

[0044] In a specific embodiment, the obtained dialogue dataset is used in the instruction fine-tuning stage of the large language model to improve the model's ability to reply to specific instructions. Among them, the specific task can be a question-and-answer task in the field of astronomy or a knowledge question-and-answer in the medical field. The present application does not limit the specific task. In addition, the pre-constructed initial large model can be a GPT (Generative Pre-trained Transformer) series, a LLaMA (Large Language Model Application) series, or any large model based on the Transformer architecture. The present application does not limit the initial large model.

[0045] In an optional embodiment, the dialogue dataset can be in Json format. Each initial dialogue sample in the dialogue dataset is abstracted into a dictionary type. Among them, the key-value definition of each field includes "system" for defining the system role or background setting, "user" for the content input by the user, and "assistant" for the content of the system's answer. A dictionary is a data structure commonly used to store key-value pairs (key-value). The key is unique, and the value can be any type of data. For example: {"name": "Zhang San", "age": 25}, where "name" and "age" are keys, and "Zhang San" and 25 are values.

[0046] Figure 2 The data schematic diagram of a dialogue dataset provided by an embodiment of the present application is as Figure 2 shown. The dialogue dataset mainly includes two parts: input Instruction and output Answer. Among them, the input Instruction includes "system" for defining the system role or background setting and "user" for the content input by the user, and the output Answer refers to the content of the system's answer "assistant".

[0047] In an example, as Figure 2 shown, "system" is "Yow are a helpful assistant", "user" is "Compare the advantages of optimistic lock and pessimistic lock?", and "assistant" is "Optimistic lock and pessimistic lock are in database concurrency control...". In a specific embodiment, the number of initial dialogue samples in the dialogue dataset is m, where m is a positive integer. The dialogue samples can be constructed manually or obtained by collecting user historical data. The present application does not limit this.

[0048] S11: Concatenate the initial dialogue samples in the dialogue dataset to obtain multiple concatenated samples;

[0049] S12: Adjust the attention module in the initial large model based on the spliced ​​samples so that when the attention module calculates the spliced ​​samples, the information between different initial dialogue samples in each spliced ​​sample is isolated from each other;

[0050] Furthermore, the initial conversation samples in the acquired conversation dataset are concatenated to obtain multiple concatenated samples. It is important to note that during the concatenation process, to ensure the semantic integrity of different conversation samples, concatenation is performed on an initial conversation sample basis. Furthermore, to ensure the training efficiency of the initial large model, the string lengths included in the concatenated samples are smaller than a threshold. It should be noted that the acquired initial conversation samples can be concatenated arbitrarily, as long as the initial conversation samples are not segmented and the length of the concatenated sample string is smaller than a threshold.

[0051] It should be noted that before splicing the initial conversation samples, in order to further improve the efficiency of model training, the initial conversation samples are first filtered. The filtering includes but is not limited to filtering out low-quality samples and deleting redundant character strings in the samples.

[0052] After obtaining the spliced ​​sample, the initial large model's attention module is adjusted based on information such as the number of initial dialogue samples included in the spliced ​​sample and the number of characters in each initial dialogue sample. In essence, this can be understood as adjusting the initial large model's attention mechanism based on the spliced ​​sample so that when calculating the spliced ​​sample, the initial large model isolates the information between the initial dialogue samples included in the spliced ​​sample. In other words, it controls which information in the spliced ​​sample is visible to each other and which information needs to be isolated.

[0053] S13: By splicing samples, the adjusted initial large model is iteratively fine-tuned to obtain the target large model.

[0054] After splicing the initial dialogue samples and adjusting the attention mechanism in the initial large model, the adjusted initial large model is iteratively fine-tuned through the spliced ​​samples to obtain a target large model that can respond quickly and accurately to specific tasks.

[0055] In an optional embodiment, when calculating the spliced ​​samples, the initial large model calculates mutually visible information within the spliced ​​samples independently. Independent information can be calculated in parallel, thereby improving the efficiency of model fine-tuning training. Specifically, information between different initial conversation samples is isolated, while information within the same initial conversation sample is visible to each other. Therefore, in this specific embodiment, when calculating each spliced ​​sample, different initial conversation samples are calculated independently and in parallel, while different information within the same initial conversation sample is calculated sequentially.

[0056] Therefore, by splicing the conversation samples, we avoid introducing a large number of meaningless padding tokens through sample padding strategies. This reduces the number of invalid tokens in the samples, thereby improving the efficiency of model fine-tuning and training. Furthermore, the attention module in the large model is adjusted based on the spliced ​​samples to isolate the different initial conversation samples in the spliced ​​samples, preventing deviations in contextual attention. This ensures model accuracy while improving training efficiency.

[0057] In an optional embodiment, the initial conversation samples in the conversation dataset are spliced, including:

[0058] Segment the character strings in the initial conversation sample into word units;

[0059] The initial dialogue samples are taken as the splicing units and the number of word units included in the splicing samples is less than a threshold, so as to splice the initial dialogue samples; wherein adjacent initial dialogue samples are separated by separators.

[0060] In order to ensure the efficiency of model fine-tuning training, the length of the spliced ​​samples cannot exceed the maximum training length of the model. Therefore, in a specific embodiment, the character string in the initial conversation sample can be segmented into word units (Tokens) by a Tokenizer. After segmentation, in an optional embodiment, the sample is preprocessed, wherein the preprocessing includes but is not limited to removing punctuation marks, converting to uppercase and lowercase, and processing special characters, thereby reducing the computational complexity of the model. It should be noted that the token after segmentation can be a phrase, a punctuation mark, an English word, etc., and this application does not limit this.

[0061] Furthermore, the initial conversation samples are spliced ​​based on pre-set splicing rules. Specifically, to ensure model training efficiency, the number of tokens included in the spliced ​​samples cannot exceed a threshold, where the threshold is the maximum number of tokens allowed in the initial large model.

[0062] Furthermore, to ensure that semantic information within the same initial conversation sample is not fragmented, thereby ensuring effective model training, in this specific embodiment, the initial conversation sample is used as the unit for splicing. In other words, the initial conversation sample in each spliced ​​sample remains the original sample, and the structure of the original sample is not destroyed. For ease of understanding, the following example illustrates this.

[0063] For example, the threshold is 800, that is, the maximum allowed number of tokens in the initial large model is 800. The initial conversation sample includes sample A1, sample A2, sample A3, sample A4 and sample A5. After the initial conversation sample is segmented, sample A1 includes 200 tokens, sample A2 includes 200 tokens, sample A3 includes 300 tokens, sample A4 includes 500 tokens, and sample A5 includes 100 tokens.

[0064] Based on the threshold of 800, when the four initial samples are spliced ​​together, the total number of tokens in samples A1, A2, and A3 is 700. If sample A4 is added, the total number of tokens exceeds the threshold. Therefore, samples A1, A2, and A3 are spliced ​​together to form a spliced ​​sample. Samples A4 and A5 are spliced ​​together to form another spliced ​​sample.

[0065] It is worth noting that, during the splicing process, the splicing can be performed in the order of the initial samples obtained, or it can be performed randomly, and this application does not limit this.

[0066] In an alternative embodiment, it is understood that ensuring that the number of tokens included in each spliced ​​sample is as close to a threshold as possible can maximize model training efficiency. Therefore, in a specific embodiment, the initial conversation samples are first sorted according to the number of tokens included in each initial conversation sample, and then spliced ​​based on the sorting results.

[0067] That is to say, based on the above embodiment, as an optional embodiment, the splicing rule for splicing the initial conversation samples further includes splicing based on a sorting result, where the sorting result is a sorting of the initial conversation samples according to the number of tokens included.

[0068] It should be noted that, in a specific embodiment, the initial conversation samples can be sorted in ascending or descending order when they are spliced ​​together, and this application does not impose any restrictions on this. In an optional embodiment, in order to ensure that multiple initial conversation samples with a small number of tokens are spliced ​​into a spliced ​​sample and that the continuity of the contextual semantics is maintained as much as possible, the initial conversation samples can be sorted in ascending order.

[0069] After ascending sorting, the initial conversation samples are further spliced ​​in sequence according to the sorting results. For ease of understanding, the following example will be used.

[0070] For example, the threshold is 500, that is, the maximum allowed number of tokens in the initial large model is 500. The initial conversation sample includes sample A1, sample A2, sample A3, sample A4 and sample A5. After the initial conversation sample is segmented, sample A1 includes 200 tokens, sample A2 includes 80 tokens, sample A3 includes 50 tokens, sample A4 includes 100 tokens, and sample A5 includes 400 tokens.

[0071] After sorting the above samples in ascending order, they are sample A3, sample A2, sample A4, sample A1, and sample A5. Based on the sorting results, samples A3, A2, A4, and A1 are concatenated into a single concatenated sample, which contains 430 tokens. Sample A5 is used as a separate concatenated sample for fine-tuning the initial large model.

[0072] Therefore, the initial conversation samples are sorted first, and then the samples are spliced ​​based on the sorting results, which further reduces the number of invalid tokens in the spliced ​​samples, thereby further improving the model training efficiency.

[0073] It is understandable that sorting the initial samples may disrupt the continuity between samples, that is, the randomness between samples, which in turn affects the semantic information between samples and reduces the accuracy of the model. Therefore, in order to maximize the accuracy of the model, based on the above embodiment, as an optional embodiment, the splicing rules also include: splicing based on the original order of the initial conversation samples. In other words, when splicing the initial conversation samples, the splicing is performed according to the original order in which the samples were acquired, with the initial conversation samples as the splicing units, and the number of word units included in the spliced ​​samples is less than a threshold.

[0074] For example, the initial conversation samples contained in the original conversation dataset in Json format are:

[0075]

[0076] Among them, i represents the content of the input instruction, that is, Figure 2 As shown, it includes system for defining system roles or background settings and user input content user. o is the output Answer, that is, Figure 2 The system's response content is shown as assistant.

[0077] t is a token, and l is used to represent the number of tokens included in the initial conversation sample. In fact, it can also be understood as the length of the initial conversation sample. For example, l1 represents the number of tokens included in the first initial conversation sample in the spliced ​​sample. Therefore, l mis the number of tokens included in the mth initial conversation sample in the splicing sample, Indicates that the number of tokens in the mth initial conversation sample is l m Specific Token.

[0078] After splicing the above initial conversation samples according to the splicing rules provided in the embodiment of the present application, the following splicing samples are obtained:

[0079] Splicing sample 1:

[0080] Splicing sample 2:

[0081] …

[0082] Splicing sample H:

[0083]

[0084] in, T is the threshold value. The formula indicates that for each completed splicing sample, the sum of all the tokens in the initial conversation sample is less than the threshold value (the maximum number of tokens allowed in the initial large model). n It is used to represent the sequence number (sort position) of the last initial conversation sample in the current spliced ​​sample in the original conversation dataset. For example, p1=5 means that the current spliced ​​sample splices 5 initial conversation samples, and the next spliced ​​sample starts splicing from the 6th initial conversation sample, that is, p1+1=6.

[0085] is the set of tokens included in an initial conversation sample. The upper right script represents the sequence number (sort position) of the initial conversation sample in the original conversation dataset, and the lower right script represents the sequence number of the token. An initial conversation sample consists of l m Tokens. EOS is the separator between initial conversation samples. n is the number of initial conversation samples included in the spliced ​​sample, and H is the number of spliced ​​samples after m initial conversation samples are spliced ​​together.

[0086] Therefore, by splicing the original initial conversation samples in the original order, the proportion of invalid tokens can be reduced while avoiding the shift of context attention, ensuring data diversity while avoiding destroying the intrinsic connection between samples.

[0087] In an optional embodiment, adjusting the attention module in the initial large model according to the spliced ​​samples includes:

[0088] The first word unit after each delimiter is taken as the starting position, and the position ID is assigned to each word unit in the spliced ​​sample starting from the starting position; different starting positions correspond to the same position ID;

[0089] According to the position ID, determine the cumulative length of the position of the spliced ​​sample;

[0090] Generate the block diagonal mask in the attention module based on the concatenated samples, separators, and position cumulative length.

[0091] In this specific embodiment, it is understood that after concatenating the acquired conversation dataset, the calculation of each concatenated sample requires contextual isolation between different initial conversation samples within the concatenated sample, while information within the same initial conversation sample must be mutually visible. Therefore, after obtaining the concatenated sample, the attention mechanism in the initial large model must be adjusted based on the concatenated sample. Specifically, the attention module must be adjusted so that different initial conversation samples within the same concatenated sample are mutually isolated, while information within the same initial conversation sample is mutually visible.

[0092] Specifically, the position of the token in each initial sample is first marked, that is, the position ID is configured for the token. In the above embodiment, EOS is used as a delimiter to isolate different initial conversation samples. Therefore, in a specific embodiment, the EOS delimiter can be identified, and the first token after the EOS delimiter is used as the starting position, and the position ID of each starting position is the same. It should be noted that the position ID of the starting position can start from 0 or from 1, and this application does not limit this. For ease of understanding, the following examples will be given.

[0093] For example, the concatenated sample is [A, B]EOS[C, D]EOS[E], where A and B are tokens from the first initial conversation sample, C and D are tokens from the second initial conversation sample, and E is a token from the third initial conversation sample. The first token after the EOS delimiter is used as the starting position. Therefore, A, C, and E are the starting position tokens, and their corresponding position ID is 0. Accordingly, the position ID of the concatenated sample can be expressed as [0, 1, 0, 1, 0].

[0094] It's understandable that by assigning position IDs to the tokens in different initial conversation samples within the concatenated sample, we can determine the length of each initial conversation sample and the cumulative position length of the concatenated sample. Specifically, the sum of the position IDs of the last token in each initial conversation sample is the cumulative position length of the concatenated sample.

[0095] Furthermore, the block diagonal mask (BDM) in the attention module is generated based on the spliced ​​samples, delimiters, and cumulative position length. It should be noted that in the attention mechanism, the attention mask (AttentionMask) is used to control which positions can "see" each other (i.e., participate in the attention calculation). In the spliced ​​samples, it is necessary to ensure that different initial dialogue samples cannot "see" each other, that is, they are isolated from each other to avoid confusion of contextual information.

[0096] The Block Diagonal Mask generated from the concatenated samples consists of multiple blocks of all-ones arranged diagonally, with all remaining zeros. This mask allows token information within a block to be visible to each other while maintaining complete isolation between blocks. It is suitable for scenarios where contextual information within subsequences is fully interactive while information between sequences is completely isolated. For ease of understanding, the following example illustrates this.

[0097] For example, the stitched sample after stitching is [A, B, C]EOS[D, E]EOS[F], and the position ID of the stitched sample is [0, 1, 2, 0, 1, 0]. The Block Diagonal Mask generated based on this stitched sample is:

[0098]

[0099] Among them, the causal mask of the first initial dialogue sample [A, B, C] is The Causal Mask of the second initial dialogue sample [A, B, C] is The CausalMask of the second initial dialogue sample [A, B, C] is [1]. It should be noted that the Causal Mask can make the information in the same initial dialogue sample visible to each other.

[0100] In this specific embodiment, because the attention calculations for different initial dialogue samples do not interfere with each other, that is, they are isolated from each other, they can be processed in parallel. Specifically, as an optional embodiment, the Variable-Length Attention module of the Flash Attention library can be used to implement attention calculations for spliced ​​sequences.

[0101] Therefore, based on the parallel computing characteristics of variable-length operators, efficient mask-based parallel processing is implemented in the attention module, effectively reducing the computational interference between different initial dialogue samples in the spliced ​​samples.

[0102] In an optional embodiment, before iteratively fine-tuning the adjusted initial large model by splicing samples, the method further includes:

[0103] Determining the number of target dialogue samples included in the target concatenated sample and the number of word units included in the target dialogue sample; wherein the target concatenated sample is the sample currently input into the initial large model;

[0104] Collect the loss function of the initial large model after each iterative fine-tuning training to obtain a loss function sequence;

[0105] Based on the number of samples, the number of word units, and the loss function sequence, the target loss function in the current iterative fine-tuning training of the initial large model is determined.

[0106] In a specific embodiment, in order to adapt the initial large model to the changes brought about by the splicing of the initial conversation samples, it is necessary to make corresponding adjustments to the training configuration, that is, to adjust the parameters of the initial large model. Parameters include but are not limited to the learning rate, global batch size, gradient clipping, etc., to ensure normal model convergence.

[0107] In an optional embodiment, the loss function of the large model is continuously optimized. Specifically, the loss function sequence is obtained based on historical iterative fine-tuning training. The number of target dialogue samples included in the target spliced ​​samples currently input to the initial large model and the number of word units included in the target dialogue samples are calculated. For details, see formula (1):

[0108]

[0109] in, Normalized Loss is the target loss function in the current iterative fine-tuning training of the initial large model, k is the number of target dialogue samples included in the target spliced ​​sample (for example, 3 dialogue samples are spliced), L i is the number of tokens included in the i-th target dialogue sample, is the loss function sequence, i.e., the loss function obtained after each iterative fine-tuning training of the initial large model. i For the i-th target dialogue sample in the loss function sequence The starting position in .

[0110] It should be noted that formula (1) describes the loss function sequence of each target dialogue sample, which is first divided by its own length, that is, normalized based on length, and then summed and divided by the number of subsequences. In an optional embodiment, the above formula (1) can also be expressed as: Normalized Loss = torch.sum(losses.view(-1)*loss_mask / loss_token_num) / pack_cnt.

[0111] Here, torch.sum is the summation function in the torch library, losses.view(-1) flattens the loss tensor of the target dialogue sample into a one-dimensional vector. "-1" indicates that the dimension size is automatically calculated to keep the total number of elements unchanged. loss_mask is a mask tensor with the same shape as the flattened loss, usually composed of 0s and 1s, which is used to specify which positions of the loss should be included (specified by 1) or ignored (specified by 0). loss_token_num is the number of tokens included in each target dialogue sample, and pack_cnt is the number of target dialogue samples spliced ​​in a target spliced ​​sample.

[0112] In an optional embodiment, before iteratively fine-tuning the adjusted initial large model by splicing samples, the method further includes:

[0113] Obtaining a first number of initial conversation samples and a second number of spliced ​​samples;

[0114] The global batch size of the initial large model is adjusted according to the first quantity and the second quantity; wherein, when the second quantity is larger, the global batch size is larger.

[0115] In a specific embodiment, in addition to the loss function, the global batch size (global-batch-size) needs to be adjusted. Specifically, a first number of all initial conversation samples in the conversation dataset and a second number of spliced ​​samples are counted, and the global-batch-size is further calculated based on the first number and the second number. For details, see formula (2):

[0116]

[0117] Among them, Adjusted Global Batch Size is the adjusted global batch size, origin_gbs is the second quantity, num_packs is the initial global batch size of the initial large model, and num_samples_in_dataset is the first quantity.

[0118] Therefore, through the embodiments of the present application and the above embodiments, after adjusting the loss function and the global batch size, the supervised fine-tuning process of the large language model can be accelerated on the hardware device by splicing samples.

[0119] Figure 3 This is a flow chart of a control method for model supervision fine-tuning provided in another embodiment of the present application. In order to make those skilled in the art more clear about the technical solution provided in this application, the following will be combined with Figure 3 For explanation. Figure 3As shown, the control method provided in this application mainly includes five parts, namely, obtaining a dialogue data set, initial dialogue sample splicing, model attention module adjustment, model parameter adjustment and model fine-tuning training.

[0120] In a specific embodiment, a large dataset of conversations is obtained, and the initial conversation samples in the conversation dataset are spliced ​​together based on the splicing rules provided in the above embodiment. Furthermore, the attention mechanism of the initial large model is adjusted based on the spliced ​​samples to ensure that different initial conversation samples in the spliced ​​samples are isolated from each other and that information within the same initial conversation sample is visible to each other. Furthermore, the hyperparameters of the initial large model (including but not limited to the global batch size, loss function, and learning rate) need to be adjusted. The resulting spliced ​​samples are then input into the parameter-adjusted large model, and the large model is iteratively fine-tuned to obtain a target large model that can answer specific tasks.

[0121] In an optional embodiment, the control method for model supervised fine-tuning provided in this application is compared with a baseline method. The baseline method refers to the fine-tuning process of the llama3.1 8B model trained using the original dataset based on the Pai-Megatron-Patch framework.

[0122] Table 1 is a comparative table of training results for different control methods with different parameter configurations. As shown in Table 1, the model supervised fine-tuning control method provided by this application, which uses the same training framework, training hyperparameters (except for the global batch size), number of training cards, and initial model as the baseline method, can reduce the number of training samples to 7.3% of the original data set, thereby achieving a significant reduction in end-to-end training time. Among them, the training time of the baseline method is 56 hours and 17 minutes. The training time of the control method of this application is related to the set global batch size, and can be reduced to 4 hours and 2 minutes at the lowest, with an average speed increase of about 13 times.

[0123] Table 1 Comparison of training results of different control methods corresponding to different parameter configurations

[0124]

[0125]

[0126] Table 2 shows a comparison of the training results of different control methods for different dialogue datasets. As shown in Table 2, the control method provided in this application scored on 14 benchmarks of the OpenCompass evaluation method. OpenCompass is an open source toolbox that focuses on providing functions for evaluating large-scale language models (LLMs). Benchmarking refers to the process of standardized evaluation of software, hardware, or model performance.

[0127] As shown in Table 2, the inference test scores obtained by the baseline method and the control method provided by this application show that the control method provided by this application maintains excellent model accuracy in all 14 benchmark tests. Compared with the baseline method, the control method provided by this application surpasses the baseline model in 8 tasks (with the highest improvement of 10.9%), and the average score of 14 tasks is up to 0.14% higher than the baseline.

[0128] The reliability of the training results of the control method provided in this application was verified in 14 benchmark tests covering areas such as mathematical reasoning, code generation, and common sense logic. Among them, the splicing method with a global batch size of 48 was used to train a model that surpassed the baseline model in 8 benchmark tests, with an average score improvement of 0.14% on 14 tasks. The accuracy of the resulting model of the splicing method with some global batch size configurations was at most about 0.8% lower than the baseline model, and did not significantly affect the model performance.

[0129] Table 2 Comparison of training results of different control methods corresponding to different dialogue datasets

[0130]

[0131]

[0132]

[0133] Therefore, the model-supervised fine-tuning control method provided in this application preprocesses the collected conversation dataset and concatenates the initial conversation samples to reduce the number of invalid tokens in the samples. Furthermore, during the actual attention module calculation, the operator's characteristic of variable sequence length is utilized to parallelize the calculations of each concatenated sequence, reducing the mutual influence between sequences and overcoming the time-consuming training problem of introducing a large number of invalid tokens in the initial conversation samples. This significantly reduces the number of data samples, improving training efficiency while maintaining the accuracy of the training results.

[0134] In the above embodiments, the control method for model supervised fine-tuning is described in detail. The present application also provides an embodiment corresponding to a control device for model supervised fine-tuning.

[0135] Figure 4 A schematic diagram of the structure of a control device for model supervision fine-tuning provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the device includes:

[0136] An acquisition module 40 is used to acquire a conversation dataset and a pre-built initial large model;

[0137] A splicing module 41 is used to splice the initial conversation samples in the conversation dataset to obtain multiple spliced ​​samples;

[0138] An adjustment module 42 is configured to adjust the attention module in the initial large model according to the spliced ​​samples so that when the attention module calculates the spliced ​​samples, information between different initial conversation samples in each spliced ​​sample is isolated from each other;

[0139] The training module 43 is used to iteratively fine-tune the adjusted initial large model by splicing samples to obtain a target large model.

[0140] In addition, the control device for model supervision fine-tuning provided in the embodiment of the present application further includes:

[0141] A segmentation module, used to segment the character strings in the initial conversation sample into word units;

[0142] The splicing submodule is used to splice the initial dialogue samples using the initial dialogue samples as splicing units and according to a splicing rule in which the number of word units included in the splicing samples is less than a threshold; wherein adjacent initial dialogue samples are isolated by separators.

[0143] A configuration module is used to take the first word unit after each delimiter as the starting position, and sequentially configure a position ID for each word unit in the spliced ​​sample starting from the starting position; different starting positions correspond to the same position ID;

[0144] The cumulative length determination module is used to determine the cumulative length of the spliced ​​sample according to the position ID;

[0145] The mask generation module is used to generate the block diagonal mask in the attention module based on the splicing samples, separators and position cumulative length.

[0146] A first determination module is configured to determine the number of target dialogue samples included in the target spliced ​​sample and the number of word units included in the target dialogue sample; wherein the target spliced ​​sample is the sample currently input into the initial large model;

[0147] The collection module is used to collect the loss function of the initial large model after each iterative fine-tuning training to obtain a loss function sequence;

[0148] The second determination module is used to determine the target loss function in the current iterative fine-tuning training of the initial large model based on the number of samples, the number of word units and the loss function sequence.

[0149] a quantity determination module, configured to obtain a first quantity of initial conversation samples and a second quantity of spliced ​​samples;

[0150] The third determination module is used to adjust the global batch size of the initial large model according to the first quantity and the second quantity; wherein, when the second quantity is larger, the global batch size is larger.

[0151] Figure 5 A schematic diagram of the structure of a control device for model supervision fine-tuning provided in another embodiment of the present application is shown as follows: Figure 5 As shown, the control device for model supervision fine-tuning includes: a memory 50 for storing a computer program;

[0152] The processor 51 is configured to implement the steps of the model-supervised fine-tuning control method mentioned in the above embodiment when executing a computer program.

[0153] The control device for model supervision fine-tuning provided in this embodiment may include but is not limited to a laptop computer or a desktop computer.

[0154] Among them, the processor 51 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 51 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), and a programmable logic array (PLA). The processor 51 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 51 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 51 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0155] The memory 50 may include one or more computer-readable storage media, which may be non-transitory. The memory 50 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 50 is at least used to store the following computer program 501, wherein, after the computer program is loaded and executed by the processor 51, it can implement the relevant steps of the model-supervised fine-tuning control method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 50 may also include an operating system 502 and data 503, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 502 may include Windows, Unix, Linux, etc. The data 503 may include but is not limited to the relevant data involved in the model-supervised fine-tuning control method, etc.

[0156] In some embodiments, the control device for model supervision fine-tuning may further include a display screen 52 , an input / output interface 53 , a communication interface 54 , a power supply 55 , and a communication bus 56 .

[0157] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the control device for model supervision fine-tuning, and may include more or fewer components than shown in the figure.

[0158] The control device for model-supervised fine-tuning provided in an embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the control method for model-supervised fine-tuning in the above embodiment.

[0159] It should be noted that although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

Claims

1. A control method for model supervised fine-tuning, characterized in that: The method comprises: Obtain a conversation dataset and a pre-built initial large model; splicing the initial conversation samples in the conversation dataset to obtain a plurality of spliced ​​samples; Adjusting the attention module in the initial large model according to the spliced ​​samples so that when the attention module calculates the spliced ​​samples, information between different initial dialogue samples in each spliced ​​sample is isolated from each other; The adjusted initial large model is iteratively fine-tuned and trained using the spliced ​​samples to obtain a target large model.

2. The control method of model-supervised fine-tuning according to claim 1, characterized in that: Splicing the initial conversation samples in the conversation dataset, including: Segmenting the character string in the initial conversation sample into word units; The initial dialogue samples are spliced ​​according to a splicing rule in which the initial dialogue samples are used as splicing units and the number of word units included in the splicing samples is less than a threshold; wherein adjacent initial dialogue samples are separated by separators.

3. The control method of model supervised fine-tuning according to claim 2, characterized in that: The splicing rule further includes: performing splicing based on a sorting result, wherein the sorting result is a sorting of the initial dialogue samples according to the number of included word units.

4. The control method of model supervised fine-tuning according to claim 2, characterized in that: The splicing rule further includes: performing splicing based on the original order of the initial conversation samples.

5. The control method of model-supervised fine-tuning according to claim 2, wherein: Adjusting the attention module in the initial large model according to the spliced ​​sample includes: The first word unit after each delimiter is used as the starting position, and a position ID is sequentially configured for each word unit in the spliced ​​sample starting from the starting position; different starting positions correspond to the same position ID; Determining the cumulative length of the position of the spliced ​​sample according to the position ID; A block diagonal mask in the attention module is generated according to the spliced ​​samples, the separator and the position cumulative length.

6. The control method of model-supervised fine-tuning according to claim 2, characterized in that: Before iterative fine-tuning training is performed on the adjusted initial large model using the spliced ​​samples, the following steps are also included: Determining the number of target dialogue samples included in the target splicing sample and the number of word units included in the target dialogue sample; wherein the target splicing sample is the sample currently input into the initial large model; Collecting the loss function of the initial large model after each iterative fine-tuning training to obtain a loss function sequence; The target loss function in the current iterative fine-tuning training of the initial large model is determined according to the number of samples, the number of word units and the loss function sequence.

7. The control method of model-supervised fine-tuning according to claim 1, characterized in that: Before iterative fine-tuning training is performed on the adjusted initial large model using the spliced ​​samples, the following steps are also included: Obtaining a first number of the initial conversation samples and a second number of the spliced ​​samples; The global batch size of the initial large model is adjusted according to the first number and the second number; wherein, when the second number is larger, the global batch size is larger.

8. A control device for model-supervised fine-tuning, characterized in that: The device comprises: The acquisition module is used to obtain the dialogue dataset and the pre-built initial large model; a splicing module, configured to splice the initial conversation samples in the conversation dataset to obtain a plurality of spliced ​​samples; an adjustment module, configured to adjust an attention module in the initial large model according to the spliced ​​samples, so that when the attention module calculates the spliced ​​samples, information between different initial dialogue samples in each of the spliced ​​samples is isolated from each other; The training module is used to iteratively fine-tune the adjusted initial large model through the spliced ​​samples to obtain a target large model.

9. A control device for model-supervised fine-tuning, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the model-supervised fine-tuning control method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the model-supervised fine-tuning control method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Long context key value cache optimization method and device, equipment and medium

    CN121188081A