Method and device for preprocessing training data, equipment and medium

By serializing and concatenating the training data and optimizing it according to the maximum sequence length of the model, the problem of excessive useless data in the training data is solved, and the training efficiency and resource utilization of large models are improved.

CN120687742APending Publication Date: 2025-09-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510856706.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing training data preprocessing methods are filled with too much useless data, resulting in low efficiency in training large models.

Method used

By obtaining multiple pieces of initial training data and serializing them, several initial training data sequences are selected according to the maximum sequence length of the model to be trained, and they are spliced ​​into a spliced ​​data sequence with a length less than or equal to the maximum sequence length, thereby reducing the use of invalid tokens.

Benefits of technology

It improves the training efficiency of large models, reduces computing costs and resource waste, and shortens training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687742A_ABST
    Figure CN120687742A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method and device for preprocessing training data, equipment and a medium. The scheme comprises the following steps: acquiring multiple pieces of initial training data; each piece of initial training data comprises first answer information representing preference answers and second answer information representing non-preference answers; performing serialization processing on each piece of initial training data to obtain a plurality of initial training data sequences; selecting a plurality of initial training data sequences from the plurality of initial training data sequences according to the maximum sequence length which can be recognized by a to-be-trained model; splicing the plurality of selected initial training data sequences to obtain a spliced data sequence; the length of the spliced data sequence is smaller than or equal to the maximum sequence length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for preprocessing training data. Background Art

[0002] With the development of artificial intelligence, more and more large models are emerging in our lives that can answer user questions and meet user needs. Generally, large models need to be trained based on training data before they can be put into use. To facilitate training of large models, the training data is preprocessed. However, existing methods of processing training data require padding with large amounts of irrelevant data, which reduces the efficiency of large model training.

[0003] Therefore, the technical problem of how to preprocess data to improve the training efficiency of large models is in urgent need of solution. Summary of the Invention

[0004] The embodiments of this specification provide a method, apparatus, device, and medium for preprocessing training data to solve the problem of excessive filling of useless data in existing methods for preprocessing training data, which affects the efficiency of large model training.

[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows: The present invention provides a method for preprocessing training data, including: Acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; Serialize each piece of initial training data to obtain multiple initial training data sequences; Selecting a plurality of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the to-be-trained model; Several selected initial training data sequences are spliced ​​together to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0006] An embodiment of this specification provides a device for preprocessing training data, including: An initial training data acquisition module is used to acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; A serialization processing module is used to serialize each piece of initial training data to obtain multiple initial training data sequences; A sequence selection module, configured to select a number of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the model to be trained; The splicing module is used to splice the selected initial training data sequences to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0007] An embodiment of this specification provides a device for preprocessing training data, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; Serialize each piece of initial training data to obtain multiple initial training data sequences; Selecting a plurality of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the to-be-trained model; Several selected initial training data sequences are spliced ​​together to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0008] An embodiment of this specification provides a computer-readable medium having computer-readable instructions stored thereon. The computer-readable instructions can be executed by a processor to implement a method for preprocessing training data.

[0009] At least one embodiment of the present disclosure can achieve the following beneficial effects: by obtaining multiple pieces of initial training data containing questions and corresponding preferred and non-preferred answers; selecting, based on the maximum sequence length recognizable by the model to be trained, several initial training data sequences from the multiple initial training data sequences obtained after serializing each piece of initial training data; and concatenating the several initial training data sequences to obtain a concatenated data sequence having a length less than or equal to the maximum sequence length. Using valid training data for concatenation can avoid the use of excessive useless data for padding, thereby avoiding excessive computation of useless data when training a large model using preprocessed data, thereby improving the efficiency of training the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this application. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0011] Figure 1 This is a schematic diagram of an existing batch-filled training data sequence; Figure 2 is a schematic diagram of a training data sequence provided in an embodiment of this specification; Figure 3 This is a flow chart of a method for preprocessing training data provided in an embodiment of this specification; Figure 4 The embodiment of this specification provides a schematic diagram of an attention mask; Figure 5 This specification provides the corresponding embodiment Figure 3 A schematic structural diagram of a device for preprocessing training data; Figure 6 This specification provides the corresponding embodiment Figure 3 A structural diagram of a device for preprocessing training data. DETAILED DESCRIPTION

[0012] To make the purpose, technical solutions, and advantages of one or more embodiments of this specification more clear, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of one or more embodiments of this specification.

[0013] In order to facilitate understanding of the technical solutions provided in the embodiments of this specification, several terms are now explained: Direct Preference Optimization (DPO) is a method for optimizing model output. It can adjust the model by directly using human preference data so that the model can produce results that are more in line with user expectations.

[0014] Preferred (Chosen) answers are those that, in the opinion of human evaluators, are more expected or higher-quality answers to a given prompt or question. These answers are more relevant to the question, more accurate, or in some way more consistent with human preferences and values.

[0015] Rejected responses: These are responses that, in the opinion of human evaluators, do not meet the expected response for a given prompt or question. These responses are irrelevant, inaccurate, inappropriate, or of low quality.

[0016] Prompt: This is text input into the model to guide or instruct it to produce a specific output. Prompts can be questions, instructions, partial sentences, or any other form of text. They stimulate the model to produce a relevant response or continue completing a specified task.

[0017] Reward model: This model assigns reward points to the model's Chosen answers and penalty points to the model's Reject answers. It is primarily used to evaluate the quality of the model's generated answers relative to human preferences.

[0018] Attention Mask: refers to a technology used to control the model's attention mechanism during deep learning.

[0019] Token: refers to a basic unit in the input text. This unit can be a word, subword, character, or any other meaningful language fragment.

[0020] Position Id: It is used to identify the position of each token in the sequence in the model.

[0021] Attention mechanism: is a resource allocation mechanism that allows the model to dynamically focus on the importance of different parts of the sequence when processing sequence data.

[0022] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0023] When training large models using DPO, all training samples need to be padded to the same length. The existing "batch padding" approach converts the chosen and rejected responses into sequences in two batches for each training sample, then pads multiple batches to the same length using invalid tokens.

[0024] To make the sequence in the prior art clearer, Figure 1 This is a schematic diagram of an existing batch-filled training data sequence.

[0025] like Figure 1 As shown in the figure, for two original training samples, sample a and sample b, sample a and sample b can be split into four sequences, each containing one answer message. Each sequence needs to be padded, and the padded sequences contain a large number of invalid tokens. Because invalid tokens do not participate in attention calculations, this practice leads to a significant waste of GPU memory, reducing memory utilization and significantly increasing the risk of memory overflow during model training. Furthermore, because the space occupied by invalid tokens that need to be padded is reduced, the number of valid samples that can be accommodated in each input batch is reduced, further increasing model training time.

[0026] In order to solve the defects in the prior art, this solution provides the following embodiments.

[0027] Figure 2 This is a schematic diagram of a training data sequence provided in the embodiment of this specification. Figure 2 As shown in the figure, for two original training samples, sample c and sample d, samples c and d can be split into four sequences, each containing one answer message. The four sequences are then concatenated together as a whole sequence. If the length of the concatenated four sequences does not meet the model's input length requirements, a small amount of invalid tokens can be used for padding to meet the model's sequence length requirements. If the length of the concatenated four sequences does meet the model's input length requirements, no additional token data is needed for padding. This reduces the number of invalid tokens required for padding, lowers the computational cost of model training, and improves model training efficiency.

[0028] Next, a method for preprocessing training data provided in an embodiment of the specification will be described in detail with reference to the accompanying drawings.

[0029] Figure 3 This is a flow chart of a method for preprocessing training data provided by an embodiment of this specification. From a program perspective, the execution body of the process can be a program installed on a server or terminal.

[0030] like Figure 3 As shown, the process may include the following steps: Step 302: Acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer.

[0031] In the embodiments of this specification, the first answer information and the second answer information may be answers to the same question. An initial training data may also include question information to which the first answer information and the second answer information are directed.

[0032] Initial training data can be obtained from a pre-set sample database. The sample database can store multiple pieces of initial training data. The multiple pieces of initial training data in the training sample database can be based on expert experience, historical questions and answers, manually written questions and answers, model-generated questions and answers, etc.

[0033] In the embodiments of this specification, the data stored in the sample database may be in the first form of question + first answer + second answer; or may be in the second form of question + first answer, and question + second answer. In actual applications, one question may correspond to multiple first answers, or may correspond to multiple second answers. If the data stored in the sample database is in the second form, then a first data may be randomly selected, and based on the question information contained in the first data, second data with the question information may be selected from the database, and the answer information in the selected second data may be spliced ​​together with the first data to obtain initial training data. For example, randomly select data 1: question 1 + first answer 1, based on question 1, question 1 + first answer 2, and question 1 + second answer 1 may be obtained from the sample database, then the initial training data question 1 + first answer 1 + first answer 2 + second answer 1 may be obtained. Alternatively, based on question 1, only question 1 + second answer 1 may be obtained from the sample database, then the initial training data question 1 + first answer 1 + second answer 1 may be obtained.

[0034] In actual applications, the initial training data can be data from scenarios such as medical diagnosis scenarios, legal and regulatory consultation scenarios, customer service scenarios, after-sales maintenance scenarios, and question-and-answer scenarios. Multiple pieces of initial training data may include training data from one or more scenarios. For example, multiple pieces of initial training data can all be training data from medical diagnosis scenarios; or, they can be training data from multiple scenarios such as medical diagnosis scenarios, customer service scenarios, and after-sales maintenance scenarios. As for which scenario the initial training data is obtained from, it can be obtained based on the actual model training needs and is not limited here.

[0035] Step 304: Serialize each piece of initial training data to obtain multiple initial training data sequences.

[0036] In the embodiments of this specification, an initial training data sequence can be obtained based on an initial training data, which can be in the form of question + preferred answer + non-preferred answer; or in the form of question + preferred answer + question + non-preferred answer. It is also possible to obtain an initial training data sequence containing two sequences based on an initial training data sequence, such as a sequence of question + preferred answer, and a sequence of question + non-preferred answer. The preferred answer in the initial training data sequence can be one or more; the non-preferred answer can be one or more, which can be specifically determined based on the number of preferred answers and non-preferred answers contained in the initial training data.

[0037] In an embodiment of the present specification, if the initial training data sequence includes two subsequences, the initial training data can be split to obtain a first sub-initial training data of question + preferred answer, and a second sub-initial training data of question + non-preferred answer. The split sub-initial training data can be serialized to obtain two subsequences corresponding to the initial training data. Alternatively, the initial training data can be serialized first. After the serialization is completed, the data sequence of non-preferred answers can be split out, and the data sequence corresponding to the question can be copied to form a subsequence with the split data sequence; the remaining data sequence of question + preferred answer forms another subsequence, thereby obtaining an initial training data sequence containing two subsequences. If the initial training data sequence contains only one sequence, the two subsequences can be spliced ​​together to obtain an initial training data sequence containing question + preferred answer + non-preferred answer.

[0038] In the initial training data sequence in the embodiments of this specification, the data sequence corresponding to the question may be before the data sequence corresponding to the answer, and the data sequence corresponding to the question may also be after the data sequence corresponding to the answer.

[0039] In the embodiments of this specification, serialization processing can mean converting text into a digital sequence. Specifically, the text in the initial training data can be segmented to obtain token information such as single words, words, or punctuation marks, and then the token information can be converted into a digital sequence using a preset conversion method. The digital sequence can be arranged in the order of the single words, words, or punctuation marks in the initial training data to obtain the initial training data sequence. The length of the digital sequence after the segmentation information is converted can be different.

[0040] In practical applications, the initial training data may be serialized using at least one of the following methods: ASCII code conversion, Unicode code point conversion, custom encoding, and hash encoding.

[0041] Step 306: Select a number of the initial training data sequences from the multiple initial training data sequences according to the maximum sequence length that can be recognized by the model to be trained.

[0042] In the embodiments of this specification, the maximum sequence length can represent the sequence length that can be input into the model to be trained during the DPO training process. Alternatively, the maximum sequence length can be determined based on the processing power of the computer. The specific method for determining the maximum sequence length can be found in related art and will not be further described here.

[0043] The model to be trained can be a deep learning model with a model framework but no prior training, a large model that has undergone fine-tuning, or a model already in use that requires regular training. A large model can refer to a model in the fields of machine learning, deep learning, or artificial intelligence that has a large number of parameters and a complex structure, such as at least one of the GPT series models, Claude 4 models, Inception models, and WaveNet models. These models can be applied to fields such as natural language processing, computer vision processing, and speech recognition.

[0044] In the embodiments of this specification, a preset algorithm can be used to select several initial training data sequences from multiple training data sequences to obtain a better allocation strategy. The preset algorithm can be any one of a knapsack algorithm, a greedy algorithm, and a dynamic programming algorithm.

[0045] In actual applications, two initial training data sequences with the same problem can both be selected, or only one of them can be selected. Whether both initial training data sequences with the same problem need to be selected can be set based on actual needs and is not specifically limited here.

[0046] Step 308: Splicing the selected initial training data sequences to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0047] In the embodiments of this specification, the initial training data sequences that need to be spliced ​​together can be spliced ​​together using a preset algorithm based on multiple optimal initial training data sequences obtained based on the maximum sequence length, so that the length of the spliced ​​data sequence can be equal to the maximum sequence length, or as close to the maximum sequence length as possible, so that a small number of invalid tokens can be used for filling, thereby reducing the number of invalid tokens contained in the processed training data sequence, further reducing the number of invalid tokens processed by the model, and increasing the number of valid data sequences processed by the model, thereby reducing the waste of model computing resources and the training time, and improving the efficiency of model training.

[0048] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.

[0049] Figure 3 The method in

[15] obtains multiple pieces of initial training data containing questions and corresponding preferred and non-preferred answers. Based on the maximum sequence length recognizable by the model to be trained, several initial training data sequences are selected from the multiple initial training data sequences obtained after serializing each piece of initial training data. These initial training data sequences are then concatenated to produce a concatenated data sequence whose length is less than or equal to the maximum sequence length. Using valid training data for concatenation avoids the use of excessive useless data for padding, thereby preventing excessive computation of useless data when training large models using preprocessed data, thereby improving the efficiency of training large models.

[0050] based on Figure 3 The present specification also provides some specific implementation plans of the method, which are described below.

[0051] Optionally, as an implementation manner, in the embodiment of this specification, if the length of the concatenated data sequence is less than the maximum sequence length, the method may further include: The concatenated data sequence is padded to obtain a padded sequence; the length of the padded sequence is equal to the maximum sequence length.

[0052] In the embodiments of this specification, the padded sequence may include several selected pieces of initial training data, and may also include invalid tokens used to pad the sequence, such as the [PAD] token and the [UNK] token. If the length of the concatenated data sequence is equal to the maximum sequence length, no padding is required for the concatenated data sequence.

[0053] In practical applications, invalid tokens do not carry any valid information and are only used to fill the sequence so that the length of the filled sequence meets the model input length requirement. For example, if the model requires an input length of 100 and the length of the concatenated data sequence is 90, then invalid tokens can be used to fill 10 sequence lengths so that the length of the filled sequence is 100, which can meet the input length requirement of the model, so that the model can recognize and process the padded sequence and successfully complete the model training task.

[0054] In the embodiments of this specification, the maximum sequence length can be determined based on the performance of the model and the performance of the hardware device that carries the model. For example, factors such as the model architecture, the number of parameters of the model, the computing resources occupied by the model, the memory capacity and the processor can all affect the maximum sequence length.

[0055] Optionally, as an implementation manner, the method of selecting several initial training data sequences from the multiple initial training data sequences based on the maximum sequence length recognizable by the model to be trained in the embodiments of this specification may specifically include: Determining a packing strategy using a knapsack algorithm based on the sequence length of each of the initial training data sequences and the maximum sequence length; According to the packaging strategy, several initial training data sequences are selected from the multiple initial training data sequences.

[0056] The backpack algorithm in the embodiments of this specification is a technology for solving combinatorial optimization problems. The server can use the backpack algorithm to package several initial training data sequences into a sequence set. It is understandable that in order to avoid the waste of computing power, the 0 / 1 backpack algorithm can be used for packaging, so that each initial training data sequence can only be selected once, so that the initial training data sequences in the packaged sequence set are all different, reducing the repeatability of the training data input to the model, increasing the diversity of the training data, and thus improving the robustness of the trained model. In practical applications, other backpack algorithms can also be used to process multiple initial training data sequences, such as a complete backpack algorithm, a multiple backpack algorithm, or a grouped backpack algorithm.

[0057] In the embodiments of this specification, each initial training data sequence can be marked to obtain initial sequence identification information representing each initial training data sequence, and then a backpack algorithm can be performed based on the initial sequence identification information to obtain a packing strategy containing the initial sequence identification information, so that the corresponding initial training data sequence can be selected from the initial training data sequence based on the initial sequence identification information in the packing strategy. The initial sequence identification information contained in the packing strategy obtained by the backpack algorithm can be one or more, and accordingly, the initial training data sequence selected based on the packing strategy can be one or more. This allows the sum of the lengths of the several initial training data sequences selected after packaging to be less than or equal to the maximum sequence length, while also reducing the number and length of invalid token padding.

[0058] In practical applications, a knapsack algorithm can be used to select several initial training data sequences from multiple initial training data sequences. Then, the knapsack algorithm can be used to select several initial training data sequences from the remaining initial training data sequences until all the multiple initial training data sequences generated based on the initial training data are packaged, resulting in multiple different sequence sets. The initial training data sequences in one sequence set can be completely different from those in another sequence set; one sequence set can contain some of the same initial training data sequences as another sequence set. Alternatively, the knapsack algorithm can be used to package all the multiple initial training data sequences generated based on the initial training data, resulting in multiple different sets. The specific packaging method can be set based on the performance of the hardware device and user needs and is not specifically limited here.

[0059] To improve the accuracy of the model, in the embodiments of this specification, the initial training data sequence can be split into two subsequences so that the model can learn separately and reduce the interference caused by multiple answers to the same question. Optionally, as an implementation method, each piece of initial training data described in the embodiments of this specification also includes question information; the serialization of each piece of initial training data to obtain multiple initial training data sequences can specifically include: For any piece of initial training data, perform sequence processing on the piece of initial training data to obtain a first subsequence and a second subsequence corresponding to the piece of initial training data; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The first subsequence and the second subsequence are concatenated to obtain the initial training data sequence.

[0060] In the embodiments of this specification, the question information in the initial training data can be used as prompt word information for the input model to prompt the model to answer the question.

[0061] As an implementation method, the initial training data can be split to obtain first sub-initial training data containing question information and first answer information, and second sub-initial training data containing question information and second answer information; the first sub-initial training data and the second sub-initial training data are serialized respectively to obtain a first subsequence and a second subsequence.

[0062] As another implementation, the initial training data may be serialized to obtain a total sequence containing question information, first answer information, and second answer information; the total sequence may be separated to obtain a first subsequence of the sequence containing question information and first answer information, and a second subsequence of the sequence containing question information and second answer information.

[0063] As another implementation, the question information, the first answer information, and the second answer information can be separated to obtain the corresponding question sequence, the first answer sequence, and the second answer sequence. The question sequence can be spliced ​​with the first answer sequence and the second answer sequence respectively to obtain the first subsequence and the second subsequence.

[0064] In the embodiments of this specification, the initial training data sequence including the first subsequence and the second subsequence can be used as a sequence for model training, so that each answer has a corresponding question, which can improve the accuracy of model training.

[0065] In actual applications, the server can also serialize the initial training data to obtain a total sequence containing question information, first answer information and second answer information. The total sequence can be used as the initial training data sequence without separation, so that the initial training data sequence contains a sequence of question information, first answer information and second answer information, so that when the model processes the same initial training data, it processes less data and reduces costs.

[0066] Optionally, in the embodiments of this specification, the concatenation of the first subsequence and the second subsequence to obtain the initial training data sequence may specifically include: splicing the head of the second subsequence to the tail of the first subsequence to obtain the initial training data sequence; Alternatively, the head of the first subsequence is spliced ​​to the tail of the second subsequence to obtain the initial training data sequence.

[0067] In the embodiments of this specification, the order of the first subsequence can be that the sequence of question information comes first and the sequence of first answer information comes second; or the sequence of first answer information comes first and the sequence of question information comes second. The splicing order of the sequence of question information and the sequence of first answer information in the second subsequence can be the same as the splicing order of the first subsequence, or it can be different from the splicing order of the first subsequence. For example, the initial training data contains question 1, the Chosen answer corresponding to question 1, and Reject1 corresponding to question 1. The first subsequence can be the sequence of question 1+Chosen1, and the second subsequence can be the sequence of Reject1+question 1; or the first subsequence is the sequence of question 1+Chosen1, and the second subsequence is the sequence of question 1+Reject1.

[0068] To clearly illustrate the initial training data sequence, continuing with the above example, the first subsequence is the sequence of Question 1 + Chosen 1, and the second subsequence is the sequence of Question 1 + Reject 1. The initial training data can be a concatenation of the sequence of Question 1 + Chosen 1 + Question 1 + Reject 1; or a concatenation of the sequence of Question 1 + Reject 1 + Question 1 + Chosen 1. This avoids splicing a subsequence between the question and answer sequences of another subsequence, thereby improving the accuracy of model recognition and the accuracy of model training.

[0069] In the embodiments of this specification, several selected training data sequences can also be spliced ​​together. Specifically, the head of an initial training data sequence can be spliced ​​with the tail of another initial training data sequence, or a subsequence in an initial training data sequence can be spliced ​​between two subsequences in another initial training data sequence. For example, the selected initial training data sequences include sequence 1: the sequence of question 1 + chosen 1 + question 1 + reject 1; sequence 2: the sequence of question 2 + chosen 2 + question 2 + reject 2. The spliced ​​sequence can be the sequence of question 1 + chosen 1 + question 1 + reject 1 + question 2 + chosen 2 + question 2 and reject 2, or the sequence of question 1 + chosen 1 + question 2 and chosen 1 + question 1 + reject 2 + question 2 and reject 2.

[0070] In practical applications, the splicing of initial training data sequences can be orderly splicing, such as splicing one initial training data sequence as a whole with another initial training data sequence, or splicing together the sequences containing the first answer information in each initial training data sequence, splicing together the sequences containing the second answer information, and then splicing the two sequences together; it can also be random and disordered splicing, which can be performed according to actual needs or settings. There is no specific limitation on the splicing method of the initial training data sequences here.

[0071] To further improve the accuracy of model training, optionally, the method described in the embodiments of this specification may further include: For any initial training data sequence among the initial training data sequences included in the spliced ​​data sequence, generating a sequence position identifier corresponding to the any initial training data sequence according to the length of the any initial training data sequence; The sequence position identifiers corresponding to the initial training data sequences contained in the spliced ​​data sequence are arranged according to the order of the initial training data sequences in the spliced ​​data sequence to obtain sequence position identifiers corresponding to the spliced ​​data sequence. This allows the model to determine the starting positions of different initial training data sequences based on the sequence position identifiers, allowing the model to know the positions of each sample contained in a spliced ​​data sequence and avoid identifying a spliced ​​data sequence as a single sample.

[0072] In the embodiments of this specification, the sequence position identifier of the initial training data sequence is used to identify the position information of each token sequence in the initial training data sequence.

[0073] In practical applications, in the process of serializing the initial training data, a sequence position identifier corresponding to the initial training data sequence can be generated to obtain an initial training data sequence containing a sequence position identifier corresponding to the initial training data sequence. Then, in the process of splicing a number of selected initial training data sequences, the sequence position identifier corresponding to the initial training data sequence can be spliced ​​together so that the spliced ​​initial training data sequence and the sequence position identifier are in corresponding positions, so that the model can identify the various training sample information contained in the spliced ​​data sequence. The server can also generate a sequence position identifier corresponding to the initial training data sequence after generating the initial training data sequence, and can also establish a corresponding relationship between the sequence position identifier and the initial training data sequence, so that after splicing the initial training data sequence, the sequence position identifier can be spliced ​​in the splicing order to improve the accuracy of model training.

[0074] In practical applications, it is also possible to generate a sequence position identifier corresponding to the invalid sequence based on the length of the invalid sequence used for padding contained in the padded sequence, so that the sequence position identifiers of each initial training data sequence and the sequence position identifiers of the invalid sequence can be spliced ​​based on the splicing order of each sequence in the padded sequence to obtain a sequence position identifier corresponding to the padded sequence.

[0075] Optionally, as an implementation manner, each piece of initial training data described in the embodiments of this specification further includes question information; the initial training data sequence includes a first subsequence and a second subsequence; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; Generating a sequence position identifier corresponding to any initial training data sequence according to the length of any initial training data sequence may specifically include: generating, according to the length of the first subsequence included in any one of the initial training data sequences, a first subsequence position identifier corresponding to the first subsequence; generating a second subsequence position identifier corresponding to the second subsequence according to a length of the second subsequence included in any one of the initial training data sequences; The first subsequence position identifier and the second subsequence position identifier are concatenated to obtain a sequence position identifier corresponding to any initial training data sequence.

[0076] In the embodiments of this specification, one subsequence corresponds to one subsequence position identifier. The subsequence position identifiers may be concatenated in the order in which the subsequences are concatenated into the initial training data sequence to obtain the sequence position identifier corresponding to the initial training data sequence.

[0077] In practical applications, if the first subsequence is a sequence of questions and Chosen answers, and the second subsequence is a sequence of questions and Reject answers, then the first subsequence position identifier is used to identify the sequence of questions and Chosen answers, and the second subsequence position identifier is used to identify the sequence of questions and Reject answers. In this way, different subsequence position identifiers can be used to identify different subsequences, which makes it easier for the model to distinguish the first subsequence and the second subsequence in different initial training data sequences, thereby improving the accuracy of model training.

[0078] As an implementation manner, optionally, the sequence position identifier in the embodiment of this specification is a character string; the first character of the sequence position identifier corresponding to each of the initial training data sequences is the same.

[0079] In the embodiments of this specification, the first character of the sequence position identifier corresponding to each initial training data sequence can be set based on the model architecture and user needs, for example, to a value such as 0. The first character of the first sequence position identifier corresponding to the first subsequence in the initial training data sequence is the same as the first character of the second subsequence position identifier corresponding to the second subsequence. This allows the position of each initial training data sequence to be determined based on the first character of the sequence position identifier, thereby improving the accuracy of the model in processing training data.

[0080] As an implementation manner, in the embodiments of this specification, generating a sequence position identifier corresponding to any initial training data sequence according to the length of any initial training data sequence may specifically include: Based on the preset first character, the position identifier corresponding to each unit length sequence in any initial training data sequence is determined according to a preset increasing rule to obtain a sequence position identifier corresponding to any initial training data sequence.

[0081] In the embodiments of this specification, one token corresponds to one position number, and a sequence position identifier corresponding to an initial training data sequence is composed of the position numbers corresponding to the various tokens that make up the initial training data sequence. The preset increment rule can represent increasing the value of the preset step length based on the first character according to the preset step length value to obtain the position number corresponding to the second token, and then increasing the value of the preset step length again to obtain the position number corresponding to the third token, and so on, to obtain the position numbers corresponding to the various tokens in the initial training data sequence, and then obtain the sequence position identifier corresponding to the initial training data sequence. For example, if an initial training data sequence contains 5 tokens, the first character is 0, and the preset step length value is 1, then the sequence position identifier corresponding to the initial training data sequence can be obtained as

[01234] .

[0082] Optionally, as an implementation manner, the method described in the embodiments of this specification may further include: Based on the length information of each of the initial training data sequences contained in the spliced ​​data sequence and the position information of each of the initial training data sequences in the spliced ​​data sequence, an attention mask corresponding to the spliced ​​data sequence is generated; the attention mask value between different initial training data sequences in the attention mask is a mask value representing a mask identifier.

[0083] In the embodiments of this specification, an attention mask can be generated based on an attention mechanism and can be used to indicate the model's level of attention to the input initial training data sequence. This can also isolate the visibility between subsequences or initial training data sequences, reducing the probability of mutual influence between the initial training data sequences in the spliced ​​data sequence. The attention mask can use two different mask values ​​to identify the parts that require or do not require the model's attention, preventing the model from focusing too much on irrelevant parts. This can reduce the consumption of model resources such as computing power and memory, thereby reducing model training costs.

[0084] In practical applications, a concatenated data sequence can be generated first, and then a corresponding attention mask can be generated based on the generated concatenated data sequence. After the concatenated data sequence is input into the model, the model can determine the part of the initial training data sequence that needs attention based on the attention mask, thereby reducing the attention to invalid tokens and non-initial training data sequences, further reducing resource consumption.

[0085] Optionally, as an implementation manner, the attention mask in the embodiment of this specification is an NxN matrix; wherein N is the length of the spliced ​​data sequence; the spliced ​​data sequence includes a plurality of the initial training data sequences, and one of the initial training data sequences includes a first subsequence and a second subsequence, the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The mask value for the i-th row and j-th column in the matrix satisfies the following conditions:

[0086] in, represents the mask value of the i-th row and j-th column in the matrix; is the mask value representing the non-masked identifier; is the mask value representing the mask identifier; represents the first total length of the first k subsequences contained in the concatenated data sequence; represents the second total length of the first k-1 subsequences contained in the concatenated data sequence, , r represents the total number of subsequences contained in the initial training data series in the spliced ​​data sequence; Indicates the length of the first subsequence contained in the concatenated data sequence.

[0087] In the embodiment of this specification, if the length of the spliced ​​data sequence is less than the maximum sequence length and is filled with invalid tokens, the mask value of the attention mask of the filled invalid token part can be set to , so that the model does not pay attention to the parts filled with invalid tokens. and The value of can be set based on expert experience or according to user needs. For example, Set to 0, Setting it to negative infinity makes the model not pay attention to the areas marked with negative infinity, nor does it process the areas marked with negative infinity, which reduces resource consumption and avoids the impact of invalid areas on model training results.

[0088] As another implementation, sub-mask blocks can be generated based on each sub-sequence, and then the sub-mask blocks are spliced ​​according to the splicing order of the spliced ​​data sequence. In the embodiment of this specification, the attention mask corresponding to the spliced ​​data sequence is generated based on the length information of each of the initial training data sequences contained in the spliced ​​data sequence and the position information of each of the initial training data sequences in the spliced ​​data sequence, which can specifically include: For any subsequence, according to the length of any subsequence, generate a sub-attention mask corresponding to the subsequence; the sub-attention mask is a matrix of axa, where a is the length of any subsequence, in, It can represent the mask value of the i-th row and j-th column in the matrix, is the mask value representing the non-masked identifier; is the mask value representing the mask identifier; According to the order of each subsequence in the spliced ​​data sequence, the sub-attention masks corresponding to each subsequence are spliced ​​in the diagonal direction to obtain the attention mask corresponding to the spliced ​​data sequence; the mask values ​​at other positions outside the sub-attention masks in the attention mask corresponding to the spliced ​​data sequence are mask values ​​representing the mask identifier.

[0089] In the embodiment of this specification, the subsequences can be spliced ​​along the diagonal line i=j according to the order of each subsequence in the spliced ​​data sequence to obtain the attention mask corresponding to the spliced ​​data sequence. The value area is set to value.

[0090] In order to more clearly illustrate the attention mask of the concatenated data sequence, Figure 4 This is a schematic diagram of an attention mask provided in the embodiment of this specification. Figure 4As shown, the area indicated by 41 is the spliced ​​data sequence, wherein the first chosen answer and the first rejected answer can represent an initial training data sequence, and the corresponding second chosen answer and the second rejected answer can represent another initial training data sequence. It can be understood that both the chosen answer and the rejected answer contain sequences corresponding to the question information, but they are not shown in the figure; 42 can represent a sequence position identifier corresponding to the spliced ​​data sequence; 43 can represent an attention mask corresponding to the spliced ​​data sequence, wherein the white area represents the masked area, and the dark area represents the non-masked area.

[0091] Optionally, as an implementation manner, the method described in the embodiments of this specification may further include: The sequence position identifier and the spliced ​​data sequence are provided to the model to be trained.

[0092] In the embodiments of this specification, the model to be trained can be a model that has been fine-tuned with a large amount of data, or a model that only has a model framework, or a model in use that needs to be updated, etc.

[0093] In the embodiment of this specification, the sequence position identifier corresponding to the padded sequence and the padded sequence may also be provided to the model to be trained.

[0094] In practical applications, the model to be trained can be trained multiple times, and the sequence position identifiers and the concatenated data sequences can be divided into multiple batches and input into the model to be trained. The number of initial training data sequences contained in a batch can be determined based on the data processing capability of the model. The concatenated data sequences or padded data sequences contained in a batch can be partially identical or completely different. For example, if 2000 concatenated data sequences and padded sequences are generated in the above manner, and it is assumed that a batch can contain 400 data items, up to 400 sequences can be selected from the 2000 to form a batch and input into the model to be trained.

[0095] Optionally, as an implementation manner, the method described in the embodiments of this specification may further include: The attention mask corresponding to the spliced ​​data sequence and the spliced ​​data sequence are provided to the model to be trained.

[0096] In the embodiment of this specification, the attention mask corresponding to the padded sequence and the padded sequence can also be provided to the model to be trained.

[0097] In practical applications, it is also possible to select training samples that can be input into the model to be trained from the attention masks corresponding to the multiple spliced ​​data sequences and the multiple spliced ​​data sequences based on the preset rounds of training and the number of sequences allowed to be input in each round or batch of the model, and input these training samples into the model to be trained to train the model to be trained.

[0098] In practical applications, the sequence position identifier, the attention mask corresponding to the concatenated data sequence, and the concatenated data sequence can also be provided to the model to be trained so that the model to be trained can be trained based on the input data. Alternatively, if padding is required for the concatenated sequence, the sequence position identifier corresponding to the padded sequence, the attention mask corresponding to the padded sequence, and the padded sequence can also be provided to the model to be trained.

[0099] In the embodiment of this specification, the model to be trained can also be trained using the DPO direct preference optimization method.

[0100] In the embodiments of this specification, after the training model is trained using the DPO direct preference optimization method, a trained model can be obtained. The trained model can also be evaluated based on a preset evaluation method to determine whether the trained model meets preset requirements, thereby determining whether to continue training the trained model. Specifically, the trained model can be evaluated based on the validation samples in the validation set to determine whether the loss value of the trained model is less than or equal to the preset loss value. If it is determined that the loss value of the trained model is less than or equal to the preset loss value, the training is completed and a trained model is obtained.

[0101] In practical applications, the reward model can also be trained using the above-mentioned sequence position identifiers, the attention mask corresponding to the spliced ​​data sequence, and the spliced ​​data sequence to improve the performance and accuracy of the reward model. The reward model can also be used to evaluate the trained model to determine the pros and cons of the answers generated by the trained model relative to human preferences. Specifically, the reward model can be used to evaluate the answers output by the trained model to obtain an evaluation score. If the evaluation score is greater than or equal to a preset score, a trained model is obtained. The reward model can assign reward points to preferred answers output by the trained model and penalty points to non-preferred answers output by the trained model. By calculating the reward and penalty scores, an evaluation score can be obtained to evaluate the pros and cons of the trained model.

[0102] Through the above method, on the first hand, the present application splices the sequences corresponding to the question + Chosen answer and the question + Reject answer of each initial training data in the initial training data set, and packages and splices several initial training data sequences, thereby significantly reducing the filling of invalid tokens, improving the utilization rate of GPU video memory, and evolving the data input process, reducing the number of model training times and training time, and by reducing data fragmentation, optimizing the overall structure of the data, providing a more coherent training input for the model.

[0103] Secondly, the present application can calculate the packing strategy through the backpack algorithm to splice several initial training data sequences in the initial training data set into a continuous long sequence, which not only enables the model to process several initial training data at the same time in a single forward reasoning process, reducing the overall training time, but also can further reduce the use of invalid tokens and improve the utilization of video memory. At the same time, it can also speed up the training speed and improve the model's ability to process data, making the training process more efficient.

[0104] Thirdly, this application also uses a customized attention mask, which can cleverly isolate the visibility between each subsequence, effectively avoiding the problem of cross-attention, and thus can precisely control the attention allocation of the model to ensure that the model can focus on relevant data during training, thereby improving the training effect and final performance of the model. It can also provide the model with a clearer data view, optimize the model's learning process, and enable the model to more accurately understand and respond to training data.

[0105] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 5 The embodiments of this specification provide corresponding Figure 3 A schematic diagram of the structure of a device for preprocessing training data. Figure 5 As shown, the device may include: An initial training data acquisition module 502 is configured to acquire a plurality of pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; A serialization processing module 504 is used to serialize each piece of initial training data to obtain multiple initial training data sequences; A sequence selection module 506 is configured to select a plurality of initial training data sequences from the plurality of initial training data sequences based on a maximum sequence length recognizable by the model to be trained; The splicing module 508 is configured to splice the selected initial training data sequences to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0106] based on Figure 5 The present specification also provides some specific implementation plans of the method, which are described below.

[0107] Optionally, if the length of the concatenated data sequence is less than the maximum sequence length, the apparatus further includes a padding module, which may be specifically configured to: The concatenated data sequence is padded to obtain a padded sequence; the length of the padded sequence is equal to the maximum sequence length.

[0108] Optionally, the sequence selection module may be specifically used to: Determining a packing strategy using a knapsack algorithm based on the sequence length of each of the initial training data sequences and the maximum sequence length; According to the packaging strategy, several initial training data sequences are selected from the multiple initial training data sequences.

[0109] Optionally, each piece of initial training data further includes question information; the sequence selection module may be specifically configured to: For any piece of initial training data, perform sequence processing on the piece of initial training data to obtain a first subsequence and a second subsequence corresponding to the piece of initial training data; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The first subsequence and the second subsequence are concatenated to obtain the initial training data sequence.

[0110] Optionally, the sequence selection module may also be used to: splicing the head of the second subsequence to the tail of the first subsequence to obtain the initial training data sequence; Alternatively, the head of the first subsequence is spliced ​​to the tail of the second subsequence to obtain the initial training data sequence.

[0111] Optionally, the apparatus may further include a sequence position identifier generation module, which may be specifically configured to: For any initial training data sequence among the initial training data sequences included in the spliced ​​data sequence, generating a sequence position identifier corresponding to the any initial training data sequence according to the length of the any initial training data sequence; The sequence position identifiers corresponding to the initial training data sequences contained in the spliced ​​data sequence are arranged according to the order of the initial training data sequences in the spliced ​​data sequence to obtain the sequence position identifiers corresponding to the spliced ​​data sequence.

[0112] Optionally, each piece of initial training data further includes question information; the initial training data sequence includes a first subsequence and a second subsequence; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The sequence position identifier generation module can be specifically used to: generating, according to the length of the first subsequence included in any one of the initial training data sequences, a first subsequence position identifier corresponding to the first subsequence; generating a second subsequence position identifier corresponding to the second subsequence according to a length of the second subsequence included in any one of the initial training data sequences; The first subsequence position identifier and the second subsequence position identifier are concatenated to obtain a sequence position identifier corresponding to any initial training data sequence.

[0113] Optionally, the sequence position identifier is a character string; the first character of the sequence position identifier corresponding to each of the initial training data sequences is the same.

[0114] Optionally, the apparatus may further include an attention mask generation module, which may be specifically configured to: Based on the length information of each of the initial training data sequences contained in the spliced ​​data sequence and the position information of each of the initial training data sequences in the spliced ​​data sequence, an attention mask corresponding to the spliced ​​data sequence is generated; the attention mask value between different initial training data sequences in the attention mask is a mask value representing a mask identifier.

[0115] Optionally, the attention mask is an NxN matrix; wherein N is the length of the concatenated data sequence; the concatenated data sequence includes a plurality of the initial training data sequences, wherein one of the initial training data sequences includes a first subsequence and a second subsequence, wherein the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The mask value for the i-th row and j-th column in the matrix satisfies the following conditions:

[0116] in, represents the mask value of the i-th row and j-th column in the matrix; is the mask value representing the non-masked identifier; is the mask value representing the mask identifier; represents the first total length of the first k subsequences contained in the concatenated data sequence; represents the second total length of the first k-1 subsequences contained in the concatenated data sequence, , r represents the total number of subsequences contained in the initial training data series in the spliced ​​data sequence; Indicates the length of the first subsequence contained in the concatenated data sequence.

[0117] Optionally, the device can also be used for: The sequence position identifier and the spliced ​​data sequence are provided to the model to be trained.

[0118] Optionally, the device can also be used for: The attention mask corresponding to the spliced ​​data sequence and the spliced ​​data sequence are provided to the model to be trained.

[0119] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.

[0120] Figure 6 The embodiments of this specification provide corresponding Figure 3 A schematic diagram of the structure of a device for preprocessing training data. Figure 6 As shown, the device 600 may include: at least one processor 610; and, A memory 630 in communication with the at least one processor; wherein, The memory 630 stores instructions 620 that can be executed by the at least one processor 610. The instructions are executed by the at least one processor 610 to enable the at least one processor 610 to: Acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; Serialize each piece of initial training data to obtain multiple initial training data sequences; Selecting a plurality of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the to-be-trained model; Several selected initial training data sequences are spliced ​​together to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

[0121] Based on the same idea, the embodiments of this specification also provide a computer-readable medium corresponding to the above method. The computer-readable medium stores computer-readable instructions, which can be executed by a processor to implement the above method of preprocessing training data: The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. Figure 6 As for the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0122] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0123] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0124] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.

[0125] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0126] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0127] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0129] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0131] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0132] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0133] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0134] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0135] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0136] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0137] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for preprocessing training data, comprising: Obtain multiple pieces of initial training data; Each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; Serialize each piece of initial training data to obtain multiple initial training data sequences; Selecting a plurality of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the to-be-trained model; Several selected initial training data sequences are spliced ​​together to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

2. The method according to claim 1, further comprising: if the length of the concatenated data sequence is less than the maximum sequence length; Performing padding processing on the spliced ​​data sequence to obtain a padded sequence; The length of the padded sequence is equal to the maximum sequence length.

3. The method according to claim 1, wherein the selecting of a plurality of initial training data sequences from the plurality of initial training data sequences based on the maximum sequence length recognizable by the model to be trained comprises: Determining a packing strategy using a knapsack algorithm based on the sequence length of each of the initial training data sequences and the maximum sequence length; According to the packaging strategy, several initial training data sequences are selected from the multiple initial training data sequences.

4. The method according to claim 1, wherein each piece of initial training data further includes question information; and wherein the serializing each piece of initial training data to obtain multiple initial training data sequences specifically comprises: For any piece of initial training data, perform sequence processing on the piece of initial training data to obtain a first subsequence and a second subsequence corresponding to the piece of initial training data; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The first subsequence and the second subsequence are concatenated to obtain the initial training data sequence.

5. The method according to claim 4, wherein the step of concatenating the first subsequence and the second subsequence to obtain the initial training data sequence comprises: splicing the head of the second subsequence to the tail of the first subsequence to obtain the initial training data sequence; Alternatively, the head of the first subsequence is spliced ​​to the tail of the second subsequence to obtain the initial training data sequence.

6. The method according to claim 1, further comprising: For any initial training data sequence among the initial training data sequences included in the spliced ​​data sequence, generating a sequence position identifier corresponding to the any initial training data sequence according to the length of the any initial training data sequence; The sequence position identifiers corresponding to the initial training data sequences contained in the spliced ​​data sequence are arranged according to the order of the initial training data sequences in the spliced ​​data sequence to obtain the sequence position identifiers corresponding to the spliced ​​data sequence.

7. The method according to claim 6, wherein each piece of initial training data further includes question information; the initial training data sequence includes a first subsequence and a second subsequence; the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; Generating a sequence position identifier corresponding to any initial training data sequence according to the length of any initial training data sequence specifically includes: generating, according to the length of the first subsequence included in any one of the initial training data sequences, a first subsequence position identifier corresponding to the first subsequence; generating a second subsequence position identifier corresponding to the second subsequence according to a length of the second subsequence included in any one of the initial training data sequences; The first subsequence position identifier and the second subsequence position identifier are concatenated to obtain a sequence position identifier corresponding to any initial training data sequence.

8. The method according to claim 6, wherein the sequence position identifier is a character string; the first character of the sequence position identifier corresponding to each of the initial training data sequences is the same.

9. The method according to claim 1, further comprising: Generate an attention mask corresponding to the spliced ​​data sequence according to the length information of each of the initial training data sequences contained in the spliced ​​data sequence and the position information of each of the initial training data sequences in the spliced ​​data sequence; The attention mask value between different initial training data sequences in the attention mask is a mask value representing a mask identifier.

10. The method according to claim 9, wherein the attention mask is an NxN matrix; wherein N is the length of the concatenated data sequence; the concatenated data sequence includes a plurality of the initial training data sequences, wherein each initial training data sequence includes a first subsequence and a second subsequence, wherein the first subsequence is a sequence including the question information and the first answer information, and the second subsequence is a sequence including the question information and the second answer information; The mask value for the i-th row and j-th column in the matrix satisfies the following conditions: in, represents the mask value of the i-th row and j-th column in the matrix; is the mask value representing the non-masked identifier; is the mask value representing the mask identifier; represents the first total length of the first k subsequences contained in the concatenated data sequence; represents the second total length of the first k-1 subsequences contained in the concatenated data sequence, , r represents the total number of subsequences contained in the initial training data series in the spliced ​​data sequence; Indicates the length of the first subsequence contained in the concatenated data sequence.

11. The method according to claim 6, further comprising: The sequence position identifier and the spliced ​​data sequence are provided to the model to be trained.

12. The method according to claim 9, further comprising: The attention mask corresponding to the spliced ​​data sequence and the spliced ​​data sequence are provided to the model to be trained.

13. A device for preprocessing training data, comprising: An initial training data acquisition module is used to acquire multiple pieces of initial training data; Each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; A serialization processing module is used to serialize each piece of initial training data to obtain multiple initial training data sequences; A sequence selection module, configured to select a number of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the model to be trained; The splicing module is used to splice the selected initial training data sequences to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

14. A device for preprocessing training data, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire multiple pieces of initial training data; each piece of initial training data includes first answer information indicating a preferred answer and second answer information indicating a non-preferred answer; Serialize each piece of initial training data to obtain multiple initial training data sequences; Selecting a plurality of the initial training data sequences from the plurality of initial training data sequences according to a maximum sequence length recognizable by the to-be-trained model; Several selected initial training data sequences are spliced ​​together to obtain a spliced ​​data sequence; the length of the spliced ​​data sequence is less than or equal to the maximum sequence length.

15. A computer-readable medium having computer-readable instructions stored thereon, wherein the computer-readable instructions can be executed by a processor to implement the method for pre-processing training data according to any one of claims 1 to 12.