Method and device for determining training sample data, equipment and medium

By determining the confusion score of the training sample data and conducting ablation experiment, selecting data sets with excellent evaluation index values to cooperate as the target training sample, the problem of insufficient inference ability of large models in the existing technology is solved, and more efficient screening of training sample data is achieved, and the model's inference ability is improved.

CN120354134APending Publication Date: 2025-07-22ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510496173.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing screening of training sample data through text mining cannot effectively improve the inference ability of large models.

Method used

The first large model determines the confusion score of the training sample data, divides it into different data sets, and uses these sets to perform ablation experiments on the second large model. The data set with the evaluation index value is better than other sets are selected as the target training sample data, and is used to train the third large model.

Benefits of technology

The inference ability of the large model is improved, and the performance of the model is enhanced by selecting training sample data that is more in line with the model's learning needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354134A_ABST
    Figure CN120354134A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method and device for determining training sample data, equipment and a medium. The scheme can comprise the following steps: acquiring a plurality of pieces of training sample data; on the basis of the first large model, determining confusion degree scores of the plurality of pieces of training sample data; dividing the plurality of pieces of training sample data into at least two data sets based on the confusion score; the confusion score of the training sample data contained in the first data set in the two data sets is greater than the confusion score of the training sample data contained in the second data set; performing an ablation experiment on the second large model by utilizing each data set, and determining an evaluation index value of the second large model in each data set; determining training sample data contained in a target data set of which the evaluation index value is superior to the evaluation index values of other data sets in each data set as target training sample data for training a third model; the magnitude of the third large model is greater than that of the first large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment and medium for determining training sample data. Background Art

[0002] With the continuous development of science and technology, more and more big models are being applied to actual business, such as intelligent customer service, intelligent question-answering systems, etc. During or before the use of the big model, the quality of the text input into the big model can be evaluated through text mining, such as fasttext or the open source classification model Fineweb-edu classifier Bert, etc., low-quality texts can be filtered, and high-quality texts can be input into the big model so that the big model can learn. However, text mining can only filter the quality of the text to train the big model, which is to filter the training data from the perspective of the quality of the text itself. For the big model, the text that may be filtered out cannot well improve the ability of the model.

[0003] Therefore, how to provide a training sample that can more effectively improve the reasoning ability of large models is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of this specification provide a method, apparatus, device and medium for determining training sample data to solve the problem that the existing determination of training sample data cannot well improve the reasoning ability of large models.

[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows.

[0006] A method for determining training sample data provided by an embodiment of the present specification includes: obtaining a plurality of training sample data;

[0007] Based on the first largest model, the perplexity scores of the plurality of training sample data are determined; based on the perplexity scores, the plurality of training sample data are divided into at least two data sets; the perplexity score of the training sample data contained in the first data set of the two data sets is greater than the perplexity score of the training sample data contained in the second data set; ablation experiments are performed on the second largest model using each data set to determine the evaluation index value of the second largest model in each data set; the training sample data contained in the target data set whose evaluation index value in each data set is better than the evaluation index value of other data sets is determined as the target training sample data for training the third largest model; the magnitude of the third largest model is greater than the magnitude of the first largest model.

[0008] An apparatus for determining training sample data provided by an embodiment of this specification includes: a sample acquisition module configured to acquire a plurality of pieces of training sample data; a score determination module configured to determine the perplexity scores of the plurality of pieces of training sample data based on a first large model; a set partitioning module configured to partition the plurality of pieces of training sample data into at least two data sets based on the perplexity scores; the perplexity scores of the training sample data included in the first data set among the two data sets are greater than those of the training sample data included in the second data set; an index value determination module configured to perform ablation experiments on a second large model using each data set to determine the evaluation index values of the second large model in each data set; a target sample determination module configured to determine the training sample data included in a target data set whose evaluation index value is better than those of other data sets as the target training sample data for training a third large model; the scale of the third large model is greater than that of the first large model.

[0009] An apparatus for determining training sample data provided by an embodiment of this specification includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: acquire a plurality of pieces of training sample data; determine the perplexity scores of the plurality of pieces of training sample data based on a first large model; partition the plurality of pieces of training sample data into at least two data sets based on the perplexity scores; the perplexity scores of the training sample data included in the first data set among the two data sets are greater than those of the training sample data included in the second data set; perform ablation experiments on a second large model using each data set to determine the evaluation index values of the second large model in each data set; determine the training sample data included in a target data set whose evaluation index value is better than those of other data sets as the target training sample data for training a third large model; the scale of the third large model is greater than that of the first large model.

[0010] A computer-readable medium provided by an embodiment of this specification, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement a method for determining training sample data.

[0011] At least one embodiment of this specification can achieve the following beneficial effects: The perplexity scores of several pieces of training sample data obtained can be determined through a first large model. Based on the perplexity scores, the several pieces of training sample data are divided into at least two data sets. The ablation experiment of a second large model is carried out using each data set, and the evaluation index values of the second large model in each data set are determined. The data set with evaluation index values superior to other data sets is used as the target data set, and the training samples in the target data set are determined as the target training sample data for training a third large model. Among them, the target training sample data for model training can be determined based on the perplexity scores and the evaluation indexes of the large model under each training data, so that the target training sample data that can better improve the inference ability of the third large model can be determined by combining the capabilities of the large model. In practical applications, the third large model can be trained with the target training sample data to improve the inference ability of the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 It is a schematic diagram of an application scenario of a method for determining training sample data provided by an embodiment of this specification;

[0014] Figure 2 It is a schematic flowchart of a method for determining training sample data provided by an embodiment of this specification;

[0015] Figure 3 It is a schematic diagram of training sample data and corresponding evaluation index values provided by an embodiment of this specification;

[0016] Figure 4 It is a schematic diagram of training sample data and corresponding evaluation index values provided by an embodiment of this specification;

[0017] Figure 5 It is a schematic diagram of the overall architecture of a method for determining training sample data provided by an embodiment of this specification;

[0018] Figure 6 It is a schematic diagram of the structure of a device for determining training sample data provided by an embodiment of this specification;

[0019] Figure 7 It is a schematic diagram of the structure of a device for determining training sample data provided by an embodiment of this specification. Specific Embodiments

[0020] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by one or more embodiments of this specification.

[0021] The terms used in one or more embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit one or more embodiments of this application. The singular forms "a", "the", and "said" used in one or more embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this application refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0023] To clearly illustrate the implementation manners of the embodiments of this specification, the following explanations are given for some terms.

[0024] Large Model: Refers to a machine learning model with a huge number of parameters and extremely strong computing power. It is usually based on deep learning technology, can process massive amounts of data, and learn complex patterns from it. It can perform tasks such as natural language processing, image generation, and scientific computing, and can also handle some complex problems through its reasoning ability, and is widely used in multiple fields. The large model may include at least one of models such as the GPT series, BERT, PaLM, and Stable Diffusion.

[0025] Base pre-training: It refers to the initial large-scale unsupervised learning process of a language model. In this process, the model self-learns the language structure and patterns from a large amount of text data without being optimized for specific tasks. Through base pre-training, the model can learn extensive language knowledge, including grammar, semantics, and to some extent, common sense reasoning ability. After the model completes base pre-training, it can be further fine-tuned to adapt to specific downstream tasks, such as text classification, question answering systems, etc.

[0026] Perplexity (PPL): It is an important metric for evaluating the performance of a language model, which can reflect the model's ability to predict text sequences. The lower the perplexity, the better the model's prediction effect on the given dataset, that is, the model can capture the probability distribution of the language more accurately. Perplexity can be understood as a measure of the average uncertainty of the model for each word; it is calculated based on probability and is usually used to evaluate the effectiveness of a language model when generating or predicting the next word.

[0027] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the accompanying drawings. To address the deficiencies in the prior art, the following embodiments are provided:

[0028] Figure 1 It is a schematic diagram of the application scenario of a method for determining training sample data provided by an embodiment of this specification. As Figure 1 shown, the solution may include several training sample data 1, a cloud server 2, and target training sample data 3. The cloud server 2 may contain multiple server nodes or may deploy multiple large models. The cloud server 2 can obtain several training sample data, determine the perplexity scores of the several training sample data through a first large model, divide the several training sample data into multiple data sets based on the perplexity scores, conduct ablation experiments on a second large model using the data sets, determine the evaluation metric values of the second large model, determine the target data set based on the evaluation metric values, and the cloud server 2 can use each training sample data in the target data set as the target training sample data 3. In practical applications, the cloud server 2 can also use the target training sample data 3 to train a third large model to improve the reasoning ability of the third large model.

[0029] Among them, the first large model, the second large model, and the third large model can be the same large language model or different large language models. As an implementation, the model scale or the order of the number of parameters included in the first large model and the second large model can be smaller than the model scale or the order of the number of parameters included in the third large model. For example, the first large model can be a large model with 1B, 2B, etc. parameters, the second large model can be a large model with 5B, 10B, etc. parameters, and the third large model can be a large model with 80B, 70B parameters.

[0030] As an implementation, the first large model, the second large model, and the third large model can be general inference large models, which can handle a variety of tasks and applications without being restricted by a specific field. By training on multi-domain datasets, they learn extensive knowledge and skills, thus possessing cross-domain generalization capabilities. General inference large models can handle various tasks, such as natural language processing, computer vision, speech recognition, etc., and are applicable to different industries and scenarios. General inference large models can handle various data types, such as text, images, speech, etc., to achieve cross-modal understanding and generation. General inference large models are usually based on deep learning architectures, such as Transformer, and capture complex data patterns through multi-layer neural networks. General inference large models can be adapted to new tasks through fine-tuning or transfer learning without having to train from scratch.

[0031] In practical applications, the model scale or the order of the number of parameters included in the third large model can be larger than those of the first large model and the second large model used to screen target training samples. Using the first large model and the second large model with a smaller model scale or order of the number of parameters to screen out the target training samples for training the third large model can also improve the screening efficiency.

[0032] The model types of the first large model, the second large model, and the third large model are the same. For example, they are all large language models, or all strong base models, etc.; using the first large model and the second large model with the same model type to screen the training samples for training the third large model can make the samples more in line with the training samples that the third large model actually lacks or omits in learning, meet the training needs of the actual situation of the third large model, and comprehensively improve the inference ability of the large model by training the third large model with the samples.

[0033] In practical applications, the first large model, the second large model, and the third large model can also be models of the same scale or parameters. For example, they can all be large language models with 80B parameters, or all strong base models with 70B parameters, etc. The cloud server can be a virtual server based on cloud computing technology and can be remotely accessed and managed through the network. The cloud server pools resources such as the CPU, memory, and storage of physical servers through virtualization technology to achieve dynamic allocation and elastic scaling, and can provide services such as cloud computing and cloud storage. The cloud server can contain one or more server nodes, and the server can include, but is not limited to, any device, equipment, platform, device cluster, etc. with computing and processing capabilities. It can be connected to one or more data processing platforms through local area network connections, wide area network connections, Internet connections, or other types of data networks. In practical applications, if the amount of data to be processed is small, a conventional server can also be used for processing.

[0034] Next, a method for determining training sample data provided in the embodiments of the specification will be specifically described in conjunction with the accompanying drawings. Figure 2 It is a schematic flowchart of a method for determining training sample data provided in the embodiments of the present specification. From a program perspective, the execution subject of the process can be a program running on a server or a data processing platform. As Figure 2 shown, the process can include the following steps:

[0035] Step 202: Obtain a number of training sample data.

[0036] In the embodiments of the present specification, the number of training sample data can include one or more text forms of training sample data such as web page text, news articles, etc. The number of training sample data can include data from one or more different fields, such as multiple fields including science and technology, culture, entertainment, biology, chemical industry, computer, education, medicine, and history, or can also be data including only one field. The number of training sample data can be obtained from one or more different data sources, or can also be obtained from a database storing data from one or more different fields. The training sample data can also be data in some known training sample libraries, such as data in some open source datasets. Here, the source and specific content of the training sample data are not specifically limited.

[0037] In practical applications, several pieces of training sample data can be pre-processed training sample data. Specifically, the training sample data can be securely processed, such as removing sensitive words, etc.; then the securely processed training sample data can be tokenized to obtain the respective token sequences corresponding to the training sample data, thereby improving the security of the data and the model. Of course, if the training sample data does not have content that needs to be pre-processed, the training sample data does not have to be pre-processed either. The pre-processing step is an optional step.

[0038] To improve processing efficiency, the training sample data containing multiple tokenized words can also be organized into Lazy file shards, and then the training sample data can be sharded or grouped. For example, it can be divided into 0 - 31 encoded index indexes to obtain 32 index indexes. The sharded Lazy files can also be stored in a distributed cloud, such as a Cloud Parallel File System (CPFS), Simple Storage Service (S3), IBM Cloud Object Storage, etc., so that when using the training sample data, the pre-processed training sample data can be obtained from the cloud storage, improving the data processing efficiency and also reducing the data storage cost.

[0039] Step 204: Based on the first large model, determine the perplexity score of the several pieces of training sample data.

[0040] In the embodiments of this specification, the first large model is a large language model. Specifically, it can be a large model of the O1 series, such as models like OpenAI o1, Mureka O1, and Huatuo - o1, or it can be a large model of the GPT series, such as models like GPT - 4, GPT - 3, etc. The first large model can be a large model of magnitudes such as 1B, 2B, 3B, etc.

[0041] The perplexity score can be used to represent the accuracy of the large model's inference on the training sample data. The higher the score of the perplexity score, the less sufficient the large model's learning of the training sample data, the lower the accuracy of the large model's inference on the training sample data, and the less accurately the large model can understand this type of data. The lower the score of the perplexity score, the more sufficient the large model's learning of the training sample data, the higher the accuracy of the large model's inference on the training sample data, indicating that the large model is applicable to the inference of this type of data.

[0042] The perplexity scores of each training sample data can be obtained by calculating with the perplexity calculation formula. Specifically, for a training sample or the token sequence corresponding to this training sample, the large model can predict the (N + 1)-th token based on the first N tokens, and predict the prediction results of each token in this training sample. Based on the prediction results of each token, the PPL perplexity score is determined. For details, reference can be made to the materials related to the calculation of the PPL value, which will not be elaborated here.

[0043] Step 206: Based on the perplexity scores, divide the several pieces of training sample data into at least two data sets.

[0044] Among them, the perplexity scores of the training sample data included in the first data set in the two data sets are greater than the perplexity scores of the training sample data included in the second data set.

[0045] In the embodiments of this specification, the server can divide several pieces of training sample data into multiple data sets based on the perplexity scores through clustering algorithms, classification algorithms, sorting, etc. It can be understood that the perplexity scores of multiple pieces of training sample data belonging to the same data set can be the same or belong to a preset score interval. The perplexity scores of multiple pieces of training sample data belonging to different data sets are different, and the gap can be relatively large. In practical applications, after dividing several pieces of training sample data, the obtained data sets can be greater than or equal to two. For example, the training data can be divided into two data sets, three data sets, four data sets, etc. Specifically, the number of data sets can be determined based on actual needs.

[0046] As an implementation method, if it is necessary to divide into two sets, multiple pieces of training sample data with perplexity scores greater than the first preset score can be divided into the first set; multiple pieces of training sample data with perplexity scores less than or equal to the first preset score can be divided into the second set. As another implementation method, if it is necessary to divide into three sets, multiple pieces of training sample data with perplexity scores less than the first preset score can be divided into the first set; multiple pieces of training sample data with perplexity scores greater than the second preset score can be divided into the second set; multiple pieces of training sample data with perplexity scores between the first preset score and the second preset score can be divided into the third set. According to a similar logic, the training data can be divided into a preset number of data sets, which will not be listed one by one here.

[0047] In the embodiments of this specification, the perplexity score of any training sample data included in the first data set can be greater than the preset perplexity score; the perplexity score of any training sample data included in the second data set can be less than or equal to the preset perplexity score. Further, the perplexity score of any training sample data in the first data set is greater than the perplexity score of any training sample data in the second data set.

[0048] In practical applications, the first data set can include multiple training sample data in different fields; the second data set can also include multiple training sample data in different fields. The fields to which the training sample data included in the first data set belong can be the same as or different from the fields to which the training sample data included in the second data set belong. The number of training sample data included in the first data set can be the same as or different from the number of training sample data included in the second data set.

[0049] Step 208: Use each data set to perform an ablation experiment on the second largest model to determine the evaluation index values of the second largest model in each data set.

[0050] In the embodiments of this specification, the second largest model can be the same model as the first largest model, or a model that is different from the first largest model but of the same type. For example, both the first largest model and the second largest model are models based on the Transformer architecture. Specifically, they can be large models of the O1 series, or they can also be large models of the GPT series. The first largest model and the second largest model can be models of the same type but different scales. For example, the first largest model is a model with 2B parameters, and the second largest model is a model with 10B parameters.

[0051] An ablation experiment is a key method in machine learning and deep learning research to verify the effectiveness of model design. It can be achieved by removing or adjusting a certain component or parameter of the large model, or by inputting different training data into the model to determine the recognition situation of the model for different data, and observing the impact on the performance of the large model. In the embodiments of this specification, an ablation experiment can be performed on the second largest model through the data set to obtain the evaluation index value.

[0052] In practical applications, the same second largest model can be used to perform ablation experiments separately with different data sets to obtain the evaluation index values corresponding to each data set. The evaluation index value can be used to reflect the performance of the second largest model on different data sets. The same second largest model can refer to a single large model, or multiple large models with the same model parameters and architecture, or it can also represent multiple large models copied from a single large model, etc. There is no limitation on the model form here, as long as an ablation experiment can be performed.

[0053] In the embodiments of this specification, the evaluation metric value can represent a metric for evaluating the performance of the second largest model, such as at least one of metrics like accuracy, precision, recall, and F1-score. Accuracy can represent the ratio of the correct prediction results obtained by the second largest model for reasoning on each data set to all the training sample data. Specifically, it can represent the prediction accuracy rate. Precision can represent the ratio of the number of correctly predicted positive samples to the total number of samples predicted as positive classes. Specifically, it can represent the ratio of true positive samples among the samples predicted as positive classes in each data set by the second largest model. Recall can represent the ratio of the number of correctly predicted positive samples to the total number of true positive samples in each data set. Specifically, it can represent the ability of the second largest model to find true positive samples. F1-score can represent the harmonic mean of precision and recall, balancing the contradiction between the two.

[0054] Step 210: Determine the training sample data included in the target data set whose evaluation metric value in each data set is better than that of other data sets as the target training sample data for training the third largest model; the scale of the third largest model is larger than that of the first largest model.

[0055] In the embodiments of this specification, the third largest model can be a general inference large model, and the obtained target training sample data can be used to improve the inference ability of the general inference large model. As an implementation, the third largest model can be a model different from the first largest model or the second model, but the type of the third largest model can be the same as that of the first largest model and the second largest model. The scale of the third largest model can be larger than that of the first largest model and the second largest model. For example, the first largest model, the second largest model, and the third largest model are all large models of the GPT series. The third largest model can be a large model with an 80B scale, the second largest model can be a large model with a 10B scale, and the first largest model can be a large model with a 1B scale.

[0056] In the embodiments of this specification, that the evaluation metric of the target data set is better than that of other data sets can indicate that the training result obtained by training the second largest model with the training sample data in the target data set is more accurate than the training result obtained by training the second largest model with the training sample data in other data sets. The second largest model has a better understanding ability for the target training sample data in the target data set. The target training sample data in the target data set can be used to train the third largest model, thereby improving the model performance of the third largest model. Among them, since the model types of the second largest model and the third largest model are the same, the experimental results of the second largest model can be used to determine the target training sample data that can be used to train the third largest model.

[0057] In practical applications, the third large model can be trained using the target training sample data. Alternatively, based on the target training sample data, data with a high similarity to the target training sample data can be obtained from various data sources as training data to train the third large model. Alternatively, a known large model can be used to expand the target training sample data to obtain multiple pieces of data similar to the target training sample data, and the third large model can be trained using the obtained larger number of training data.

[0058] In the embodiments of this specification, the first large model, the second large model, and the third large model can all be base models. A base model can represent a general model pre-trained on a large scale and has the ability to adapt to a wide range of downstream tasks. The training of the third large model can also be the pre-training of the third large model, without labeling the training data, but only allowing the third large model to learn language knowledge in order to improve the knowledge reasoning ability of the third large model for data in various fields.

[0059] In practical applications, if the quality of the training sample data is high, instead of performing ablation experiments on the second large model, the first data set with a high perplexity score can be used as the target data set to train the third large model to improve the model ability of the third large model.

[0060] In practical applications, if the evaluation metric value of the first data set is better than that of the second data set, the first data set can be used as the target data set; or, if the evaluation metric value of the second data set is better than that of the first data set, the second data set can be used as the target data set; or, if the evaluation metric value of the first data set is not better than that of the second data set, the target data set may not be selected from the first data set and the second data set, and training sample data from other fields can be used for processing to determine the target data set in other fields.

[0061] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some of the steps can also be omitted or deleted.

[0062] Figure 2In the method, the perplexity scores of a number of training sample data obtained are determined by the first large model. The number of training sample data can be divided into at least two data sets based on the perplexity scores. Ablation experiments are performed on the second large model using each data set to determine the evaluation index values of the second large model in each data set. The data set with evaluation index values superior to other data sets is used as the target data set, and the training samples in the target data set are determined as the target training sample data for training the third large model. Among them, the target training sample data for model training can be determined based on the perplexity scores and the evaluation indexes of the large model under each training data, so that the target training sample data that can better improve the inference ability of the third large model can be determined in combination with the capabilities of the large model. In practical applications, the third large model can be trained with the target training sample data to improve the inference ability of the large model.

[0063] Based on Figure 2 For the method, some specific implementation schemes of this method are also provided in the embodiments of this specification, which will be described below.

[0064] The server can use the first large model to process a number of training sample data and determine the perplexity scores based on the processing results. Optionally, in the embodiments of this specification, determining the perplexity scores of the number of training sample data based on the first large model may specifically include: using the first large model to perform training inference on the number of training sample data to obtain an inference result; determining the perplexity scores of the number of training sample data based on the inference result.

[0065] In the embodiments of this specification, when the first large model performs training inference on the training sample data, it may be that the first large model infers the second token or segment based on the first token or segment in the training sample data, infers the third word based on the inferred second word, and so on until the last word is inferred to complete the inference of the training sample data and obtain an inference result. Thus, the perplexity scores of the training sample data can be determined based on each word information in the inference result.

[0066] In the embodiments of this specification, the perplexity score can represent the probability that the first large model selects the correct candidate word from the candidate words to be selected during the process of predicting the training sample data. It can be understood that the lower the perplexity score, the better the model fits the training sample data and the more accurate the prediction result. Further, when determining the perplexity scores of the same model for different training sample data, if the perplexity score is higher, it can indicate that the prediction accuracy of the model for the training sample data with a high perplexity score is worse. Or, when determining the perplexity scores of different models for the same training sample data, if the perplexity score is higher, it can be determined that the prediction performance of the model for this training sample data is worse.

[0067] In the related art, generally, those with high perplexity scores are deleted to avoid affecting the model. In this embodiment, the perplexity scores can be used to screen the training samples instead of directly deleting the training data with high perplexity scores. Since the data with high perplexity scores can be regarded as the data that the model does not correctly understand, if the target training sample data selected is the data with high perplexity scores, this data can also be used to train the third large model, so that the third large model can learn the knowledge that has not been learned or missed before, improving the performance of the model.

[0068] In practical applications, the inference results of the first large model on several training sample data can be used to determine the perplexity of each training sample data. Furthermore, based on the perplexity, the sufficiency of the first large model's learning of each training sample can be determined. Thus, the training samples with low perplexity and high prediction accuracy of the first large model, that is, the samples for which the model has fully learned this type of knowledge, can be determined as a data set, and the training samples with high perplexity and low prediction accuracy of the first large model, that is, the samples for which the model has not fully learned this type of knowledge, can be determined as a data set. Therefore, the sufficiency of the model's learning of each training sample data can be determined based on the perplexity scores, and the insufficiently learned ones can be divided into one data set, and the sufficiently learned ones can be divided into one data set. Furthermore, the data set with insufficient learning can be used to train the third large model to improve the model performance of the third large model.

[0069] To improve the convenience of method implementation, the perplexity scores can be recorded in units of sample data. Among them, for each training sample data, the perplexity score corresponding to the training sample data can be determined by using the perplexity scores corresponding to each word segment included in the sample data. This also facilitates screening the target training data in units of sample data, which can better meet the actual needs. It can also reduce the amount of data saved, and there is no need to save the perplexity score corresponding to each word segment for each word segment. Optionally, in the embodiments of this specification, determining the perplexity scores of the several training sample data may specifically include: for any one of the several training sample data, determining the perplexity scores of each word segment included in the any one of the training sample data; based on the perplexity scores of each word segment, determining the perplexity score of the any one of the training sample data.

[0070] In the embodiments of this specification, a word segment can represent each word segment obtained during the tokenization process of a training sample data. Among them, assuming that a training sample data is divided into N word segments, for the i-th word segment among them, the perplexity score of the word segment can represent the accuracy of the large model in predicting the i-th word segment based on the previous i - 1 word segments. It can be through the formula Calculate the perplexity scores of each token in a training sample data. Among them, N can represent the number of tokens in a training sample data; i can represent the i-th token in a training sample data; P(ω i |ω1,ω2,...,ω i-1 ) can represent the prediction accuracy of the i-th token by the model. exp can represent the natural exponential function, specifically, it can represent the exponential operation with the natural constant e as the base. In practical applications, the method of calculating the perplexity score can also refer to some public technologies, which will not be elaborated here.

[0071] After calculating the perplexity scores of each token, the perplexity scores of each token can be summed and averaged, and the obtained average value is used as the perplexity score of this training sample data. Or, it can also be the perplexity score of this training sample data obtained by weighted summation of the perplexity scores of each token. Among them, the weight value can be set according to actual needs. For example, the same weight value can be set, or different weight values can be set. For example, the weight of the token earlier in the training sample data can be set smaller, and the weight of the token later can be set larger; or, the weight of the token earlier in the training sample data can be set larger, and the weight of the token later can be set smaller.

[0072] To improve the processing efficiency of each training sample data, the distributed training method can be used to process each training data. Optionally, the method described in the embodiments of this specification may further include: serializing the several training sample data to obtain several data sequences; packing the several data sequences to obtain a packed sequence; dividing the packed sequence into batches to obtain several batch data; sharding the several batch data to obtain several sharded data; respectively providing each of the sharded data to each first large model, and using each first large model for distributed training inference.

[0073] In the embodiments of this specification, serializing several training sample data can represent performing tokenization processing on several training sample data to obtain the respective token values corresponding to each training sample data. The tokens belonging to the same training sample data can be concatenated to obtain a data sequence, which can also be called a token sequence. Tokenization processing can be to decompose a training sample data into units that the model can understand and process, such as a character, a word, a punctuation mark, etc. Tokenization processing can be the process of converting the original text into discrete units (tokens) that the model can process. For example, the text of a training sample data is tokenized, and the tokenized training sample is encoded to obtain a data sequence corresponding to the training sample data.

[0074] For example, at the word level, tokens such as the training data "this is an apple" can be tokenized to obtain ["this", "is", "an", "apple"]. Another example is that the training samples can be divided into character-level tokens. For example, "apple" can be split into ["a", "p", "p", "l", "e"]. After tokenizing the training sample data, the tokens are serialized. For example, when the training sample is divided into word-level tokens and the word "hello" is obtained, after serializing "hello", the serialized sequence [104, 101, 108, 108, 111] is obtained.

[0075] In the embodiments of this specification, the server can splice multiple data sequences together for packaging to obtain a packaged sequence of data sequences containing multiple training sample data. Multiple packaged sequences can also be batch-divided to obtain multiple batches of training data. One batch can contain multiple training sample data, which can further increase the amount of data input into the large model at one time and improve the processing efficiency of the large model. The length of the data in one batch can be preset according to the model processing ability or the computer processing ability. The data lengths of the batches obtained by batch division can not exceed the preset length, which can avoid the problem that the large model cannot recognize. In practical applications, based on the preset length of the batch, several packaged sequences can be divided or spliced together once to obtain several batches of data that meet the batch data length.

[0076] Generally, in order to make the training sample data input into the model meet the text recognition length of the model, it is necessary to fill each training sample data with useless tokens so that the length of the filled training sample data can meet the text recognition length of the model. However, this method will cause the model to recognize a large number of useless tokens, wasting model resources and reducing the efficiency of the model in processing training sample data.

[0077] In the embodiments of this specification, in order to reduce the number of useless tokens in the batch data input into the model, a batch of data can be obtained by splicing multiple data. If a batch of data does not meet the text recognition length of the model, a small amount of useless tokens can be filled. Thereby, the proportion of useless tokens in the training sample data input into the model can be reduced, avoiding too many useless tokens from occupying the model computing power, and also improving the model processing efficiency.

[0078] In the embodiments of this specification, sharding may mean dividing a number of batches of data into sharded data that requires different nodes for processing. A processing node may obtain one or more sharded data. The server may distribute a number of sharded data to different nodes of the first large model for distributed training. The first large model may perform parallel training on a number of sharded data through distributed training, improving the training efficiency and thus increasing the rate of obtaining the processing results.

[0079] In the embodiments of this specification, distributed training may be divided into parallel training methods such as data parallelism, pipeline parallelism, and tensor parallelism. Distributed training may mean storing the first large model distributively so that the first large model can be distributed across different nodes, and distributing a number of sharded data to each node where the first large model is distributed, and each first large model performs training based on the obtained sharded data. Alternatively, distributed training may also mean allocating different network layers of the first large model to different computing devices, distributing a number of sharded data to different network layers, and different network layers perform training based on the obtained sharded data. Or, the intra-layer parameters of the first large model can be split onto different computing devices, and a number of sharded data are distributed to the computing devices containing the intra-layer parameters to achieve intra-layer parallel training. The server may perform parallel training on sharded data through distributed training methods to improve the training efficiency.

[0080] By grouping data sequences and performing packaging processing using sequence sets, the packaging efficiency is improved. Optionally, the method described in the embodiments of this specification may further include: grouping the number of data sequences into multiple sequence sets; the step of performing packaging processing on the number of data sequences to obtain a packaged sequence may specifically include: for any one of the multiple sequence sets, performing packaging processing on each data sequence in the any one sequence set to obtain a packaged sequence.

[0081] In the embodiments of this specification, the number of sequence sets may be preset. A number of data sequences may be grouped based on the preset number of sets, and the number of data sequences may be divided into sequence sets with the preset number of sets; further, the server may perform encoding identification on each sequence set to obtain a lazy file corresponding to the preset number of sets.

[0082] In practical applications, the obtained training sample data may be tens of thousands of data or more. To improve the packaging efficiency, the training sample data can be broken into smaller parts. Specifically, the original number of data sequences can be divided into multiple sequence sets, and packaging processing is performed on sequence sets of a smaller magnitude, improving the packaging processing efficiency. In addition, the packaging operation can be performed in parallel for each sequence set, which can also improve the processing efficiency.

[0083] Moreover, encoding identifiers corresponding to different sequence sets can be assigned. After packing each sequence set, the encoding identifier of the corresponding sequence set can be assigned to the packed sequences. As a result, when searching for data, the corresponding packed sequences or sequence sets can be quickly found based on the encoding identifiers.

[0084] The server can calculate the perplexity score for each batch and can count the perplexity scores of the data in each batch, so as to be able to screen the target training sample data in batches subsequently. Therefore, when the data volume is large, the screening efficiency of the target training sample data can be improved. Optionally, in the embodiments of this specification, determining the perplexity scores of the several training sample data may specifically include: for any batch of data among the several batches of data, determining the perplexity scores corresponding to each data sequence included in the batch of data; according to the perplexity scores corresponding to each data sequence included in the batch of data, determining the perplexity score corresponding to the batch of data.

[0085] In the embodiments of this specification, after tokenizing the training sample data, the perplexity scores of each token in a data sequence can be calculated using the above formula. The mean value processing can be performed on each perplexity score to obtain a first mean value, and the obtained first mean value can be used as the perplexity score of the data sequence; alternatively, the weighted sum of the perplexity scores of each token can be calculated, and the result of the weighted sum can be determined as the perplexity score of the data sequence.

[0086] In the embodiments of this specification, the perplexity scores corresponding to each data sequence included in the same batch can be determined in the above manner. The weighted sum processing is performed based on the perplexity scores of each data sequence in the same batch to obtain the perplexity score corresponding to the batch; alternatively, the mean value processing is performed based on the perplexity scores of each data sequence in the same batch to obtain a second mean value, and the second mean value can be used as the perplexity score corresponding to the batch. As a result, the perplexity score corresponding to a batch can be obtained, improving the accuracy of data set division for the batch.

[0087] In the embodiments of this specification, the server can divide several batches into multiple data sets based on the perplexity scores of each batch in batches; one data set can include one or more batches. Therefore, when the data volume is large, the data sets can be divided in batches according to the perplexity scores of the batch data, improving the efficiency of screening the target training sample data based on the data sets.

[0088] By storing the correspondence between the data indices and perplexity scores of each training sample data, it is possible to use the data index to replace the original training sample data during storage, without storing the original training sample data, which can reduce the amount of data stored and further reduce the data storage cost. Optionally, the method described in the embodiments of this specification may further include: determining the sequence index corresponding to each data sequence in any batch of data; the sequence index represents the position information of each data sequence included in any batch of data; saving the first correspondence between the sequence index and the perplexity score corresponding to any batch of data.

[0089] In the embodiments of this specification, the sequence index may be an index such as a packed index, a hash index, a partition index, an inverted index, etc. The sequence index can be used to represent the start position and end position of each training sample data in a batch. The position information may include the start position and end position of the data sequence corresponding to each training sample data. During the process of the large model processing multiple data sequences included in each batch, the start position and end position of each data sequence represented in the sequence index can be used to determine the complete data sequence corresponding to a training sample data, avoiding confusion with other training sample data.

[0090] In the embodiments of this specification, the perplexity score in the first correspondence may represent the perplexity score corresponding to any training sample data in this batch. For the convenience of statistics, each piece of data in the same batch can be recorded as the same perplexity score. Alternatively, the perplexity score in the first correspondence can be determined as the perplexity score of a batch of data, and the perplexity score is stored in units of batches, which is convenient for subsequent division of the data set based on the perplexity score of the batch. It is possible to avoid data screening in units of training sample data, which can improve the processing efficiency and reduce the processing cost.

[0091] By iterating the second large model multiple times with the data set, multiple evaluation indicators corresponding to each data set are obtained, improving the accuracy of the finally determined target training sample data. As an implementation manner, optionally, the evaluation indicator values described in the embodiments of this specification may include multiple evaluation indicator values generated in multiple different rounds of iteration; the training sample data included in the target data set whose evaluation indicator values in each data set are better than those of other data sets is determined as the target training sample data for training the third large language model, which may specifically include: if the number of evaluation indicator values in the multiple evaluation indicator values corresponding to the first data set that are better than those in the multiple evaluation indicator values corresponding to the second data set is greater than or equal to a preset threshold, then the training sample data included in the first data set is determined as the target training sample data.

[0092] In the embodiments of this specification, during the iterative training of the large model, the evaluation metric values of the large model after each round or some rounds of iteration can be obtained. If the second large model is iteratively trained in multiple rounds using the first data set, the second large model after each round of iteration can be evaluated to obtain the evaluation metric values of the second large model for each round of iteration based on the first data set; or, after the second large model is iteratively trained for a preset number of rounds using the first data set, the second large model after the preset number of rounds of iteration is evaluated to obtain the evaluation metric values corresponding to the second large model for the preset number of iterative rounds.

[0093] Alternatively, for any data set in the first data set or the second data set, a first preset number of rounds of iteration can be performed on the second large model by obtaining training sample data of a first data volume from the any data set to determine the evaluation metric values corresponding to the any data set of the first data volume. Training sample data of a second data volume is obtained from the any data set and the second large model is similarly iteratively trained for a second preset number of rounds to determine the evaluation metric values corresponding to the any data set of the second data volume. Furthermore, multiple evaluation metric values corresponding to the second large model trained using the any data set can be obtained. The first preset number of rounds can be the same as the second preset number of rounds, or the first preset number of rounds can be different from the second preset number of rounds.

[0094] In practical applications, the training sample data in the first data set can be given to the second large model, and the second large model is trained for a preset number of rounds using the first data set to obtain the trained second large model. The validation set or the evaluation set is used to determine the evaluation metric values of the trained second large model. To clearly illustrate the method for obtaining the evaluation metric values, here, taking the example of performing a preset number of rounds of iteration on the second large model using different amounts of data to determine the evaluation metric values of the second large model under different amounts of data: A first amount of training data can be sampled from the first data set to perform multiple rounds of iterative training on the second large model, and then the validation set is used to validate the second large model after multiple rounds of iterative training to obtain the first evaluation metric value; A second amount of training data is sampled from the first data set to perform multiple rounds of iterative training on the original second large model or the second large model that has been trained using the first amount of data, and then the validation set is used to validate the second large model after multiple rounds of iterative training using the second amount of training data to obtain the second evaluation metric value, and so on. Multiple evaluation metric values of the second large model for the first data set can be obtained.

[0095] Similarly, multiple evaluation metric values of the second largest model for the second data set can be determined based on this method; thus, the target data set can be determined from the two data sets based on the multiple evaluation metric values. It can be understood that the number of training sample data obtained in different rounds can be different, the training sample data in different rounds can be different, or part of the training sample data in different rounds can be the same. Through the above method, multiple evaluation metric values can be obtained, and then the target training sample data can be determined through the multiple evaluation metric values, improving the accuracy of the obtained target training sample data.

[0096] The training sample data with different or the same data volume used in the above training process of the second largest model can also be data similar to the training sample data in any data set obtained from the data source.

[0097] In the embodiments of this specification, the preset threshold can be determined based on expert experience, or can be determined based on actual requirements or model performance. If the evaluation metric value of the first data set is better than the evaluation metric value corresponding to the second data set, it can be indicated that the difference between the evaluation metric value of the first data set and the evaluation metric value corresponding to the second data set is greater than the preset metric difference. For example, assuming the preset metric difference is 1%, the evaluation metric value of the first data set is 30%, and the evaluation metric value of the second data set is 26%, then it can be determined that the difference between the two is 4% > 1%, and it can be determined that the evaluation metric value of the first data set is better than the evaluation metric value of the second data set.

[0098] The server can count the number of evaluation metric values of the first data set that are better than the evaluation metric values of the second data set for the obtained multiple evaluation metric values corresponding to the first data set and the multiple evaluation metric values corresponding to the second data set; if the counted number is greater than or equal to the preset number, the training sample data in the first data set can be used as the target training sample data, and the third largest model can be trained using the target training sample data to reduce the perplexity of the third largest model for this type of data and improve the inference accuracy of the third largest model.

[0099] To clearly illustrate the process for determining the target training sample data provided in the above embodiments, Figure 3 is a schematic diagram of the training sample data and the corresponding evaluation metric values provided by the embodiments of this specification. Assume that the model is trained using the Massive Multitask Language Understanding (MMLU) test set, and the evaluation metric values of the large model are obtained. As Figure 3As shown, the schematic diagram can be obtained by performing an ablation experiment after processing the data in the social science field in the MMLU test set in the above manner. The horizontal axis can represent the number of training sample data in the social science field used during the ablation experiment, and the vertical axis can represent the corresponding evaluation metric values. The fineweb_middle_9_15 shown by the green broken line can represent obtaining the evaluation metric values by performing an ablation experiment on the second largest model using different numbers of training sample data in data set 1 within the PPL value range of 9 - 15 in the data for the social science field; the fineweb_middle_5_9 shown by the yellow broken line can represent obtaining the evaluation metric values corresponding to each round by performing a preset number of rounds of ablation experiments on the second largest model using different numbers of training sample data in data set 2 within the PPL value range of 5 - 9 in the data for the social science field; the fineweb_middle_base shown by the blue broken line can represent obtaining the evaluation metric values corresponding to different numbers of training sample data by performing a preset number of rounds of ablation experiments on the second largest model using different numbers of training sample data in data set 3 containing all the data in the social science field without a PPL value limit. Among them, it can be clearly seen that the higher the data where the evaluation metric values represented by the green broken line are located, the higher the metrics compared to the yellow broken line and the blue broken line, which can indicate that the influence degree of data set 1 on the prediction accuracy of the second largest model is higher, and the training sample data in data set 1 can be used as the target training sample data. The number of evaluation metric values represented by the green broken line that are both better than the evaluation metric values of data set 2 represented by the yellow broken line and better than the evaluation metric values of data set 3 represented by the blue broken line is greater than the preset number four. It can be determined that data set 1 corresponding to the green broken line is the target data set, and further, the training sample data in data set 1 can be determined as the target training sample data. Green is more conducive to improving the model performance. On the other hand, the high PPL value of the green broken line can indicate that the model has not been fully learned before. The present application can use the training sample data in data set 1 with a high PPL value and a good improvement in the model performance of the second largest model as the training data for training the third largest model, so as to identify the data that the model has missed and can improve the performance, and accurately identify and fill the gaps.

[0100] Through the above Figure 3 It can be intuitively found that the model trained with data in the high PPL score segment shows significantly better overall performance in various metrics than the model trained with data in the low PPL value segment. It can be understood that Figure 3 This is only an example and not a specific limitation.

[0101] In practical applications, the training sample data that may be obtained is data that the model has already fully learned, and it may not be necessary to use such data to train the model. Figure 4It is a schematic diagram of training sample data and corresponding evaluation index values provided in the embodiments of this specification. Assume that the large model is trained using the Massive Multitask Language Understanding (MMLU) test set, and the evaluation index values of the large model are obtained. As Figure 4 shown, this schematic diagram can be obtained by performing an ablation experiment after processing the data in the historical domain in the MMLU test set in the above manner. The horizontal axis can represent the number of training sample data in the historical domain used during the ablation experiment, and the vertical axis can represent the corresponding evaluation index values. The fineweb_middle_9_15 shown by the green broken line can represent performing an ablation experiment on the second large model for a preset number of rounds using different numbers of training sample data in data set 4 with a PPL value in the range of 9 - 15 for the training data in the historical domain, and obtaining the evaluation index values corresponding to each round; the fineweb_middle_5_9 shown by the yellow broken line can represent performing an ablation experiment on the second large model for a preset number of rounds using different numbers of training sample data in data set 5 with a PPL value in the range of 5 - 9 for the training data in the historical domain, and obtaining the evaluation index values corresponding to different numbers of training samples; the fineweb_middle_base shown by the blue broken line can represent performing an ablation experiment on the second large model for a preset number of rounds using different numbers of training sample data in data set 6 containing all the data in the historical domain without restricting the PPL value for the training data in the historical domain, and obtaining the evaluation index values corresponding to each round. Among them, it can be clearly seen that there are many intersections between the green broken line, the yellow broken line, and the blue broken line, and the numerical values of the evaluation index values corresponding to the three data sets are relatively close. It can be determined that the influence degrees of the three data sets on the second large model are all relatively small. Furthermore, it can be determined that the second large model has a deep learning degree for the data in the historical domain, and the target data set cannot be determined from data set 4, data set 5, and data set 6 representing the data in the historical domain. It can be understood that Figure 4 This is only an example and is not a specific limitation.

[0102] In practical applications, a number of training sample data obtained may include data in multiple fields or disciplines. After obtaining the data, the training sample data can be classified according to the field or discipline. Then, for the training sample data in each field, the target training sample data in each field can be screened respectively in the above manner. Further, the third large model can be trained using the target training sample data to improve the model performance of the third large model. Among them, if model training is performed on the data of a field, the full amount of field data can be divided into a training set, a validation set, and a test set according to a preset ratio. After processing the second large model through different data sets divided based on the training set, the performance of the model on the same test set or validation set can be observed, and the differences between the second large models obtained by training with data sets representing different PPL score segments can be compared.

[0103] As another implementation manner, optionally, the evaluation index values described in the embodiments of this specification may include multiple evaluation index values generated through multiple different rounds of iteration; determining the training sample data included in the target data set with evaluation index values superior to those of other data sets in each data set as the target training sample data for training the third large language model specifically includes: if the evaluation index values generated by the preset consecutive rounds of iteration corresponding to the first data set are all superior to the evaluation index values generated by the preset consecutive rounds of iteration corresponding to the second data set, then determine the training sample data included in the first data set as the target training sample data.

[0104] In the embodiments of this specification, samples with different data volumes can be used for training a preset number of times. The number of training sample data used in each training can be incremented, that is, the number of training sample data used in the subsequent round is more than that used in the previous round. If the evaluation index values generated in the preset consecutive rounds corresponding to the first data set are all better than the evaluation index values generated in the preset consecutive rounds corresponding to the second data set, it can be indicated that the evaluation index value of the second largest model corresponding to any one of the consecutive several rounds in the first data set is greater than the evaluation index value of the second data set in the corresponding round. For example, assume that the preset consecutive rounds are 3 times, and 5 trainings are performed based on the first data set and the second data set; in the order of training, the evaluation index values of the first data set are 22%, 21.5%, 23%, 26%, and 28.7% respectively; in the order of training, the evaluation index values of the second data set are 22%, 22.5%, 21.8%, 22.7%, and 23% respectively; then it can be determined that the evaluation index value of the third training of the first data set is greater than the evaluation index value of the third training of the second data set, the evaluation index value of the fourth training of the first data set is greater than the evaluation index value of the fourth training of the second data set, and the evaluation index value of the fifth training of the first data set is greater than the evaluation index value of the fifth training of the second data set, and they are consecutive and 3 times, then several pieces of training sample data in the first data set can be determined as the target training sample data. Through the above method, the target training sample data that has a greater impact on the performance of the second largest model can be determined, and then the target training sample data that has a greater impact on the performance of the third largest model of the same type can be determined, so as to be able to train the third largest model based on the target training sample data and improve the reasoning ability of the third largest model for this type of data such as the target training sample data.

[0105] Before determining the target training samples, the training sample data can also be preprocessed to ensure the quality of the training sample data and reduce the impact of the noise of the training sample data on the processing results. Optionally, the method in the embodiments of this specification may further include: obtaining several pieces of initial training sample data; performing data preprocessing on the initial training sample data to obtain the several pieces of training sample data; the preprocessing includes at least one of a process for removing data noise, a process for correcting grammar or semantics, and a process for removing sensitive words.

[0106] In the embodiments of this specification, the initial training sample data may represent the sample data that needs to be stored in the database and is obtained from multiple data sources. After obtaining the data from multiple data sources, the initial training sample data can be preprocessed, and the preprocessed initial training sample data can be stored in the database, so as to directly obtain the preprocessed data from the database as several pieces of training sample data. Alternatively, the initial training sample data may also represent several pieces of training sample data obtained. Before determining the perplexity score using the first large model, the several pieces of training sample data can be preprocessed, and based on the first large model, the perplexity score corresponding to the preprocessed several pieces of training sample data can be determined.

[0107] In the embodiments of this specification, removing data noise may mean removing the data in the initial training sample data that affects the recognition of the large model, such as irrelevant characters, HTML characters, special symbols, extra spaces, duplicate texts, etc. Correcting grammar or semantics may mean correcting the grammar, word order, etc. problems existing in the text of the initial training sample data. Removing sensitive words may mean removing the relatively sensitive words in the initial training sample data, such as sensitive words indicating violations, etc.

[0108] In the embodiments of this specification, by preprocessing the initial training sample data, the initial training sample data can be cleaned, and data with a higher perplexity score can be selected from the high-quality data as the target training sample data, without conducting an ablation experiment on the second large model to determine the target training sample data, which can promote the third large model to learn knowledge that has not been learned before, identify and fill in the gaps, and improve the performance of the third large model.

[0109] Generally, our intuitive understanding is that the lower the PPL value, the higher the quality of the text usually means. Therefore, the data in the low score segment can provide higher-quality input information for model training, making it easier for the model to learn effective language patterns and semantic relationships, and thus can have better performance in various evaluation indicators. While the data in the high PPL score segment may have more noise, grammar or semantic problems, which affect the extraction and learning of effective knowledge by the model, resulting in relatively lower performance of the final large model. And after processing the initial training sample data through the above method, the data quality is relatively high. In this case, if the PPL value is still high, it can be determined that the data with a high PPL may be the knowledge that the large model has not fully learned or missed.

[0110] Since the amount of data contained in each data set is large, in order to improve the efficiency of the ablation experiment, part of the data can be obtained from the data set to conduct the ablation experiment on the second largest model. Optionally, in the embodiments of this specification, using each data set to conduct the ablation experiment on the second largest model may specifically include: for any one of the data sets, selecting a preset number of training sample data from the any one of the data sets; using the preset number of training sample data to conduct the ablation experiment on the second largest model.

[0111] In the embodiments of this specification, the preset number can be determined based on expert experience, or can be determined based on the performance of the second largest model, or can be determined based on the needs of the experiment user conducting the ablation experiment on the second largest model, and no specific limitation is made here. The preset number can be less than or equal to the number of training sample data contained in any data set. For example, assuming that the number of training sample data contained in a data set is 60,000, the preset number can be set to values such as 50,000, 20,000, 35,000, etc.

[0112] In the embodiments of this specification, the second largest model can be used for multiple ablation experiments with different data sets to obtain the evaluation index values of the second largest model corresponding to different data sets, and determine the influence of different data sets on the second largest model. Among them, the training sample data in the data set with a higher degree of influence can represent the data that plays a stronger role in the second largest model, which is more beneficial to improving the model performance, or the training sample data that the second largest model cannot correctly identify or has insufficient learning and cannot accurately process. The data set can be used as the target data set, and the training sample data in the data set can be used as the target training sample data to train the third largest model of the same type, so that the third largest model can learn the target training sample data, solve the problem of low learning degree of the target training sample data, and further improve the reasoning ability of the third largest model to process various texts.

[0113] In practical applications, the server can store each data set through distributed cloud storage, and the each data set obtained from the cloud storage is used for the ablation experiment. The cloud storage can be a Cloud Parallel File System (CPFS), an Object Storage Service (OSS), a Network Attached Storage (NAS), etc. In order to reduce resource occupancy and improve the throughput rate of the overall system, the sequence index corresponding to the data sequence of the training sample data and the perplexity score corresponding to the sequence index can be stored in the preset database, and the training sample data is not stored.

[0114] As an implementation manner, optionally, saving the first corresponding relationship between the sequence index and the perplexity score corresponding to any batch of data in the embodiments of this specification may specifically include: saving the first corresponding relationship between the sequence index and the perplexity score corresponding to any batch of data to the distributed file system CPFS.

[0115] In the embodiments of this specification, a sequence index can be used to identify the start position and the end position of each piece of training sample data in a batch. A sequence index can uniquely identify a batch of data. The perplexity score corresponding to a batch of data can represent the perplexity scores of each piece of training sample data in this batch. The original training sample data corresponding to the sequence index may not be saved in CPFS, and the sequence index and the perplexity score corresponding to the batch of data are saved.

[0116] In the embodiments of this specification, the file system CPFS can be a Cloud Parallel File System (CPFS), which can refer to a distributed storage solution for high-performance computing, big data analysis, and applications that require shared file storage. CPFS can support multiple devices to access the unified file system in parallel. Furthermore, multiple data sets can be obtained in parallel using CPFS, and parallel ablation experiments can be performed on the second largest model using multiple data sets to improve the efficiency of determining the evaluation index value of the second largest model. Further, the efficiency of determining the target training sample data can be improved. Moreover, storing data using CPFS can improve data throughput and reduce latency, etc.

[0117] As another implementation manner, optionally, the method in the embodiments of this specification may further include: saving the first corresponding relationship saved to the distributed file system CPFS to the object storage service system OSS.

[0118] In the embodiments of this specification, the object storage service system (OSS) can be a cloud storage solution based on an object storage architecture, which can encapsulate the first corresponding relationship in CPFS into an independent data object; the data object can include information such as the data itself, metadata, and unique identifiers contained in the first corresponding relationship. On the one hand, OSS can store the encapsulated data object for a long time, improving the data preservation duration; on the other hand, it can also store unstructured data.

[0119] In the embodiments of this specification, the data stored in CPFS can be copied to OSS, or it can be moved from CPFS to OSS. If it is moved from CPFS to OSS, the data in CPFS will be deleted after it is saved in OSS. OSS can store data through multiple nodes based on a distributed architecture, reducing storage costs. In practical applications, network attached storage (NAS) can also be used to perform cloud storage on the first correspondence between the sequence index and the confusion score corresponding to any batch of data. Specifically, file-level storage services can be provided through a local area network. NAS can connect storage devices to a network so that multiple computers can share and access storage resources.

[0120] The server may also save the batch data, the confusion score of the batch data, and the corresponding relationship between the index sequence for subsequent data processing and analysis. Optionally, the method described in the embodiment of this specification may also include: saving the second corresponding relationship between the sequence index, the training sample data corresponding to the index, and the confusion score corresponding to any batch data.

[0121] In the embodiments of the present specification, the sequence index may have a corresponding relationship with the perplexity score corresponding to any batch of data; the training sample data corresponding to the index may be the training sample data contained in any batch of data having a corresponding relationship with the sequence index. In practical applications, after the first corresponding relationship is stored in CPFS or OSS, the server may associate the data sequence of each training sample data in the batch data with the perplexity score of the batch data through the sequence index in an artificial intelligence platform such as the AI Security Testing Platform (AIS) of Ant Detection, and obtain a second corresponding relationship; specifically, the artificial intelligence platform may obtain the first corresponding relationship from CPFS or OSS, and determine the data sequence of multiple training sample data belonging to the batch data corresponding to the perplexity score identified by the sequence index based on the perplexity score of the batch data in the first corresponding relationship, and then establish a second corresponding relationship between the data sequence of the multiple training sample data, the sequence index and the perplexity score of the batch data, so that based on the second corresponding relationship, the data sequence of multiple training sample data corresponding to each batch of data belonging to the same data set and the sequence index corresponding to each batch of data may be input into the second largest model for ablation experiment.

[0122] To improve the throughput and acquisition rate of data, the second corresponding relationship can also be stored in the Open Data Processing Service. Optionally, in the embodiments of this specification, saving the second corresponding relationship of the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data may specifically include: saving the second corresponding relationship of the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data to the Open Data Processing Service ODPS.

[0123] In the embodiments of this specification, the Open Data Processing Service (ODPS) is a large-scale, fully managed data warehouse and analysis platform that can be used for storing, computing, and analyzing massive amounts of data and can be widely applied to scenarios such as data warehouses, machine learning, and log analysis, and can be used for distributed training and model deployment. In actual applications, storing the second corresponding relationship in ODPS completes the shelving and cloud processing of the training sample data with perplexity scores and sequence indexes. PAI can be integrated with ODPS to perform distributed table storage of the second corresponding relationship, storing the corresponding relationship among the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data in a data table to provide data support for subsequent processing.

[0124] As an implementation, the second corresponding relationship can also be stored in CPFS. Optionally, in the embodiments of this specification, the method may further include: synchronously saving the second corresponding relationship to the distributed file system CPFS.

[0125] In the embodiments of this specification, the second corresponding relationship can be synchronously saved to a distributed file system, and the distributed file system can provide the training sample data for ablation experiments for the second large model. Thus, there is no need to obtain the training sample data from the data source, improving the efficiency of obtaining the training sample data and making the ablation experiment more efficient.

[0126] In actual applications, after obtaining the batch data with perplexity scores in CPFS, each batch can be divided into multiple data sets according to the perplexity score of the batch, so as to perform ablation experiments on the second large model according to the data sets. Or, during the process of storing in CPFS, multiple data sets can be determined based on the perplexity score of the batch, and the second corresponding relationship can be stored in the form of data sets. One file in CPFS can store one data set, and one data set can contain multiple second corresponding relationships and the perplexity scores of the data sequences, training sample data, and batch data with the second corresponding relationship, etc., facilitating obtaining data from CPFS in units of data sets to perform ablation experiments on the second large model.

[0127] In practical applications, several pieces of source data can be obtained from multiple data sources, such as data sources like public examinations and textbooks, online education platforms, professional field resources, and academic papers and datasets. After performing domain classification, quality filtering, and balanced distribution processing on the source data, an MMLU test set can be obtained. The MMLU test set can be used to evaluate the language model's ability to understand multi-task and multi-domain knowledge. The MMLU test set can contain 57 subject areas, such as subject areas like humanities, social sciences, medicine, law, and so on. The server can determine the processing ability of the second-largest model for natural language-type data without performing labeling processing on the data in the MMLU test set, and further determine the processing ability of the third-largest model of the same type.

[0128] In practical applications, several pieces of training sample data obtained can be data from multiple domains obtained from the MMLU test set; the several pieces of training sample data can be classified into data of different domains, and ablation experiments can be performed on the second-largest model for the data of each domain to obtain multiple evaluation metric values representing the processing ability of the second-largest model for the data of each domain. The data of the domain with better evaluation metric values can be used as target training data to train the third-largest model. If the perplexity score values of the data in a domain belong to different perplexity score value segments, the evaluation metric values corresponding to the data in different perplexity score value segments in this domain can be determined. If there is an obvious stratification among the evaluation metric values and they do not belong to an intersecting state, it can be determined that the data in this domain is data that the third-largest model needs to learn. All the data in this domain can be determined as target training sample data to train the third-largest model, or the data corresponding to the better evaluation metric values of different perplexity score value segments in this domain can be determined as target training sample data to train the third-largest model, improving the third-largest model's processing ability for the data in this domain.

[0129] In practical applications, several pieces of training sample data obtained can be data from one domain obtained from the MMLU test set. The perplexity score values of the data in this one domain can be determined, and multiple data sets can be divided based on the perplexity score values. The evaluation metric values representing the processing ability of the second-largest model for the data in each data set can be determined using the data sets. Furthermore, based on the evaluation metric values, the data in this domain that the large model lacks in learning can be determined from multiple data sets, and the training sample data in the target data set that is superior to other data sets can be used as target training sample data to train the third-largest model, enabling the third-largest model to eliminate the perplexity for the target training sample data after learning this part of the target training sample data and improving the large model's ability to process natural language.

[0130] The various technical features in the above embodiments can be combined arbitrarily as long as there are no conflicts or contradictions between the features. However, due to space limitations, not all combinations are described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope disclosed in this specification.

[0131] According to the above description, in order to clearly illustrate the embodiments of the present application and the corresponding means, Figure 5 is the overall schematic diagram of a method for determining training sample data provided by an embodiment of this specification.

[0132] Step 502: Store the original tokenized data set into the cloud distributed file system CPFS.

[0133] In the embodiments of this specification, several initial training sample data can be obtained from multiple trusted data sources, such as education platforms, knowledge popularization platforms, paper data platforms, etc. The several initial training sample data can be general data.

[0134] In practical applications, several initial training sample data can be tokenized to obtain an original tokenized data sequence, and the data sequence can also be divided into lazy data sets to obtain multiple data sets, which can be stored in the CPFS database for storing data obtained from external data sources.

[0135] The CPFS database is a distributed storage solution designed for high-performance computing, big data analysis, and applications that require shared file storage, and can be used to store data sets for a long time. In practical applications, if the server needs to process data, it can directly obtain several initial training sample data quickly from this CPFS database without searching from the data source.

[0136] Before tokenizing the data, it is also possible to perform processing procedures such as procurement and disinfection on the data obtained from the source data. Specifically, for example, removing sensitive words, correcting grammar or semantics, removing noise, etc., which can improve the data quality of the obtained source data. The above-mentioned original tokenized data set can represent the data sequence after tokenization after performing processing procedures such as procurement and disinfection; or, it can also represent the data sequence after tokenization after performing processing procedures such as procurement and disinfection, and after dividing the data sequence into lazy data sets, the obtained multiple lazy data sets.

[0137] Step 504: Pre-train data processing.

[0138] In the embodiments of this specification, pre-training data processing may refer to partitioning the data in the dataset for distributed pre-training inference. Specifically, it may include processing such as packing the data in the dataset, batch partitioning, and sharding. Data preprocessing may refer to performing on the initial training sample data.

[0139] In the embodiments of this specification, to improve data processing efficiency, data processing may be performed in a distributed manner. The server nodes used for pre-training can read the stored Lazy dataset from the cloud storage CPFS. For any Lazy dataset, preprocessing may include operations such as packing (Packed), batching, and sharding the data. For example, multiple data sequences in the dataset can be packed to obtain several packed sequences, and at least one data sequence is included in one packed sequence; then, several packed sequences are grouped to obtain several batches of data; one batch of data contains at least one packed sequence, and the sequence lengths of different batches of data can be the same, or the sequence lengths of different batches of data can also be different. After batch partitioning, several batches of data can also be sharded to obtain multiple sharded data. So that the server nodes can use the sharded data for the subsequent distributed pre-training inference for the first large model.

[0140] Step 506: Pre-training inference and PPL scoring.

[0141] Pre-training inference may refer to performing distributed training inference on the first large model using multiple sharded data. Specifically, the first large model can be deployed to multiple server nodes in the distributed server, and multiple shards are input into the first large model of different server nodes, so that the first large model can perform distributed training inference based on multiple shards. In the embodiments of this specification, PPL scoring may refer to performing PPL scoring on each batch of data included in each shard based on the inference result of the distributed training inference of the first large model. Specifically, the perplexity score of the batch data may first calculate the first perplexity score for each token in the batch data, and then determine the second perplexity score of this training sample data by calculating the average of the first perplexity scores of the tokens belonging to the same training sample data; the second perplexity scores of each training sample data belonging to the same batch of data are averaged to obtain the third perplexity score corresponding to one batch of data. In this way, the perplexity scores of each batch of data can be calculated. For the convenience of subsequent data processing, the perplexity score of one batch of data can be determined as the perplexity score of the training sample data included in this batch of data.

[0142] Step 508: Store the data index and PPL value of the original data into CPFS.

[0143] The sequence index can be used to identify the start position and end position of each data sequence in the batch data, so that after the batch data is input into the large model, the large model can identify each data sequence based on the sequence index. The server node can index the original tokenized data sequences of the original data contained in each batch of data using a preset index to obtain the data index of the original data corresponding to a batch of data, or obtain the data index corresponding to each original data. The meaning of the data index here is the same as that of the above-mentioned sequence index.

[0144] During the pre-training process, the server node can generate the index of the original tokenized data sequence so that the first large model can distinguish different training samples based on the index during the training process. To facilitate the orderly screening of target training samples, the server node can store the first correspondence between the sequence index and the PPL value of the batch data in CPFS. Do not store the data sequences of each training sample data corresponding to the batch data; thus, the throughput efficiency of the overall system can be improved.

[0145] In practical applications, each server node can use the same CPFS database to store the first correspondence generated by each server node, or can use different CPFS databases to store the first correspondence generated by different server nodes respectively. For example, one server node corresponds to one CPFS database. Here, the first correspondence is stored. This CPFS can be a different database from the CPFS database storing the initial training data above. This CPFS database can be used to store the data generated by the nodes in the distributed server. And the above CPFS is used to store the data obtained from the data source.

[0146] Step 510: Store the data index and PPL value of the original data into the object storage service system OSS.

[0147] To further improve the data processing efficiency and reduce the data storage cost, in the embodiments of this specification, the first correspondence between the data index and the PPL value of the original data stored in CPFS can be transferred and stored in OSS storage. OSS storage can be used as an intermediate transfer medium to ensure the efficiency of data reading and writing. OSS can obtain the first correspondence from CPFS for storage to avoid the problem that the pressure of storing data in CPFS is too large due to the relatively fast data processing speed of the node.

[0148] Step 512: Associate the data sequences of the original data based on the data index and PPL value obtained from OSS.

[0149] In the embodiments of this specification, the original data can be obtained from CPFS. The original data can be the original training sample data, such as multiple texts; or, what is stored in CPFS is the word segmentation sequence or token sequence corresponding to the original training sample data. A corresponding relationship is established between the obtained original data or the word segmentation sequence or token sequence of the original data, the data index, and the PPL value. It can also be understood as saving the corresponding relationship among the data index, PPL value of the original data, and the corresponding original data or the word segmentation sequence or token sequence of the original data. Among them, the word segmentation of the original data stored in CPFS is in the same standardized data format as the packaged data after pre-training data processing. For example, the format after tokenizing the original training sample data and the format for establishing the index are correspondingly consistent, and the order of the corresponding training samples in both is the same, avoiding the inconsistency between the data sequence and the sequence order of the index identifier, which is beneficial for subsequent searches.

[0150] In the embodiments of this specification, a second corresponding relationship among the data sequence of the original data, the sequence index, and the perplexity score of the batch data can be established on the AIS artificial intelligence platform. Among them, AIS can obtain the first corresponding relationship from OSS and obtain the data sequence corresponding to the sequence index from CPFS by looking up the table, and establish a second corresponding relationship among the data sequence, the sequence index, and the perplexity score of the batch data.

[0151] Step 514: Upload and store the data sequence, data index, and PPL value of the associated original data to the cloud and on the shelves.

[0152] In the embodiments of this specification, uploading and storing the data sequence, data index, and PPL value of the associated original data to the cloud and on the shelves can mean storing the second corresponding relationship obtained after processing the data on the artificial intelligence platform in ODPS to complete offline storage. ODPS can be used to store the data processed by each node, complete the aggregation of the data, and avoid the problem that the data is too scattered and needs to obtain data from different databases or nodes.

[0153] To facilitate data reading, step 516 can also be executed: Store the data sequence, data index, and PPL value of the associated original data in CPFS.

[0154] In the embodiments of this specification, the second corresponding relationship among the data sequence, data index, and PPL value of the associated original data can be stored in CPFS, so that CPFS stores batch data with perplexity scores and sequence indexes, so as to provide efficient data support for subsequent ablation experiments. The CPFS database storing the second corresponding relationship can be a database arranged locally. It is a database of the same type as the above CPFS but arranged on different nodes.

[0155] Step 518: Divide the data in CPFS into data sets and sample from the data sets.

[0156] In the embodiments of this specification, after obtaining several batches of data with perplexity scores and sequence indexes from CPFS, the perplexity scores of each batch of data can be used to divide the data into multiple data sets. The divided data sets can be determined based on actual requirements. The server node can perform sample-level sampling from each data set, where sampling from each data set can mean obtaining a preset number of batches of data from any of the multiple data sets.

[0157] Step 520: Obtain the training sample data set for the first score segment, the training sample data set for the second score segment, and the training sample data set for the third score segment.

[0158] In the embodiments of this specification, the training sample data set for the first score segment, the training sample data set for the second score segment, and the training sample data set for the third score segment can be obtained by sampling from the first data set, the second data set, and the third data set respectively. As an implementation manner, the PPL score of any batch of data in the first data set is higher than the PPL score of any batch of data in the second data set; the PPL score of any batch of data in the second data set is higher than the PPL score of any batch of data in the third data set.

[0159] It should be noted that the training sample data for the above three score segments are only illustrative. In actual applications, the division of samples and the number of divided sets can be determined according to actual requirements, which are not limited here.

[0160] Step 522: Ablation experiment evaluation and comparison.

[0161] An ablation experiment can be performed based on the training sample data sets for different score segments obtained in the above steps.

[0162] Specifically, the server node can perform an ablation experiment on the second largest model for a preset number of rounds using the training sample data for the first score segment to obtain each first evaluation index value. Similarly, the training sample data for the second score segment can be used to perform an ablation experiment on the second largest model for a preset number of rounds to obtain each second evaluation index value; the training sample data for the third score segment can also be used to perform an ablation experiment on the second largest model for a preset number of rounds to obtain each third evaluation index value.

[0163] Furthermore, the target training samples for training the third largest model can be screened based on each evaluation index value. The specific ablation experiment process and sample screening process can refer to the foregoing embodiments and will not be elaborated here.

[0164] In practical applications, if the evaluation index value of a certain data set among multiple data sets is better than that of other data sets, then this certain data set can be determined as the target data set, the training sample data in this certain data set can be determined as the target training sample data, and the third largest model can be trained using the target training sample data to make up for the deficiency that the third largest model has insufficient learning of the training sample data in this target data set, and improve the model performance of the third largest model.

[0165] The above steps have similar or identical technical features to the method for determining training sample data described above, and reference can be made to the foregoing description, which will not be elaborated here one by one. It should be understood that in the embodiments of this specification, the order of some steps can be adjusted according to actual needs, or some steps can be omitted.

[0166] Through the above method, on the one hand, the first largest model can be used to perform distributed training on each piece of training sample data to obtain perplexity scores, and based on the perplexity scores, the training sample data can be divided at the sample-level granularity, and ablation experiments can be performed on the second largest model respectively based on the division results to improve the accuracy of the ablation experiment results.

[0167] On the second hand, according to the influence of data with different perplexity score segments on the performance of the large model, through multiple rounds of ablation experiments, sample data that the large model itself has not yet learned thoroughly or omitted can be newly discovered, or new ablation data that is more suitable for the large model's generalization reasoning and comprehensive improvement of capabilities in vertical fields can be constructed.

[0168] On the third hand, the quality of the text can be evaluated in combination with the training and inference framework of the base large model, so as to be able to comprehensively improve the model inference ability more pertinently according to the actual situation of the large model itself.

[0169] On the fourth hand, this specification provides a new dimensional perspective. In the subsequent process of pre-training data screening and use, data with higher PPL scores can be preferentially considered to improve the model training effect, newly discover sample data that the model itself has not yet learned thoroughly or omitted, and train the large model based on such samples to improve the model performance of the large model.

[0170] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 6 It is a schematic structural diagram of a device for determining training sample data provided by the embodiments of this specification. As Figure 6 shown, the device may include:

[0171] A sample acquisition module 602, configured to acquire a number of pieces of training sample data;

[0172] The score determination module 604 is configured to determine the perplexity scores of the several pieces of training sample data based on the first large model;

[0173] The set partitioning module 606 is configured to partition the several pieces of training sample data into at least two data sets based on the perplexity scores; the perplexity scores of the training sample data included in the first data set among the two data sets are greater than the perplexity scores of the training sample data included in the second data set;

[0174] The metric value determination module 608 is configured to perform ablation experiments on the second large model using each data set to determine the evaluation metric values of the second large model in each data set;

[0175] The target sample determination module 610 is configured to determine the training sample data included in the target data set with evaluation metric values superior to those of other data sets in each data set as the target training sample data for training the third large model; the scale of the third large model is larger than the scale of the first large model.

[0176] Based on Figure 6 For the apparatus, embodiments of this specification also provide some specific implementation schemes of this method, which will be described below.

[0177] Optionally, the score determination module may specifically be configured to: an inference unit, which is configured to perform training inference on the several pieces of training sample data using the first large model to obtain an inference result; a score determination unit, which is configured to determine the perplexity scores of the several pieces of training sample data based on the inference result.

[0178] Optionally, the score determination unit may specifically be configured to: for any piece of training sample data among the several pieces of training sample data, determine the perplexity scores of each word segment included in the any piece of training sample data; based on the perplexity scores of each word segment, determine the perplexity score of the any piece of training sample data.

[0179] Optionally, the apparatus may further include a distributed training inference module, which may specifically be configured to: perform serialization processing on the several pieces of training sample data to obtain several data sequences; perform packaging processing on the several data sequences to obtain a packaged sequence; perform batch partitioning on the packaged sequence to obtain several batch data; perform sharding processing on the several batch data to obtain several sharded data; respectively provide each of the sharded data to each first large model, and perform distributed training inference using each first large model.

[0180] Optionally, the device can also be used to: group the several data sequences into multiple sequence sets; the process of packing the several data sequences to obtain a packed sequence specifically includes: for any one of the multiple sequence sets, packing each data sequence in the any one sequence set to obtain a packed sequence.

[0181] Optionally, the score determination unit can specifically be used to: for any one batch of data among the several batches of data, determine the perplexity scores corresponding to the respective data sequences included in the any one batch of data; and determine the perplexity score corresponding to the any one batch of data according to the perplexity scores corresponding to the respective data sequences included in the any one batch of data.

[0182] Optionally, the device can be used to: determine the sequence indices corresponding to the respective data sequences in the any one batch of data; the sequence index represents the position information of the respective data sequences included in the any one batch of data; and save the first correspondence between the sequence index and the perplexity score corresponding to the any one batch of data.

[0183] Optionally, the evaluation metric values include multiple evaluation metric values generated through multiple different rounds of iteration; the target sample determination module can specifically be used to: if the number of evaluation metric values of the first data set that are better than those of the second data set is greater than or equal to a preset threshold, determine the training sample data included in the first data set as the target training sample data.

[0184] Optionally, the evaluation metric values include multiple evaluation metric values generated through multiple different rounds of iteration; the target sample determination module can specifically be used to: if each evaluation metric value generated by the preset consecutive rounds of iteration of the first data set is better than each evaluation metric value generated by the second data set corresponding to the preset consecutive rounds of iteration, determine the training sample data included in the first data set as the target training sample data.

[0185] Optionally, the device further includes a data preprocessing module, which can specifically be used to: obtain several pieces of initial training sample data; perform data preprocessing on the initial training sample data to obtain the several pieces of training sample data; the preprocessing includes at least one of a process for removing data noise, a process for correcting grammar or semantics, and a process for removing sensitive words.

[0186] Optionally, the metric value determination module can specifically be used to: for any one of the respective data sets, select a preset number of training sample data from the any one data set; and perform an ablation experiment on the second largest model by using the preset number of training sample data.

[0187] Optionally, the device can also be used to: save the first correspondence between the sequence index and the perplexity score corresponding to any batch of data to the distributed file system CPFS.

[0188] Optionally, the device can also be used to: save the first correspondence saved to the distributed file system CPFS to the object storage service system OSS.

[0189] Optionally, the device can also be used to: save the second correspondence between the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data.

[0190] Optionally, the device can also be used to: save the second correspondence between the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data to the open data processing service ODPS.

[0191] Optionally, the device can also be used to: synchronously save the second correspondence to the distributed file system CPFS.

[0192] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.

[0193] Figure 7 It is a schematic structural diagram of a device for determining training sample data provided by the embodiments of this specification. As Figure 7 shown, the device 700 may include: at least one processor 710; and a memory 730 communicatively connected to the at least one processor; wherein, the memory 730 stores instructions 720 executable by the at least one processor 710, and the instructions are executed by the at least one processor 710 to enable the at least one processor 710 to: obtain a number of training sample data; based on a first large model, determine the perplexity scores of the number of training sample data; based on the perplexity scores, divide the number of training sample data into at least two data sets; the perplexity scores of the training sample data included in the first data set in the two data sets are greater than the perplexity scores of the training sample data included in the second data set; use each data set to perform an ablation experiment on a second large model to determine the evaluation index values of the second large model in each data set; determine the training sample data included in the target data set whose evaluation index value in each data set is better than the evaluation index values of other data sets as the target training sample data for training a third large model; the scale of the third large model is greater than the scale of the first large model.

[0194] Based on the same idea, the embodiments of this specification also provide a computer-readable medium corresponding to the above method. Computer-readable instructions are stored on the computer-readable medium, and the computer-readable instructions can be executed by a processor to implement the method of the device for determining training sample data as described above.

[0195] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, Figure 7 for the device shown, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.

[0196] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards in the relevant regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0197] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0198] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function by logically programming the method steps so that the controller is in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0199] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices. For the convenience of description, when describing the above devices, the functions are divided into various units and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0200] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0201] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks. In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0202] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0203] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0204] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0205] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0206] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for determining training sample data, comprising: Obtaining a number of training sample data; Based on a first large model, determining the perplexity scores of the number of training sample data; Based on the perplexity scores, dividing the number of training sample data into at least two data sets; the perplexity scores of the training sample data included in the first data set in the two data sets are greater than the perplexity scores of the training sample data included in the second data set; Using each data set to perform ablation experiments on a second large model, and determining the evaluation index values of the second large model in each data set; Determining the training sample data included in the target data set whose evaluation index value in each data set is better than the evaluation index values of other data sets as the target training sample data for training a third large model; the scale of the third large model is larger than the scale of the first large model.

2. The method according to claim 1, wherein the determining the perplexity scores of the number of training sample data based on the first large model specifically comprises: Using the first large model to perform training inference on the number of training sample data to obtain an inference result; Based on the inference result, determining the perplexity scores of the number of training sample data.

3. The method according to claim 2, wherein the determining the perplexity scores of the number of training sample data specifically comprises: For any one of the number of training sample data, determining the perplexity scores of each word segment included in the any one of the training sample data; Based on the perplexity scores of each word segment, determining the perplexity score of the any one of the training sample data.

4. The method according to claim 1, further comprising: Performing serialization processing on the number of training sample data to obtain a number of data sequences; Performing packaging processing on the number of data sequences to obtain a packaged sequence; Performing batch division on the packaged sequence to obtain a number of batch data; Performing sharding processing on the number of batch data to obtain a number of sharded data; Providing each of the sharded data to each first large model respectively, and performing distributed training inference using each first large model.

5. The method according to claim 3, further comprising: Grouping the number of data sequences into multiple sequence sets; The performing packaging processing on the number of data sequences to obtain a packaged sequence specifically comprises: For any one of the multiple sequence sets, performing packaging processing on each data sequence in the any one of the sequence sets to obtain a packaged sequence.

6. The method according to claim 3, wherein the determining the perplexity scores of the number of training sample data specifically comprises: For any one of the number of batch data, determining the perplexity scores corresponding to each data sequence included in the any one of the batch data; According to the perplexity scores corresponding to each data sequence included in the any one of the batch data, determining the perplexity score corresponding to the any one of the batch data.

7. The method according to claim 6, further comprising: Determine the sequence index corresponding to each data sequence in any batch of data; The sequence index represents the position information of each data sequence included in any batch of data; Save the first correspondence between the sequence index and the perplexity score corresponding to any batch of data.

8. The method according to claim 1, wherein the evaluation index values include multiple evaluation index values generated through multiple different rounds of iteration; the step of determining the training sample data included in the target data set with evaluation index values superior to those of other data sets in each data set as the target training sample data for training the third large language model specifically includes: If the number of evaluation index values of the first data set that are superior to those of the second data set is greater than or equal to a preset threshold, then determine the training sample data included in the first data set as the target training sample data.

9. The method according to claim 1, wherein the evaluation index values include multiple evaluation index values generated through multiple different rounds of iteration; the step of determining the training sample data included in the target data set with evaluation index values superior to those of other data sets in each data set as the target training sample data for training the third large language model specifically includes: If each evaluation index value generated by the preset consecutive rounds of iteration of the first data set is superior to each evaluation index value generated by the preset consecutive rounds of iteration of the second data set, then determine the training sample data included in the first data set as the target training sample data.

10. The method according to claim 1, further comprising: Obtain a number of initial training sample data; Perform data preprocessing on the initial training sample data to obtain the number of training sample data; The preprocessing includes at least one of a process for removing data noise, a process for correcting grammar or semantics, and a process for removing sensitive words.

11. The method according to claim 1, wherein the ablation experiment on the second large model using each data set specifically includes: For any one of the data sets in each data set, select a preset number of training sample data from the any one of the data sets; Perform an ablation experiment on the second large model using the preset number of training sample data.

12. The method according to claim 7, wherein the step of saving the first correspondence between the sequence index and the perplexity score corresponding to any batch of data specifically includes: Save the first correspondence between the sequence index and the perplexity score corresponding to any batch of data to the distributed file system CPFS.

13. The method according to claim 12, further comprising: Save the first correspondence saved to the distributed file system CPFS to the object storage service system OSS.

14. The method according to claim 7, further comprising: Save the second correspondence between the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data.

15. The method according to claim 14, wherein the step of saving the second corresponding relationship between the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data specifically includes: Saving the second corresponding relationship between the sequence index, the training sample data corresponding to the index, and the perplexity score corresponding to any batch of data to the Open Data Processing Service ODPS.

16. The method according to claim 15, further comprising: Synchronously saving the second corresponding relationship to the Distributed File System CPFS.

17. An apparatus for determining training sample data, comprising: A sample acquisition module, configured to acquire a plurality of pieces of training sample data; A score determination module, configured to determine the perplexity scores of the plurality of pieces of training sample data based on a first large model; A set partitioning module, configured to partition the plurality of pieces of training sample data into at least two data sets based on the perplexity scores; the perplexity scores of the training sample data included in the first data set among the two data sets are greater than the perplexity scores of the training sample data included in the second data set; An index value determination module, configured to perform ablation experiments on a second large model using each data set, and determine the evaluation index values of the second large model in each data set; A target sample determination module, configured to determine the training sample data included in the target data set whose evaluation index value is better than the evaluation index values of other data sets as the target training sample data for training a third large model; the scale of the third large model is greater than the scale of the first large model.

18. An apparatus for determining training sample data, comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: Acquire a plurality of pieces of training sample data; Determine the perplexity scores of the plurality of pieces of training sample data based on a first large model; Partition the plurality of pieces of training sample data into at least two data sets based on the perplexity scores; the perplexity scores of the training sample data included in the first data set among the two data sets are greater than the perplexity scores of the training sample data included in the second data set; Perform ablation experiments on a second large model using each data set, and determine the evaluation index values of the second large model in each data set; Determine the training sample data included in the target data set whose evaluation index value is better than the evaluation index values of other data sets as the target training sample data for training a third large model; the scale of the third large model is greater than the scale of the first large model.

19. A computer-readable medium, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement the method for determining training sample data according to any one of claims 1 to 16.

Citation Information

Cited By

  • Screening method and equipment of instruction data, medium and product

    CN120653995A

  • Network security alarm research and judgment model training method and device

    CN121418127A