Training data acquisition method, system and apparatus, storage medium and program

By evaluating the initial data set with multiple processing methods and training for small models, the optimal processing method is determined, which solves the problem of large-scale training data and low efficiency, and achieves efficient and low-loss training data acquisition.

WO2025139738A1PCT designated stage expired Publication Date: 2025-07-03SHANGHAI XIYU JIZHI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/137908
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-30
Filing Date
2024-12-09
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Large model training requires large amount of data and low processing efficiency. The prior art can easily lead to useful information loss and resource consumption during data processing.

Method used

By obtaining the initial data set, the data is processed using multiple candidate processing methods to obtain multiple data sets, and these processing methods are trained and evaluated using small machine learning models to determine the optimal processing method for acquisition of large-model training data.

Benefits of technology

It reduces the loss of useful information in the training data of large models, improves data processing efficiency, reduces resource consumption, and ensures that large models achieve better performance with fewer training steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137908_03072025_PF_FP_ABST
    Figure CN2024137908_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a training data acquisition method, system and apparatus, a storage medium and a program. The training data acquisition method comprises: acquiring an initial data set having data of the same type, and using a plurality of candidate processing modes to process the data in the initial data set so as to obtain a plurality of data sets; on the basis of the plurality of data sets, training a small model to obtain a trained small model; and, on the basis of the trained small model, evaluating the plurality of candidate processing modes to determine a target processing mode, the target processing mode being configured to process pre-trained data so as to obtain training data of a large model, and the size of the small model being smaller than that of the large model. The present disclosure uses the different candidate processing modes to obtain the data sets for small model training, and evaluates the performance of the small model to select an optimal data processing mode for determining the training data for the large model, thereby obtaining training data which is easy for the large model to learn and remains intact, and allowing the large model to achieve optimal performance with fewer training steps.
Need to check novelty before this filing date? Find Prior Art

Description

A training data acquisition method, system, device, storage medium and program

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to Chinese patent application number 2023118596009, filed with the Chinese Patent Office on December 30, 2023, entitled “A training data acquisition method, system, device, storage medium and program,” the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0003] The present disclosure relates to the field of artificial intelligence, and in particular to a method, system, device, storage medium, and program for acquiring training data, as well as a distributed system for acquiring training data. Background Art

[0004] The amount of data required for large model training is very large. Before training a large model, the raw data usually needs to be processed to obtain training data. In the process of processing the raw data to obtain training data, information in the raw data that is useful for large model training may be damaged. For example, different types of data may be suitable for different processing methods. When the raw data is processed using an inappropriate processing method, the generated training samples will be unfavorable for large model training. For example, information useful for large model training will be damaged, thereby affecting the accuracy or quality of the large model. In addition, due to the large amount of data required for large model training, processing the raw data into training data requires a large amount of computing power, for example, more than one thousand cores, and actually about four to five thousand cores, in order to process the raw data into the required training data as quickly as possible, which leads to low data processing efficiency and high resource consumption.

[0005] Therefore, the present invention proposes a training data acquisition method that can determine the optimal processing method for different types of data, so as to process the original data into training data that is easy for large models to learn, reduce the loss of useful information in the data, and at the same time reduce resource consumption. Summary of the Invention

[0006] One or more embodiments of the present disclosure provide a method for acquiring training data. The method comprises: acquiring an initial data set, the initial data set including data of the same type; processing data in the initial data set using multiple candidate processing methods to obtain multiple data sets; training a first machine learning model based on the multiple data sets to obtain a trained first machine learning model; and evaluating the multiple candidate processing methods based on the trained first machine learning model to determine a target processing method, the target processing method being configured to process pre-trained data to obtain training data for a second machine learning model, wherein the size of the first machine learning model is smaller than the size of the second machine learning model.

[0007] One or more embodiments of the present disclosure provide a distributed system for acquiring training data, the distributed system comprising a computing cluster comprising a plurality of computing devices; a manager configured to acquire raw data, classify the raw data based on data types to obtain a plurality of initial data sets; and send the plurality of initial data sets to the plurality of computing devices, respectively, wherein any computing device among the plurality of computing devices is configured to execute a training data acquisition method based on the received initial data sets.

[0008] One or more embodiments of the present disclosure provide a training data acquisition system, including an acquisition module, a processing module, a training module and an evaluation module; the acquisition module is configured to acquire an initial data set, and the initial data set includes data of the same type; the processing module is configured to use multiple candidate processing methods to process the data in the initial data set to obtain multiple data sets, and each data set in the multiple data sets corresponds to one candidate processing method among the multiple candidate processing methods; the training module is configured to train a first machine learning model based on each data set to obtain a trained first machine learning model; and the evaluation module is configured to use the trained first machine model to evaluate the multiple candidate processing methods to determine a target processing method, and the target processing method is configured to process pre-trained data to obtain training data for a second machine learning model, wherein the size of the first machine learning model is smaller than the size of the second machine learning model.

[0009] One or more embodiments of the present disclosure provide a training data acquisition device, including a processor, wherein the processor is configured to execute a training data acquisition method.

[0010] One or more embodiments of the present disclosure provide a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the training data acquisition method.

[0011] One or more embodiments of the present disclosure provide a computer program product, including a computer program or computer-executable instructions, wherein when the computer program or computer-executable instructions are executed by a processor, a training data acquisition method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present disclosure will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures.

[0013] FIG1 is an exemplary flowchart of a method for acquiring training data according to some embodiments of the present disclosure.

[0014] FIG2 is an exemplary flowchart of a method for acquiring training data according to some embodiments of the present disclosure.

[0015] FIG3 is an exemplary flowchart of a method for acquiring training data according to some embodiments of the present disclosure.

[0016] FIG4 is a schematic diagram of a method for acquiring training data according to some embodiments of the present disclosure.

[0017] FIG5 is an exemplary structural diagram of a training data acquisition system according to some embodiments of the present disclosure.

[0018] FIG6 is an exemplary structural diagram of a distributed processing system for acquiring training data according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of the present disclosure. Those skilled in the art can apply the present disclosure to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0020] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.

[0021] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0022] Flowcharts are used in this disclosure to illustrate the operations performed by systems according to embodiments of the present disclosure. It should be understood that the preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0023] FIG1 is an exemplary flow chart of a training data acquisition method according to some embodiments of the present disclosure. In some embodiments, process 100 may be performed by a processing device. In some embodiments, process 100 may be performed by the training method acquisition system shown in FIG5 or the distributed system shown in FIG6.

[0024] Step 110: Acquire an initial data set, which includes data of the same type.

[0025] The initial data set refers to the data that has been divided into different types.

[0026] In some embodiments, raw data (also referred to as raw data packets) may be obtained. The raw data may be preliminarily classified to obtain a plurality of different types of first data sets, and any one of the plurality of different types of first data sets may be designated as an initial data set. For example, the raw data may be classified into a numerical data set, a categorical data set, a binary data set, a temporal data set, a text data set, an image data set, an audio data set, a video data set, etc.

[0027] In some embodiments, the first dataset can be further sub-classified to obtain multiple different types of second datasets. Any one of the multiple different types of second datasets can be designated as the initial dataset. For example, a text dataset can be classified into a webpage dataset, a document dataset, a forum dataset, a topic dataset, etc.

[0028] In some embodiments, the second data set may be further classified to obtain a plurality of third data sets of different types, and one of the plurality of third data sets of different types may be designated as the initial data set.

[0029] For example, the second data set can be further classified according to data format to obtain multiple third data sets of different types. For example, the web page data in the web page data set can be divided into Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript, JavaScript Object Notation (JSON), Extensible Markup Language (XML), and multimedia files such as images, videos, and audio.

[0030] For another example, the second dataset can be further categorized based on data content to obtain multiple third datasets of different types. For example, a document dataset can be categorized into scientific and technological document datasets, literary and fiction document datasets, educational document datasets, financial document datasets, and life guide document datasets. Another example is categorized into PDF documents, Word documents, and Text documents based on data storage format.

[0031] For another example, the second dataset can be further categorized based on the data's location to obtain multiple different types of third datasets. For example, forum datasets can be categorized into the Asia-Pacific Forum dataset, the European Forum dataset, the North American Forum dataset, the South American Forum dataset, and so on.

[0032] For another example, the second dataset can be further classified according to the data representation to obtain multiple different types of third datasets, such as dividing the question dataset into true / false question dataset, single-choice question dataset, multiple-choice question dataset, and text answer question dataset.

[0033] Step 120: Process the data in the initial dataset using multiple candidate processing methods to obtain multiple datasets. The data in the same initial dataset is processed using multiple candidate processing methods to obtain multiple datasets. Each of the multiple datasets corresponds to a candidate processing method, and each dataset is obtained by processing the initial dataset using the corresponding candidate processing method.

[0034] Among them, data processing methods include one or more combinations of data cleaning, data splicing, data conversion, feature selection, data reorganization, data labeling, data enhancement, etc.

[0035] Data cleaning refers to removing errors, incomplete or inconsistent parts in the data to improve the quality and accuracy of the data. Data cleaning can include missing value processing, outlier processing, data deduplication, format unification, data integration, data dimensionality reduction, etc. In some embodiments, the initial data set can be a web page data set, which includes multiple web page data, and data cleaning processing can be performed on the web page data. There are a large number of hyperlinks (URLs) in the web page data, some of which are useful data, and some are useless data (such as hyperlinks of advertisements, etc.). In addition, there is generally a lot of redundant information on the web page, such as the header and footer of the web page will have repeated hyperlink data, and the only valid data information is the title and text in the web page (for example, news headlines and text). Therefore, the web page data can be deduplicated to extract the useful hyperlink data information, and the invalid, repeated and interfering data information can be removed.

[0036] Data concatenation can include row concatenation, column concatenation, merging, and the like. In some embodiments, the initial dataset can include a dataset of PDF documents, and PDF documents with the same subject matter can be merged. In some embodiments, the data concatenation process includes the use of data concatenation tools, such as the concat and merge functions in the Pandas library and the JOIN operation in SQL. Using different tools can correspond to different data processing methods.

[0037] Data conversion may include data type conversion, data format conversion, data normalization, data discretization, data transformation (e.g., logarithmic transformation, exponential transformation, etc.). In some embodiments, the initial data set may include a PDF document data set, and the PDF document may be converted into a text document using a data format conversion process. In some embodiments, the data processing process may include a conversion technology corresponding to the data conversion. For example, the data processing process may include using a model (e.g., a trained machine learning model) or a tool (e.g., Adobe Acrobat Pro, PDFMiner, etc.) to convert a PDF document into a text document.

[0038] Step 130: Based on multiple data sets, train the first machine learning model to obtain a trained first machine learning model.

[0039] As described herein, the first machine learning model is a machine learning model having a parameter count less than or equal to a first threshold. The first threshold can be any value less than or equal to 100 megabytes, such as 90 megabytes. For example, the first machine learning model has a parameter count of 100,000, 1 megabyte, 10 megabytes, or 100 megabytes. The first machine learning model can also be referred to as a small model.

[0040] In some embodiments, the first machine learning model may include a neural network model, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), etc. In some embodiments, the first machine learning model may be an unsupervised learning model. For example, the first machine learning model may include an autoencoder model, a generative latent semantic analysis (LSA) model, a hidden Markov model (HMM), an ensemble clustering model, a variational autoencoder (VAE) model, etc.

[0041] In some embodiments, multiple data sets can be input into the same first machine learning model to train the first machine learning model. For example, the training data in multiple data sets can be input into the first machine learning model in sequence to obtain the trained first machine learning model. Optionally, after the training data in one of the data sets are input into the first machine learning model, the training data in the next data set can be input for training until the training data in all data sets are used to train the first machine learning model. For another example, the training data in each data set can be divided into multiple data sets. Then, one piece of training data in each data set is input into the first machine learning model in sequence to train the first machine learning model. Then, another piece of training data in each data set is input in sequence until the first machine learning model converges, completing the training of the first machine learning model.

[0042] In some embodiments, the first machine learning model may include multiple sub-models. The multiple sub-models may be independent machine learning models or integrated into the first machine learning model. The multiple sub-models may be the same machine learning model. The same machine learning model means that the model structure and size are the same. The multiple sub-models may each have their own input layer. The training data in the multiple data sets may be input into the multiple sub-models for training. Each sub-model corresponds to a data set, and the sub-model is trained based on the corresponding data set to obtain a trained sub-model. For example, the multiple data sets include data set S1, data set S2, ...., data set Sn, and the sub-models include sub-model M1, sub-model M2, ..., sub-model Mn. Data set S1 is used to train sub-model M1 to obtain trained sub-model M1, data set S2 is used to train sub-model M2 to obtain trained sub-model M2, and so on. Data set Sn is used to train sub-model Mn to obtain trained sub-model Mn. Based on this, different sub-models can correspond to a data set and also correspond to a data processing method, that is, each sub-model can be trained by processing the initial data set using one of multiple candidate processing methods.

[0043] In some embodiments, the data in each dataset can be divided into training data and test data (or validation data). The training data is used to train the first machine learning model, and the test data or validation data can be used to test or validate the trained first machine learning model. For more information on testing the trained first machine learning model, please refer to the subsequent description.

[0044] In some embodiments, the first machine learning model may be trained using a model training method. Typical model training methods may include a gradient descent algorithm, a Bayesian optimization algorithm, an adaptive learning rate method, and the like.

[0045] Step 140, evaluate multiple candidate processing methods based on the trained first machine learning model to determine a target processing method. The target processing method is configured to process the pre-training data to obtain training data for the second machine learning model, wherein the size of the first machine learning model is smaller than the size of the second machine learning model. In some embodiments, the size of the model can be defined by the number of parameters of the model. The smaller size of the first machine learning model than the size of the second machine learning model can also be referred to as the smaller number of parameters of the first machine learning model than the number of parameters of the second machine learning model. In some embodiments, the size of the model can also be defined by the required computing power of the model. The greater the required computing power, the larger the model size. The type of the pre-training data is the same as the data type in the initial data set. For example, if the initial data set includes documents in PDF format, the pre-training data can be documents in PDF format.

[0046] As described herein, the second machine learning model is a machine learning model having a parameter count greater than a second threshold. The second threshold can be any value greater than 100 megabytes, such as 1000 megabytes. For example, the second machine learning model has a parameter count of 1000 megabytes. The second machine learning model can also be referred to as a large model.

[0047] In some embodiments, the second machine learning model may include a neural network model, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), etc. In some embodiments, the second machine learning model may be an unsupervised learning model. For example, the second machine learning model may include an autoencoder model, a generative latent semantic analysis (LSA) model, a hidden Markov model (HMM), an ensemble clustering model, a variational autoencoder (VAE) model, etc.

[0048] In some embodiments, evaluating multiple candidate processing methods based on the trained first machine learning model to determine the target processing method includes: obtaining the performance evaluation results generated during the training of the first machine learning model based on each data set in the multiple data sets. For example, each data set can be input into the first machine learning model, and the first machine learning model can generate a prediction result, and the prediction result is statistically analyzed to determine the performance evaluation result corresponding to each data set. The final target processing method is determined by comparing the performance evaluation results corresponding to each data set. For example, the test data in the multiple data sets are respectively input into the first machine learning model for training to generate multiple sets of output data (i.e., multiple sets of prediction result sets). For example, the multiple data sets may include data set S1, data set S2, ..., data set Sn, and the test data in data set S1, data set S2, ..., data set Sn are sequentially input into the trained first machine learning model to generate prediction result set D1, prediction result set D2, ..., prediction result set Dn. The prediction results in prediction result set D1, prediction result set D2, ..., prediction result set Dn can be statistically calculated to determine the values ​​V1, V2, ..., Vn of the evaluation index (e.g., cross entropy). From the values ​​V1, V2, ..., Vn of the evaluation index (e.g., cross entropy), a value that satisfies a preset condition (e.g., a maximum value, a minimum value, a value less than a preset threshold value, a value greater than a preset threshold value, or a value within a preset interval) is selected. The candidate processing method corresponding to the value of the evaluation index that satisfies the preset condition is designated as the target processing method. The performance evaluation result of the first neural network model under the training of the data set obtained by the target processing method satisfies the preset condition.

[0049] In some embodiments, evaluating multiple candidate processing methods based on the trained first machine learning model to determine a target processing method includes: obtaining test data for each of multiple data sets, processing the test data in each data set using the trained first machine learning model to obtain a prediction result, determining a performance evaluation result (e.g., a performance parameter value) of the first machine learning model under training on a data set obtained by processing an initial data set by each of the multiple candidate processing methods based on the prediction result, and determining a target processing method from the multiple candidate processing methods based on the performance evaluation result. The performance evaluation result of the trained first machine learning model under the test data obtained by the target processing method meets the preset conditions. For more description of the method for determining the target processing method, please refer to the detailed description of Figure 2.

[0050] In some embodiments, the first machine learning model includes multiple sub-models of different sizes, and based on each data set in the multiple data sets, training the first machine learning model to obtain the trained first machine learning model includes: using the same training data in each data set to train the multiple sub-models respectively to obtain multiple trained sub-models. Using the trained first machine model to evaluate multiple candidate processing methods to determine the target processing method includes: obtaining the performance evaluation results of the multiple sub-models under the training of the data set obtained by each candidate processing method in the multiple candidate processing methods, each candidate processing method corresponds to a performance evaluation result, predicting the performance evaluation results of the second machine learning model under the training data obtained by each candidate processing method based on the performance evaluation results of the multiple sub-models, determining the target processing method based on the performance evaluation results of the second machine learning model under the training data obtained by each candidate processing method, and the performance evaluation results of the second machine learning model under the training of the data set obtained by the target processing method meet the preset conditions.

[0051] Predicting the performance evaluation results of the second machine learning model under the training data obtained by each candidate processing method based on the performance evaluation results of multiple sub-models includes: fitting the performance evaluation results of multiple sub-models under the data set training obtained by each candidate processing method to generate a correspondence between the model parameters and performance parameters under the data set training obtained by each candidate processing method; and determining the performance evaluation results of the second machine learning model under the training data obtained by each candidate processing method based on the correspondence between the model parameters and performance parameters under the data set training obtained by each candidate processing method.

[0052] Determining the performance evaluation results of the second machine learning model under training on the training data obtained by each candidate processing method based on the correspondence between the model parameters and the performance parameters includes: obtaining the model parameters of the second machine learning model, the model parameters including at least one of the structural parameters and training parameters of the second machine learning model; determining the estimated performance parameters of the second machine learning model under training on the training data obtained by each candidate processing method based on the correspondence between the model parameters and the performance parameters and the model parameters of the second machine learning model.

[0053] In some embodiments, there are multiple initial data sets, each containing different types of data. The processing device may perform steps 110-140 on each initial data set to determine a target processing method (also referred to as an optimal processing method) for each type of data. In some embodiments, different types of data may be processed differently. In some embodiments, different types of data may have the same target processing method.

[0054] In some embodiments of the present invention, different processing methods are used to process the same type of initial data set to obtain multiple data sets, a small model is trained based on the multiple data sets to obtain a trained small model, and different processing methods are evaluated based on the performance evaluation results of the small model to select a data processing method suitable for large model training. Based on this, the optimal data processing method can be selected to reduce the damage rate of useful information in the process of obtaining large model training data, and the training data most suitable for large model training can be obtained, so that the large model can achieve better performance with fewer training steps. At the same time, evaluating candidate processing methods through small models can reduce resource consumption and improve data processing efficiency compared to directly using large models to evaluate candidate processing methods.

[0055] FIG2 is an exemplary flowchart of training data acquisition according to some embodiments of the present disclosure. As shown in FIG2 , process 200 includes the following steps. In some embodiments, process 200 may be executed by the training data acquisition system (e.g., evaluation module 540) shown in FIG5 or by some components (e.g., computing device 610-1, computing device 610-2, ..., or computing device 610-N) in distributed processing system 600 shown in FIG6 .

[0056] Step 210: Acquire test data from each of the multiple data sets.

[0057] As described in this article, the test data can be configured to evaluate the performance of the trained model and can also be called a test set.

[0058] The data type in each data set is the same. Each data set is obtained by processing the initial data set by one of multiple candidate processing methods. For example, multiple data sets may include data set S1, data set S2, ...., data set Sn, and the candidate processing methods may include candidate processing method P1, candidate processing method P2, ...., candidate processing method Pn, which are respectively configured to process the same initial data set to obtain data set S1, data set S2, ...., data set Sn. For more description of the initial data set and the candidate processing methods, please refer to steps 110 and 120 in Figure 1. For example, the data types in the initial data sets are the same, and the initial data set can be obtained by classifying the original data (for example, the original data packet). For another example, the candidate processing methods include data cleaning, data conversion, data splicing, data conversion, feature selection, data reorganization, data labeling, data enhancement, etc.

[0059] In some embodiments, each data set can be divided into training data (i.e., training set) and test data (i.e., test set). For example, before training a first machine learning model, part of the data in each data set can be used as test data, and another part of the data can be used as training data. The training data can be used to train the first machine learning model. The test data can be stored in a storage device. The training data is different from the test data.

[0060] In some embodiments, a certain number or proportion of data can be randomly selected from each data set to form a test set. For example, 100 pieces of data can be randomly selected from 10,000 pieces of data as the test set of the data set. The test set is independently stored in a specific storage space and does not participate in any model training, thereby ensuring the independence of the test data set. It also ensures that the first machine learning model has no training memory and training traces of the test data set, thereby improving the accuracy of the verification results.

[0061] In some embodiments, the test set can be backed up to prevent contamination of the test set by model training. For example, dedicated storage space can be allocated on the server with read-only permissions to ensure that the test set is free of data contamination that could distort the model performance evaluation results.

[0062] Step 220: Process the test data in each dataset using the trained first machine learning model to obtain a prediction result. The prediction result includes output data generated by the trained first machine learning model after processing the test data in each dataset. The output data generated by processing the test set in the same dataset by the trained first machine learning model can constitute a prediction result set corresponding to each dataset (or test set).

[0063] Using the trained first machine learning model to process the test data in each data set means taking each test data as the input of the trained first machine learning model to obtain the output data of the trained first machine model. The test data in each data set is input into the trained first machine learning model to generate a set of output data (i.e., a set of prediction results). The test data in multiple data sets are respectively input into the trained first machine learning model to generate multiple sets of output data (i.e., multiple sets of prediction results). For example, multiple data sets may include data set S1, data set S2, ...., data set Sn, and the test data in data set S1, data set S2, ...., data set Sn are sequentially input into the trained first machine learning model to generate prediction result set D1, prediction result set D2, ...., prediction result set Dn.

[0064] Step 230: Determine, based on the prediction results, the performance evaluation results of the trained first machine learning model under the training of the data set obtained under each candidate processing method.

[0065] In some embodiments, the performance of the trained first machine learning model can be evaluated based on the prediction results to obtain a performance evaluation result of the trained first machine learning model. The performance evaluation result may include the values ​​of one or more performance parameters or evaluation indicators.

[0066] Performance evaluation indicators or performance parameters may include accuracy, precision, recall, F1 score, AUC (Area Under the Curve), mean squared error (MSE), R2 score, cross-entropy, etc.

[0067] For example, a statistical calculation can be performed on the prediction results in each prediction result set to determine the value of the performance parameter of the trained first machine learning model corresponding to the prediction result set. Alternatively, a statistical calculation can be performed on the prediction results in prediction result set D1, prediction result set D2, ...., and prediction result set Dn to determine the value V1, V2, ...., Vn of at least one evaluation indicator (e.g., cross entropy).

[0068] Step 240 : Determine a target processing method from multiple candidate processing methods based on the performance evaluation result.

[0069] Based on the above description, each data set is obtained by processing the initial data set by one of the multiple candidate processing methods. It can be understood that each data set corresponds to a candidate processing method. Each prediction result set is obtained by processing the test data in a data set by the trained first machine learning model. It can be understood that each prediction result set corresponds to a data set, that is, it also corresponds to a candidate processing method. Each value of the performance evaluation index is obtained by counting the prediction results in a prediction result set. Therefore, it can be understood that the value of each evaluation index can correspond to a prediction result set, that is, it also corresponds to a data set and a candidate processing method. The value of the performance evaluation index can be used to measure the performance of the first machine learning model under the training of the data set obtained by different candidate processing methods. For example, the evaluation index value V1 represents the performance of the first machine learning model under the training of the data set S1 obtained by processing the initial data set by the candidate processing method P1; the evaluation index value V2 represents the performance of the first machine learning model under the training of the data set S2 obtained by processing the initial data set by the candidate processing method P2; and so on. The evaluation index value Vn represents the performance of the first machine learning model under the training of the data set Sn obtained by processing the initial data set by the candidate processing method Pn.

[0070] Therefore, the values ​​of the performance evaluation indicators corresponding to multiple candidate processing methods can be compared, and the values ​​of all performance evaluation indicators that meet the preset conditions can be determined. The candidate processing method corresponding to the performance evaluation indicator value that meets the preset conditions is determined as the target processing method, thereby representing that the performance parameters (i.e., evaluation indicators) of the first machine learning model under the evaluation of the test data obtained by the target processing method meet the preset conditions. The preset condition can be that the value of the performance evaluation indicator is the maximum value or the minimum value of all performance evaluation indicator values, or is greater than a preset threshold, or is less than a preset threshold, or is within a preset range. For example, when the performance evaluation indicator is cross entropy, the candidate processing method corresponding to the maximum value of all cross entropy values ​​is the target processing method.

[0071] In this embodiment, the degree of adaptability of each candidate processing method to the initial data set is determined by comparing performance parameters to determine the optimal processing method for that type of data. For example, a sentence is input to a machine learning model, and based on the context and corpus, it associates or infers the next sentence and outputs it. By comparing the cross-entropy values ​​of the machine learning model's output data, the candidate processing method corresponding to the test data with the lowest cross-entropy value is determined as the optimal processing method for the current test set.

[0072] In some embodiments, steps 210-240 may be performed for each type of data set to determine a target processing method (also referred to as an optimal processing method) for each type of data. In some embodiments, the target processing methods may be different for different types of data. In some embodiments, the target processing methods may be the same for different types of data.

[0073] In some embodiments, the first machine learning model trained in step 220 may include multiple trained sub-models. The sub-models have the same size. Each trained sub-model is obtained by training a sub-model with training data from one of the multiple data sets. Step 230 may include inputting the test data from the multiple data sets into the corresponding trained sub-models to obtain the prediction results of each trained sub-model. Then, step 240 is performed. For example, the test data from data sets S1, S2, ...., and Sn may be input into the trained sub-model M1, the trained sub-model M2, ...., and the trained sub-model Mn to obtain prediction result sets D1, D2, ..., and Dn. The prediction result sets D1, D2, ..., and Dn correspond to the data sets S1, S2, ...., and Sn, respectively. The data sets S1, S2, ...., and Sn correspond to the candidate processing methods P1, P2, ..., and Pn, respectively. Performance evaluation index (e.g., cross entropy) values ​​V1, V2, ..., and Vn of the trained sub-models can be determined based on prediction result set D1, prediction result set D2, ..., and prediction result set Dn, respectively. The performance evaluation index (e.g., cross entropy) values ​​V1, V2, ..., and Vn can be compared to determine an optimal performance evaluation index. A target prediction result set corresponding to the optimal performance evaluation index can be determined. A corresponding target data set can be determined based on the target prediction result set corresponding to the optimal performance parameter. The candidate processing method corresponding to the target data set is the target processing method.

[0074] By determining the optimal processing method for this type of data, we can determine the training data for the large model. This method makes the training data more suitable for large-scale model training, and the processing of the training data is less likely to cause information loss in the training data, thereby improving the quality of the training data for the large model. Because large models are more computationally and resource-intensive, using a small model, approximately 1 / 10 or even 1% of the size of the large model, to evaluate the quality of the data can reduce resource consumption.

[0075] FIG3 is an exemplary flowchart of training data acquisition according to some embodiments of the present disclosure. As shown in FIG3 , process 300 includes the following steps. In some embodiments, process 300 may be executed by the training data acquisition system (e.g., evaluation module 540) shown in FIG3 or by some components (e.g., computing device 610-1, computing device 610-2, ..., or computing device 610-N) in distributed processing system 600 in FIG6 .

[0076] In step 310, a plurality of sub-models of different sizes are trained respectively using the same training data in the same data set to obtain a plurality of trained sub-models. The training data may be at least part of the data in the data set. The data set may be obtained by processing the initial data set in one of a plurality of candidate methods. The initial data set includes data of the same type. For more descriptions of the initial data set and the candidate processing methods, see steps 110 and 120 in FIG1 . For example, the data types in the initial data set are the same, and the initial data set can be obtained by classifying the original data (e.g., the original data packet). For another example, the candidate processing methods include one or more combinations of data cleaning, data conversion, data splicing, data conversion, feature selection, data reorganization, data labeling, data enhancement, etc.

[0077] The multiple sub-models can be small models, that is, the multiple sub-models of different sizes can be machine learning models with a size of less than 100 megabytes (i.e., the number of parameters is less than 100 megabytes). For example, the multiple sub-models of different sizes can be machine learning models with sizes of 100K, 1 megabyte, 10 megabyte, 100 megabytes, etc. For another example, the multiple sub-models of different sizes can be machine learning models with sizes of 100K, 0.5 megabytes, 1.5 megabytes, 15 megabytes, etc.

[0078] The same training data is input into multiple sub-models of different sizes for training to obtain multiple trained sub-models. In some embodiments, sub-models of different sizes may have different numbers of training steps. The number of training steps can be determined based on the size of the sub-model. For example, the smaller the sub-model size, the larger the number of training steps. In some embodiments, sub-models of different sizes may have the same number of training steps. Model training can be performed with a fixed number of training steps or a variable number of training steps for sub-models of different sizes. The number of training steps refers to the number of times the model updates the parameters during the training process. The process in which the sub-model learns from the training data each time and updates the model parameters according to the loss function is called one-step training. In deep learning, optimization algorithms such as stochastic gradient descent (SGD) can be used to update parameters. A small batch of training data (mini-batch) is used each time to calculate the gradient and update the parameters. Such a parameter update is a one-step training.

[0079] Step 320 , respectively obtain performance evaluation results of trained sub-models of different sizes.

[0080] In some embodiments, for any sub-model, during the training process of the sub-model, at each step of training, the sub-model generates a prediction result based on a batch of input training data, and based on the prediction result, the sub-model in the training step is subjected to a performance evaluation to generate performance evaluation data (e.g., the value of a performance evaluation index), and the sub-model parameters are updated based on the performance evaluation data. For example, the performance of the sub-model can be evaluated using a cross-entropy loss function based on the prediction result generated by a batch of input training data, and a cross-entropy value (i.e., a statistical value) can be generated. For multiple training steps, the sub-model can generate multiple prediction results and multiple performance evaluation data (e.g., multiple values ​​of a performance evaluation index, such as a cross-entropy value). The multiple performance evaluation data are related to the number of training steps and the size of the sub-model. Linear fitting can be performed on the multiple performance evaluation data to obtain the relationship between the number of training steps, model size, and performance evaluation index (e.g., cross entropy). In some embodiments, the relationship between the number of training steps, model size, and performance evaluation index (e.g., cross entropy) corresponding to any sub-model can be understood as the performance evaluation result of the sub-model (e.g., a relationship curve between performance parameter-model size-training step number). In some embodiments, the number of training steps, model size, and performance evaluation data (e.g., cross entropy) corresponding to any sub-model may be collectively referred to as the performance evaluation result of the sub-model.

[0081] In some embodiments, for any sub-model, the performance of the trained sub-model can be evaluated using the test data in the data set used to train the sub-model to obtain a performance evaluation result. For example, for any trained sub-model, the test data is input into the trained sub-model to generate a prediction result (i.e., a prediction result set), and the performance of the trained sub-model is evaluated based on each prediction result in the prediction result set to generate performance evaluation data. The performance evaluation data, model size, and the number of training steps of the model constitute the performance evaluation result. For example, the prediction results in the prediction result set can be statistically analyzed to generate performance evaluation data (e.g., cross entropy value). For another example, the model performance can be evaluated based on the prediction results using indicators such as mean square error (MSE) and mean absolute error (MAE).

[0082] Step 330: predict the performance evaluation results of the second machine learning model under the training data obtained by each candidate processing method based on the performance evaluation results of the multiple sub-models.

[0083] Predicting the performance evaluation results of the second machine learning model under training on the training data obtained by each candidate processing method based on the performance evaluation results of multiple sub-models includes: fitting the performance evaluation results of multiple sub-models under training on the data set obtained by each candidate processing method to generate a correspondence between model parameters and performance parameters under training on the data set obtained by each candidate processing method; and determining the performance evaluation results of the second machine learning model under training on the training data obtained by each candidate processing method based on the correspondence between the model parameters and the performance parameters.

[0084] Model parameters may include model structure parameters (e.g., the total number of model parameters) and / or training parameters (e.g., the number of training steps). Performance parameters (performance evaluation metrics) may include statistics such as cross entropy, mean squared error (MSE), and mean absolute error (MAE).

[0085] The correspondence between model parameters and performance parameters can be expressed as a functional relationship between the model parameters and the performance parameters. In this functional relationship, the model parameters are independent variables, and the performance parameters are dependent variables. For example, the correspondence between model parameters and performance parameters can be a functional relationship between model size, number of training steps, and cross entropy. Model size and number of training steps are the independent variables of this functional relationship, and cross entropy is the dependent variable of this functional relationship.

[0086] Based on step 320, multiple performance evaluation results of multiple sub-models of different sizes trained on the same training data can be obtained. Each size of the sub-model can correspond to a performance evaluation result. For example, the performance evaluation result can be expressed as the value of the performance evaluation index of the sub-model of a specific size under different training steps. By fitting the performance of sub-models of different sizes under different training steps, the corresponding relationship between the model size, training steps and performance parameters can be obtained. For another example, the performance evaluation result can be expressed as the relationship between the number of training steps, model size and performance evaluation index (e.g., cross entropy) corresponding to any sub-model (e.g., a relationship curve between performance parameters-model size-training steps). By fitting the relationship between the number of training steps, model size and performance evaluation index (e.g., cross entropy) of sub-models of different sizes (e.g., a relationship curve between performance parameters-model size-training steps), the corresponding relationship between the model size, training steps and performance parameters can be obtained; typical fitting algorithms may include least squares method, gradient descent method, etc.

[0087] Determining the performance evaluation results of the second machine learning model under training on the training data obtained by each candidate processing method based on the correspondence between the model parameters and the performance parameters can include: obtaining the model parameters of the second machine learning model, the model parameters including at least one of the structural parameters and the training parameters of the model; based on the correspondence between the model parameters and the performance parameters and the model parameters of the second machine learning model, determining the estimated performance parameters of the second machine learning model under training on the training data obtained by each candidate processing method, that is, the performance evaluation results. For example, the correspondence between the model parameters and the performance parameters is a functional relationship between model size-training steps-performance parameters. The size and training steps of the second machine learning model can be input into the functional relationship between model size-training steps-performance parameters to determine the estimated performance parameters of the second machine learning model under training on the training data obtained by each candidate processing method.

[0088] Step 340: Determine a target processing method based on the performance evaluation results of the second machine learning model trained on the training data obtained from each candidate processing method. The performance evaluation results of the second machine learning model trained on the data set obtained from the target processing method meet the preset conditions.

[0089] The target processing mode is configured to process the pre-training data to obtain training data for the second machine learning model. For more description of the second machine learning model, please refer to the detailed description in Figure 1.

[0090] Multiple data sets are obtained by processing the same initial data set using multiple candidate processing methods. Each data set is obtained by processing the initial data set using one of the multiple candidate processing methods. Steps 310-340 are performed on each of the multiple data sets to obtain a corresponding relationship between the model parameters and the performance parameters corresponding to each data set.

[0091] The size and / or preset number of training steps of the second machine learning model can be obtained in advance. The preset number of training steps can be determined based on actual needs, for example, it can be determined based on computing resources. A smaller number of steps can be set for fewer computing resources. Based on the size and / or preset number of training steps of the second machine learning model and the correspondence between multiple model parameters and performance parameters, the performance parameters of the second machine learning model under the size and preset number of training steps are determined. For example, for each correspondence between a model parameter and a performance parameter, a performance parameter of a second machine learning model can be determined based on the size and / or preset number of training steps of the second machine learning model. The performance parameters of multiple second machine learning models are compared to determine the optimal performance parameters. The candidate processing method of the data set corresponding to the optimal performance parameter is determined as the target processing method.

[0092] In some embodiments, for each type of initial data set among multiple types of initial data sets, different candidate processing methods can be used to process the initial data set of that type to obtain multiple data sets of that type, and each data set can be processed separately through steps 310-340 to obtain the target processing method for that type of data. For example, in Figure 4, after processing the initial data set using candidate processing method Pi, data set Si can be obtained, and data set Si is input into sub-model M1, sub-model M2, sub-model M3, ..., sub-model Mn to train sub-model M1, sub-model M2, sub-model M3, ..., sub-model Mn. The training of each sub-model includes a certain number of training steps. The number of training steps for multiple sub-models can be the same or different. In each step of training of each sub-model, sub-model M1, sub-model M2, sub-model M3, ..., sub-model Mn will generate prediction result 1, prediction result 2, prediction result 3, ..., prediction result N, respectively. The performance evaluation results 1, performance evaluation results 2, performance evaluation results 3, ..., performance evaluation results N corresponding to each sub-model can be generated based on the prediction results 1, prediction results 2, prediction results 3, ..., prediction results N generated by each training step in all the training steps. For example, the performance parameters of sub-model 1 after each training step can be determined based on the prediction results 1 generated by each training step of sub-model 1, and then the training steps, size, and performance parameters of sub-model 1 after all training steps of sub-model 1 are fitted to generate the performance evaluation result 1 of sub-model 1. The performance evaluation results of other sub-models can also be obtained. Optionally, the performance evaluation results 1, performance evaluation results 2, performance evaluation results 3, ..., performance evaluation results N can be fitted to finally generate the fitting relationship fi between the model parameters and the performance parameters. The fitting relationship fi between the model parameters and the performance parameters represents the corresponding relationship between the performance parameters and the model parameters of the model under the training of the data set obtained by the candidate processing mode Pi.

[0093] Similarly, the initial data set can be processed using the candidate processing methods P1, P2, ..., Pn to obtain data sets S1, S2, ..., Sn. Then, multiple sub-models are trained based on the data sets S1, S2, ..., Sn to obtain the corresponding relationships f1, f2, ..., fn between the performance parameters of the model under the training of the data sets obtained by processing the candidate processing methods P1, P2, ..., Pn and the model parameters. The performance parameters of the second machine learning model can be calculated by using the size and number of training steps of the second machine learning model and the corresponding relationships f1, f2, ..., fn between the performance parameters and the model parameters, so as to obtain the performance parameters of the second machine learning model under the training of the data sets obtained by processing the candidate processing methods P1, P2, ..., Pn. The optimal performance parameters are determined by comparing the performance parameters of the second machine learning model under the training of the data sets obtained by processing the candidate processing methods P1, P2, ..., Pn, so as to determine the candidate processing method corresponding to the optimal performance parameters (i.e., the target processing method).

[0094] In this embodiment, the performance evaluation results of a large model after training with different data processing methods are used to predict the performance evaluation results of a large model after training with different data processing methods using small models of different sizes. This is used to determine the data processing method that is suitable for training the large model. Because large models are more computationally and resource-intensive, using a small model that is only about 1 / 10 or even 1% the size of the large model to evaluate the quality of the data can reduce resource consumption.

[0095] Figure 5 is a schematic diagram of a training data acquisition system according to some embodiments of the present disclosure. As shown in Figure 5, training data acquisition system 500 may include an acquisition module 510, a processing module 520, a training module 530, and an evaluation module 540. Each module in training data acquisition system 500, such as the acquisition module, the processing module, the training module, and the evaluation module, may be executed by a computing device or a processing device.

[0096] The acquisition module 510 is configured to acquire an initial data set, which includes data of the same type. For more description on how to acquire the initial data set, please refer to step 110 in FIG1 .

[0097] Processing module 520 is configured to process the data in the initial dataset using multiple candidate processing methods to obtain multiple datasets, each of which corresponds to one of the multiple candidate processing methods. For more information on obtaining the datasets and the candidate processing methods, see step 120 in Figure 1.

[0098] The training module 530 is configured to train the first machine learning model based on the multiple data sets to obtain a trained first machine learning model. For more description of the training of the first machine learning model, please refer to step 130 in Figure 1.

[0099] Evaluation module 540 is configured to evaluate multiple candidate processing approaches using the trained first machine model to determine a target processing approach, where the target processing approach is configured to process pre-trained data to obtain training data for a second machine learning model, wherein the size of the first machine learning model is smaller than the size of the second machine learning model. For more information on determining the target candidate processing approach, see step 140 in FIG. 1 .

[0100] One or more embodiments of the present disclosure provide a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, a training data acquisition method is implemented.

[0101] Figure 6 is a schematic diagram of a distributed processing system according to some embodiments of the present disclosure. As shown in Figure 6, distributed processing system 600 may include a computing cluster 610, a network 620, a memory 630, a load balancer 650, and a manager 640. Distributed processing system 600 can be used in fields such as natural language processing, video processing, autonomous driving, healthcare, financial risk control, etc. The components in distributed processing system 600 can be connected in one or more different ways. By way of example only, one or more computing devices in computing cluster 610 may be connected to manager 640 and / or memory 630 via network 620. As another example, one or more computing devices in computing cluster 610 may be directly connected to memory 630, as shown in Figure 6. Memory 630 may also be connected to manager 640 via network 620. Network 620 may include any suitable network (e.g., a wide area network (WAN), a local area network (LAN), wired and / or wireless network access points, etc.) that can facilitate information and / or data exchange for distributed processing system 600. In some embodiments, network 620 may be the Internet, an intranet, or a combination thereof.

[0102] The computing cluster 610 may include at least two computing devices, such as computing device 610-1, computing device 610-2, ..., computing device 610-N, etc. In some embodiments, the at least two computing devices may be connected and / or communicate with each other via a wireless connection (e.g., network 620), a wired connection, or a combination thereof. For example, computing device 610-1 may receive data processed by computing device 610-2 via network 620. As another example, one of the at least two computing devices (e.g., computing device 610-1) may be directly connected to and / or communicate with the other computing device (e.g., computing device 610-2).

[0103] As used herein, a computing device may also be referred to as a computing node. Each of the at least two computing devices may include at least one processor and at least one storage device. In some embodiments, the computing device (e.g., computing device 610-1) may be any suitable computer, such as a laptop computer, a tablet computer, a desktop computer, etc. The at least one processor may include a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), or any other circuit or processor capable of executing one or more functions or any combination thereof. At least one storage device may store data, instructions, and / or any other information. In some embodiments, at least one storage device of the computing device may store data obtained from the load balancer 650 and / or the manager 640. In some embodiments, the storage device may store algorithms and / or instructions, and at least one processor may execute or use the algorithms and / or instructions to perform the training data acquisition method described in this application. For example, the processor may execute the training data acquisition method in processes 100, 200, and 300. For further example, the processor may obtain an initial data set from the load balancer 650, process the data in the initial data set using multiple candidate processing methods to obtain multiple data sets; train a first machine learning model based on each of the multiple data sets to obtain a trained first machine learning model; and evaluate the multiple candidate processing methods based on the trained first machine learning model to determine a target processing method.

[0104] In some embodiments, the processor may further process the received data set based on a determined target processing method to generate training data for training the large model.

[0105] Manager 640 can receive raw data packets and divide them into multiple initial data sets based on data type. Manager 640 can distribute the multiple initial data sets to computing devices in computing cluster 610 via load balancer 650. For more information on data classification, see the detailed description of FIG1 . Based on the received initial data sets, the computing devices can determine the optimal processing method (i.e., target processing method) corresponding to the initial data sets. The computing devices can also process the data sets of that type based on the optimal processing method to obtain training data for training the large model.

[0106] The load balancer 650 allocates at least two computing devices to acquire training data based on a load balancing algorithm or strategy. Exemplary load balancing algorithms may include a round-robin algorithm, a weighted round-robin algorithm, an automatic backup algorithm, a minimum number of connections algorithm, a random algorithm, a weighted random algorithm, and the like. In some embodiments, the load balancer 650 may allocate computing devices to acquire model training data based on the computing device's resources and the amount of data in each initial data set. In some embodiments, the load balancer 650 may be integrated with the manager 640.

[0107] Memory 630 may store data, instructions, and / or any other information. In some embodiments, memory 630 may store data obtained from load balancer 650 and / or manager 640. For example, memory 630 may store configuration information for each of at least two computing devices. In some embodiments, memory 630 may store data and / or instructions that manager 640 may execute or use to perform the exemplary methods / systems described herein.

[0108] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit the present disclosure. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and revisions to the present disclosure. Such modifications, improvements, and revisions are suggested in the present disclosure and remain within the spirit and scope of the exemplary embodiments of the present disclosure.

[0109] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of the present disclosure are approximate values, in specific embodiments, the settings of such numerical values ​​are as accurate as possible within the feasible range.

[0110] Finally, it should be understood that the embodiments described in this disclosure are intended only to illustrate the principles of the embodiments of the present disclosure. Other variations may also fall within the scope of this disclosure. Therefore, by way of example and not limitation, alternative configurations of the embodiments of the present disclosure may be considered consistent with the teachings of this disclosure. Accordingly, the embodiments of the present disclosure are not limited to the embodiments explicitly described and illustrated in this disclosure. Industrial Applicability

[0111] By adopting the above scheme, the small model is trained by using the data sets obtained by different candidate processing methods, and the performance of the small model is evaluated to screen the optimal data processing method, which is used to determine the training data of the large model, thereby obtaining training data that is easy for the large model to learn and the data is not damaged, so that the large model can achieve better performance with fewer training steps.

Claims

1. A method for obtaining training data, comprising: Obtaining an initial data set, where the initial data set includes data of the same type; Processing the data in the initial data set using multiple candidate processing methods to obtain multiple data sets; Training a first machine learning model based on the multiple data sets to obtain a trained first machine learning model; And Evaluating the multiple candidate processing methods based on the trained first machine learning model to determine a target processing method, where the target processing method is configured to process pre-training data to obtain training data for a second machine learning model, and wherein the size of the first machine learning model is smaller than the size of the second machine learning model.

2. The training data acquisition method according to claim 1, wherein, Evaluating the multiple candidate processing methods based on the trained first machine model to determine a target processing method includes: Obtaining test data in each of the multiple data sets; Processing the test data in each data set using the trained first machine learning model to obtain prediction results; Determining a performance evaluation result of the first machine learning model trained with each data set obtained under each candidate processing method among the multiple candidate processing methods based on the prediction results; and Determining the target processing method from the multiple candidate processing methods based on the performance evaluation results, where the performance evaluation result of the trained first machine learning model under the test data processed by the target processing method meets a preset condition.

3. The training data acquisition method according to claim 1, wherein, The first machine learning model includes multiple sub-models of different sizes. Training a first machine learning model based on each of the multiple data sets to obtain a trained first machine learning model includes: Training the multiple sub-models respectively using the same training data in each data set to obtain multiple trained sub-models; and Evaluating the multiple candidate processing methods using the trained first machine model to determine a target processing method includes: Obtaining a performance evaluation result of each of the multiple sub-models trained with the data set obtained under each candidate processing method among the multiple candidate processing methods, where each candidate processing method corresponds to a performance evaluation result; Predicting a performance evaluation result of the second machine learning model trained with the training data obtained under each candidate processing method based on the performance evaluation results of the multiple sub-models; and Determining the target processing method based on the performance evaluation result of the second machine learning model trained with the training data obtained under each candidate processing method, where the performance evaluation result of the second machine learning model trained with the data set obtained under the target processing method meets a preset condition.

4. The training data acquisition method according to claim 3, wherein, Predicting a performance evaluation result of the second machine learning model trained with the training data obtained under each candidate processing method based on the performance evaluation results of the multiple sub-models includes: Fitting the performance evaluation results of the multiple sub-models trained with the data set obtained under each candidate processing method to generate a correspondence between the model parameters and the performance parameters under the data set obtained under each candidate processing method; and Determine the performance evaluation result of the second machine learning model trained with the training data obtained under each candidate processing method based on the correspondence between the model parameters and the performance parameters.

5. The training data acquisition method according to claim 4, wherein, Determining the performance evaluation result of the second machine learning model trained with the training data obtained under each candidate processing method based on the correspondence between the model parameters and the performance parameters includes: Obtain the model parameters of the second machine learning model, where the model parameters include at least one of the structural parameters and training parameters of the model; Based on the correspondence between the model parameters and the performance parameters and the model parameters of the second machine learning model, determine the estimated performance parameters of the second machine learning model trained with the training data obtained under each candidate processing method.

6. The training data acquisition method according to claim 1, wherein, Training the first machine learning model based on each dataset in the multiple datasets to obtain the trained first machine learning model includes: Sequentially input the training data in the multiple datasets into the first machine learning model to obtain the trained first machine learning model.

7. The training data acquisition method according to claim 1, wherein, The first machine learning model includes multiple sub-models of the same size. Training the first machine learning model based on each dataset in the multiple datasets to obtain the trained first machine learning model includes: Input the training data in the multiple datasets into the multiple sub-models respectively to obtain the trained first machine learning model, and each dataset corresponds to one sub-model.

8. The training data acquisition method according to claim 1, wherein, Obtain an initial dataset, where the initial dataset includes data of the same type, including: Obtain the original data; Classify the original data based on the data type to obtain multiple initial datasets; and Send the multiple initial datasets to multiple computing devices respectively to obtain the initial datasets, and the multiple computing devices are nodes in a distributed system.

9. The training data acquisition method according to claim 3, wherein, The performance evaluation result of each sub-model includes: the performance evaluation data of each sub-model, the model size of each sub-model, and the training steps of each sub-model. The performance evaluation data of each sub-model is obtained by inputting the test data in the dataset of each sub-model into the trained each sub-model and evaluating the prediction results generated by each sub-model.

10. The training data acquisition method according to claim 5, wherein, The structural parameter of the model is the size of the model, and the training parameter of the model is the training steps of the model.

11. A distributed system for training data acquisition, comprising: A computing cluster, the computing cluster includes multiple computing devices; A manager, configured to obtain the original data, classify the original data based on the data type to obtain multiple initial datasets; And send the multiple initial datasets to the multiple computing devices respectively, and any computing device in the multiple computing devices is configured to execute any one of the training data acquisition methods as claimed in claims 1-10 based on the received initial dataset.

12. A training data acquisition system, comprising an acquisition module, a processing module, a training module and an evaluation module; The obtaining module is configured to obtain an initial data set, where the initial data set includes data of the same type; The processing module is configured to process the data in the initial data set by using a plurality of candidate processing methods to obtain a plurality of data sets, and each data set in the plurality of data sets corresponds to one of the plurality of candidate processing methods; The training module is configured to train a first machine learning model based on the plurality of data sets to obtain a trained first machine learning model; and The evaluation module is configured to evaluate the plurality of candidate processing methods by using the trained first machine model to determine a target processing method, where the target processing method is configured to process pre-training data to obtain training data for a second machine learning model, and the size of the first machine learning model is smaller than the size of the second machine learning model.

13. A training data acquisition device, including a processor, where the processor is configured to execute the training data acquisition method according to any one of claims 1 to 10.

14. A computer-readable storage medium, where the storage medium stores computer instructions, and when a computer reads the computer instructions in the storage medium, the computer executes the training data acquisition method according to any one of claims 1 to 10.

15. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer executable instructions are executed by a processor, the training data acquisition method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Model training method and device, equipment and storage medium

    CN113222149A

  • Data processing method and device, equipment and computer storage medium

    CN113469204A

  • Method and system for selecting data processing model

    CN115688932A

  • Training data acquisition method, system and device, storage medium and program

    CN117910603A

  • Model training program, model training method, and information processing device

    EP4160492A1

Cited By

  • High-speed wire harness low-resistance welding method and system based on local microenvironment control

    CN121696520A