Pre-training data processing method and device

Through the pre-trained multi-dimensional data evaluation model, the training data is evaluated and screened in multi-dimensional, which solves the problem that low-quality data in the existing technology affects the model performance, and achieves higher quality and applicability training data, improving model performance and training efficiency.

CN120011767APending Publication Date: 2025-05-16SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510110920.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively screen low-quality data during model training, resulting in the impact of model performance and the inefficient manual cleaning and single-dimensional machine learning scoring methods.

Method used

A pre-trained multi-dimensional data evaluation model is used to evaluate the data in the training dataset in a multi-dimensional manner, evaluate the evaluation parameters, and filter out the target training data based on these parameters. The model combines large language model and multi-layer perceptron to improve data quality and applicability through multi-dimensional evaluation and weighting operations.

Benefits of technology

Through multi-dimensional evaluation and screening, the quality and applicability of training data are improved, the impact of noise data is reduced, and the performance and training efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011767A_ABST
    Figure CN120011767A_ABST
Patent Text Reader

Abstract

The invention provides a pre-training data processing method and device, and relates to the technical field of data processing. The method comprises the steps that a first training data set of a target training task is acquired, a pre-trained multi-dimensional data evaluation model is adopted to evaluate training data in the first training data set, and evaluation parameters of the training data in the first training data set in multiple dimensions are obtained, and finally, screening target training data of the target training task from the first training data set according to the evaluation parameters of the training data in the multiple dimensions in the first training data set. According to the method provided by the invention, through multi-dimensional evaluation and screening, the quality and applicability of the training data are improved, and better data support is provided for subsequent target training tasks, so that the performance of a training model and the final application effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a pre-training data processing method and device. Background Art

[0002] When training a model, the data quality in the training dataset is crucial to the model performance. However, low-quality data exists in the existing dataset, which affects the model performance.

[0003] Currently, in order to screen out low-quality data in existing data sets, low-quality data is cleaned manually, but it is time-consuming and labor-intensive and difficult to process massive data, or a single-dimensional machine learning scoring method is used, but it is difficult to comprehensively measure data quality. Summary of the invention

[0004] The purpose of the present invention is to provide a pre-training data processing method and device to address the deficiencies in the above-mentioned prior art, so as to improve the quality and applicability of training data through multi-dimensional evaluation and screening, and provide better data support for subsequent target training tasks.

[0005] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for processing pre-training data, the method comprising:

[0007] Obtain a first training data set for a target training task;

[0008] Using a pre-trained multi-dimensional data evaluation model, evaluate each training data in the first training data set to obtain evaluation parameters of each training data in the first training data set in multiple dimensions;

[0009] Target training data for the target training task are screened from the first training data set according to evaluation parameters of the plurality of dimensions of the respective training data in the first training data set.

[0010] In an optional implementation, the pre-trained multi-dimensional data evaluation model is used to evaluate each training data in the training data set to obtain evaluation parameters of each training data in the training data set in multiple dimensions, including:

[0011] Generate evaluation prompt words for each training data according to at least one of the description information of the multiple dimensions, the association relationship between different dimensions in the multiple dimensions, the scoring rules of the multiple dimensions under the target training task, the data type of each training data and the task type of the target training task;

[0012] According to the evaluation prompt words, the multi-dimensional data evaluation model is used to evaluate the various training data to obtain evaluation parameters of the various training data in the multiple dimensions.

[0013] In an optional embodiment, the multi-dimensional data evaluation model includes: a first preset large language model and a multi-layer perceptron;

[0014] The step of evaluating each training data according to the evaluation prompt word using the multi-dimensional data evaluation model to obtain evaluation parameters of each training data in the multiple dimensions includes:

[0015] According to the evaluation prompt word, the first preset large language model is used to process each training data to obtain embedded information of each training data;

[0016] The multilayer perceptron is used to process the embedded information of each training data to obtain evaluation parameters of each training data in the multiple dimensions.

[0017] In an optional implementation, before the multilayer perceptron is used to process the embedded information of each training data to obtain the evaluation parameters of each training data in the multiple dimensions, the method further includes:

[0018] Obtain a second training data set for the multi-dimensional evaluation task;

[0019] The multi-layer perceptron is trained using the second training data set to obtain the multi-dimensional data evaluation model.

[0020] In an optional implementation, screening target training data for the target training task from the first training data set according to the evaluation parameters of the training data in the first training data set in the multiple dimensions includes:

[0021] Based on the weight parameters of the multiple dimensions under the target training task, weighted operations are performed on the evaluation parameters of the multiple dimensions of each training data to obtain target evaluation parameters of each training data;

[0022] According to the target evaluation parameters of each training data, training data that meets an evaluation threshold is determined from the first training data set as the target training data.

[0023] In an optional embodiment, the method further comprises:

[0024] According to the data type of each training data and / or the task type of the target training task, based on preset weight parameters and / or historical weight parameters, the weight parameters of the multiple dimensions under the target training task are determined.

[0025] In an optional embodiment, the method further comprises:

[0026] Determining initial weight parameters of the multiple dimensions under the target training task according to the data type of the respective training data and / or the task type of the target training task;

[0027] Determining initial target training data based on the initial weight parameters and the initial evaluation threshold;

[0028] Based on the initial target training data, training a second preset large language model to obtain a target second model;

[0029] Evaluate the target second model based on at least one evaluation index to obtain at least one evaluation index parameter of the target second model;

[0030] According to at least one of the evaluation index parameters, the initial weight parameter and / or the initial evaluation threshold is adjusted or determined to obtain the weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task.

[0031] In an optional embodiment, the method further comprises:

[0032] Determining a plurality of initial weight parameters of the plurality of dimensions under the target training task according to a data type of the respective training data and / or a task type of the target training task;

[0033] Based on the multiple initial weight parameters and at least one initial evaluation threshold, respectively determine corresponding initial target training data;

[0034] Based on the initial target training data, training the second preset large language models respectively to obtain multiple target second models;

[0035] Evaluate the plurality of target second models based on at least one evaluation index to obtain at least one evaluation index parameter of the plurality of target second models;

[0036] According to at least one of the evaluation index parameters of the multiple target second models, weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task are obtained.

[0037] In an optional embodiment, obtaining the weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task according to at least one evaluation index parameter of the multiple target second models includes:

[0038] Based on at least one of the evaluation index parameters of the plurality of target second models, the initial weight parameter and / or initial evaluation threshold corresponding to the target second model whose evaluation index parameter meets the preset conditions is determined as the weight parameter and / or evaluation threshold of the plurality of dimensions under the target training task;

[0039] Alternatively, based on at least one of the evaluation index parameters of the plurality of target second models, the initial weight parameter and / or initial evaluation threshold corresponding to the target second model whose evaluation index parameter meets a preset performance target is determined as the weight parameter and / or evaluation threshold of the plurality of dimensions under the target training task;

[0040] Alternatively, fitting is performed based on at least one of the evaluation index parameters of the plurality of target second models and the corresponding plurality of initial weight parameters and at least one initial evaluation threshold to generate a performance weight fitting function;

[0041] Based on the performance weight fitting function, weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task are obtained.

[0042] In a second aspect, an embodiment of the present application further provides a pre-training data processing device, the device comprising:

[0043] An acquisition module, used to acquire a first training data set for a target training task;

[0044] An evaluation module, configured to evaluate each training data in the first training data set by using a pre-trained multi-dimensional data evaluation model, and obtain evaluation parameters of each training data in the first training data set in multiple dimensions;

[0045] A screening module is used to screen the target training data of the target training task from the first training data set according to the evaluation parameters of the training data in the first training data set in the multiple dimensions.

[0046] The beneficial effects of this application are:

[0047] The embodiment of the present application provides a pre-training data processing method and device, the method comprising: obtaining a first training data set of a target training task, using a pre-trained multi-dimensional data evaluation model to evaluate each training data in the first training data set, obtaining evaluation parameters of each training data in the first training data set in multiple dimensions, and finally screening target training data of the target training task from the first training data set according to the evaluation parameters of each training data in the first training data set in multiple dimensions. The method of the present application aims to improve the quality and applicability of training data through multi-dimensional evaluation and screening, and provide better data support for subsequent target training tasks, so as to improve the performance of the training model and the final application effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 One of the flowcharts of a pre-training data processing method provided in an embodiment of the present application;

[0050] Figure 2 A second flowchart of a pre-training data processing method provided in an embodiment of the present application;

[0051] Figure 3 A third flowchart of a pre-training data processing method provided in an embodiment of the present application;

[0052] Figure 4 A fourth flowchart of a pre-training data processing method provided in an embodiment of the present application;

[0053] Figure 5 A fifth flowchart of a pre-training data processing method provided in an embodiment of the present application;

[0054] Figure 6 A sixth flowchart of a pre-training data processing method provided in an embodiment of the present application;

[0055] Figure 7 A flowchart of a pre-training data processing method provided in an embodiment of the present application is shown in FIG7;

[0056] Figure 8 A schematic diagram of the functional modules of a pre-training data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0058] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0059] In the description of the present application, it should be noted that if the terms "upper", "lower", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings, or is the orientation or position relationship in which the product of the application is usually placed when used. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application.

[0060] In addition, the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0061] It should be noted that, in the absence of conflict, the features in the embodiments of the present application may be combined with each other.

[0062] In order to screen out target training data that better meets the requirements of the target training task, an embodiment of the present application provides a pre-training data processing method, by obtaining a first training data set of the target training task, using a pre-trained multi-dimensional data evaluation model, evaluating each training data in the first training data set, obtaining evaluation parameters of each training data in the first training data set in multiple dimensions, and finally screening the target training data of the target training task from the first training data set according to the evaluation parameters of each training data in the first training data set in multiple dimensions. By performing a multi-dimensional evaluation model on each training data to evaluate multiple dimensions, the target training data that meets the requirements of the target training task can be screened out more accurately without manual cleaning, which can not only ensure the quality of the screened target training data, but also improve the screening efficiency.

[0063] A pre-training data processing method provided in an embodiment of the present application is explained in detail below by using specific examples in conjunction with the accompanying drawings. A pre-training data processing method provided in an embodiment of the present application can be implemented by a computer device pre-installed with: a pre-training data processing algorithm or detection software, by running the algorithm or software. The computer device can be, for example, a server or a terminal, and the terminal can be a user computer. Figure 1 One of the flow charts of a pre-training data processing method provided in an embodiment of the present application; Figure 1 As shown, the method includes:

[0064] S101: Obtain a first training data set for a target training task.

[0065] In this embodiment, before processing the pre-training data, it is first necessary to collect or prepare a data set for the target training task. The first training data set can come from multiple channels and contain a large amount of raw data, which may be text, images, audio or other forms of data, depending on the nature of the target training task. For example, if the target training task is to train a natural language processing model, the first training data set may contain a large number of text files, and if it is an image recognition task, the first training data set may contain a large number of image files.

[0066] S102: using a pre-trained multi-dimensional data evaluation model to evaluate each training data in the first training data set, and obtaining evaluation parameters of each training data in the first training data set in multiple dimensions.

[0067] Specifically, the multi-dimensional data evaluation model is pre-trained, and its purpose is to evaluate each training data in the first training data set from multiple dimensions. Multi-dimensional means not only looking at the training data from a single perspective, but also taking multiple factors into consideration. For example, for text data, the dimensions that may be considered include the source of the text data, the length of the data, the vocabulary, the semantic complexity, the emotional tendency, the information density, etc.; for image data, the dimensions that may be considered include the resolution of the image, the color distribution, the number of objects, the clarity of the objects, the noise level of the image, etc.

[0068] Each piece of training data in the first training data set is input into the pre-trained multi-dimensional data evaluation model. The multi-dimensional data evaluation model will analyze and calculate each piece of training data according to its internal algorithm model and pre-trained parameters, thereby generating multi-dimensional evaluation parameters for each piece of training data.

[0069] S103: Filter target training data for a target training task from the first training data set according to evaluation parameters of each training data in multiple dimensions in the first training data set.

[0070] Specifically, according to the evaluation parameters of multiple dimensions, a preset screening strategy or a screening strategy generated according to the target training task is used to select target training data suitable for the target training task from the first training data set.

[0071] By screening the first training data set, the quality and relevance of the target training data can be improved, and some irrelevant, low-quality or unrepresentative data can be avoided from being included in the training process, thereby improving the performance and efficiency of subsequent training models. This helps the model better learn the characteristics and laws of the data, reduces the interference of noise data, and makes the trained model more accurate and stable.

[0072] In summary, the embodiment of the present application provides a pre-training data processing method, which includes: obtaining a first training data set of a target training task, using a pre-trained multi-dimensional data evaluation model to evaluate each training data in the first training data set, and obtaining evaluation parameters of each training data in the first training data set in multiple dimensions, and finally screening the target training data of the target training task from the first training data set according to the evaluation parameters of each training data in the first training data set in multiple dimensions. The method of the present application aims to improve the quality and applicability of training data through multi-dimensional evaluation and screening, and provide better data support for subsequent target training tasks, so as to improve the performance of the training model and the final application effect.

[0073] The present application embodiment also provides another possible implementation of the pre-training data processing method. Figure 2 A second flow chart of a pre-training data processing method provided in an embodiment of the present application is as follows: Figure 2 As shown, a pre-trained multi-dimensional data evaluation model is used to evaluate each training data in the training data set to obtain evaluation parameters of each training data in the training data set in multiple dimensions, including:

[0074] S201, generating evaluation prompt words for each training data according to at least one of the description information of multiple dimensions, the correlation relationship between different dimensions in the multiple dimensions, the scoring rules of multiple dimensions under the target training task, the data type of each training data and the task type of the target training task.

[0075] S202: According to the evaluation prompt words, each training data is evaluated using a multi-dimensional data evaluation model to obtain evaluation parameters of each training data in multiple dimensions.

[0076] In this embodiment, for example, the target training task is text training, and the multiple dimensions may include dimensions such as the source of training data, the knowledge of training data, and whether the training data contains illegal words. The source of training data can represent the authority of the training data. For example, the training data source is a journal or a conference, which is more authoritative than the training data source is a web page. There may be certain connections between different dimensions. For example, if the training data source is a journal or a conference, the training data is more knowledgeable.

[0077] Different target training tasks may have different evaluation criteria and scoring rules for each dimension. For example, when training text, the training data needs to be more authoritative. The description information of the training data sources in multiple dimensions can be: the training data source is a journal with a higher authority score than the source is a web page, but the maximum score is not more than 10 points, etc. The scoring rules for dimensions such as the knowledge of the training data and whether the training data contains illegal words are similar. These rules will be incorporated into the generation of evaluation prompt words so that the data can be accurately evaluated according to the specific requirements of the task.

[0078] The data type (such as text, image, audio) and task type (such as classification, regression, generation task) generally also affect the generation of evaluation prompts. For the classification task of text data, the evaluation prompts may focus on the category characteristics of the text; for the generation task of image data, they may pay more attention to the details and overall composition of the image. At the same time, the focus and method of evaluation will vary depending on the data type and task type.

[0079] When generating evaluation prompts, specific training data will be taken into account. For example, for a specific text, the evaluation prompt may mention the data source, knowledge, whether there are illegal words, etc. of the text, so that the evaluation prompt is more in line with the actual data.

[0080] The generated evaluation prompt words are passed as part of the input to the multi-dimensional data evaluation model. The evaluation prompt words provide the model with specific evaluation directions and requirements, guiding the model on how to evaluate the data. The multi-dimensional data evaluation model analyzes and processes each training data according to the evaluation prompt words. For text data, the model may analyze the data source, knowledge, whether there are illegal words, etc. of the text; through the evaluation, the corresponding evaluation parameters are output for each training data in multiple dimensions. For text data, evaluation parameters such as "data source: 0.8 (assuming that 0 to 1 means that the authority is getting higher and higher), knowledge: 0.3 (assuming that 0 to 1 means that the knowledge is getting higher and higher)" may be obtained.

[0081] In the method provided in the embodiment of the present application, evaluation prompt words for each training data are generated based on the descriptive information of multiple dimensions, the correlation between different dimensions in the multiple dimensions, the scoring rules of multiple dimensions under the target training task, the data type of each training data and at least one of the task types of the target training task; based on the evaluation prompt words, a multi-dimensional data evaluation model is used to evaluate each training data to obtain evaluation parameters of each training data in multiple dimensions, and evaluation prompt words are generated in combination with multiple information. Based on the evaluation prompt words, the multi-dimensional data evaluation model is used to evaluate the training data to finally obtain evaluation parameters, which helps to evaluate the training data more scientifically and more in line with the requirements of the target task, thereby better screening out high-quality data, providing a more valuable data foundation for subsequent training processes, and improving the quality and performance of the training model.

[0082] The embodiment of the present application also provides another possible implementation of the pre-training data processing method, wherein the multi-dimensional data evaluation model includes: a first preset large language model and a multi-layer perceptron; Figure 3 A third flow chart of a pre-training data processing method provided in an embodiment of the present application is as follows: Figure 3 As shown, according to the evaluation prompt words, a multi-dimensional data evaluation model is used to evaluate each training data to obtain evaluation parameters of each training data in multiple dimensions, including:

[0083] S301 . According to the evaluation prompt word, each training data is processed using a first preset large language model to obtain embedding information of each training data.

[0084] In this embodiment, the first preset large language model can be GPT (Generative Pretrained Transformer Decode Only Transformer), which is a model based on deep learning, usually trained on a large amount of text data, and can learn rich language knowledge and semantic information. In this multi-dimensional data evaluation model, the first preset large language model is used as the first link in processing training data. It can handle various types of text inputs and has powerful text understanding and representation capabilities. For example, for natural language processing tasks, it can understand the semantics, grammar, emotional tendencies and other information of the text.

[0085] The evaluation prompt word provides the first preset large language model with the direction and specific requirements for the evaluation. It can guide the large language model to focus on specific aspects of the training data. For example, if the evaluation prompt word is "focus on evaluating the authority of the text", the first preset large language model will focus on evaluating the data source of the text based on this prompt.

[0086] First, the large language model takes the training data and evaluation prompt words as input. For text training data, the large language model will convert its results into an internal representation after processing, that is, embedded information. This embedded information is a high-dimensional, dense vector representation of the training data, which contains the information of the text in the semantic space learned by the large language model.

[0087] S302: Using a multi-layer perceptron, the embedded information of each training data is processed to obtain evaluation parameters of each training data in multiple dimensions.

[0088] Multilayer Perceptron (MLP) is a classic neural network architecture consisting of multiple neuron layers, including input layer, hidden layer and output layer. It can perform nonlinear transformation on input data and learn complex mapping relationships between input and output. In the multi-dimensional data evaluation model, the multilayer perceptron is used as a subsequent processing unit, taking the output of the large language model as input to further process and transform the data.

[0089] The multilayer perceptron receives the embedded information output by the large language model as input. Since the embedded information is a high-dimensional vector, the multilayer perceptron can process this high-dimensional input. The multilayer perceptron performs nonlinear transformations on the embedded information layer by layer through its internal neurons and weights. In each layer, the neurons calculate based on the input and weights, and generate outputs through activation functions (such as ReLU, Sigmoid, etc.), and finally convert the input embedded information into the required evaluation parameters, thereby obtaining the evaluation parameters of each training data in multiple dimensions.

[0090] In the method provided in the embodiment of the present application, according to the evaluation prompt words, the first preset large language model is used to process each training data to obtain the embedded information output after the processing of each training data; the embedded information of each training data is processed by a multi-layer perceptron to obtain the evaluation parameters of each training data in multiple dimensions. In practical applications, for different target training tasks, this evaluation model can perform targeted evaluation of the training data according to the evaluation prompt words. For example, when training a text classification model, authoritative texts can be screened out through this evaluation model to better train the classifier. By using a multi-dimensional data evaluation model that combines the first preset large language model and the multi-layer perceptron, the training data is processed according to the evaluation prompt words, and evaluation parameters that are similar to manual evaluation, more targeted and practical can be obtained, which provides strong data support for subsequent data screening and target training tasks, and helps to improve the performance and effect of the training model.

[0091] The present application embodiment also provides another possible implementation of the pre-training data processing method. Figure 4A fourth flow chart of a pre-training data processing method provided in an embodiment of the present application is as follows: Figure 4 As shown, a multi-layer perceptron is used to process the embedded information of each training data to obtain evaluation parameters of each training data in multiple dimensions. The method also includes:

[0092] S401. Obtain a second training data set for a multi-dimensional evaluation task.

[0093] S402: Use the second training data set to train the multilayer perceptron to obtain a multi-dimensional data evaluation model.

[0094] In this embodiment, the second training data set is a data set specially prepared for training a multi-layer perceptron. It contains a large number of samples, which should be data with known evaluation parameters to provide supervision information for the multi-layer perceptron. For different types of tasks, the specific content of the second training data set will be different. For example, if a multi-dimensional evaluation is performed on text data, the second training data set may contain a series of text samples, and each sample has been labeled with the corresponding multi-dimensional evaluation parameters manually or in other ways. The diversity and representativeness of the second training data set are very important. It should cover a variety of different situations so that the multi-layer perceptron can learn a wide range of data features and evaluation rules.

[0095] Specifically, the embedded information obtained after processing each sample in the second training data set is used as the input of the multilayer perceptron. These embedded information can be obtained in the same or similar manner as the first preset large language model mentioned above, ensuring that the input data format and feature representation are consistent. The multi-dimensional evaluation parameters annotated in the second training data set are used as supervisory information. For each sample's embedded information, the multilayer perceptron will try to predict its corresponding evaluation parameter, and then compare the predicted result with the known annotated result. By calculating the error between the predicted result and the annotated result (for example, using a loss function such as mean square error or cross entropy), and adjusting the weights and parameters of the multilayer perceptron according to the error. The optimization algorithm (such as stochastic gradient descent) will update the weights of the multilayer perceptron according to the gradient information of the loss function, so that the predicted result gradually approaches the annotated result. This process will be iterated continuously until a certain training goal is reached, such as reaching a predetermined training round or the value of the loss function is lower than a certain threshold.

[0096] After multiple rounds of training, the multi-layer perceptron has learned how to accurately predict multi-dimensional evaluation parameters based on the input embedded information. At this point, the multi-layer perceptron becomes part of the multi-dimensional data evaluation model, and together with the first preset large language model, they form a complete multi-dimensional data evaluation model. This multi-dimensional data evaluation model can accept evaluation prompts and training data for new training data. The first preset large language model first converts and processes the training data into embedded information, and then the trained multi-layer perceptron converts the embedded information into multi-dimensional evaluation parameters, thereby achieving multi-dimensional evaluation of the training data.

[0097] In the method provided in the embodiment of the present application, a second training data set for a multidimensional evaluation task is obtained; the second training data set is used to train the multilayer perceptron to obtain a multidimensional data evaluation model. Different target training tasks may require different evaluation dimensions and evaluation criteria. By training the multilayer perceptron, the multidimensional data evaluation model can be adapted to the specific needs of different tasks. For example, for different natural language processing tasks, the corresponding second training data set can be used for training so that the evaluation model can accurately evaluate the dimensions required for different tasks. For example, in machine translation tasks and text summarization tasks, the evaluation dimensions and criteria will be different. Through training, the model can be adapted to these different requirements. Before using the multilayer perceptron to process the embedded information obtained from each training data to obtain the evaluation parameters, it is crucial to train it using a dedicated second training data set. This step ensures that the multilayer perceptron can accurately convert the embedded information into the required multidimensional evaluation parameters, thereby improving the performance and adaptability of the entire multidimensional data evaluation model, and providing better data evaluation services for subsequent data screening and target training tasks.

[0098] The present application embodiment also provides another possible implementation of the pre-training data processing method. Figure 5 A fifth flow chart of a pre-training data processing method provided in an embodiment of the present application is as follows: Figure 5 As shown, according to the evaluation parameters of each training data in the first training data set in multiple dimensions, target training data of the target training task are screened from the first training data set, including:

[0099] S501: Based on the weight parameters of multiple dimensions under the target training task, weighted operations are performed on the evaluation parameters of each training data in multiple dimensions to obtain target evaluation parameters of each training data.

[0100] In this embodiment, in the target training task, different evaluation dimensions may have different importance to the final result, so a corresponding weight parameter is assigned to each dimension. For example, in a text classification training task, if the authoritative evaluation dimension is considered to be crucial for model training, then the authoritative evaluation dimension may be given a higher weight; and the dimension of whether there are illegal words, if it is relatively not the most critical factor, may be given a lower weight. These weight parameters reflect the relative importance of each dimension in the target training task. The weight parameters can be set manually or determined by some automatic optimization algorithm.

[0101] Then, by performing weighted operations on the evaluation parameters of each training data in multiple dimensions, the target evaluation parameters of each training data are obtained.

[0102] S502: According to the target evaluation parameters of each training data, determine the training data that meets the evaluation threshold from the first training data set as the target training data.

[0103] Specifically, the evaluation threshold can be a pre-set value or determined by some automatic optimization algorithm to filter training data. Its setting is usually based on the specific requirements of the task and the expectations of the data. For example, in a classification task, if high-quality data needs to be filtered out, the evaluation threshold may be set higher, such as 0.8; while for a more relaxed task, the evaluation threshold can be set to 0.5. The evaluation threshold can be determined based on experience, experiments, or statistical analysis. You can first conduct a small-scale experiment on part of the data to observe the quality of the data filtered out under different thresholds and the subsequent training effect, and then select the most appropriate evaluation threshold.

[0104] After calculating the target evaluation parameter of each training data in the first training data set, compare it with the evaluation threshold. For those training data whose target evaluation parameters are greater than or equal to the evaluation threshold, they are determined as target training data. For example, if the evaluation threshold is 0.7, the text training data with a target evaluation parameter of 0.7 calculated will be selected as the target training data; and the training data with a target evaluation parameter lower than 0.7 will be excluded from the target training data. Such a screening process ensures that the target training data finally selected meets certain quality or feature requirements, which is conducive to improving the efficiency of subsequent training tasks and the performance of the model.

[0105] In the method provided in the embodiment of the present application, based on the weight parameters of multiple dimensions under the target training task, the evaluation parameters of each training data in multiple dimensions are weighted, and the target evaluation parameters of each training data are obtained; according to the target evaluation parameters of each training data, the training data that meets the evaluation threshold is determined from the first training data set as the target training data. Through the above-mentioned weighted operation and screening process, the training data in the first training data set can be screened more finely according to the requirements of the target training task. First, the evaluation parameters of each dimension are weighted and integrated using the weight parameters to obtain target evaluation parameters that better meet the task requirements; then, the most suitable training data is selected as the target training data according to the evaluation threshold, providing better quality and more targeted data for subsequent training tasks, and improving the training effect and the quality of the model.

[0106] The present application embodiment also provides another possible implementation of the pre-training data processing method, which further includes:

[0107] According to the data type of each training data and / or the task type of the target training task, based on preset weight parameters and / or historical weight parameters, weight parameters of multiple dimensions under the target training task are determined.

[0108] In this embodiment, different data types will affect the determination of weight parameters of multiple dimensions, and the task type also plays an important role in the determination of weight parameters. For example, in a classification task, more attention may be paid to the dimensions that can distinguish the characteristics of different categories; in a regression task, more emphasis may be placed on the dimensions related to continuous variables; in a generation task, more attention may be paid to dimensions such as the integrity and richness of the data.

[0109] These can be pre-set weight parameters before starting training, usually based on experience or domain knowledge. They can serve as an initial reference to provide a rough importance distribution for different dimensions. Historical weight parameters are weight parameters that have been used previously in similar or the same task. If similar training tasks have been completed before, the weight parameters used in previous tasks can be used as a reference. These parameters may have been verified and, to some extent, reflect the effective weight distribution of evaluation dimensions for this type of task and data type.

[0110] When determining the weight parameters of multiple dimensions under the target training task, the data type of the training data, the task type of the target training task, the preset weight parameters, and the historical weight parameters are comprehensively considered. For different data types and task types, you can first select an initial weight distribution from the preset weight parameters, and then adjust it according to the historical weight parameters. For example, if this task is a new text classification task, you may first refer to the historical weight parameters of the previous text classification task, and then adjust it according to the current data type and the special requirements of the task, combined with the preset weight parameters. The weight parameters may be fine-tuned according to the specific goals of the current task and the specific situation of the data. For example, if the current text data contains a large number of professional vocabulary, the weight of the vocabulary dimension may be appropriately increased.

[0111] This method of determining weight parameters has a certain degree of flexibility and dynamism. According to new data types, task types, and new discoveries and needs, weight parameters can be adjusted dynamically instead of being fixed. For example, during the training process, if it is found that the evaluation results of some dimensions have little effect on the final training effect, while other dimensions have a greater impact, then the weight parameters of these dimensions can be reduced or increased accordingly to optimize the subsequent data screening and training process.

[0112] In the method provided in the embodiment of the present application, by considering the data type of the training data, the task type of the target training task, and combining the preset weight parameters and the historical weight parameters, the weight parameters of multiple dimensions under the target training task can be determined more reasonably. This helps to more accurately evaluate the training data, perform weighted operations on the evaluation parameters of different dimensions according to different tasks and data characteristics, and then screen out the target training data that is more suitable for the target training task, thereby improving the effectiveness of the entire training process and the performance of the training model.

[0113] The present application embodiment also provides another possible implementation of the pre-training data processing method. Figure 6 A flowchart of a pre-training data processing method provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the method also includes:

[0114] S601: Determine initial weight parameters of multiple dimensions under a target training task according to the data type of each training data and / or the task type of the target training task.

[0115] In this embodiment, different data types have different characteristics, which will affect the importance of each dimension in the evaluation process. For example, for text data, it may be evaluated based on dimensions such as data source, vocabulary, semantic complexity, and knowledge; for image data, dimensions such as image resolution, color distribution, object shape, and texture may be considered. When determining the initial weight parameters, the characteristics of the data type need to be considered. For example, for text data in natural language processing tasks, if more emphasis is placed on semantic understanding, a higher initial weight may be assigned to the semantic dimension; for image data in image recognition tasks, a higher initial weight may be assigned to the object shape or texture dimension.

[0116] The type of task is also an important consideration. For example, classification tasks may focus more on whether the data can clearly distinguish different categories, regression tasks may focus more on the numerical accuracy of the data, and generation tasks may focus more on the coherence and diversity of the data. Therefore, different initial weight parameters will be assigned according to different task types.

[0117] S602: Determine initial target training data based on initial weight parameters and initial evaluation thresholds.

[0118] Specifically, the initial evaluation threshold is a pre-set value used for preliminary screening of training data. It can be set based on experience or preliminary testing, and the initial evaluation result is obtained by weighting the evaluation parameters of each training data in multiple dimensions in the first training data set with the initial weight parameter. The initial evaluation result is compared with the initial evaluation threshold, and data greater than or equal to the threshold will be selected as the initial target training data.

[0119] S603: Based on the initial target training data, train the second preset large language model to obtain a target second model.

[0120] Specifically, the initial target training data is used as input for training the second preset large language model. These data contain screened data with certain quality and characteristics, which can provide more representative and targeted information for the model. The second preset large language model is a model with the same structure as the target training model to be trained, but the size (parameter quantity) is smaller than the size of the target model, and the training data requirement will also be proportionally reduced compared to the target training model, that is, if the parameter quantity of the target training model is 100G, the required training data quantity is 100T, the parameter quantity of the second preset large language model is 1G, the required training data quantity is 1T, and the required computing power will be much smaller, but the performance of the second preset large language model obtained is almost the same or highly similar. Therefore, the target training data obtained after screening is used to train the second preset large language model first, obtain the target second model, and evaluate its model performance under the initial weights and initial evaluation thresholds, and then determine the evaluation parameters and / or evaluation thresholds of multiple dimensions, or their optimization direction.

[0121] S604: Evaluate the target second model based on at least one evaluation index to obtain at least one evaluation index parameter of the target second model.

[0122] Among them, the evaluation indicators can be of various types, which vary according to different tasks and data types. For classification tasks, accuracy, recall, F1 value, etc. may be used; for regression tasks, mean square error, mean absolute error, etc. may be used; for generation tasks, BLEU score (text generation), SSIM score (image generation), etc. may be used. These indicators are used to measure the performance of the target second model, reflecting the fit degree of the target second model to the target training data and its prediction ability for new data.

[0123] The target second model is evaluated based on at least one evaluation index to obtain at least one evaluation index parameter of the target second model.

[0124] S605: Adjust or determine an initial weight parameter and / or an initial evaluation threshold according to at least one evaluation index parameter, and obtain weight parameters and / or evaluation thresholds of multiple dimensions under the target training task.

[0125] Specifically, the evaluation index parameters are analyzed. If it is found that the performance of the model does not meet expectations, the initial weight parameters and / or initial evaluation thresholds can be adjusted according to the feedback of the evaluation index parameters. For example, if it is found that the evaluation index parameters of the model in a certain dimension are low, it may indicate that the dimension has not received enough attention during the screening of training data or the training process, so the weight parameters of the dimension can be increased accordingly. For the evaluation threshold, if it is found that the initial target training data screened out is too much or too little, resulting in overfitting or underfitting of the model, the evaluation threshold can be increased or lowered accordingly. For example, if it is found that too much initial target training data causes the model to overfit, the evaluation threshold may be increased to screen out better quality data; conversely, if too little data causes underfitting, the evaluation threshold may be lowered.

[0126] This adjustment process is an iterative optimization process. Based on the initial weight parameters and the initial evaluation threshold, the initial target training data is determined, and a cyclic verification is performed until the evaluation indicators meet the requirements, and then the weight parameters and / or evaluation threshold are obtained.

[0127] In the method provided in the embodiment of the present application, the initial weight parameters and the initial evaluation threshold are determined by taking the data type and the task type as the starting point, the initial target training data is screened out, the second preset large language model is trained, the performance of the model is evaluated, and the initial parameters are adjusted according to the evaluation results to form a closed-loop optimization process, and finally the weight parameters and evaluation thresholds that are more suitable for the target training task are obtained to improve the effect of the entire training process and the performance of the trained model.

[0128] The present application embodiment also provides another possible implementation of the pre-training data processing method. Figure 7 A flowchart of a pre-training data processing method provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the method also includes:

[0129] S701: Determine multiple initial weight parameters of multiple dimensions under the target training task according to the data type of each training data and / or the task type of the target training task.

[0130] In this embodiment, different data types have different characteristics, which will affect the importance of each dimension in the evaluation process. For example, for text data, it may be evaluated based on dimensions such as data source, vocabulary, semantic complexity, and knowledge; for image data, dimensions such as image resolution, color distribution, object shape, and texture may be considered. When determining the initial weight parameters, the characteristics of the data type need to be considered. For example, for text data in natural language processing tasks, if more emphasis is placed on semantic understanding, a higher initial weight may be assigned to the semantic dimension; for image data in image recognition tasks, a higher initial weight may be assigned to the object shape or texture dimension.

[0131] The type of task is also an important consideration. For example, classification tasks may focus more on whether the data can clearly distinguish different categories, regression tasks may focus more on the numerical accuracy of the data, and generation tasks may focus more on the coherence and diversity of the data. Therefore, different initial weight parameters will be assigned according to different task types.

[0132] S702: Based on multiple initial weight parameters and at least one initial evaluation threshold, respectively determine corresponding initial target training data.

[0133] The initial evaluation threshold is a criterion for screening training data. There is at least one initial evaluation threshold in order to select initial target training data from different perspectives or using different screening criteria.

[0134] For each initial weight parameter combination, a weighted operation is performed on it and the evaluation parameters of each training data in multiple dimensions in the first training data set. For each set of initial weight parameters, there is a different weighted operation method to obtain a different comprehensive evaluation result.

[0135] Then, these comprehensive evaluation results are compared with the corresponding initial evaluation thresholds, and the data that meets the corresponding threshold conditions will be selected as the initial target training data. For example, for a set of initial weight parameters, if the comprehensive evaluation result of a certain training data calculated based on it is higher than the corresponding initial evaluation threshold, then this data will be included in the initial target training data set corresponding to the set of initial weight parameters. This means that multiple sets of initial target training data will be generated, each corresponding to a different combination of initial weight parameters and initial evaluation thresholds.

[0136] S703: Based on the initial target training data, train the second preset large language models respectively to obtain multiple target second models.

[0137] Among them, each group of initial target training data will be used as input to train the second preset large language model separately. Because different groups of initial target training data have different characteristics and screening criteria, they will have different training effects on different second preset large language models. During the training process, different second preset large language models will adjust their own parameters according to different groups of initial target training data and learn different features and patterns. Through multiple iterations and parameter updates, multiple target second models will be obtained, and each model is trained specifically according to different initial target training data.

[0138] S704: Evaluate the multiple target second models based on at least one evaluation index to obtain at least one evaluation index parameter of the multiple target second models.

[0139] Different evaluation indicators will reflect the effect of the target second model from different aspects. For each target second model, the test data is input into it, and its performance is calculated according to the corresponding evaluation indicator. The output of the model is compared with the true label or reference standard to obtain the corresponding evaluation indicator parameters.

[0140] S705. Obtain weight parameters and / or evaluation thresholds of multiple dimensions under the target training task according to at least one evaluation index parameter of the multiple target second models.

[0141] By analyzing the evaluation index parameters of multiple target second models, we can understand the model performance under different combinations of initial weight parameters and initial evaluation thresholds. Based on these performance indicators, we can determine which combination is better, and then adjust or determine the final weight parameters and evaluation thresholds.

[0142] For example, if the evaluation index parameters of a target second model indicate that it performs poorly in a certain dimension, it may mean that the corresponding initial weight parameters need to be adjusted; if the threshold for screening a set of initial target training data causes the model to overfit or underfit, the initial evaluation threshold can be adjusted. The ultimate goal is to find the optimal combination of weight parameters and evaluation thresholds based on the evaluation results of multiple models and taking into account various factors, so that the trained model has better performance on the target training task.

[0143] In the method provided in the embodiment of the present application, multiple initial weight parameters are determined by utilizing data type and task type, multiple groups of initial target training data are screened according to these parameters and initial evaluation thresholds, multiple target second models are trained, these models are evaluated, and weight parameters and evaluation thresholds are optimized based on the evaluation results, ultimately finding the most suitable combination of parameters and thresholds for the target training task, thereby improving the overall performance and effect of the training model.

[0144] The embodiment of the present application further provides another possible implementation of the pre-training data processing method, which obtains weight parameters and / or evaluation thresholds of multiple dimensions under the target training task according to at least one evaluation index parameter of multiple target second models, including:

[0145] Based on at least one evaluation index parameter of multiple target second models, the initial weight parameters and / or initial evaluation thresholds corresponding to the target second models whose evaluation index parameters meet preset conditions are determined as the weight parameters and / or evaluation thresholds of multiple dimensions under the target training task.

[0146] In this embodiment, first, multiple target second models are evaluated using at least one evaluation index, and corresponding evaluation index parameters are obtained. These evaluation index parameters can reflect the performance of each target second model on different tasks. Then, find the target second model with the best performance that meets the preset conditions. The best performance here may be based on different evaluation criteria. For example, for classification tasks, if the main focus is on comprehensive performance, it may be the model with the highest F1 value; if more attention is paid to recall rate, it may be the model with the highest recall rate. Once the best-performing target second model is determined, its corresponding initial weight parameters and / or initial evaluation thresholds are determined as the final weight parameters and / or evaluation thresholds of multiple dimensions under the target training task. This means that the training data screening and training process of the model is considered to be the most effective, and the initial weight parameters and initial evaluation thresholds used can bring the best performance to the target training task, so these parameters are directly used as the final result.

[0147] Alternatively, based on at least one evaluation index parameter of multiple target second models, the initial weight parameters and / or initial evaluation thresholds corresponding to the target second models whose evaluation index parameters meet the preset performance targets are determined as the weight parameters and / or evaluation thresholds of multiple dimensions under the target training task.

[0148] Specifically, the preset performance target is a performance standard set in advance according to the specific requirements of the task. For example, in an image recognition task, the accuracy rate may be set to 90% as the performance target; in a text generation task, the BLEU score may be set to 0.8 as the performance target.

[0149] For each target second model, its evaluation index parameter is compared with the preset performance target. When the evaluation index parameter of a target second model meets the preset performance target, the initial weight parameter and / or initial evaluation threshold corresponding to the model is determined as the final weight parameter and / or evaluation threshold. This method is relatively flexible. As long as the predetermined performance requirements are met, the corresponding initial parameters can be considered appropriate without necessarily pursuing the optimal performance, thereby improving processing efficiency.

[0150] Alternatively, fitting is performed based on at least one evaluation index parameter of multiple target second models and corresponding multiple initial weight parameters and at least one initial evaluation threshold to generate a performance weight fitting function.

[0151] Based on the performance weight fitting function, weight parameters and / or evaluation thresholds of multiple dimensions under the target training task are obtained.

[0152] Specifically, first, collect the evaluation index parameters of multiple target second models, as well as their corresponding initial weight parameters and initial evaluation thresholds. For example, for different target second models, there will be different evaluation index parameters (such as accuracy, recall, etc.), and record their initial weight parameters (such as weights of different dimensions) and initial evaluation thresholds used during training. Then, use these data for fitting. Fitting is a mathematical modeling method that aims to find a functional relationship, here is the performance weight fitting function, which can describe the relationship between the evaluation index parameters and the initial weight parameters and initial evaluation thresholds. Linear regression, nonlinear regression or other fitting methods can be used, depending on the distribution and relationship of the data.

[0153] Based on the performance weight fitting function, according to the specific optimization goal or task requirements, by solving the maximum value of the function or the value that meets certain conditions, the weight parameters and / or evaluation thresholds of multiple dimensions under the target training task are determined. For example, by finding the maximum value of the performance weight fitting function, the combination of weight parameters and evaluation thresholds that optimizes the performance can be found; or according to other constraints, a combination of parameters that meets certain performance requirements can be found.

[0154] In the method provided in the embodiment of the present application, the weight parameters and / or evaluation thresholds of multiple dimensions under the target training task are determined according to the evaluation index parameters of multiple target second models. Each method has its unique advantages and applicable scenarios. Selecting an appropriate method according to the specific situation can optimize the training data screening and model training process, and ultimately improve the performance of the target training task.

[0155] The following is a corresponding explanation of the pre-training data processing device provided by any of the above embodiments of the present application. Its specific implementation process and the technical effects produced are the same as those of the corresponding method embodiments mentioned above. For the sake of brief description, the parts not mentioned in this embodiment can refer to the corresponding contents in the method embodiments.

[0156] Figure 8 A schematic diagram of the functional modules of a pre-training data processing device provided in an embodiment of the present application. Figure 8 As shown, the pre-training data processing device 100 includes:

[0157] An acquisition module 110 is used to acquire a first training data set for a target training task;

[0158] An evaluation module 120 is used to evaluate each training data in the first training data set using a pre-trained multi-dimensional data evaluation model to obtain evaluation parameters of each training data in the first training data set in multiple dimensions;

[0159] The screening module 130 is used to screen target training data of a target training task from the first training data set according to evaluation parameters of each training data in multiple dimensions in the first training data set.

[0160] The above-mentioned device is used to execute the method provided by the aforementioned embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0161] The above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), or one or more microprocessors, or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0162] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A pre-training data processing method, characterized in that: The method comprises: Obtain a first training data set for a target training task; Using a pre-trained multi-dimensional data evaluation model, evaluate each training data in the first training data set to obtain evaluation parameters of each training data in the first training data set in multiple dimensions; Target training data for the target training task are screened from the first training data set according to evaluation parameters of the plurality of dimensions of the respective training data in the first training data set.

2. The method according to claim 1, characterized in that The pre-trained multi-dimensional data evaluation model is used to evaluate each training data in the training data set to obtain evaluation parameters of each training data in the training data set in multiple dimensions, including: Generate evaluation prompt words for each training data according to at least one of the description information of the multiple dimensions, the association relationship between different dimensions in the multiple dimensions, the scoring rules of the multiple dimensions under the target training task, the data type of each training data and the task type of the target training task; According to the evaluation prompt words, the multi-dimensional data evaluation model is used to evaluate the various training data to obtain evaluation parameters of the various training data in the multiple dimensions.

3. The method according to claim 2, characterized in that The multi-dimensional data evaluation model includes: a first preset large language model and a multi-layer perceptron; The step of evaluating each training data according to the evaluation prompt word using the multi-dimensional data evaluation model to obtain evaluation parameters of each training data in the multiple dimensions includes: According to the evaluation prompt word, the first preset large language model is used to process each training data to obtain embedded information of each training data; The multilayer perceptron is used to process the embedded information of each training data to obtain evaluation parameters of each training data in the multiple dimensions.

4. The method according to claim 3, characterized in that Before the embedded information of each training data is processed by the multilayer perceptron to obtain the evaluation parameters of each training data in the multiple dimensions, the method further includes: Obtain a second training data set for the multi-dimensional evaluation task; The multi-layer perceptron is trained using the second training data set to obtain the multi-dimensional data evaluation model.

5. The method according to claim 1, characterized in that The step of screening target training data for the target training task from the first training data set according to the evaluation parameters of the training data in the first training data set in the multiple dimensions includes: Based on the weight parameters of the multiple dimensions under the target training task, weighted operations are performed on the evaluation parameters of the multiple dimensions of each training data to obtain target evaluation parameters of each training data; According to the target evaluation parameters of each training data, training data that meets an evaluation threshold is determined from the first training data set as the target training data.

6. The method according to claim 5, characterized in that The method further comprises: According to the data type of each training data and / or the task type of the target training task, based on preset weight parameters and / or historical weight parameters, the weight parameters of the multiple dimensions under the target training task are determined.

7. The method according to claim 5, characterized in that The method further comprises: Determining initial weight parameters of the multiple dimensions under the target training task according to the data type of the respective training data and / or the task type of the target training task; Determining initial target training data based on the initial weight parameter and the initial evaluation threshold; Based on the initial target training data, training a second preset large language model to obtain a target second model; Evaluate the target second model based on at least one evaluation index to obtain at least one evaluation index parameter of the target second model; According to at least one of the evaluation index parameters, the initial weight parameter and / or the initial evaluation threshold is adjusted or determined to obtain the weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task.

8. The method according to claim 5, characterized in that The method further comprises: Determining a plurality of initial weight parameters of the plurality of dimensions under the target training task according to a data type of the respective training data and / or a task type of the target training task; Based on the multiple initial weight parameters and at least one initial evaluation threshold, respectively determine corresponding initial target training data; Based on the initial target training data, training the second preset large language models respectively to obtain multiple target second models; Evaluate the plurality of target second models based on at least one evaluation index to obtain at least one evaluation index parameter of the plurality of target second models; According to at least one of the evaluation index parameters of the multiple target second models, weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task are obtained.

9. The method according to claim 8, characterized in that The step of obtaining the weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task according to at least one evaluation index parameter of the multiple target second models includes: Based on at least one of the evaluation index parameters of the plurality of target second models, the initial weight parameter and / or initial evaluation threshold corresponding to the target second model whose evaluation index parameter meets the preset conditions is determined as the weight parameter and / or evaluation threshold of the plurality of dimensions under the target training task; or Based on at least one of the evaluation index parameters of the plurality of target second models, the initial weight parameter and / or initial evaluation threshold corresponding to the target second model whose evaluation index parameter meets the preset performance target is determined as the weight parameter and / or evaluation threshold of the plurality of dimensions under the target training task; or Based on at least one of the evaluation index parameters of the plurality of target second models and the corresponding plurality of initial weight parameters and at least one initial evaluation threshold, fitting is performed to generate a performance weight fitting function; Based on the performance weight fitting function, weight parameters and / or evaluation thresholds of the multiple dimensions under the target training task are obtained.

10. A pre-training data processing device, characterized in that: The device comprises: An acquisition module, used to acquire a first training data set for a target training task; An evaluation module, configured to evaluate each training data in the first training data set by using a pre-trained multi-dimensional data evaluation model, and obtain evaluation parameters of each training data in the first training data set in multiple dimensions; A screening module is used to screen the target training data of the target training task from the first training data set according to the evaluation parameters of the training data in the first training data set in the multiple dimensions.

Citation Information

Cited By

  • Reinforcement learning model training method and device based on dynamic evaluation indexes

    CN120875086A

  • A method and apparatus for training reinforcement learning models based on dynamic evaluation metrics

    CN120875086B

  • Power industry language model training data dynamic selection method, system and equipment

    CN121412671A

  • Method and system for screening high-scientific-value corpora from large-scale webpages

    CN121683780A