Processing method and device of model fine tuning data set, electronic equipment and readable medium

By dividing the model fine-tuning data set into sub-data sets and determining the target training steps, the model overfitting or underfitting caused by inappropriate data set types is solved, and the model's performance and overall training effect are improved under different tasks.

CN119939242APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411920003.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the process of fine-tuning of the model, the data set type proportion is inappropriate, which can easily lead to overfitting or underfitting of the model, affecting the training effect.

Method used

By dividing the model fine-tuning dataset into several subdatasets, each subdataset contains different types of training samples. These subdatasets are used to iteratively train the preset model, determine the target training steps of each subdataset, and adjust the overall training process of the model based on these steps.

Benefits of technology

Through the division of sub-data sets and the determination of target training steps, the model can perform well under different types of tasks, avoid overfitting and underfitting, and improve the overall training effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939242A_ABST
    Figure CN119939242A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model fine tuning data set processing method and device, electronic equipment and a readable medium. The method comprises the following steps: acquiring a model fine tuning data set; the model fine tuning data set comprises a plurality of sub-data sets, and each sub-data set comprises at least two types of training samples; for any sub-data set, adopting the sub-data set to carry out iteration training of a preset step number on a preset model, and obtaining an intermediate model of non-synchronization number training; determining a target training step number corresponding to the sub-data set based on the sub-data set and an intermediate model corresponding to the sub-data set; and if the difference value of the target training step numbers corresponding to the sub-data sets in the model fine tuning data set is not greater than a preset threshold value, taking the model fine tuning data set as a target model fine tuning data set. According to the method, the data set is divided into the different sub-data sets, the different sub-data sets are adjusted to complete model training under the approximate training step number, and therefore the model can have good performance under different types of tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of processing model fine-tuning datasets, and in particular to a method for processing a model fine-tuning dataset, a device for processing a model fine-tuning dataset, an electronic device, and a computer-readable medium. Background Art

[0002] Generally speaking, pre-trained models can be fine-tuned for specific tasks to improve their processing capabilities. At the same time, the model's data set can be divided into multiple different types based on different sources, different technical fields, different text styles, etc. The proportion of different types of data in the training process can affect the training effect of the model. If the proportion of data types is not appropriate, it may lead to overfitting and underfitting of the model, resulting in poor training effect of the model. Summary of the invention

[0003] The embodiment of the present invention provides a method, device, electronic device and computer-readable storage medium for processing a model fine-tuning data set to improve the training effect of the model.

[0004] The embodiment of the present invention discloses a method for processing a model fine-tuning data set, comprising:

[0005] Acquire a model fine-tuning dataset; the model fine-tuning dataset includes a plurality of sub-datasets, and the sub-datasets include at least two types of training samples;

[0006] For any of the sub-data sets, the sub-data set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps;

[0007] Determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset;

[0008] If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the model fine-tuning data set is used as the target model fine-tuning data set.

[0009] Optionally, the step of obtaining a model fine-tuning dataset includes:

[0010] Acquire an initial model data set, wherein the initial model data set includes at least two types of training samples;

[0011] Inputting the training sample into a pre-training model, and determining a forward activation value and a back propagation value corresponding to the training sample based on an output of the pre-training model for the training sample;

[0012] Based on the forward activation values ​​and back propagation values ​​corresponding to the training samples, the plurality of training samples are clustered to obtain a model fine-tuning data set including a plurality of sub-data sets.

[0013] Optionally, for any of the sub-data sets, the step of using the sub-data set to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps includes:

[0014] For any of the sub-datasets, dividing the sub-dataset into a sub-training set and a sub-validation set;

[0015] The sub-training set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

[0016] Optionally, the step of determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset includes:

[0017] Using the validation set, respectively determining the perplexities corresponding to the intermediate models of different steps;

[0018] The target number of training steps corresponding to the sub-dataset is determined according to the perplexity corresponding to the intermediate model of different steps.

[0019] Optionally, the step of using the validation set to respectively determine the perplexities corresponding to the intermediate models of different steps includes:

[0020] For the intermediate model of any step number, the training samples in the validation set are used to calculate the perplexity corresponding to the training samples respectively;

[0021] The mean of the perplexity corresponding to the training samples in the validation set is taken as the perplexity corresponding to the intermediate model of the current step.

[0022] Optionally, if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the step of using the model fine-tuning data set as the target model fine-tuning data set includes:

[0023] For the intermediate model of any step number, determining the weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset;

[0024] Determining the overall target number of training steps for the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers;

[0025] If the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than the second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than the first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset.

[0026] Optionally, the method further comprises:

[0027] If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, the proportion of different types of training samples in the sub-datasets is adjusted, and the steps of iteratively training a preset model for a preset number of steps using the sub-dataset for any of the sub-datasets to obtain an intermediate model trained with different steps, and determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset, are re-executed until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

[0028] An embodiment of the present invention further provides a processing device for a model fine-tuning data set, comprising:

[0029] A data set acquisition module, used to acquire a model fine-tuning data set; the model fine-tuning data set includes a plurality of sub-data sets, and the sub-data sets include at least two types of training samples;

[0030] A training module, for iteratively training a preset model with a preset number of steps using any of the sub-data sets, to obtain an intermediate model trained with different number of steps;

[0031] A training step confirmation module, used to determine a target training step corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset;

[0032] The target model fine-tuning data set determination module is used to use the model fine-tuning data set as the target model fine-tuning data set if the difference in the target number of training steps corresponding to the sub-data sets in the model fine-tuning data set is not greater than a first preset threshold.

[0033] Optionally, the data set acquisition module includes:

[0034] An initial data set acquisition submodule is used to acquire an initial model data set, wherein the initial model data set includes at least two types of training samples;

[0035] A model output submodule, used for inputting the training sample into the pre-training model, and determining the forward activation value and the back propagation value corresponding to the training sample based on the output of the pre-training model for the training sample;

[0036] The clustering submodule is used to cluster the plurality of training samples based on the forward activation values ​​and the back propagation values ​​corresponding to the training samples to obtain a model fine-tuning data set including a plurality of sub-data sets.

[0037] Optionally, the training module includes:

[0038] A data set division submodule, used for dividing any of the sub-data sets into a sub-training set and a sub-validation set;

[0039] The intermediate model acquisition submodule is used to use the sub-training set to iteratively train the preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

[0040] Optionally, the training step confirmation module includes:

[0041] A perplexity determination submodule, used to use the verification set to respectively determine the perplexities corresponding to the intermediate models of different steps;

[0042] The training step confirmation submodule determines the target number of training steps corresponding to the sub-data set according to the perplexity corresponding to the intermediate model of different step numbers.

[0043] Optionally, the perplexity determination submodule includes:

[0044] A sample perplexity calculation unit, for calculating the perplexity corresponding to each training sample using the training samples in the validation set for the intermediate model of any step number;

[0045] The perplexity calculation unit is used to use the mean of the perplexities corresponding to the training samples in the validation set as the perplexity corresponding to the intermediate model of the current step number.

[0046] Optionally, the target model fine-tuning data set determination module includes:

[0047] A weighted average perplexity calculation submodule, for determining, for an intermediate model of any number of steps, a weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset;

[0048] An overall target training step number determination submodule, used to determine the overall target training step number of the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers;

[0049] The target model fine-tuning dataset determination submodule is used to use the model fine-tuning dataset as the target model fine-tuning dataset if the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than a second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than a first preset threshold.

[0050] Optionally, the device further comprises:

[0051] A proportion adjustment module is used to adjust the proportions of different types of training samples in the sub-datasets if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, and re-execute the steps of iteratively training a preset model with the sub-dataset for a preset number of steps to obtain an intermediate model trained with different steps, and determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset, until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

[0052] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0053] The memory is used to store computer programs;

[0054] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.

[0055] The embodiment of the present invention further discloses one or more computer-readable media on which instructions are stored. When executed by one or more processors, the processors execute the method as described in the embodiment of the present invention.

[0056] The embodiments of the present invention include the following advantages:

[0057] Through the processing method of the model fine-tuning dataset provided by the embodiment of the present invention, a model fine-tuning dataset is obtained; the model fine-tuning dataset includes several sub-datasets, and the sub-datasets include at least two types of training samples; for any of the sub-datasets, the sub-dataset is used to iteratively train the preset model for a preset number of steps to obtain an intermediate model trained with different steps; based on the sub-dataset and the intermediate model corresponding to the sub-dataset, the target number of training steps corresponding to the sub-dataset is determined; if the difference in the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than a first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset. Thus, by dividing the dataset into different sub-datasets, adjusting different sub-datasets to complete the training of the model under close training steps, the model can have better performance under different types of tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of a method for processing a model fine-tuning data set provided in an embodiment of the present invention;

[0059] Figure 2 is a structural block diagram of a processing device for a model fine-tuning data set provided in an embodiment of the present invention;

[0060] Figure 3 is a block diagram of an electronic device provided in an embodiment of the present invention;

[0061] Figure 4 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0062] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0063] Reference Figure 1 , shows a flowchart of a method for processing a model fine-tuning dataset provided in an embodiment of the present invention, which may specifically include the following steps:

[0064] Step 101, obtaining a model fine-tuning dataset; the model fine-tuning dataset includes a plurality of sub-datasets, and the sub-datasets include at least two types of training samples;

[0065] In an embodiment of the present invention, a model fine-tuning dataset can be used to fine-tune a training model. The model fine-tuning dataset can be used to fine-tune the model. The model can be a model that can be used to process various tasks such as text translation, dialogue interaction, image recognition, etc., for example, a large language model, a convolutional neural network model, a recurrent neural network model, an encoder-decoder model, etc., and the present invention is not limited to this.

[0066] Model fine-tuning training can refer to further training on a specific task or dataset based on a pre-trained model to improve the model's performance on that task. For example, for a model that can handle multiple different tasks such as text translation, conversation interaction, email writing, and weekly report writing at the same time, further training can be performed on text translation to improve the model's performance on text translation tasks.

[0067] In an embodiment of the present invention, the model fine-tuning dataset may include at least two types of training samples. The types of samples may be divided based on the source of the samples, technical field, text style, etc., which is not limited in the present invention.

[0068] At the same time, in the embodiment of the present invention, in order to reduce the difficulty of adjusting the data set, the model fine-tuning data set can be divided into several sub-data sets, and the sub-data sets include at least two types of training samples. Therefore, in the subsequent training process, the more appropriate number of training steps can be determined for different sub-data sets, so that compared with adjusting the proportion of different types of training samples in the complete model fine-tuning data set at the same time, the adjustment difficulty can be reduced to a certain extent.

[0069] Step 102, for any of the sub-data sets, using the sub-data set to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps;

[0070] In an embodiment of the present invention, for any sub-data set, the sub-data set can be used to perform iterative training of a preset number of steps on a preset model. The training sample in the sub-data set can record the input data that needs to be input into the model, and the standard output corresponding to the input data. Thus, the input data in the training sample can be input into the model, and the output of the model can be obtained. The similarity between the model output and the standard output in the training sample is compared, and the preset model is trained again for the sub-data set based on the comparison result. So that the output of the model can be close to the standard output.

[0071] The number of steps for model training can be set in advance, and the model can be iteratively trained for the preset number of steps. During each step of training, an intermediate model corresponding to the step can be obtained, so that after completing the iterative training for the preset number of steps, intermediate models trained for different steps can be obtained.

[0072] In a specific implementation, when the number of iterations is small, the intermediate model obtained by training is usually underfitting, that is, the similarity between the model output and the standard output in the training sample is low, and when the number of iterations is high, the intermediate model obtained by training is usually overfitting, that is, the similarity between the model output and the standard output of the training sample in the current data set is too high, resulting in the model being unable to perform well in other data sets. Therefore, it is not necessary to obtain the intermediate models corresponding to all the steps, but it is possible to obtain the intermediate models of specific steps, such as the 10th to 20th rounds, the 10th to 30th rounds, the 15th to 25th rounds, etc., and the present invention does not limit this.

[0073] At the same time, since the intermediate models obtained between adjacent iteration steps can have similar model performances, it is also possible to extract several specific steps of intermediate models from the preset step range and retain them according to actual needs. For example, in the 10th to 20th rounds, the intermediate models of the 10th, 12th, 14th, 16th, 18th, and 20th rounds are retained, and the present invention does not limit this.

[0074] Step 103, determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset;

[0075] In a specific implementation, based on the sub-dataset and the intermediate model corresponding to the sub-dataset, the intermediate model that performs well on the sub-dataset, that is, the intermediate model with a high similarity between the model output and the standard output in the training sample, can be analyzed, and the number of steps corresponding to the intermediate model can be used as the target number of training steps for the sub-dataset. Under the target number of training steps, the model can achieve better performance on the sub-dataset.

[0076] Step 104: If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the model fine-tuning data set is used as a target model fine-tuning data set.

[0077] In the embodiment of the present invention, in order to avoid overfitting or underfitting of some types of training samples during the training process, it is necessary to make the sub-datasets in the model fine-tuning data set have similar target training steps. In this case, the model can obtain better performance on different sub-datasets, and can avoid the situation where too many data sets appear overfitting or underfitting, so that the model can obtain better performance on the model fine-tuning data set as a whole. In this case, fine-tuning the model using the model fine-tuning data set can make the fine-tuned model have better performance on specific tasks.

[0078] Thus, a first preset threshold can be preset, and the difference between the target training steps corresponding to different sub-datasets in the model fine-tuning dataset can be calculated. If there is no case where the difference between the target training steps is greater than the preset threshold, it can be considered that the target training steps corresponding to the sub-datasets are relatively close at this time, and the model has achieved better performance in the model fine-tuning dataset as a whole. At this time, the ratio of different types of training samples in the sub-datasets is appropriate, and the current model fine-tuning dataset can be used as the target model fine-tuning dataset.

[0079] If the difference in the target training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than the first preset threshold, it can be considered that the target training steps corresponding to the sub-datasets have a certain gap. At this time, using the model fine-tuning data set to train the model may cause underfitting or overfitting of some types of data. At this time, it is necessary to adjust the sub-datasets with a large difference in target training steps from other sub-datasets, modify the proportion of different types of training samples contained therein, so that the target training steps of the sub-datasets can be closer, and obtain a model fine-tuning data set suitable for fine-tuning training.

[0080] Through the processing method of the model fine-tuning dataset provided by the embodiment of the present invention, a model fine-tuning dataset is obtained; the model fine-tuning dataset includes several sub-datasets, and the sub-datasets include at least two types of training samples; for any of the sub-datasets, the sub-dataset is used to iteratively train the preset model for a preset number of steps to obtain an intermediate model trained with different steps; based on the sub-dataset and the intermediate model corresponding to the sub-dataset, the target number of training steps corresponding to the sub-dataset is determined; if the difference in the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than a first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset. Thus, by dividing the dataset into different sub-datasets, adjusting different sub-datasets to complete the training of the model under close training steps, the model can have better performance under different types of tasks.

[0081] In one embodiment of the present invention, the step of obtaining a model fine-tuning dataset includes:

[0082] S11, obtaining an initial model data set, where the initial model data set includes at least two types of training samples;

[0083] Generally speaking, when the model parameters are adjusted, inputting training samples from different data sets into the model may produce the same output results as the original ones, or may produce different output results.

[0084] At the same time, when the same model is used for different data sets, if there is a data set that adjusts the ratio of the types of training samples it contains, it may cause changes in the model's parameters, which may further affect the performance of the model on other data sets, causing the model to perform worse on other data, making it difficult to adjust the ratio of the types of training samples in the data set.

[0085] To this end, in the process of dividing the model fine-tuning dataset into sub-datasets, the embodiments of the present invention can divide the sub-datasets based on the correlation between the sub-datasets so that the sub-datasets do not affect each other as much as possible. That is, when the type ratio of the training samples included in some datasets is changed, the performance of the model on other sub-datasets will not be affected, so that the type ratio of the training samples in different sub-datasets can be adjusted independently, thereby reducing the difficulty of adjusting the type ratio of the training samples in the model fine-tuning dataset.

[0086] Thus, an initial model data set may be obtained first. The initial model data set may include training samples of at least two categories, and sub-data sets have not yet been divided.

[0087] S12, inputting the training sample into a pre-training model, and determining a forward activation value and a back propagation value corresponding to the training sample based on an output of the pre-training model for the training sample;

[0088] The input data in the training sample can be input into the pre-training model to obtain the output of the pre-training model for the training sample. Thereafter, the forward activation value and the back propagation value corresponding to the training sample can be determined.

[0089] Specifically, the model can usually be a multi-layer structure. For example, the model can include an input layer, an embedding layer, a fully connected layer, a convolutional layer, a loop layer, a pooling layer, etc. The different layers in the model can be connected in sequence. After the data is input into the model, the first layer in the model can process the input data to obtain a feature data, and each subsequent layer can process the feature data output by the previous layer in sequence to obtain the model output. Afterwards, the model output can be compared with the standard output in the training sample, and a loss function for measuring the difference between the model output and the standard output is calculated, and the model parameters in the model are updated based on the loss function to reduce the difference between the model output and the standard output.

[0090] The forward activation value corresponding to the training sample may be the output value of each layer in the model during the forward propagation process of the model. The reverse gradient value corresponding to the training sample may refer to the gradient of the model parameters of each layer relative to the loss function during the process of updating the model parameters through the back propagation of the model.

[0091] S13, clustering the plurality of training samples based on the forward activation values ​​and the back propagation values ​​corresponding to the training samples to obtain a model fine-tuning data set including a plurality of sub-data sets.

[0092] In the embodiment of the present invention, it is possible to determine whether the training samples will affect each other based on the forward activation values ​​and back propagation values ​​corresponding to the training samples, and cluster the training samples.

[0093] Specifically, the distance between different training samples can be calculated based on the forward activation value and the back propagation value corresponding to the training samples. If the distance is small, it is believed that the training samples may affect each other, and if the distance is large, it is believed that the training samples have little influence on each other.

[0094] Training samples that may affect each other are placed in the same class, and training samples that have less impact on each other are placed in different classes. This allows the training samples to be divided into multiple sub-datasets with less impact between different sub-datasets.

[0095] Specifically, for the training sample x j , and its forward activation value is a j , the backward gradient value is g j . Training sample x j and x k The distance metric d(x j ,x k ) can be expressed as:

[0096]

[0097] Here, dot() is the dot product. Based on this distance metric, the Kmeans clustering algorithm is used to cluster all training samples into K categories, denoted as D i , (i=1,...K). The clustered sub-datasets can minimize the mutual influence.

[0098] This is because, remember x j ,x k Input data to the model, take the lth layer of the model, and abstract it into a linear matrix W l , assuming that the input of this layer is a j =F(x j ),a k =F(x k ), where F is the calculation function of the previous layer, then the output of the model at layer L is y j =W L a j ,y k =W L a k , assuming W L In the data xj The gradient is calculated and updated on for:

[0099]

[0100] In this embodiment of the present invention, the updated model is used for the data x k The influence of needs to be as small as possible, which is equivalent to the original L-th layer input a k After the update, the Lth layer W L The output remains the same as before, so we have:

[0101]

[0102] Then we can get ΔW L a j ≈0, that is, when the data x k The update direction of the Lth layer weight is related to another data x j When the input of layer i is orthogonal, change the data x k The ratio of the data x j The smaller the impact.

[0103] In one embodiment of the present invention, the step of iteratively training a preset model with a preset number of steps using the sub-dataset for any of the sub-datasets to obtain an intermediate model trained with different number of steps includes:

[0104] S21, for any of the sub-data sets, dividing the sub-data set into a sub-training set and a sub-validation set;

[0105] S22, using the sub-training set to iteratively train the preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

[0106] In the specific implementation, for any sub-dataset, the sub-dataset can be first divided into a sub-training set and a sub-validation set. The sub-training set is used to perform iterative training of a preset model for a preset number of steps to obtain an intermediate model trained with different steps. The sub-validation set can be used to verify the training effect of the model.

[0107] The proportion of the sub-training set and the sub-validation set can be determined according to actual needs. In a specific example of the present invention, 99% of the training samples in the sub-data set can be set as the sub-training set, and 1% of the training samples can be set as the sub-validation set.

[0108] In one embodiment of the present invention, the step of determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset includes:

[0109] S31, using the verification set, respectively determining the perplexities corresponding to the intermediate models of different steps;

[0110] S32, determining the target number of training steps corresponding to the sub-dataset according to the perplexity corresponding to the intermediate model of different step numbers.

[0111] Based on the sub-dataset and the intermediate model corresponding to the sub-dataset, we can analyze the intermediate model that performs well on the sub-dataset, that is, the model output has a high similarity with the standard output in the training sample, and use the number of steps corresponding to the intermediate model as the target number of training steps for the sub-dataset. Under the target number of training steps, the model can achieve good performance on the sub-dataset.

[0112] In the specific implementation, perplexity can be used to represent the performance of different intermediate models on the sub-datasets. The perplexity can be a quantification of the model's prediction ability. The lower the perplexity, the higher the similarity between the model output and the standard output in the training sample, and the more accurate the model's prediction of the test data.

[0113] Thus, the training samples in the validation set can be input into the intermediate model, and the perplexities corresponding to the intermediate models of different steps can be determined respectively. And according to the perplexities corresponding to the intermediate models of different steps, the intermediate model with lower perplexity than other intermediate models is selected as the better model, and the corresponding number of training steps is used as the target number of training steps corresponding to the sub-data set.

[0114] In one embodiment of the present invention, the step of using the validation set to respectively determine the perplexities corresponding to the intermediate models of different steps includes:

[0115] S41, for the intermediate model of any step number, using the training samples in the validation set, respectively calculating the perplexity corresponding to the training samples;

[0116] S42, taking the mean of the perplexities corresponding to the training samples in the validation set as the perplexity corresponding to the intermediate model of the current step number.

[0117] In a specific implementation, for an intermediate model of any number of steps, the perplexity corresponding to the training samples in the validation set can be first calculated respectively.

[0118] Specifically, the perplexity corresponding to the training sample can be expressed as:

[0119]

[0120] Among them, w i represents the i-th token in the training sample, and N is the number of tokens in the training sample.

[0121] Thereafter, the mean of the perplexities corresponding to the training samples in the validation set may be used as the perplexity corresponding to the intermediate model of the current step.

[0122] Specifically, the mean perplexity of the jth intermediate model corresponding to the tth round under the i-th validation set can be expressed as:

[0123]

[0124] Where i∈{1, 2, …, n}, j∈{1, 2, …, m}, and L(·) represents the number of data items in the data set.

[0125] For intermediate models with different steps, this method can be used to determine the perplexity corresponding to the intermediate model, so that for a sub-data set, the perplexity corresponding to the intermediate model with different steps can be obtained.

[0126] In one embodiment of the present invention, if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning dataset is not greater than a first preset threshold, the step of using the model fine-tuning dataset as the target model fine-tuning dataset includes:

[0127] S51, for an intermediate model of any step number, determining a weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset;

[0128] In an embodiment of the present invention, in order to enable the sub-datasets in the model fine-tuning data set to complete training at the same time and obtain better training results, the overall training degree of the model fine-tuning data set at different numbers of steps can be determined first.

[0129] Therefore, for any number of steps, the weighted average perplexity of the model fine-tuning dataset can be determined according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset. The weighted average perplexity is calculated as:

[0130]

[0131] S52, determining the overall target number of training steps for the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers;

[0132] In the embodiment of the present invention, in order to allow the sub-datasets in the model fine-tuning data set to complete training at the same time and obtain a good training effect, it is necessary to make the overall perplexity of the model fine-tuning data set smaller. Therefore, according to the weighted average perplexity of the model fine-tuning data set corresponding to different step numbers, the number of steps with a smaller weighted average perplexity can be selected as the overall target training step number of the model fine-tuning data set. At this time, for the model fine-tuning data set, its overall training degree is better.

[0133] In the specific implementation, the overall target number of training steps can be expressed as:

[0134]

[0135] S53: If the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than the second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than the first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset.

[0136] Afterwards, in order to enable the model to obtain better training results, the target training steps corresponding to different sub-datasets can be close to the overall target training steps of the model fine-tuning dataset.

[0137] Thus, the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset can be calculated. If the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than the second preset threshold, it can be considered that the target number of training steps corresponding to the sub-dataset is relatively close to the overall target number of training steps of the model fine-tuning dataset.

[0138] At the same time, a first preset threshold can be preset, and the difference between the target training steps corresponding to different sub-datasets in the model fine-tuning dataset can be calculated. If there is no case where the difference between the target training steps is greater than the preset threshold, it can be considered that the target training steps corresponding to the sub-datasets are relatively close at this time, and the model has achieved better performance in the model fine-tuning dataset as a whole. At this time, the ratio of different types of training samples in the sub-datasets is appropriate, and the current model fine-tuning dataset can be used as the target model fine-tuning dataset.

[0139] In one embodiment of the present invention, the method further comprises:

[0140] S61: If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, the proportion of different types of training samples in the sub-datasets is adjusted, and the steps of iteratively training a preset model for a preset number of steps using the sub-datasets for any of the sub-datasets are re-executed to obtain intermediate models trained with different numbers of steps, and determining the target number of training steps corresponding to the sub-datasets based on the sub-datasets and the intermediate models corresponding to the sub-datasets, until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

[0141] If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than the first preset threshold, it can be considered that the target number of training steps corresponding to the sub-datasets has a certain gap. At this time, using the model fine-tuning data set to train the model may cause underfitting or overfitting of some types of data. At this time, it is necessary to adjust the sub-datasets and modify the proportion of different types of training samples contained therein so that the target number of training steps of the sub-datasets can be closer, and a model fine-tuning data set suitable for fine-tuning training is obtained.

[0142] In the specific implementation, in the experimental round t, the set of perplexities of each group of models with different synchronization numbers in the validation set of the i-th sub-dataset is Use cubic spline interpolation to fit a curve of perplexity with respect to the number of training steps Let the lowest point of the curve be This lowest point represents the target number of training steps for the best model fit on the i-th sub-dataset. Similarly, the set of perplexities for each model on the full dataset is Fit a curve and set the lowest point of the curve to

[0143] Secondly, calculate the proportion of tokens in each sub-dataset in the total training data set:

[0144]

[0145] Where T(·) represents the number of tokens in the dataset. satisfy

[0146] In order to make the model gather at the position of the overall target training step number at the lowest point of the sub-validation set confusion corresponding to each sub-dataset, the ratio can be corrected according to the iterative steps of the lowest point of the validation set confusion of each sub-dataset in the tth round of experiments, so that the trough of the validation set confusion curve of each data set As close as possible to the trough of the overall validation set perplexity curve Overlap, that is, "trough convergence", the ratio correction formula is as follows:

[0147]

[0148] Among them, k=10 and μ=15000 are generally taken.

[0149] Normalized ratio,

[0150]

[0151] Back-calculate the weight of each sub-dataset

[0152]

[0153] In this embodiment, in order to prevent the data set with fewer data items from losing too much data when synthesizing the full training set due to the lower new weight, the weight can be increased proportionally according to the size of the data set with the least data.

[0154]

[0155] The new weight of the sub-dataset with the smallest amount of data is 1.

[0156] Finally, let t = t + 1 and perform a new round of fine-tuning training.

[0157] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0158] Reference Figure 2 , shows a structural block diagram of a processing device for a model fine-tuning dataset provided in an embodiment of the present invention, which may specifically include the following modules:

[0159] The data set acquisition module 201 is used to acquire a model fine-tuning data set; the model fine-tuning data set includes a plurality of sub-data sets, and the sub-data sets include at least two types of training samples;

[0160] A training module 202 is used for iteratively training a preset model with a preset number of steps using any of the sub-data sets to obtain an intermediate model trained with different number of steps;

[0161] A training step confirmation module 203 is used to determine a target training step corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset;

[0162] The target model fine-tuning dataset determination module 204 is configured to use the model fine-tuning dataset as the target model fine-tuning dataset if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning dataset is not greater than a first preset threshold.

[0163] Optionally, the data set acquisition module includes:

[0164] An initial data set acquisition submodule is used to acquire an initial model data set, wherein the initial model data set includes at least two types of training samples;

[0165] A model output submodule, used for inputting the training sample into the pre-training model, and determining the forward activation value and the back propagation value corresponding to the training sample based on the output of the pre-training model for the training sample;

[0166] The clustering submodule is used to cluster the plurality of training samples based on the forward activation values ​​and the back propagation values ​​corresponding to the training samples to obtain a model fine-tuning data set including a plurality of sub-data sets.

[0167] Optionally, the training module includes:

[0168] A data set division submodule, used for dividing any of the sub-data sets into a sub-training set and a sub-validation set;

[0169] The intermediate model acquisition submodule is used to use the sub-training set to iteratively train the preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

[0170] Optionally, the training step confirmation module includes:

[0171] A perplexity determination submodule, used to use the verification set to respectively determine the perplexities corresponding to the intermediate models of different steps;

[0172] The training step confirmation submodule determines the target number of training steps corresponding to the sub-data set according to the perplexity corresponding to the intermediate model of different step numbers.

[0173] Optionally, the perplexity determination submodule includes:

[0174] A sample perplexity calculation unit, for calculating the perplexity corresponding to each training sample using the training samples in the validation set for the intermediate model of any step number;

[0175] The perplexity calculation unit is used to use the mean of the perplexities corresponding to the training samples in the validation set as the perplexity corresponding to the intermediate model of the current step number.

[0176] Optionally, the target model fine-tuning data set determination module includes:

[0177] A weighted average perplexity calculation submodule, for determining, for an intermediate model of any number of steps, a weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset;

[0178] An overall target training step number determination submodule, used to determine the overall target training step number of the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers;

[0179] The target model fine-tuning dataset determination submodule is used to use the model fine-tuning dataset as the target model fine-tuning dataset if the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than a second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than a first preset threshold.

[0180] Optionally, the device further comprises:

[0181] A proportion adjustment module is used to adjust the proportions of different types of training samples in the sub-datasets if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, and re-execute the steps of iteratively training a preset model with the sub-dataset for a preset number of steps to obtain an intermediate model trained with different steps, and determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset, until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

[0182] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0183] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0184] Memory 303, used for storing computer programs;

[0185] The processor 301 is used to execute the program stored in the memory 303 to implement the following steps:

[0186] Acquire a model fine-tuning dataset; the model fine-tuning dataset includes a plurality of sub-datasets, and the sub-datasets include at least two types of training samples;

[0187] For any of the sub-data sets, the sub-data set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps;

[0188] Determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset;

[0189] If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the model fine-tuning data set is used as the target model fine-tuning data set.

[0190] Optionally, the step of obtaining a model fine-tuning dataset includes:

[0191] Acquire an initial model data set, wherein the initial model data set includes at least two types of training samples;

[0192] Inputting the training sample into a pre-training model, and determining a forward activation value and a back propagation value corresponding to the training sample based on an output of the pre-training model for the training sample;

[0193] Based on the forward activation values ​​and back propagation values ​​corresponding to the training samples, the plurality of training samples are clustered to obtain a model fine-tuning data set including a plurality of sub-data sets.

[0194] Optionally, for any of the sub-data sets, the step of using the sub-data set to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps includes:

[0195] For any of the sub-datasets, dividing the sub-dataset into a sub-training set and a sub-validation set;

[0196] The sub-training set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

[0197] Optionally, the step of determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset includes:

[0198] Using the validation set, respectively determining the perplexities corresponding to the intermediate models of different steps;

[0199] The target number of training steps corresponding to the sub-dataset is determined according to the perplexity corresponding to the intermediate model of different steps.

[0200] Optionally, the step of using the validation set to respectively determine the perplexities corresponding to the intermediate models of different steps includes:

[0201] For the intermediate model of any step number, the training samples in the validation set are used to calculate the perplexity corresponding to the training samples respectively;

[0202] The mean of the perplexity corresponding to the training samples in the validation set is taken as the perplexity corresponding to the intermediate model of the current step.

[0203] Optionally, if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the step of using the model fine-tuning data set as the target model fine-tuning data set includes:

[0204] For the intermediate model of any step number, determining the weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset;

[0205] Determining the overall target number of training steps for the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers;

[0206] If the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than the second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than the first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset.

[0207] Optionally, the method further comprises:

[0208] If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, the proportion of different types of training samples in the sub-datasets is adjusted, and the steps of iteratively training a preset model for a preset number of steps using the sub-dataset for any of the sub-datasets to obtain an intermediate model trained with different steps, and determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset, are re-executed until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

[0209] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0210] The communication interface is used for communication between the above terminal and other devices.

[0211] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0212] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0213] like Figure 4 As shown, in another embodiment provided by the present invention, a computer-readable storage medium 401 is also provided, in which instructions are stored. When the computer-readable storage medium 401 is run on a computer, the computer executes the method for processing the model fine-tuning data set described in the above embodiment.

[0214] In another embodiment provided by the present invention, a computer program product containing instructions is also provided, which, when executed on a computer, enables the computer to execute the method for processing the model fine-tuning dataset described in the above embodiment.

[0215] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0216] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0217] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0218] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A method for processing a model fine-tuning dataset, characterized in that: include: Acquire a model fine-tuning dataset; the model fine-tuning dataset includes a plurality of sub-datasets, and the sub-datasets include at least two types of training samples; For any of the sub-data sets, the sub-data set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps; Determining a target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset; If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than a first preset threshold, the model fine-tuning data set is used as the target model fine-tuning data set.

2. The method according to claim 1, characterized in that The step of obtaining a model fine-tuning dataset includes: Acquire an initial model data set, wherein the initial model data set includes at least two types of training samples; Inputting the training sample into a pre-training model, and determining a forward activation value and a back propagation value corresponding to the training sample based on an output of the pre-training model for the training sample; Based on the forward activation values ​​and back propagation values ​​corresponding to the training samples, the plurality of training samples are clustered to obtain a model fine-tuning data set including a plurality of sub-data sets.

3. The method according to claim 1, characterized in that The step of iteratively training a preset model with a preset number of steps using any of the sub-data sets to obtain an intermediate model trained with different number of steps includes: For any of the sub-datasets, dividing the sub-dataset into a sub-training set and a sub-validation set; The sub-training set is used to iteratively train a preset model for a preset number of steps to obtain an intermediate model trained with different number of steps.

4. The method according to claim 3, characterized in that The step of determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset includes: Using the validation set, respectively determining the perplexities corresponding to the intermediate models of different steps; The target number of training steps corresponding to the sub-dataset is determined according to the perplexity corresponding to the intermediate model of different steps.

5. The method according to claim 4, characterized in that The step of using the validation set to respectively determine the perplexity corresponding to the intermediate models of different steps includes: For the intermediate model of any step number, the training samples in the validation set are used to calculate the perplexity corresponding to the training samples respectively; The mean of the perplexity corresponding to the training samples in the validation set is taken as the perplexity corresponding to the intermediate model of the current step.

6. The method according to claim 4, characterized in that The step of using the model fine-tuning dataset as the target model fine-tuning dataset if the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning dataset is not greater than a first preset threshold comprises: For the intermediate model of any step number, determining the weighted average perplexity of the model fine-tuning dataset according to the perplexity corresponding to the sub-dataset in the model fine-tuning dataset; Determining the overall target number of training steps for the model fine-tuning dataset according to the weighted average perplexity of the model fine-tuning dataset corresponding to different step numbers; If the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset and the overall target number of training steps of the model fine-tuning dataset is not greater than the second preset threshold, and the difference between the target number of training steps corresponding to the sub-dataset in the model fine-tuning dataset is not greater than the first preset threshold, the model fine-tuning dataset is used as the target model fine-tuning dataset.

7. The method according to claim 1, characterized in that The method further comprises: If the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is greater than a first preset threshold, the proportion of different types of training samples in the sub-datasets is adjusted, and the steps of iteratively training a preset model for a preset number of steps using the sub-dataset for any of the sub-datasets to obtain an intermediate model trained with different steps, and determining the target number of training steps corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset, are re-executed until the difference in the target number of training steps corresponding to the sub-datasets in the model fine-tuning data set is not greater than the first preset threshold.

8. A processing device for a model fine-tuning data set, characterized in that: include: A data set acquisition module, used to acquire a model fine-tuning data set; the model fine-tuning data set includes a plurality of sub-data sets, and the sub-data sets include at least two types of training samples; A training module, for iteratively training a preset model with a preset number of steps using any of the sub-data sets, to obtain an intermediate model trained with different number of steps; A training step confirmation module, used to determine a target training step corresponding to the sub-dataset based on the sub-dataset and the intermediate model corresponding to the sub-dataset; The target model fine-tuning data set determination module is used to use the model fine-tuning data set as the target model fine-tuning data set if the difference in the target number of training steps corresponding to the sub-data sets in the model fine-tuning data set is not greater than a first preset threshold.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Training method of business data model

    CN120372298A