Model training method and device, equipment, storage medium and program product
By optimizing model parameters through noise processing and predictive label quality screening of pretrained models, the problem of overfitting pretrained models is solved, and the generalization ability and accuracy of the model on specific tasks is improved.
Patent Information
- Application Number
- CN202510444299.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
The performance of existing pretrained models on specific tasks is easily affected by overfitting training data, resulting in poor generalization ability and low model accuracy.
By noise processing of the training samples of the target task, input the initial task processing model, adjust the model parameters, combine unsupervised training and predicted label quality screening, the model is optimized to improve generalization ability and accuracy.
It reduces the model's overfitting of the training data, improves the generalization ability of the model when facing new inputs and the quality of the output results, and comprehensively improves the model's performance in the target tasks.
Smart Images

Figure CN120354969A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a model training method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] With the development of computer technologies, pre-trained models have emerged. A pre-trained model is an artificial intelligence model pre-trained on a large-scale dataset and has the ability to learn general features and adapt to multiple tasks. For example, a pre-trained language model is an artificial intelligence model trained based on a large-scale text dataset.
[0003] To improve the performance of a pre-trained model on a specific task, the pre-trained model can be fine-tuned based on the dataset corresponding to the specific task. Currently, the pre-trained model is usually fine-tuned on a labeled dataset so that the pre-trained model can provide more accurate results in a specific task. However, the current training method is likely to cause the model to overfit the training data, resulting in poor generalization ability of the model when facing new inputs and low model accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide a model training method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve model accuracy for the above technical problems.
[0005] This application provides a model training method, including:
[0006] Obtain a first training sample and a second training sample corresponding to a target task, and a training label corresponding to the first training sample;
[0007] Input the first training sample after noise processing into an initial task processing model corresponding to the target task to obtain a first predicted label corresponding to the first training sample;
[0008] Based on the first predicted label and the training label corresponding to the first training sample, adjust the model parameters of the initial task processing model to obtain an updated task processing model;
[0009] Input the second training sample into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample;
[0010] Based on the label quality corresponding to the second predicted labels, determine positive predicted labels and negative predicted labels from the multiple second predicted labels;
[0011] Based on the positive prediction labels and the negative prediction labels, adjust the model parameters of the updated task processing model to obtain the target task processing model corresponding to the target task.
[0012] This application also provides a model training device, including:
[0013] A training data acquisition module, configured to acquire a first training sample and a second training sample corresponding to a target task, and training labels corresponding to the first training sample;
[0014] A first training module, configured to input the first training sample into an initial task processing model corresponding to the target task after noise processing to obtain a first prediction label corresponding to the first training sample;
[0015] The first training module is further configured to adjust the model parameters of the initial task processing model based on the first prediction label and the training labels corresponding to the first training sample to obtain an updated task processing model;
[0016] A second training module, configured to input the second training sample into the updated task processing model to obtain a plurality of second prediction labels corresponding to the second training sample;
[0017] The second training module is further configured to determine positive prediction labels and negative prediction labels from the plurality of second prediction labels based on the label quality corresponding to the second prediction labels;
[0018] The second training module is further configured to adjust the model parameters of the updated task processing model based on the positive prediction labels and the negative prediction labels to obtain the target task processing model corresponding to the target task.
[0019] This application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the above model training method when executing the computer program.
[0020] This application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the above model training method when being executed by a processor.
[0021] This application also provides a computer program product, including a computer program, and the computer program implements the steps of the above model training method when being executed by a processor.
[0022] The above model training method, device, computer device, storage medium and computer program product obtain a first training sample and a second training sample corresponding to a target task, and a training label corresponding to the first training sample; input the first training sample after noise processing into an initial task processing model corresponding to the target task to obtain a first prediction label corresponding to the first training sample; based on the first prediction label and the training label corresponding to the first training sample, adjust the model parameters of the initial task processing model to obtain an updated task processing model; input the second training sample into the updated task processing model to obtain multiple second prediction labels corresponding to the second training sample; based on the label quality corresponding to the second prediction labels, determine positive prediction labels and negative prediction labels from the multiple second prediction labels; based on the positive prediction labels and the negative prediction labels, adjust the model parameters of the updated task processing model to obtain a target task processing model corresponding to the target task. In this way, first, based on the first training sample corresponding to the target task and the training label corresponding to the first training sample, perform supervised training on the initial task processing model corresponding to the target task to obtain an updated task processing model. During the supervised training process, the first training sample is input into the model after noise processing. The addition of noise helps reduce the overfitting of the model to the training data, preventing the model from overlearning specific patterns and details in the training data, thereby improving the generalization ability of the model when facing the target task. Further, perform unsupervised training on the updated task processing model based on the second training sample to obtain a target task processing model corresponding to the target task. In the unsupervised training, input the second training sample into the updated task processing model to obtain multiple second prediction labels, select the second prediction labels with higher label quality as positive prediction labels, select the second prediction labels with lower label quality as negative prediction labels, and fine-tune the updated task processing model based on the positive prediction labels and the negative prediction labels, which helps the model learn how to make better decisions and improve the quality of the model output results. Combining the addition of noise and the selection of prediction labels can comprehensively improve the performance of the model in the target task and improve the accuracy of the model. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 It is an application environment diagram of the model training method in an embodiment;
[0025] Figure 2 It is a flowchart of the model training method in an embodiment;
[0026] Figure 3 Schematic diagram of training a language model in one embodiment;
[0027] Figure 4 Schematic diagram of training a task processing model in one embodiment;
[0028] Figure 5 Schematic diagram of training a language model in another embodiment;
[0029] Figure 6 Schematic diagram of training a task processing model in another embodiment;
[0030] Figure 7 Schematic diagram of training a question - answering model in one embodiment;
[0031] Figure 8 Schematic diagram of training a content review model in one embodiment;
[0032] Figure 9 Schematic diagram of the process of model application in one embodiment;
[0033] Figure 10 Schematic diagram of training a language model in another embodiment;
[0034] Figure 11 Schematic diagram of the news review interaction process in one embodiment;
[0035] Figure 12 Structural block diagram of a model training device in one embodiment;
[0036] Figure 13 Internal structure diagram of a computer device in one embodiment;
[0037] Figure 14 Internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0039] The model training method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed in the cloud or other devices.
[0040] Specifically, the server 104 obtains the first training sample and the second training sample corresponding to the target task, and the training label corresponding to the first training sample, and trains the initial task processing model corresponding to the target task based on these data to obtain the target task processing model corresponding to the target task. First, the server 104 performs supervised training on the initial task processing model corresponding to the target task to obtain an updated task processing model. The first training sample is input into the initial task processing model corresponding to the target task after noise processing to obtain the first predicted label corresponding to the first training sample. Based on the first predicted label and the training label corresponding to the first training sample, the model parameters of the initial task processing model are adjusted to obtain the updated task processing model. Further, the server 104 performs unsupervised training on the updated task processing model to obtain the target task processing model corresponding to the target task. The second training sample is input into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample. Based on the label quality corresponding to the second predicted labels, positive predicted labels and negative predicted labels are determined from the multiple second predicted labels. Based on the positive predicted labels and the negative predicted labels, the model parameters of the updated task processing model are adjusted to obtain the target task processing model corresponding to the target task.
[0041] Subsequently, the terminal 102 may send the task data to be processed corresponding to the target task to the server 104. The server 104 inputs the task data to be processed into the target task processing model corresponding to the target task to obtain the task processing result corresponding to the task data to be processed. The server 104 may return the task processing result to the terminal 102.
[0042] Among them, the terminal 102 may be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device may be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0043] It can be understood that both the terminal and the server can be used alone to execute the model training method provided in the embodiments of the present application. The terminal and the server can also be used in cooperation to execute the model training method provided in the embodiments of the present application.
[0044] In one embodiment, as Figure 2As shown in the figure, a model training method is provided. Taking the application of this method to a computer device as an example, the computer device can be a terminal or a server. It can be understood that this method can be executed independently by the terminal or the server itself, or can be implemented through the interaction between the terminal and the server. Among them:
[0045] Step S202: Obtain the first training sample, the second training sample corresponding to the target task, and the training label corresponding to the first training sample for the target task.
[0046] Among them, the target task is a specific and targeted task. The target task can be set according to actual needs. For example, the target task can be a question-and-answer task; the target task can be a dialogue generation task; the target task can be an instruction-following task; the target task can be a translation task; the target task can be a content review task.
[0047] The training data corresponding to the target task is the data prepared for training the model for processing the target task. The training data corresponding to the target task includes the first training sample, the second training sample corresponding to the target task, and the training label corresponding to the first training sample. The training sample corresponding to the target task is the input data of the model, and the training label corresponding to the training sample is the expected output data of the model. The first training sample and the second training sample corresponding to the target task can be the same training sample or different training samples. The first training sample is a labeled training sample, and the training label corresponding to the first training sample is the labeling result of the first training sample. For example, if the target task is a question-and-answer task, the first training sample is the question to be answered, and the training label corresponding to the first training sample is the correct answer to the question; if the target task is a translation task, the first training sample is the text to be translated, and the training label corresponding to the first training sample is the correct translation result of the text. The second training sample can be a labeled training sample or an unlabeled training sample.
[0048] Specifically, the computer device can obtain the training data corresponding to the target task locally or from other devices, and perform model training on the initial task processing model corresponding to the target task based on the training data to obtain the target task processing model corresponding to the target task.
[0049] Step S204: Input the first training sample after noise processing into the initial task processing model corresponding to the target task to obtain the first prediction label corresponding to the first training sample.
[0050] Among them, noise processing refers to adding noise to data. The initial task processing model corresponding to the target task is the task processing model to be trained corresponding to the target task. It can be understood that the initial task processing model corresponding to the target task can be a pre-trained model. For example, the initial task processing model corresponding to the target task can be a pre-trained language model. Of course, the initial task processing model corresponding to the target task can be a model obtained by training the pre-trained model. For example, obtain the third training sample corresponding to the target task, input the third training sample into the pre-trained model, obtain the predicted label corresponding to the third training sample, and based on the predicted label and the training label corresponding to the third training sample, adjust the pre-trained model to obtain the initial task processing model corresponding to the target task.
[0051] It can be understood that the predicted label corresponding to the training sample is the prediction result output by the model for the training sample. The training label corresponding to the training sample is the result expected to be output by the model for the training sample.
[0052] Step S206: Based on the first predicted label and the training label corresponding to the first training sample, adjust the model parameters of the initial task processing model to obtain an updated task processing model.
[0053] Specifically, based on the first training sample and the training label corresponding to the first training sample, perform supervised training on the initial task processing model corresponding to the target task to obtain an updated task processing model. The computer device can perform noise processing on the first training sample, input the first training sample after noise processing into the initial task processing model corresponding to the target task, and the initial task processing model corresponding to the target task performs data processing on the input data and outputs the first predicted label corresponding to the first training sample. The computer device can calculate the model loss based on the first predicted label and the training label corresponding to the first training sample, backpropagate the model loss to adjust the model parameters of the initial task processing model, and through multiple rounds of model iterative training until the convergence condition is met to obtain an updated task processing model.
[0054] Step S208: Input the second training sample into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample.
[0055] Specifically, the update task processing model is unsupervised trained based on the second training sample to obtain the target task processing model corresponding to the target task. The computer device can input the second training sample into the update task processing model, and the update task processing model processes the input data and outputs the second prediction label corresponding to the second training sample. For the second training sample, the update task processing model can output multiple second prediction labels corresponding to the second training sample. For example, the computer device can input the second training sample into the update task processing model multiple times to obtain multiple second prediction labels corresponding to the second training sample, and the update task processing model outputs at least one second prediction label each time. The update task processing model can control the randomness of the output result through hyperparameters, thereby affecting the quality and diversity of the output result. For example, the update task processing model can control the randomness of the output result through the temperature parameter, so that even if the same training sample is input into the update task processing model multiple times, the update task processing model can output multiple different prediction labels. Another example is that the computer device can input the second training sample into the update task processing model to obtain the second prediction label corresponding to the second training sample, and input the second training sample after noise processing into the update task processing model to obtain a new second prediction label corresponding to the second training sample.
[0056] Step S210: Determine the positive prediction label and the negative prediction label from the multiple second prediction labels based on the label quality corresponding to the second prediction label.
[0057] Among them, the label quality corresponding to the second prediction label refers to the degree of superiority or inferiority of the second prediction label in at least one dimension. For example, the label quality corresponding to the second prediction label refers to the degree of superiority or inferiority of the second prediction label in at least one dimension such as relevance, coherence, creativity, etc. It can be understood that the positive prediction label represents the second prediction label with higher label quality. The negative prediction label represents the second prediction label with lower label quality.
[0058] Specifically, after obtaining multiple second prediction labels corresponding to the second training sample, the computer device can determine the positive prediction label and the negative prediction label from the multiple second prediction labels based on the label quality corresponding to the second prediction label, take the second prediction label with higher label quality as the positive prediction label, and take the second prediction label with lower label quality as the negative prediction label.
[0059] Step S212: Adjust the model parameters of the update task processing model based on the positive prediction label and the negative prediction label to obtain the target task processing model corresponding to the target task.
[0060] Among them, the target task processing model corresponding to the target task refers to the task processing model that has completed training corresponding to the target task.
[0061] Specifically, after determining the positive prediction label and the negative prediction label, the computer device can calculate the model loss based on the positive prediction label and the negative prediction label, backpropagate the model loss to adjust and update the model parameters of the task processing model, and through multiple rounds of model iterative training until the convergence condition is met, to obtain the target task processing model corresponding to the target task.
[0062] In the above model training method, first, based on the first training samples corresponding to the target task and the training labels corresponding to the first training samples, supervised training is performed on the initial task processing model corresponding to the target task to obtain an updated task processing model. During the supervised training process, the first training samples are input into the model after noise processing. The addition of noise helps to reduce the overfitting of the model to the training data, prevent the model from overlearning specific patterns and details in the training data, and thus can improve the generalization ability of the model when facing the target task. Further, unsupervised training is performed on the updated task processing model based on the second training samples to obtain the target task processing model corresponding to the target task. In the unsupervised training, the second training samples are input into the updated task processing model to obtain multiple second prediction labels. The second prediction labels with higher label quality are selected as positive prediction labels, and the second prediction labels with lower label quality are selected as negative prediction labels. Fine-tuning the updated task processing model based on the positive prediction labels and the negative prediction labels helps the model learn how to make better decisions and improve the quality of the model output results. Combining the addition of noise and the selection of prediction labels can comprehensively improve the performance of the model in the target task and improve the model accuracy.
[0063] In one embodiment, the model training method further includes:
[0064] Perform feature transformation on the first training samples to obtain training embedding features;
[0065] Obtain the target scaling factor corresponding to the target task;
[0066] Based on the target scaling factor, perform scaling processing on the initial noise features to obtain target noise features;
[0067] Fuse the target noise features and the training embedding features to obtain the first training samples after noise processing.
[0068] Among them, feature transformation refers to converting the source data into a feature vector representation. For example, the first training samples are converted into training embedding features through an embedding layer (embedding layer), and the embedding layer is used to convert discrete data into a continuous vector representation. Another example is to convert the first training samples into training embedding features through one-hot encoding.
[0069] The target scaling factor corresponding to the target task is the scaling factor that needs to be used during the training of the initial task processing model corresponding to the target task, and is used to perform noise processing on the training samples corresponding to the target task. It can be understood that different tasks can correspond to the same scaling factor or different scaling factors. Moreover, the scaling factor can be a fixed value. For example, a preset value is used as the scaling factor. The scaling factor can also be a dynamic value. For example, the scaling factor is related to the number of training rounds of the model, and different scaling factors are used for different training rounds, and the scaling factor decreases as the number of training rounds increases.
[0070] The initial noise feature is a feature vector representing noise. The initial noise feature can be a randomly generated feature vector. The initial noise feature can be a feature vector randomly sampled from a probability distribution. For example, the initial noise feature is randomly sampled from a uniform distribution; the initial noise feature is randomly sampled from a Gaussian distribution. The target noise feature is a feature vector obtained by performing scaling processing on the initial noise feature. The scaling processing is used to adjust the scale of the feature vector.
[0071] Specifically, when performing noise processing on the first training sample, the computer device can first perform feature transformation on the first training sample to convert the first training sample into a training embedding feature. Furthermore, the computer device can obtain the target scaling factor corresponding to the target task, and perform scaling processing on the initial noise feature based on the target scaling factor to obtain the target noise feature. For example, dividing the target scaling factor by the initial noise feature to obtain the target noise feature. Another example is multiplying the target scaling factor and the initial noise feature to obtain the target noise feature. Finally, the computer device can fuse the target noise feature and the training embedding feature to obtain the first training sample after noise processing. For example, adding the target noise feature and the training embedding feature to obtain the first training sample after noise processing.
[0072] In one embodiment, the target scaling factor is adjusted based on the feature sequence length and feature dimension of the training embedding feature, and the initial noise feature is scaled based on the adjusted target scaling factor to obtain the target noise feature.
[0073] In a specific application, the calculation formula for performing noise processing on the training sample is as follows:
[0074]
[0075]
[0076]
[0077]
[0078] Among them, for the training sample Perform feature transformation to obtain training embedded features . The training samples can be text sequences. Let L denote the sequence length. For the words in the text sequence perform feature transformation to obtain embedding vectors . Add a random noise vector to each embedding vector. The random noise vector is randomly drawn from a uniform distribution Uniform(−1,1) and scaled according to the dimension d of the embedding vector and the sequence length L of the text sequence to ensure that the size of the noise vector is appropriate. denotes the random noise vector drawn from the uniform distribution for the embedding vector . is the noise vector obtained by scaling . is a hyperparameter used to control the intensity of the noise. Add the scaled noise vector to the original embedding vector to obtain the noisy embedding vector. denotes the training sample after noise processing.
[0079] Take the target task as a text processing task and the task processing model as a language model as an example to illustrate the supervised training process. Refer to Figure 3 , obtain the training text and the training labels corresponding to the training text. Input the training text into the embedding layer. Through the embedding layer, convert the words in the training text into embedding vectors. Generate noise vectors for each embedding vector through the noise layer. Add the noise vectors to the embedding vectors to obtain the noisy embedding vectors. Combine the noisy embedding vectors and input them into the pre-trained language model (i.e., the initial task processing model). The language model outputs the predicted labels corresponding to the training text. Calculate the model loss through the cross-entropy loss function based on the training labels and the predicted labels corresponding to the training text to obtain the cross-entropy loss. Backpropagate the cross-entropy loss to adjust the model parameters of the language model. Through multiple rounds of model iterative training until the convergence condition is met, obtain the updated language model (i.e., the updated task processing model).
[0080] It can be understood that during the fine-tuning process of the language model, adding noise to the embedding vectors can prevent the model from overfitting to specific patterns in the training data and enhance the generalization ability of the model.
[0081] In the above embodiment, perform feature transformation on the first training sample to convert the first training sample into training embedded features that are easier to calculate. Obtain the target scaling factor corresponding to the target task, and perform scaling processing on the initial noise feature based on the target scaling factor to obtain the personalized target noise feature. Fuse the target noise feature and the training embedded features, and embed the target noise feature into the training embedded features to obtain the first training sample after noise processing.
[0082] In one embodiment, obtaining a target scaling factor corresponding to a target task includes:
[0083] Based on the noise content of the original task data corresponding to the target task, a target scaling factor corresponding to the target task is determined; the target scaling factor and the noise content are positively correlated.
[0084] Among them, the original task data corresponding to the target task refers to the task data actually used in the scenario of the target task. For example, if the target task is a news task, the original task data corresponding to the target task refers to the original news, which the user can consult in daily life. For another example, if the target task is a question-and-answer task, the original task data corresponding to the target task is the original, actual questions and answers. For another example, if the target task is a content review task, the original task data corresponding to the target task is the original, actual content to be reviewed and the review results.
[0085] The noise content of the original task data refers to the proportion or intensity of the noise information in the original task data. The noise information in the original task data refers to the irrelevant, erroneous or interfering information in the original task data. For example, the noise information can be typos, incorrect sentences, garbled characters, advertisements and other information.
[0086] For example, the noise content of the original task data corresponding to highly professional tasks such as news tasks and academic tasks is less than the noise content of the original task data corresponding to highly entertaining and life-related tasks such as social tasks and novel tasks. It can be understood that the words used in highly professional tasks usually pursue accuracy, objectivity and standardization, so the noise content is relatively small. The words used in tasks with strong entertainment and life-related characteristics are usually free, casual and jumpy, so the noise content is relatively large.
[0087] Specifically, the computer device can obtain the noise content of the original task data corresponding to the target task, and determine the target scaling factor corresponding to the target task based on the noise content of the original task data corresponding to the target task. The target scaling factor is positively correlated with the noise content, that is, the greater the noise content of the original task data corresponding to the target task, the greater the target scaling factor corresponding to the target task. For example, scaling factors corresponding to various noise content intervals are pre-set, and the scaling factor corresponding to the noise content interval to which the noise content corresponding to the target task belongs is used as the target scaling factor corresponding to the target task. For another example, a mapping formula between noise content and scaling factor is designed, and the noise content corresponding to the target task is substituted into the mapping formula to calculate the target scaling factor corresponding to the target task.
[0088] It is understood that the noise content of the original task data corresponding to the target task can be preset, for example, the noise content of the original task data corresponding to each task can be set manually according to the characteristics of each task. The noise content of the original task data corresponding to the target task can also be calculated in advance, for example, a large amount of original task data corresponding to the target task is collected, the noise content of each original task data is calculated, and the average value of all noise contents is used as the final noise content corresponding to the target task.
[0089] In one embodiment, the target task is a text task. The original task data corresponding to the text task refers to the original text data. The noise content of the original task data corresponding to the text task refers to the noise content of the original text data. The noise content of the original text data can be the proportion or intensity of irrelevant, erroneous or interfering information present in the original text data. For example, the noise content of the original text data is determined based on grammatical noise such as typos and incorrect sentences in the original text data, as well as contextual noise such as advertisements and copyright notices.
[0090] In one embodiment, the first training sample and the second training sample corresponding to the target task may be preprocessed training samples. Preprocessing refers to processing noise, missing values, and abnormal values in the data to ensure data consistency.
[0091] In one embodiment, based on the noise content of the original task data corresponding to the target task, the target scaling factor corresponding to the target task in the initial training round is determined. The target scaling factor corresponding to the target task in the remaining training rounds decreases as the current training round increases. Wherein, the initial training round refers to at least one initial training round. For example, a total of 100 rounds of training are required, and the initial training rounds can be the first 10 rounds. The remaining training rounds refer to the remaining training rounds after the initial training round. For example, a total of 100 rounds of training are required, if the initial training rounds are the first 10 rounds, then the remaining training rounds are the remaining 90 rounds. In this way, based on the noise content of the original task data corresponding to the target task, the target scaling factor corresponding to the target task in the initial training round is determined, which can ensure that the noise matching the target task is added in the initial training round. In addition, the target scaling factor corresponding to the target task in the remaining training rounds decreases as the current training round increases, thereby gradually reducing the addition of noise when the model learns more and more knowledge, so that the model can learn more accurate knowledge.
[0092] In the above embodiments, based on the noise content of the original task data corresponding to the target task, the target scaling factor corresponding to the target task is determined, so that different tasks can obtain different scaling factors due to different noise contents, and subsequent personalized noise processing can be realized. The target scaling factor corresponding to the target task is positively correlated with the noise content of the original task data corresponding to the target task. Therefore, the training samples after noise processing can imitate the characteristics of the actual task data. Training the model based on the training samples after noise processing helps to improve the data processing accuracy of the model when facing the actual task data.
[0093] In one embodiment, obtaining the target scaling factor corresponding to the target task includes:
[0094] Determining the target scaling factor corresponding to the target task based on the current training round in which the initial task processing model is located;
[0095] Wherein, the target scaling factor decreases as the current training round increases, and the decreasing speed first decreases, then increases, and then decreases as the current training round increases.
[0096] Specifically, based on the training samples corresponding to the target task and the training labels corresponding to the training samples, supervised training is performed on the initial task processing model corresponding to the target task. During the supervised training process, after multiple rounds of training until the convergence condition is met, an updated task processing model is obtained. It can be understood that multiple training samples are used in each round of training, and different training samples can be used between different rounds, or the same training samples can be used.
[0097] The current training round refers to the current training round in which the model is located. For example, if a total of 100 rounds of training are required, in the 50th round of training, the current training round is the 50th round. The target scaling factor corresponding to the target task is related to the current training round in which the model is located. The target scaling factor is a dynamic value and changes with the change of the current training round. The computer device can determine the current training round in which the initial task processing model is located, and based on the current training round in which the initial task processing model is located, determine the target scaling factor corresponding to the target task. The target scaling factor decreases as the current training round increases. For example, if a total of 100 rounds of training are required, the target scaling factor in the 2nd round is less than that in the 1st round, and the target scaling factor in the 3rd round is less than that in the 2nd round, and so on. The target scaling factor in the 100th round is less than that in the 99th round. And, the decreasing speed of the target scaling factor first decreases, then increases, and then decreases as the current training round increases.
[0098] In one embodiment, the target scaling factor decreases as the current training round increases, and the decreasing speed first decreases, then increases, and then decreases again as the current training round increases, until the target scaling factor decreases to the target value. It can be understood that the target value can be set as needed.
[0099] In one embodiment, based on the noise content of the original task data corresponding to the target task, the target scaling factor corresponding to the target task in the initial training round is determined. The target scaling factor corresponding to the target task in the remaining training rounds decreases as the current training round increases, and the decreasing speed first decreases, then increases, and then decreases again as the current training round increases.
[0100] In the above embodiment, the target scaling factor decreases as the current training round increases, that is, the noise added to the training samples gradually decreases, enabling the model to gradually learn knowledge from more accurate data. Moreover, the decreasing speed of the target scaling factor first decreases, then increases, and then decreases again as the current training round increases. Performing noise processing through the periodic target scaling factor helps to prevent the model from falling into a local optimal solution.
[0101] In one embodiment, the model training method further includes:
[0102] Based on the second training sample, perform a correlation analysis on the second predicted label to obtain the correlation degree corresponding to the second predicted label;
[0103] Perform a coherence analysis on the second predicted label to obtain the coherence degree corresponding to the second predicted label;
[0104] Perform a creativity analysis on the second predicted label to obtain the creativity degree corresponding to the second predicted label;
[0105] Fuse the correlation degree, coherence degree, and creativity degree corresponding to the second predicted label to obtain the label quality corresponding to the second predicted label.
[0106] Among them, the correlation analysis refers to analyzing the degree of correlation between the input data and output data of the model. Performing a correlation analysis on the predicted label to obtain the correlation degree corresponding to the predicted label. The coherence analysis refers to analyzing the coherence degree of the output data of the model in terms of content. Performing a coherence analysis on the predicted label to obtain the coherence degree corresponding to the predicted label. The creativity analysis refers to analyzing the innovation degree of the output data of the model in terms of content. Performing a creativity analysis on the predicted label to obtain the creativity degree corresponding to the predicted label.
[0107] Specifically, for any second prediction label, the computer device can perform a correlation analysis on the second prediction label based on the second training samples to obtain the correlation degree corresponding to the second prediction label. For example, feature extraction is performed on the second training samples to obtain the semantic features corresponding to the second training samples, feature extraction is performed on the second prediction label to obtain the semantic features corresponding to the second prediction label, and the similarity between the semantic features corresponding to the second training samples and the semantic features corresponding to the second prediction label is calculated as the correlation degree corresponding to the second prediction label. Another example is to input the second training samples and the second prediction label into a correlation analysis model, and the correlation analysis model outputs the correlation degree corresponding to the second prediction label.
[0108] For any second prediction label, the computer device can perform a coherence analysis on the second prediction label to obtain the coherence degree corresponding to the second prediction label. For example, input the second prediction label into a coherence analysis model, and the coherence analysis model outputs the coherence degree corresponding to the second prediction label. Another example is to divide the second prediction label into multiple segments, extract the themes of each segment. If the themes of each segment are consistent, then determine that the coherence degree corresponding to the second prediction label is the first coherence degree. If the themes of each segment are inconsistent, then determine that the coherence degree corresponding to the second prediction label is the second coherence degree, and the first coherence degree is greater than the second coherence degree. Another example is to divide the second prediction label into multiple segments, calculate the semantic similarity between adjacent segments, and determine the coherence degree corresponding to the second prediction label based on the semantic similarity between adjacent segments. The coherence degree is positively correlated with the semantic similarity.
[0109] For any second prediction label, the computer device can perform a creativity analysis on the second prediction label to obtain the creativity degree corresponding to the second prediction label. For example, input the second prediction label into a creativity analysis model, and the creativity analysis model outputs the creativity degree corresponding to the second prediction label. Another example is to calculate the content repetition rate corresponding to the second prediction label. The content repetition rate corresponding to the second prediction label refers to the proportion of repeated or similar content in the second prediction label in the second prediction label, and determine the creativity degree corresponding to the second prediction label based on the content repetition rate. The creativity degree is negatively correlated with the content repetition rate.
[0110] For any second prediction label, the computer device can fuse the correlation degree, coherence degree, and creativity degree corresponding to the second prediction label to obtain the label quality corresponding to the second prediction label. For example, take the average value of the correlation degree, coherence degree, and creativity degree corresponding to the second prediction label to obtain the label quality corresponding to the second prediction label. Another example is to perform a weighted sum of the correlation degree, coherence degree, and creativity degree corresponding to the second prediction label to obtain the label quality corresponding to the second prediction label. The weights corresponding to the correlation degree, coherence degree, and creativity degree can be set as needed.
[0111] In the above embodiments, a correlation analysis is performed on the second predicted label to obtain the correlation degree corresponding to the second predicted label, a coherence analysis is performed on the second predicted label to obtain the coherence degree corresponding to the second predicted label, a creativity analysis is performed on the second predicted label to obtain the creativity degree corresponding to the second predicted label, and the correlation degree, coherence degree, and creativity degree corresponding to the second predicted label are fused to obtain the label quality corresponding to the second predicted label. By comprehensively considering the performance of the second predicted label in dimensions such as correlation, coherence, and creativity, the label quality corresponding to the second predicted label can be obtained, effectively ensuring the accuracy of the label quality.
[0112] In one embodiment, performing a coherence analysis on the second predicted label to obtain the coherence degree corresponding to the second predicted label includes:
[0113] Inputting the second predicted label into a coherence analysis model to obtain the coherence degree corresponding to the second predicted label;
[0114] Among them, the coherence analysis model is a model obtained through supervised training based on coherence data and incoherence data.
[0115] Among them, the coherence analysis model is an artificial intelligence model used to perform a coherence analysis on input data. Coherence data and incoherence data are the training data of the coherence analysis model. Coherence data and incoherence data are data that have been labeled. Coherence data represents data that meets the coherence requirements, and incoherence data represents data that does not meet the coherence requirements. For example, text with coherent sentences is obtained as coherence data, and text with incoherent sentences is obtained as incoherence data. Coherence data and incoherence data can be collected and labeled manually.
[0116] Specifically, a coherence analysis is performed with the help of an artificial intelligence model. The computer device can obtain a trained coherence analysis model, input the second predicted label into the coherence analysis model, and the coherence analysis model outputs the coherence degree corresponding to the second predicted label.
[0117] The coherence analysis model is a model obtained through supervised training based on coherence data and incoherence data. Inputting the coherence data into the coherence analysis model to be trained to obtain the predicted label corresponding to the coherence data, inputting the incoherence data into the coherence analysis model to be trained to obtain the predicted label corresponding to the incoherence data, and based on the predicted label corresponding to the coherence data and the training label, and the predicted label corresponding to the incoherence data and the training label, adjusting the model parameters of the coherence analysis model. Through multiple rounds of model iterative training until the convergence condition is met, the trained coherence analysis model is obtained. It can be understood that the training label corresponding to the coherence data is a positive label, and the training label corresponding to the incoherence data is a negative label.
[0118] In the above embodiments, the coherence analysis model is a model obtained by supervised training based on coherence data and incoherence data, which can effectively distinguish coherence data from incoherence data, thereby achieving accurate coherence analysis. By inputting the second prediction label into the coherence analysis model, the coherence analysis model outputs the coherence degree corresponding to the second prediction label, and an accurate coherence degree can be obtained through the coherence analysis model.
[0119] In one embodiment, determining positive prediction labels and negative prediction labels from multiple second prediction labels based on the label quality corresponding to the second prediction labels includes:
[0120] From multiple second prediction labels, obtaining the second prediction labels with label quality greater than the first quality threshold as positive prediction labels, and obtaining the second prediction labels with label quality less than the second quality threshold as negative prediction labels; wherein, the first quality threshold is greater than or equal to the second quality threshold.
[0121] Wherein, the first quality threshold and the second quality threshold are quality thresholds set for the label quality. The first quality threshold is used to screen out positive prediction labels, the second quality threshold is used to screen out negative prediction labels, and the first quality threshold is greater than or equal to the second quality threshold.
[0122] Specifically, positive prediction labels and negative prediction labels are determined from multiple second prediction labels based on the quality thresholds. The computer device can compare the label quality corresponding to the second prediction label with the first quality threshold, and from multiple second prediction labels, obtain the second prediction labels with label quality greater than the first quality threshold as positive prediction labels. Then compare the label quality corresponding to the second prediction label with the second quality threshold, and from multiple second prediction labels, obtain the second prediction labels with label quality less than the second quality threshold as negative prediction labels.
[0123] In the above embodiments, positive prediction labels and negative prediction labels can be quickly determined from multiple second prediction labels based on the first quality threshold and the second quality threshold.
[0124] In one embodiment, the updated task processing model also outputs the prediction probabilities corresponding to multiple second prediction labels respectively.
[0125] Adjusting the model parameters of the updated task processing model based on the positive prediction labels and the negative prediction labels to obtain the target task processing model corresponding to the target task includes:
[0126] Based on the prediction probability corresponding to the positive prediction label and the prediction probability corresponding to the negative prediction label, obtaining a target loss; the target loss is negatively correlated with the prediction probability corresponding to the positive prediction label, and the target loss is positively correlated with the prediction probability corresponding to the negative prediction label;
[0127] Based on the target loss, adjust the model parameters of the update task processing model until the convergence condition is met, and obtain the target task processing model corresponding to the target task.
[0128] Among them, input the second training sample into the update task processing model, and the update task processing model outputs the second predicted label corresponding to the second training sample and the prediction probability corresponding to the second predicted label. The prediction probability corresponding to the second predicted label refers to the probability that the model outputs the second predicted label for the second training sample. It can be understood that when inputting the second training sample into the update task processing model, the update task processing model can generate multiple candidate predicted labels, and each candidate predicted label has a corresponding prediction probability. Finally, the update task processing model can select at least one candidate predicted label from multiple candidate predicted labels as the second predicted label for output.
[0129] The convergence condition is the condition for judging whether the model has converged. The convergence condition includes but is not limited to at least one of the model loss being greater than a preset loss value, the number of model training rounds being greater than a preset number of rounds, the change rate of the model loss being less than a preset change rate, etc.
[0130] Specifically, the computer device can calculate the model loss based on the prediction probability corresponding to the positive predicted label and the prediction probability corresponding to the negative predicted label to obtain the target loss. Among them, the target loss is negatively correlated with the prediction probability corresponding to the positive predicted label, and the target loss is positively correlated with the prediction probability corresponding to the negative predicted label. The training goal of the model is to make the target loss smaller and smaller. That is, increase the prediction probability corresponding to the positive predicted label and decrease the prediction probability corresponding to the negative predicted label, so that the model can output high-quality results when facing new inputs. The computer device can backpropagate the target loss to adjust the model parameters of the update task processing model, and through multiple rounds of model iterative training, until the convergence condition is met, to obtain the target task processing model corresponding to the target task.
[0131] In a specific application, the calculation formula of the target loss is as follows:
[0132]
[0133] Among them, represents the target loss under the model parameters . chosen represents the positive predicted label, and reject represents the negative predicted label. represents the probability of generating the output when the input X is given under the model parameters . represents the probability of generating the output when the input X is given under the model parameters The probability. For "chosen", the goal of the model is to increase its generation probability; for "reject", the goal of the model is to decrease its generation probability.
[0134] In the above embodiments, based on the prediction probability corresponding to the positive prediction label and the prediction probability corresponding to the negative prediction label, the target loss is obtained. The target loss is negatively correlated with the prediction probability corresponding to the positive prediction label, and the target loss is positively correlated with the prediction probability corresponding to the negative prediction label. Based on such a target loss, adjusting and updating the model parameters of the task processing model can enable the model to gradually learn how to output high-quality results and improve the model accuracy.
[0135] In one embodiment, inputting the second training sample into the updated task processing model to obtain multiple second prediction labels corresponding to the second training sample, including:
[0136] Inputting the second training sample into the updated task processing model after noise processing to obtain multiple second prediction labels corresponding to the second training sample.
[0137] Specifically, noise processing can also be added during the unsupervised training process. The computer device can input the second training sample into the updated task processing model after noise processing. The updated task processing model processes the input data and outputs the second prediction labels corresponding to the second training sample. For the second training sample, the updated task processing model can output multiple second prediction labels corresponding to the second training sample.
[0138] In a specific application, referring to Figure 4 , the model training is divided into two stages. In the first stage, obtaining the first training sample corresponding to the target task and the training label corresponding to the first training sample, inputting the first training sample into the initial task processing model corresponding to the target task after noise processing to obtain the first prediction label corresponding to the first training sample, calculating the model loss based on the first prediction label and the training label corresponding to the first training sample, backpropagating the model loss to adjust the model parameters of the initial task processing model, and through multiple rounds of model iterative training until the first convergence condition is met to obtain the updated task processing model. In the second stage, obtaining the second training sample corresponding to the target task, inputting the second training sample into the updated task processing model to obtain multiple second prediction labels corresponding to the second training sample, determining the positive prediction label and the negative prediction label from the multiple second prediction labels based on the label quality corresponding to the second prediction label, calculating the model loss based on the positive prediction label and the negative prediction label, backpropagating the model loss to adjust the model parameters of the updated task processing model, and through multiple rounds of model iterative training until the second convergence condition is met to obtain the target task processing model corresponding to the target task. It can be understood that the first convergence condition and the second convergence condition can be the same convergence condition or different convergence conditions.
[0139] In a specific application, referring to Figure 5 , obtain a second training text, input the second training text into an embedding layer, convert the words in the second training text into embedding vectors through the embedding layer, generate noise vectors for each embedding vector through a noise layer, add the noise vectors to the embedding vectors to obtain noisy embedding vectors. Combine the respective noisy embedding vectors and input them into a supervised-trained language model (i.e., the updated task processing model). The language model outputs multiple second prediction labels corresponding to the second training text. Calculate the label quality corresponding to the second prediction labels, determine positive prediction labels and negative prediction labels from the second prediction labels based on the label quality. The positive prediction labels represent the second prediction labels that are not affected by noise and have higher quality, and the negative prediction labels represent the second prediction labels that are affected by noise and have lower quality. Calculate the model loss based on the relevant data of the positive prediction labels and negative prediction labels, backpropagate the model loss to adjust the model parameters of the language model, and through multiple rounds of model iterative training until the convergence condition is met, obtain the trained language model (i.e., the target task processing model).
[0140] In the above embodiment, during the unsupervised training process, the second training sample is input into the updated task processing model after noise processing. The addition of noise helps to improve the diversity of the model output results and also helps to reduce the overfitting of the model to the training data.
[0141] In one embodiment, inputting the second training sample into the updated task processing model to obtain multiple second prediction labels corresponding to the second training sample includes:
[0142] Input the second training sample into the updated task processing model;
[0143] Through the updated task processing model, perform data processing on the second training sample to obtain initial output features, perform noise processing on the initial output features to obtain target output features, and based on the target output features, obtain the second prediction labels corresponding to the second training sample and the prediction probabilities corresponding to the second prediction labels.
[0144] Specifically, when applying noise processing during the unsupervised training process, noise processing can be performed at the input stage of the model or at the output stage of the model. If noise processing is performed at the input stage of the model, then the second training sample is input into the updated task processing model after noise processing to obtain multiple second prediction labels corresponding to the second training sample and the prediction probabilities corresponding to the second prediction labels. If noise processing is performed at the output stage of the model, then noise processing is performed at the output layer of the model to obtain multiple second prediction labels corresponding to the second training sample and the prediction probabilities corresponding to the second prediction labels.
[0145] The computer device inputs the second training sample into the update task processing model. Through the update task processing model, data processing is performed on the second training sample to obtain initial output features. The initial output features are the input features of the output layer in the update task processing model. The initial output features are input into the output layer after noise processing, and the output layer outputs the second prediction label and the prediction probability corresponding to the second prediction label. In the output layer, noise processing is performed on the initial output features to obtain target output features. Based on the target output features, the second prediction label corresponding to the second training sample and the prediction probability corresponding to the second prediction label are obtained.
[0146] In the above embodiment, performing noise processing in the output stage of the model helps to improve the diversity of the model output results and also helps to reduce the overfitting of the model to the training data.
[0147] In one embodiment, the second training sample is input into the update task processing model to obtain multiple second prediction labels corresponding to the second training sample. Based on the label quality corresponding to the second prediction labels, positive prediction labels and negative prediction labels are determined from the multiple second prediction labels, including:
[0148] The second training sample is input into the update task processing model to obtain the positive prediction label corresponding to the second training sample;
[0149] The second training sample is input into the update task processing model after noise processing to obtain the negative prediction label corresponding to the second training sample.
[0150] Specifically, the computer device can input the second training sample into the update task processing model to obtain the positive prediction label corresponding to the second training sample, and input the second training sample into the update task processing model after noise processing to obtain the negative prediction label corresponding to the second training sample. That is, if the input data of the update task processing model is the second training sample without noise processing, it is considered that the second prediction label output by the update task processing model is a high-quality second prediction label, and the second prediction label output by the update task processing model is used as the positive prediction label. If the input data of the update task processing model is the second training sample after noise processing, it is considered that the second prediction label output by the update task processing model is a low-quality second prediction label, and the second prediction label output by the update task processing model is used as the negative prediction label.
[0151] It can be understood that the second training sample is input into the update task processing model to obtain at least one positive prediction label corresponding to the second training sample. The second training sample is input into the update task processing model after noise processing to obtain at least one negative prediction label corresponding to the second training sample.
[0152] In a specific application, refer to Figure 6, the model training is divided into two stages. In the first stage, the first training sample corresponding to the target task and the training label corresponding to the first training sample are obtained, the first training sample is input into the initial task processing model corresponding to the target task after noise processing, the first prediction label corresponding to the first training sample is obtained, the model loss is calculated based on the first prediction label and the training label corresponding to the first training sample, the model loss is back-propagated to adjust the model parameters of the initial task processing model, and multiple rounds of model iterative training are performed until the first convergence condition is met, and the updated task processing model is obtained. In the second stage, the second training sample corresponding to the target task is obtained, the second training sample is input into the updated task processing model, the positive prediction label corresponding to the second training sample is obtained, the second training sample is input into the updated task processing model after noise processing, the negative prediction label corresponding to the second training sample is obtained, the model loss is calculated based on the positive prediction label and the negative prediction label, the model loss is back-propagated to adjust the model parameters of the updated task processing model, and multiple rounds of model iterative training are performed until the second convergence condition is met, and the target task processing model corresponding to the target task is obtained.
[0153] In the above embodiment, the second training sample is input into the update task processing model to obtain the positive prediction label corresponding to the second training sample, and the second training sample is input into the update task processing model after noise processing to obtain the negative prediction label corresponding to the second training sample. Thus, the positive prediction label and the negative prediction label can be quickly distinguished through the input data, which helps to improve the model training efficiency.
[0154] In one embodiment, inputting the second training sample into the update task processing model to obtain a positive prediction label corresponding to the second training sample includes:
[0155] The second training sample is input into the update task processing model to obtain a plurality of second prediction labels corresponding to the second training sample, and a positive prediction label is determined from the plurality of second prediction labels based on label qualities corresponding to the second prediction labels.
[0156] Specifically, the second training sample that has not been processed by noise is input into the update task processing model to obtain multiple second prediction labels corresponding to the second training sample, the label qualities corresponding to the multiple second prediction labels are calculated, and the positive prediction label is determined from the multiple second prediction labels based on the label quality. For example, the second prediction label whose label quality is greater than the first quality threshold is used as the positive prediction label. For another example, multiple second prediction labels are sorted from large to small according to label quality, and the second prediction label with the highest sorting is used as the positive prediction label. In this way, the positive prediction label is further screened out by label quality from the output results that are not affected by noise, so as to select the best from the best, which helps to improve the accuracy of the positive prediction label and thus improve the quality of model training.
[0157] In one embodiment, the second training sample is input into the update task processing model after noise processing to obtain the negative prediction label corresponding to the second training sample, including:
[0158] The second training sample is input into the update task processing model after noise processing to obtain multiple second prediction labels corresponding to the second training sample, and the negative prediction label is determined from the multiple second prediction labels based on the label quality corresponding to the second prediction labels.
[0159] Specifically, the second training sample after noise processing is input into the update task processing model to obtain multiple second prediction labels corresponding to the second training sample, the label quality corresponding to each of the multiple second prediction labels is calculated, and the negative prediction label is determined from the multiple second prediction labels based on the label quality. For example, the second prediction label with a label quality less than the second quality threshold is used as the negative prediction label. For another example, the multiple second prediction labels are sorted in ascending order of label quality, and the second prediction label with a higher ranking is used as the negative prediction label. In this way, the negative prediction label is further screened out from the output results affected by noise through the label quality, which helps to improve the accuracy of the negative prediction label and further improve the model training quality.
[0160] In one embodiment, the target task is a question-and-answer task, the first training sample and the second training sample are training questions, the training label corresponding to the first training sample is the known answer corresponding to the first training sample, and the target task processing model is the target question-and-answer model.
[0161] The model training method further includes:
[0162] The question to be processed is input into the target question-and-answer model to obtain the prediction label corresponding to the question to be processed;
[0163] Based on the prediction label corresponding to the question to be processed, the target answer corresponding to the question to be processed is determined.
[0164] Among them, the question-and-answer task is a task of giving an answer according to a question and answer. The task processing model corresponding to the question-and-answer task is a question-and-answer model. The question is input into the question-and-answer model, and the question-and-answer model outputs an answer. The training sample corresponding to the question-and-answer task is a training question, and the known answer corresponding to the training question is the training label corresponding to the training sample.
[0165] Specifically, the method of the present application can be applied to a question-and-answer task to improve the model accuracy for the question-and-answer task. In the model training phase, the computer device obtains a first training sample and a second training sample corresponding to the question-and-answer task, and a training label corresponding to the first training sample; inputs the first training sample after noise processing into the initial question-and-answer model corresponding to the question-and-answer task to obtain a first predicted label corresponding to the first training sample; adjusts the model parameters of the initial question-and-answer model based on the first predicted label and the training label corresponding to the first training sample to obtain an updated question-and-answer model; inputs the second training sample into the updated question-and-answer model to obtain multiple second predicted labels corresponding to the second training sample; determines a positive predicted label and a negative predicted label from the multiple second predicted labels based on the label quality corresponding to the second predicted label; and adjusts the model parameters of the updated question-and-answer model based on the positive predicted label and the negative predicted label to obtain the target question-and-answer model corresponding to the question-and-answer task.
[0166] In the model application phase, the computer device obtains a question to be processed, inputs the question to be processed into the target question-and-answer model, the target question-and-answer model processes the input data, outputs a predicted label corresponding to the question to be processed, and the computer device determines a target answer corresponding to the question to be processed based on the predicted label corresponding to the question to be processed. For example, if the predicted label is a predicted answer, the predicted answer output by the target question-and-answer model is used as the target answer corresponding to the question to be processed.
[0167] In a specific application, referring to Figure 7 , the model training for the question-and-answer model is divided into two phases. In the first phase, a first training question corresponding to the question-and-answer task and a training answer corresponding to the first training question are obtained, the first training question after noise processing is input into the initial question-and-answer model corresponding to the question-and-answer task to obtain a first predicted answer corresponding to the first training question, the model loss is calculated based on the first predicted answer and the training answer corresponding to the first training question, and the model loss is backpropagated to adjust the model parameters of the initial question-and-answer model. Through multiple rounds of model iterative training until the first convergence condition is met, an updated question-and-answer model is obtained. In the second phase, a second training question corresponding to the question-and-answer task is obtained, the second training question is input into the updated question-and-answer model to obtain multiple second predicted answers corresponding to the second training question, the answer quality corresponding to the second predicted answers is calculated, a positive predicted answer and a negative predicted answer are determined from the multiple second predicted answers based on the answer quality, the model loss is calculated based on the relevant data of the positive predicted answer and the negative predicted answer, and the model loss is backpropagated to adjust the model parameters of the updated question-and-answer model. Through multiple rounds of model iterative training until the second convergence condition is met, the target question-and-answer model corresponding to the target task is obtained.
[0168] In the above embodiments, the method of the present application can be applied to a question-and-answer task. By training the question-and-answer model with the method of the present application, the model accuracy of the question-and-answer model can be effectively improved.
[0169] In one embodiment, the target task is a content review task, the first training sample and the second training sample are training contents, the training label corresponding to the first training sample is the known review result corresponding to the first training sample, and the target task processing model is the target content review model.
[0170] The model training method further includes:
[0171] Input the content to be reviewed into the target content review model to obtain the predicted label corresponding to the content to be reviewed;
[0172] Based on the predicted label corresponding to the content to be reviewed, determine the target review result corresponding to the content to be reviewed.
[0173] Among them, the content review task is a task of reviewing content. For example, the content review task can be a news review task, a video review task, etc. It can be understood that the content can be in the form of text, image, video, audio, etc. The task processing model corresponding to the content review task is the content review model. Input the content to be reviewed into the content review model, and the content review model outputs the review result. The training sample corresponding to the content review task is the training content, and the known review result corresponding to the training content is the training label corresponding to the training sample.
[0174] Specifically, the method of the present application can be applied to the content review task to improve the model accuracy for the content review task. In the model training stage, the computer device obtains the first training sample and the second training sample corresponding to the content review task, and the training label corresponding to the first training sample; input the first training sample after noise processing into the initial content review model corresponding to the content review task to obtain the first predicted label corresponding to the first training sample; based on the first predicted label and the training label corresponding to the first training sample, adjust the model parameters of the initial content review model to obtain the updated content review model; input the second training sample into the updated content review model to obtain multiple second predicted labels corresponding to the second training sample; based on the label quality corresponding to the second predicted label, determine the positive predicted label and the negative predicted label from the multiple second predicted labels; based on the positive predicted label and the negative predicted label, adjust the model parameters of the updated content review model to obtain the target content review model corresponding to the content review task.
[0175] In the model application stage, the computer device obtains the content to be reviewed, inputs the content to be reviewed into the target content review model, the target content review model performs data processing on the input data, outputs the predicted label corresponding to the content to be reviewed, and the computer device determines the target review result corresponding to the content to be reviewed based on the predicted label corresponding to the content to be reviewed. For example, the predicted label is the predicted review result, and the predicted review result output by the target content review model is used as the target review result corresponding to the content to be reviewed.
[0176] In a specific application, referring to Figure 8 , the model training for the content review model is divided into two stages. In the first stage, the first training content corresponding to the content review task and the training review results corresponding to the first training content are obtained. The first training content is input into the initial content review model corresponding to the content review task after noise processing to obtain the first predicted review results corresponding to the first training content. The model loss is calculated based on the first predicted review results and the training review results corresponding to the first training content, and the model loss is backpropagated to adjust the model parameters of the initial content review model. Through multiple rounds of model iterative training, until the first convergence condition is met, an updated content review model is obtained. In the second stage, the second training content corresponding to the content review task is obtained. The second training content is input into the updated content review model to obtain multiple second predicted review results corresponding to the second training content. The review quality corresponding to the second predicted review results is calculated, and the positive predicted review results and negative predicted review results are determined from the multiple second predicted review results based on the review quality. The model loss is calculated based on the relevant data of the positive predicted review results and negative predicted review results, and the model loss is backpropagated to adjust the model parameters of the updated content review model. Through multiple rounds of model iterative training, until the second convergence condition is met, the target content review model corresponding to the target task is obtained.
[0177] In the above embodiments, the method of the present application can be applied to content review tasks. By training the content review model with the method of the present application, the model accuracy of the content review model can be effectively improved.
[0178] In one embodiment, as Figure 9 shown, the model training method further includes:
[0179] Step S902, obtaining the task data to be processed corresponding to the target task.
[0180] Among them, the target task is a specific and targeted task. The target task can be set according to actual needs. For example, the target task can be a question-answering task; the target task can be a dialogue generation task; the target task can be an instruction-following task; the target task can be a translation task; the target task can be a content review task.
[0181] The task data to be processed corresponding to the target task refers to the task data that needs to be processed corresponding to the target task, specifically the data that needs to be processed by the task processing model corresponding to the target task.
[0182] Step S904, inputting the task data to be processed into the target task processing model corresponding to the target task to obtain the task processing results corresponding to the task data to be processed.
[0183] Among them, the target task processing model corresponding to the target task refers to the task processing model that has completed training for the target task.
[0184] It can be understood that the target task processing model is the model obtained according to the aforementioned model training method. Regarding the model training process, it will not be elaborated here.
[0185] Specifically, the computer device can obtain the task data to be processed corresponding to the target task locally or from other devices, obtain the target task processing model corresponding to the target task, input the task data to be processed corresponding to the target task into the target task processing model, and the target task processing model processes the input data and outputs the task processing result corresponding to the task data to be processed.
[0186] For example, if the target task is a question-and-answer task and the task data to be processed is the question to be processed, input the question to be processed into the target question-and-answer model corresponding to the question-and-answer task to obtain the target answer corresponding to the question to be processed.
[0187] For example, if the target task is a content review task and the task data to be processed is the content to be reviewed, input the content to be reviewed into the target content review model corresponding to the content review task to obtain the target content review result corresponding to the content to be reviewed.
[0188] For example, if the target task is a translation task and the task data to be processed is the text to be translated, input the text to be translated into the target translation model corresponding to the translation task to obtain the target translation result corresponding to the text to be translated.
[0189] In the above embodiments, the target task processing model trained by the aforementioned model training method has high model accuracy. Input the task data to be processed corresponding to the target task into the target task processing model to obtain the task processing result corresponding to the task data to be processed. Correspondingly, the task processing result has high accuracy.
[0190] In one embodiment, the target task is a content review task.
[0191] Obtaining the task data to be processed corresponding to the target task includes:
[0192] Obtaining the content to be reviewed uploaded in the content application;
[0193] The model training method further includes:
[0194] When it is determined based on the task processing result that the content to be reviewed passes the content review, publish the content to be reviewed in the content application.
[0195] Among them, the content review task is a task for reviewing content. For example, the content review task can be a news review task, a video review task, etc. It can be understood that the content can be presented in the form of text, image, video, audio, etc.
[0196] The content application is an application program for publishing and browsing content. For example, the content application can be a social application, and users can browse content such as friends' dynamics, live broadcasts, and short videos in the social application, and publish their own dynamics, live broadcasts, short videos, etc. For another example, the content application can be a news application, and users can browse news in the news application and publish news. It can be understood that the application program in the method of this application can be a client installed on the terminal, or a non-installed application program, such as a small program, or a web application opened through a browser.
[0197] Specifically, the method of this application can be applied to the content review task in the content application, and the target content review model trained by the method of this application is used to review the content to be published in the content application. If the content review is passed, the publication is allowed; if the content review is not passed, the publication is not allowed.
[0198] The computer device obtains the content to be reviewed uploaded by the current user in the content application, inputs the content to be reviewed into the target content review model corresponding to the content review task, obtains the content review result corresponding to the content to be reviewed, and determines whether the content to be reviewed passes the content review based on the content review result. If the content to be reviewed passes the content review, the computer device can publish the content to be reviewed in the content application, so that other users can browse the content published by the current user in the content application. If the content to be reviewed does not pass the content review, the computer device can return a message indicating that the content review is not passed to the current user to prompt that the content uploaded by the current user cannot be published because it does not pass the content review.
[0199] For example, the content application can be a news application, and the content review task can be a news review task. The current user uploads the news to be published in the news application, and the server corresponding to the news application inputs the news to be published into the target news review model corresponding to the news review task to obtain the news review result corresponding to the news to be published. If the news review result indicates that the news to be published passes the news review, the server publishes the news to be published in the news application, so that other users can browse the news published by the current user in the news application.
[0200] In one embodiment, content review models corresponding to content review tasks of different themes can be trained respectively, so as to further improve the accuracy of content review.
[0201] In the above embodiments, the method of the present application can be applied to the content review task in content applications, allowing the approved content to be published to reduce the content review pressure of the content application and improve the content review efficiency.
[0202] In a specific embodiment, the method of the present application can be applied to the model training scenario for language models. The present application proposes a language model fine-tuning method that combines noise embedding and output data screening, which can solve the overfitting and data-dependence problems existing in the existing language model fine-tuning methods and obtain a language model that performs better on specific tasks.
[0203] Refer to Figure 10 , the model training is divided into two stages. In the first stage, the pre-trained language model is subjected to supervised training. Specifically, various text data are obtained, and based on the text data, the first training samples and the training labels corresponding to the first training samples are determined. For example, multiple question-and-answer pairs are constructed based on various text data, the questions in the question-and-answer pairs are used as the first training samples, and the answers in the question-and-answer pairs are used as the training labels corresponding to the first training samples. The first training samples are input into the pre-trained language model (i.e., the initial task processing model) after noise processing to obtain the first predicted labels corresponding to the first training samples. Based on the first predicted labels and the training labels corresponding to the first training samples, the model parameters of the pre-trained language model are adjusted, and through multiple rounds of model iterative training until the first convergence condition is met, an updated language model (i.e., the updated task processing model) is obtained. For noise processing, the words in the first training samples are first converted into embedding vectors, corresponding noise vectors are generated for each embedding vector, the noise vectors are added to the embedding vectors to obtain the noisy embedding vectors, and the noisy embedding vectors together form the first training samples after noise processing.
[0204] In the second stage, the model obtained from the first-stage training is further trained unsupervised. Specifically, various text data are acquired, and second training samples are determined based on the text data. For example, questions are constructed based on various text data, and the questions are used as the second training samples. The second training samples are input into the updated language model multiple times to obtain multiple second prediction labels corresponding to the second training samples and the occurrence probabilities (i.e., prediction probabilities) respectively corresponding to the multiple second prediction labels. Based on criteria such as relevance, coherence, and creativity corresponding to the second prediction labels, higher-quality second prediction labels are selected from the multiple second prediction labels as selection preferences (i.e., positive prediction labels), and lower-quality second prediction labels are selected from the multiple second prediction labels as rejection preferences (i.e., negative prediction labels). That is, given the input X, the model generates multiple responses Y1, Y2, ……, Yn, and based on the quality of the responses, several better responses are selected as selection preferences, and several worse responses are selected as rejection preferences. The model loss is calculated based on the prediction probability corresponding to the positive prediction label and the prediction probability corresponding to the negative prediction label, and the model parameters of the updated language model are adjusted based on the model loss. Through multiple rounds of model iterative training until the second convergence condition is met, the target language model (i.e., the target task processing model) is obtained. The training objective of the model is to maximize the occurrence probability of the selection preference while minimizing the occurrence probability of the rejection preference.
[0205] It can be understood that the method of this application has obvious advantages in the following aspects:
[0206] 1. Efficiency and scalability: Through noise embedding and output data screening, the training time and the dependence on high-quality labeled data are reduced, and the model performance can be effectively improved with limited computing resources. At the same time, it has good scalability and can easily adapt to different types of language models and task requirements.
[0207] 2. Reducing overfitting: Through noise embedding, the model can be prevented from overfitting to a specific dataset, and the generalization ability of the model in unseen data and complex tasks can be improved. The addition of noise makes the model more flexible and can generate more diverse responses.
[0208] 3. Reinforcement learning optimization: The quality of the generated responses can be improved through output data screening (i.e., determining selection preferences and rejection preferences), and the problems of complex reward design and unstable training in traditional reinforcement learning are reduced.
[0209] 4. Improving generation quality: After noise embedding and output data screening, in specific tasks, the quality of the responses generated by the model is significantly improved, and it can generate more natural, more coherent, and creative content to meet more complex task requirements.
[0210] 5. Flexible task adaptability: Supports training for various task types, including instruction following, dialogue generation, question-answering tasks, etc. By adjusting the intensity of noise embedding and the details of output data screening, it can adapt to the requirements of different tasks and ensure optimal performance.
[0211] In summary, the method of this application provides an efficient, stable, and flexible solution for language model training by integrating noise embedding and output data screening. It can not only significantly improve the generation ability of the model but also effectively reduce the training cost and reduce the dependence on a large amount of labeled data.
[0212] Moreover, the models trained by the method of this application have shown significant improvements in performance on various datasets. On the AlpacaEval test set, the accuracy of the model trained by the method of this application is 64.7%, while the accuracy of the model trained by the traditional method is 29.8%, with an accuracy improvement of 34.9%. On the Evol-Instruct dataset, the accuracy of the model trained by the method of this application is 79.6%, while the accuracy of the model trained by the traditional method is 70.3%, with an accuracy improvement of 9.3%. On the ShareGPT dataset, the accuracy of the model trained by the method of this application is 76.3%, while the accuracy of the model trained by the traditional method is 68.7%, with an accuracy improvement of 7.5%. On the OpenPlatypus dataset, the accuracy of the model trained by the method of this application is 70.6%, while the accuracy of the model trained by the traditional method is 62.0%, with an accuracy improvement of 8.6%.
[0213] In dialogue generation and open-ended question answering tasks, the models trained by the method of this application have shown significant improvements in the following aspects:
[0214] 1. Relevance: When generating content, the optimized model answers users' questions more accurately, and the relevance score has been significantly improved.
[0215] 2. Coherence: The responses generated by the model are more fluent and consistent, and the coherence score has been significantly improved.
[0216] The method of this application is suitable for application in small datasets and tasks with high noise interference, and can show significant improvements in the following aspects:
[0217] 1. Reduction in overfitting rate: The optimized model shows stable performance on the test set, and the gap between the training loss and the validation loss has been significantly reduced, indicating that overfitting has been effectively suppressed.
[0218] 2. Task execution accuracy: The optimized model shows stable performance on diverse tasks and has improved adaptability to new tasks and complex instructions.
[0219] In a specific embodiment, the method of the present application can be applied to news applications to reduce the news review pressure of news applications and improve the efficiency and accuracy of news review. Figure 11 , the server of the news application can obtain training data from the database, and obtain a news review model based on the training data through the aforementioned model training method. The user can log in to the news application on the user terminal, enter the news editing interface in the news application, and edit the news to be published in the news editing interface. The editing operations include filling in the news title, filling in the text, inserting pictures and other operations. After the editing is completed, click the "Publish" control. If the user clicks the "Publish" control, the user terminal requests the server of the news application to review the news to be published. The news application inputs the news to be published into the news review model, and the news review model outputs the review result. The server of the news application can return the review result to the user terminal. If the review result is passed, the server will publish the news edited by the user in the news application, so that other users can browse the news in the news application.
[0220] It can be understood that during the model training process, by introducing noise, the overfitting of the model to the training data is reduced. The addition of noise prevents the model from over-learning specific patterns and details in the data, thereby improving the model's generalization ability when faced with new, unseen inputs. By introducing noise, the model becomes more flexible and can generate more creative and coherent responses to diverse inputs. By introducing noise, the training effect of the model on a small amount or incompletely labeled data can be effectively improved, reducing dependence on a large amount of high-quality data. At the same time, during the model training process, combined with the screening of output data, the adaptability and efficiency of the model can be further improved. The model can gradually optimize its performance in dynamic adjustment without the support of a large amount of labeled data, thereby reducing data dependence and training costs.
[0221] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0222] Based on the same inventive concept, an embodiment of the present application further provides a model training apparatus for implementing the model training method involved above. The solution provided by this apparatus for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model training apparatus provided below can refer to the limitations on the model training method in the foregoing, and will not be elaborated herein.
[0223] In one embodiment, as Figure 12 shown, a model training apparatus is provided, including: a training data acquisition module 1202, a first training module 1204, and a second training module 1206, where:
[0224] The training data acquisition module 1202 is configured to acquire a first training sample, a second training sample corresponding to a target task, and a training label corresponding to the first training sample.
[0225] The first training module 1204 is configured to input the first training sample after noise processing into an initial task processing model corresponding to the target task, and obtain a first predicted label corresponding to the first training sample.
[0226] The first training module 1204 is further configured to adjust the model parameters of the initial task processing model based on the first predicted label and the training label corresponding to the first training sample, and obtain an updated task processing model.
[0227] The second training module 1206 is configured to input the second training sample into the updated task processing model, and obtain a plurality of second predicted labels corresponding to the second training sample.
[0228] The second training module 1206 is further configured to determine a positive predicted label and a negative predicted label from the plurality of second predicted labels based on the label quality corresponding to the second predicted label.
[0229] The second training module 1206 is further configured to adjust the model parameters of the updated task processing model based on the positive predicted label and the negative predicted label, and obtain a target task processing model corresponding to the target task.
[0230] In one embodiment, the first training module 1204 is further configured to:
[0231] perform feature transformation on the first training sample to obtain training embedded features;
[0232] acquire a target scaling factor corresponding to the target task;
[0233] perform scaling processing on the initial noise features based on the target scaling factor to obtain target noise features;
[0234] fuse the target noise features and the training embedded features to obtain the first training sample after noise processing.
[0235] In one embodiment, the first training module 1204 is further configured to:
[0236] Determine a target scaling factor corresponding to the target task based on the noise content of the original task data corresponding to the target task; the target scaling factor is positively correlated with the noise content.
[0237] In one embodiment, the first training module 1204 is further configured to:
[0238] Determine a target scaling factor corresponding to the target task based on the current training round in which the initial task processing model is located;
[0239] wherein, the target scaling factor decreases as the current training round increases, and the decreasing speed first decreases, then increases, and then decreases as the current training round increases.
[0240] In one embodiment, the second training module 1206 is further configured to:
[0241] Perform a correlation analysis on the second predicted label based on the second training sample to obtain the correlation degree corresponding to the second predicted label;
[0242] Perform a coherence analysis on the second predicted label to obtain the coherence degree corresponding to the second predicted label;
[0243] Perform a creativity analysis on the second predicted label to obtain the creativity degree corresponding to the second predicted label;
[0244] Fuse the correlation degree, coherence degree, and creativity degree corresponding to the second predicted label to obtain the label quality corresponding to the second predicted label.
[0245] In one embodiment, the second training module 1206 is further configured to:
[0246] Input the second predicted label into the coherence analysis model to obtain the coherence degree corresponding to the second predicted label;
[0247] wherein, the coherence analysis model is a model obtained by supervised training based on coherence data and incoherence data.
[0248] In one embodiment, the second training module 1206 is further configured to:
[0249] From multiple second predicted labels, obtain the second predicted labels with label quality greater than the first quality threshold as positive predicted labels, and obtain the second predicted labels with label quality less than the second quality threshold as negative predicted labels; wherein, the first quality threshold is greater than or equal to the second quality threshold.
[0250] In one embodiment, the update task processing model also outputs prediction probabilities corresponding to a plurality of second prediction labels. The second training module 1206 is further configured to:
[0251] Based on the prediction probability corresponding to the positive prediction label and the prediction probability corresponding to the negative prediction label, obtain a target loss; the target loss is negatively correlated with the prediction probability corresponding to the positive prediction label, and the target loss is positively correlated with the prediction probability corresponding to the negative prediction label;
[0252] Based on the target loss, adjust the model parameters of the update task processing model until the convergence condition is satisfied, and obtain the target task processing model corresponding to the target task.
[0253] In one embodiment, the second training module 1206 is further configured to:
[0254] Input the second training sample into the update task processing model after noise processing to obtain a plurality of second prediction labels corresponding to the second training sample.
[0255] In one embodiment, the second training module 1206 is further configured to:
[0256] Input the second training sample into the update task processing model to obtain a positive prediction label corresponding to the second training sample;
[0257] Input the second training sample into the update task processing model after noise processing to obtain a negative prediction label corresponding to the second training sample.
[0258] In one embodiment, the second training module 1206 is further configured to:
[0259] Input the second training sample into the update task processing model to obtain a plurality of second prediction labels corresponding to the second training sample, and determine a positive prediction label from the plurality of second prediction labels based on the label quality corresponding to the second prediction label.
[0260] The second training module 1206 is further configured to:
[0261] Input the second training sample into the update task processing model after noise processing to obtain a plurality of second prediction labels corresponding to the second training sample, and determine a negative prediction label from the plurality of second prediction labels based on the label quality corresponding to the second prediction label.
[0262] The above model training device first performs supervised training on the initial task processing model corresponding to the target task based on the first training samples corresponding to the target task and the training labels corresponding to the first training samples, to obtain an updated task processing model. During the supervised training process, the first training samples are input into the model after noise processing. The addition of noise helps reduce the overfitting of the model to the training data, preventing the model from overlearning specific patterns and details in the training data, thereby improving the generalization ability of the model when facing the target task. Further, unsupervised training is performed on the updated task processing model based on the second training samples to obtain the target task processing model corresponding to the target task. During the unsupervised training, the second training samples are input into the updated task processing model to obtain multiple second prediction labels. The second prediction labels with higher label quality are selected as positive prediction labels, and the second prediction labels with lower label quality are selected as negative prediction labels. Fine-tuning the updated task processing model based on the positive prediction labels and negative prediction labels helps the model learn how to make better decisions and improve the quality of the model output results. Combining the addition of noise and the selection of prediction labels can comprehensively improve the performance of the model in the target task and improve the model accuracy.
[0263] Each module in the above model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0264] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 13 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data involved in the model training method. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a model training method.
[0265] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 14 . The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a model training method. The display unit of the computer device is used to form a visually visible picture, which may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0266] Those skilled in the art can understand that Figure 13 , Figure 14 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0267] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0268] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0269] In one embodiment, a computer program product is provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0270] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0271] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., without limitation.
[0272] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0273] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A model training method, characterized in that, The method includes: Obtaining a first training sample and a second training sample corresponding to a target task, and a training label corresponding to the first training sample; Inputting the first training sample after noise processing into an initial task processing model corresponding to the target task to obtain a first predicted label corresponding to the first training sample; Adjusting model parameters of the initial task processing model based on the first predicted label and the training label corresponding to the first training sample to obtain an updated task processing model; Inputting the second training sample into the updated task processing model to obtain a plurality of second predicted labels corresponding to the second training sample; Determining a positive predicted label and a negative predicted label from the plurality of second predicted labels based on the label quality corresponding to the second predicted label; Adjusting model parameters of the updated task processing model based on the positive predicted label and the negative predicted label to obtain a target task processing model corresponding to the target task.
2. The method according to claim 1, characterized in that, The method further includes: Performing feature transformation on the first training sample to obtain training embedding features; Obtaining a target scaling factor corresponding to the target task; Performing scaling processing on initial noise features based on the target scaling factor to obtain target noise features; Fusing the target noise features and the training embedding features to obtain a first training sample after noise processing.
3. The method according to claim 2, wherein The obtaining of the target scaling factor corresponding to the target task includes: Determining the target scaling factor corresponding to the target task based on the noise content of the original task data corresponding to the target task; the target scaling factor is positively correlated with the noise content.
4. The method according to claim 2, wherein The obtaining of the target scaling factor corresponding to the target task includes: Determining the target scaling factor corresponding to the target task based on the current training round in which the initial task processing model is located; wherein, the target scaling factor decreases as the current training round increases, and the decreasing speed first decreases, then increases, and then decreases as the current training round increases.
5. The method according to claim 1, characterized in that, The method further includes: Performing correlation analysis on the second predicted labels based on the second training sample to obtain a correlation degree corresponding to the second predicted labels; Performing coherence analysis on the second predicted labels to obtain a coherence degree corresponding to the second predicted labels; Performing creativity analysis on the second predicted labels to obtain a creativity degree corresponding to the second predicted labels; Fusing the correlation degree, coherence degree, and creativity degree corresponding to the second predicted labels to obtain the label quality corresponding to the second predicted labels.
6. The method according to claim 5, wherein The performing of coherence analysis on the second predicted labels to obtain the coherence degree corresponding to the second predicted labels includes: Inputting the second predicted labels into a coherence analysis model to obtain the coherence degree corresponding to the second predicted labels; wherein, the coherence analysis model is a model obtained by supervised training based on coherence data and incoherence data.
7. The method according to claim 1, characterized in that, The determining of the positive predicted label and the negative predicted label from the plurality of second predicted labels based on the label quality corresponding to the second predicted labels includes: From the multiple second predicted labels, obtain the second predicted labels with label quality greater than the first quality threshold as positive predicted labels, and obtain the second predicted labels with label quality less than the second quality threshold as negative predicted labels; wherein, the first quality threshold is greater than or equal to the second quality threshold.
8. The method according to claim 1, characterized in that, The updated task processing model also outputs the prediction probabilities corresponding to the multiple second predicted labels respectively; Adjusting the model parameters of the updated task processing model based on the positive predicted labels and the negative predicted labels to obtain the target task processing model corresponding to the target task, including: Obtaining a target loss based on the prediction probability corresponding to the positive predicted label and the prediction probability corresponding to the negative predicted label; the target loss is negatively correlated with the prediction probability corresponding to the positive predicted label, and the target loss is positively correlated with the prediction probability corresponding to the negative predicted label; Based on the target loss, adjust the model parameters of the updated task processing model until the convergence condition is satisfied to obtain the target task processing model corresponding to the target task.
9. The method according to claim 1, characterized in that, Inputting the second training sample into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample, including: Input the second training sample into the updated task processing model after noise processing to obtain multiple second predicted labels corresponding to the second training sample.
10. The method according to claim 1, wherein Inputting the second training sample into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample, and determining positive predicted labels and negative predicted labels from the multiple second predicted labels based on the label quality corresponding to the second predicted labels, including: Input the second training sample into the updated task processing model to obtain the positive predicted label corresponding to the second training sample; Input the second training sample into the updated task processing model after noise processing to obtain the negative predicted label corresponding to the second training sample.
11. The method according to claim 10, characterized in that, Inputting the second training sample into the updated task processing model to obtain the positive predicted label corresponding to the second training sample, including: Input the second training sample into the updated task processing model to obtain multiple second predicted labels corresponding to the second training sample, and determine positive predicted labels from the multiple second predicted labels based on the label quality corresponding to the second predicted labels; Inputting the second training sample into the updated task processing model after noise processing to obtain the negative predicted label corresponding to the second training sample, including: Input the second training sample into the updated task processing model after noise processing to obtain multiple second predicted labels corresponding to the second training sample, and determine negative predicted labels from the multiple second predicted labels based on the label quality corresponding to the second predicted labels.
12. A model training device, characterized in that, The device includes: A training data acquisition module, configured to acquire a first training sample and a second training sample corresponding to a target task, and a training label corresponding to the first training sample; A first training module, configured to input the first training sample into an initial task processing model corresponding to the target task after noise processing to obtain a first predicted label corresponding to the first training sample; The first training module is further configured to adjust model parameters of the initial task processing model based on the first prediction labels and training labels corresponding to the first training samples, so as to obtain an updated task processing model; The second training module is configured to input the second training samples into the updated task processing model to obtain a plurality of second prediction labels corresponding to the second training samples; The second training module is further configured to determine positive prediction labels and negative prediction labels from the plurality of second prediction labels based on the label quality corresponding to the second prediction labels; The second training module is further configured to adjust model parameters of the updated task processing model based on the positive prediction labels and the negative prediction labels to obtain a target task processing model corresponding to the target task.
13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.