Data Processing Method and Device

By acquiring and training associated sample data and annotating target data of multiple labels, the problem of insufficient training data of multi-task deep learning models is solved, and the expansion of data sets and the improvement of model training effects is achieved.

CN114358313BActive Publication Date: 2025-07-18SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210006294.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-04
Publication Date
2025-07-18
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

When multitasking deep learning models are insufficient training data and high acquisition costs, it is difficult to train and expand the training data set effectively, resulting in poor model training results.

Method used

By obtaining the first sample data and the second sample data with a business association relationship therewith, the first business model and the second business model are trained respectively, and the sample data is input into each model to obtain target data marked with multiple labels, and the training data set is constructed.

Benefits of technology

The training data set is expanded, which reduces the cost of training data acquisition and model training difficulty, and improves the training effect of multi-task learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358313B_ABST
    Figure CN114358313B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method and apparatus. The method includes: obtaining first sample data and second sample data having a business association relationship with the first sample data; training a first business model based on the first sample data and a first sample label, and training a second business model based on the second sample data and a second sample label; inputting the first sample data into the second business model, and inputting the second sample data into the first business model; obtaining first target data output by the second business model and second target data output by the first business model; constructing a training data set based on the first target data and the second target data. By using multi-stage pre-training and using the first business model and the second business model for annotation, the problems of partial label missing and inconsistent definition between data sets are solved, the training data of the target business model is expanded, and the learning and training effect of the target business model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a data processing method. This application also relates to a data processing device, a computing device, and a computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence technology, multi-task deep learning models are increasingly applied. For example, in the field of face recognition, a person's identity can be recognized based on attributes such as the nose, eyes, hairstyle, etc. During the training process of a multi-task deep learning model, the multi-task deep learning model often requires a large amount of data with all labeled tags. However, due to the difficulties in collecting training data with all labeled tags and the high acquisition cost, the quantity of training data for the multi-task learning model is insufficient, resulting in difficult model training and poor training effects. Therefore, in the case of a small quantity of training data for the multi-task learning model, how to expand the quantity of training data so as to better train the multi-task learning model and reduce the difficulty of model training is an urgent problem to be solved currently. Summary of the Invention

[0003] In view of this, embodiments of this application provide a data processing method. This application also relates to a data processing device, a computing device, and a computer-readable storage medium to solve the problems of insufficient training data and high acquisition cost in the prior art.

[0004] According to a first aspect of embodiments of this application, a data processing method is provided, including:

[0005] Obtain first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with a first sample tag, and the second sample data is labeled with a second sample tag;

[0006] Train a first business model based on the first sample data and the first sample tag, and train a second business model based on the second sample data and the second sample tag;

[0007] Input the first sample data into the second business model, and input the second sample data into the first business model;

[0008] Obtain first target data output by the second business model and second target data output by the first business model, where both the first target data and the second target data are labeled with the first sample tag and the second sample tag;

[0009] Construct a training data set based on the first target data and the second target data.

[0010] According to the second aspect of the embodiments of the present application, another data processing method is provided, including:

[0011] Obtain at least two initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with a corresponding training label;

[0012] Train a corresponding initial business model according to each initial sample set;

[0013] Process each initial sample set through each initial business model based on a preset rule;

[0014] Construct a training data set according to the processing results of each initial business model.

[0015] According to the third aspect of the embodiments of the present application, a data processing device is provided, including:

[0016] A first acquisition module configured to acquire first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with a first sample label and the second sample data is labeled with a second sample label;

[0017] A training module configured to train a first business model according to the first sample data and the first sample label, and train a second business model according to the second sample data and the second sample label;

[0018] An input module configured to input the first sample data into the second business model and input the second sample data into the first business model;

[0019] A second acquisition module configured to acquire first target data output by the second business model and second target data output by the first business model, where both the first target data and the second target data are labeled with a first sample label and a second sample label;

[0020] A construction module configured to construct a training data set based on the first target data and the second target data.

[0021] According to the fourth aspect of the embodiments of the present application, another data processing device is provided, including:

[0022] An acquisition module configured to acquire at least two initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with a corresponding training label;

[0023] A training module configured to train a corresponding initial business model according to each initial sample set;

[0024] A processing module, configured to process each initial sample set through each initial business model based on a preset rule;

[0025] A construction module, configured to construct a training data set according to the processing result of each initial business model.

[0026] According to a fifth aspect of the embodiments of the present application, there is provided a computing device, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the computer instructions, the steps of the data processing method are implemented.

[0027] According to a sixth aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the steps of the data processing method are implemented.

[0028] The data processing method provided by the present application includes: obtaining first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label; training a first business model according to the first sample data and the first sample label, and training a second business model according to the second sample data and the second sample label; inputting the first sample data into the second business model, and inputting the second sample data into the first business model; obtaining first target data output by the second business model and second target data output by the first business model, where both the first target data and the second target data are labeled with the first sample label and the second sample label; constructing a training data set based on the first target data and the second target data.

[0029] An embodiment of the present application realizes that by inputting the first sample data into the second business model and obtaining the first target data output by the second business model, the first target data is labeled with both the first sample label and the second sample label, inputting the second sample data into the first business model, and obtaining the second target data output by the first business model, the second target data is labeled with both the first sample label and the second sample label, thereby expanding the training data set of the target business model and reducing the training data acquisition cost and the model training difficulty. Description of the Drawings

[0030] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present application;

[0031] Figure 2 is a processing flowchart of a data processing method applied to a text recognition model provided by an embodiment of the present application;

[0032] Figure 3 It is the training architecture diagram of the second service model provided by an embodiment of the present application;

[0033] Figure 4 It is the flowchart of another data processing method provided by an embodiment of the present application;

[0034] Figure 5 It is the structural schematic diagram of a data processing device provided by an embodiment of the present application;

[0035] Figure 6 It is the structural schematic diagram of another data processing device provided by an embodiment of the present application;

[0036] Figure 7 It is the structural block diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0037] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0038] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.

[0039] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0040] First, the noun terms related to one or more embodiments of the present application are explained.

[0041] Multitask Learning: Multitask learning involves learning multiple tasks with business relevance together. That is, when multiple objective loss functions are learned simultaneously, it is called multitask learning. For example, in face recognition technology, it is usually necessary to predict information in multiple dimensions such as eyes, nose, and mouth in a face. The prediction tasks of these multiple dimensions of information can be learned by multiple models for single-task learning, or by a single model for multitask learning. Multitask learning has better results in model learning compared to single-task learning and can better train the model.

[0042] The purpose of multitask learning is to use a single model to output multiple labels and for the model to share some weights. This does not exist in traditional linear learning models, but there are usually branches in multitask learning models. The part before the branch point belongs to the shared weights, and the same loss function is used for each different task. After the branch point, each task has its own independent weights, and different tasks can use different loss functions.

[0043] Before the branch point, all tasks share the same weights, so it only needs to be calculated once. Compared with using multiple single-task learning models for learning and training, a large amount of computing time can be saved. And most of the computing time is spent on feature extraction. The closer the branch point is to the front, the fewer shared weights there are, and the closer the branch point is to the back, the more shared weights there are. Thus, by sharing weights, the purpose of reducing the number of model parameters, saving model space, and reducing the computational amount can be achieved. Moreover, during multitask learning, there is business relevance between each task, so the purpose of mutual supervised learning to improve the model learning effect can be achieved.

[0044] Transfer learning: Transfer learning is a technique commonly used in deep learning. Its purpose is to modify and retrain a model adapted to task A so that it can adapt to task B, and it can shorten the training time, and sometimes it can even be better than not using transfer learning. Its principle is that on a larger dataset A, the model can better learn features. If we only use dataset B and dataset B is small, the model features will be learned poorly. So we can keep the feature layer and weights of the model trained on dataset A, replace the output layer with the output result we want, and then retrain.

[0045] Learning without Forgetting: Learning without Forgetting is an enhancement of multitask learning. Currently, Learning without Forgetting is achieved through model annotation. First, learn one type of label and use it to annotate another type of label. The explanation for doing this is that the purpose of using the label annotated by the model is to prevent the model from forgetting the previously learned effects.

[0046] At present, when transfer learning is applied to multi-task learning, it is unable to output the labels of both A and B simultaneously. After learning B, the model will completely lose its memory of A and can no longer be used for A, which is a functional deficiency. Moreover, the training of a multi-task learning model requires a large amount of data with all the output labels annotated. For example, in face recognition technology, if one wants to output the key point coordinates of both the eye sockets and the pupils, among the existing common data, there are only a large number of data containing the key point coordinates of the eye sockets and a large number of data containing the key point coordinates of the pupils, but there is less data that simultaneously contains the key point coordinates of both the eye sockets and the pupils. At this time, there are two training methods:

[0047] The first method: Only use the data that simultaneously contains the key point coordinates of the eye sockets and the pupils for model training. At this time, a large amount of data with only partial labels will be wasted.

[0048] The second method: Only use a large amount of data with partial labels. However, each time only training partial labels may, due to the data distribution, result in only being able to learn one label each time, thus leading to poor learning effects.

[0049] In addition, because there are still differences in the definition of the training dataset, when using different datasets, the model may not be able to learn the expected results. For example, in the training dataset containing the key point coordinates of the eye sockets, some data contains the key point coordinates outside the eye sockets, and some data contains the key point coordinates outside the eye sockets. Therefore, the model cannot determine which definition to use, resulting in the problem of poor learning effects of the model.

[0050] In the memory learning method, the currently used model does not use an annotation model but uses the model to be produced for annotation. The disadvantage of this is that the annotation is not the limit that can be achieved with the current dataset. Multi-task learning has the potential to improve the effect, so use labels closer to the limit for learning. And this method will, due to some labels being manually annotated and some labels being model-annotated, lead to inconsistent domains, resulting in the problem of poor model training effects.

[0051] Based on this, in this application, a data processing method is provided. This application also involves a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0052] Figure 1 The flowchart of a data processing method provided according to an embodiment of the present application is shown, which specifically includes the following steps:

[0053] Step 102: Obtain first sample data and second sample data that has a business association relationship with the first sample data. Among them, the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label.

[0054] Among them, the first sample data can be understood as the training data when the model is trained. For example, in a face recognition model, the first sample data can be a face photo. In a spam detection model, the first sample data can be an email. The first sample label can be understood as the thing predicted by the model. The first sample data can be labeled using the first sample label. For example, if the first sample label is to identify an apple, and the first sample data is a picture containing multiple fruits, labeling the first sample data with the first sample label means marking the apple in the picture containing multiple fruits. Another example, if the first sample label is to identify the position of the mouth in a face, and the first sample data is a face picture, using the first sample label to label the first sample data means marking the key point coordinates of the mouth in the face picture.

[0055] In practical applications, the first sample label can also be understood as the correct result that the model expects to output. After the model outputs a prediction result, the prediction result can be compared with the first sample label to determine whether the prediction result is correct.

[0056] The second sample data can be understood as the training data when the model is trained. The second sample data and the first sample data have a business association relationship. Among them, having a business association relationship means having a relationship of the same business field and similar business tasks. For example, in the face recognition business field, both the first sample data and the second sample data are pictures containing face information. The first sample data is a picture with the key point coordinates of the eyes marked, and the business task corresponding to the first sample data is "identifying the eyes"; the second sample data is a picture with the key point coordinates of the pupils marked, and the business task corresponding to the second sample data is "identifying the pupils".

[0057] The first sample data and the second sample data can be the same original file. For example, taking "Figure A is a face picture" as an example, the first sample data is Figure A with the key point coordinates of the eyes marked, and the second sample data is Figure A with the key point coordinates of the pupils marked. The first sample data and the second sample data can also be different original files. Taking "Figure A is a face picture and Figure B is a full-body photo" as an example, the first sample data is Figure A with the key point coordinates of the pupils marked, and the second sample data is Figure B with the key point coordinates of the eyes marked.

[0058] In the field of text information recognition services, both the first sample data and the second sample data are natural texts. The first sample data is a natural text with verbs marked, and the corresponding service task of the first sample data is "identifying verbs". The second sample data is a natural text with nouns marked, and the corresponding service task of the second sample data is "identifying nouns".

[0059] The first sample data and the second sample data can be the same original file. For example, there is a natural text A. The first sample data is the natural text A with verbs marked, and the second sample data is the natural text A with nouns marked. The original files of both the first sample data and the second sample data are the natural text A. The first sample data and the second sample data can also be different original files. For example, there is a natural text A and a natural text B. The first sample data is the natural text A with verbs marked, and the second sample data is the natural text B with nouns marked.

[0060] The second sample label is also different from the first sample label according to different tasks. For example, for task A of identifying eyes from a human face, the first sample label is obtaining the key point coordinates of the eyes. For task B of identifying pupils from a human face, the second sample label is obtaining the key point coordinates of the pupils. Another example is that for task A of obtaining nouns from a natural text, the first sample label is identifying nouns. For task B of obtaining verbs from a natural text, the second sample label is identifying verbs. Although the tasks of the first sample label and the second sample label are different, the two tasks have a similar business relationship.

[0061] In a specific embodiment of the present application, the first sample data is obtained, and the first sample data is marked with the first sample label; the second sample data is obtained, and the second sample data is marked with the second sample label. Among them, the first sample data is 100 pictures with the positions of the mouths marked, and the second sample data is 100 pictures with the positions of the noses marked.

[0062] In another specific embodiment of the present application, the first sample data is obtained, and the first sample data is marked with the first sample label; the second sample data is obtained, and the second sample data is marked with the second sample label. Among them, the first sample data is 100 emails marked with the sending time, and the second sample data is 100 emails marked with the sending address.

[0063] Step 104: Train a first service model based on the first sample data and the first sample label, and train a second service model based on the second sample data and the second sample label.

[0064] Among them, the first business model can be understood as a trained model that can handle business normally, which can input the first sample data and output the first sample label. For example, the first business model is a face recognition model, and its purpose is to identify the position of the mouth in the face. After inputting a face picture into the first business model, the output result of the first business model is a face picture with the position of the mouth marked.

[0065] The second business model can also be understood as a trained model that can sort out business normally, which can input the second sample data and output the second sample label. For example, the second business model is a text recognition model, and its purpose is to detect nouns in natural text. After inputting natural text into the second business model, the output result of the second business model is the nouns in the natural text.

[0066] In practical applications, although the second business model and the first business model handle different businesses, the businesses they handle are related. For example, the business handled by the first business model is to identify the position of the mouth in the face, and the business handled by the second business model is to identify the position of the eyes in the face. Another example is that the business handled by the first business model is to identify verbs in natural text, and the business handled by the second business model is to identify nouns in natural text.

[0067] In a specific embodiment of the present application, following the above example, 100 pictures with the position of the mouth marked are sequentially input into the first initial model, the prediction results output by the first initial model are obtained, the loss value of the first initial model is calculated according to the prediction results and the first sample label, the parameters of the first initial model are adjusted according to the loss value, and the input training data continues to be input until the parameters are adjusted until the first initial model can correctly output the expected result. At this time, the first initial model has been trained into the first business model. 100 pictures with the position of the nose marked are sequentially input into the second initial model, the prediction results output by the second initial model are obtained, the loss value of the second initial model is calculated according to the prediction results and the second sample label, and the parameters of the second initial model are adjusted according to the loss value until the parameters are adjusted until the second initial model can correctly output the expected result. At this time, the second initial model has been trained into the second business model.

[0068] Step 106: Input the first sample data into the second business model and input the second sample data into the first business model.

[0069] Among them, inputting the first sample data into the second business model can be understood as annotating the first sample data with the second sample label; inputting the second sample data into the first business model can be understood as annotating the second sample data with the first sample label.

[0070] In practical applications, since the first sample data is labeled with the first sample label and the second sample data is labeled with the second sample label, but for the final model training, training data with all label outputs is required. It is possible to manually label the first sample data and the second sample data, but this method is labor-intensive and costly in terms of human resources and annotation costs. Therefore, a model annotation method can be adopted, that is, the sample label is annotated to other data that does not have this sample label annotated.

[0071] For example, if the existing A training dataset is labeled with the a sample label and the B training dataset is labeled with the b sample label, the B training dataset is annotated with the a sample label to generate a B training dataset that is labeled with both the a sample label and the b sample label. The A training dataset is labeled with the b sample label to generate an A training dataset that is labeled with both the a sample label and the b sample label. This not only expands the training data with all labels, solves the waste of resources in some training data, but also solves the problem that the model training effect is not good due to training the model with data with only some labels.

[0072] In a specific embodiment of the present application, following the above example, 100 pictures with the position of the mouth marked are sequentially input into the second service model to obtain the output result of the second service model. 100 pictures with the position of the nose marked are sequentially input into the first service model to obtain the output result of the first service model.

[0073] In another specific embodiment of the present application, following the above example, 100 emails with the sending time marked as the first sample data are input into the second service model to obtain the output result of the second service model. The output result of the second service model is the sending time and sending address of each of the 100 emails. 100 emails with the sending address marked as the second sample data are input into the first service model to obtain the output result of the first service model. The output result of the first service model is the sending address and sending time of each of the 100 emails.

[0074] Step 108: Obtain the first target data output by the second service model and the second target data output by the first service model, where both the first target data and the second target data are labeled with the first sample label and the second sample label.

[0075] Among them, the first target data can be understood as the data output by the second service model based on the first sample data, and the second target data can be understood as the data output by the first service model based on the second sample data.

[0076] In practical applications, since the second business model is already a trained business model, it is possible to perform business processing on the input training data, that is, label the training data. Then, the output result, the first target data, is labeled with both the label carried by the input training data and the label labeled by the second business model. Similarly, the second target data is also labeled with two labels. Thus, the training data with only partial labels is expanded into training data with all labels, achieving the purpose of expanding the training data set, providing sufficient training data for the training of the target business model, and solving the problem of poor training effect caused by sparse training data.

[0077] In a specific embodiment of the present application, following the above example, 100 pictures with the position of the mouth marked are sequentially input into the second business model to obtain the first target data output by the second business model. The first target data is 100 pictures with both the position of the mouth and the position of the nose marked; 100 pictures with the position of the nose marked are sequentially input into the first business model to obtain the second target data output by the first business model. The second target data is 100 pictures with both the position of the nose and the position of the mouth marked.

[0078] In another specific embodiment of the present application, following the above example, 100 emails marked with the sending time are used as the first sample data and input into the second business model to obtain the first target data output by the second business model. The first target data is 100 emails marked with the sending time and the sending address; 100 emails marked with the sending address are used as the second sample data and input into the first business model to obtain the second target data output by the first business model. The second target data is 100 emails marked with the sending time and the sending address.

[0079] Step 110: Construct a training data set based on the first target data and the second target data.

[0080] Among them, the training data set can be understood as the training data of the target business model. For example, if the first target data is 20 pictures marked with eyes and mouth, and the second target data is 30 pictures marked with eyes and mouth, then the training data set is 50 pictures marked with eyes and mouth. When training the target business model, this training data set can be used.

[0081] In practical applications, for business models in certain fields, due to the particularity of the fields, it is difficult to collect relevant training data. Therefore, the data processing method provided in the present application can be adopted to expand the training data with partial labels into training data with all labels, provide sufficient training data for the business model, and ensure the training effect of the business model.

[0082] In a specific embodiment of the present application, following the previous example, 100 pictures labeled with mouths and noses and 100 pictures labeled with mouths and noses are combined to construct 200 pictures labeled with noses and mouths.

[0083] In another specific embodiment of the present application, following the previous example, 100 emails labeled with sending time and sending address and 100 emails labeled with sending address and sending time are combined to construct a training dataset, and the training dataset is 200 emails labeled with sending address and sending time.

[0084] A data processing method provided by the present application includes: obtaining first sample data and second sample data having a business association relationship with the first sample data, wherein the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label; training a first business model according to the first sample data and the first sample label, and training a second business model according to the second sample data and the second sample label; inputting the first sample data into the second business model, and inputting the second sample data into the first business model; obtaining first target data output by the second business model and second target data output by the first business model, wherein both the first target data and the second target data are labeled with the first sample label and the second sample label; constructing a training dataset based on the first target data and the second target data. By first training the first business model and the second business model with the first sample data and the second sample data respectively, and then inputting the first sample data into the second business model to obtain the first target data, and inputting the second sample data into the first business model to obtain the second target data, the first sample data and the second sample data originally labeled with only partial labels are expanded into the first target data and the second target data labeled with all labels, and the training dataset of the target business model is expanded by using the first sample data and the second sample data, avoiding waste of training data with only partial sample labels, and at the same time enabling the target business model to have sufficient training data during training, improving the training and learning effect of the target business model.

[0085] The following combines the attached Figure 2 , taking the application of the data processing method provided by the present application in a text recognition model as an example, to further illustrate the data processing method. Among them, Figure 2 shows a processing flowchart of a data processing method applied to a text recognition model provided by an embodiment of the present application, specifically including the following steps:

[0086] Step 202: Obtain first sample data and second sample data that has a business association relationship with the first sample data. Among them, the first sample data is labeled with a first sample label, and the second sample data includes second sample reference data and a second sample reference label corresponding to the second sample reference data, second sample target data, and a second sample target label corresponding to the second sample target data.

[0087] Among them, the second sample target data is the training data to be augmented, and the second sample reference data is the training data that is differently defined from the second sample target data. For example, the second sample target data is to obtain the coordinate points of the upper lip of the mouth, and the second sample reference data is to obtain the coordinate points of the lower lip of the mouth. The definitions of the second sample target label and the second sample reference label are also different.

[0088] In a specific embodiment of the present application, following the above example, obtain first sample data. The first sample data is labeled with a first sample label. The first sample data is 30 natural text sentences, which are labeled with nouns in the natural text sentences; obtain second sample data. The second sample data is 20 natural text sentences, among which 10 are second sample reference data, labeled with verbs in the natural text sentences, and 10 are second sample target reference data, labeled with central verbs in the natural text sentences.

[0089] Step 204: Train a first business model based on the first sample data and the first sample label, and train a second business model based on the second sample data and the second sample label.

[0090] Training a second business model based on the second sample data and the second sample label further includes:

[0091] Train a second pre-trained business model based on the second sample reference data and the second sample reference label;

[0092] Train the second pre-trained business model based on the second sample target data and the second sample target label to obtain a second business model.

[0093] Among them, the second pre-trained business model can be understood as a model trained by reference training data and reference labels. The annotation of this pre-trained model is different from that of the second business model. For example, the second pre-trained business model is to identify verbs in natural text sentences. Input "Xiaoming ran back and opened the door" into the second pre-trained business model, and the output result of the second pre-trained business model is "ran, opened". The second business model is to identify the central verb in natural text sentences. Input "Xiaoming ran back and opened the door" into the second business model, and the output result of the second business model is "opened".

[0094] In practical applications, a second sample reference data that is different from the second sample target data can be used to pre-train a business model, and then the learning rate is reduced and this business model is used for transfer learning of the second sample target data to obtain a second business model. Since the dataset defined in advance is used in the second training, the annotation of the final model will conform to the definition required by us. As Figure 3 shown, Figure 3 FIG. 5 is a training architecture diagram of a second business model provided by an embodiment of the present application. Among them, the second sample reference data and the second sample reference label are first input into the second initial business model to train and obtain a second pre-trained business model, and then the second sample target data and the second sample target label are used to train the second pre-trained business model to obtain a second business model.

[0095] In a specific embodiment of the present application, following the above example, a first business model is trained according to 30 natural text sentences of the first sample data and the first sample label. A second business pre-trained model is trained according to 10 natural text sentences of the second sample reference data and the second sample reference label. The second business pre-trained model can only recognize the verbs in the natural text sentences. A second business model is trained according to 10 natural texts of the second sample target data and the second sample target label. The second business model can recognize the central verbs in the natural text sentences.

[0096] Before training the second pre-trained business model according to the second sample target data and the second sample target label, the method further includes:

[0097] Receiving a parameter adjustment instruction;

[0098] Responding to the parameter adjustment instruction to adjust the target parameters of the second pre-trained business model.

[0099] Among them, the parameter adjustment instruction can be understood as an instruction issued by the user for the model. Through this instruction, the parameters of the model can be adjusted. The target parameters can include model parameters such as the learning rate parameter and batch size of the model.

[0100] In practical applications, the user can adjust the parameters of the model to train a model that meets the user's expectations more efficiently and accurately.

[0101] In a specific example of the present application, following the above example, receive a parameter adjustment instruction, and then respond to the parameter adjustment instruction to adjust the target parameters of the second pre-trained business model.

[0102] Specifically, responding to the parameter adjustment instruction to adjust the target parameters of the second pre-trained business model includes:

[0103] Adjust the learning rate parameter of the second pre-trained service model in response to the parameter adjustment instruction to reduce the learning rate of the second pre-trained service model.

[0104] Among them, the learning rate parameter can be understood as the learning efficiency of the model. The learning rate, as an important hyperparameter in supervised learning and deep learning, determines whether the objective function can converge to the local minimum and when it converges to the minimum. A suitable learning rate can make the objective function converge to the local minimum within an appropriate time.

[0105] In practical applications, since the second pre-trained model has learned the second sample reference data, when training the second and training models based on the second sample target data, there is no need to start learning from scratch. An overly high learning rate is likely to cause too high a loss of the model, affecting the results we have learned well before.

[0106] In a specific embodiment of the present application, following the above example, reduce the learning rate of the second pre-trained service model in response to the parameter adjustment instruction.

[0107] Step 206: Input the first sample data into the second service model, and input the second sample data into the first service model.

[0108] In a specific embodiment of the present application, following the above example, input 30 natural text sentences labeled with the first sample label into the second service model, and input 10 natural text sentences labeled with the second sample reference label and 10 natural text sentences labeled with the second sample target reference label into the first service model.

[0109] Step 208: Obtain the first target data output by the second service model and the second target data output by the first service model, where both the first target data and the second target data are labeled with the first sample label and the second sample label.

[0110] In a specific embodiment of the present application, following the above example, obtain the first target data output by the second service model. The first target data is 30 natural text sentences labeled with nouns and central verbs, and obtain the second target data output by the first service model. The second target data is 20 natural text sentences labeled with central verbs and nouns.

[0111] Step 210: Construct a training data set based on the first target data and the second target data.

[0112] In a specific embodiment of the present application, following the above example, construct 50 natural text sentences labeled with nouns and central verbs from 30 natural text sentences labeled with nouns and central verbs and 20 natural text sentences labeled with central verbs and nouns.

[0113] After constructing a training data set based on the first target data and the second target data, the method further includes:

[0114] Training a target business model based on the training data set.

[0115] The training data set includes target data, and the target data is labeled with a first sample label and a second sample label;

[0116] Among them, the target business model can be understood as the business model that the user ultimately wants. The target business model can output results with all labels.

[0117] In practical applications, the training data set includes first target data generated by annotation of a second business model and second target data generated by annotation of a first business model. Both the first target data and the second target data are labeled with a first sample label and a second sample label. Based on the training data set, a target business model can be trained.

[0118] In a specific embodiment of the present application, following the above example, the training data set composed of the first target data and the second target data is input into the target business model to train the target business model.

[0119] Specifically, training a target business model based on the training data set includes:

[0120] Inputting the target data into the target business model;

[0121] Obtaining a first predicted label and a second predicted label output by the target business model;

[0122] Calculating a model loss value based on the first predicted label, the first sample label, the second predicted label, and the second sample label;

[0123] Adjusting the model parameters of the target business model according to the model loss value, and continuing to train the target business model until a model training stop condition is reached.

[0124] Among them, the first predicted label can be understood as the output result of the target business model after the first target data is input; the second predicted label can be understood as the output result of the target business model after the second target data is input; the first predicted label and the second predicted label may be inconsistent with the model output result expected by the user. Therefore, it is necessary to calculate the model loss value and adjust the model parameters according to the model loss value to improve the output accuracy of the model.

[0125] In a specific embodiment of the present application, following the previous example, 50 natural text sentences already labeled with sample labels are input into the target business model to obtain the first predicted label and the second predicted label output by the target business model. The loss value of the model is calculated based on the first predicted label and the first sample label, and the second predicted label and the second sample label. The model parameters are adjusted according to the loss value until the model training stop condition is reached.

[0126] Specifically, reaching the model training stop condition includes:

[0127] The model loss value is less than the preset loss value threshold; and / or

[0128] The number of training rounds reaches the preset number of training rounds.

[0129] Among them, the preset loss value threshold can be understood as the expected loss value set by the user. When it is less than this preset loss value threshold, it means that the current model has been trained and meets the standard expected by the user.

[0130] The number of training rounds can be understood as the number of times the model uses sample data for training; the preset number of training rounds can be understood as the number of times the model uses sample data set by the user. After the model uses sample data to reach the preset number of training rounds, the model stops training.

[0131] In a specific embodiment of the present application, taking the example of stopping the training of the target business model by the loss value being less than the preset loss value threshold, if the preset loss value threshold is 0.5, then when the calculated Loss value is less than 0.5, it is determined that the training of the target business model is completed.

[0132] In another specific embodiment of the present application, taking the example of stopping the training of the target business model by the preset number of training rounds, if the preset number of training rounds is 20 rounds, when the number of training rounds of the sample data reaches 20 rounds, it is determined that the training of the target business model has been completed.

[0133] In practical applications, since the definition of the training data labeled by the model may be different from that of manual labeling, the output results of the target business model trained by the training data set labeled by the model may have errors. To further improve the output accuracy of the target business model, target training data with all labels manually labeled can also be added to the training data set. The purpose is to obtain a target business model with better model output while making the best use of as much data as possible. Since the target training data is manually labeled, the business model trained by the target training data will have better effects. Therefore, the training data set other than the target training data can be used for pre-training to obtain the target pre-trained business model. After training is completed, the learning rate is adjusted to a smaller value, and then the trained model is used for transfer learning of the target training data with all labels. Because the labeling of the target training data is complete and correct, the output results of the final model will be very good, and will be better than the output results of models with inconsistent label definitions or sparse data sets.

[0134] In a specific embodiment of the present application, a target pre-trained business model is trained based on 50 natural text sentences, a first sample label, and a second sample label. The learning rate parameter of the target pre-trained business model is adjusted and decreased, and then the existing natural text sentences manually labeled with the first sample label and the second sample label are input into the target pre-trained business model to train a target business model with better output effects and more accurate results.

[0135] A data processing method provided by the present application includes: obtaining first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with a first sample label, and the second sample data includes second sample reference data and a second sample reference label corresponding to the second sample reference data, second sample target data, and a second sample target label corresponding to the second sample target data. A first business model is trained according to the first sample data and the first sample label, and a second business model is trained according to the second sample data and the second sample label. The first sample data is input into the second business model, and the second sample data is input into the first business model. First target data output by the second business model and second target data output by the first business model are obtained, where both the first target data and the second target data are labeled with the first sample label and the second sample label. A training data set is constructed based on the first target data and the second target data. By using multi-stage pre-training and using the first business model and the second business model for labeling, the problems of partial label loss and inconsistent definitions between data sets are solved, the training data of the target business model is expanded, and the learning and training effects of the target business model are improved.

[0136] Figure 4 The flowchart of another data processing method provided according to an embodiment of the present application is shown, which specifically includes the following steps:

[0137] Step 402: Obtain at least three initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with a corresponding training label.

[0138] Among them, there is also a business association relationship between the three initial sample sets. The business association relationship can be understood as that each initial sample set is a relationship of the same data type. For example, the initial sample set A is a face picture, the initial sample set B is an eye picture, and the initial sample set C is a pupil picture. The training label is the sample label corresponding to the sample data.

[0139] In practical applications, the number of initial sample sets can be many. The data processing method provided by the present application is not only applicable to sample data with two different labels, but also applicable to sample data with two or more different labels. In the case of obtaining at least three initial sample sets, the data processing method provided by the present application can also be used to expand the training data set based on the initial sample sets.

[0140] In a specific embodiment of the present application, continuing with the above example, obtain the initial sample set A, the initial sample set B, and the initial sample set C. The initial sample set A includes A sample data and the corresponding A training label, the initial sample set B includes B sample data and the corresponding B training label, and the initial sample set C includes C sample data and the corresponding C training label.

[0141] Step 404: Train a corresponding initial business model according to each initial sample set.

[0142] Among them, the initial business model can be understood as inputting a kind of sample data and outputting the output result of a sample label corresponding to the sample data. For example, the initial business model A inputs A sample data, and the output result is labeled with the output result of the A training label corresponding to the A sample data.

[0143] In practical applications, the data processing method provided by the present application is not limited to two annotations, and its extensibility has no limit. Just add the number of initial business models as needed, and then train the corresponding initial business models according to each different sample data and its corresponding training label.

[0144] In a specific embodiment of the present application, continuing with the above example, train the initial business model A according to the A sample data and the A training label, train the initial business model B according to the B sample data and the B training label, and train the initial business model C according to the C sample data and the C training label.

[0145] Specifically, training a corresponding initial business model according to each initial sample set includes:

[0146] Determining a target initial sample set from the at least three initial sample sets;

[0147] Training a corresponding target initial business model according to the target initial sample set.

[0148] Among them, the target initial sample set can be understood as an initial sample set selected from multiple initial sample sets. For example, there are 4 initial sample sets A, B, C, and D. If the initial sample set A is selected, then the initial sample set A is the target initial sample set. The target initial business model can be understood as a business model obtained by training according to the target initial sample set.

[0149] In a specific embodiment of the present application, continuing with the above example, among the initial sample set A, the initial sample set B, and the initial sample set C, the initial sample set B is selected as the target initial sample set, and the B initial business model is trained according to the B sample data and the B training labels in the initial sample set B. The B initial business model is the target initial business model.

[0150] Specifically, training a corresponding initial business model for any one initial sample set includes:

[0151] Determining a target initial sample set from the at least three initial sample sets;

[0152] Inputting the target sample data in the target initial sample set into the target initial business model;

[0153] Obtaining the target prediction label output by the target initial business model;

[0154] Calculating a model loss value based on the target prediction label and the target training label corresponding to the target sample data;

[0155] Adjusting the model parameters of the target initial business model according to the model loss value, and continuing to train the target initial business model until the model training stop condition is reached.

[0156] Among them, the target initial business model can be understood as the business model to be trained. The target sample data of the target initial business model can be understood as the sample data in the target initial sample set, the target training label can be understood as the training label in the target initial sample set, and the target prediction label can be understood as the output result of the target initial business model.

[0157] In practical applications, the training methods of each initial business model are the same, but the sample data and training labels used in the training of each initial business model are different.

[0158] In a specific embodiment of the present application, following the above example, a target initial sample set is determined among the initial sample set A, the initial sample set B, and the initial sample set C, and the target initial sample set is the initial sample set B. The B sample data in the initial sample set B is input into the target initial business model, and the B target prediction label output by the target initial business model is obtained. The model loss value of the target initial business model is calculated based on the B target prediction label and the B target training label. The model parameters of the target initial business are adjusted according to the model loss value, and the B sample data is continuously used to train the target initial business model until the model training stop condition is reached.

[0159] Step 406: Process each initial sample set through each initial business model based on a preset rule.

[0160] Processing each initial sample set through each initial business model based on a preset rule includes:

[0161] Determine the target initial sample set;

[0162] The target initial sample set is sequentially input into each initial business model except the target initial business model corresponding to the target initial sample set, and the processing results output by each initial business model are obtained.

[0163] Among them, the preset rule can be understood as inputting the initial sample set into the initial business model obtained by training with other initial sample sets.

[0164] In practical applications, in order to expand the training data of the final target business model, it can be achieved by mutual annotation between each initial sample set.

[0165] In a specific embodiment of the present application, after training the corresponding initial business model through each initial sample set, the initial business model A, the initial business model B, and the initial business model C are obtained. The initial sample set A is used as the target initial sample set, and the initial sample set A is sequentially input into the initial business model B and the initial business model C. The output result of the initial business model B is the B+A sample data, and the B+A sample data is labeled with the training label A and the training label B. The output result of the initial business model C is the C+A sample data, and the C+A sample data is labeled with the training label A and the training label C.

[0166] Similarly, finally, the A+B initial sample set and the A+C initial sample set output by the initial business model A, the B+C initial sample set and the B+A initial sample set output by the initial business model B, and the C+A initial sample set and the C+B initial sample set output by the initial business model C can be obtained.

[0167] It is also possible to input the initial sample set of A + B into the initial business model C to obtain the initial sample set of C + A + B output by the initial business model C, and input the initial sample set of A + C into the initial business model B to obtain the initial sample set of B + A + C output by the initial business model B. The other two initial business models can also output the initial sample set of A + B + C with three training labels.

[0168] Specifically, input the target initial sample set into each initial business model except the target initial business model corresponding to the target initial sample set in turn, and obtain the processing results output by each initial business model, including:

[0169] Determine the number n of initial business models except the target initial business model corresponding to the target initial sample set, where n ≥ 2;

[0170] Input the target initial sample set into the first initial business model to obtain the first target initial sample set output by the first initial business model;

[0171] Input the (i - 1)-th target initial sample set output by the (i - 1)-th initial business model into the i-th initial business model to obtain the i-th target initial sample set output by the i-th initial business model, where 2 ≤ i ≤ n;

[0172] Increment i by 1 and determine whether i is greater than n. If not, continue to perform the operation of inputting the (i - 1)-th target initial sample set output by the (i - 1)-th initial business model into the i-th initial business model. If so, obtain the processing results output by each initial business model.

[0173] Among them, the initial business model can be understood as the business model except the target initial business model. For example, there are four initial business models A, B, C, and D. Determine that the initial business model D is the target initial business model. At this time, the number of initial business models is 3, namely the initial business models A, B, and C.

[0174] In practical applications, in order to expand the training data of the final target business model, sample data with at least two different training labels can be selected, or sample data with all training labels can be selected.

[0175] In a specific embodiment of the present application, following the previous example, taking n = 2 and i = 2 as an example. The initial sample set A is input into the initial service model B, and the B+A initial sample set output by the initial service model B is obtained. The B+A initial sample set is input into the initial service model C, and the C+B+A initial sample set output by the initial service model C is obtained. Through the same steps as above, the A+B+C initial sample set, A+C+B initial sample set output by the initial service model A, the B+A+C initial sample set, B+C+A initial sample set output by the initial service model B can be obtained.

[0176] According to the above data processing method, all initial sample sets can be maximally utilized to provide a large amount of training data for the target service model.

[0177] Step 408: Construct a training data set according to the processing results of each initial service model.

[0178] In practical applications, since multiple initial sample sets containing partial labels and multiple initial sample sets containing all labels are generated, eligible initial sample sets can be selected for training according to the actual situation.

[0179] In a specific embodiment of the present application, a training data set is constructed according to the A+B+C initial sample set, A+B+C initial sample set, B+A+C initial sample set, B+C+A initial sample set, C+A+B initial sample set, C+B+A initial sample set. The target service model is trained based on the training data set.

[0180] In another specific embodiment of the present application, a training data set is constructed according to the A+B initial sample set, A+C initial sample set, B+A initial sample set, B+C initial sample set, C+A initial sample set, C+B initial sample set, A+B+C initial sample set, A+B+C initial sample set, B+A+C initial sample set, B+C+A initial sample set, C+A+B initial sample set, C+B+A initial sample set. The target service model is trained based on the training data set.

[0181] A data processing method provided by the present application includes: obtaining at least three initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with corresponding training labels; training corresponding initial business models according to each initial sample set; processing each initial sample set through each initial business model based on preset rules; and constructing a training data set according to the processing results of each initial business model. In the case of having at least three initial sample sets, by processing each initial sample set through preset rules, multiple initial sample sets with at least two training labels are obtained, expanding the training data set of the target business model, maximizing the utilization of each initial sample set, and solving the problem of difficult model training caused by sparse training data sets.

[0182] Corresponding to the above method embodiment, the present application also provides an embodiment of a data processing device. Figure 5 The structure diagram of a data processing device provided by an embodiment of the present application is shown. As Figure 5 shown, the device includes:

[0183] A first acquisition module 502, configured to acquire first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label;

[0184] A training module 504, configured to train and obtain a first business model according to the first sample data and the first sample label, and train and obtain a second business model according to the second sample data and the second sample label;

[0185] An input module 506, configured to input the first sample data into the second business model, and input the second sample data into the first business model;

[0186] A second acquisition module 508, configured to acquire first target data output by the second business model and second target data output by the first business model, where both the first target data and the second target data are labeled with a first sample label and a second sample label;

[0187] A construction module 510, configured to construct a training data set based on the first target data and the second target data.

[0188] The training module 504 is further configured that the second sample data includes second sample reference data and a second sample reference label corresponding to the second sample reference data, second sample target data, and a second sample target label corresponding to the second sample target data;

[0189] Train a second pre-trained service model based on the second sample reference data and the second sample reference labels;

[0190] Train the second pre-trained service model according to the second sample target data and the second sample target labels to obtain a second service model.

[0191] The device further includes:

[0192] An adjustment module, configured to receive a parameter adjustment instruction;

[0193] In response to the parameter adjustment instruction, adjust the target parameters of the second pre-trained service model

[0194] The adjustment module is further configured to adjust the learning rate parameter of the second pre-trained service model in response to the parameter adjustment instruction to reduce the learning rate of the second pre-trained service model.

[0195] The device further includes:

[0196] An acquisition module, configured to train and obtain a target service model based on the training data set.

[0197] The acquisition module is further configured to input the target data into the target service model;

[0198] Obtain a first prediction label and a second prediction label output by the target service model;

[0199] Calculate a model loss value based on the first prediction label, the first sample label, the second prediction label, and the second sample label;

[0200] Adjust the model parameters of the target service model according to the model loss value, and continue to train the target service model until the model training stop condition is reached.

[0201] The acquisition module is further configured to the model loss value is less than a preset loss value threshold; and / or

[0202] The number of training rounds reaches a preset number of training rounds.

[0203] A data processing device provided by the present application includes a first acquisition module configured to acquire first sample data and second sample data that has a business association relationship with the first sample data. Among them, the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label; a training module configured to train and obtain a first business model according to the first sample data and the first sample label, and train and obtain a second business model according to the second sample data and the second sample label; an input module configured to input the first sample data into the second business model and input the second sample data into the first business model; a second acquisition module configured to acquire first target data output by the second business model and second target data output by the first business model. Among them, both the first target data and the second target data are labeled with the first sample label and the second sample label; a construction module configured to construct a training data set based on the first target data and the second target data. By using multi-stage pre-training and using the first business model and the second business model for annotation, the problems of partial label loss and inconsistent definition between data sets are solved, the training data of the target business model is expanded, and the learning and training effect of the target business model is improved.

[0204] Corresponding to the above method embodiment, the present application also provides an embodiment of a data processing device. Figure 6 The structural schematic diagram of another data processing device provided by an embodiment of the present application is shown. As Figure 6 shown, the device includes:

[0205] An acquisition module 602 configured to acquire at least two initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with a corresponding training label;

[0206] A training module 604 configured to train a corresponding initial business model according to each initial sample set;

[0207] A processing module 606 configured to process each initial sample set through each initial business model based on a preset rule;

[0208] A construction module 608 configured to construct a training data set according to the processing result of each initial business model.

[0209] The training module 604 is further configured to:

[0210] Determine a target initial sample set among the at least three initial sample sets; train a corresponding target initial business model according to the target initial sample set.

[0211] The training module 604 is further configured to:

[0212] Determine a target initial sample set from the at least three initial sample sets;

[0213] Input the target sample data in the target initial sample set into a target initial business model;

[0214] Obtain a target prediction label output by the target initial business model;

[0215] Calculate a model loss value based on the target prediction label and a target training label corresponding to the target sample data;

[0216] Adjust model parameters of the target initial business model according to the model loss value, and continue to train the target initial business model until a model training stop condition is reached.

[0217] The processing module 606 is further configured to:

[0218] Determine a target initial sample set;

[0219] Input the target initial sample set into each initial business model except the target initial business model corresponding to the target initial sample set in sequence, and obtain processing results output by each initial business model.

[0220] The processing module 606 is further configured to:

[0221] Determine the number n of initial business models except the target initial business model corresponding to the target initial sample set, where n≥2;

[0222] Input the target initial sample set into the first initial business model, and obtain a first target initial sample set output by the first initial business model;

[0223] Input the (i - 1)-th target initial sample set output by the (i - 1)-th initial business model into the i-th initial business model, and obtain an i-th target initial sample set output by the i-th initial business model, where 2≤i≤n;

[0224] Increment i by 1, and determine whether i is greater than n. If not, continue to perform the operation of inputting the (i - 1)-th target initial sample set output by the (i - 1)-th initial business model into the i-th initial business model. If so, obtain processing results output by each initial business model.

[0225] A data processing device provided by the present application includes: an acquisition module configured to acquire at least two initial sample sets, where there is a business association relationship between each initial sample set, and the sample data in each initial sample set is labeled with a corresponding training label; a training module configured to train a corresponding initial business model according to each initial sample set; a processing module configured to process each initial sample set through each initial business model based on a preset rule; and a construction module configured to construct a training data set according to the processing results of each initial business model. By processing each initial sample set through the preset rule, multiple initial sample sets with at least two training labels are obtained, expanding the training data set of the target business model, maximizing the utilization of each initial sample set, and solving the problem of difficult model training caused by sparse training data sets.

[0226] The above is a schematic solution of the data processing device in this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.

[0227] Figure 7 FIG. shows a structural block diagram of a computing device 700 according to an embodiment of the present application. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.

[0228] The computing device 700 further includes an access device 740, and the access device 740 enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a worldwide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and so on.

[0229] In an embodiment of the present application, the above components of the computing device 700 and Figure 7 other components not shown in the figure may also be connected to each other, for example, through a bus. It should be understood that Figure 7 the shown structural block diagram of the computing device is only for illustrative purposes and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.

[0230] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smart phones), wearable computing devices (e.g., smart watches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 700 can also be a mobile or stationary server.

[0231] Wherein, when the processor 720 executes the computer instructions, the steps of the data processing method are implemented.

[0232] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above data processing method.

[0233] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the steps of the data processing method described above are implemented.

[0234] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above data processing method.

[0235] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0236] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0237] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0238] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0239] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. According to the content of the present application, many modifications and variations can be made. The present application selects and specifically describes these embodiments to better explain the principle and practical application of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A data processing method, characterized in that, Including: Obtain first sample data and second sample data that has a business association relationship with the first sample data. Among them, the first sample data is labeled with a first sample label, and the second sample data is labeled with a second sample label; Train a first business model based on the first sample data and the first sample label, and train a second business model based on the second sample data and the second sample label; Input the first sample data into the second business model, and input the second sample data into the first business model; Obtain first target data output by the second business model and second target data output by the first business model. Among them, both the first target data and the second target data are labeled with a first sample label and a second sample label; Construct a training data set based on the first target data and the second target data; After constructing the training data set based on the first target data and the second target data, the method further includes: Train a target business model based on the training data set; The training data set includes target data, and the target data is labeled with a first sample label and a second sample label; Training a target business model based on the training data set includes: Input the target data into the target business model; Obtain a first prediction label and a second prediction label output by the target business model; Calculate a model loss value based on the first prediction label, the first sample label, the second prediction label, and the second sample label; Adjust the model parameters of the target business model according to the model loss value, and continue to train the target business model until a model training stop condition is reached.

2. The data processing method according to claim 1, wherein The second sample data includes second sample reference data and a second sample reference label corresponding to the second sample reference data, second sample target data, and a second sample target label corresponding to the second sample target data; Training a second business model based on the second sample data and the second sample label further includes: Train a second pre-trained business model based on the second sample reference data and the second sample reference label; Train the second pre-trained business model according to the second sample target data and the second sample target label to obtain a second business model.

3. The data processing method according to claim 2, wherein Before training the second pre-trained business model according to the second sample target data and the second sample target label, the method further includes: Receive a parameter adjustment instruction; Respond to the parameter adjustment instruction to adjust the target parameters of the second pre-trained business model.

4. The data processing method according to claim 3, wherein Responding to the parameter adjustment instruction to adjust the target parameters of the second pre-trained business model includes: Respond to the parameter adjustment instruction to adjust the learning rate parameter of the second pre-trained business model to reduce the learning rate of the second pre-trained business model.

5. The data processing method according to claim 1, wherein Reaching a model training stop condition includes: The model loss value is less than a preset loss value threshold; and / or The number of training rounds reaches a preset number of training rounds.

6. The data processing method according to any one of claims 1-5, characterized in that, The first sample data includes face images; the first sample labels include orbital key point coordinates; the second sample data includes face images; the second sample labels include pupil key point coordinates.

7. A data processing method, characterized in that, including: Obtain at least three initial sample sets, where there is a business association relationship among each of the initial sample sets, and the sample data in each initial sample set is labeled with corresponding training labels; Train corresponding initial business models according to each initial sample set; Process each initial sample set through each initial business model based on preset rules; Construct a training data set according to the processing results of each initial business model; Processing each initial sample set through each initial business model based on preset rules includes: Determine a target initial sample set, where the target initial sample set is one of the at least three initial sample sets; Sequentially input the target initial sample set into each initial business model except the target initial business model corresponding to the target initial sample set, and obtain the processing results output by each initial business model; Training corresponding initial business models according to each initial sample set includes: Determine a target initial sample set among the at least three initial sample sets; Train the corresponding target initial business model according to the target initial sample set; For any one of the initial sample sets to train the corresponding initial business model includes: Input the target sample data in the target initial sample set into the target initial business model; Obtain the target prediction labels output by the target initial business model; Calculate the model loss value based on the target prediction labels and the target training labels corresponding to the target sample data; Adjust the model parameters of the target initial business model according to the model loss value, and continue to train the target initial business model until the model training stop condition is reached.

8. The data processing method according to claim 7, wherein Sequentially input the target initial sample set into each initial business model except the target initial business model corresponding to the target initial sample set, and obtain the processing results output by each initial business model, including: Determine the number n of initial business models except the target initial business model corresponding to the target initial sample set, where n≥2; Input the target initial sample set into the first initial business model, and obtain the first target initial sample set output by the first initial business model; Input the (i - 1)th target initial sample set output by the (i - 1)th initial business model into the ith initial business model, and obtain the ith target initial sample set output by the ith initial business model, where 2≤i≤n; Increment i by 1, and determine whether i is greater than n. If not, continue to perform the operation of inputting the (i - 1)th target initial sample set output by the (i - 1)th initial business model into the ith initial business model. If so, obtain the processing results output by each initial business model.

9. A data processing device, characterized in that, including: A first acquisition module, configured to acquire first sample data and second sample data having a business association relationship with the first sample data, where the first sample data is labeled with first sample labels, and the second sample data is labeled with second sample labels; A training module, configured to train and obtain a first service model according to the first sample data and the first sample label, and train and obtain a second service model according to the second sample data and the second sample label; An input module, configured to input the first sample data into the second service model, and input the second sample data into the first service model; A second acquisition module, configured to acquire first target data output by the second service model and second target data output by the first service model, wherein both the first target data and the second target data are labeled with the first sample label and the second sample label; A construction module, configured to construct a training data set based on the first target data and the second target data, where the training data set includes target data, and the target data is labeled with the first sample label and the second sample label; An acquisition module, configured to train and obtain a target service model based on the training data set; The acquisition module is further configured to: input the target data into the target service model; obtain a first prediction label and a second prediction label output by the target service model; calculate a model loss value based on the first prediction label, the first sample label, the second prediction label, and the second sample label; adjust model parameters of the target service model according to the model loss value, and continue to train the target service model until a model training stop condition is reached.

10. A data processing device, characterized in that, Including: An acquisition module, configured to acquire at least two initial sample sets, where there is a service association relationship between each pair of initial sample sets, and the sample data in each initial sample set is labeled with a corresponding training label; A training module, configured to train a corresponding initial service model according to each initial sample set; A processing module, configured to process each initial sample set through each initial service model based on a preset rule; A construction module, configured to construct a training data set according to the processing results of each initial service model; The processing module is further configured to determine a target initial sample set, where the target initial sample set is one of the at least three initial sample sets; input the target initial sample set into each initial service model except the target initial service model corresponding to the target initial sample set in sequence, and obtain processing results output by each initial service model; A training module, configured to: determine a target initial sample set among the at least three initial sample sets; train a corresponding target initial service model according to the target initial sample set; The training module is further configured to: determine a target initial sample set among the at least three initial sample sets; input target sample data in the target initial sample set into the target initial service model; obtain a target prediction label output by the target initial service model; calculate a model loss value based on the target prediction label and the target training label corresponding to the target sample data; adjust model parameters of the target initial service model according to the model loss value, and continue to train the target initial service model until a model training stop condition is reached.

11. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the computer instructions, the steps of the method according to any one of claims 1-6 or 7-8 are implemented.

12. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method according to any one of claims 1-6 or 7-8 are implemented.

Citation Information

Patent Citations

  • Semi-supervised learning method and system for text classification

    CN112528030A

  • Method and device for training model

    CN113361621A