Data processing method and device, terminal equipment and computer readable storage medium
By directive labeling and filtering of unlabeled data, high-quality labeled data are generated for training of pre-trained language models, which solves the problem of poor results in downstream tasks and significantly improves the performance of the model.
Patent Information
- Application Number
- CN202311448392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
The pre-trained language model has poor results when executing downstream tasks, mainly due to the lack of labeling of the unlabeled data used during training.
By directive annotation of unlabeled data, labeled data with instruction labels are generated, and these data are filtered, high-quality target annotation data is obtained for model training.
The performance of the model in downstream tasks is improved, and the expressiveness and accuracy of the model are enhanced by training using high-quality labeled data.
Smart Images

Figure CN119939235A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, terminal device and computer-readable storage medium. Background Art
[0002] The current training method of pre-trained language models is to train with a large amount of corpus (usually unlabeled data) to obtain a general language representation model. If the pre-trained model is used to perform some downstream tasks, the model will perform poorly when performing specific downstream tasks because the pre-trained data is unlabeled. Therefore, it is necessary to solve the problem of poor performance of the model when performing downstream tasks. Summary of the invention
[0003] The present application provides a data processing method, which can label unlabeled data used for model training, and then train the model through the labeled data, thereby improving the performance of the model.
[0004] In a first aspect, the present application provides a data processing method, the method comprising:
[0005] Get unlabeled datasets;
[0006] Performing instruction labeling on each data in the unlabeled data set to obtain each labeled data with an instruction label;
[0007] The labeled data are screened to obtain target labeled data for use in training the template model.
[0008] In some embodiments of the present application, the step of performing instruction labeling on each data in the unlabeled data set to obtain each labeled data with an instruction label includes:
[0009] Performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label;
[0010] A second instruction labeling is performed according to each of the initially labeled data to obtain each labeled data having an instruction label.
[0011] In some embodiments of the present application, the step of performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label includes:
[0012] Perform a first instruction labeling on each data in the unlabeled data set, remove the unlabeled data set that does not obtain the initial instruction label, and obtain each initial labeled data with the initial instruction label.
[0013] In some embodiments of the present application, the step of performing second instruction labeling according to each of the initially labeled data to obtain each labeled data having an instruction tag includes:
[0014] Deleting the instruction labels respectively marked on the initially marked data to obtain the data to be marked;
[0015] Perform a second instruction labeling on each of the to-be-labeled data to obtain each labeled data having an instruction label.
[0016] In some embodiments of the present application, the step of performing second instruction labeling according to each of the initially labeled data to obtain each labeled data having an instruction tag includes:
[0017] The second instruction labeling is performed on each of the initially labeled data, and the initially labeled data without the instruction label of the second instruction labeling is removed to obtain each labeled data with the instruction label.
[0018] In some embodiments of the present application, the screening of the labeled data to obtain the target labeled data includes:
[0019] Scoring each of the labeled data to determine a scoring score corresponding to each of the labeled data;
[0020] According to the scoring scores corresponding to the labeled data, the labeled data are screened to obtain the target labeled data.
[0021] In some embodiments of the present application, after screening the labeled data to obtain each target labeled data, the method further includes:
[0022] Determine each of the target labeled data as positive sample data;
[0023] Determine that the unlabeled data in the unlabeled data set excluding the target labeled data is negative sample data;
[0024] A target model is trained according to the positive sample data and the negative sample data.
[0025] In a second aspect, the present application further provides a data processing device, the device comprising:
[0026] Acquisition module, used to obtain unlabeled data sets;
[0027] A labeling module, used for labeling each data in the unlabeled data set with an instruction to obtain each labeled data with an instruction label;
[0028] The screening module is used to screen the labeled data to obtain target labeled data for training the template model.
[0029] In some embodiments of the present application, the annotation module is specifically used to:
[0030] Performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label;
[0031] A second instruction labeling is performed according to each of the initially labeled data to obtain each labeled data having an instruction label.
[0032] In some embodiments of the present application, the annotation module is further used to:
[0033] Perform a first instruction labeling on each data in the unlabeled data set, remove the unlabeled data set that does not obtain the initial instruction label, and obtain each initial labeled data with the initial instruction label.
[0034] In some embodiments of the present application, the annotation module is further used to:
[0035] Deleting the instruction labels respectively marked on the initially marked data to obtain the data to be marked;
[0036] Perform a second instruction labeling on each of the to-be-labeled data to obtain each labeled data having an instruction label.
[0037] In some embodiments of the present application, the annotation module is further used to:
[0038] The second instruction labeling is performed on each of the initially labeled data, and the initially labeled data without the instruction label of the second instruction labeling is removed to obtain each labeled data with the instruction label.
[0039] In some embodiments of the present application, the screening module is specifically used for:
[0040] Scoring each of the labeled data to determine a scoring score corresponding to each of the labeled data;
[0041] According to the scoring scores corresponding to the labeled data, the labeled data are screened to obtain the target labeled data.
[0042] In some embodiments of the present application, the data processing device further includes a determination module, and the determination module is specifically configured to:
[0043] Determine each of the target labeled data as positive sample data;
[0044] Determine that the unlabeled data in the unlabeled data set excluding the target labeled data is negative sample data;
[0045] A target model is trained according to the positive sample data and the negative sample data.
[0046] In a third aspect, the present application also provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in any one of the data processing methods described.
[0047] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps in any one of the data processing methods described.
[0048] The data processing method provided by the present application can obtain high-quality labeled data for the model by labeling the unlabeled data with instruction labels and then screening the labeled data. Based on this, when the model is trained with the high-quality labeled data, the performance of the model can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 is a scenario schematic diagram of a data processing system provided in an embodiment of the present application;
[0051] Figure 2 This is a flow chart of an embodiment of a data processing method in an embodiment of the present application;
[0052] Figure 3 is a functional module schematic diagram of a data processing device in an embodiment of the present application;
[0053] Figure 4 It is a schematic diagram of the structure of the terminal device in the embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0055] In the description of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the feature. In the description of the present application, "plurality" means two or more, unless otherwise clearly and specifically defined.
[0056] In this application, the word "exemplary" is used to mean "used as an example, illustration or description". Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. At the same time, it is to be understood that in the specific implementation of this application, when user information, user data and other related data are involved, when the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data shall comply with relevant laws, regulations and standards of relevant countries and regions.
[0057] In order to enable any person skilled in the art to implement and use the present application, the following description is provided. In the following description, details are listed for the purpose of explanation. It should be understood that those of ordinary skill in the art will recognize that the present application can be implemented without using these specific details. In other examples, known structures and processes will not be elaborated in detail to avoid unnecessary details that make the description of the present application obscure. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest range of principles and features disclosed in the present application.
[0058] The present application provides a data processing method, apparatus, device and storage medium, which are described in detail below.
[0059] See also Figure 1 , Figure 1 Schematic diagram of a data processing system provided in an embodiment of the present application. The data processing system may include a terminal device 100 and a storage device 200. The storage device 200 may transmit data to the terminal device 100. Figure 1 The terminal device 100 in the embodiment can obtain the logs stored in the storage device 200 to execute the data processing method in the present application.
[0060] In the embodiment of the present application, the terminal device 100 includes but is not limited to a desktop computer, a portable computer, a network server, a PDA (Personal Digital Assistant), a tablet computer, a wireless terminal device, an embedded device, etc.
[0061] In an embodiment of the present application, communication between the terminal device 100 and the storage device 200 can be achieved through any communication method, including but not limited to mobile communications based on the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), and Worldwide Interoperability for Microwave Access (WiMAX), or computer network communications based on the TCP / IP Protocol Suite (TCP / IP) and User Datagram Protocol (UDP).
[0062] It should be noted that Figure 1 The scenario diagram of the data processing system shown is merely an example. The data processing system and scenario described in the embodiments of the present application are intended to more clearly illustrate the technical solution of the embodiments of the present application, and do not constitute a limitation on the technical solution provided in the embodiments of the present application. A person of ordinary skill in the art will appreciate that with the evolution of the data processing system and the emergence of new business scenarios, the technical solution provided in the embodiments of the present application is equally applicable to similar technical problems.
[0063] like Figure 2 As shown, Figure 2 This is a flow chart of an embodiment of a data processing method in an embodiment of the present application. The data processing method may include the following steps 201 to 203:
[0064] 201. Obtain an unlabeled dataset.
[0065] The unlabeled data in the unlabeled data set in the embodiment of the present application can be any type of data, such as text data, image data, etc., and the specific embodiment of the present application is not limited. The unlabeled data set is a data set formed by combining various unlabeled data. Among them, the way to obtain the unlabeled data set can be to read the unlabeled data set stored in the memory, or to directly obtain the unlabeled data set sent by the user or other terminal, or to obtain the unlabeled data set from the public Internet, and the specific embodiment of the present application is not limited. It should be noted that in the embodiment of the present application, each unlabeled data in the unlabeled data set can be pre-processed, including deleting duplicate data, filtering the length of short text, and deleting potential low-quality fragments through some heuristic methods (such as html URLs, etc.), and finally forming an unlabeled data set.
[0066] 202. Perform instruction labeling on each data in the unlabeled data set to obtain each labeled data with an instruction label.
[0067] In the embodiment of the present application, the obtained unlabeled data set does not have instruction labels, so in order to ensure the high performance of the subsequent model. Therefore, after obtaining the unlabeled data set, the instruction labels of each unlabeled data in the unlabeled data set can be predicted first, so that each unlabeled data in the unlabeled data set carries a label, and the set of labeled data with instruction labels is regarded as the data set for model training.
[0068] Specifically, in the embodiment of the present application, the open source model 1 can be used to annotate the instruction labels of each unlabeled data in the unlabeled data set. For example, chatglm-6b is used as the open source model 1, and the model is used to annotate the instruction labels of each unlabeled data in the unlabeled data set by constructing a suitable prompt, so that each unlabeled data can have an instruction label. In the embodiment of the present application, the annotation model for the instruction label can pre-set multiple label sets, and each label set has multiple associated labels. For example: when the unlabeled data set is obtained, the annotation model can feature each data set in the unlabeled data set, and perform feature analysis, and determine the specific instruction label according to the analysis results. For example, if the analysis result determines that the current unlabeled data matches the function related to playback, the annotation model can select a specific instruction label from the playback-related label set for annotation, including: "play music", "stop playing", "turn up the volume" and other different instruction labels. Of course, it should be noted that other types of instruction labels are also included in actual situations, and the specific embodiment of the present application is not limited.
[0069] In addition, in an embodiment of the present application, an open source model for label annotation can be trained so that the open source model can perform instruction label annotation. For example, a labeled data set can be obtained, and the labeled data set can be a data set with manually labeled labels and manually labeled scores. The labeled data set includes multiple labeled data, and one labeled data corresponds to one manually labeled label and\or one manually labeled score. After that, the labeled data set can be input into the open source model, and the open source model is used to predict labels for each labeled data in the labeled data set, and determine whether the predicted labels of the labeled data are different from the true labels. If there is a situation where the predicted label is different from the true label, adjust the model parameters of the open source model until the predicted labels of each labeled data obtained by the open source model are the same as the corresponding true labels. Or determine the loss value of each labeled data in the labeled data set, determine whether the loss value of each labeled data is greater than a loss value threshold, and if there is at least one labeled data whose loss value is greater than the loss value threshold, adjust the model parameters of the open source model until the loss value of each labeled data obtained by the open source model is less than the loss value threshold. It should be noted that if the model is adjusted based on the loss value, a specific loss function can be selected according to the actual situation, and the specific embodiment of the present application is not limited. At this point, the training of the open source model can be completed, so that the open source model can effectively label the unlabeled data with instruction labels.
[0070] 203. Screen the labeled data to obtain target labeled data for training the template model.
[0071] In the embodiment of the present application, after obtaining data marked with instruction labels, since there are many types of instruction labels. Therefore, in order to train a specific target model, data screening can be performed for each data with instruction labels. For example, data of a specific label type can be screened to select data of a specific label type. The screened data can help the target model to improve the performance of specific aspects in a targeted manner and avoid interference of the target model with data of other instruction label types.
[0072] The data processing method provided by the present application can obtain high-quality labeled data for the model by labeling the unlabeled data with instruction labels and then screening the labeled data. Based on this, when the model is trained with the high-quality labeled data, the performance of the model can be effectively improved.
[0073] In order to better implement the embodiment of the present application, in one embodiment of the present application, each data in the unlabeled data set is labeled with instructions to obtain each labeled data with instruction labels, including:
[0074] Perform a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label; perform a second instruction labeling on each initial labeled data to obtain each labeled data with an instruction label.
[0075] The above embodiment provides a solution for labeling unlabeled data with instructions through an open source model, so that the unlabeled data has instruction labels. In order to further determine the correctness of the labeling of the instruction label, the embodiment of the present application also provides a double labeling method.
[0076] Specifically, in the embodiment of the present application, the first instruction labeling of each unlabeled data in the unlabeled data set can be performed by labeling each unlabeled data in the unlabeled data set through the first open source model that has been trained to obtain each initial labeled data with an initial instruction label. Among them, the first instruction labeling of each unlabeled data in the unlabeled data set is similar to the instruction labeling described in the above embodiment, and will not be repeated here.
[0077] Afterwards, after obtaining each initial labeled data with an initial instruction label, each initial labeled data with an initial instruction label can be sent to the second open source model for labeling. At this time, the second open source model can re-label the data based on each initial labeled data, thereby forming a second labeled instruction label for each initial labeled data. At this time, each initial labeled data corresponds to two instruction labels. Next, it can be determined whether the two instruction labels corresponding to each initial labeled data are the same, and then the initial labeled data with the same two instruction labels can be selected as the actual labeled data. By comparing whether the two labels are the same, the reliability of the model's data labeling can be verified.
[0078] It should be noted that the second open source model in the embodiment of the present application can be the same as the first open source model. If the second open source model is the same as the first open source model, the training method of the second open source model is different from that of the first open source model. For example: the first open source model can be trained by comparing whether the predicted label is the same as the true label; the second open source model can be trained by adjusting the loss value, which is similar to the scheme described in the above embodiment and will not be described in detail here. If the second open source model is different from the first open source model, the training method of the first open source model and the second open source model is not limited.
[0079] In order to better implement the embodiment of the present application, in one embodiment of the present application, each data in the unlabeled data set is labeled with a first instruction to obtain each initial labeled data with an initial instruction label, including:
[0080] The first instruction labeling is performed on each data in the unlabeled data set, and the unlabeled data set that has not obtained the initial instruction label is removed to obtain each initial labeled data with the initial instruction label.
[0081] In the above embodiment, a solution is provided for labeling each data in the unlabeled data set according to the model. However, in actual situations, there may be a situation where some unlabeled data cannot obtain instruction labels. In order to prevent the data that does not obtain instruction labels from affecting the subsequent model training, it is necessary to remove the data that does not obtain instruction labels.
[0082] Specifically, the removal method may include detecting each data after being labeled by the first instruction, determining the initial labeled data with the label and the unlabeled data without the label, and then deleting the unlabeled data without the label.
[0083] In order to better implement the embodiment of the present application, in one embodiment of the present application, a second instruction labeling is performed according to each initially labeled data to obtain each labeled data with an instruction label, including:
[0084] The instruction labels of each initially labeled data are deleted to obtain the data to be labeled; and the second instruction labeling is performed on each data to be labeled to obtain each labeled data with the instruction label.
[0085] The above embodiment provides an implementation scheme for continuing to perform a second instruction labeling on the basis of the initial labeled data with the first instruction label. However, when the unlabeled data acquires the first instruction label, if the second labeling is performed on the initial labeled data with the first instruction label, it may interfere with the labeling of the second labeling model, so the label of the first label needs to be deleted. The initial labeled data after the first instruction label is deleted is changed to the data to be labeled. Next, the second instruction labeling can be performed on the data to be labeled. Among them, the method of performing the second instruction labeling is the same as that in the above embodiment, and will not be repeated here.
[0086] It should be noted that, if the data to be labeled is labeled with the second instruction, since the first instruction label has been deleted, label screening can be performed according to the second instruction label after the second instruction label.
[0087] In order to better implement the embodiment of the present application, in one embodiment of the present application, a second instruction labeling is performed according to each initially labeled data to obtain each labeled data with an instruction label, including:
[0088] The second instruction labeling is performed on each initially labeled data, and the initially labeled data without the instruction label of the second instruction labeling is removed to obtain each labeled data with the instruction label.
[0089] According to the above embodiment, in actual situations, after the first instruction label is labeled, there may be data that does not get the first instruction label. Similarly, when the second instruction label is performed, there may also be data that does not get the second instruction label. Similarly, in order to avoid the data that does not get the second instruction label affecting the subsequent model training, it is necessary to remove the data that does not get the second instruction label, so as to obtain the actual labeled data.
[0090] In order to better implement the embodiment of the present application, in one embodiment of the present application, each labeled data is screened to obtain each target labeled data, including:
[0091] Scoring each labeled data to determine the scoring score corresponding to each labeled data; screening each labeled data according to the scoring score corresponding to each labeled data to obtain each target labeled data.
[0092] In the above embodiment, a scheme is provided for filtering each labeled data by label type to obtain target labeled data. In order to improve the filtering effect, the embodiment of the present application also provides a scheme for filtering the labeled data by model scoring.
[0093] Specifically, in the embodiment of the present application, when scoring each labeled data, any scoring model can be used for scoring. For example: there are different types of reference data in the scoring model. After each labeled data is input into the scoring model, the scoring model can calculate the feature similarity between each labeled data and different reference data, so as to score according to the feature similarity, and then score each labeled data according to the degree of similarity. For example: the score of labeled data with a similarity of 0%-20% can be 1 point; the score of labeled data with a similarity of 21%-40% can be 2 points; the score of labeled data with a similarity of 41%-60% can be 3 points; the score of labeled data with a similarity of 61%-80% can be 4 points; the score of labeled data with a similarity of 81%-100% can be 5 points. In this way, the scoring results of each labeled data can be determined.
[0094] Alternatively, multiple types of reference features can be set in the scoring model. After each labeled data is input into the scoring model, one of the labeled data is used for explanation. The scoring model can extract the data features in the labeled data, and then perform feature decomposition on the extracted data features to obtain each sub-feature. After that, the feature type of each sub-feature is determined, and the feature type of the sub-feature is compared with the feature type of each reference feature. The more the same feature type is, the higher the score of the current labeled data is. For example: if there is one same feature type, it is 1 point; if there are two same feature types, it is 2 points, and so on, which will not be repeated here.
[0095] Among them, the reference data or reference features involved in the above description are used to determine the quality of each labeled data. The higher the score, the closer the labeled data is to the reference feature or reference data, and the higher the quality. Therefore, in actual situations, specific reference data or reference features can be set according to actual conditions. In addition, the scoring methods listed in the above description can be adjusted according to actual conditions, and the specific embodiments of the present application are not limited. Therefore, after obtaining the scores of each labeled data, the labeled data with higher scores in the scoring results can be selected as target labeled data. For example: Continuing the above-described scheme for explanation, a score threshold of 3 can be set. If the score in the scoring result exceeds the score threshold 3, the labeled data exceeding the score threshold 3 is determined as the target labeled data. It should be noted that the score threshold can be set according to actual conditions, and the specific embodiments of the present application are not limited. After the target labeled data is obtained, the filtered target labeled data can be used for specific model training.
[0096] In order to better implement the embodiment of the present application, in one embodiment of the present application, after screening each labeled data to obtain each target labeled data, the method further includes:
[0097] Determine each target labeled data as positive sample data; determine the unlabeled data in the unlabeled data set excluding the target labeled data as negative sample data; and train the target model based on the positive sample data and the negative sample data.
[0098] In the above embodiment, a scheme for model training is provided for the target labeled data obtained after screening. However, in actual situations, there may be a situation where there are fewer target labeled data. Therefore, after the target labeled data is obtained, the number of target labeled data can be determined. If the number of target labeled data is less than the preset number threshold, the non-target labeled data that was previously filtered out can be used as training samples of the model at the same time, that is, the data different from the target labeled data in the unlabeled data set is input into the model training. At this time, the target labeled data can be used as positive sample data, and the unlabeled data in the unlabeled data set except the target labeled data can be used as negative sample data. Specifically, during the training process, the model can focus on learning similar features in the same label based on the positive sample data, so that input samples with similar features can be better identified in the subsequent recognition process. And the model can focus on learning different features in different labels based on the negative sample data, so that input samples with different features can be better identified in the subsequent recognition process, thereby further improving the performance of the target model.
[0099] In summary, the process of the embodiment of the present application may include:
[0100] (1) is to use the open source model 1 to generate instructions for the data in the unlabeled data A. For example, chatglm-6b is used as the open source model 1, and the model is used to generate instructions for the unlabeled data A by constructing a suitable prompt. Not all data can generate suitable instructions. The present invention allows the model to generate None instructions for data that cannot generate instructions in the prompt. When the instruction is generated as None, it means that the data cannot generate suitable instructions. Through the instruction generation of the open source model 1, we can generate instruction alignment data A1 with instructions from the unlabeled data A, A1∈A.
[0101] (2) Since the capability of the model of the generated data A1 is relatively weak, some inappropriate instructions will be generated in the instruction generation. If these instructions are added to the training process of the pre-trained language model, it may have a negative impact. Therefore, based on A1, the generated instructions are deleted, and a model with stronger capabilities and larger parameters is used to generate instructions for this part of the data, such as the LLAMA2-13B model. Similarly, this model is used to generate instructions for data A1 by constructing a suitable prompt. Not all data can generate suitable instructions. The present invention allows the model to generate None instructions for data that cannot generate instructions in the prompt. When the instruction is generated as None, it means that the data cannot generate suitable instructions. Through the instruction generation of the open source model 2, we can generate instruction alignment data A2 with instructions in the data A1, A2∈A1∈A.
[0102] (3) Select high-quality training data pairs as training data. Not all training data pairs (instruction-output pairs) are of high quality. If all of them are used for training, it may not be beneficial. Therefore, it is necessary to filter the quality of data A2. The present invention still constructs a scoring prompt, and scores the data of A2 (1-5 points) through the open source language model 3. The data with a score higher than 3 points is retained, and the data with a score lower than 3 points is filtered. At this time, the open source language model 3 can select the model used in the previous step or reselect a general language model with scoring capabilities. Finally, high-quality instruction alignment data A3 is obtained, A3∈A2∈A1∈A.
[0103] (4) Use the unlabeled data A0 (A-A3) and the high-quality instruction-aligned data A3 as positive sample data for the training of the pre-trained language model. Through the above steps, we can obtain the high-quality instruction-aligned data A3 with instructions from the data A. The amount of A3 data is relatively small compared to the original unlabeled data A. Therefore, we still need to add the unlabeled data as negative sample data to the training of the pre-trained language model, but the data added to the training at this time must exclude the instruction data A3 that already has high-quality instructions, that is, A0 = A1-A3. At the same time, the high-quality instruction-aligned data A3 also needs to be added to the training of the pre-trained language model.
[0104] In order to better implement the data processing method in the embodiment of the present application, in addition to the data processing method, the embodiment of the present application also provides a data processing device, such as Figure 3 As shown, the device 300 includes:
[0105] An acquisition module 301 is used to acquire an unlabeled data set;
[0106] The labeling module 302 is used to label each data in the unlabeled data set with an instruction to obtain each labeled data with an instruction label;
[0107] The screening module 303 is used to screen the labeled data to obtain target labeled data for use in training the template model.
[0108] The data processing device provided by the present application can first obtain an unlabeled data set through the acquisition module 301, then label the unlabeled data with instruction labels through the labeling module 302, and finally filter the labeled data through the screening module 303, thereby obtaining high-quality labeled data for the model. Based on this, when the model is trained with the high-quality labeled data, the performance of the model can be effectively improved.
[0109] In some embodiments of the present application, the marking module 302 is specifically used to:
[0110] Performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label;
[0111] A second instruction labeling is performed according to each initially labeled data to obtain each labeled data with an instruction label.
[0112] In some embodiments of the present application, the marking module 302 is further used to:
[0113] The first instruction labeling is performed on each data in the unlabeled data set, and the unlabeled data set that has not obtained the initial instruction label is removed to obtain each initial labeled data with the initial instruction label.
[0114] In some embodiments of the present application, the marking module 302 is further used to:
[0115] Delete the instruction labels of each initially labeled data to obtain the data to be labeled;
[0116] Perform the second instruction labeling on each piece of data to be labeled to obtain each piece of labeled data with an instruction label.
[0117] In some embodiments of the present application, the marking module 302 is further used to:
[0118] The second instruction labeling is performed on each initially labeled data, and the initially labeled data without the instruction label of the second instruction labeling is removed to obtain each labeled data with the instruction label.
[0119] In some embodiments of the present application, the screening module 303 is specifically used to:
[0120] Scoring each labeled data and determining the scoring score corresponding to each labeled data;
[0121] According to the scoring scores corresponding to each labeled data, each labeled data is screened to obtain each target labeled data.
[0122] In some embodiments of the present application, the device further includes a determination module, which is specifically configured to:
[0123] Determine each target labeled data as positive sample data;
[0124] Determine that the unlabeled data excluding the target labeled data in the unlabeled data set is negative sample data;
[0125] The target model is trained based on the positive sample data and the negative sample data.
[0126] The present application also provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of any data processing method in the present application. The terminal device integrates any data processing method provided in the present application. Figure 4 As shown, it shows a schematic diagram of the structure of the terminal device involved in the embodiment of the present application, specifically:
[0127] The terminal device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will appreciate that Figure 4 The terminal device structure shown in the figure does not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0128] The processor 401 is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, the processor 401 executes various functions of the terminal device and processes data, thereby monitoring the terminal device as a whole. Optionally, the processor 401 may include one or more processing cores; the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, the user interface and the application program, etc., and the modem processor mainly processes wireless communication. It is understandable that the above-mentioned modem processor may not be integrated into the processor 401.
[0129] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0130] The terminal device also includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to manage charging, discharging, power consumption management and other functions through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0131] The terminal device may further include an input unit 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0132] Although not shown, the terminal device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 401 in the terminal device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, thereby realizing various functions, such as:
[0133] Get unlabeled datasets;
[0134] Perform instruction labeling on each data in the unlabeled data set to obtain each labeled data with instruction labels;
[0135] The labeled data are screened to obtain the target labeled data for training the template model.
[0136] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0137] To this end, an embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any data processing method provided in the embodiment of the present application. For example, the computer program is loaded by a processor to execute the following steps:
[0138] Get unlabeled datasets;
[0139] Perform instruction labeling on each data in the unlabeled data set to obtain each labeled data with instruction labels;
[0140] The labeled data are screened to obtain the target labeled data for training the template model.
[0141] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.
[0142] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.
[0143] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0144] The above is a detailed introduction to a data processing method and device provided in an embodiment of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Get unlabeled datasets; Performing instruction labeling on each data in the unlabeled data set to obtain each labeled data with an instruction label; The labeled data are screened to obtain target labeled data for use in training the template model.
2. The data processing method according to claim 1, characterized in that: The step of performing instruction labeling on each data in the unlabeled data set to obtain each labeled data with an instruction label includes: Performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label; A second instruction labeling is performed according to each of the initially labeled data to obtain each labeled data having an instruction label.
3. The data processing method according to claim 2, characterized in that: The step of performing a first instruction labeling on each data in the unlabeled data set to obtain each initial labeled data with an initial instruction label includes: Perform a first instruction labeling on each data in the unlabeled data set, remove the unlabeled data set that does not obtain the initial instruction label, and obtain each initial labeled data with the initial instruction label.
4. The data processing method according to claim 3, characterized in that: The step of performing second instruction labeling according to each of the initially labeled data to obtain each labeled data having an instruction label includes: Deleting the instruction labels respectively marked on the initially marked data to obtain the data to be marked; Perform a second instruction labeling on each of the to-be-labeled data to obtain each labeled data having an instruction label.
5. The data processing method according to claim 2, characterized in that: The step of performing second instruction labeling according to each of the initially labeled data to obtain each labeled data having an instruction label includes: The second instruction labeling is performed on each of the initially labeled data, and the initially labeled data without the instruction label of the second instruction labeling is removed to obtain each labeled data with the instruction label.
6. The data processing method according to claim 1, characterized in that: The step of screening the labeled data to obtain the target labeled data includes: Scoring each of the labeled data to determine a scoring score corresponding to each of the labeled data; According to the scoring scores corresponding to the labeled data, the labeled data are screened to obtain the target labeled data.
7. The data processing method according to claim 1, characterized in that: After screening the labeled data to obtain the target labeled data, the method further includes: Determine each of the target labeled data as positive sample data; Determine that the unlabeled data in the unlabeled data set excluding the target labeled data is negative sample data; A target model is trained according to the positive sample data and the negative sample data.
8. A data processing device, characterized in that: The device comprises: Acquisition module, used to obtain unlabeled data sets; A labeling module, used for labeling each data in the unlabeled data set with an instruction to obtain each labeled data with an instruction label; The screening module is used to screen the labeled data to obtain target labeled data for training the template model.
9. A terminal device, characterized in that: The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps in the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in the data processing method according to any one of claims 1 to 7.