Meta-learning Data Augmentation Framework
The data augmentation system addresses the challenge of limited training data in natural language processing by generating diverse data using token manipulation and semi-supervised learning, enhancing model performance and reducing training costs.
Patent Information
- Application Number
- JP2021096983
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-30
- Filing Date
- 2021-06-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-06-10
AI Technical Summary
Existing natural language processing systems face challenges in understanding opinions expressed in input text with limited training data, requiring large labeled datasets and static operators that generate non-diverse additional data, often inappropriate for the task.
A data augmentation system generates diverse training data using a set of data augmentation operators, including token deletion, insertion, replacement, and swap, and inverse operators, to enhance machine learning models for tasks like entity matching and error detection, utilizing unlabeled data and semi-supervised learning techniques.
The system effectively expands training data, improving natural language processing tasks by generating diverse and relevant data without extensive labeled datasets, enhancing performance and reducing time and cost in model training.
Smart Images

Figure 0007697824000001 
Figure 0007697824000002 
Figure 0007697824000003
Abstract
Description
Background Art
[0001] Background
[0001] It is a difficult task to realize a natural language processing system that enables a computer to respond to natural language input. The task becomes even more difficult when a machine attempts to understand the opinions expressed in the input text and extract classification information based on limited training data. There is a need for techniques and systems that can respond to the requirements of the latest natural language systems in a time- and cost-efficient manner.
Summary of the Invention
Means for Solving the Problems
[0002] Summary
[0002] Certain embodiments of the present disclosure relate to a non-transitory computer-readable storage medium storing instructions executable by a data augmentation system including one or more processors to cause the data augmentation system to perform a method for generating training data for a machine learning model. The method can include accessing a machine learning model from a machine learning model repository, identifying a data set associated with the machine learning model, generating a set of data augmentation operators using the data set, selecting a token sequence associated with the machine learning model, generating at least one token sequence by applying at least one data augmentation operator of the set of augmentation operators to the selected token sequence, selecting a subset of the token sequences from the generated at least one token sequence, storing the subset of the token sequences in a training data repository, and providing the subset of the token sequences to the machine learning model.
[0003]
[0003] According to some embodiments of the present disclosure, generating a set of data augmentation operators using a dataset further includes selecting one or more data augmentation operators, generating a sequentially formatted input token sequence of the identified dataset, applying one or more data augmentation operators to at least one token sequence among the sequentially formatted input token sequences to generate at least one modified token sequence, and determining a set of augmentation operators for restoring the at least one modified token sequence to the corresponding sequentially formatted input token sequence.
[0004]
[0004] According to some embodiments of the present disclosure, the accessed machine learning model is a sequence-to-sequence machine learning model.
[0005]
[0005] According to some embodiments of the present disclosure, selecting a subset of token sequences further includes filtering at least one token sequence from the generated at least one token sequence using a filtering machine learning model, determining the weight of at least one sequence token within the at least one filtered token sequence using a weighting machine learning model, and applying the weight to at least one token sequence among the at least one filtered token sequences.
[0006]
[0006] The weight of at least one token sequence is determined based on the importance of the token sequence when training the machine learning model.
[0007]
[0007] According to some embodiments of the present disclosure, the importance of at least one token sequence is determined by calculating the validation loss of the machine learning model when trained using the at least one token sequence.
[0008]
[0008] According to some embodiments of the present disclosure, the filtering machine learning model is trained using the validation loss.
[0009]
[0009] According to some embodiments of the present disclosure, the weighted machine learning model is trained until the validation loss reaches a threshold.
[0010]
[0010] According to some embodiments of the present disclosure, at least one data augmentation operator includes at least one of a token deletion operator, a token insertion operator, a token replacement operator, a token swap operator, a span deletion operator, a span shuffle operator, a column shuffle operator, a column deletion operator, an entity swap operator, an inverse translation operator, a class generator operator, and an inverse data augmentation operator.
[0011]
[0011] According to some embodiments of the present disclosure, the inverse data augmentation operator is a combination of a plurality of data augmentation operators.
[0012]
[0012] According to some embodiments of the present disclosure, at least one data augmentation operator is context-dependent.
[0013]
[0013] According to some embodiments of the present disclosure, providing a subset of the token sequence as input to the machine learning model further includes accessing unlabeled data from an unlabeled data repository, generating an extended unlabeled token sequence of the accessed unlabeled data, determining a soft label of the extended unlabeled token sequence, and providing the extended unlabeled token sequence together with the associated soft label as input to the machine learning model.
[0014]
[0014] Certain embodiments of the present disclosure are executable by a data augmentation system including one or more processors to cause the data augmentation system to perform a method for generating a data augmentation operator for generating an augmented token sequence. The method includes accessing unlabeled data from an unlabeled data repository, preparing one or more token sequences of the accessed unlabeled data, transforming the one or more token sequences to generate at least one corrupted sequence, providing the one or more token sequences and the at least one corrupted sequence as inputs to a sequence-to-sequence model of the data augmentation system, executing the sequence-to-sequence model to determine at least one operation required to revert the at least one corrupted sequence to a sequence within the one or more token sequences used to generate the at least one corrupted sequence, and generating an inverse data augmentation operator based on the determined one or more operations for reverting the at least one corrupted sequence.
[0015]
[0015] According to some embodiments of the present disclosure, preparing one or more token sequences of the accessed unlabeled data comprises converting each row in a database table into a token sequence, the token sequence including indicators for the start and end of column values.
[0016]
[0016] According to some embodiments of the present disclosure, transforming the one or more token sequences to generate at least one corrupted sequence further comprises selecting a token sequence from the one or more token sequences, selecting a data augmentation operator from a set of data augmentation operators, and applying the data augmentation operator to the selected token sequence.
[0017]
[0017] According to some embodiments of the present disclosure, generating at least one corrupted sequence further includes generating a collective corrupted sequence by sequentially applying multiple data augmentation operators to the selected token sequence.
[0018]
[0018] Certain embodiments of the present disclosure relate to a non-transitory computer-readable storage medium having stored thereon instructions executable by a data augmentation system including one or more processors to cause the data augmentation system to perform a method for extracting classification information from input data. The method may include adding a task-specific layer to a machine learning model to generate a modified network, initializing the machine learning model and the modified network of the added task-specific layer, selecting input data entries, where the selecting includes serializing the input data entries, providing the serialized input data entries to the modified network, and extracting the classification information using the task-specific layer of the modified network.
[0019]
[0019] According to some embodiments of the present disclosure, the machine learning model further includes generating augmented data using at least one inverse data augmentation operator and pre-training the machine learning model using the augmented data.
[0020]
[0020] According to some embodiments of the present disclosure, serializing the input data entry further includes identifying class tokens and other tokens within the input data entry and marking the class tokens and other tokens with different markers representing the start and end of each token.
[0021] BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles of the present disclosure. [Brief description of the drawings]
[0022]
Figure 1
[0022] Block diagram showing an exemplary data augmentation system according to an embodiment of the present disclosure.
Figure 2A
[0023] Block diagram showing an exemplary data augmentation operator generator according to an embodiment of the present disclosure.
Figure 2B
[0024] Exemplary tabular form and sequential representation of data for an entity matching classification task according to an embodiment of the present disclosure are shown.
Figure 2C
[0025] Exemplary tabular form and sequential representation of data for an error detection classification task according to an embodiment of the present disclosure are shown.
Figure 3
[0026] Block diagram showing an exemplary meta - learning policy framework according to an embodiment of the present disclosure.
Figure 4A
[0027] Exemplary data flow diagram of the meta - learning data augmentation system of FIG. 1 used for a sequence classification task according to an embodiment of the present disclosure is shown.
Figure 4B
[0028] Exemplary backpropagation techniques for fine - tuning a machine learning model and an exemplary policy for managing training data for a machine learning model according to an embodiment of the present disclosure are shown.
Figure 5
[0029] Block diagram of an exemplary computing device according to an embodiment of the present disclosure.
Figure 6
[0030] Exemplary language machine learning model for a sequence classification task according to an embodiment of the present disclosure is shown.
Figure 7
[0031] Exemplary sequence list for entity matching as a sequence classification task according to an embodiment of the present disclosure is shown.
Figure 8
[0032] Exemplary sequence list for error detection as a sequence classification task according to an embodiment of the present disclosure is shown.
Figure 9
[0033] Exemplary sequence list text classification as a sequence classification task according to an embodiment of the present disclosure is shown.
Figure 10
[0034] It is a diagram showing an exemplary language machine learning model having a specific layer according to an embodiment of the present disclosure.
Figure 11
[0035] It is a flowchart showing an exemplary method for data augmentation operations according to an embodiment of the present disclosure.
Figure 12
[0036] It is a flowchart showing an exemplary method for data augmentation operations according to an embodiment of the present disclosure.
Figure 13
[0037] It is a flowchart showing an exemplary method for generating an inverse data augmentation operator according to an embodiment of the present disclosure.
Figure 14
[0038] It is a flowchart showing an exemplary method for generating training data for a machine learning model according to an embodiment of the present disclosure.
Figure 15
[0039] It is a flowchart showing an exemplary method for a sequence classification task for extracting classification information from an input sequence according to an embodiment of the present disclosure.
Mode for Carrying Out the Invention
[0023] Detailed Description
[0040] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the exemplary embodiments of the present disclosure. It will be understood by those skilled in the art that the principles of the exemplary embodiments can be practiced without all of the specific details. The disclosed embodiments are illustrative and are not intended to disclose all possible embodiments in accordance with the claims and the disclosure. Well-known methods, procedures, and components have not been described in detail so as not to obscure the principles of the exemplary embodiments. Unless explicitly stated otherwise, the exemplary methods and processes described herein are not limited to a particular order or sequence, nor to a particular system configuration. Additionally, some of the described embodiments or some of their elements can occur, or be performed, simultaneously, at the same time, or in parallel.
[0024]
[0041] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations except in cases where execution is not possible. For example, when it is stated that a component can include A or B, then, unless otherwise specifically stated or execution is not possible, the component can include A, or B, or A and B. As a second example, when it is stated that a component can include A, B, or C, then, unless otherwise specifically stated or execution is not possible, the component can include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0025]
[0042] Next, reference is made in detail to the embodiments of the present disclosure illustrated in the accompanying drawings. Unless explicitly stated otherwise, sending and receiving, as used herein, are understood to have a broad meaning that includes sending or receiving in response to a particular request or without such a particular request. Therefore, these terms encompass both the active and passive forms of sending and receiving.
[0026]
[0043] The embodiments described herein provide techniques and methodologies for data integration, data cleaning, and text classification for extracting classification information based on limited training data using natural language techniques by a computing system.
[0027]
[0044] The described embodiments provide distinct advantages over existing techniques for natural language processing. Existing natural language processing systems may require a large set of labeled training data to operate in an unsupervised manner. Alternatively, existing systems may apply a static set of operators to limited labeled data and generate non-diverse additional data from the limited labeled data. Further, the labels of the limited labeled data may not be appropriate for the generated additional data. Therefore, there is a need to generate operators that can generate additional data with minimal supervision.
[0028]
[0045] The embodiments disclosed herein can help perform various natural language processing tasks in a semi-supervised manner by generating data augmentation operators and using the generated data augmentation operators to generate diverse training data. The described embodiments can help natural language processing tasks such as entity matching, error detection, data cleaning, and text classification, to name a few. The embodiments of the present disclosure sequentially transform data into a representation for using the same operator to generate data for training for different natural language processing tasks listed above. This can provide significant advantages in natural language processing systems that may need to say the same thing but in different ways, respond to different individuals or questions. By enabling semi-supervised, data augmentation operators, and sequential data generation, the embodiments disclosed herein can improve the ability to use natural language processing in various industries and specific contexts without requiring a time-consuming and costly pre-training process.
[0029]
[0046] FIG. 1 is a block diagram showing an exemplary data augmentation system 100 according to an embodiment of the present disclosure. The data augmentation system 100 can increase the data available for natural language processing tasks by generating additional data and augmenting the available data. The data augmentation system 100 can achieve the augmentation of additional data by updating a portion of the available data of the sentence. The data augmentation system 100 can use data augmentation operators that can help perform the update of the available data. The data augmentation system 100 can generate augmented data using a static set of data augmentation operators that can add, delete, replace, and swap words within a sentence in the available data. The data augmentation system 100 can also have the ability to generate new data augmentation operators that perform more complex operations on the available data to generate augmented data.
[0030]
[0047] The data augmentation system 100 can include a data augmentation operator (DAO) generator 110 for generating augmentation operators that can generate data used to train a machine learning model. The DAO generator 110 can be a set of software functions or an entire software program (s) applied to a sequence (e.g., a text sentence) to generate a transformation sequence. A processor (e.g., the CPU 520 of FIG. 5 described later) can execute software functions and programs representing one or more components of the data augmentation system 100, including the DAO generator 110. The processor can be a virtual or physical processor of a computing device. The computing device that executes the software function or program can include a single processor or core or multiple processors or cores, or can be multiple computing devices spread across a distributed computing environment, network, cloud, or virtualized computing environment.
[0031]
[0048] The transformation sequence can include the original sequence along with updates to one or more words of the original sequence. The updates can include adding, deleting, replacing, or swapping words or phrases within the original sequence. The transformation sequence can help to augment the original sequence as training data for a machine learning model. The software function of the DAO generator 110 can include a data augmentation operator used to generate a transformation sequence for training a machine learning model. The DAO generator 110 can generate a data augmentation operator that can be stored for later use in a data augmentation operator (DAO) repository 120.
[0032]
[0049] The DAO repository 120 can organize data augmentation operators by a machine learning model intended to use the augmented sequences generated by the data augmentation operators. For example, the DAO repository 120 can include a separate database for managing each set of data augmentation operators that generate a training set for training a machine learning model. In some embodiments, the data augmentation operator can have a reference to all machine learning models for which training data of the augmented sequence was generated using the data augmentation operator. In some embodiments, the DAO generator 110 can obtain an existing data augmentation operator in the DAO repository 120 as an input for generating a new data augmentation operator. The DAO repository 120 can store relationship information between a previous data augmentation operator and a new data augmentation operator generated using the previous data augmentation operator as an input. In some embodiments, the DAO repository 120 can include the difference between the new and previous data augmentation operators. The DAO repository 120 can function as a version management that manages multiple versions of the data augmentation operators generated and updated by the DAO generator 110.
[0033]
[0050] As shown in FIG. 1, the data augmentation system 100 can also include a data repository, i.e., an unlabeled data repository 130 and a text corpus repository 160, which can be used to train a machine learning model. The unlabeled data repository 130 can also be provided as an input to the DAO generator 110 to generate data augmentation operators stored in the DAO repository 120.
[0034]
[0051] The unlabeled data repository 130 can include unannotated data (e.g., data that has not been labeled or annotated by a human or other process). The unlabeled data repository 130 can be an RDBMS, an NRDBMS, or other types of data stores. The unlabeled data repository 130 can provide a large amount of data that has not been annotated by a human or other process and is difficult to use for supervised learning of a natural language processing system for training. The data augmentation system 100 can use the data augmentation operator of the DAO repository 120 to generate additional unlabeled data for training a machine learning model (e.g., the target model 141). The data augmentation system 100 can encode the unlabeled data in the unlabeled data repository 130 and infer labels using a MixMatch method adjusted for natural language processing. The MixMatch method can infer low-entropy labels to be assigned to the unlabeled data generated using the data augmentation operator of the DAO repository 120. The MixMatch method can receive feedback regarding the inferred labels to improve the inference in future iterations. The MixMatch method can improve its label inference ability based on the use of the unlabeled data with the inferred labels in downstream language processing tasks. The downstream language processing tasks can evaluate the unlabeled data in training and use by the machine learning model and provide feedback regarding the inferred labels to the MixMatch method. The data augmentation system 100 can connect the unlabeled data in the unlabeled data repository 130 with the inferred labels to additional data generated to meet the training data requirements of the target model 141. A detailed description of using the unlabeled data repository 130 to generate additional data is presented below in connection with FIG. 4A and its corresponding description.
[0035]
[0052] The data augmentation system 100 can include a machine learning (ML) model repository 140 that provides a machine learning model for generating data augmentation operators used to generate training data for other machine learning models within the ML model repository 140 as a result. The ML model repository 140 can also include ML models, such as a target model 141, that can be trained using additional training data to extract classification information. The data augmentation system 100 can use the data stored in both the text corpus repository 160 and the unlabeled data repository 130 as inputs to train the target model 141 to improve the extraction of classification information from input sentences.
[0036]
[0053] The data augmentation system 100 can use the data augmentation operators in the DAO repository 120 to generate additional data for training the target model 141 to improve the extraction of classification information from input sentences. The target model 141 is a machine learning model and can include multiple layers. The multiple layers can include fully connected layers or partially connected layers. The target model 141 can transform the data in the text corpus repository 160 and the unlabeled data repository 130 before other layers of the target model 141 use the data. The target model 141 can be a language model that can use an embedding layer 611 (as described in FIG. 6 below) to transform the data in the text corpus repository 160 and the unlabeled data repository 130. In some embodiments, the target model 141 can be pre-trained. The transformation of the data in the text corpus repository 160 and the unlabeled data repository 130 is presented below in relation to FIG. 6 and its corresponding description.
[0037]
[0054] The ML model repository 140 can provide a target model 141 that can assist in extracting classification information of the input sentence. The target model 141 can include an encoding layer (e.g., the embedding layer 611 and the encoding layer 612 in FIG. 6) for converting data from the text corpus repository 160 and the unlabeled data repository 130. The target model 141 can be a modified neural network architecture such as, for example, BERT, ELMO, etc. The conversion of data using the target model 141 is presented in FIGS. 7-9 and their corresponding descriptions below. The classification layer of the target model 141 used to extract classification information is presented in relation to FIG. 10 and its corresponding description below.
[0038]
[0055] In some embodiments, the ML model repository 140 can provide a machine learning model as an input to the DAO generator 110 to generate a data augmentation operator for generating additional training data. The ML model repository 140 can provide a sequence-to-sequence model as an input to the DAO generator 110 to generate a data augmentation operator. In some embodiments, the sequence-to-sequence model can be a standard sequence-to-sequence model such as the T5 model from Google. A detailed description of the sequence-to-sequence model used to generate the data augmentation operator is presented in relation to FIGS. 2A-B and their descriptions below.
[0039]
[0056] FIG. 2A is a block diagram showing an exemplary data augmentation operator (DAO) generator 110 (as shown in FIG. 1) according to an embodiment of the present disclosure. The DAO generator 110 can operate independently of the rest of the components of the data augmentation system 100. The DAO generator 110 can generate a data augmentation operator regardless of the request to generate additional training data using the generated data augmentation operator. The DAO generator 110 can also directly receive a request to generate a data augmentation operator.
[0040]
[0057] As shown in FIG. 2A, the DAO generator 110 can include a sequence-to-sequence model 220 for generating a new data augmentation operator for generating training data. The training data generated using the new data augmentation operator can be used to train a machine learning model used for natural language processing tasks. Examples of natural language processing tasks are presented below in connection with FIG. 10 and its corresponding description. In some embodiments, the sequence-to-sequence model 220 can be trained using the training data generated using the new data augmentation operator to improve its ability to generate data augmentation operators.
[0041]
[0058] As shown in FIG. 2A, the sequence-to-sequence model 220 interacts with unlabeled data 231-232 to generate a data augmentation operator. The interface of the sequence-to-sequence model 220 with the unlabeled data 231-232 can include converting the input data and generating a token sequence. The sequence-to-sequence model 220 can interact with the unlabeled data repository 130 (as shown in FIG. 1) to generate a sequential representation of the data as a token sequence. The sequence-to-sequence model 220 can search the unlabeled data repository 130 for unlabeled data 231. In some embodiments, the DAO generator 110 can fill the unlabeled data 231 with data from the unlabeled data repository 130. When the data augmentation system 100 receives a request for either generating a new data augmentation operator or training a machine learning model, it can share the data from the unlabeled data repository 130 with the DAO generator 110. The DAO generator can store the data received from the data augmentation system 100 as unlabeled data 231. In some embodiments, the DAO generator 110 can receive labeled data from the text corpus repository 160 (as shown in FIG. 1) for use in generating a data augmentation operator.
[0042]
[0059] The unlabeled data 231 can be a sequential representation of data generated by the sequence-to-sequence model 220 using data from the unlabeled data repository 130. In some embodiments, the unlabeled data 231 can be generated using labeled data within the text corpus repository 160, or a mixture of labeled and unlabeled data. The sequential data representation can include identifying individual tokens within the text and including markers indicating the start and end of each token. The sequential data can include the markers and tokens as a character string. In some embodiments, the sequence-to-sequence model 220 can introduce only a marker for the start of a token and can use the same marker as an end marker for the immediately preceding token. The input text from the unlabeled data 231 can be sequentially represented by including the class-type token marker "[CLS]" and another type of token marker "[SEP]". For example, "The room was modern room" can be sequentially formatted as "[CLS] The room was modern [SEP] room", indicating two tokens, the class token ("the room was modern") and another token ("room"), identified using the markers "[CLS]" and "[SEP]". The format of the sequential representation of the data can depend on the text classification task for which the unlabeled data 231 is used as training data to train the data augmentation system 100. A data augmentation system (e.g., the data augmentation system 100 of FIG. 1) trained for an intent classification task can include markers within the sentence sequence to indicate the start or end of a sequence. For example, the input sequence "where is the orange bowl?" used as input data for training for an intent classification task can be sequentially formatted as "[CLS] where is the orange bowl? [SEP]".
[0043]
[0060] In some embodiments, the sequence-to-sequence model 220 can convert tabular input data into a sequential representation before using it to generate data augmentation operators. In some embodiments, the data augmentation operator can include the functionality of serializing data before converting the serialized data using the augmentation functionality in the data augmentation operator. The sequence-to-sequence model 220 can convert tabular data into sequential data by converting each row of the data to include markers for each column cell and its content value using the markers "[COL]" and "[VAL]". An exemplary row in a table having contact information (in three columns, name, address, and phone) can be sequentially represented as follows: "[COL] Name [Val] Apple Inc. [COL] Address [VAL] 1 Infinity Loop [COL] Phone [VAL] 408-000-0000". In some embodiments, additional markers, such as "[SEP]" placed between tokens representing two rows, can be used to combine multiple rows of table data into a single sequence. A detailed description of the conversion of tabular data to sequential representation is presented below in connection with FIG. 2B and its corresponding description.
[0044]
[0061] FIG. 2B shows exemplary tabular and sequential representations of data for an entity matching classification task according to an embodiment of the present disclosure. As shown in FIG. 2B, tables 250 and 260 can represent training data for two entities that may be required to be matched. The DAO generator 110 can convert the input data in tables 250 and 260 into serialized forms as data 271 and 272, respectively, before converting them and generating data augmentation operators. Missing columns in table 260, such as "Model", can be presented in serialized form by including a missing column name with a null value.
[0045]
[0062] As shown in FIG. 2B, the two entities represented by tables 250 and 260 can be input data for an entity matching task. The target model 141 can return a match for representing the same entity (the book with the title "Instant Immersion Spanish Deluxe 2.0"). The target model 141 can return a match based on pre-training of the model using the serialized form of one or both of the tables 250 and 260 as training data. During the training of the target model 141, the data augmentation system 100 can serialize the table 250 as data 271 and generate additional training data 272 by applying a data augmentation operator. For example, the data 271 can be converted to data 272 using a data augmentation operator that performs a token replacement operation to replace the value "topics entertainment" with "NULL". Such a transformation for generating training data can help train the target model 141 to determine that the entities represented by the tables 250 and 260 represent the same entity (the book titled "Instant Immersion Spanish Deluxe 2.0"). A detailed description of training a machine learning model to perform an error detection classification task is presented below in connection with FIG. 8 and its corresponding description.
[0046]
[0063] In some embodiments, the serialization process can depend on the purpose of generating the data augmentation operator. A detailed description of the serialization of data specific to the purpose of error detection / data cleaning is presented below in connection with FIG. 2C and its description.
[0047]
[0064] FIG. 2C shows an exemplary tabular and sequential representation of data for an error detection classification task according to an embodiment of the present disclosure. The data 291 can represent the sequential presentation of rows within the table 280. As described in FIG. 2A, strings with markers "[COL]" and "[VAL]" can be used to indicate the start of each column name within the table row and the value for the column within the table row.
[0048]
[0065] For example, data 291 represents row 281 having three columns "Name", "Address", and "Phone" using the "[COL]" marker, the same column name following it, and the "[VAL]" marker, and the values existing within row 281 for each of the columns following it. Unlike in FIG. 2B where a complete table is serialized as a sequence for the entity matching task, in the error detection task, a single row 281 is serialized using data 291.
[0049]
[0066] As shown in FIG. 2C, the serialized data 291 can include an additional part indicating that the column values that need to be error-corrected are presented using the "[SEP]" marker. The additional part represents the cells 282 within row 281 of table 280 that are being examined for error detection and data cleaning. A detailed description of training a machine learning model to perform the error detection classification task is presented below in relation to FIG. 8 and its corresponding description.
[0050]
[0067] Referring back to FIG. 2A, the sequence-to-sequence model 220 can save the serialized unlabeled data to unlabeled data 231. In some embodiments, the DAO generator 110 can hold the serialized representations of the data in temporary memory and discard them when generating data augmentation operators.
[0051]
[0068] The DAO generator 110 can include unlabeled damaged data 232 that is used as an input to the sequence-to-sequence model 220 to generate data augmentation operators. The DAO generator 110 can generate unlabeled damaged data 232 by applying existing data augmentation operators to the unlabeled data 231. The DAO generator 110 can access existing data augmentation operators from the DAO repository 120 (as shown in FIG. 1). In some embodiments, the data augmentation system 100 can provide the DAO generator 110 with existing data augmentation operators along with the unlabeled data 231. Existing data augmentation operators can transform tokens or spans or complete token sequences of a serialized sequence of unlabeled data. For example, a token deletion data augmentation operator can eliminate a single token within a token sequence. In another scenario, a span replacement data augmentation operator can replace a sequence of words within a sentence represented by a token sequence. In yet another scenario, a back translation data augmentation operator can translate a token sequence representing a sentence into a different language and then back translate it into the original language of the sentence that can result in a modified token sequence. The token sequences updated using the existing data augmentation operators provided to the DAO generator 110 are stored within the unlabeled damaged data 232.
[0052]
[0069] The data within the unlabeled data 232 without damaged labels can be generated by applying a plurality of existing data augmentations of the operator to each sequence within the unlabeled data 231. In some embodiments, the DAO generator 110 can apply different sets of existing data augmentation operators to each token sequence within the unlabeled data 231. In some embodiments, the DAO generator 110 can apply the same set or subset of existing data augmentation operators to each sequence in a different order. The data augmentation operator can be randomly selected from an existing set of data augmentation operators to achieve the application of different data augmentation operators. The DAO generator 110 can select the data augmentation operators in a different order for applying to each token sequence to produce a random effect on the token sequence. The set of data augmentation operators applied to the token sequences within the unlabeled data 231 can be based on the topic of the unlabeled data 231. In some embodiments, the set of data augmentation operators applied can depend on the classification task performed by a machine learning model (e.g., the target model 141 of FIG. 1) that can utilize the training data generated using the newly generated data augmentation operators as a result.
[0053]
[0070] The data 232 without damaged labels can include a relationship to the original data within the label-free data 231 that has been transformed using existing data augmentation operators. The relationship can include a reference to the original token sequence that has been damaged using a series of existing data augmentation operators. The DAO generator 110 can use the data 232 without damaged labels as an input to the sequence-to-sequence model 220 to generate data augmentation operators. The sequence-to-sequence model 220 can generate data augmentation operators by determining a set of data transformation operations required to revert the transformed token sequence within the data 232 without damaged labels back to the associated original token sequence. The sequence-to-sequence model 220 can be pre-trained to learn the transformation operations required to revert the transformed token sequence within the damaged sequence back to the original sequence. The sequence-to-sequence model 220 can determine the transformation operations used to construct the data augmentation operators. Such data augmentation operators constructed by reverting the result of the transformation back are called inverse data augmentation operators. A detailed description of inverse data augmentation construction is presented below in connection with FIG. 13 and its corresponding description. In some embodiments, the DAO generator 110 can maintain a record of all existing data augmentation operators applied to the original sequence within the label-free data 231 to generate the transformed sequence within the data 232 without damaged labels. The DAO generator 110 can associate the tracked existing data augmentation operators with the transformed sequence and store them within the data 232 without damaged labels. The tracked set of existing data augmentation operators applied during the transformation can be combined to generate data augmentation operators.
[0054]
[0071] The DAO generator 110 can transfer the generated inverse data augmentation operators to the DAO repository 120. In some embodiments, the DAO generator 110 can notify the data augmentation system 100 about the generated inverse data augmentation operators. The data augmentation system 100 can obtain the inverse data augmentation operators from the DAO generator 110 and store them in the DAO repository 120. In some embodiments, the DAO generator or the data augmentation system 100 can supply the inverse data augmentation operators to the ML model platform 150 (as shown in FIG. 1) to generate new training data. The DAO generator 110 can generate inverse data augmentation operators at regular intervals. In some embodiments, the DAO generator 110 can generate inverse data augmentation operators when new data is added to the unlabeled data repository 130 and / or the text corpus repository 160.
[0055]
[0072] Referring back to FIG. 1, in a natural language processing system such as the data augmentation system 100, opinions can be conveyed using different words or groups of words that have a similar meaning. The data augmentation system 100 can identify the opinion of an input sentence by extracting classification information using a machine learning model in the machine learning model repository 140. The data augmentation system 100 can extract classification information by pre-training the target model 141 using limited training data in the text corpus repository 160. Using the labeled data in the text corpus repository 160, the data augmentation system 100 can generate additional data (e.g., using the ML model platform 150, which is described in more detail below) to generate multiple sentences that convey related opinions. The generated sentences can be used to create new sentences. The sentences themselves can be complete sentences.
[0056]
[0073] By generating additional clauses according to the described embodiments, the data augmentation system 100 can extract classification information cost-effectively and efficiently. Further, the data augmentation system 100, outlined above and described in more detail below, can generate additional data from limited labeled data, which other existing systems may consider an insufficient amount of data. The data augmentation system 100 can utilize unlabeled data within the unlabeled data repository 130 that is considered unusable by existing systems.
[0057]
[0074] The data augmentation system 100 can include a machine learning (ML) model platform 150 that can be used to train a machine learning model. The ML model platform 150 can access a machine learning model from the machine learning (ML) model repository 140 for training. The ML model platform 150 can train a model from the ML model repository 140 using data from the text corpus repository 160 and / or the unlabeled data repository 130 as input. The ML model platform 150 can also obtain data augmentation operators from the DAO repository 120 as input to generate additional training data for training the machine learning model.
[0058]
[0075] In some embodiments, the ML model platform 150 can connect to a meta-learning policy framework 170 to determine whether additional training data generated using data augmentation operators from the DAO repository 120 can be used to train a machine learning model. A detailed description of the components of the meta-learning policy framework 170 is presented below in connection with FIG. 3 and its corresponding description.
[0059]
[0076] FIG. 3 is a block diagram showing exemplary components of a meta-learning policy framework 170 according to an embodiment of the present disclosure. The meta-learning policy framework 170 can assist in training a machine learning model trained using an ML model platform 150 (as shown in FIG. 1). The meta-learning policy framework 170 can assist in the learning of a machine learning model by fine-tuning the training data of the machine learning model filled with data augmentation operators. The meta-learning policy framework 170 can learn to identify the most important training data and learn the effectiveness of the important training data when training the machine learning model (e.g., the target model 141 in FIG. 1), thereby fine-tuning the training data. Fine-tuning can include adjusting the weights of each token sequence in the training data. The meta-learning policy framework 170 can use a machine learning model for learning to identify important training data from data generated using data augmentation operators in a DAO repository 120 (as shown in FIG. 1) applied to the training data in a text corpus repository 160 (as shown in FIG. 1).
[0060]
[0077] The components of the meta - learning policy framework 170 can include a machine - learning model for identifying important token sequences for use as training data for a machine - learning model, a filtering model 371, and a weighting model 372. The meta - learning policy framework 170 can also include a learning ability for improving the identification of token sequences. The meta - learning policy framework 170 can help improve the training of a machine - learning model through repeated improved identification of important token sequences used as input to the machine - learning model. The meta - learning policy framework 170 also includes a loss function 373 that can assist the filtering model 371 and the weighting model 372 in improving their identification of important token sequences for training the machine - learning model. The loss function 373 can help improve the identification ability of the filtering model 371 and the weighting model 372 by evaluating the effectiveness of a machine - learning model trained using the data identified by the filtering model 371 and the weighting model 372. The loss function 373 can help educate the filtering model 371 and the weighting model 372 to learn how to identify important data for training the machine - learning model. The use of the loss function 373 for improving the filtering model 371 and the weighting model 372 is presented below in connection with FIG. 4A and its corresponding description.
[0061]
[0078] Filtering model 371 can help filter out and remove undesirable sequences generated using data augmentation operators applied by the ML model platform 150 to sequences within the text corpus repository 160. The filtering model 371 can be a binary classifier that determines whether to consider or discard the augmented sequences generated by applying the data augmentation operators stored within the DAO repository 120 (generated by the DAO generator 110 of FIG. 1) by the ML model platform 150. The filtering model 371 can be a machine learning model that can be trained using the same sequences generated for training a machine learning model (e.g., the target model 141). The binary classification task within the filtering model 371 can be context-dependent based on the classification task of the machine learning model. The filtering model 371 can classify based on the specificity of the augmented sequences, such as new tokens introduced by the data augmentation operators to the original token sequences (within the text corpus repository 160).
[0062]
[0079] Filtering model 371 can be used to quickly filter the extended sentences when the number of extended sentences exceeds a specific threshold. In some embodiments, filtering model 371 can filter the extended examples based on specific predetermined conditions. A user (e.g., user 190 in FIG. 1) can supply the conditions and is configured using a configuration file provided to data expansion system 100 as part of a request (e.g., request 195 in FIG. 1). In some embodiments, filtering model 371 can intelligently filter specific extended sentences based on the data expansion operators applied to the existing sequences to generate the expansion sequences. The selection of specific data expansion operators and the expansion sequences created by them can be based on the topic of the existing sequences. In some embodiments, filtering model 371 can be applied to the data expansion operators instead of the extended sentences generated by the data expansion operators.
[0063]
[0080] Weighting model 372 can determine the importance of the selected examples by assigning weights to the expansion sequences for calculating the loss of training a machine learning model using the expansion sequences. In some embodiments, weighting model 372 can be directly applied to the extended examples generated using the data expansion operators of DAO repository 120. In some embodiments, weighting model 372 can determine which data expansion operators can be used more than others by applying the weights to the data expansion operators instead of the expansion sequences.
[0064]
[0081] The loss function 373 can be used to train the filtering model 371 and the weighting model 372. In some embodiments, the loss function 373 can be a layer of a target machine learning model (e.g., target model 141) that helps evaluate the validation loss value when executing the target machine learning model to classify a validation sequence set. The calculation of the validation loss and the backpropagation of the loss value are presented below in connection with FIG. 4B and its corresponding description.
[0065]
[0082] Referring back to FIG. 1, the ML model platform 150 can use the data augmentation operator in the DAO repository 120 to process the data in the text corpus repository 160 and generate additional training data for the machine learning model. The data augmentation operator can process the data in the text corpus repository 160 and generate additional labeled data by updating one or more parts of an existing text sentence. In some embodiments, the data augmentation operator can receive some or all of the sentences directly as input from a user (e.g., user 190) instead of loading them from the text corpus repository 160.
[0066]
[0083] The ML model platform 150 can store additional data (e.g., in the form of sentences) in the text corpus repository 160 for later use. In some embodiments, the additional data is temporarily stored in memory and supplied to the DAO generator 110 to generate additional training data. The text corpus repository 160 can receive and store the additional data generated by the ML model platform 150.
[0067]
[0084] The ML model platform 150 can select different data augmentation operators for application to the input data selected from the text corpus repository 160 to generate additional data. The ML model platform 150 can select different data augmentation operators for each input data sentence. The ML model platform 150 can also select data augmentation operators based on a predefined criterion or in a random manner. In some embodiments, the ML model platform 150 can apply the same data augmentation operator over a set of sentences or a set period of time.
[0068]
[0085] The text corpus repository 160 can be pre-filled using a corpus of sentences. In some embodiments, the text corpus repository 160 stores a set of input sentences supplied by a user before passing them to other components of the data augmentation system 100. In other embodiments, the sentences within the text corpus repository 160 can be supplied by a separate system. For example, the text corpus repository 160 can include sentences supplied by user input, other systems, other data sources, or feedback from the data augmentation system 100 or its components. As described above with reference to the unlabeled data repository 130, the text corpus repository 160 can be a relational database management system (RDBMS) (e.g., an Oracle database, a Microsoft SQL server, MySQL, PostgreSQL, or an IBM DB2). The RDBMS can be designed to efficiently return data for rows or entire records from the database with as few operations as possible. The RDBMS can store data by serializing each row of data within the data structure. Within the RDBMS, the data associated with a record can be stored serially such that the data associated with all categories of the record can be accessed in a single operation. Further, the RDBMS can efficiently enable access to related records stored within different tables. For example, in an RDBMS, tables can be linked by reference columns, and the RDBMS can join the tables together to obtain data for the data structure. In some embodiments, the text corpus repository 160 can be a non-relational database system (NRDBMS) (e.g., XML, Cassandra, CouchDB, MongoDB, an Oracle NoSQL database, FoundationDB, or Redis).Non-relational database systems can store data using various data structures, such as key-value stores, document stores, graphs, and tuple stores, among others. For example, a non-relational database using a document store could combine all of the data associated with a particular identifier into a single document encoded using XML. The text corpus repository 160 can also be an in-memory database such as Memcached. In some embodiments, the contents of the text corpus repository 160 can exist in both a persistent storage database and an in-memory database, such as is possible in Redis. In some embodiments, the text corpus repository 160 can be stored on the same database as the unlabeled data repository 130.
[0069]
[0086] The data augmentation system 100 can receive requests for various tasks, including generating data augmentation operators for generating augmented data for training machine learning models. Requests to the data augmentation system 100 can include generating augmented data for training a machine learning model and extracting classification information using the machine learning model. The data augmentation system 100 can receive various requests through the network 180. The network 180 can be a local network, the Internet, or a cloud network. The user 190 can send requests to the data augmentation system 100 through the network 180 for various tasks enumerated above. The user 190 can interact with the data augmentation system 100 through a tablet, laptop, or portable computer using a web browser or an installed application. The user 190 can send requests 195 to the data augmentation system 100 through the network 180 for generating augmented data and operators for training a machine learning model and extracting classification information using the machine learning model.
[0070]
[0087] The components of the data augmentation system 100 can be executed on a single computer or can be distributed across multiple computers or processors. Different components of the data augmentation system 100 can communicate through a network (e.g., LAN or WAN) 180 or the Internet. In some embodiments, each component can be executed on multiple compute instances or processors. An instance of each component of the data augmentation system 100 can be a part of a connected network such as a cloud network (e.g., Amazon AWS, Microsoft Azure, Google Cloud). In some embodiments, some or all of the components of the data augmentation system 100 are executed within a virtualized environment such as a hypervisor or virtual machine.
[0071]
[0088] FIG. 4A shows an exemplary data flow diagram of the meta-learning data augmentation system 100 used for a sequence classification task, according to an embodiment of the present disclosure. The components shown in FIG. 4A refer again to the components described in FIGS. 1-3, and where appropriate, the labels for those components use the same label numbers as used in the previous figures.
[0072]
[0089] As shown in FIG. 4A, the data flow diagram shows an exemplary data flow between the components of the data augmentation system from the DAO generator 110 in stage 1 to the target model 141 and the loss function 373 in stage 5. The data flow can result in the generation of new training data and the training of various machine learning models for use in training machine learning models designed for various natural language processing tasks.
[0073]
[0090] In stage 1, the DAO generator 110 can receive the unlabeled data 231 as input and generate a set of data augmentation operators 421. The data augmentation operators 421 can be different from the basic data augmentation operators, such as operators for inserting, deleting, and swapping tokens or spans. The basic data augmentation operators used to generate the inverse data augmentation operators are presented above in relation to FIG. 2A and its corresponding description. As described in FIG. 2A, the DAO generator 110 can generate data augmentation operators that can create any extended sentence that has a very different structure from the original sentence but can retain the topic of the original sentence. Basic data augmentation operators that act on tokens or spans on the sentence sequence cannot create such diverse and meaningful extended sentences. For example, a simple token swap basic operator, when applied to the original sentence "where is the Orange Bowl?", can swap synonyms and create "Where is the Orangish Bowl". Instead, the application of the inverse data augmentation operator to the original sentence can create diverse extended sentences, such as "Where is the Indianapolis Bowl in New Orleans?" or "Where is the Syracuse University Orange Bowl?", that still maintain the topic (football). In another scenario, the basic token deletion operator applied to the original sentence can result in a structurally lost and grammatically incorrect "is the Orange Bowl?". By utilizing the inverse data augmentation operator, the data augmentation system 100 can create a diverse set of extended sentences for training the machine learning model in a later stage, as described below.
[0074]
[0091] In stage 2, the newly generated inverse data augmentation operator stored as data augmentation operator 421 can be applied to the existing training data 411 of the text corpus repository 160 to create augmented data 412. A subset of the data augmentation operators of data augmentation operator 421 can create augmented data 412 by applying each data augmentation operator to generate a plurality of augmented sentences. In some embodiments, data augmentation system 100 can skip applying some data augmentation operators to specific sentences within training data 411. The selective application of data augmentation operators to the original sentences of training data 411 can be preconfigured or based on a configuration provided to data augmentation system 100 by a user (e.g., user 190 of FIG. 1) beforehand or with a request (e.g., request 195 of FIG. 1).
[0075]
[0092] In some embodiments, unlabeled data 231 can be used to generate additional training data. Unlabeled data 231 needs to be associated with labels before being used as additional training data. As shown in FIG. 4A, data augmentation system 100 can generate soft labels 481 to be associated with unlabeled data 231 and additional augmented unlabeled data 432 generated using unlabeled data 231.
[0076]
[0093] The unlabeled data 231 and the extended unlabeled data 432 can have soft labels 481 applied using a close guess algorithm. The close guess algorithm can be based on the proximity of sequences within the unlabeled data 231 and the training data 411. The proximity of the sequences can be determined based on the proximity of the vector-encoded representations of the sequences within the unlabeled data 231 and the training data 411. In some embodiments, the label determination process can include averaging multiple versions of the labels generated by a machine learning model. The averaging process can include averaging the vectors representing the multiple labels. This label determination process can be part of the MixMatch method as described in FIG. 1 above. The soft label 481 is a character string that can be obtained using one of the methods described above.
[0077]
[0094] Unlabeled data 231 can be associated with soft label 481 applied using a similarity inference algorithm. The similarity inference algorithm can be based on the proximity of the unlabeled sequence of unlabeled data 231, the labeled sequence of training data 411, and the extended sequence of extended data 412. The proximity of the sequences can be determined based on the proximity of the vectors of the sequences. In some embodiments, the label determination process can include averaging multiple versions of the label. The averaging process can include averaging the vectors representing the multiple labels. The soft label 481 associated with the sequence of unlabeled data 231 can also be associated with the sequence of extended unlabeled data 432. The association of soft label 481 with extended unlabeled data 432 can be based on the relationship between the sequences of unlabeled data 231 and extended unlabeled data 432. In some embodiments, the sequence of unlabeled data 231 can share the same soft label with the extended sequence of extended unlabeled data 432 generated therefrom (by applying data expansion operator 421). Data expansion system 100 can generate a soft label for associating the extended sequence of extended unlabeled data 432 using the similarity inference algorithm described above. Data expansion system 100 uses policy model 470 at a later stage to identify training data that satisfies the policy. Policy model 470 can help set the quality of the training data for training a machine learning model (e.g., target model 141). Policy model 470 can be implemented using filtering model 371 and weighting model 372 of the machine learning model.
[0078]
[0095] In stage 3, a new set of training data present within the extended data 412 and the data without extended labels 432 is examined to identify high-quality training data using the filtering model 371. The filtering model 371 is involved in making a precise determination as to whether a particular sequence of the extended data 412 and the data without extended labels 432 should be included. In some embodiments, the filtering model 371 can apply different filtering strategies to filter the extended data 412 and the data without extended labels 432. A detailed description of the filtering model 371 is presented above in connection with FIG. 3 and its corresponding description.
[0079]
[0096] In stage 4, the data augmentation system 100 can further identify the most important training data from the filtered set of training data using the weighting model 372. Different from the precise determination method of the filtering model 371, the weighting model 372 reduces the effect of a particular sequence when training a machine learning model by associating weights with the sequence. A machine learning model that consumes sequences for training can ignore the results if low weights are applied to the sequences by the weighting model 372. A detailed description of the weighting model 372 is presented above in connection with FIG. 3 and its corresponding description.
[0080]
[0097] In stage 5, the target model 141 can be trained using the target batch 413 of sequences identified by the configured policy model 470, which is implemented by the filtering model 371 and the weighting model 372. The target model 141 can also be fine-tuned using the loss function 373.
[0081]
[0098] The loss function 373 can fine-tune the behavior of the target model 141 for other machine learning models such as the provided target batch 413 and the filtering model 371 and the weighting model 372 when determining the sequences to be included in the target batch 413. The loss function 373 can fine-tune the machine learning model by determining the amount of deviation between the original data (e.g., training data 411) and the data generated through the process represented by stages 1 to 4. The loss function 373 can be a cross-entropy loss or an L2 loss function. The loss function 373 can be used to compare the probabilistic output of the machine learning model from the softmax layer (as described later in FIG. 10) to calculate a score regarding the dissimilarity between the output and the label.
[0082]
[0099] (Shown as an arrow between the loss function 373 and the machine learning model, i.e., the target model 141, the filtering model 371, and the weighting model 372 in FIGS. 4A and 4B) The backpropagation step of stage 5 can update the linear layers 1031 to 1033 (shown in FIG. 10 described later) and other layers of the target model 141 to assist in the classification task. The update can result in the generation of more accurate predictions for future use of the model and the layers. Various techniques such as the Adam algorithm, Stochastic Gradient Descent (SGD), and SGD with Momentum can be used to update the parameters in the layers within the target model 141 and other layers.
[0083]
[0100] FIG. 4B shows an exemplary backpropagation technique for fine-tuning a machine learning model and a policy for managing training data for the machine learning model according to an embodiment of the present disclosure. As shown in FIG. 4B, the two-phase backpropagation process shows the data flow for co-training the policy model 470 and the target model 141.
[0084]
[0101] In Phase 1, the loss function 373 can backpropagate the success level of the target model 141 regardless of the behavior of the policy model 470. When receiving the success level, the target model 141 can determine whether to use the current training data (e.g., training data 411) or request an updated training data set.
[0085]
[0102] In Phase 2, the backpropagation process can optimize the policy model 470, including the filtering model 371 and the weighting model 372, so that they can generate a more effective extended sequence for training the target model 141. The policy model 470 can generate a batch of selected extended sequences using the training data 411 supplied to the target model 141 to be trained and generate an updated target model 441. Next, using the updated target model 441, the loss can be calculated using the validation data 414. The calculated loss can be used to update the policy model 470 by backpropagating the loss value to the policy model 470. The policy model 470 can examine the calculated loss to improve and reduce the loss from the updated target model 441. The second phase of the backpropagation process can train the data augmentation policy defined by the policy model 470 so that the trained model can achieve good performance on the validation data 414.
[0086]
[0103] FIG. 5 is a block diagram of an exemplary computing device 500 in accordance with an embodiment of the present disclosure. In some embodiments, computing device 500 can be a dedicated server that provides the functionality described herein. In some embodiments, components of data augmentation system 100, such as DAO generator 110, ML model platform 150, and meta-learning policy framework 170 of FIG. 1, can be implemented using computing device 500, or a plurality of computing devices 500 operating in parallel. In some embodiments, a plurality of repositories of data augmentation system 100, such as DAO repository 120, unlabeled data repository 130, ML model repository 140, and text corpus repository 160, can be implemented using computing device 500. In some embodiments, models stored within ML model repository 140, such as target model 141, sequence-to-sequence model 220, filtering model 371, and weighting model 372, can be implemented using computing device 500. Further, computing device 500 can be a second device that provides the functionality described herein, or receives information from a server to provide at least a portion of the described functionality. Further, computing device 500 can be an additional device or group of devices that stores or provides data in accordance with an embodiment of the present disclosure, and in some embodiments, computing device 500 can be a virtual computing device such as a virtual machine, a plurality of virtual machines, or a hypervisor.
[0087]
[0104] Computing device 500 can include one or more central processing units (CPUs) 520 and system memory 521. Computing device 500 can also include one or more graphics processing units (GPUs) 525 and graphics memory 526. In some embodiments, computing device 500 can be a headless computing device that does not include the GPU(s) 525 or graphics memory 526.
[0088]
[0105] The CPU 520 can be a single or multiple microprocessors, field programmable gate arrays, or digital signal processors that have the ability to execute a set of instructions stored in a memory (e.g., system memory 521), cache (e.g., cache 541), or register (e.g., one of registers 540). The CPU 520 can include, among other things, one or more registers (e.g., registers 540) for storing various types of data including data, instructions, floating point values, conditional values, memory addresses for locations in a memory (e.g., system memory 521 or graphic memory 526), pointers, and counters. The CPU registers 540 can include dedicated registers used for storing data associated with executing instructions such as an instruction pointer, instruction counter, or memory stack pointer. The system memory 521 can include tangible or non-transitory computer-readable media such as a flexible disk, hard disk, compact disk read-only memory (CD-ROM), magneto-optical (MO) drive, digital versatile disk random-access memory (DVD-RAM), solid-state disk (SSD), flash drive or flash memory, processor cache, memory register, or semiconductor memory. The system memory 521 can be one or more memory chips having the ability to store data and enable direct access by the CPU 520. The system memory 521 can be any type of random-access memory (RAM) or other available memory chip having the ability to operate as described herein.
[0089]
[0106] The CPU 520 can sometimes communicate with the system memory 521 via a system interface 550, which is sometimes referred to as a bus. In embodiments that include a GPU 525, the GPU 525 can be any type of dedicated circuitry that can operate on and modify a memory (e.g., graphic memory 526) to provide or accelerate the creation of images. The GPU 525 can have a highly parallel structure optimized for processing large parallel blocks of graphical data more efficiently than the general-purpose CPU 520. Additionally, the functionality of the GPU 525 can be included within a chipset of a dedicated processing device or coprocessor.
[0090]
[0107] The CPU 520 can execute programming instructions stored in the system memory 521 or other memory, act on data stored in a memory (e.g., the system memory 521), and communicate with the GPU 525 through a system interface 550 that bridges communication between various components of the computing device 500. In some embodiments, the CPU 520, GPU 525, system interface 550, or any combination thereof is integrated within a single chipset or processing device. The GPU 525 can execute a set of instructions stored in a memory (e.g., the system memory 521) and operate on graphical data stored in the system memory 521 or the graphic memory 526. For example, the CPU 520 can provide instructions to the GPU 525, and the GPU 525 can process the instructions and render the graphics data stored in the graphic memory 526. The graphic memory 526 can be any memory space accessible by the GPU 525, including local memory, system memory, on-chip memory, and hard disks. The GPU 525 can enable the display of graphical data stored in the graphic memory 526 on the display device 524, or process the graphical information and provide the information to connected devices through the network interface 518 or the I / O device 530.
[0091]
[0108] Computing device 500 can include a display device 524 and an input / output (I / O) device 530 (e.g., a keyboard, mouse, or pointing device) connected to an I / O controller 523. The I / O controller 523 can communicate with other components of the computing device 500 via a system interface 550. Here, it should be understood that the CPU 520 can also communicate with the system memory 521 and other devices in ways other than through the system interface 550, such as through serial communication or direct point-to-point communication. Similarly, the GPU 525 can communicate with the graphics memory 526 and other devices in ways other than through the system interface 550. In addition to receiving input, the CPU 520 can provide output via the I / O device 530 (e.g., through a printer, speaker, bone conduction, or other output device).
[0092]
[0109] Further, computing device 500 can include a network interface 518 for interfacing and connecting to a LAN, WAN, MAN, or the Internet through various connections, including but not limited to standard telephone lines, LAN or WAN links (e.g., 802.21, T1, T3, 56 kb, X.25), broadband connections (e.g., ISDN, frame relay, ATM), wireless connections (e.g., in particular, those compliant with 802.11a, 802.11b, 802.11b / g / n, 802.11ac, Bluetooth, Bluetooth LTE, 3GPP, or WiMax standards), or any combination of any or all of the above. Network interface 518 can include an embedded network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem, or any other device suitable for interfacing computing device 500 to any type of network having communication capabilities and performing the operations described herein.
[0093]
[0110] FIG. 6 shows an exemplary machine learning model, such as a target model 141 for a sequence classification task, according to an embodiment of the present disclosure. The target model 141 can be a pre-trained machine learning model such as BERT, DistilBERT, RoBERTa, etc. The target model 141 can be used for various classification tasks that obtain classification information embedded in the input sentence. The classification information can be a label of a class of a group of labels given to the input sentence. The classification information can include sentiment, topic, and intention conveyed in the input sentence. For example, a product review input sentence can be supplied to a language machine learning model to obtain sentiment classification information. The sentiment classification information can be a part of a class of values defined by a user of the target model 141 (e.g., user 190 in FIG. 1). For example, sentiment analysis of a product review input sentence can result in values +1, -1, and 0 representing positive, negative, and neutral sentiment, respectively. Depending on the embodiment, the input to the target model 141 can be input text of multiple sentences. Depending on the embodiment, the input to the target model 141 can be a question or a statement.
[0094]
[0111] As shown in FIG. 6, the target model 141 can obtain the input in a sequential format. For example, the input sentence "The room was modern room" can be converted into a token sequence as "[CLS] The room was modern [SEP] room", using "[CLS]" to indicate the start of the token sequence and other tokens separated by the "[SEP]" symbol. Depending on the embodiment, the format of the sequential representation of the data may depend on the text classification task for which the input sequence is used as training data. A data augmentation system trained for an intent classification task (e.g., the data augmentation system 100 of FIG. 1) can include markers within the sentence sequence to indicate the start or end of the sequence. For example, the input sequence "where is the orange bowl?" used as input data for training for an intent classification task can be represented in sequential format as "[CLS] where is the orange bowl? [SEP]".
[0095]
[0112] As shown in FIG. 1, when the product review sentence 620 is input, the target model 141 can generate an output 630. The output 630 is a set of sentiment values represented by the set of numerical values {+1, -1}. The product review sentence 620 can result in more than one sentiment value from the set of sentiment values. Depending on the embodiment, the sentiment value output can be the sum of multiple sentiment values. The total sentiment value can be a simple sum of all sentiment values. Depending on the embodiment, weights can be applied to each sentiment value before summing to generate the composite sentiment value of the input sentence 610. The input token sequence (e.g., the product review sentence 620) can be generated using the data augmentation operator of the DAO repository 120. The data augmentation operator used to generate the input token sequence can be the inverse data augmentation operator generated using the DAO generator 110.
[0096]
[0113] FIG. 7 shows an exemplary sequence list for entity matching as a sequence classification task according to an embodiment of the present disclosure. The list includes the original sequence 710 and the modified sequences 720 generated using data augmentation operators 711-715. The basic data augmentation operators 711-712 perform simple token changes, resulting in either a syntactically incorrect sequence (e.g., extended sequence 721) or a semantically incorrect sequence (e.g., extended sequence 722). The data augmentation operators 711-715 can be independent of a specific task performed by a machine learning model. The data augmentation system 100 can enable training data for various classification tasks (as described in FIGS. 2A-2C) represented in the same serialized format used to train a machine learning model (e.g., the target model 141 of FIG. 1). The data augmentation system 100 can uniformly represent different data in a serialized format by applying the same data augmentation operators. The application of the same data augmentation operators 711-715 to training data used to train a machine learning model for error detection and text classification is given in the description of FIGS. 8 and 9 below.
[0097]
[0114] FIG. 8 shows an exemplary sequence list for error detection as a sequence classification task according to an embodiment of the present disclosure. The list includes the original sequence 810 and the extended sequences 820 generated using data augmentation operators 711-715. The basic data augmentation operators 711-712 can perform simple token changes to the original sequence 810, resulting in either a syntactically incorrect sequence (e.g., extended sequence 822) or a sequence that is too semantically close to the original sequence and thus lacks diversity (e.g., extended sequence 821).
[0098]
[0115] The data augmentation operators can be applied to different training data sets (e.g., the original sequence 810) to generate new extended sequences 820 for training a machine learning model for different classification tasks (e.g., error detection).
[0099]
[0116] Figure 9 shows an exemplary sequence list text classification as a sequence classification task according to an embodiment of the present disclosure. A list, the original sequence 910, and the extended sequences 920 generated using the data augmentation operators 711-715. The basic data augmentation operators 711-712 perform simple token changes, resulting in either a syntactically incorrect sequence (e.g., extended sequence 922) or a sequence that is too semantically close to the original sequence and thus lacks diversity (e.g., extended sequence 921).
[0100]
[0117] The data augmentation operators applied to the original sequence 710 can be applied to different training data sets (e.g., the original sequence 910) to generate new extended sequences 920 for training machine learning models for different classification tasks (e.g., error detection). Further, the same inverse data augmentation operator, such as "invdDA1" 713, can result in different operations being applied when applied to different sequences (e.g., the original sequences 710, 810, 910). For example, "invDA1" 713 applied to the original sequences 710 and 810 results in a token insertion operation at the end of the extended sequences, as shown in extended sequences 723 and 823. The same operator applied to the original sequence 910 results in two token insertions, as shown in extended sequence 923. The different operations applied by the same inverse data augmentation operator (e.g., "invDA1" 713) can be context-based, such as for a classification task.
[0101]
[0118] Figure 10 shows an exemplary language machine learning model having specific layers, according to an embodiment of the present disclosure. The target model 141 can include or be connected to additional task-specific layers 1030 and 1040. The target model 141 can be used for various classification tasks by updating the task-specific layers 1030 and 1040. As shown in FIG. 10, the task-specific layer 1030 can include linear layers 1031-1033, and the task-specific layer 1040 can include softmax layers 1041-1042.
[0102]
[0119] The linear layers 1031-1032 can help extract various types of classification information. For example, the linear layer 1031 can extract sentiment classification information. As described in FIGS. 6-8, the sequential format of the tabular form and the sequential representation of the key-value information can help handle entity matching and error detection tasks. The classification information for such alternative tasks can be match or mismatch values or clean or dirty values.
[0103]
[0120] Softmax layers 1041-1042 can help understand the probabilities of each encoding and related labels. Softmax layers 1041-1042 can understand the probability of each encoding by converting the class prediction score into a probability prediction of the input instance (e.g., input sequence 1020). The conversion can include converting a vector representing the classes of the classification task into probability percentages. For example, an input sequence 1020 having a vector (1.6, 0.0, 0.8) representing various sentiment class values can be provided as an input to softmax layer 1041 to generate the probabilities of the positive, neutral, and negative values of the sentiment class. The output vector of the probabilities generated by softmax layer 1041 can approximate a one-hot distribution. The proximity of the output of softmax layer 1041 to the percentage distribution can be based on the characteristics of the softmax function used by softmax layer 1041. The one-hot distribution can be a multi-dimensional vector. One of the dimensions of the multi-dimensional vector can contain the value 1, and the other dimensions can contain 0 as the value. The ML model can utilize the multi-dimensional vector to predict one of the classes with a 100% probability.
[0104]
[0121] Figure 11 is a flowchart showing an exemplary method for data augmentation operations according to an embodiment of the present disclosure. The steps of method 1100 can be performed on computing device 500 of FIG. 5 or, alternatively, by a system (e.g., data augmentation system 100 of FIG. 1) using its features, for purposes of illustration. It is understood that the exemplary method 1100 can be modified to change the order of the steps and to include additional steps.
[0105]
[0122] In step 1110, the system can access a machine learning model (e.g., the target model 141 in FIG. 1) from a machine learning model repository (e.g., the ML model repository 140 in FIG. 1). The system can access the machine learning model (e.g., the target model 141) based on a task request received by the system, such as a classification task. For example, the system can receive a request for a text classification task for performing sentiment analysis on a set of comments. In some embodiments, the system can receive a request, particularly for generating an extended token sequence for providing training data to the machine learning model. In some embodiments, the system can receive a request for training the machine learning model. In some embodiments, the system can automatically trigger a data augmentation process at regular intervals.
[0106]
[0123] In step 1120, the system can identify a dataset of the accessed machine learning model (e.g., the training data 411 in FIG. 4A). The system can identify the dataset based on various factors, such as the requested task. For example, a text classification task for performing sentiment analysis can result in the selection of a set of review comments for training the machine learning model using a text classification task layer. In some embodiments, the dataset can be selected based on the requested machine learning model.
[0107]
[0124] In step 1130, the system can generate a set of data augmentation operators using the identified dataset. The system can select one or more operators for the identified dataset, generate a modified sequence, and generate a set of data augmentation operators by determining an operation to revert the modified sequence back to the identified dataset. The identified operation can result in the determination of the augmentation operator. In some embodiments, the system can select an input sequence from the identified dataset (e.g., training data 411) for applying the data augmentation operator. A further description of the process of generating a data augmentation operator for generating a new token sequence is presented below in connection with FIG. 13 and its description.
[0108]
[0125] In step 1140, the system can apply at least one data augmentation operator from the set of data augmentation operators (e.g., data augmentation operator 421) to the selected input token sequence (e.g., training data 411) to generate at least one token sequence (e.g., augmented data 412). The system can select the input token sequence based on a policy criterion. The user can supply the policy criterion as part of the request sent to the system in step 1110. In some embodiments, all of the training data previously associated with the accessed machine learning model (e.g., target model 141) (e.g., training data 411) is selected as the input token sequence. In some embodiments, selecting the token sequence can be based on converting the training data into a token sequence by serializing the training data. Serialization of training data in different formats is presented above in connection with FIGS. 2A - C and their corresponding descriptions.
[0109]
[0126] The system can select a data augmentation operator (e.g., data augmentation operator 421) from a data augmentation operator (DAO) repository 120 that includes a set of data augmentation operators generated in step 1130. The system can select a data augmentation operator based on the task requested in step 1110. In some embodiments, the data augmentation operator can be selected based on the machine learning model accessed in step 1110 or an available training set (e.g., training data 411). For example, the training set for a machine learning model is limited, resulting in the selection of a larger number of data augmentation operators to generate a large number of token sequences for training the machine learning model. The system can apply a plurality of identified data augmentation operators to each of the selected input sequences. In some embodiments, a particular data augmentation operator can be skipped based on a policy provided by the user or determined during a previous execution. A policy for selecting important new token sequences and, thus, the data augmentation operators that create the new token sequences are presented above in connection with FIG. 4B and its corresponding description.
[0110]
[0127] In step 1150, the system can filter at least one token sequence using a filtering model (e.g., filtering model 371 of FIG. 3). A detailed description of filtering token sequences is presented above in connection with FIGS. 4A - B and their corresponding descriptions.
[0111]
[0128] In step 1160, the system can determine the weight of at least one token sequence within the at least one filtered token sequence. The system can use a weighting model 372 to determine the weight of each token sequence. A detailed description of calculating the weight of a token sequence is presented above in connection with FIGS. 4A - B and their corresponding descriptions.
[0112]
[0129] In step 1170, the system can apply weights to at least one token sequence among at least one token sequence. The system can apply weights only to the sequences filtered by the filtering model in step 1160. The system can apply the calculated weights by associating the weights with the token sequences. In some embodiments, the association or the sequences and the associated weights can be stored in a database (e.g., text corpus repository 160).
[0113]
[0130] In step 1180, the system can identify a subset of token sequences (e.g., the target batch 413 in FIG. 4A) from the generated at least one token sequence (e.g., the augmented data 412). The system can select sequences based on the weights applied to the sequences. The system can calculate a validation loss for determining the selection of the subset of sequences. In some embodiments, the data augmentation system can repeat steps 1160 and 1180 until the validation loss reaches a threshold. The filtering model 371 and the weighting model 372 can be adjusted by the system using the validation loss to help identify a subset of sequences.
[0114]
[0131] In step 1190, the system can provide the selected subset of token sequences as input to a machine learning model (e.g., the target model 141). In some embodiments, the system can store the selected subset of sequences in a training data repository (e.g., text corpus repository 160) before providing them to the training machine learning model. When step 1190 is completed, the system completes executing method 1100 on computing device 500 (step 1199).
[0115]
[0132] Figure 12 is a flowchart showing an exemplary method for data augmentation operations according to an embodiment of the present disclosure. The steps of method 1200 may be performed on computing device 500 of FIG. 5 or, alternatively, by a system (e.g., data augmentation system 100 of FIG. 1) that utilizes its features, for purposes of illustration. It is understood that the exemplified method 1200 may be modified to change the order of steps and to include additional steps.
[0116]
[0133] In step 1210, the system can access unlabeled data 231 from unlabeled data repository 130. The system can access unlabeled data 412 when it is configured to function in a semi-supervised learning mode. A user (e.g., user 190 of FIG. 1) can configure the system to operate in a semi-supervised mode by providing, as part of a request (e.g., request 195 of FIG. 1), configuration parameters.
[0117]
[0134] In step 1220, the system can generate an extended unlabeled token sequence (extended unlabeled data 432 of FIG. 4A) of the accessed unlabeled data 231. The system can apply a data augmentation operator (e.g., data augmentation operator 421) to the unlabeled data 231 to generate the extended unlabeled token sequence. The data augmentation system can select a data augmentation operator from DAO repository 120. The system can also generate the data augmentation operators before applying them to the unlabeled data 231. The process of generating and selecting data augmentation operators is presented above in connection with FIG. 11 and its corresponding description.
[0118]
[0135] In step 1230, the system can determine the soft labels of the token sequences without extended labels (e.g., data 432 without extended labels). The data augmentation system can determine the soft labels by inferring the labels (e.g., soft label 481) for the unlabeled data 231. The system can apply the inferred soft labels to all token sequences without extended labels generated from the token sequences in the unlabeled data 231. In some embodiments, the system can also generate soft labels for the token sequences without extended labels by inferring the labels.
[0119]
[0136] In step 1240, the system can associate the soft labels with the token sequences without extended labels. The association of the labels can include the system storing the association in a database (e.g., the unlabeled data repository 130).
[0120]
[0137] In step 1250, the system can provide the token sequences without extended labels, along with the associated soft labels, as input to a machine learning model (e.g., the target model 141). When step 1250 is completed, the system completes executing method 1200 on the computing device 500 (step 1299).
[0121]
[0138] FIG. 13 is a flowchart showing an exemplary method for generating an inverse data augmentation operator according to an embodiment of the present disclosure. The steps of method 1300 can be executed on the computing device 500 of FIG. 5 for illustrative purposes, or otherwise, by a data augmentation operator (DAO) generator (e.g., the data augmentation operator generator 110 of FIG. 1) that uses its features. In some embodiments, other components of the data augmentation system 100 can execute method 1300. It is understood that the illustrated method 1300 can be modified to change the order of the steps and to include additional steps.
[0122]
[0139] In step 1310, the DAO generator can access the unlabeled data (e.g., unlabeled data 231 in FIG. 4A) from the unlabeled data repository 130. In some embodiments, the DAO generator can access the unlabeled data 231 through the network 180 (as shown in FIG. 1). A user (e.g., user 190 in FIG. 1) can provide unlabeled data through the network 180.
[0123]
[0140] In step 1320, the DAO generator can check whether the accessed unlabeled data is formatted as a database table. If the answer to the question is yes, then at this time, the DAO generator can proceed to step 1330.
[0124]
[0141] In step 1330, the DAO generator can convert each row in the database table into a token sequence. Serializing tabular data into a token sequence is presented above in connection with FIGS. 2A - C and their corresponding descriptions.
[0125]
[0142] If the answer to the question in step 1320 is no, that is, if the unlabeled data 231 is not formatted as a database table, then at this time, the DAO generator can jump to step 1340. In step 1340, the DAO generator can prepare one or more token sequences of the accessed unlabeled data. The DAO generator can prepare the token sequences by serializing the unlabeled data. Serializing the unlabeled data can require including markers that separate the tokens of the text within the unlabeled data. The tokens within the text can be the words within the sentence identified by the separating space character.
[0126]
[0143] In step 1350, the DAO generator can convert one or more prepared token sequences and generate at least one damaged sequence (e.g., damaged label-free data 232). The DAO generator can generate damaged sequences using basic data augmentation operators (e.g., data augmentation operators 711 - 712). The basic data augmentation operators can include operators for tokens (e.g., token substitution, swap, insertion, deletion), spans containing multiple tokens (e.g., span substitution, swap), or the entire sequence (e.g., back translation). The DAO generator can generate damaged sequences by applying multiple basic data augmentation operators within the sequence. The DAO generator can randomly select basic data augmentation operators to apply to the token sequence to generate damaged sequences. The generated damaged sequences are associated with the original token sequences. In some embodiments, a single token sequence can be associated with multiple damaged sequences generated by applying different sets of basic data augmentation operators, or the same set of basic data augmentation operators applied in different orders.
[0127]
[0144] In step 1360, the DAO generator can provide one or more token sequences and at least one generated damaged sequence as inputs to a sequence-to-sequence model 220 (as shown in FIG. 2).
[0128]
[0145] In step 1370, the DAO generator can execute the sequence-to-sequence model 220 and determine one or more operations required to revert at least one damaged sequence back to its respective sequence within the one or more token sequences.
[0129]
[0146] In step 1380, the DAO generator can generate data augmentation operators (e.g., the inverse data augmentation operators 713-715 in FIG. 7) based on one or more determined operations to reverse at least one damaged sequence. When step 1380 is completed, the DAO generator completes executing method 1300 on computing device 500 (step 1399).
[0130]
[0147] FIG. 14 is a flowchart showing an exemplary method for generating training data for a machine learning model according to an embodiment of the present disclosure. The steps of method 1400 may be executed on computing device 500 of FIG. 5 for illustrative purposes, or otherwise by a data augmentation operator (DAO) generator (e.g., data augmentation operator generator 110 of FIG. 1) that uses its features. In some embodiments, other components of data augmentation system 100 may execute method 1400. It is understood that the illustrated method 1400 may be modified to change the order of steps and to include additional steps.
[0131]
[0148] In step 1410, the DAO generator can select a token sequence (e.g., a text sentence) from one or more token sequences (e.g., unlabeled data 231 in FIG. 2, training data 411 in FIG. 4A). The DAO generator can select sequences serially. In some embodiments, the DAO generator can select based on selection criteria that depend on a request received from a user. The received request can include a classification task that requires a particular type of data.
[0132]
[0149] In step 1420, the sequence-to-sequence model 220 can select a data augmentation operator from a set of data augmentation operators existing in the DAO repository 120. The data augmentation system 100 can pre-select a set of data augmentation operators available to the DAO generator.
[0133]
[0150] In step 1430, the sequence-to-sequence model 220 can apply a data augmentation operator to the selected token sequence to generate a transformed token sequence. The transformed token sequence can include updated tokens, token spans. The transformation can depend on the data augmentation operator applied to the token sequence selected in step 1410. When step 1430 is completed, the sequence-to-sequence model 220 completes executing method 1400 on computing device 500 (step 1499).
[0134]
[0151] FIG. 15 is a flowchart showing an exemplary method for extracting classification information from an input sequence (e.g., text classification, entity matching, error detection) for a sequence classification task according to an embodiment of the present disclosure. The steps of method 1500 can be executed on computing device 500 of FIG. 5 for illustrative purposes, or otherwise by a system using its features (e.g., data augmentation system 100 of FIG. 1). It is understood that the illustrated method 1500 can be modified to change the order of steps and include additional steps.
[0135]
[0152] In step 1510, the system can generate augmented data (augmented data 412) using at least one inverse data augmentation operator (e.g., inverse data augmentation operators 713-715 of FIG. 7).
[0136]
[0153] In step 1520, the system can pre-train a machine learning model (e.g., target model 141) using the augmented data as training data. Training a machine learning model using augmented data is presented above in connection with FIG. 4A and its corresponding description.
[0137]
[0154] In step 1530, the system can add task-specific layers (e.g., linear layers 1031-1033, softmax layers 1041-1043) to the machine learning model (e.g., target model 141) to generate a modified network 1000 (as shown in FIG. 10).
[0138]
[0155] In step 1540, the system can initialize the modified network. The system can initialize the network by copying it to memory and running it on a computing device (e.g., computing device 500).
[0139]
[0156] In step 1550, the system can identify class tokens and other tokens within the input data entry. The system can use a machine learning model intended for natural language tasks for the identification tokens. For example, a sentiment analysis mining machine learning model can identify tokens by identifying the topic within a text example and the opinion sentences taking up the topic.
[0140]
[0157] In step 1560, the system can mark class tokens and other tokens using different markers that represent the start and end of each token. The data augmentation system can mark various tokens within the input sequence using markers such as "[CLS]" and "[SEP]". In some embodiments, only the start of a token can be identified using a marker that can serve the role of the end marker of the previous token.
[0141]
[0158] In step 1570, the system can serialize the input data entry using a sequence-to-sequence data augmentation model.
[0142]
[0159] In step 1580, the system can provide the serialized input data entry to the modified network. The system can associate a reference with a classification layer (e.g., linear layers 1031-1032, softmax layers 1041-1043) that can extract relevant classification information from the serialized input data. The associated reference can be based on a request from a user (e.g., user 190) or the content of the input data entry. For example, if the input data entry has an intended question structure, then the system can associate a reference to the opinion extraction and sentiment analysis classification layer with the input data entry.
[0143]
[0160] In step 1590, the system can extract classification information using the task-specific layer of the modified network. Further explanation of the extraction of classification information is presented above in connection with FIG. 10 and its corresponding description. When step 1590 is completed, the system completes executing method 1500 on computing device 500 (step 1599).
[0144]
[0161] Exemplary embodiments have been described above with reference to flowchart diagrams or block diagrams of methods, apparatus (systems), and computer program products. It will be understood that each block of the flowchart diagrams or block diagrams, and combinations of blocks in the flowchart diagrams or block diagrams, can be implemented by computer program products, or instructions on a computer program product. These computer program instructions can be provided to a processor of a computer or other programmable data processing apparatus to produce means for causing the instructions executed via the processor of the computer or other programmable data processing apparatus to implement the functions / acts specified in the block or blocks of the flowchart or block diagram.
[0145]
[0162] These computer program instructions can also be stored in a computer-readable medium that causes a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer-readable medium form a manufacture including instructions for implementing the functions / acts specified in the blocks or groups of blocks of a flowchart or block diagram.
[0146]
[0163] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / acts specified in the blocks or groups of blocks of a flowchart or block diagram, causing a series of operational steps to be performed on the computer, other programmable apparatus, or other device, thereby creating a computer-implemented process.
[0147]
[0164] Any combination of one or more computer-readable media (singular or plural) can be utilized. The computer-readable media can be a non-transitory computer-readable storage medium. In the context of this specification, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0148]
[0165] The program code incorporated on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, fiber optic cable, RF, IR, etc., or any suitable combination of the foregoing.
[0149]
[0166] The calculation, for example, the computer program code for implementing the embodiments, can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or the like. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, executed partially on the user's computer and partially on a remote computer, or executed entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection to an external computer can be made (e.g., through the Internet using an Internet service provider). The computer program code can be compiled into object code executable by a processor, or partially compiled into intermediate object code, or interpreted in an interpreter, just-in-time compiler, or virtual machine environment intended to execute the computer program code.
[0150]
[0167] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in a flowchart or block diagram can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function(s). Note also that in some alternative implementations, the functions noted in a block can occur out of the order noted in the drawings. For example, two blocks shown in succession can, in fact, be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order, depending on the functionality involved. Also note that each block of the block diagrams or flowchart diagrams, as well as combinations of blocks in the block diagrams or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or combinations of dedicated hardware and computer instructions.
[0151]
[0168] The above-described embodiments are not mutually exclusive, and it is understood that elements, components, materials, or steps described with respect to one exemplary embodiment can be combined with or excluded from other embodiments in a manner suitable for achieving the desired design objective.
[0152]
[0169] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Specific adaptations and modifications of the above-described embodiments can be made. From the consideration of this specification and the practice of the invention disclosed herein, other embodiments may become apparent to those skilled in the art. This specification and examples are intended to be considered only as exemplary. It is also intended that the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to be limited to any particular arrangement of steps. Therefore, those skilled in the art can understand that these steps can be performed in a different order while implementing the same method.
Description of Reference Signs
[0153] 100 Data Augmentation System 110 Data Augmentation Operator Generator 120 Data Augmentation Operator Repository 130 Unlabeled Data Repository 140 Machine Learning Model Repository 141 Target Model 150 Machine Learning Model Platform 160 Text Corpus Repository 170 Meta - learning Policy Framework 180, 1000 Network 220 Sequence - to - Sequence Model 231 Unlabeled Data 232 Damaged Unlabeled Data 250, 260, 280 Tables 271, 272, 291 Data 281 Row 371 Filtering Model 372 Weighting Model 373 Loss Function 411 Training Data 412 Augmented Data 413 Target Batch 414 Validation Data 421 Data Augmentation Operator 432 Augmented Unlabeled Data 441 Updated Target Model 470 Policy Model 481 Soft Label 500 Computing Device 710, 810, 910 Original Sequence 711, 712 Basic Data Augmentation Operators 713 - 715 Inverse Data Augmentation Operators 720 Changed Sequence 721, 722, 723, 820, 821, 822, 823, 920, 921, 922, 923 Augmented Sequences 1020 Input sequence 1030, 1040 Task-specific layer
Claims
1. A non - transitory computer - readable storage medium storing instructions executable to cause a data augmentation system including one or more processors to perform a method for generating training data for a machine learning model, the method comprising: accessing a machine learning model from a machine learning model repository; identifying a dataset associated with the machine learning model; generating a set of data augmentation operators using the dataset; selecting a token sequence associated with the machine learning model; generating at least one token sequence by applying at least one data augmentation operator of the set of data augmentation operators to the selected token sequence; selecting a subset of token sequences from the at least one generated token sequence; storing the subset of token sequences in a training data repository; providing the subset of token sequences to the machine learning model; A non - transitory computer - readable storage medium including the above.
2. Generating a set of data augmentation operators using the dataset further includes: selecting one or more data augmentation operators; generating input token sequences formatted sequentially for the identified dataset; applying the one or more data augmentation operators to at least one token sequence of the sequentially formatted input token sequences to generate at least one modified token sequence; determining the set of augmentation operators for reverting the at least one modified token sequence back to the corresponding sequentially formatted input token sequence; The non - transitory computer - readable storage medium according to claim 1, further including the above.
3. The non - transitory computer - readable storage medium according to claim 1, wherein the accessed machine learning model is a sequence - to - sequence machine learning model.
4. Selecting a subset of token sequences includes: filtering at least one token sequence from the at least one generated token sequence using a filtering machine learning model; Using a weighted machine learning model to determine the weight of at least one sequence token within the at least one filtered token sequence; Applying the weight to at least one token sequence among the at least one filtered token sequences; The non - transient computer - readable storage medium according to claim 1, further comprising.
5. The non - transient computer - readable storage medium according to claim 3, wherein the weight of the at least one token sequence is determined based on the importance of the token sequence when training the machine learning model.
6. The non - transient computer - readable storage medium according to claim 5, wherein the importance of the at least one token sequence is determined by calculating the validation loss of the machine learning model when trained using the at least one token sequence.
7. The non - transient computer - readable storage medium according to claim 6, wherein the filtering machine learning model is trained using the validation loss.
8. The non - transient computer - readable storage medium according to claim 6, wherein the weighted machine learning model is trained using the validation loss.
9. The non - transient computer - readable storage medium according to claim 8, wherein the weighted machine learning model is trained until the validation loss reaches a threshold.
10. The non - transient computer - readable storage medium according to claim 1, wherein the at least one data augmentation operator includes at least one of a token deletion operator, a token insertion operator, a token replacement operator, a token swap operator, a span deletion operator, a span shuffle operator, a column shuffle operator, a column deletion operator, an entity swap operator, an inverse translation operator, a class generator operator, an inverse data augmentation operator.
11. The non - transient computer - readable storage medium according to claim 10, wherein the inverse data augmentation operator is a combination of a plurality of data augmentation operators.
12. The non - transient computer - readable storage medium according to claim 1, wherein the at least one data augmentation operator is context - dependent.
13. Providing the subset of the token sequence as an input to the machine learning model is Accessing unlabeled data from an unlabeled data repository; Generate an extended label-free token sequence of the accessed label-free data and determine a soft label for the extended label-free token sequence and provide the extended label-free token sequence along with the associated soft label as an input to the machine learning model The non-transitory computer-readable storage medium according to claim 1, further comprising **Claim 14** A non-transitory computer-readable storage medium storing instructions executable to cause a data augmentation system including one or more processors to perform a method for generating a data augmentation operator for generating an augmented token sequence, the method comprising accessing label-free data from a label-free data repository preparing one or more token sequences of the accessed label-free data transforming the one or more token sequences to generate at least one corrupted sequence providing the one or more token sequences and the at least one corrupted sequence as inputs to a sequence-to-sequence model of the data augmentation system executing the sequence-to-sequence model to determine at least one operation required to revert the at least one corrupted sequence to the sequence within the one or more token sequences used to generate the at least one corrupted sequence generating an inverse data augmentation operator based on the determined one or more operations for reverting the at least one corrupted sequence The non-transitory computer-readable storage medium comprising **Claim 15** Preparing the one or more token sequences of the accessed label-free data is converting each row in a database table into a token sequence, the token sequence including indicators for the start and end of column values, and further comprising converting, the non-transitory computer-readable storage medium according to claim 14 **Claim 16** Transforming the one or more token sequences to generate at least one corrupted sequence is selecting a token sequence from the one or more token sequences selecting a data augmentation operator from a set of data augmentation operators applying the data augmentation operator to the selected token sequence; The non - transient computer - readable storage medium according to claim 14, further comprising. **Claim 17** generating at least one damaged sequence; The non - transient computer - readable storage medium according to claim 16, further comprising generating a set of damaged sequences by sequentially applying a plurality of data augmentation operators to the selected token sequence. **Claim 18** A non - transient computer - readable storage medium storing instructions executable to cause a data augmentation system including one or more processors to perform a method for extracting classification information from input data, the method comprising: adding a task - specific layer to a first machine learning model to generate a modified network; initializing the modified network of the first machine learning model and the added task - specific layer; selecting an input data entry, the selecting including serializing the input data entry; providing the serialized input data entry to the modified network; extracting classification information using the task - specific layer of the modified network; converting the input data entry by applying the same data augmentation operator to the serialized input data entry to generate different augmented data entries, wherein the operations in generating the different augmented data entries are determined based on tasks related to different second machine learning models; filtering the generated different augmented data entries based on the second machine learning model; training the first machine learning model using the filtered augmented data entries; The non - transient computer - readable storage medium comprising. **Claim 19** wherein the method further comprises: generating augmented data using at least one inverse data augmentation operator; pre - training the first machine learning model using the augmented data; The non - transient computer - readable storage medium according to claim 18, further comprising. **Claim 20** wherein serializing the input data entry comprises: identifying class tokens and other tokens within the input data entry; marking the class token and the other tokens using different markers that represent the start and end of each token; The non-transitory computer-readable storage medium according to claim 18, further comprising.
Citation Information
Patent Citations
Data augmentation in training deep neural network (DNN) based on genetic model
JP2020187734A
Learning data augmentation policies
WO2019222734A1