Method for augmenting natural-language data by using generative adversarial learning

By employing generative adversarial learning to create an embedding space for natural language, the system addresses the challenges of backpropagation and data memorization in existing models, enabling the generation of new sentences without prior training and maintaining model efficiency.

WO2025127397A1PCT designated stage expired Publication Date: 2025-06-19I-BRICKS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016936
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-10-31
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing natural language generation models face challenges with backpropagation due to discrete token generation and suffer from data memorization, limiting their ability to synthesize new sentences.

Method used

The approach involves using generative adversarial learning to create an embedding space for natural language, allowing for unsupervised learning and generation of new sentences without prior training, using a system comprising an embedding space interpretation model, a generator module, and discriminator modules.

Benefits of technology

This method enables the generation of completely new sentences using only noise, overcoming the limitations of backpropagation and data memorization, while maintaining a relatively lightweight model for efficient sentence synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016936_19062025_PF_FP_ABST
    Figure KR2024016936_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a system and method that applies a generative adversarial learning algorithm for the augmentation and synthesis of natural-language sentence data. The present invention proposes a method for synthesizing, without referencing training data, completely new natural-language text not in the training data, through unsupervised learning. Furthermore, to perform unsupervised learning, the present invention generates an embedding space for natural language rather than generating conventional discrete tokens, thereby addressing the issue of data memorization while also solving the problem that backpropagation has previously been impossible.
Need to check novelty before this filing date? Find Prior Art

Description

Natural language data augmentation using generative adversarial learning

[0001] The present invention relates to natural language processing technology.

[0002] Typically, natural language generation uses autoregressive methods, a type of supervised learning, to generate discrete natural language tokens. Existing natural language sentence generation and synthesis algorithms generate discrete tokens, making backpropagation using deep learning impossible. Therefore, pre-training with training data was necessary for the generator model directly involved in sentence generation. Consequently, the problem of memorizing data, which reproduces existing natural language exactly when synthesizing actual sentences, arises. Models that memorize and reproduce the same data in this way cannot be considered to synthesize new natural language sentences, and therefore pose a critical problem for generative models.

[0003] To address these issues, the inventors of the present invention adopted an unsupervised learning approach, rather than supervised learning, a machine learning method for inferring a function from training data. In other words, the inventors of the present invention discussed and collaborated to propose a method for synthesizing completely new natural language not found in the training data, without referencing training data through unsupervised learning. The present invention, instead of generating discrete tokens to perform unsupervised learning, generates an embedding space for natural language, thereby resolving the problem of the impossibility of conventional backpropagation and simultaneously solving the problem of data memorization.

[0004] As mentioned above, conventional natural language generation and synthesis models require pre-training of the generator model because they generate discrete tokens. This process makes backpropagation impossible for deep learning model training, and the problem of data memorization arises by reproducing the same sentences. The purpose of the present invention is to solve the problem of data memorization by generating an embedding space of natural language sentences instead of generating discrete tokens. Furthermore, the problem of data memorization is solved by synthesizing sentences through unsupervised learning.

[0005] Another purpose of the present invention is to enable generation of completely new sentences using only noise without prior training of a generator model through model learning as described above.

[0006] Meanwhile, other unspecified purposes of the present invention will be additionally considered within the scope that can be easily inferred from the detailed description and effects thereof below.

[0007] The first aspect of the present invention to achieve the above-mentioned task is a natural language data augmentation system using adversarial generative learning:

[0008] A natural language data augmentation system device comprising a new natural language sentence generator configured as a computer and executed by a processor; and

[0009] An online network terminal that communicates with the above natural language data augmentation system device and receives sentences generated by the new natural language sentence generator,

[0010] The above novel natural language sentence generator includes an embedding space interpretation model, a generator module, a first discriminator module, and a second discriminator module which is a language model different from the first discriminator module.

[0011] The above embedding space interpretation model is a model that predicts words that follow actual data and is trained using actual natural language sentences of a dataset, and the generator module continuously generates a fake embedding space, and the first discriminator module and the second discriminator module continuously determine the fake embedding space generated by the generator module based on the grammatical meaning of the natural language sentence so as to prevent the fake embedding space generated by the generator module from passing through the embedding space interpretation model, and when the first discriminator module and the second discriminator module determine that the embedding space generated by the generator module is real, the embedding space is passed through the embedding space interpretation model only for the corresponding embedding space and output as a generated sentence.

[0012] In addition, in a natural language data augmentation system using adversarial generative learning according to a preferred embodiment of the present invention, it is preferable that the actual natural language sentence for training the embedding space interpretation model is a multi-turn natural language conversation sentence.

[0013] In addition, in a natural language data augmentation system using adversarial generative learning according to a preferred embodiment of the present invention, it is preferable that the first discriminator module uses a BERT (Bidirectional Encoder Representations from Transformers) model, and the second discriminator module uses an LSTM (Long Short Time Memory) model.

[0014] A second aspect of the present invention is a natural language data augmentation method using adversarial generative learning, wherein a new natural language sentence generator executed by a processor:

[0015] We train an embedding space interpretation model that predicts words that follow real data using natural language sentences.

[0016] Including a step of generating and outputting a new natural language sentence from the embedding space generated by the generator module only when the above embedding space interpretation model is passed.

[0017] The generator module continuously generates a fake embedding space, and the first discriminator module and the second discriminator module using a different language model from the first discriminator module continuously determine the fake embedding space generated by the generator module based on the grammatical meaning of a natural language sentence, so that the fake embedding space generated by the generator module does not pass the embedding space interpretation model, and when the first discriminator module and the second discriminator module determine that the embedding space generated by the generator module is real, it is preferable to pass the embedding space interpretation model only for the corresponding embedding space so as to output a generated sentence.

[0018] The present invention is a system that generates new sentences using only noise. It synthesizes diverse sentences, addresses the impossibility of backpropagation in existing deep learning algorithms, and resolves the problem of data memorization. While the model training process requires a relatively complex algorithm using adversarial learning, after training, it only requires a generator and an embedding space interpretation model, enabling the generation of diverse sentences with a relatively lightweight model.

[0019] Meanwhile, even if the effect is not explicitly mentioned herein, it is added that the effect and its provisional effect described in the following specification expected by the technical features of the present invention are treated as described in the specification of the present invention.

[0020] Figure 1 shows a schematic configuration of a natural language data augmentation system according to a preferred embodiment of the present invention.

[0021] FIG. 2 schematically illustrates the configuration of a new natural language sentence generator (101) included in a natural language augmentation system device (100) in a preferred embodiment of the present invention.

[0022] Figure 3 schematically illustrates the entire process of a natural language data augmentation method using a new natural language sentence generator (101) as described above.

[0023] Figure 4 schematically illustrates the learning process of an embedding space interpretation model according to a preferred embodiment of the present invention.

[0024] FIG. 5 illustrates an antagonistic relationship between a generator module (120) and two discriminator modules (130, 140) according to an embodiment of the present invention.

[0025] Figure 6 illustrates steps S130 and S140 of Figure 3 in more detail.

[0026] It is to be understood that the attached drawings are provided for reference only to help understand the technical concept of the present invention, and the scope of the present invention is not limited thereby.

[0027] Hereinafter, the present invention will be described in detail with reference to the drawings illustrating the configuration of the present invention. In describing the present invention, detailed descriptions of related known functions that are obvious to those skilled in the art and that may unnecessarily obscure the gist of the present invention will be omitted.

[0028]

[0029] Figure 1 shows a schematic configuration of a natural language data augmentation system according to a preferred embodiment of the present invention.

[0030] The natural language data augmentation system device of the present invention is executed by a processor (10). The processor does not simply refer to a CPU or GPU. It can also represent the functionality of a server device constructed with one or more computer devices. Specifically, the processor (10) executes a novel natural language sentence generator described below, which processes the methods of the present invention by executing data and program commands residing in a memory module.

[0031] The input unit (13) inputs data required for executing the language model of the processor into the processor (10). Then, the output unit (14) outputs a sentence generated by a new natural language sentence generator.

[0032] The storage unit (15) may consist of one or more memories and databases. The storage unit (15) stores resources and data required for the language model. It may also be an external storage unit accessed by the process (10) via communication.

[0033] The language model of the natural language data augmentation system device of the present invention generates and predicts new words in response to input sentences. This effectively augments natural language data and can be provided as a service. The online network terminal (1) communicates with the natural language data augmentation system and receives sentences generated by the new natural language sentence generator.

[0034] Figure 2 schematically illustrates the configuration of a new natural language sentence generator (101) included in the natural language augmentation system device (100) of the present invention.

[0035] The novel natural language sentence generator (101) of the present invention includes an embedding space interpretation module (110) and an adversarial generative learning model. A generator module (120), a first discriminator module (130), and a second discriminator module (140) constitute the adversarial generative learning model.

[0036] The embedding space interpretation model (110) is a model that predicts words that will follow actual data and learns using actual natural language sentences in the dataset. Actual sentences cannot be learned by a computer. They must be converted into vectors, which are numerical representations that a computer can read, and such embedding vectors are generated on a word-by-word basis. The embedding space refers to a set of embedding vectors. Generally, words with similar meanings have close distances between their embedding vectors, and the more dissimilar the words, the further the distances between their embedding vectors become. The language model is designed so that the embedding space of words generated by adversarial learning executed between the generator module (120) and the discriminator modules (130, 140) must pass through the embedding space interpretation model (110) to be able to generate sentences.

[0037] The generator module (120) continuously generates a fake embedding space. In other words, it creates fake data that follows the presented sentence. If the fake data cannot pass the embedding space interpretation model (110), the classification error compared to the real data is large and the probability of it being real data is low, so it is merely noise. However, even if it was initially noise, if the classification error is reduced through the discrimination process of the discriminator module (130, 140), i.e., if the discriminator module (130, 140) is unable to distinguish between fake and real sentences, it is recognized as real data.

[0038] The discriminator module (130, 140) continuously determines whether the data generated by the generator module (120) is real or fake. It retrieves real data from the training dataset and classifies the fake data generated by the generator module (120). It calculates the error in the embedding space and blocks fake data from passing the embedding space interpretation model.

[0039] In a preferred embodiment, the discriminator modules (130, 140) of the present invention use different language models. The first discriminator module (130) uses a BERT (Bidirectional Encoder Representations from Transformers) model, and the second discriminator module (130) uses an LSTM (Long Short Time Memory) model, specifically a bidirectional LSTM model. The BERT model is known to those skilled in the art as a language model created by stacking the encoder part of the transformer model architecture. The LSTM model is an improved model that considers past data as well as recent data to predict future data in order to solve the long-term dependency problem in which learning becomes difficult as the input length of an RNN (Recurrent Neural Network) increases. It is known to those skilled in the art as a language model that adds memory cells, forget cells, etc. In order to configure the discrimination of the discriminator module more precisely, different language models are used. The BERT model is advantageous in capturing the structural features of sentences. The LSTM model was advantageous in considering the order of words within a sentence. Meanwhile, the distinction between "first" and "second" was used to indicate that they were different models.

[0040]

[0041] Figure 3 schematically illustrates the entire process of a natural language data augmentation method using a new natural language sentence generator (101) as described above.

[0042] This is a process of a step (S100) in which a user terminal inputs a natural language sentence and a step (140) in which a new natural language sentence generator generates and outputs a subsequent sentence.

[0043] The new natural language sentence generator (101) first trains an embedding space interpretation model (S110).

[0044] It's best to learn using actual natural language sentences. Multi-turn natural language conversational sentences are even better. A multi-turn natural language conversational sentence refers to a natural language conversational sentence with three or more exchanges. The following conversation is a single-turn conversational sentence.

[0045] A: It's so cold today.

[0046] B: That's right.

[0047] The following dialogue is a multi-turn dialogue sentence.

[0048] A: Please give me a cup of warm coffee.

[0049] B: How many spoons of prim should I put in?

[0050] A: Two spoons.

[0051] P: Can I have a cup of coffee?

[0052] Q: Here it is.

[0053] P: Thank you.

[0054]

[0055] Figure 4 outlines the learning process of the embedding space interpretation model. First, the multi-turn natural language conversation sentence described above is input into the model (S200).

[0056] A new natural language sentence generator (101) converts multi-turn natural language sentences into a computer-interpretable embedding space (S210). The embedding space has been described above.

[0057] Next, an embedding space interpretation model is created (S220). It is preferable to use a model that provides the BLEU (BiLingual Evaluation Understudy) index for this task. In this invention, GPT-2 was used. The BLEU index is an index that evaluates the quality of text translated from one language to another, and has a value between 0 and 1. A higher index indicates better performance. The index is calculated by calculating the degree of overlap between the actual correct answer and the translated result and dividing it by the total number of correct words. This model is used solely for interpreting the embedding space.

[0058] Embedding space interpretation models learn to predict the next word after a set of words. For example, if the sentence "Please give me a cup of coffee," they learn a rule that leads to the sentence "Here it is. Thank you."

[0059] Next, an adversarial generative model is trained (S120). This is performed on both the generator and discriminator training phases. The generator module is based on a one-dimensional convolutional neural network (CNN) and is used to create a fake embedding space. The discriminator module uses two different language models, as described above. Step S120 is described in more detail in Figure 5.

[0060] Figure 5 illustrates the antagonistic relationship between a generator module (120) and two discriminator modules (130, 140).

[0061] The discriminator module (130, 140) continuously performs the task of distinguishing between the real embedding space and the fake embedding space by using the real embedding space (113) extracted from the real sentence (103) and the fake embedding space (115) generated by the generator module (120).

[0062] The 1D-CNN-based generator module (120) continuously performs a fake embedding space to deceive the discriminator module (130, 140). In the early stages of learning, the generator module (120) generates strange words, but as learning progresses, it produces sentences that resemble real ones.

[0063] The main goal of this learning of the present invention is to produce a good generator module (120). During the learning process, the discriminator module (130, 140) learns to distinguish between real and fake sentences, and at the same time, the generator module (120) learns to deceive the discriminator module (130, 140). Therefore, ideally, the discriminator module (130, 140) ultimately becomes unable to distinguish between real (embedding space of actual sentences) and fake (embedding space of sentences generated by the generator).

[0064] The goal of the above adversarial learning is to enable the generator to create a plausible embedding space. First, two discriminator modules (130, 140) determine whether the embedding space extracted from the sentence is a real embedding space or a fake embedding space generated by the generator. The two discriminator loss functions (L ) are expressed in Equation 1 below. D ) is used.

[0065]

[0066] (Equation 1)

[0067]

[0068] To help the generator converge, two additional loss functions are used. The first loss function is L in Equation 2. SDPIt is expressed as , and uses the Kullback-Leibler Divergence function, which is used to calculate the difference between two probability distributions. In the mathematical formula, f_theta is the embedding space interpretation model, and sigma is the softmax function. The fake embedding space and the real embedding space are each passed through the embedding space interpretation model, and then the softmax function is applied to calculate the Kullback-Leibler divergence.

[0069]

[0070] (Equation 2)

[0071]

[0072] And the second loss function is L in Equation 3 SFP It is expressed as , which uses the mean square error and mean absolute error loss functions. H real -H fake The part is the mean absolute error, μ r -μ f The part refers to the mean square error. The former calculates the difference between the <average of the absolute values ​​of the differences between all word vectors> of real and fake data, respectively, and the latter calculates the average of the vectors of all words and then the sum of the squares of the vectors.

[0073]

[0074] (Equation 3)

[0075]

[0076] Adversarial learning is performed using two loss functions that assist the generator as shown above.

[0077] In this way, the discriminator module (130, 140) learns to identify real data as real and fake data as fake. The generator module (120) continuously generates fake data and learns to ultimately identify the fake data as real. As a result of this adversarial learning, the generator can ultimately generate a fake embedding space that is very similar to the real embedding space extracted from actual natural language.

[0078] Then, the process returns to step S130 of Figure 3. The fake embedding space created by the generator module passes through the embedding space interpretation model (S130). Then, the new natural language sentence generator generates and outputs new natural language sentences from the fake data passed through the embedding space interpretation model (S140).

[0079] Figure 6 illustrates steps S130 and S140 of Figure 3 in more detail.

[0080] The generator module completes learning through an adversarial learning process with the discriminator module (130, 140) (S300).

[0081] At this point, the fake embedding space generated by the trained generator module exceeds the pre-set discriminator module's classification error criteria for real and fake (S310) and passes the embedding space interpretation model (S320). The fake embedding space is then converted into natural language and output as a generated sentence (S330). This sentence is the newly synthesized sentence. Through this process, the novel natural language sentence generator of the present invention augments natural language data.

[0082]

[0083] For reference, the natural language data augmentation method using adversarial generative learning according to one embodiment of the present invention may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the medium may be those specifically designed and constructed for the present invention, or may be known and usable by those skilled in the art of computer software.

[0084] Examples of computer-readable media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter or the like. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the present invention, and vice versa.

[0085] Meanwhile, the scope of protection of the present invention is not limited to the description and expression of the embodiments explicitly described above. Various substitutions and modifications are possible within the scope of protection of the present invention. Furthermore, it should be noted that the scope of protection of the present invention may not be limited due to obvious modifications or substitutions made in the technical field to which the present invention pertains.

Claims

1. A natural language data augmentation system using adversarial generative learning: A natural language data augmentation system device comprising a new natural language sentence generator, the new natural language sentence generator comprising a computer and executed by a processor; and An online network terminal for communicating with the above natural language data augmentation system device and receiving sentences generated by the new natural language sentence generator, The above novel natural language sentence generator comprises an embedding space interpretation model, a generator module, a first discriminator module, and a second discriminator module which is a language model different from the first discriminator module. The above embedding space interpretation model is a model that predicts words that follow actual data and is learned using actual natural language sentences of the dataset, the generator module continuously generates a fake embedding space, the first discriminator module and the second discriminator module continuously determine the fake embedding space generated by the generator module based on the grammatical meaning of the natural language sentence so as to prevent the fake embedding space generated by the generator module from passing the embedding space interpretation model, and when the first discriminator module and the second discriminator module determine the embedding space generated by the generator module to be real, the embedding space interpretation model is passed only for the corresponding embedding space and output as a generated sentence, a natural language data augmentation system using generative adversarial learning.

2. In paragraph 1, A natural language data augmentation system using generative adversarial learning, wherein the actual natural language sentences that train the above embedding space interpretation model are multi-turn natural language sentences.

3. In paragraph 1, A natural language data augmentation system using generative adversarial learning, wherein the first discriminator module uses a BERT (Bidirectional Encoder Representations from Transformers) model, and the second discriminator module uses an LSTM (Long Short Time Memory) model.

4. A natural language data augmentation method using adversarial generative learning, wherein a new natural language sentence generator executed by a processor: We train an embedding space interpretation model that predicts words that follow real data using natural language sentences. Including a step of generating and outputting the embedding space generated by the generator module as a new natural language sentence only when the above embedding space interpretation model is passed, A natural language data augmentation method using generative adversarial learning, characterized in that the generator module continuously generates a fake embedding space, the first discriminator module and the second discriminator module using a different language model from the first discriminator module continuously determine the fake embedding space generated by the generator module based on the grammatical meaning of natural language sentences so as to prevent the fake embedding space generated by the generator module from passing the embedding space interpretation model, and when the first discriminator module and the second discriminator module determine that the embedding space generated by the generator module is real, only the corresponding embedding space is passed through the embedding space interpretation model so as to be output as a generated sentence.

5. In paragraph 4, A natural language data augmentation method using generative adversarial learning, wherein the natural language sentences for training the above-mentioned embedding space interpretation model are multi-turn natural language sentences.

6. In paragraph 4, A method for augmenting natural language data using generative adversarial learning, wherein the first discriminator module uses a BERT (Bidirectional Encoder Representations from Transformers) model, and the second discriminator module uses an LSTM (Long Short Time Memory) model.

Citation Information

Patent Citations

  • Text synthesis method and device and electronic equipment

    CN114549698A

  • Apparatus for guiding teeth

    KR1020240111231A

  • Carrot juice containing baranul jeju water

    KR1020240163255A

  • Method and apparatus for generating and editing images using contrasitive learning and generative adversarial network

    KR102477700B1