Learning Data Expansion Device
The learning data expansion device addresses the issue of word relevance in generating augmentation data by processing learning sentences with multiple paraphrasing degrees and using threshold values to determine sentence inclusion, resulting in improved data quality and relevance for learning models.
Patent Information
- Application Number
- JP2023561462
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-16
- Filing Date
- 2022-10-14
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2042-10-14
AI Technical Summary
Existing methods for generating augmentation data for learning data do not adequately consider the degree of relevance between words, leading to inappropriate inclusion of sentences in augmentation data.
A learning data expansion device that generates expanded sentences by processing learning sentences with multiple degrees of paraphrasing and determines the relevance of word pairs with syntactic relationships, using predetermined threshold values to decide whether to include these sentences in the augmentation data.
This approach allows for the appropriate generation of expansion data, considering the degree of relevance between words, thereby improving the quality and relevance of augmentation data for learning models.
Smart Images

Figure 0007689198000003 
Figure 0007689198000004 
Figure 0007689198000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to a learning data augmentation device that generates augmentation data for augmenting learning data.
Background Art
[0002] In recent years, the progress of artificial intelligence technologies such as deep learning has been remarkable. In particular, artificial intelligence that discovers certain rules from a large amount of data and realizes recognition and prediction is known. The capabilities of such artificial intelligence are determined by the quantity and quality of the learning data used to train the model. Therefore, for the purpose of augmenting learning data, a technique for generating augmentation data by processing with multiple degrees of augmentation for each augmentation method is described in Patent Document 1.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Assuming that the learning data includes a plurality of sentences expressed in a certain language, when generating the augmentation data, for example, whether a sentence obtained by replacing a word in a sentence included in the learning data with another word can be added to the augmentation data varies depending on the degree of association between the words. For example, consider the sentence "Explain the stage of a certain point service." included in the learning data. Regarding the sentence "Explain the stage of a certain point service." obtained by replacing the word "stage" in the sentence with another word "stage", in this case, since the degree of association between "point service" and "stage" is considered to be high, it is not desirable to replace "stage" with "stage". Therefore, it should be determined that the sentence rewritten as above should not be added to the augmentation data.
[0005] However, in the above Patent Document 1, the points to note in the expansion of the learning data as described above are not mentioned, and it is eagerly awaited to appropriately generate expansion data in consideration of the degree of relevance between words.
[0006] The present disclosure has been made to solve the above problems, and an object thereof is to appropriately generate expansion data in consideration of the degree of relevance between words.
Means for Solving the Problems
[0007] The learning data expansion device according to the present disclosure includes an expansion sentence generation unit that generates a plurality of expanded sentences by processing learning sentences included in pre-given learning data according to a plurality of degrees of paraphrasing, and derives the degree of relevance in a word pair having a syntactic relationship in each expanded sentence, and based on the comparison result between the obtained degree of relevance and a plurality of predetermined threshold values at multiple levels, determines for each level whether to add the expanded sentence to the expansion data for expanding the learning data, and an expansion data generation unit that generates expansion data at multiple levels with the expanded sentences determined to be added.
[0008] In the above learning data expansion device, the expansion sentence generation unit generates a plurality of expanded sentences by processing learning sentences included in pre-given learning data according to a plurality of degrees of paraphrasing, and the expansion data generation unit derives the degree of relevance in a word pair having a syntactic relationship in each generated expanded sentence, and based on the comparison result between the obtained degree of relevance and a plurality of predetermined threshold values at multiple levels, determines for each level whether to add the expanded sentence to the expansion data for expanding the learning data, and generates expansion data at multiple levels with the expanded sentences determined to be added. For example, when there is a word pair whose degree of relevance in a word pair having a syntactic relationship in the expanded sentence is equal to or lower than the threshold value at a certain level, it is determined that the expanded sentence is not added to the expansion data at that level, and the expansion data at that level is generated with the expanded sentences determined to be added. In this way, it is possible to appropriately generate expansion data in consideration of the degree of relevance between words.
Advantages of the Invention
[0009] According to the present disclosure, extended data can be appropriately generated in consideration of the degree of association between words.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0011] Hereinafter, an embodiment of a learning data expansion device according to the present disclosure will be described with reference to the drawings.
[0012] As shown in FIG. 1, the learning data expansion device 10 includes an extended text generation unit 11, an extended data generation unit 12, a model accuracy derivation unit 13, and a data determination unit 14. Hereinafter, the functions of each unit will be described. However, the detailed functions and processing contents will be described in detail later according to the flowchart of FIG. 2.
[0013] The extended text generation unit 11 is a functional unit that generates a plurality of extended texts by processing the learning texts included in the pre-provided learning data 20 according to a plurality of degrees of paraphrasing.
[0014] The extended data generation unit 12 derives the degree of relevance in the word pairs having a syntactic relationship in each generated extended text, and based on the comparison result between the obtained degree of relevance and a plurality of levels of predetermined thresholds, determines for each level whether to add the extended text to the extended data for expanding the learning data, and generates multiple levels of extended data A, B,... Z (hereinafter collectively referred to as "extended data 30") with the extended texts determined to be added.
[0015] The model accuracy derivation unit 13 is a functional unit that derives the accuracy of each model based on the accuracy of the output results obtained by inputting the pre-prepared test data 40 to each of the models (models A to Z, model 0, etc. in FIG. 1) obtained when learning the learning data 20 and the extended data 30 at each level together, or when learning only the learning data 20. More specifically, the model accuracy derivation unit 13 includes a learning unit 13A and a verification unit 13B. The learning unit 13A learns only the learning data 20, holds the obtained model 0, learns the learning data 20 and the extended data 30 at each level together, holds the models A to Z obtained therefrom, and learns the learning data 20 and all the extended data 30 together, and holds the obtained model ALL. The verification unit 13B inputs the test data 40 to each of the models 0, models A to Z, and model ALL held by the learning unit 13A to obtain the respective output results, and derives the accuracy of each model based on the accuracy of each obtained output result. Note that the learning unit 13A may learn the learning data 20 together with two or more of the extended data 30 at each level and hold the obtained models. As a representative example of such a pattern, in this embodiment, an example in which the learning unit 13A learns the learning data and all the extended data together and holds the obtained model ALL will be described.
[0016] The data determination unit 14 is a functional unit that determines, as the optimal extended data, the extended data at the stage where the accuracy of the model is higher than the accuracy when only the learning data 20 is used for learning (the accuracy of model 0) and is the highest accuracy. Note that various data such as the learning data 20, extended data 30, test data 40, and the PMI model 35 described later shown in FIG. 1 are stored in an arbitrary memory of the learning data extension device 10. However, it is not essential that it be inside the learning data extension device 10, and an external memory of the learning data extension device 10 may be used.
[0017] Next, the processing executed in the learning data extension device 10 will be described with reference to the flowchart of FIG. 2.
[0018] First, the extended sentence generation unit 11 receives the learning data 20 (step S1), and generates a plurality of extended sentences by processing the learning sentences included in the learning data 20 according to a plurality of paraphrasing degrees, for example, as shown in FIG. 3 (step S2). In the example of FIG. 3, extended strengths 1 to 5 corresponding to a plurality (for example, five levels) of paraphrasing degrees are predetermined. For example, the extended strength 5 means "no words that are not paraphrased", and it is the stage with the highest paraphrasing degree where all paraphrasable words are paraphrased. The extended strength 4 is the stage where "only proper nouns are not paraphrased", the extended strength 3 is the stage where "only proper nouns and common nouns are not paraphrased", and so on, and they are set so that the paraphrasing degree gradually decreases.
[0019] According to the above-described expansion intensity, an example of generating a plurality of expanded sentences is explained by processing a learning sentence "Please tell me about the stage of the d-point club (registered trademark)", which includes the proper noun "d-point club (registered trademark)" and the common noun "stage", at each of the expansion intensities of 3 to 5. As shown in FIG. 3, in the expansion with an expansion intensity of 5, since there are "no words that are not paraphrased", the proper noun "d-point club" is paraphrased into "d-point section", and the common noun "stage" is paraphrased into "stage", respectively, and the expanded sentence "Please tell me about the stage of the d-point section" is generated. In the expansion with an expansion intensity of 4, since "only proper nouns are not paraphrased", the proper noun "d-point club" is not paraphrased, and the common noun "stage" is paraphrased into "stage", and the expanded sentence "Please tell me about the stage of the d-point club" is generated. Further, in the expansion with an expansion intensity of 3, since "only proper nouns and common nouns are not paraphrased", neither the proper noun "d-point club" nor the common noun "stage" is paraphrased, and as a result, the same expanded sentence as before the expansion, "Please tell me about the stage of the d-point club", is generated.
[0020] FIG. 5 shows the flow of expanded sentence generation and expanded data generation. As shown in FIG. 5, the expanded sentence generation unit 11 generates, by step S2 in FIG. 2, an expanded sentence group 25A consisting of a plurality of expanded sentences expanded at an expansion intensity of 1, an expanded sentence group 25B consisting of a plurality of expanded sentences expanded at an expansion intensity of 2, an expanded sentence group 25C consisting of a plurality of expanded sentences expanded at an expansion intensity of 3, an expanded sentence group 25D consisting of a plurality of expanded sentences expanded at an expansion intensity of 4, and an expanded sentence group 25E consisting of a plurality of expanded sentences expanded at an expansion intensity of 5 (these are collectively referred to as the "expanded sentence group 25"), and transfers them to the expanded data generation unit 12.
[0021] Returning to FIG. 2, in the next step S3, the extended data generation unit 12 derives the degree of relevance in the word pairs having a syntactic relationship in each extended sentence, and based on the comparison result between the degree of relevance and multiple levels of threshold values, determines for each level whether to add the extended sentence to the extended data, and generates multiple levels of extended data with the extended sentences determined to be added. Here, for example, as shown in FIG. 4(a), the extended data generation unit 12 extracts word pairs consisting of proper nouns and any of nouns, adjectives, and verbs having a syntactic relationship with the proper nouns from the extended sentence group 25 (step S31), and derives the degree of relevance of the extracted word pairs (step S32). Here, for example, point-wise mutual information (hereinafter referred to as "PMI") is used as the "degree of relevance", and by inputting the word pair into the PMI model 35 obtained in advance by machine learning using articles of FAQ (Frequently Asked Questions), Internet encyclopedia sites (Wikipedia (registered trademark)), etc., the degree of relevance of the word pair is derived as its output. Further, the extended data generation unit 12 compares the derived degree of relevance with each of the multiple levels of threshold values (step S33). Here, for example, if there is a word pair in the extended sentence whose degree of relevance is equal to or lower than the threshold value of a certain level, it is determined not to add the extended sentence to the extended data of that level. Therefore, if there is a word pair whose degree of relevance is equal to or lower than the threshold value of a certain level, the extended sentence is not added to the extended data of that level, while if there is no word pair whose degree of relevance is equal to or lower than the threshold value of a certain level, the extended sentence is added to the extended data of that level.
[0022] The PMI used as the "degree of relevance" in step S33 above is a measure indicating the degree of relevance between word pairs (two words), and the PMI(x, y) of a word pair consisting of word x and word y is defined by the following formula.
Equation
[0023] In addition, as the "threshold value" used in step S33 above, for example, "threshold value 1", "threshold value 2",... "threshold value 5" with a division number of 5 and values set to the values shown in Fig. 4(b) are used. Also, as shown in Fig. 4(c), when using "threshold value 3 (value is "0")" in Fig. 4(b) as the threshold value, assuming that a relevance of "-0.3" is derived for a word pair consisting of the proper noun "d-point club" and the noun "stage" that has a dependency relationship with the proper noun, in the extended sentence to be judged, since there is a word pair (d-point club, stage) whose relevance is below the threshold value, it is determined that the extended sentence is not added.
[0024] By the processing of step S3 in Fig. 2 as described above, as shown in Fig. 5, for each of the extended sentence groups 25A, 25B,..., 25E, "judgment on whether to add an extended sentence" using five levels of threshold values 1 to 5 is performed, and a total of 25 types of extended data 30 are generated. For the sake of convenience, Fig. 5 shows a part of the above 25 types of extended data 30, that is, extended data 30A generated using extension strength 1 and threshold value 1, extended data 30B generated using extension strength 1 and threshold value 2, extended data 30L generated using extension strength 2 and threshold value 1, extended data 30M generated using extension strength 2 and threshold value 2, extended data 30Y generated using extension strength 5 and threshold value 4, and extended data 30Z generated using extension strength 5 and threshold value 5.
[0025] Return to FIG. 2. In the next step S4, the model accuracy derivation unit 13 derives the accuracy of the model based on the accuracy of the output results obtained by inputting the previously prepared test data 40 to each of the models obtained when (1) only the training data 20 is used for training, (2) the training data 20 and the extended data 30 at each stage are combined for training, and (3) the training data 20 and all the extended data 30 are combined for training (step S4). More specifically, for example, as shown in FIG. 6, the model accuracy derivation unit 13 first obtains model 0 by training only the training data 20, and combines the training data 20 with each stage of the extended data 30 (that is, combines the training data 20 with extended data A, combines the training data 20 with extended data B,..., combines the training data 20 with extended data Z) to obtain models A to Z. Then, the training data 20 and all the extended data 30 are combined for training to obtain model ALL. Then, the model accuracy derivation unit 13 inputs the test data 40 to each of the obtained models 0, model ALL, and models A to Z to obtain the output results for each model, and based on the accuracy of the output results for each obtained model, derives the accuracy of the model as follows, for example.
[0026] Here, as shown in FIG. 7, when assuming a text classification model for classifying whether a movie review article has positive (affirmative) content or negative (negative) content, the score of the classification accuracy is derived as the accuracy of the model (denoted as "score" in FIGS. 6 and 7). By inputting the "text" in the test data 40 into the text classification model obtained by training the combined learning data 20 and extended data 30 shown in FIG. 7, for the test sentence "I was touched by the last scene.", an estimated result of "negative" is obtained, and for the test sentence "It was not worth watching.", an estimated result of "negative" is also obtained. In this case, when collating with the correct category of the test data, although the estimated result for the test sentence "It was not worth watching." is correct, the estimated result for the test sentence "I was touched by the last scene." is incorrect. Therefore, out of the two classification times, there is one correct answer, and a classification accuracy score of 0.5 is derived as the accuracy of the model.
[0027] Returning to FIG. 2, in the next step S5, the data determination unit 14 determines the extended data at the stage where the accuracy of the model is higher than the accuracy when only the learning data is learned and is the highest accuracy as the optimal extended data (step S5). In the example of FIG. 6, the score of the extended data M, which is "0.7", is higher than the accuracy "0.35" when only the learning data is learned and is the highest accuracy. Therefore, the extended data M is determined as the optimal extended data. Thereafter, for example, the data obtained by extending the learning data 20 with the extended data M determined as the optimal extended data is appropriately used as the data replacing the learning data 20 or as additional data for the learning data 20.
[0028] In step S5, when there are multiple pieces of extended data at a stage where the accuracy of the model is higher than the accuracy when only the training data 20 is used for training and is the highest, the corresponding multiple pieces of extended data may be determined as the optimal extended data, or one piece of extended data selected by an arbitrary method from the multiple pieces of extended data may be determined as the optimal extended data. Further, when the accuracy is the highest when only the training data 20 is used for training, it can be determined that there is no optimal extended data, so the determination of the optimal extended data is avoided.
[0029] According to the embodiment described above, by the processes of steps S2 and S3 in FIG. 2, multiple stages of extended data 30A, 30B,..., 30Z for selecting extended data with an appropriate degree of relevance can be appropriately generated in consideration of the degree of relevance between words of a word pair. Accordingly, when the training data includes a plurality of sentences expressed in a certain language, the cost for generating the extended data can be reduced, and the change in the degree of word extension according to the context can be flexibly accommodated.
[0030] Also, as the "degree of relevance" between words of a word pair, by using the pointwise mutual information (PMI) which has high reliability and is widely and generally used in language data, multiple stages of extended data 30A, 30B,..., 30Z can be appropriately generated based on an appropriate "degree of relevance".
[0031] Further, the extended data generation unit 12 sets, as a word pair having a dependency relationship, a word pair composed of a proper noun and any one of a noun, an adjective, and a verb having a dependency relationship with the proper noun as a target for deriving the degree of relevance, so that an appropriate degree of relevance can be derived while emphasizing the dependency relationship between the proper noun and other words.
[0032] In addition, for an extended sentence in which there are a plurality of word pairs having a dependency relationship, if there is a word pair among the plurality of word pairs whose degree of relevance is equal to or lower than a threshold value at a certain stage, the extended data generation unit 12 determines not to add the extended sentence to the extended data at that stage (step S33 in Fig. 4(a)). Thereby, if there is even one word pair whose degree of relevance is equal to or lower than the threshold value, addition of the extended sentence to the extended data at that stage can be avoided, and it is possible to avoid addition of inappropriate extended sentences from a safer side.
[0033] Alternatively, instead of the above, for an extended sentence in which there are a plurality of word pairs having a dependency relationship, the extended data generation unit 12 may determine not to add the extended sentence to the extended data at that stage when the degrees of relevance of all of the plurality of word pairs are equal to or lower than a threshold value at a certain stage. In this case, while avoiding addition of extended sentences in which the degrees of relevance of all included word pairs are equal to or lower than the threshold value, it is possible to actively promote addition of extended sentences to the extended data.
[0034] Furthermore, by the processes of steps S4 and S5 in Fig. 2, the model accuracy derivation unit 13 derives the accuracy (score) of each model based on the accuracy of the output results obtained by inputting the test data 40 to each of the models obtained when (1) only the learning data 20 is learned, (2) the learning data 20 and the extended data 30 at each stage are learned together, and (3) the learning data 20 and all of the extended data 30 are learned together, and determines the extended data at the stage where the accuracy of the model is higher than the accuracy when only the learning data is learned and is the highest accuracy as the optimal extended data. Thereby, it becomes possible to realize not merely extension of the learning data but extension of the learning data using the optimal extended data. At that time, it is not essential to target the model obtained in the case of (3) above, and even without making (3) essential, it is possible to realize extension of the learning data using the optimal extended data.
[0035] Of course, the model accuracy derivation unit 13 may derive the accuracy (score) of each model by using, as a further basis, the accuracy of the output result obtained by inputting the test data 40 to the model obtained when two or more of the extended data 30 at each stage are combined with the learning data 20 for learning. This includes the case of (3) above, and can target a large number of models obtained from further combination patterns such as "learning data 20 + extended data A + extended data B", "learning data 20 + extended data A + extended data M", "learning data 20 + extended data A + extended data Z", etc., and can determine the optimal extended data from a wider range of extended data candidates.
[0036] (Explanation of terms, explanation of hardware configuration (Fig. 8), etc.) Note that the block diagrams used in the description of the above embodiments and modifications show blocks of functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Also, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized by using one physically or logically combined device, or may be realized by directly or indirectly (for example, using wired, wireless, etc.) connecting two or more physically or logically separated devices and using these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.
[0037] Functions include, but are not limited to, judgment, decision-making, determination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, solution, selection, selection determination, establishment, comparison, assumption, expectation, presumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), assigning, etc. For example, a functional block (component) that enables transmission is called a transmitting unit or a transmitter. As described above, the implementation method is not particularly limited.
[0038] For example, the learning data augmentation device in an embodiment of the present disclosure may function as a computer that performs the processing in this embodiment. FIG. 8 is a diagram showing a hardware configuration example of a learning data augmentation device 10 according to an embodiment of the present disclosure. The above-described learning data augmentation device 10 may physically be configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
[0039] Note that in the following description, the term "device" can be read as a circuit, a device, a unit, etc. The hardware configuration of the learning data augmentation device 10 may be configured to include one or more of each device shown in the figure, or may be configured without including some of the devices.
[0040] Each function in the learning data augmentation device 10 is realized by causing the processor 1001 to perform operations by loading a predetermined software (program) onto hardware such as the processor 1001 and the memory 1002, and controlling communication by the communication device 1004, or controlling at least one of reading and writing data in the memory 1002 and the storage 1003.
[0041] Processor 1001 controls the entire computer by operating, for example, an operating system. Processor 1001 may be composed of a central processing unit (CPU: Central Processing Unit) including an interface with peripheral devices, a control device, an arithmetic device, registers, and the like.
[0042] Also, processor 1001 reads programs (program codes), software modules, data, etc. from at least one of storage 1003 and communication device 1004 into memory 1002 and executes various processes according to these. As the program, a program for causing a computer to execute at least a part of the operations described in the above-described embodiments is used. Although it has been described that the above-described various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. Processor 1001 may be implemented by one or more chips. Note that the program may be transmitted from a network via a telecommunication line.
[0043] Memory 1002 is a computer-readable recording medium and may be composed of, for example, at least one of ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), and the like. Memory 1002 may be referred to as a register, a cache, a main memory (main storage device), and the like. Memory 1002 can store programs (program codes), software modules, etc. executable for implementing the wireless communication method according to an embodiment of the present disclosure.
[0044] Storage 1003 is a computer-readable recording medium and may be composed of, for example, at least one of an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital versatile disc, a Blu-ray (registered trademark) disc), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-described storage medium may be, for example, a database including at least one of the memory 1002 and the storage 1003, or other appropriate media.
[0045] The communication device 1004 is hardware (a transmission / reception device) for performing communication between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc.
[0046] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives an external input. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that performs an output to the outside. Note that the input device 1005 and the output device 1006 may have an integrated configuration (e.g., a touch panel). Also, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus or may be configured using different buses for each device.
[0047] Each aspect / embodiment described in the present disclosure may be used alone, in combination, or switched for use during execution. Also, the notification of predetermined information (e.g., the notification of "being X") is not limited to being explicitly performed, and may be performed implicitly (e.g., by not performing the notification of the predetermined information).
[0048] As described above in detail, it is obvious to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented in the form of modifications and variations without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is for illustrative purposes only and does not have any limiting meaning for the present disclosure.
[0049] The processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in the present disclosure may be reordered as long as there is no contradiction. For example, regarding the methods described in the present disclosure, the elements of various steps are presented using an exemplary order and are not limited to the specific order presented.
[0050] The input / output information, etc. may be stored in a specific location (e.g., memory) or may be managed using a management table. The input / output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.
[0051] In the present disclosure, the description "based on" does not mean "only based on" unless otherwise specified. In other words, the description "based on" means both "only based on" and "at least based on".
[0052] In the present disclosure, when the terms "include", "including" and their variants are used, these terms are intended to be inclusive, similar to the term "comprising". Furthermore, the term "or" used in the present disclosure is not intended to be an exclusive disjunction.
[0053] In the present disclosure, for example, when articles are added by translation, such as a, an, and the in English, the present disclosure may include that the nouns following these articles are in the plural form.
[0054] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other". Note that the term may also mean "A and B are each different from C". Terms such as "separate" and "coupled" may also be interpreted in the same way as "different".
Description of Reference Numerals
[0055] 10…Learning data augmentation device, 11…Expanded text generation unit, 12…Expanded data generation unit, 13…Model accuracy derivation unit, 13A…Learning unit, 13B…Verification unit, 14…Data determination unit, 20…Learning data, 25…Expanded text group, 30…Expanded data, 35…PMI model, 40…Test data, 1001…Processor, 1002…Memory, 1003…Storage, 1004…Communication device, 1005…Input device, 1006…Output device, 1007…Bus.
Claims
1. An augmented sentence generation unit that generates a plurality of augmented sentences by processing sentences for learning included in pre-provided learning data according to a plurality of degrees of paraphrasing; An augmented data generation unit that derives the degree of association in word pairs having a syntactic relationship in each augmented sentence, and determines for each stage whether to add the augmented sentence to augmented data for augmenting the learning data based on the comparison result between the obtained degree of association and a plurality of pre-determined thresholds at multiple levels, and generates augmented data at multiple levels with the augmented sentences determined to be added; A learning data augmentation device comprising the above.
2. The learning data augmentation device according to Claim 1, wherein the augmented data generation unit uses point-wise mutual information as the degree of association.
3. The augmented data generation unit uses, as the word pairs having a syntactic relationship, a word pair consisting of a proper noun and any one of a noun, an adjective, and a verb having a syntactic relationship with the proper noun, as the object for deriving the degree of association. The learning data augmentation device according to Claim 1.
4. For an augmented sentence in which there are a plurality of word pairs having a syntactic relationship, the augmented data generation unit determines not to add the augmented sentence to the augmented data at that stage if there is a word pair among the plurality of word pairs whose degree of association is equal to or lower than a threshold at a certain stage. The learning data augmentation device according to Claim 1.
5. For an augmented sentence in which there are a plurality of word pairs having a syntactic relationship, the augmented data generation unit determines not to add the augmented sentence to the augmented data at that stage if all the degrees of association in the plurality of word pairs are equal to or lower than a threshold at a certain stage. The learning data augmentation device according to Claim 1.
6. A model accuracy derivation unit that derives the accuracy of the model based on the accuracy of the output results obtained by inputting predetermined test data to each of the models obtained when the learning data and the augmented data at each stage are combined for learning and when only the learning data is learned; A data determination unit that determines the augmented data at the stage where the accuracy of the model is higher than the accuracy when only the learning data is learned and is the highest accuracy as the optimal augmented data. The learning data augmentation device according to Claim 1, further comprising the above.
7. The model accuracy derivation unit Based on the accuracy of the output result obtained by inputting predetermined test data into the model obtained when training by combining two or more of the learning data and the augmented data at each stage, the accuracy of the model is derived. The learning data augmentation device according to claim 6.
Citation Information
Patent Citations
Training data expansion apparatus, method, and program
JP2020140466A
Adversarial Training Data Augmentation for Generating Related Responses
US20200227030A1
Method and apparatus for processing language based on trained network model
US20200372217A1