Learning assistance device, method for operating learning assistance device, and program for operating learning assistance device

By deriving and utilizing indices to select and edit relationships between input and correct answer components, the learning support device enhances language model performance by generating diverse and relevant training data, addressing the issue of insufficient learning data in existing models.

WO2025150363A1PCT designated stage expired Publication Date: 2025-07-17FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/044661
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2024-12-17
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing language models face performance issues due to insufficient learning data, particularly when new data is generated without considering the varying degrees of relationship between input and output sentences, leading to potential low performance in generating expected output sentences.

Method used

A learning support device that derives an index representing the relationship between input and correct answer components, selects relevant pairs, and performs editing processes such as separation, integration, order change, or deletion to generate new learning data, enhancing the performance of language models by providing diverse and relevant training data.

Benefits of technology

The solution enables the generation of higher-performance language models by ensuring that the training data reflects meaningful relationships between input and output sentences, reducing the risk of overfitting and improving overall model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024044661_17072025_PF_FP_ABST
    Figure JP2024044661_17072025_PF_FP_ABST
Patent Text Reader

Abstract

This learning assistance device comprises a processor. The processor: acquires existing learning data of a language model configured from input data and correct answer data, at least one of which is a sentence, the language model including input data that includes a plurality of first constituent components and correct answer data that includes a plurality of second constituent components; derives an index representing the relatedness between the first constituent components and the second constituent components; selects a set of a first constituent component and a second constituent component for which the index satisfies a preset condition; and generates new learning data by applying an editing process related to the selected set to the existing learning data.
Need to check novelty before this filing date? Find Prior Art

Description

Learning support device, operation method for learning support device, and operation program for learning support device

[0001] The technology of the present disclosure relates to a learning assistance device, an operating method for a learning assistance device, and an operating program for a learning assistance device.

[0002] Recently, language models such as BERT (Bidirectional Encoder Representations from Transformers) and ChatGPT (Chat Generative Pre-trained Transformer) have been attracting attention. Learning a language model requires a large amount of training data (also called teacher data or training data), which is often insufficient. Therefore, for example, "Bohan Li, et al. "Data augmentation approaches in natural language processing: A survey," AI Open Volume 3, 2022, Pages 71-90. (hereinafter referred to as Non-Patent Document 1), data augmentation (DA) is performed to generate new training data from existing training data. Non-Patent Document 1 introduces various data augmentation techniques, such as paraphrasing words or sentences, and inserting, substituting, or deleting random words or sentences.

[0003] For example, between an input sentence and a summary sentence that summarizes it, the degree of relationship between the input sentence and the output sentence that is converted from it varies depending on the combination of sentences included. More specifically, when the output sentence is a summary sentence, sentences in the input sentence whose content is included in a sentence in the summary sentence have a high relationship with the sentence in the summary sentence. On the other hand, sentences in the input sentence whose content is not included in a sentence in the summary sentence have a low relationship with the sentence in the summary sentence. For this reason, if new training data is generated by simply deleting sentences indiscriminately without taking into account the different degrees of relationship between such sentences and then used for training, there is a risk that a language model with low performance will be generated that does not produce the expected output sentence.

[0004] One embodiment of the technique of the present disclosure provides a learning assistance device, an operating method for a learning assistance device, and an operating program for a learning assistance device that can contribute to the generation of a language model with higher performance than conventional models.

[0005] The learning support device of the present disclosure includes a processor, which acquires existing training data of a language model consisting of input data and correct answer data, at least one of which is a sentence, the input data including a plurality of first components and correct answer data including a plurality of second components, derives an index representing the relationship between the first components and the second components, selects pairs of first components and second components whose index satisfies a predetermined condition, and generates new training data by performing an editing process related to the selected pair on the existing training data.

[0006] It is preferable that the index is at least one of the following: the superficial similarity between the first component and the second component; the semantic similarity between the first component and the second component; the keyword similarity between the first component and the second component; a numerical value corresponding to the result of topic classification of the first component and the second component; a numerical value corresponding to the result of determining the implication relationship between the first component and the second component; and the coverage rate of one of the first component and the second component relative to the other.

[0007] When the indicator is coverage, it is preferable that the processor performs a first selection to select a pair of a first component and a second component whose coverage satisfies a first selection condition, derives an integrated coverage rate between a common component, which is one of the first component and the second component that is selected in common with two or more other components of the first component and the second component, and an integrated component that integrates the other two or more components, selects a pair of a first component and a second component whose coverage rate satisfies a second selection condition that is stricter than the first selection condition, and performs a second selection to select a pair of a common component and an integrated component whose integrated coverage rate satisfies the second selection condition.

[0008] Preferably, the processor derives an index of one first component and one second component, and an index of two or more first components and one second component.

[0009] Preferably, the first and second components are sentences.

[0010] It is preferable that the processor performs at least one of the following editing processes: a first editing process for generating new learning data using only a pair of selected first and second components from the existing learning data; a second editing process for integrating pairs of selected first and second components from multiple existing learning data to generate new learning data; a third editing process for deleting pairs of selected first and second components from the existing learning data to generate new learning data; and a fourth editing process for changing the order of pairs of selected first and second components from the existing learning data to generate new learning data.

[0011] When the processor performs at least one of the second editing process and the fourth editing process to generate new training data, it is preferable that the processor uses a sentence order alignment model that arranges the order of sentences in the input text.

[0012] The processor preferably selects data to be officially adopted as training data from among a plurality of new training data candidates generated by the editing process.

[0013] It is preferable that the processor makes the selection based on a first degree of deviation between the target new training data and the existing training data that is the source of the target new training data and new training data other than the target new training data, and a second degree of deviation between the target new training data and the set of existing training data.

[0014] The processor preferably makes the selection using a sentence semantic decision model that outputs a score representing the clarity of the meaning of the input sentence.

[0015] A language model is preferably a model tasked with transforming input sentences into output sentences of different styles.

[0016] Preferably, the processor trains the language model using at least the new training data.

[0017] Preferably, the processor trains the language model using new training data and then trains the language model using existing training data.

[0018] There are a plurality of types of editing processes, and it is preferable that the processor divides new learning data generated for each of the plurality of types of editing processes in stages and provides the data to the language model for learning.

[0019] It is preferable that the processor performs learning by providing new training data to the language model in order of the relatively low learning load.

[0020] The method of operating the learning assistance device disclosed herein includes acquiring existing training data of a language model consisting of input data including a plurality of first components and correct answer data including a plurality of second components, at least one of which is a sentence, deriving an index representing the relationship between the first components and the second components, selecting a pair of the first components and the second components whose index satisfies a predetermined condition, and generating new training data by performing an editing process related to the selected pair on the existing training data.

[0021] The operating program of the learning support device disclosed herein causes a computer to execute processes including: acquiring existing training data of a language model consisting of input data including a plurality of first components and correct answer data including a plurality of second components, at least one of which is a sentence; deriving an index representing the relationship between the first components and the second components; selecting a pair of the first components and the second components whose index satisfies a predetermined condition; and generating new training data by performing an editing process related to the selected pair on the existing training data.

[0022] According to the technology of the present disclosure, it is possible to provide a learning assistance device, an operating method for a learning assistance device, and an operating program for a learning assistance device that can contribute to the generation of a language model with higher performance than conventional methods.

[0023] 1 is a diagram illustrating a learning support system. FIG. 1 is a diagram illustrating a task of a language model. FIG. 2 is a diagram illustrating existing training data. FIG. 3 is a diagram illustrating new training data. FIG. 4 is a block diagram illustrating a computer constituting a data extension device and a learning device. FIG. 5 is a block diagram illustrating a processing unit of a CPU of the data extension device. FIG. 6 is a diagram illustrating a division result. FIG. 7 is a diagram illustrating a target from which an index is to be derived. FIG. 8 is a diagram illustrating specific contents of an index. FIG. 9 is a diagram illustrating a derivation result. FIG. 10 is a diagram illustrating processing by a set selection unit. FIG. 11 is a diagram illustrating a first editing process. FIG. 12 is a diagram illustrating another example of the first editing process. FIG. 13 is a diagram illustrating a third editing process. FIG. 14 is a diagram illustrating a fourth editing process. FIG. 15 is a block diagram illustrating a processing unit of a CPU of the learning device. FIG. 16 is a diagram illustrating a learning procedure for a language model using existing training data. FIG. 17 is a diagram illustrating a learning procedure for a language model using new training data. A flowchart illustrating a processing procedure for the data extension device. A flowchart illustrating a processing procedure for the learning device. FIG. 18 is a diagram illustrating another example of a target from which an index is to be derived. A diagram illustrating an example of narrowing down targets from which an index is to be derived based on a combination condition. A diagram illustrating another example of a derivation result. FIG. 19 is a diagram illustrating a fifth editing process. FIG. 20 is a diagram illustrating a sixth editing process. FIG. 21 is a diagram illustrating another example of the third editing process. A diagram illustrating a derivation result of a mode in which the recall of existing input data sentences relative to existing supervised data sentences is derived as an index. 2 is a diagram showing the processing of the set selection unit when the index is recall. FIG. 3 is a block diagram showing the processing unit of a mode in which first selection and second selection are performed. FIG. 4 is a diagram showing the integration processing. FIG. 5 is a diagram showing the processing of deriving the integrated recall. FIG. 6 is a diagram showing the task of the sentence order rearrangement model. FIG. 7 is a diagram showing an example of a sentence. FIG. 8 is a diagram showing an example of use of the sentence order rearrangement model. FIG. 9 is a block diagram showing the processing unit of the 2_1 embodiment. FIG. 10 is a diagram showing the detailed configuration of the selection unit. FIG. 11 is a graph for explaining the processing of the first calculation unit. FIG. 12 is a graph for explaining the processing of the second calculation unit. FIG. 13 is a diagram showing selection conditions of the 2_1 embodiment. FIG. 14 is a block diagram showing the processing unit of the 2_2 embodiment. FIG. 15 is a diagram showing the task of the sentence meaning determination model. FIG. 16 is a diagram showing how new learning input data and new correct answer data are input to the sentence meaning determination model, scores are output from the sentence meaning determination model, and an average score is calculated.FIG. 1 is a diagram showing selection conditions for a second_2 embodiment; FIG. 2 is a diagram showing a learning schedule for a language model; FIG. 3 is a diagram showing the learning load of new learning data; FIG. 4 is a diagram showing how new learning data is divided into stages and given to a language model for learning; FIG. 5 is a diagram showing tasks of a pair selection determination model; FIG. 6 is a flowchart showing the procedure of processing using a pair selection determination model; FIG. 7 is a diagram showing an example where input data is numerical data; and FIG. 8 is a diagram showing an example where input data is speech recognition data.

[0024] [First Embodiment] As shown in FIG. 1 as an example, a learning support system 2 is a system that supports learning of a language model 15 (see FIG. 2 ) and is installed at a company that develops the language model 15, etc. The learning support system 2 is composed of a data extension device 10 and a learning device 11. The data extension device 10 and the learning device 11 are, for example, desktop personal computers, and are operated by a developer DV of the language model 15. The data extension device 10 and the learning device 11 are connected to each other so that they can communicate with each other via a computer network such as a local area network (LAN).

[0025] The data extension device 10 and the learning device 11 are both examples of a "learning support device" according to the technology of the present disclosure. In this manner, the "learning support device" according to the technology of the present disclosure may be composed of multiple computers. Note that the functions of the data extension device 10 and the learning device 11 may be assigned to a single computer, and the single computer may operate as a "learning support device" according to the technology of the present disclosure.

[0026] The developer DV collects data that he / she considers useful for learning in consideration of the task of the language model 15, and inputs the collected data as existing training data 12E to the data expansion device 10. As a result, the data expansion device 10 acquires an existing training data group 12EG, which is a collection of multiple existing training data 12E.

[0027] The data expansion device 10 performs data expansion to generate new training data 12N (see FIG. 4 ) from existing training data 12E. Through this data expansion, the data expansion device 10 has a new training data group 12NG, which is a collection of multiple new training data 12N. The data expansion device 10 distributes and outputs a training data group 12G, which is composed of the existing training data group 12EG and the new training data group 12NG, to the training device 11. The training device 11 trains a language model 15 using the existing training data 12E and the new training data 12N. Note that, hereinafter, the existing training data 12E and the new training data 12N may be collectively referred to as training data 12.

[0028] As an example, as shown in FIG. 2 , the language model 15 is a model responsible for the text summarization task of converting input data 16, which is an input text, into output data 17, which is a summary text. In other words, the language model 15 is an example of a "model responsible for the task of converting input text into output text of a different style" according to the technology of the present disclosure. The language model 15 is, for example, BERT, ChatGPT, etc. Here, in this specification, a text is defined as a collection of multiple sentences. A sentence is also defined as a collection of multiple words separated by periods. However, sentences of 200 characters or more (compound sentences, complex sentences, overlapping sentences, etc.) are exceptionally defined as sentences. A sentence is an example of a "component" according to the technology of the present disclosure. Note that FIG. 2 illustrates Dazai Osamu's "Run, Melos" as input data 16 and its summary as output data 17, but these are merely examples. A detailed factory maintenance report may be the input data 16, and a text summarizing abnormalities and their causes may be the output data 17.

[0029] As an example, as shown in Figure 3, existing training data 12E is composed of existing training input data 16E and existing correct answer data 17ECA. Also, as an example, as shown in Figure 4, new training data 12N is composed of new training input data 16N and new correct answer data 17NCA. The existing training input data 16E and existing correct answer data 17ECA, as well as the new training input data 16N and new correct answer data 17NCA, are also texts containing multiple sentences. The existing correct answer data 17ECA is a summary of the existing training input data 16E, and the new correct answer data 17NCA is a summary of the new training input data 16N.

[0030] The existing learning input data 16E and the new learning input data 16N are data input to the language model 15 during learning. The existing correct answer data 17ECA and the new correct answer data 17NCA are data for checking answers, so to speak, to be compared with the existing learning output data 17E (see FIG. 18 ) and the new learning output data 17N (see FIG. 19 ) output from the language model 15 in response to the input of the existing learning input data 16E and the new learning input data 16N. The existing learning input data 16E is an example of "input data" according to the technology of the present disclosure. Furthermore, the existing correct answer data 17ECA is an example of "correct answer data" according to the technology of the present disclosure.

[0031] 5, the computer that constitutes the data extension device 10 and the learning device 11 includes a storage 20, a memory 21, a CPU (Central Processing Unit) 22, a communication unit 23, a display 24, and an input device 25. These components are interconnected via a bus line 26.

[0032] The storage 20 is a hard disk drive built into the computer that constitutes the data expansion device 10 and the learning device 11, or connected via a cable or network. Alternatively, the storage 20 is a disk array consisting of multiple hard disk drives. The storage 20 stores control programs such as an operating system, various application programs, and various data associated with these programs. Note that a solid state drive may be used instead of a hard disk drive.

[0033] The memory 21 is a work memory for the CPU 22 to execute processing. The CPU 22 loads a program stored in the storage 20 into the memory 21 and executes processing in accordance with the program. In this way, the CPU 22 comprehensively controls each part of the computer. The CPU 22 is an example of a "processor" according to the technology of the present disclosure. The memory 21 may be built into the CPU 22.

[0034] The communication unit 23 is a network interface that controls the transmission of various information via the network. The display 24 displays various screens. Each screen is equipped with an operation function using a GUI (Graphical User Interface). The computers that make up the data extension device 10 and the learning device 11 accept operation instructions input from an input device 25 via each screen. The input device 25 is a keyboard, mouse, touch panel, microphone for voice input, etc.

[0035] In the following explanation, the parts of the computer that make up the data expansion device 10 (storage 20 and CPU 22) are distinguished by adding the suffix "A" to their symbols, and the parts of the computer that make up the learning device 11 (storage 20 and CPU 22) are distinguished by adding the suffix "B" to their symbols.

[0036] As an example, as shown in Fig. 6, an operating program 30 is stored in the storage 20A of the data extension device 10. The operating program 30 is an application program for causing a computer to function as the data extension device 10. In other words, the operating program 30 is an example of an "operating program for a learning assistance device" according to the technology of the present disclosure. The storage 20A also stores selection conditions 31 and the like. The selection conditions 31 are an example of a "preset condition" according to the technology of the present disclosure.

[0037] When the operating program 30 is started, the CPU 22A of the computer constituting the data extension device 10 cooperates with the memory 21 and the like to function as a read / write (hereinafter abbreviated as RW) control unit 35, a sentence division unit 36, an index derivation unit 37, a set selection unit 38, an editing unit 39, and a distribution control unit 40.

[0038] The RW control unit 35 controls the storage of various data in the storage 20A and the reading of various data from the storage 20A. For example, the RW control unit 35 stores existing training data 12E in the storage 20A. The RW control unit 35 also reads an existing training data group 12EG from the storage 20A and outputs the read existing training data group 12EG to the sentence dividing unit 36, the editing unit 39, and the distribution control unit 40. The RW control unit 35 also reads selection conditions 31 from the storage 20A and outputs the read selection conditions 31 to the set selecting unit 38.

[0039] The sentence dividing unit 36 ​​divides the sentences of the existing learning input data 16E and the existing correct answer data 17ECA of each existing learning data 12E in the existing learning data group 12EG into sentences one by one. The sentence dividing unit 36 ​​outputs the sentence division results 45 to the index derivation unit 37.

[0040] As an example, as shown in FIG. 7 , the segmentation result 45 is obtained by assigning sentence IDs to sentences included in the existing learning input data 16E and the existing correct answer data 17ECA for each existing learning data ID (Identification Data) for uniquely identifying the existing learning data 12E (for each existing learning data 12E). The sentence ID for a sentence included in the existing learning input data 16E (a sentence classified as "input data" in FIG. 7 ) is obtained by adding a number indicating the order of appearance after SS, which is an abbreviation for Source Sentence. The sentence ID for a sentence included in the existing correct answer data 17ECA (a sentence classified as "correct answer data" in FIG. 7 ) is obtained by adding a number indicating the order of appearance after TS, which is an abbreviation for Target Sentence. In the following, a sentence included in the existing learning input data 16E will be referred to as an existing input data sentence ESS. Furthermore, a sentence included in the existing supervised answer data 17ECA will be referred to as an existing supervised answer data sentence ETS. The existing input data sentence ESS is an example of a "first component" according to the technology of the present disclosure. The existing supervised answer data sentence ETS is also an example of a "second component" according to the technology of the present disclosure.

[0041] 6, the index derivation unit 37 derives an index INX (see FIG. 9) representing the relationship between the existing input data sentence ESS and the existing supervised data sentence ETS for each existing training data 12E. The index derivation unit 37 outputs the derived result 46 of the index INX to the pair selection unit 38.

[0042] As an example, as shown in Fig. 8, the index derivation unit 37 derives the index INX for a plurality of existing input data sentences ESS and a plurality of existing supervised data sentences ETS in a brute force manner. Fig. 8 illustrates a case where the existing input data sentences ESS have four sentence IDs "SS0001" to "SS0004" and the existing supervised data sentences ETS have two sentence IDs "TS0001" and "TS0002". In this case, as shown by the dashed lines and below the arrows, the index derivation unit 37 derives indexes INX for "SS0001-TS0001," "SS0001-TS0002," "SS0002-TS0001," "SS0002-TS0002," "SS0003-TS0001," "SS0003-TS0002," "SS0004-TS0001," and "SS0004-TS0002." Note that the "-" connecting the sentence ID of the existing input data sentence ESS and the sentence ID of the existing supervised data sentence ETS means that the pair of the existing input data sentence ESS and the existing supervised data sentence ETS of the sentence ID is the target for deriving the index INX.

[0043] 9, the index INX is obtained by weighting the surface similarity SPS between the existing input data sentence ESS and the existing super-answer data sentence ETS, and the numerical value VTC corresponding to the result of topic classification of the existing input data sentence ESS and the existing super-answer data sentence ETS, respectively, with α and (1-α). α is a value arbitrarily set to control the balance between the surface similarity SPS and the numerical value VTC, and is in the range of 0<α<1. The surface similarity SPS and the numerical value VTC corresponding to the result of topic classification are both numerical values ​​with a minimum value of 0 and a maximum value of 1. Note that the semantic similarity SMS between the existing input data sentence ESS and the existing super-answer data sentence ETS, which will be described later, and the keyword similarity KM between the existing input data sentence ESS and the existing super-answer data sentence ETS are also numerical values ​​with a minimum value of 0 and a maximum value of 1.

[0044] The superficial similarity SPS is derived based on the degree of agreement between the words constituting the existing input data sentence ESS and the words constituting the existing super-answer data sentence ETS, the degree of agreement between the order of the words, etc. The closer the superficial similarity SPS is to 0, the less superficially similar the existing input data sentence ESS and the existing super-answer data sentence ETS are, and the closer it is to 1, the more superficially similar the existing input data sentence ESS and the existing super-answer data sentence ETS are. The index derivation unit 37 derives, as the superficial similarity SPS, well-known evaluation indices such as ROUGE (Recall-Oriented Understudy for Gisting Evaluation), BLEU (BiLingual Evaluation Understudy), CIDEr (Consensus-based Image Description Evaluation), and METEOR (Metric for Evaluation of Translation with Explicit Ordering).

[0045] While the superficial similarity SPS is, so to speak, the similarity in appearance between sentences, the semantic similarity SMS is a similarity that goes deeper into the meaning of the existing input data sentence ESS and the existing supervised data sentence ETS. The closer the semantic similarity SMS is to 0, the less semantic similarity there is between the existing input data sentence ESS and the existing supervised data sentence ETS, and the closer it is to 1, the more semantic similarity there is between the existing input data sentence ESS and the existing supervised data sentence ETS. The index derivation unit 37 derives the semantic similarity SMS using a well-known model such as SentenceBERT or BERTScore.

[0046] The keyword similarity KM is a superficial similarity limited to nouns and proper nouns among the words that make up the existing input data sentence ESS and the existing supervised data sentence ETS.

[0047] The index derivation unit 37 derives a numerical VTC according to the results of the topic classification, for example, as follows. That is, the index derivation unit 37 performs clustering, such as Latent Dirichlet Allocation (LDA), on the existing input data sentence ESS and the existing correct answer data sentence ETS to identify the topics inherent in the existing input data sentence ESS and the existing correct answer data sentence ETS. Next, the index derivation unit 37 derives a numerical VTC according to the degree of agreement between the topics of the existing input data sentence ESS and the existing correct answer data sentence ETS. To facilitate topic identification, the developer DV may manually register keywords for each topic. Note that a topic refers to an issue discussed in a sentence. For example, the sentence "The President of the United States and the Prime Minister of Japan played golf" contains three inherent topics: politics, international affairs, and sports.

[0048] The index derivation unit 37 derives a numerical value VIR according to the determination result of the entailment relationship, for example, as follows. That is, the index derivation unit 37 inputs the existing input data sentence ESS and the existing supervised data sentence ETS into a well-known textual entailment recognition (RTE) model or a natural language inference (NLI) model. Then, the model outputs the probability that the input existing input data sentence ESS and the existing supervised data sentence ETS are in an entailment relationship, and this probability is used as the numerical value VIR according to the determination result of the entailment relationship. Note that sentences being in an entailment relationship means, for example, that if the sentence "Bach, a composer representing Baroque music, published 'Air on the G String' around 1730," is true, then "The composer of 'Air on the G String' is Bach." This refers to a relationship in which the sentence " is also true.

[0049] The index INX is not limited to a weighted sum of the surface similarity SPS and the numerical value VTC corresponding to the topic classification results, but may also be an arithmetic average or weighted average of these. Furthermore, the index INX may be calculated by adding, for example, 0.1 to the arithmetic average or weighted average of the surface similarity SPS and the numerical value VTC corresponding to the topic classification results if an entailment relationship is determined to exist, or by subtracting, for example, 0.1 if an entailment relationship is determined to exist. The index INX may also include, in addition to the surface similarity SPS and the numerical value VTC corresponding to the topic classification results, the semantic similarity SMS, the keyword similarity KM, and the numerical value VIR corresponding to the entailment relationship determination results. The index INX may include at least one of the surface similarity SPS, the semantic similarity SMS, the keyword similarity KM, the numerical value VTC corresponding to the topic classification results, and the numerical value VIR corresponding to the entailment relationship determination results.

[0050] As an example, as shown in FIG. 10, the derivation result 46 is a registered index INX for each pair of existing input data sentence ESS and existing correct answer data sentence ETS for each existing learning data ID (each existing learning data 12E).

[0051] 6, the pair selection unit 38 selects pairs of existing input data sentences ESS and existing correct answer data sentences ETS whose index INX satisfies the selection condition 31. The pair selection unit 38 outputs the pair selection results 47 to the editing unit 39.

[0052] 11 , the selection condition 31 may be, for example, that the index INX is 0.3 or greater. Therefore, the set selection unit 38 extracts pairs of an existing input data sentence ESS and an existing supervised data sentence ETS having an index INX of 0.3 or greater from the derivation result 46, thereby obtaining a selection result 47. In other words, the set selection unit 38 removes pairs of an existing input data sentence ESS and an existing supervised data sentence ETS having an index INX less than 0.3 from the derivation result 46, thereby obtaining the selection result 47. Note that the selection condition 31 may be, for example, that the index INX is within the top 30%. Alternatively, the index INX may be 0.3 or greater and within the top 30%.

[0053] 6 , the editing unit 39 generates new learning data 12N by performing editing processing on the existing learning data 12E regarding the pair of the existing input data sentence ESS and the existing correct answer data sentence ETS registered in the selection result 47. The editing unit 39 generates multiple pieces of new learning data 12N and outputs a new learning data group 12NG, which is a collection of the multiple pieces of new learning data 12N, to the RW control unit 35. The RW control unit 35 stores the new learning data group 12NG in the storage 20A.

[0054] The RW control unit 35 reads the existing learning data group 12EG and the new learning data group 12NG from the storage 20A in response to a distribution instruction for the learning data group 12G from the developer DV via the input device 25. Then, the RW control unit 35 outputs the read existing learning data group 12EG and the new learning data group 12NG to the distribution control unit 40. The distribution control unit 40 distributes and outputs the learning data group 12G consisting of the existing learning data group 12EG and the new learning data group 12NG to the learning device 11 specified in the distribution instruction.

[0055] 12 to 16, the editing unit 39 performs four types of editing processes: a first editing process, a second editing process, a third editing process, and a fourth editing process. In Figures 12 to 16, the existing input data sentence ESS and the existing supervised data sentence ETS connected by a thin solid line are pairs of the existing input data sentence ESS and the existing supervised data sentence ETS selected by the pair selection unit 38. This also applies to the subsequent figures.

[0056] 12 and 13 show the first editing process. The first editing process is a process of separating a pair of a selected existing input data sentence ESS and an existing correct answer data sentence ETS from the existing training data 12E, and generating new training data 12N using only the pair of the selected existing input data sentence ESS and the existing correct answer data sentence ETS. In other words, the first editing process is a process of deleting everything other than the pair of the selected existing input data sentence ESS and the existing correct answer data sentence ETS.

[0057] 12 illustrates an example in which a set of existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and an existing correct answer data sentence ETS with sentence ID "TS0002" is selected. In this case, the editing unit 39 deletes the existing input data sentences ESS with sentence IDs "SS0001" and "SS0002" and the existing correct answer data sentence ETS with sentence ID "TS0001," and generates new training data 12N using only the set of existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and the existing correct answer data sentence ETS with sentence ID "TS0002." More specifically, new training input data 16N is generated using the existing input data sentences ESS with sentence IDs "SS0003" and "SS0004." Furthermore, new correct answer data 17NCA is generated using the existing correct answer data sentence ETS with sentence ID "TS0002." As can be seen from the example in FIG. 12, multiple existing input data sentences ESS may be selected for one existing correct answer data sentence ETS. Conversely, multiple existing correct answer data sentences ETS may be selected for one existing input data sentence ESS. Note that, hereinafter, a sentence included in new learning input data 16N will be referred to as a new input data sentence NSS. Furthermore, a sentence included in new correct answer data 17NCA will be referred to as a new correct answer data sentence NTS.

[0058] FIG. 13 illustrates an example in which a pair of an existing input data sentence ESS with sentence ID “SS0001” and an existing correct answer data sentence ETS with sentence ID “TS0001”, a pair of an existing input data sentence ESS with sentence ID “SS0003” and an existing correct answer data sentence ETS with sentence ID “TS0002”, and a pair of an existing input data sentence ESS with sentence ID “SS0004” and an existing correct answer data sentence ETS with sentence ID “TS0003” are selected. In this case, the editing unit 39 deletes the existing input data sentence ESS with the isolated sentence ID "SS0002" and generates new training data 12N using only the existing input data sentences ESS with sentence IDs "SS0001," "SS0003," and "SS0004" and the existing correct answer data sentences ETS with sentence IDs "TS0001," "TS0002," and "TS0003." More specifically, new training input data 16N is generated using the existing input data sentences ESS with sentence IDs "SS0001," "SS0003," and "SS0004." Furthermore, new correct answer data 17NCA is generated using the existing correct answer data sentences ETS with sentence IDs "TS0001," "TS0002," and "TS0003." As can be seen from this example, one of the new learning input data 16N and the new correct answer data 17NCA may be a direct copy of the existing learning input data 16E or the existing correct answer data 17ECA.

[0059] 14 shows the second editing process. The second editing process is a process of integrating pairs of selected existing input data sentences ESS and existing supervised answer data sentences ETS of multiple existing learning data 12E to generate new learning data 12N. The multiple existing learning data 12E integrated in the second editing process are, for example, existing learning data 12E whose average value of the brute-force index INX between the existing input data sentences ESS is 0.7 or greater.

[0060] 14 illustrates an example in which sentences of two existing training data 12E_A and 12E_B are integrated to generate new training data 12N. Also, FIG. 14 illustrates an example in which a pair of an existing input data sentence ESS_A with a sentence ID of "SS0003A" and an existing correct answer data sentence ETS_A with a sentence ID of "TS0001A" in the existing training data 12E_A, and a pair of an existing input data sentence ESS_A with a sentence ID of "SS0004A" and an existing correct answer data sentence ETS_A with a sentence ID of "TS0002A" in the existing training data 12E_A are selected. Furthermore, Figure 14 illustrates a case in which a pair of existing input data sentence ESS_B with sentence ID "SS0001B" in existing learning data 12E_B and existing correct answer data sentence ETS_B with sentence ID "TS0001B", as well as a pair of existing input data sentence ESS_B with sentence ID "SS0003B" in existing learning data 12E_B and existing correct answer data sentence ETS_B with sentence ID "TS0002B" are selected.

[0061] In this case, the editing unit 39 extracts from the existing training data 12E_A a pair of an existing input data sentence ESS_A with a sentence ID of "SS0003A" and an existing correct answer data sentence ETS_A with a sentence ID of "TS0001A," as well as a pair of an existing input data sentence ESS_A with a sentence ID of "SS0004A" and an existing correct answer data sentence ETS_A with a sentence ID of "TS0002A." The editing unit 39 also extracts from the existing training data 12E_B a pair of an existing input data sentence ESS_B with a sentence ID of "SS0001B" and an existing correct answer data sentence ETS_B with a sentence ID of "TS0001B," as well as a pair of an existing input data sentence ESS_B with a sentence ID of "SS0003B" and an existing correct answer data sentence ETS_B with a sentence ID of "TS0002B." These sentences are then integrated to generate new training data 12N. More specifically, new training input data 16N is generated using existing input data sentences ESS_A with sentence IDs "SS0003A" and "SS0004A" and existing input data sentences ESS_B with sentence IDs "SS0001B" and "SS0003B." New correct answer data 17NCA is also generated using existing correct answer data sentences ETS_A with sentence IDs "TS0001A" and "TS0002A" and existing correct answer data sentences ETS_B with sentence IDs "TS0001B" and "TS0002B."

[0062] 15 shows the third editing process. The third editing process is a process of generating new learning data 12N by deleting a pair of a selected existing input data sentence ESS and an existing supervised answer data sentence ETS from the existing learning data 12E. In other words, the third editing process is a process of separating everything except the pair of the selected existing input data sentence ESS and the existing supervised answer data sentence ETS. In other words, the third editing process is the flip side of the first editing process.

[0063] 15, like the case of FIG. 12, illustrates a case in which a set of existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and an existing correct answer data sentence ETS with sentence ID "TS0002" is selected. In this case, the editing unit 39 deletes the existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and the existing correct answer data sentence ETS with sentence ID "TS0002," and generates new training data 12N using only the set of existing input data sentences ESS with sentence IDs "SS0001" and "SS0002" and the existing correct answer data sentence ETS with sentence ID "TS0001." More specifically, new training input data 16N is generated using the existing input data sentences ESS with sentence IDs "SS0001" and "SS0002." Also, new correct answer data 17NCA is generated using the existing correct answer data sentence ETS with sentence ID "TS0001".

[0064] 16 shows the fourth editing process. The fourth editing process is a process for generating new learning data 12N by changing the order of pairs of selected existing input data sentences ESS and existing supervised data sentences ETS from the existing learning data 12E.

[0065] 16 illustrates a case where a set of an existing input data sentence ESS with sentence ID "SS0001" and an existing correct answer data sentence ETS with sentence ID "TS0001," as well as a set of existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and an existing correct answer data sentence ETS with sentence ID "TS0002" are selected. In this case, the editing unit 39 generates new training data 12N by changing the order of the existing input data sentences ESS with sentence IDs "SS0003" and "SS0004" and the existing correct answer data sentence ETS with sentence ID "TS0002" to first, and the order of the existing input data sentence ESS with sentence ID "SS0001" and the existing correct answer data sentence ETS with sentence ID "TS0001" to last.

[0066] The editing unit 39 performs all four types of editing processes, the first editing process, the second editing process, the third editing process, and the fourth editing process, on one existing training data 12E. Therefore, four new training data 12N are generated from one existing training data 12E. However, the first editing process and the third editing process cannot be performed on existing training data 12E from which all pairs of existing input data sentences ESS and existing correct answer data sentences ETS have been selected. The first editing process, the second editing process, the third editing process, and the fourth editing process may be performed randomly, such as by performing the first editing process and the fourth editing process on some existing training data 12E and all of the first editing process, the second editing process, the third editing process, and the fourth editing process on other existing training data 12E.

[0067] 17, an operating program 50 is stored in storage 20B of learning device 11. Operating program 50 is an application program for causing a computer to function as learning device 11. In other words, operating program 50, together with operating program 30, is an example of an "operating program for a learning assistance device" according to the technology of the present disclosure. Storage 20B also stores a language model 15 and the like.

[0068] When the operating program 50 is started, the CPU 22B of the computer constituting the learning device 11 functions as a read / write (hereinafter abbreviated as RW (Read Write)) control unit 55 and a learning unit 56 in cooperation with the memory 21 and the like.

[0069] The RW control unit 55 controls the storage of various data in the storage 20B and the reading of various data from the storage 20B. For example, the RW control unit 55 stores the training data group 12G from the data expansion device 10 in the storage 20B. The RW control unit 55 also reads the training data group 12G from the storage 20B and outputs the read training data group 12G to the training unit 56. The RW control unit 55 also reads the language model 15 from the storage 20B and outputs the read language model 15 to the training unit 56. The training unit 56 trains the language model 15 using the training data group 12G.

[0070] 18 , the learning unit 56 inputs existing learning input data 16E to the language model 15 and causes the language model 15 to output existing learning output data 17E. The learning unit 56 compares the existing learning output data 17E with existing correct answer data 17ECA, and performs loss calculation for the language model 15 using a loss function based on the comparison result. Then, the learning unit 56 updates the various coefficients of the language model 15 according to the result of the loss calculation, and updates the language model 15 according to the update setting.

[0071] 19 , the learning unit 56 inputs new learning input data 16N to the language model 15 and causes the language model 15 to output new learning output data 17N. The learning unit 56 compares the new learning output data 17N with new correct answer data 17NCA, and performs loss calculation for the language model 15 using a loss function based on the comparison result. Then, the learning unit 56 updates various coefficients of the language model 15 according to the result of the loss calculation, and updates the language model 15 according to the update setting.

[0072] The learning unit 56 repeatedly performs a series of processes, including inputting existing training input data 16E or new training input data 16N to the language model 15, outputting existing training output data 17E or new training output data 17N from the language model 15, calculating a loss, setting an update, and updating the language model 15, while exchanging the existing training data 12E or the new training data 12N. The learning unit 56 terminates the repetition of the series of processes when the performance of the language model 15 reaches a predetermined set level. The learning unit 56 outputs the language model 15 whose performance has reached the set level to the RW control unit 55. The RW control unit 55 re-stores the language model 15 whose performance has reached the set level in the storage 20B as a trained language model 15. Note that learning may be terminated when the series of processes have been repeated a set number of times, regardless of the performance of the language model 15. Furthermore, learning of the language model 15 may be continued even after the language model 15 is re-stored in the storage 20B.

[0073] Next, the operation of the above configuration will be described with reference to the flowcharts shown in Figures 20 and 21. As shown in Figure 6, the CPU 22A of the data extension device 10 functions as an RW control unit 35, a sentence division unit 36, an index derivation unit 37, a pair selection unit 38, an editing unit 39, and a distribution control unit 40 when the operating program 30 is started. Also, as shown in Figure 17, the CPU 22B of the learning device 11 functions as an RW control unit 55 and a learning unit 56 when the operating program 50 is started.

[0074] 20 , the existing learning data 12E is input to the data expansion device 10 by the developer DV. As shown in FIG. 20 , in the data expansion device 10, the RW control unit 35 acquires the existing learning data 12E (step ST100). The existing learning data 12E is stored in the storage 20A under the control of the RW control unit 35. By repeating this process, an existing learning data group 12EG, which is a collection of the existing learning data 12E, is stored in the storage 20A.

[0075] The RW control unit 35 reads the existing training data group 12EG from the storage 20A, and the read existing training data group 12EG is output from the RW control unit 35 to the sentence division unit 36 ​​and the like.

[0076] The sentence dividing unit 36 ​​divides the sentences of the existing learning input data 16E and the existing correct answer data 17ECA of each existing learning data 12E in the existing learning data group 12EG into sentences (step ST110). Then, the division result 45 shown in FIG. 7 is output from the sentence dividing unit 36 ​​to the index derivation unit 37.

[0077] 8 and 9, the index derivation unit 37 derives the index INX of the existing input data sentence ESS and the existing supervised data sentence ETS of each existing training data 12E (step ST120). Then, the derivation result 46 shown in FIG. 10 is output from the index derivation unit 37 to the set selection unit 38.

[0078] 11, the pair selection unit 38 selects pairs of existing input data sentences ESS and existing correct answer data sentences ETS whose index INX satisfies the selection condition 31 (step ST130). Then, the selection result 47 is output from the pair selection unit 38 to the editing unit 39.

[0079] 12 to 16, the editing unit 39 performs an editing process on the existing learning data 12E relating to the pair of existing input data sentence ESS and existing supervised data sentence ETS registered in the selection result 47, thereby generating new learning data 12N (step ST140). A plurality of new learning data 12N are generated. Then, a new learning data group 12NG, which is a collection of the plurality of new learning data 12N, is output from the editing unit 39 to the RW control unit 35. The new learning data group 12NG is stored in the storage 20A under the control of the RW control unit 35.

[0080] When a distribution instruction for the learning data group 12G is received from the developer DV via the input device 25, the RW control unit 35 reads the existing learning data group 12EG and the new learning data group 12NG from the storage 20A. The read existing learning data group 12EG and the new learning data group 12NG are output from the RW control unit 35 to the distribution control unit 40. Then, the distribution control unit 40 distributes and outputs the learning data group 12G consisting of the existing learning data group 12EG and the new learning data group 12NG to the learning device 11 specified in the distribution instruction (step ST150).

[0081] 21 , in the learning device 11, the learning data group 12G is received from the data extension device 10 and acquired by the RW control unit 55 (step ST200). The learning data group 12G is stored in the storage 20B under the control of the RW control unit 55.

[0082] The RW control unit 55 reads the learning data group 12G from the storage 20B, and the read learning data group 12G is output from the RW control unit 55 to the learning unit 56.

[0083] 18, in the learning unit 56, existing training input data 16E is input to the language model 15, which in turn outputs existing training output data 17E from the language model 15. Alternatively, as shown in Fig. 19, new training input data 16N is input to the language model 15, which in turn outputs new training output data 17N from the language model 15 (step ST210).

[0084] Then, a loss calculation is performed based on the comparison result between the existing learning output data 17E and the existing correct answer data 17ECA, or the comparison result between the new learning output data 17N and the new correct answer data 17NCA, and the language model 15 is updated according to the result of the loss calculation, and the language model 15 is updated according to the update setting (step ST220).

[0085] The processes of steps ST210 and ST220 are repeated while the training data 12 is being changed (step ST240) until the performance of the language model 15 reaches the set level (NO in step ST230).

[0086] When the performance of the language model 15 reaches the set level (YES in step ST230), the learning unit 56 outputs the language model 15 to the RW control unit 55. Then, under the control of the RW control unit 55, the language model 15 from the learning unit 56 is stored in the storage 20B as a trained language model 15 (step ST250).

[0087] As described above, the CPU 22A of the data expansion device 10 includes the RW control unit 35, the index derivation unit 37, the pair selection unit 38, and the editing unit 39. The RW control unit 35 acquires the existing training data 12E. The index derivation unit 37 derives an index INX representing the relationship between the existing input data sentence ESS constituting the existing training input data 16E and the existing supervised data sentence ETS constituting the existing supervised data 17ECA. The pair selection unit 38 selects pairs of the existing input data sentence ESS and the existing supervised data sentence ETS whose index INX satisfies the selection condition 31. The editing unit 39 generates new training data 12N by performing an editing process related to the selected pair on the existing training data 12E. Since the editing process takes into account the relationship between the existing input data sentence ESS and the existing supervised data sentence ETS, new training data 12N that contributes to improving the performance of the language model 15 can be generated. Therefore, it is possible to contribute to the generation of a language model 15 with higher performance than conventional models.

[0088] 9, the index INX includes the superficial similarity SPS between the existing input data sentence ESS and the existing supervised data sentence ETS, and a numerical value VTC according to the result of topic classification of the existing input data sentence ESS and the existing supervised data sentence ETS. This makes it possible to improve the validity of the index INX, and in turn the validity of the selection of a pair of the existing input data sentence ESS and the existing supervised data sentence ETS that satisfies the selection condition 31.

[0089] In this example, the first component is the existing input data sentence ESS and the second component is the existing correct answer data sentence ETS, both of which are sentences. Therefore, processing such as division and derivation of the index INX can be easily performed.

[0090] As shown in FIGS. 12 to 16 , the editing unit 39 performs the following editing processes: a first editing process, a second editing process, a third editing process, and a fourth editing process. The first editing process is a process of generating new learning data 12N using only pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing learning data 12E. The second editing process is a process of integrating pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from multiple existing learning data 12E to generate new learning data 12N. The third editing process is a process of deleting pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing learning data 12E to generate new learning data 12N. The fourth editing process is a process of changing the order of pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing learning data 12E to generate new learning data 12N. This allows for the generation of a wide variety of new learning data 12N.

[0091] As shown in Fig. 2, the language model 15 is a model that handles tasks such as converting input text into output text of different styles, such as a text summarization task. Such language models 15 are highly versatile and in high demand. Therefore, we can contribute to the creation of language models 15 that are highly versatile and in high demand.

[0092] 19 , the learning unit 56 uses at least the new training data 12N to learn the language model 15. Therefore, compared to when the language model 15 is learned using only the existing training data 12E, it is possible to prevent problems such as overlearning caused by a lack of training data 12.

[0093] Although the pair of one existing input data sentence ESS and one existing supervised data sentence ETS has been exemplified as the target for deriving the index INX, this is not limitative. For example, as shown in Fig. 22, a pair of a plurality of existing input data sentences ESS and one existing supervised data sentence ETS may be used as the target for deriving the index INX.

[0094] 22 illustrates an example in which the existing input data sentences ESS have three sentence IDs, "SS0001" to "SS0003," and the existing correct answer data sentences ETS have two sentence IDs, "TS0001" and "TS0002." In this case, the index derivation unit 37 derives indexes INX for "SS0001-TS0001," "SS0001-TS0002," "SS0002-TS0001," "SS0002-TS0002," "SS0003-TS0001," and "SS0003-TS0002," as indicated by the dashed lines and arrows below.

[0095] The index derivation unit 37 also derives the index INX for "SS0001+SS0002+SS0003-TS0001" and "SS0001+SS0002+SS0003-TS0002". The index derivation unit 37 also derives the index INX for "SS0001+SS0002-TS0001", "SS0001+SS0002-TS0002", "SS0001+SS0003-TS0001", "SS0001+SS0003-TS0002", "SS0002+SS0003-TS0001", and "SS0002+SS0003-TS0002". This makes it possible to generate new training data 12N that also takes into account the relationship between a plurality of existing input data sentences ESS and one existing correct answer data sentence ETS.

[0096] However, when deriving the index INX for a plurality of existing input data sentences ESS and one existing supervised data sentence ETS by exhaustive search as shown in Fig. 22, depending on the number of existing input data sentences ESS and existing supervised data sentences ETS, the number of targets for deriving the index INX may become extremely large. Therefore, as an example, the method shown in Fig. 23 may be adopted.

[0097] 23, the index derivation unit 37 first derives an index INX for one existing input data sentence ESS and one existing supervised data sentence ETS, and outputs a derivation result 46. Then, in accordance with a preset combination condition 60, it determines a plurality of existing input data sentences ESS to be combined as targets for deriving an index INX. The combination condition 60 is stored in the storage 20A, and is read from the storage 20A by the RW control unit 35 and output to the index derivation unit 37. The index derivation unit 37 derives an index INX for the plurality of existing input data sentences ESS determined in accordance with the combination condition 60 and for the corresponding one existing supervised data sentence ETS.

[0098] 23, the combination condition 60 specifies that existing input data sentences ESS having an index INX of 0.7 or greater for the same existing supervised data sentence ETS are to be combined. Therefore, the index derivation unit 37 combines the existing input data sentences ESS having sentence IDs "SS0001" and "SS0002" having indexes INX of 0.83 and 0.92, respectively, for the existing supervised data sentence ETS having sentence ID "TS0001," and derives the index INX of "SS0001+SS0002-TS0001." In addition, the index derivation unit 37 combines the existing input data sentences ESS with sentence IDs "SS0002" and "SS0003" whose indexes INX for the existing correct data sentence ETS with sentence ID "TS0002" are 0.74 and 0.85, respectively, and derives the index INX of "SS0002+SS0003-TS0002".

[0099] According to the method shown in Fig. 23, it is possible to narrow down the targets for deriving the index INX. Compared to the case of deriving the index INX for a plurality of existing input data sentences ESS and one existing correct answer data sentence ETS in a brute force manner as shown in Fig. 22, it is possible to reduce the effort required for deriving the index INX.

[0100] In addition, the method of deriving the index INX may be changed depending on the number of existing input data sentences ESS and existing correct answer data sentences ETS, for example, by using the method shown in Figure 22 when the total number of existing input data sentences ESS and existing correct answer data sentences ETS is less than a threshold, and by using the method shown in Figure 23 when the total number of existing input data sentences ESS and existing correct answer data sentences ETS is equal to or greater than a threshold.

[0101] The index INX may be derived from a set of one existing input data sentence ESS and multiple existing supervised data sentences ETS. Alternatively, the index INX may be derived from a set of multiple existing input data sentences ESS and multiple existing supervised data sentences ETS.

[0102] The object for which the index INX is derived is not limited to a pair of an existing input data sentence ESS and an existing correct answer data sentence ETS. As an example, as shown in Figure 24, the object for which the index INX is derived may be between existing input data sentences ESS and between existing correct answer data sentences ETS. Figure 24 illustrates an example in which the index INX is derived between existing input data sentences ESS with sentence ID "SS0001" and existing correct answer data sentence ETS with sentence ID "TS0001", as well as between existing input data sentences ESS with sentence ID "SS0001" and "SS0002", and between existing correct answer data sentences ETS with sentence ID "TS0001" and "TS0002".

[0103] In this case, as shown in FIG. 25 as an example, the editing unit 39 performs a fifth editing process, which deletes one of the pairs of existing input data sentences ESS or existing supervised answer data sentences ETS selected because the index INX satisfies the selection condition 31. FIG. 25 illustrates an example in which existing input data sentences ESS with sentence IDs "SS0001" and "SS0002" are selected. In this case, the editing unit 39 generates new training data 12N by deleting the existing input data sentence ESS with sentence ID "SS0002," which is one of the existing input data sentences ESS with sentence IDs "SS0001" and "SS0002." This allows new training data 12N to be generated that also takes into account the relationships between the existing input data sentences ESS and between the existing supervised answer data sentences ETS. Note that if the language model 15 is responsible for the exemplary text summarization task, it is not necessary to derive the index INX between the existing supervised answer data sentences ETS.

[0104] 26 , the editing unit 39 may perform a sixth editing process in which an isolated sentence in one existing training data 12E is inserted into another existing training data 12E, thereby generating new training data 12N with even greater variation.

[0105] 26 illustrates a case where the existing input data sentence ESS_A with the sentence ID "SS0002A" in the existing learning input data 16E_A of the existing learning data 12E_A is an isolated sentence. In this case, the editing unit 39 generates new learning data 12N by inserting the existing input data sentence ESS_A with the sentence ID "SS0002A" into the existing learning input data 16E_B of the existing learning data 12E_B.

[0106] An isolated sentence is likely to be a fixed phrase, etc., that does not significantly affect the overall content of the sentence. For this reason, even if an isolated sentence is inserted into other existing training data 12E, it is considered that the relationship between the existing input data sentence ESS and the existing correct answer data sentence ETS will not be disrupted.

[0107] 27 , if there is a pair of one existing input data sentence ESS and multiple existing supervised answer data sentences ETS that satisfies the selection condition 31, the editing unit 39 generates new training data 12N in the third editing process without deleting the pair. This makes it possible to avoid a situation in which an existing input data sentence ESS that is highly related to multiple existing supervised answer data sentences ETS is deleted, making it impossible for the language model 15 to estimate data corresponding to the multiple existing supervised answer data sentences ETS.

[0108] 27 illustrates a case in which a set of an existing input data sentence ESS with sentence ID "SS0001" and two existing correct answer data sentences ETS with sentence IDs "TS0001" and "TS0002" is selected. In this case, if the existing input data sentence ESS with sentence ID "SS0001" and the existing correct answer data sentence ETS with sentence ID "TS0001" were deleted, there would be no clue to infer the sentence corresponding to the existing correct answer data sentence ETS with sentence ID "TS0002." Therefore, the editing unit 39 generates new training data 12N without deleting the set of the existing input data sentence ESS with sentence ID "SS0001" and the existing correct answer data sentences ETS with sentence IDs "TS0001" and "TS0002."

[0109] It should be noted that the set selection unit 38 may be configured not to select a set of one existing input data sentence ESS and a plurality of existing supervised data sentences ETS even if the selection condition 31 is satisfied.

[0110] Examples of the index INX include surface similarity SPS, semantic similarity SMS, keyword similarity KM, a numerical value VTC corresponding to the result of topic classification, and a numerical value VIR corresponding to the result of determining an implication relationship, but the present invention is not limited to these. If the language model 15 is a model performing the exemplary text summarization task, the index INX may be derived as the recall (RC) of the existing input data sentence ESS relative to the existing supervised data sentence ETS, as shown in the derived result 65 of FIG. 28 . The existing input data sentence ESS is an example of the "other of the first component and the second component" according to the technology of the present disclosure. The existing supervised data sentence ETS is an example of the "one of the first component and the second component" according to the technology of the present disclosure. The recall (RC) is an example of the "coverage" according to the technology of the present disclosure.

[0111] In this case, as shown in FIG. 29 as an example, the pair selection unit 38 extracts pairs of existing input data sentences ESS and existing correct answer data sentences ETS having a reproducibility RC of 0.3 or more from the derived result 65 in accordance with a selection condition 66, for example, that the reproducibility RC, which is the index INX, is 0.3 or more, to obtain a selection result 67.

[0112] In the case of a language model 15 that handles the task of summarizing text, an index INX such as the superficial similarity SPS may misjudge the relationship between the existing input data sentence ESS and the existing correct answer data sentence ETS. For example, if the existing input data sentence ESS is "Early this morning, a pair of robbers broke into a convenience store in Shinjuku 3-Chome, stole money and valuables, and fled," and the corresponding existing correct answer data sentence ETS is "A robbery occurred," the superficial similarity SPS will not be very high, but the recall RC will be high. Therefore, if the index INX is the recall RC, the risk of misjudgement of the relationship between the existing input data sentence ESS and the existing correct answer data sentence ETS can be reduced.

[0113] When the index INX is the recurrence rate RC, the processes shown in FIGS. 30 to 32 may be performed as an example.

[0114] 30 , the CPU of the data expansion device of this aspect has two pair selection units, a first pair selection unit 701 and a second pair selection unit 702, instead of the pair selection unit 38 of the first embodiment. The first pair selection unit 701 performs a first selection to select pairs of existing input data sentences ESS and existing supervised data sentences ETS whose recall ratio RC satisfies a first selection condition 661. The first selection condition 661, like the selection condition 66 shown in FIG. 29 , is that the recall ratio RC, which is the index INX, is 0.3 or greater. The first pair selection unit 701 outputs a first selection result 671 to the second pair selection unit 702.

[0115] 31 , when there is a common sentence CS, which is an existing supervised data sentence ETS that has been selected in common from two or more existing input data sentences ESS, the second set selection unit 702 first performs an integration process to integrate the two or more existing input data sentences ESS to generate an integrated sentence IS. The common sentence CS is an example of a "common component" according to the technology of the present disclosure. The integrated sentence IS is an example of an "integrated component" according to the technology of the present disclosure.

[0116] FIG. 31 illustrates a case in which existing input data sentences ESS with sentence IDs "SS0001," "SS0002," and "SS0003" are commonly selected for an existing supervised data sentence ETS with sentence ID "TS0001." In this case, the existing supervised data sentence ETS with sentence ID "TS0001" becomes the common sentence CS. The second set selection unit 702 generates an integrated sentence IS by connecting the existing input data sentences ESS with sentence IDs "SS0001," "SS0002," and "SS0003" with a coordinate conjunction (a conjunction particle in Japanese) to form a single sentence. Next, as shown in FIG. 32 as an example, the second set selection unit 702 performs a derivation process to derive an integrated recall IRC, which is the recall between the generated integrated sentence IS and the common sentence CS. The integrated recall IRC is an example of an "integrated coverage" according to the technology of the present disclosure.

[0117] 30 , the second set selection unit 702 selects a pair of an integrated sentence IS and a common sentence CS whose integrated recall IRC satisfies the second selection condition 662. Similarly, the second set selection unit 702 selects pairs of existing input data sentences ESS and existing supervised data sentences ETS other than the common sentence CS and the integrated sentence IS (such as the existing input data sentence ESS with sentence ID “SS0004” and the existing supervised data sentence ETS with sentence ID “TS0002” in FIG. 31 ) whose recall RC satisfies the second selection condition 662. The second selection condition 662 is stricter than the first selection condition 661, requiring that the integrated recall IRC or the recall RC, which is the index INX, be 0.7 or greater. The second set selection unit 702 outputs a second selection result 672 to the editing unit 39. Note that for pairs of existing input data sentences ESS and existing correct answer data sentences ETS other than the common sentence CS and the integrated sentence IS, the first selection based on the first selection condition 661, which is looser than the second selection condition 662, is not necessary, but is necessary because the first selection is performed to extract the common sentence CS.

[0118] In this case, the editing unit 39 performs editing processing on a pair of one existing input data sentence ESS and one existing correct answer data sentence ETS from among the pairs of the second selection result 672 selected as satisfying the second selection condition 662. The editing unit 39 also performs editing processing on the pair of the integrated sentence IS and the common sentence CS from the second selection result 672 selected as satisfying the second selection condition 662.

[0119] In this way, the first set selection unit 701 performs a first selection to select pairs of existing input data sentences ESS and existing supervised data sentences ETS whose recall rates satisfy the first selection condition 661. The second set selection unit 702 derives an integrated recall IRC between a common sentence CS, which is an existing supervised data sentence ETS selected in common for two or more existing input data sentences ESS, and an integrated sentence IS obtained by integrating two or more existing input data sentences ESS. The second set selection unit 702 selects pairs of existing input data sentences ESS and existing supervised data sentences ETS whose recall rates RC satisfy a second selection condition 662 that is stricter than the first selection condition 661, and performs a second selection to select pairs of common sentences CS and integrated sentences IS whose integrated recall IRC satisfies the second selection condition 662. This makes it possible to prevent sentences that have a high recall RC when viewed individually but a low recall RC when viewed as an integrated sentence IS from being selected as a highly related sentence. More appropriate new training data 12N can be generated.

[0120] Conversely to the above embodiment, the recall RC of the existing correct answer data sentence ETS for the existing input data sentence ESS may be derived as the index INX. Furthermore, as the coverage, a precision may be derived instead of or in addition to the recall RC.

[0121] When new training data 12N is generated by performing the second and fourth editing processes as the editing processes, it is preferable to use a sentence order rearrangement model 75 shown in FIG. 33 as an example.

[0122] The sentence order sorting model 75 is a model that is responsible for the task of sorting the sentence order of input data 76, which is an input sentence, into a regular order and outputting the result as output sentences as output data 77. Figure 33 illustrates, as an example, a case in which input data 76, in which the sentence order of diary-like sentences consisting of eight sentences S1 to S8 shown in Figure 34 has been randomly changed, is input to the sentence order sorting model 75, and output data 77, in which the sentences have been sorted into the regular order of S1 to S8, is output from the sentence order sorting model 75.

[0123] 35 shows an example in which the sentence order reordering model 75 is applied to new training data 12N generated by performing the second editing process. In this case, the editing unit 39 uses the sentence order reordering model 75 to reorder the new input data sentences NSS that constitute the new training input data 16N of the new training data 12N into a regular order. The editing unit 39 also uses the sentence order reordering model 75 to reorder the new correct answer data sentences NTS that constitute the new correct answer data 17NCA of the new training data 12N into a regular order. In this way, the editing unit 39 converts the new training data 12N into reordered new training data 12AN.

[0124] 36 shows an example in which the sentence order rearrangement model 75 is applied to new training data 12N generated by performing the fourth editing process. In this case, as in the case of FIG. 35 , the editing unit 39 uses the sentence order rearrangement model 75 to rearrange the new input data sentence NSS and the new correct answer data sentence NTS into a regular order, and converts the new training data 12N into rearranged new training data 12AN.

[0125] The editing unit 39 compares the new training data 12N with the aligned new training data 12AN to derive the change rate 80. The change rate 80 is a numerical value that represents how much the arrangement of the new input data sentences NSS and the new correct answer data sentences NTS has changed between the new training data 12N and the aligned new training data 12AN. For example, if the total number of new input data sentences NSS and new correct answer data sentences NTS is 20 and the total number of new input data sentences NSS and new correct answer data sentences NTS whose arrangements have been changed is 4, the change rate 80 is (4 / 20) × 100 = 20%.

[0126] If the change rate 80 is less than a threshold, for example, less than 30%, the editing unit 39 determines that the sentence arrangement of the new training data 12N is close to the regular arrangement and adopts the new training data 12N. On the other hand, if the change rate 80 is equal to or greater than the threshold, the editing unit 39 determines that the sentence arrangement of the new training data 12N is far from the regular arrangement and does not adopt the new training data 12N.

[0127] In this way, when generating new training data 12N by performing the second editing process and the fourth editing process, the editing unit 39 uses a sentence order reordering model 75 that reorders the sentences of the input text. Therefore, as shown in Figure 35, the sentence order of the new training data 12N generated by the second editing process can be easily reordered to the correct order. Furthermore, as shown in Figure 36, it is possible to determine whether the sentence order of the new training data 12N is close to the correct order, and to adopt only new training data 12N that is determined to be close.

[0128] In the fourth editing process, multiple new training data 12N with different sentence orders are generated from one existing training data 12E. Then, the sentence order rearrangement model 75 may be applied to each of the multiple new training data 12N to output rearranged new training data 12AN, and the change rate 80 may be derived. In this case, as in the case of FIG. 36 , new training data 12N with a change rate 80 less than a threshold may be used, or only new training data 12N with the lowest change rate 80 may be used.

[0129] Although the sentence order rearrangement model 75 is applied to both the second editing process and the fourth editing process, this is not limiting. The editing process to which the sentence order rearrangement model 75 is applied may be at least one of the second editing process and the fourth editing process. Furthermore, when performing the sixth editing process shown in FIG. 26 in which an isolated sentence in some existing training data 12E is inserted into other existing training data 12E, the sentence order rearrangement model 75 may be used to determine the insertion position of the isolated sentence.

[0130] [Second_1 embodiment] As an example, as shown in FIG. 37, the CPU of the data extension device of the second_1 embodiment functions as a selection unit 85 in addition to the processing units 35 to 40 of the first embodiment (in FIG. 37, units other than the distribution control unit 40 are omitted).

[0131] The selection unit 85 receives input of the existing learning data group 12EG, the new learning data group 12NG, and selection conditions 86. The selection conditions 86 are stored in the storage 20A, read from the storage 20A by the RW control unit 35, and output to the selection unit 85. The selection unit 85 selects data to be officially adopted as learning data 12 from among multiple new learning data 12N candidates constituting the new learning data group 12NG in accordance with the selection conditions 86. The selection unit 85 outputs a selected new learning data group 12SNG, which is a collection of the selected new learning data 12N, to the distribution control unit 40. The distribution control unit 40 distributes and outputs a learning data group 12GX composed of the existing learning data group 12EG and the selected new learning data group 12SNG to the learning device 11.

[0132] As shown in FIG. 38 as an example, the selection unit 85 includes a vectorization unit 90, a first calculation unit 911, a second calculation unit 912, and a determination unit 92. The vectorization unit 90 is a repurposed vectorization unit of a well-known language model, such as BERT. The vectorization unit 90 vectorizes the existing input data sentence ESS and the existing correct answer data sentence ETS included in each existing training data 12E in the existing training data group 12EG, and the new input data sentence NSS and the new correct answer data sentence NTS included in each new training data 12N in the new training data group 12NG. Sentence vectorization is a process of converting the words and their order that make up a sentence into a multidimensional, e.g., 512-dimensional, feature vector composed of a list of multiple types of features. The vectorization unit 90 integrates the feature vectors of the existing input data sentence ESS and the existing correct answer data sentence ETS into a single representative feature vector. Similarly, the vectorization unit 90 integrates the feature vectors of the new input data sentence NSS and the new correct answer data sentence NTS into one representative feature vector. The feature vector referred to in the following description refers to this representative feature vector. The vectorization unit 90 outputs the vectorized existing training data group 12VEG, which is the vectorized existing training data group 12EG, and the vectorized new training data group 12VNG, which is the vectorized new training data group 12NG, to the first calculation unit 911 and the second calculation unit 912.

[0133] The first calculation unit 911 calculates a first deviation degree of each new learning data 12N based on the vectorized existing learning data group 12VEG and the vectorized new learning data group 12VNG. The first deviation degree is a numerical value representing the degree of deviation between the target new learning data 12N and the existing learning data 12E from which the target new learning data 12N is derived and the new learning data 12N other than the target new learning data 12N. The first calculation unit 911 outputs a first calculation result 951, which is a calculation result of the first deviation degree, to the determination unit 92. The first calculation result 951 is a first deviation degree registered for each new learning data ID that uniquely identifies the new learning data 12N.

[0134] The second calculation unit 912 calculates a second deviation degree of each new learning data 12N based on the vectorized existing learning data group 12VEG and the vectorized new learning data group 12VNG. The second deviation degree is a numerical value representing the degree of deviation between the target new learning data 12N and the set of existing learning data 12E. The second calculation unit 912 outputs a second calculation result 952, which is a calculation result of the second deviation degree, to the determination unit 92. The second calculation result 952 is a second deviation degree registered for each new learning data ID.

[0135] The determination unit 92 determines, based on the first calculation result 951 and the second calculation result 952, whether or not the target new learning data 12N may be officially adopted as learning data 12 in accordance with the selection condition 86. The determination unit 92 includes the new learning data 12N determined to be officially adopted as learning data 12 in the selected new learning data group 12SNG.

[0136] As an example, as shown in graph 100 in FIG. 39 , the first calculation unit 911 calculates a cosine similarity θ. More specifically, the first calculation unit 911 calculates, in the feature vector space 101, a cosine similarity θ1 between the feature vector of the target new training data 12N and the feature vector of the existing training data 12E from which the target new training data 12N was derived. The first calculation unit 911 also calculates, in the feature vector space 101, a cosine similarity θ2 between the feature vector of the target new training data 12N and the feature vector of new training data 12N other than the target new training data 12N. Although only one cosine similarity θ2 is calculated in FIG. 39 , the first calculation unit 911 actually calculates the cosine similarity θ2 for each feature vector of multiple new training data 12N other than the target new training data 12N. The first calculation unit 911 determines, for example, the average value θAVE of the cosine similarities θ1 and θ2 as the first deviation. That is, the average value θAVE is an example of a “first deviation” according to the technique of the present disclosure.

[0137] As an example, as shown in graph 102 in FIG. 40 , the second calculation unit 912 calculates the Euclidean distance D. More specifically, the second calculation unit 912 calculates the Euclidean distance D between the features of the target new training data 12N and the centroid features of the set 103 of features of the existing training data 12E in the feature vector space 101. The second calculation unit 912 sets this Euclidean distance D as the second deviation. That is, the Euclidean distance D is an example of the "second deviation" according to the technology of the present disclosure. Note that, for convenience of explanation, in FIGS. 39 and 40 , the dimension of the feature vector space 101 is two-dimensional, having axes D1 and D2, but the dimension of the actual feature vector space 101 is, for example, 512, as described above.

[0138] As an example, as shown in Figure 41, the selection condition 86 is that the difference between 1 and the average value θAVE of the cosine similarity is 0.7 or more (when the average value θAVE of the cosine similarity is positive), or the difference between -1 and the average value θAVE of the cosine similarity is 0.7 or more (when the average value θAVE of the cosine similarity is negative), and the Euclidean distance D is less than a predetermined distance threshold.

[0139] As described above, in the second embodiment, the selection unit 85 selects data to be officially adopted as the learning data 12 from among multiple candidates of new learning data 12N generated by the editing process. This reduces the risk that inappropriate new learning data 12N will be adopted as the learning data 12.

[0140] The selection unit 85 performs the selection based on the average value θAVE of the cosine similarities, which is a first degree of deviation between the target new training data 12N and the existing training data 12E that is the source of the target new training data 12N and the new training data 12N other than the target new training data 12N, and on the Euclidean distance D, which is a second degree of deviation between the target new training data 12N and the set of existing training data 12E.

[0141] The average value θAVE of the cosine similarity, which is the first deviation, represents the degree of variability of the target new training data 12N relative to the original existing training data 12E and other new training data 12N. If there is not much difference between the original existing training data 12E and other new training data 12N and the target new training data 12N, there is little point in officially adopting the new training data 12N as training data 12. Therefore, in the second_1 embodiment, selection is performed based on the average value θAVE of the cosine similarity, which is the first deviation, and new training data 12N that is not much different from the original existing training data 12E and other new training data 12N is excluded from being officially adopted.

[0142] On the other hand, the Euclidean distance D, which is the second deviation, represents the degree of validity of the target new training data 12N relative to the existing training data 12E. Even if the existing training data 12E and the target new training data 12N are too similar, this is undesirable from the viewpoint of variability. However, even if there is a large difference between the existing training data 12E and the target new training data 12N, this is also undesirable from the viewpoint of validity. Therefore, in the second_1 embodiment, selection is performed based on the Euclidean distance D, which is the second deviation, and new training data 12N that is significantly different from the existing training data 12E is excluded from being officially adopted. In summary, according to the second_1 embodiment, more appropriate new training data 12N can be selected as the training data 12 to be officially adopted.

[0143] Note that, instead of the cosine similarity θ, the Euclidean distance D or the Mahalanobis distance may be calculated as the first deviation. Similarly, instead of the Euclidean distance D, the Mahalanobis distance or the cosine similarity θ may be calculated as the second deviation.

[0144] [Embodiment 2_2] As an example, as shown in FIG. 42, the CPU of the data extension device of embodiment 2_2 functions as a selection unit 105 in addition to the processing units 35 to 40 of the first embodiment (in FIG. 42, units other than the distribution control unit 40 are omitted).

[0145] The selection unit 105 receives input of the new training data group 12NG, a sentence semantic determination model 106, and selection conditions 107. The sentence semantic determination model 106 and the selection conditions 107 are stored in the storage 20A, read from the storage 20A by the RW control unit 35, and output to the selection unit 105. The selection unit 105 uses the sentence semantic determination model 106 and in accordance with the selection conditions 107 to select data to be officially adopted as training data 12 from among multiple new training data 12N candidates constituting the new training data group 12NG. The selection unit 105 outputs a selected new training data group 12SNG, which is a collection of the selected new training data 12N, to the distribution control unit 40. The distribution control unit 40 distributes and outputs a training data group 12GX composed of the existing training data group 12EG and the selected new training data group 12SNG to the learning device 11.

[0146] As an example, as shown in Figures 43 and 44, the sentence meaning determination model 106 is a machine learning model that outputs a score 109 that indicates the clarity of meaning of input data 108, which is an input sentence. The score 109 takes a value between 0 and 1. As shown in Figure 43, if the input data 108 is a meaningful sentence, the sentence meaning determination model 106 outputs a score 109 that is close to 1. On the other hand, as shown in Figure 44, if the input data 108 is an meaningless sentence, the sentence meaning determination model 106 outputs a score 109 that is close to 0.

[0147] 45 as an example, the selection unit 105 inputs new training input data 16N and new supervised data 17NCA of multiple new training data 12N constituting a new training data group 12NG to a sentence semantic determination model 106 as input data 108. Then, the selection unit 105 outputs a score 109 from the sentence semantic determination model 106. The selection unit 105 calculates the arithmetic mean of the score 109 of the new training input data 16N and the score 109 of the new supervised data 17NCA, and sets this as the average score 109AVE of the new training data 12N to be compared with the selection condition 107. The selection unit 105 calculates the average score 109AVE of each of the multiple new training data 12N. FIG. 45 illustrates an example in which the score 109 of the new learning input data 16N is 0.92, the score 109 of the new correct answer data 17NCA is 0.86, and the average score 109AVE is 0.89.

[0148] 46, as an example, the selection condition 107 is that the average score 109AVE is 0.8 or more. Therefore, the new training data 12N in the example shown in FIG. 45 is selected by the selection unit 105 as data to be officially adopted.

[0149] As described above, in the second embodiment, the selection unit 105 performs selection using the sentence semantic determination model 106 that outputs the score 109 indicating the clarity of the meaning of the input data 108, which is the input sentence. This reduces the possibility that new training data 12N including meaningless sentences will be selected, which will hinder the training of the language model 15.

[0150] Note that the 2_1 embodiment and the 2_2 embodiment may be implemented in combination. For example, the 2_1 embodiment may be implemented first to narrow down the new training data 12N, and then the 2_2 embodiment may be implemented to further narrow down the new training data 12N. Furthermore, the new training data 12N to be officially adopted may be selected randomly without implementing the 2_1 embodiment and / or the 2_2 embodiment.

[0151] [Embodiment 3_1] In embodiment 3_1, the order in which existing training data 12E and new training data 12N are provided to the language model 15 in the training phase is defined. That is, as an example, as shown in Fig. 47 , the training unit 56 provides new training data 12N to the language model 15 in the first stage of the training phase, and provides existing training data 12E to the language model 15 in the second stage after the first stage.

[0152] As described above, in the third embodiment, the learning unit 56 learns the language model 15 using the new training data 12N, and then learns the language model 15 using the existing training data 12E. The new training data 12N is generated by processing the existing training data 12E, and therefore is of lower quality than the existing training data 12E. Therefore, the performance of the language model 15 can be further improved by first learning the language model 15 using the new training data 12N of relatively lower quality, improving the performance of the language model 15 to a certain extent, and then learning the language model 15 using the existing training data 12E of relatively better quality.

[0153] [Third_2nd Embodiment] As an example, as shown in FIG. 48 , new learning data 12N is classified into new learning data 12N_1 generated by a first editing process, new learning data 12N_2 generated by a second editing process, new learning data 12N_3 generated by a third editing process, and new learning data 12N_4 generated by a fourth editing process.

[0154] Of these new training data 12N_1 to 12N_4, the new training data 12N_1 generated by the first editing process has the lowest learning load on the language model 15. The new training data 12N_2 generated by the second editing process has the next lowest learning load on the language model 15, and the new training data 12N_4 generated by the fourth editing process has the next lowest learning load on the language model 15. Finally, the new training data 12N_3 generated by the third editing process has the highest learning load on the language model 15.

[0155] As shown in FIGS. 12 and 13 , the new training data 12N_1 generated by the first editing process is generated from only the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing training data 12E. In other words, the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS remain in the new training data 12N_1. Furthermore, unlike the second editing process, multiple existing training data 12E are not integrated, and the order of the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS is not changed, as in the fourth editing process. Therefore, it can be said that the training load on the language model 15 is the lowest. Furthermore, as shown in FIG. 14 , the new training data 12N_2 generated by the second editing process is generated by integrating the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from multiple existing training data 12E. That is, like the new training data 12N_1, the pair of the selected existing input data sentence ESS and the existing supervised data sentence ETS remains in the new training data 12N_2. Furthermore, the order of the pair of the selected existing input data sentence ESS and the existing supervised data sentence ETS is not changed as in the fourth editing process. For this reason, the training load on the language model 15 can be said to be the second lowest.

[0156] As shown in FIG. 16 , new training data 12N_4 generated by the fourth editing process is generated by changing the order of the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing training data 12E. Therefore, although the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS remain in the new training data 12N_4, the changed order makes prediction more difficult. Therefore, it can be said that the training load on the language model 15 is the third lowest. Also, as shown in FIG. 15 , new training data 12N_3 generated by the third editing process is generated by deleting the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS from the existing training data 12E. Therefore, unlike the new training data 12N_1, etc., the pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS do not remain in the new training data 12N_3. Therefore, it can be said that the training load on the language model 15 is the highest. Generally speaking, new training data 12N has a relatively low learning load if it contains pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS and has not been processed by integrating multiple existing training data 12E or by changing the order of pairs of selected existing input data sentences ESS and existing correct answer data sentences ETS. In other words, new training data 12N has a relatively low learning load if it maintains the relationship between the existing input data sentences ESS and existing correct answer data sentences ETS in the existing training data 12E.

[0157] As an example, as shown in FIG. 49 , the learning unit 56 divides new training data 12N_1 to 12N_4 into stages and provides them to the language model 15 for training. The learning unit 56 also provides new training data 12N to the language model 15 in order, starting with the new training data 12N with a relatively low training load, for training. That is, the learning unit 56 provides new training data 12N_1 to the language model 15 in stage 1_1 of the learning phase, and provides new training data 12N_2 to the language model 15 in stage 1_2. Next, the learning unit 56 provides new training data 12N_4 to the language model 15 in stage 1_3, and provides new training data 12N_3 to the language model 15 in stage 1_4. Finally, the learning unit 56 provides existing training data 12E to the language model 15 in stage 2.

[0158] As described above, in the third embodiment, the learning unit 56 performs learning by dividing the new training data 12N_1 to 12N_4 generated in each of the first to fourth editing processes into stages and providing the new training data 12N generated by one editing process in a concentrated manner to the language model 15. This allows the language model 15 to be trained. Compared to the case where the new training data 12N_1 to 12N_4 are randomly provided, the performance of the language model 15 can be further improved.

[0159] Furthermore, in the third embodiment, the learning unit 56 performs learning by providing new training data 12N to the language model 15 in order, starting with the new training data 12N that has a relatively low learning load. This allows the level of the language model 15 to be gradually improved. Compared to the case where new training data 12N is provided to the language model 15 in order, starting with the new training data 12N that has a relatively high learning load, the performance of the language model 15 can be further improved.

[0160] Note that the language model 15 may be trained using only the new training data 12N, without using the existing training data 12E.

[0161] Fourth Embodiment As an example, as shown in FIG. 50, in the fourth embodiment, the pair selection unit 38 selects pairs of an existing input data sentence ESS and an existing supervised data sentence ETS using a pair selection determination model 110.

[0162] The pair selection judgment model 110 outputs a judgment result 111 in response to an input of a pair of an existing input data sentence ESS and an existing supervised data sentence ETS to be judged. The judgment result 111 indicates whether or not to select the pair of the existing input data sentence ESS and the existing supervised data sentence ETS to be judged. The pair selection unit 38 outputs the judgment result 111 to the editing unit 39.

[0163] The pair selection determination model 110 is a machine learning model that outputs, for example, a numerical value between 0 and 1 as a preliminary step before outputting a determination result 111. If this numerical value is equal to or greater than a predetermined determination threshold (for example, 0.5), the pair selection determination model 110 outputs a determination result 111 indicating "select." On the other hand, if the numerical value is less than the determination threshold, the pair selection determination model 110 outputs a determination result 111 indicating "do not select." The numerical value between 0 and 1 that is output as a preliminary step before outputting the determination result 111 is an example of an "index" according to the technology of the present disclosure. Furthermore, the determination threshold is an example of a "preset condition" according to the technology of the present disclosure.

[0164] As an example, as shown in FIG. 51, in the fourth embodiment, first, as shown in FIG. 50, in the pair selection unit 38, a pair of an existing input data sentence ESS to be judged and an existing correct data sentence ETS is input to the pair selection judgment model 110, and a judgment result 111 is output from the pair selection judgment model 110 (step ST300).

[0165] If the determination result 111 indicates "select" (YES in step ST310), the editing unit 39 performs an editing process on the pair of the existing input data sentence ESS and the existing supervised data sentence ETS to be determined on the existing learning data 12E, and generates new learning data 12N (step ST320). On the other hand, if the determination result 111 indicates "do not select" (NO in step ST310), the process returns to step ST300.

[0166] Next, the learning unit 56 learns the language model 15 using the new learning data 12N generated in step ST320 (step ST330), and then the performance of the language model 15 is evaluated (step ST340).

[0167] If the performance of the language model 15 has improved since training using the new training data 12N in step ST330 (YES in step ST350), the learning unit 56 employs the process of outputting the determination result 111 of the pair selection determination model 110 in step ST300 (step ST360). On the other hand, if the performance of the language model 15 has deteriorated since training using the new training data 12N in step ST330 (NO in step ST350), the learning unit 56 does not employ the process of outputting the determination result 111 of the pair selection determination model 110 in step ST300 (step ST370). The processes of steps ST300 to ST360 or step ST370 are repeated until the performance of the language model 15 reaches the set level (NO in step ST380).

[0168] When the performance of the language model 15 reaches the set level (YES in step ST380), the learning unit 56 outputs the language model 15 to the RW control unit 55. Then, under the control of the RW control unit 55, the language model 15 from the learning unit 56 is stored in the storage 20B as a trained language model 15 (step ST390).

[0169] As described above, in the fourth embodiment, the pair selection unit 38 uses the pair selection determination model 110 to determine whether to select a pair of the existing input data sentence ESS and the existing correct answer data sentence ETS to be determined. If the determination result 111 of the pair selection determination model 110 indicates "select," the editing unit 39 performs an editing process for the pair of the existing input data sentence ESS and the existing correct answer data sentence ETS to be determined on the existing training data 12E, thereby generating new training data 12N. If the performance of the language model 15 has improved compared to before training using the new training data 12N, the learning unit 56 adopts the process for outputting the determination result 111 of the pair selection determination model 110. On the other hand, if the performance of the language model 15 has deteriorated compared to before training using the new training data 12N, the learning unit 56 does not adopt the process for outputting the determination result 111 of the pair selection determination model 110.

[0170] In this way, by repeatedly determining whether to adopt the output process of the determination result 111 of the pair selection determination model 110 depending on whether the performance of the language model 15 has improved, it is possible to gradually improve the determination performance of the pair selection determination model 110. Ultimately, it is possible to obtain a pair selection determination model 110 that can accurately select pairs of existing input data sentences ESS and existing correct answer data sentences ETS that have a relatively high relationship. This makes it possible to omit the cumbersome process of deriving an index INX such as the superficial similarity SPS in the first embodiment, thereby significantly reducing the processing load on the CPU 22A.

[0171] The pairs input to the pair selection determination model 110 are not limited to the example pair of one existing input data sentence ESS and one existing supervised data sentence ETS. They may be pairs of multiple existing input data sentences ESS and one existing supervised data sentence ETS, as shown in Fig. 22, or pairs of one existing input data sentence ESS and multiple existing supervised data sentences ETS. Furthermore, they may be pairs of existing input data sentences ESS and pairs of existing supervised data sentences ETS, as shown in Fig. 24.

[0172] In the above embodiments, the input data 16 and the output data 17 are both text, but this is not limiting. As an example, as shown in Fig. 52, the input data 16 may be numerical data consisting of a list of output values ​​of some sensor, and the output data 17 may be text stating whether the sensor output value is normal or abnormal. In this case, the first component is a list of output values ​​of the sensor.

[0173] As an example, as shown in Fig. 53, the input data 16 may be speech recognition data obtained by recognizing and transcribing speech from conference participants, and the output data 17 may be a document of minutes that summarizes the speech. Furthermore, contrary to the example of Fig. 53, the input data 16 may be a document such as a draft of a conference presentation, and the output data 17 may be audio data such as speech generated by artificial intelligence.

[0174] The first and second components are not limited to the example sentences. They may be sentences that make up a paragraph or a page. They may also be sentences written on the same date, by the same author, or with the same character style, such as font size or color. They may also be sentences that are separated by a specific delimiter, such as #.

[0175] The task performed by the language model 15 may be a translation task. Alternatively, multiple documents may be used as input data 16, and a summary of the multiple documents may be used as output data 17. Examples of the multiple documents include multiple factory maintenance reports. Examples of the multiple documents include multiple electronic medical records of a single patient, and a summary may include a discharge summary of the patient. When multiple documents are used as input data 16, the first component may be a single document.

[0176] The hardware configuration of the computers that constitute the data expansion device 10 and learning device 11 according to the technology of the present disclosure can be modified in various ways. For example, the data expansion device 10 can be configured with multiple computers separated as hardware in order to improve processing power and reliability. For example, the functions of the sentence division unit 36 ​​and the index derivation unit 37 and the functions of the pair selection unit 38 and the editing unit 39 can be distributed and performed by two computers. In this case, the data expansion device 10 is configured with two computers.

[0177] In this way, the hardware configuration of the computers of the data expansion device 10 and the learning device 11 can be changed as appropriate depending on the required performance, such as processing power, safety, and reliability. Furthermore, not only the hardware, but also the application programs such as the operating programs 30 and 50 can be duplicated or stored in multiple storage devices in order to ensure safety and reliability.

[0178] In each of the above embodiments, for example, the hardware structure of the processing units that execute various processes, such as the RW control units 35 and 55, the sentence splitting unit 36, the index derivation unit 37, the set selection unit 38, the editing unit 39, the distribution control unit 40, the learning unit 56, the first set selection unit 701, the second set selection unit 702, the selection units 85 and 105, the vectorization unit 90, the first calculation unit 911, the second calculation unit 912, and the determination unit 92, can be various processors shown below. As described above, the various processors include the CPUs 22A and 22B, which are general-purpose processors that execute software (operating programs 30 and 50) and function as various processing units, as well as programmable logic devices (PLDs) that are processors whose circuit configuration can be changed after manufacture, such as FPGAs (Field Programmable Gate Arrays), and dedicated electrical circuits that are processors having a circuit configuration designed exclusively for executing specific processing, such as ASICs (Application Specific Integrated Circuits).

[0179] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA).Furthermore, multiple processing units may be configured with a single processor.

[0180] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of the entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.

[0181] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0182] From the above description, the technology described in the following supplementary paragraphs can be understood.

[0183] [Supplementary Item 1] A learning support device comprising: a processor that acquires existing training data of a language model consisting of input data and supervised answer data, at least one of which is a sentence, the input data including a plurality of first components and supervised answer data including a plurality of second components, deriving an index representing a relationship between the first components and the second components, selecting pairs of the first components and the second components for which the index satisfies a predetermined condition, and generating new training data by performing an editing process for the selected pairs on the existing training data. [Supplementary Item 2] The learning support device according to Supplementary Item 1, wherein the index is at least one of: a superficial similarity between the first components and the second components, a semantic similarity between the first components and the second components, a keyword similarity between the first components and the second components, a numerical value corresponding to a result of topic classification of the first components and the second components, a numerical value corresponding to a result of determination of an entailment relationship between the first components and the second components, and a coverage rate of one of the first components and the second components relative to the other. [Supplementary Item 3] The learning support device according to Supplementary Item 2, wherein, when the index is the coverage rate, the processor performs a first selection to select pairs of the first and second components whose coverage rate satisfies a first selection condition, derives an integrated coverage rate of an integrated component obtained by integrating the other two or more of the first and second components, for a common component that is one of the first and second components and that is selected in common with the other two or more of the first and second components, and performs a second selection to select pairs of the first and second components whose coverage rate satisfies a second selection condition that is stricter than the first selection condition, and selects pairs of the common components and the integrated components whose integrated coverage rate satisfies the second selection condition. [Supplementary Item 4] The learning support device according to any one of Supplementary Items 1 to 3, wherein the processor derives the index for one of the first and second components and the index for two or more of the first components and one of the second components.[Supplementary Item 5] The learning assistance device according to any one of Supplementary Items 1 to 4, wherein the first constituent and the second constituent are sentences. [Supplementary Item 6] The learning assistance device according to any one of Supplementary Items 1 to 5, wherein the processor performs, as the editing process, at least one of: a first editing process of generating the new learning data using only pairs of the first constituent and the second constituent selected from the existing learning data, a second editing process of integrating pairs of the selected first constituent and the second constituent from a plurality of the existing learning data to generate the new learning data, a third editing process of deleting pairs of the selected first constituent and the second constituent from the existing learning data to generate the new learning data, and a fourth editing process of changing the order of pairs of the selected first constituent and the second constituent from the existing learning data to generate the new learning data. [Supplementary Item 7] The learning support device of Supplementary Item 6, wherein the processor uses a sentence order reordering model that reorders sentences of an input sentence when generating the new training data by performing at least one of the second editing process and the fourth editing process. [Supplementary Item 8] The learning support device of any one of Supplementary Items 1 to 7, wherein the processor selects data to be officially adopted as training data from a plurality of new training data candidates generated by the editing process. [Supplementary Item 9] The learning support device of Supplementary Item 8, wherein the processor makes the selection based on a first degree of deviation between the target new training data and the existing training data from which the target new training data was derived and new training data other than the target new training data, and a second degree of deviation between the target new training data and the set of existing training data. [Supplementary Item 10] The learning support device of Supplementary Item 8 or Supplementary Item 9, wherein the processor makes the selection using a sentence semantic determination model that outputs a score representing clarity of meaning of the input sentence. [Supplementary Item 11] The learning assistance device according to any one of Supplementary Items 1 to 10, wherein the language model is a model that is responsible for the task of converting input sentences into output sentences of different styles.[Supplementary Item 12] The learning assistance device according to any one of Supplementary Items 1 to 11, wherein the processor trains the language model using at least the new training data. [Supplementary Item 13] The learning assistance device according to Supplementary Item 12, wherein the processor trains the language model using the new training data, and then trains the language model using the existing training data. [Supplementary Item 14] The learning assistance device according to Supplementary Item 12 or 13, wherein there are a plurality of types of editing processes, and the processor divides the new training data generated for each of the plurality of types of editing processes in stages and provides the new training data to the language model for learning. [Supplementary Item 15] The learning assistance device according to Supplementary Item 14, wherein the processor provides the new training data to the language model in order of new training data with a relatively low learning load, for learning. [Supplementary Item 16] The learning support device according to any one of Supplementary Items 1 to 15, wherein the processor uses a pair selection judgment model to judge whether or not to select a pair of the first component and the second component to be judged, and when the judgment result of the pair selection judgment model is to select, generates the new training data by applying the editing process related to the pair of the first component and the second component to be judged to the existing training data, trains the language model using the generated new training data, and when performance of the language model improves compared to before training using the new training data, adopts a process for outputting the judgment result of the pair selection judgment model, and when performance of the language model deteriorates compared to before training using the new training data, does not adopt the process for outputting the judgment result of the pair selection judgment model.

[0184] The technology of the present disclosure can be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, it is not limited to the above-described embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure extends not only to programs, but also to storage media that non-temporarily store programs, and computer program products that include programs.

[0185] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0186] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0187] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A learning support device comprising a processor, the processor obtaining existing learning data of a language model composed of input data and correct answer data, at least one of which is text, the input data including a plurality of first components and the correct answer data including a plurality of second components, deriving an index representing the relationship between the first component and the second component, selecting a pair of the first component and the second component that satisfies a preset condition, and generating new learning data by performing an editing process on the existing learning data related to the selected pair.

2. The learning support device according to claim 1, wherein the index is at least any one of a surface similarity between the first component and the second component, a semantic similarity between the first component and the second component, a keyword similarity between the first component and the second component, a numerical value according to a topic classification result of the first component and the second component, a numerical value according to a determination result of an implicative relationship between the first component and the second component, and a coverage rate of the other with respect to one of the first component and the second component.

3. When the index is the coverage rate, the processor performs a first selection of selecting a pair of the first component and the second component that satisfies a first selection condition, and for a common component that is one of the first component and the second component and is commonly selected by two or more of the other components of the first component and the second component, derives an integrated coverage rate with an integrated component obtained by integrating the two or more of the other components, and performs a second selection of selecting a pair of the first component and the second component that satisfies a second selection condition that is stricter than the first selection condition, and selecting a pair of the common component and the integrated component that satisfies the second selection condition.

4. The learning support device according to claim 1, wherein the processor derives the index of one first component and one second component and the index of two or more first components and one second component.

5. The learning support device according to claim 1, wherein the first component and the second component are sentences.

6. The processor performs at least any one of the following as the editing process: a first editing process of generating the new learning data only with the selected pair of the first component and the second component among the existing learning data; a second editing process of integrating the selected pairs of the first component and the second component of the plurality of existing learning data to generate the new learning data; a third editing process of deleting the selected pair of the first component and the second component among the existing learning data to generate the new learning data; and a fourth editing process of changing the order of the selected pair of the first component and the second component among the existing learning data to generate the new learning data. The learning support device according to claim 1.

7. When the processor generates the new learning data by performing at least any one of the second editing process and the fourth editing process, the learning support device according to claim 6 uses a sentence order alignment model that arranges the order of sentences in the input text.

8. The processor selects data to be formally adopted as learning data from among a plurality of candidates for the new learning data generated by the editing process. The learning support device according to claim 1.

9. The processor makes the selection based on the first divergence between the target new learning data and the existing learning data that is the source of the target new learning data and the new learning data other than the target new learning data, and the second divergence between the target new learning data and the set of existing learning data. The learning support device according to claim 8.

10. The processor makes the selection using a sentence meaning determination model that outputs a score representing the clarity of the meaning of the input text. The learning support device according to claim 8.

11. The language model is a model that undertakes the task of converting an input text into an output text in a different style. The learning support device according to claim 1.

12. The processor learns the language model using at least the new learning data. The learning support device according to claim 1.

13. After learning the language model using the new learning data, the processor learns the language model using the existing learning data. The learning support device according to claim 12.

14. The editing process has multiple types, and the processor is the learning support device according to claim 12, which gives the newly generated learning data for each of the multiple types of the editing process to the language model step by step for learning.

15. The processor is the learning support device according to claim 14, which gives the newly generated learning data to the language model for learning in order from the newly generated learning data with a relatively low learning load.

16. A method of operating a learning support device, including: obtaining existing learning data of a language model composed of input data and correct answer data, at least one of which is text, and including input data including a plurality of first components and correct answer data including a plurality of second components; deriving an index representing the relationship between the first component and the second component; selecting a pair of the first component and the second component that satisfies a preset condition for the index; and generating new learning data by performing an editing process related to the selected pair on the existing learning data.

17. A program for operating a learning support device that causes a computer to execute a process including: obtaining existing learning data of a language model composed of input data and correct answer data, at least one of which is text, and including input data including a plurality of first components and correct answer data including a plurality of second components; deriving an index representing the relationship between the first component and the second component; selecting a pair of the first component and the second component that satisfies a preset condition for the index; and generating new learning data by performing an editing process related to the selected pair on the existing learning data.

Citation Information

Patent Citations

  • Contradiction creation device, method, and program

    JP2017054434A

  • Disclosure apparatus, disclosure method, and disclosure program

    JP2020013395A

  • Data processing device, data processing method, and data processing program

    JP2022122029A

  • Method for question-answering based on asr

    KR102552401B1

  • KR20220162097A