Learning device, sentence splitter, learning method, sentence splitting method, and program

By rearranging training data to reduce word sequence similarity and using an encoder-decoder model with a reproducer, the proposed method improves split and rephrase performance, enabling more effective division of input sentences into semantically equivalent segments.

JP2025086163APending Publication Date: 2025-06-06NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023200043
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Conventional split and rephrase methods often fail to divide input sentences into semantically equivalent sentences, even when a human can perform the division, due to the similarity in word sequences between input sentences and correct segmented sentences in the training data.

Method used

The proposed method uses a training data reconstructor to rearrange correct segmented sentences in the training data, reducing the similarity in word sequences between input sentences and correct segmented sentences, and employs an encoder-decoder model with a reproducer to learn parameters that improve sentence segmentation performance.

Benefits of technology

The method enhances the performance of split and rephrase by increasing the number of divisions made on input sentences compared to conventional methods, effectively addressing the limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025086163000001_ABST
    Figure 2025086163000001_ABST
Patent Text Reader

Abstract

To improve performance of split-and-rephrase.SOLUTION: A learning device according to one aspect of the present disclosure includes: an input unit which receives a set of training data represented by pairs of an input sentence and first correct answer split sentences each of which comprises a plurality of sentences semantically equivalent to the input sentence; a generation unit which generates second correct answer split sentences obtained by rearranging sentences included in the first correct answer split sentences included in the training data; a sentence splitting unit which generates split sentences obtained by splitting the input sentence into a plurality of sentences, by an encoder / decoder model; and a learning unit which uses the split sentences and the second correct answer split sentences to learn parameters of the encoder / decoder model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a learning device, a sentence segmentation device, a learning method, a sentence segmentation method, and a program. [Background technology]

[0002] The technique of splitting a sentence into multiple sentences that are semantically equivalent to the original sentence is called split and rephrase (Non-Patent Document 1). This technique is realized by an encoder-decoder model using a neural network. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Shashi Narayan, Claire Gardent, Shay B. Cohen, and Anastasia Shimorina, Split and rephrase. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 606-616, Summary of the Invention [Problem to be solved by the invention]

[0004] However, even if a sentence that can be divided into multiple semantically equivalent sentences is given as an input sentence, the conventional split and rephrase method sometimes fails to divide the input sentence. It is a problem that a sentence cannot be divided even if a human can divide the input sentence, which goes against the purpose of split and rephrase.

[0005] The present disclosure has been made in consideration of the above points, and aims to improve the performance of split and rephrase. [Means for solving the problem]

[0006] A learning device according to an aspect of the present disclosure includes an input unit that inputs a set of training data represented by a pair of an input sentence and a first correct segmentation sentence composed of a plurality of sentences that are semantically equivalent to the input sentence, a creation unit that creates a second correct segmentation sentence in which the sentences included in the first correct segmentation sentence included in the training data are rearranged, a sentence segmentation unit that creates a segmentation sentence obtained by dividing the input sentence into a plurality of sentences by an encoder-decoder model, and a learning unit that learns parameters of the encoder-decoder model using the segmentation sentence and the second correct segmentation sentence.

Advantages of the Invention

[0007] The performance of split and rephrase can be improved.

Brief Description of the Drawings

[0008] [Figure 1] It is a diagram for explaining an example of the proposed method. [Diagram 2] It is a diagram showing an example of the hardware configuration of the sentence segmentation device according to the present embodiment. [Diagram 3] It is a diagram showing an example of the functional configuration of the sentence segmentation device during learning. [Figure 4] It is a flowchart showing an example of the learning process according to the present embodiment. [Diagram 5] It is a diagram showing an example of the functional configuration of the sentence segmentation device during inference. [Figure 6] It is a flowchart showing an example of the sentence segmentation process according to the present embodiment.

Modes for Carrying Out the Invention

[0009] Hereinafter, an embodiment of the present invention will be described.

[0010] <Conventional methods of split and rephrase and their problems> Split and rephrase is a technique for dividing a sentence into multiple sentences that are semantically equivalent to the original sentence (Non-Patent Document 1).

[0011] For example, suppose the following sentence is given as an input sentence:

[0012] He lives in Brooklyn, New York and is married to Melissa Block. In this case, in split and rephrase, for example, the above input sentence is split into the following two sentences (hereinafter also referred to as "split sentences").

[0013] He lives in Brooklyn, New York. He is married to Melissa Block. This technology is realized by an encoder-decoder model using a neural network. For example, it can be realized by additionally training a pre-trained model such as BART or T5 using a collection of training data (hereinafter also referred to as the "training dataset") represented as pairs of an input sentence and its correct divided sentence (hereinafter also referred to as the "correct divided sentence"). Here, the input sentence is the sentence input to the model (i.e., the sentence to be divided). On the other hand, the correct divided sentence is a set of multiple divided sentences that are the correct answer when the input sentence is divided, expressed by special symbols (for example, <s>For example, the input sentence "He lives in Brooklyn, New York and is married to Melissa Block." and the correct segmented sentence "He lives in Brooklyn, New York. <s>The pair "He is married to Melissa Block." is an example of training data. The model output is also divided into multiple sentences by special symbols (e.g., <s>) are connected together.

[0014] Split and rephrase using an encoder-decoder model achieves good performance in automatic evaluation metrics (e.g., BLEU) that measure the similarity between a correct segmentation created by a human and the segmentation output by the model. However, the number of sentences contained in a segmentation tends to be smaller than the number of sentences contained in the correct segmentation, and even if an input sentence is given that can be divided into multiple semantically equivalent sentences, the conventional split and rephrase method may not be able to divide it depending on the input sentence. It is a problem that a sentence cannot be divided even if a human can divide the input sentence.

[0015] The problem with the conventional split and rephrase method, that is, "depending on the input sentence, the segmentation may not be performed", is thought to be due to the fact that the training dataset used to learn the encoder-decoder model contains many input sentences and their correct segmented sentences that are very similar in terms of word sequences.

[0016] In other words, in the encoder-decoder model, the parameters are trained so that the word sequence of the segmented sentence estimated by the model matches the word sequence of the correct segmented sentence. In this case, if a large amount of training data is given that represents pairs of input sentences and correct segmented sentences that are similar in word sequence, the model will learn the parameters to estimate the input sentence as a segmented sentence as it is. Therefore, in this case, the input sentence will often be estimated as a segmented sentence as it is.

[0017] For example, the input sentence "He lives in Brooklyn, New York and is married Melissa Block." and the correct segmented sentence created by humans "He lives in Brooklyn, New York. <s>He is married to Melissa Block.' The correct split sentence deletes the 'and' after 'New York' in the input sentence and adds '. <s>Only three tokens, "He," have been inserted, and the word sequence is very similar to the input sentence. For this reason, even if the input sentence itself is output as a segmented sentence during model training, the loss (which may be called error) is small. Therefore, it is believed that such pairs of input sentences and correct segmented sentences that are very similar in word sequence are hindering the learning of the encoder-decoder model.

[0018] In recent years, in order to improve the performance of the encoder-decoder model, a reproducer that restores the input sentence from the hidden layer vector (hidden state) when generating the output sentence in the decoder is often used during training. Since the reproducer learns parameters to generate the input sentence from the hidden state, when using a reproducer, if a pair of an input sentence and a correct segmented sentence that are very similar in terms of word sequence are given as training data, there is a high possibility that the sentence will not be segmented.

[0019] <Proposed method> The above-mentioned problem existing in the conventional split and rephrase method is caused by a large amount of input sentence and correct divided sentence pairs that are similar in word sequence to each other being given as training data, and therefore it is considered that the problem can be solved by intentionally lowering the similarity in word sequence between the input sentence and the correct divided sentence. Hereinafter, the method proposed in this embodiment (hereinafter also referred to as the "proposed method") as a solution to this problem will be described with reference to FIG. 1. FIG. 1 is a diagram for explaining an example of the proposed method. Note that FIG. 1 describes a case where a reproducer is used when learning an encoder-decoder model. However, it is not essential to use a reproducer, and a reproducer does not have to be used when learning an encoder-decoder model.

[0020] As shown in Fig. 1, the proposed method uses a training data reconstructor 1100, an encoder-decoder model 1200, and a reconstructor 1300 during learning. Note that the encoder-decoder model 1200 and the reconstructor 1300 may be the same as those in the conventional method, and in the following, they are assumed to be models based on a transformer model.

[0021] The training data reconstructor 1100 rearranges the correct divided sentences included in the training data on a sentence-by-sentence basis. This makes it possible to intentionally reduce the similarity in the form of word sequences between the input sentence and the correct divided sentences included in the training data. Hereinafter, the input sentence included in the training data is denoted as x, the correct divided sentence of the input sentence is denoted as y, and the correct divided sentence after rearrangement by the training data reconstructor 1100 is denoted as y'. In particular, the input sentence included in the i-th training data is denoted as x, i , the correct divided sentence is y i Then, the correct divided sentence y i The correct divided sentence after rearranging the training data by the training data reconstructor 1100 is called y i The input sentence x, the correct divided sentence y, and the rearranged correct divided sentence y' are all represented as word sequences (which may also be called token sequences).

[0022] The encoder-decoder model 1200 receives an input sentence x and outputs divided sentences obtained by dividing the input sentence into multiple sentences. Hereinafter, in the text of this specification, the divided sentences output by the encoder-decoder model 1200 will be represented as "^Y". In particular, i If you want to specify that the sentence is a segmented sentence that the encoder-decoder model 1200 outputs when you enter i " should be written.

[0023] Here, the encoder-decoder model 1200 receives an input sentence x and generates a hidden state h X and an encoder 1210 that outputs the hidden state h X and the split sentence y generated up to the time t-1 before time t <t The encoder 1210 also includes a self-attention layer 1211 that converts the input sentence x into a vector sequence, and a hidden state h X On the other hand, the decoder 1220 includes a feedforward layer 1212 that converts the segmented sentences y <t The self-attention layer 1221 converts the vector sequence into a vector sequence, and the vector sequence output from the self-attention layer 1221 and the hidden state h X The attention layer 1222 outputs a vector sequence from the hidden state h Y and the hidden state h Y and an output layer 1224 that generates a divided sentence ^Y from the divided sentence y <1 For example, "a" may be a special symbol that indicates the beginning of a sentence.

[0024] The replicator 1300 calculates the hidden state h at time t. Y and the reproduced sentence x generated up to the time t-1 before time t <t The encoder-decoder model 1200 receives inputs x and x, and outputs a sentence that reproduces the input sentence x input to the encoder-decoder model 1200. In the text of this specification, the sentence output by the reproducer 1300 is denoted as "^X" and will be called a reproduced sentence. In particular, i To specify that the sentence is the one that the reproducing device 1300 outputs when the encoder-decoder model 1200 inputs the sentence, use "^X i ". Note that the reproduced sentence x at time t=1 is <1 For example, "a" may be a special symbol that indicates the beginning of a sentence.

[0025] Here, the reproducing device 1300 stores the reproduced text x <t The self-attention layer 1301 converts the vector sequence output from the self-attention layer 1301 and the hidden state h at time t. Y The system includes an attention layer 1302 that outputs a vector sequence from the attention layer 1302, a feedforward layer 1303 that converts the vector sequence output from the attention layer 1302 into a hidden state, and an output layer 1304 that generates the reproduced text ^X from the hidden state output from the feedforward layer 1303.

[0026] At this time, during learning, the parameters of the encoder-decoder model 1200 and the parameters of the reproducor 1300 are trained so as to reduce the error between the segmented sentence ^Y output from the decoder 1220 and the rearranged correct segmented sentence y', and also to reduce the error between the reproduced sentence ^X output from the reproducor 1300 and the reproduced sentence x. Hereinafter, the parameters of the encoder-decoder model 1200 are denoted as θ, and the parameters of the reproducor 1300 are denoted as γ.

[0027] On the other hand, during inference (i.e., during sentence division), the encoder-decoder model 1200 uses the learned parameter θ to estimate the divided sentence ^Y from the input sentence x.

[0028] The following describes the sentence segmentation device 10 that performs training of the encoder-decoder model 1200 and the reproducer 1300 and inference (sentence segmentation) using the encoder-decoder model 1200 according to the proposed method.

[0029] <Hardware configuration example of the sentence segmentation device 10> An example of the hardware configuration of the sentence segmentation device 10 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the hardware configuration of the sentence segmentation device 10 according to this embodiment.

[0030] 2, the sentence segmentation device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.

[0031] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the sentence segmentation device 10 does not necessarily have to have at least one of the input device 101 and the display device 102, for example.

[0032] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0033] The communication I / F 104 is an interface for connecting the sentence segmentation device 10 to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. The processor 108 is various types of arithmetic devices such as a CPU (Central Processing Unit) or a GPU (Graphic Processing Unit).

[0034] 2 is an example, and the hardware configuration of the sentence segmentation device 10 is not limited to this. For example, the sentence segmentation device 10 may have multiple auxiliary storage devices 107 and multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.

[0035] During study The learning of the encoder-decoder model 1200 and the reproducer 1300 will be described below. Here, the sentence segmentation device 10 at the time of learning is provided with a training dataset D={(x i ,y i )|i=1, ,N} is given. (x i ,y i ) is the i-th training data, and N is the number of training data. Note that the sentence segmentation device 10 during learning may be called, for example, a "learning device."

[0036] <Example of functional configuration of sentence segmentation device 10 during learning> An example of the functional configuration of the sentence segmentation device 10 during learning will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the functional configuration of the sentence segmentation device 10 during learning.

[0037] As shown in Fig. 3, the sentence segmentation device 10 during learning has an input unit 201, a training data reconstruction unit 202, a sentence segmentation unit 203, a reproduction unit 204, and a learning unit 205. Each of these units is realized, for example, by a process in which one or more programs installed in the sentence segmentation device 10 are executed by the processor 108 or the like. Furthermore, the sentence segmentation device 10 during learning has a model parameter storage unit 206. The model parameter storage unit 206 is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, the model parameter storage unit 206 may also be realized, for example, by a storage area of ​​a storage device or the like communicably connected to the sentence segmentation device 10.

[0038] The input unit 201 inputs a training data set D given to the sentence segmentation device 10 during learning.

[0039] The training data reconstructing unit 202 is realized by the training data reconstructor 1100, and rearranges the correct answer divided sentences y of each training data included in the training data set D on a sentence-by-sentence basis to generate correct answer divided sentences y′.

[0040] The sentence splitting unit 203 is realized by an encoder-decoder model 1200, and receives an input sentence x included in the training data as an input, and outputs a split sentence ^Y. Here, the sentence splitting unit 203 includes an encoding unit 211 and a decoding unit 212. The encoding unit 211 is realized by an encoder 1210, and receives an input sentence x as an input, and outputs a hidden state h X The decoding unit 212 outputs the hidden state h X It takes input and outputs divided sentence ^Y.

[0041] The reproducing unit 204 is realized by the reproducing device 1300, and the hidden state h at time t obtained when the decode unit 212 generates the divided sentence ^Y is Y and the reproduced sentence x generated up to the time t-1 before the time t <t It takes as input the reproduced sentence ^X and outputs it.

[0042] The learning unit 205 uses the correct divided sentence y', the divided sentence ^Y, the input sentence x, and the reproduced sentence ^X to learn the parameter θ of the encoder-decoder model 1200 and the parameter γ of the reproducer 1300. That is, the learning unit 205 updates the parameters θ and γ stored in the model parameter storage unit 206 so as to reduce the error between the divided sentence ^Y and the correct divided sentence y' and to reduce the error between the reproduced sentence ^X and the input sentence x.

[0043] The model parameter storage unit 206 stores the parameter θ of the encoder-decoder model 1200 and the parameter γ of the reproducer 1300 .

[0044] <Learning process> The learning process according to this embodiment will be described below with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the learning process according to this embodiment.

[0045] First, the input unit 201 receives a given training data set D={(x i ,y i ) |i=1, . . . , N} is input (step S101).

[0046] Next, the training data reconstructing unit 202 rearranges the correct answer segmented sentences y of each training data included in the training data set D input in the above step S101 into sentence units to create correct answer segmented sentences y' (step S102). The training data reconstructing unit 202 may rearrange the correct answer segmented sentences y into sentence units to create correct answer segmented sentences y', for example, by the following steps 1-1 and 1-2. In the following, a certain i-th correct answer segmented sentence y i is composed of k sentences, the correct divided sentence y i Rearrange the correct divided sentence y i We will explain how to create a k sentences in y i,1 ,···,y i,k Let the i-th correct segmentation sentence be y i =y i,1 <s>y i,k Let us assume that.

[0047] Step 1-1: The training data reconstructing unit 202 reconstructs the second and subsequent sentences y i,2 ,···,y i,k Randomly select one sentence from the list.

[0048] Step 1-2: The training data reconstructing unit 202 calculates y i,1 ,···,y i,k After randomly rearranging the k-1 sentences excluding the sentence selected in step 1-1 above, the rearranged k-1 sentences are arranged after the sentence selected in step 1-1 above. As a result, the rearranged correct divided sentence y i The reason why the first sentence is removed in step 1-1 above is because the correct divided sentence y i 'But y i,1 This is to avoid being at the top.

[0049] Next, the sentence division unit 203 and the reproduction unit 204 use the parameters θ and γ to output a divided sentence ^Y and a reproduced sentence ^X for each input sentence x included in each training data (step S103). That is, for each of i=1,...,N, the sentence division unit 203 uses the parameter θ to output the divided sentence ^Y and the reproduced sentence ^X for each input sentence x included in each training data. i Take as input the split sentence ^Y i Similarly, for each of i=1, . . . , N, the reproducing unit 204 outputs the parameter γ and the divided sentence ^Y i The hidden state h obtained at time t when generating Y Using and, reproduce sentence ^X i The sentence splitting unit 203 outputs the input sentence x by, for example, the following steps 2-1 and 2-2. i Take as input the split sentence ^Y i In addition, the reproducing unit 204 may output the divided sentence ^Y by the decoding unit 212 of the sentence dividing unit 203, for example, in the following procedure 3-1. i The hidden state h obtained at time t when generating Y Using this, reproduce the sentence ^X i In the following, for a given i-th input sentence x=x i Using this, the split sentence ^Y=^Y i and the reproduced sentence ^X i The case where the following is output will be described.

[0050] Step 2-1: The encoding unit 211 of the sentence splitting unit 203 takes the input sentence x as input and converts it into a hidden state h X That is, for example, we set the self-attention layer 1211 to Φ e,1 C , the feedforward layer 1212 is Φ e,2 C Then, the encoding unit 211 will X =Φ e,2 C (Φ e,1 C (x)) creates a hidden state h X Generate and output

[0051] Step 2-2: The decode unit 212 of the sentence splitting unit 203 starts from time t=1 and repeats the following steps 2-2-1 to 2-2-2 for each time t until a special symbol indicating the end of the sentence is generated in step 2-2-2, which will be described later.

[0052] Step 2-2-1: The decoding unit 212 decodes the hidden state h X The hidden state at time t is Y That is, for example, we set the self-attention layer 1221 to Φ d,1 C , attention layer 1222 is Φ d Z , the feedforward layer 1223 is Φ d,2 C Then, the decoder 212 will Y =Φ d,2 C (Φ d,1 C (y <t )+Φ d Z (Φ d,1 C (y <t ),h X )) gives the hidden state h at time t. Y Generate and output

[0053] Step 2-2-2: The decoding unit 212 decodes the hidden state h at time t. Y That is, the decoding unit 212 generates a divided sentence ^Y at time t from the hidden state h Y is input to the output layer 1224, and y <t The probability (posterior probability) of generating a word at time t when given is calculated, and a divided sentence ^Y at time t is generated by generating a word at time t according to this posterior probability.

[0054] Step 3-1: The reproducing unit 204 starts from time t=1 and repeats the following steps 3-1-1 to 3-1-3 for each time t until a special symbol indicating the end of the sentence is generated in step 3-1-3, which will be described later.

[0055] Step 3-1-1: The reproducing unit 204 reproduces the reproduced sentence x generated up to the time t-1, which is one time before the time t. <t Take as input the reproduced sentence x <t For example, we convert the self-attention layer 1301 into a vector sequence Φ r,1 C Then, the reproducing unit 204 can obtain Φ r,1 C (x <t ) to convert it into a vector sequence.

[0056] Step 3-1-2: The reproducing unit 204 reproduces the hidden state h at time t. Y and the vector sequence obtained in step 3-1-1 above to generate a hidden state at time t. For example, the attention layer 1302 is r Z , the feedforward layer 1303 is Φ r,2 C , the hidden state is h R Then, the reproducing unit 204 is R =Φ r,2 C (Φ r,1 C (x <t ),Φ r Z (Φ r,1 C (x <t ),h Y )) gives the hidden state h at time t. R Generate and output

[0057] Step 3-1-3: The reproducing unit 204 reproduces the hidden state h at time t. R In other words, the reproducing unit 204 generates a reproduced sentence ^X at time t from the hidden state h R is input to the output layer 1304, and x <t The probability of generating a word at time t when given is calculated (posterior probability), and the reproduced sentence ^X at time t is generated by generating a word at time t according to this posterior probability.

[0058] Next, the learning unit 205 i ' and the split sentence ^Y i and input sentence x i And, the reproduced sentence ^X i The parameter θ of the encoder-decoder model 1200 and the parameter γ of the reproducer 1300 are trained using the above (step S104). i and the correct divided sentence y i ' and reproduce the sentence ^X i and input sentence x i The parameters θ and γ stored in the model parameter storage unit 206 are updated so as to reduce the error with respect to the model θ and γ. The learning unit 205 may use a known optimization method to obtain the parameters θ and γ, for example, according to the following formula (1).

[0059]

number

[0060] Next, the learning unit 205 judges whether or not to end the learning process (step S105). For example, the learning unit 205 may judge to end the learning process if a predetermined end condition is satisfied, and may judge not to end the learning process if a predetermined end condition is not satisfied. Here, the end condition may be, for example, that the parameters θ and γ have converged.

[0061] If it is determined in step S105 that the learning process is to end, the learning unit 205 returns to step S102. On the other hand, if it is determined in step S105 that the learning process is to end, the learning unit 205 ends the learning process.

[0062] ·At the time of inference Hereinafter, an explanation will be given of inference of the encoder-decoder model 1200 using the trained parameter θ. Here, an input sentence x to be segmented is provided to the sentence segmentation device 10 during inference.

[0063] <Example of Functional Configuration of Sentence Segmentation Device 10 During Inference> An example of the functional configuration of the sentence segmentation device 10 at the time of inference will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the functional configuration of the sentence segmentation device at the time of inference.

[0064] As shown in Fig. 5, the sentence segmentation device 10 at the time of inference has an input unit 201, a sentence segmentation unit 203, and an output unit 207. Each of these units is realized, for example, by a process in which one or more programs installed in the sentence segmentation device 10 are executed by the processor 108 or the like. Furthermore, the sentence segmentation device 10 at the time of inference has a model parameter storage unit 206. The model parameter storage unit 206 is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, the model parameter storage unit 206 may also be realized, for example, by a storage area of ​​a storage device connected to the sentence segmentation device 10 so as to be able to communicate with it.

[0065] The input unit 201 inputs an input sentence x given to the sentence segmentation device 10 at the time of inference.

[0066] The sentence splitting unit 203 is realized by an encoder-decoder model 1200, and receives an input sentence x as an input and outputs a split sentence ^Y. Here, the sentence splitting unit 203 includes an encoding unit 211 and a decoding unit 212. The encoding unit 211 is realized by an encoder 1210, and receives an input sentence x as an input and outputs a hidden state h X The decoding unit 212 outputs the hidden state h X It takes input and outputs divided sentence ^Y.

[0067] The output unit 207 outputs the divided sentence ^Y to a predetermined output destination. Here, examples of the output destination include the display device 102 such as a display, the auxiliary storage device 107, and a terminal device connected via a communication network.

[0068] The model parameter storage unit 206 stores the parameter θ of the trained encoder-decoder model 1200 .

[0069] <Sentence division processing> The sentence division process according to this embodiment will be described below with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the sentence division process according to this embodiment.

[0070] First, the input unit 201 receives a given input sentence x (step S201).

[0071] Next, the sentence splitting unit 203 uses the learned parameter θ to output a divided sentence ^Y for the input sentence x (step S202). That is, the sentence splitting unit 203 uses the learned parameter θ to output a divided sentence ^Y with the input sentence x as input. The sentence splitting unit 203 may output a divided sentence ^Y with the input sentence x as input, for example, according to the above steps 2-1 to 2-2.

[0072] Then, the output unit 207 outputs the divided sentence ^Y to a predetermined output destination (step S203).

[0073] <Modification> A modification of this embodiment will now be described.

[0074] Variation 1 In the above embodiment, the sentence segmentation device 10 at the time of learning has the reproducing unit 204, but the sentence segmentation device 10 at the time of learning does not have to have the reproducing unit 204. In this case, the reproduced sentence ^X is not output in step S103 of FIG. 4. Also, if the above formula (1) is used when updating the parameter θ in step S104, λlogP rd (x i |h i Y ;γ) can be omitted.

[0075] Variation 2 When creating the correct divided sentence y' in step S102 of FIG. 4, the similarity as a word sequence may not decrease much in the above steps 1-1 and 1-2. Therefore, for example, the evaluation index for measuring the similarity as a word sequence is S(·,·), and a predetermined threshold value is th, and S(x i ,y i )-S(x i ,y i ') y if and only if th i ' is adopted as the correct segmented sentence after sorting, otherwise S(x i ,y i )-S(x i ,y i Steps 1-1 and 1-2 may be repeated until ') ≥ th is satisfied. This allows the correct segmented sentence y i It is possible to create only '.

[0076] <Summary> As described above, the sentence segmentation device 10 according to this embodiment optimizes at least the parameter θ of the encoder-decoder model 1200 by using the training data reconstructing unit 202 (training data reconstructor 1100) a correct answer segmented sentence that is a rearrangement of a correct answer segmented sentence created by a human being. By rearranging the original correct answer segmented sentence, it is possible to reduce the similarity in the word sequence between the input sentence and the correct answer segmented sentence, and it is possible to increase the rate at which the input sentence is segmented compared to conventional methods.

[0077] <Experiment> We conducted an experiment to compare the proposed method with the conventional method using a dataset in which the average number of divisions of sentences when input sentences were divided by humans was 2.03. When sentences were divided using the conventional encoder-decoder model, the average number of divisions was about 1.6. In contrast, when the encoder-decoder model 1200 was trained using the proposed method and then the sentences were divided using this encoder-decoder model 1200, the average number of divisions improved to 1.92. From this result, it is believed that the number of divisions will not decrease even when the Reproducer 1300 is introduced, and that the introduction of the Reproducer 1300 is expected to achieve higher quality divisions.

[0078] The present invention is not limited to the above-described embodiments specifically disclosed, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the scope of the claims. [Explanation of symbols]

[0079] 10 Sentence segmentation device 101 Input Device 102 Display device 103 External I / F 103a Recording media 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage 108 processors 109 Bus 201 Input section 202 Training Data Reconstruction Unit 203 Sentence division part 204 Recapitulation 205 Learning Department 206 Model parameter storage unit 207 Output section 211 Encoding section 212 Decoding section 1100 Training Data Reconstructor 1200 Encoder-Decoder Model 1210 Encoder 1211 Self-attention layer 1212 Feedforward Layer 1220 Decoder 1221 Self-attention layer 1222 Attention layer 1223 Feedforward Layer 1224 Output layer 1300 Reproducer 1301 Self-attention layer 1302 Attention layer 1303 Feedforward Layer 1304 Output layer< / s> < / s> < / s> < / s> < / s> < / s>

Claims

1. an input unit for inputting a set of training data represented by a set of an input sentence and a first correct divided sentence composed of a plurality of sentences semantically equivalent to the input sentence; A creating unit that creates a second correct answer segmented sentence by rearranging sentences included in a first correct answer segmented sentence included in the training data; a sentence splitting unit that splits the input sentence into a plurality of sentences using an encoder-decoder model; a learning unit that learns parameters of the encoder-decoder model by using the divided sentence and the second correct divided sentence; A learning device having the above configuration.

2. a reproducing unit that generates a reproduced sentence by reproducing the input sentence using a hidden state of a decoder included in the encoder-decoder model, The learning unit is The learning device according to claim 1 , further comprising: a learning unit configured to learn parameters of the reproducer using the input sentence and the reproduced sentence.

3. The creation unit is The learning device according to claim 1 or 2, wherein the second correct answer segmented sentence is created so that the similarity as a word sequence between the input sentence and the second correct answer segmented sentence is lower than the similarity as a word sequence between the input sentence and the first correct answer segmented sentence.

4. an input unit for inputting an input sentence to be segmented; a sentence splitting unit that creates split sentences by splitting the input sentence to be split into a plurality of sentences using an encoder-decoder model trained using a set of an input sentence and a second correct answer split sentence obtained by rearranging a first correct answer split sentence composed of a plurality of sentences semantically equivalent to the input sentence as training data; A sentence segmentation device having the above structure.

5. an input step of inputting a set of training data represented by a set of an input sentence and a first correct divided sentence composed of a plurality of sentences semantically equivalent to the input sentence; A generation step of generating a second correct answer segmented sentence by rearranging sentences included in a first correct answer segmented sentence included in the training data; a sentence segmentation step of generating a plurality of sentences by segmenting the input sentence into a plurality of sentences using an encoder-decoder model; a learning procedure for learning parameters of the encoder-decoder model using the segmented sentence and the second correct segmented sentence; The computer executes the learning method.

6. An input procedure for inputting an input sentence to be split; a sentence division step of creating divided sentences by dividing the input sentence to be divided into a plurality of sentences by an encoder-decoder model trained using a set of an input sentence and a second correct divided sentence obtained by rearranging a first correct divided sentence composed of a plurality of sentences semantically equivalent to the input sentence as training data; A sentence segmentation method implemented by a computer.

7. A program for causing a computer to function as the learning device according to claim 1 or the sentence segmentation device according to claim 4.