Dialogue device and training device thereof

The training device uses causal relationship expressions to enhance dialogue systems by generating responses that develop topics and provide transparent explanations, addressing limitations in conventional systems.

JP7852932B2Active Publication Date: 2026-04-28NAT INST OF INFORMATION & COMM TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NAT INST OF INFORMATION & COMM TECH
Filing Date
2022-05-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional dialogue systems face limitations in generating responses that develop a topic in response to user input, lack transparency in the response generation process, and require extensive dialogue data that is difficult to collect.

Method used

A training device that includes a hypothetical input storage, causal relationship storage, and a neural network trained using causal relationship expressions to generate output sentences, allowing for the development of dialogue topics and providing transparent responses.

Benefits of technology

The system generates responses that develop dialogue topics and provide transparent explanations, using causal relationships to enhance dialogue systems' accuracy and applicability beyond simple conversation partners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852932000001
    Figure 0007852932000001
  • Figure 0007852932000002
    Figure 0007852932000002
  • Figure 0007852932000003
    Figure 0007852932000003
Patent Text Reader

Abstract

A training data generation device 54 and a training device 56 include: an assumed input storage unit 78 that stores a plurality of pieces of assumed inputs which are assumed to be input to a conversation device; an expanded causal relationship database 74 that stores a plurality of causal relationship expressions; a training data creation unit 82 that, for each of the plurality of pieces of assumed inputs stored in the assumed input storage unit 78, extracts, from the plurality of causal relationship expressions, a causal relationship expression having a prescribed relationship with the assumed input, creates a training data sample in which the assumed input is the input and the extracted causal relationship expression is the response, and stores the sample in a training data storage unit 84; and a training unit 86 for using the training data sample stored in the training data storage unit 84 to train a response generation neural network 100 designed to generate an output sentence in response to a natural language input sentence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interaction device that interacts with a user using a computer, a training device therefor, and a computer program, and more particularly to an interaction device capable of developing a topic from user input, and a training device for training such an interaction device. This application claims priority based on Japanese Application No. 2021-090300 filed on May 28, 2021, and incorporates all the descriptions set forth in the said Japanese application.

Background Art

[0002] In recent years, interest in research on dialogue systems based on deep learning has increased, and research and development of various dialogue system technologies have been promoted. Conventional dialogue system technologies can be broadly classified into the following three categories.

[0003] 1) A search-based approach that receives user input and performs information retrieval based on the information obtained from the user input from a database that may not be for dialogue, and uses the results. Deep learning techniques may be used for the selection and processing of search results.

[0004] 2) Those that automatically generate a response sentence from user input. The end-to-end method using deep learning corresponds to this. Patent Document 1 cited below adopts this method. In this method, dialogue data is obtained from dialogue logs such as online chat services. Using this dialogue data, learning mainly using deep learning techniques is performed so that the system automatically generates a response to the input.

[0005] 3) Those based on a scenario-based approach. Many so-called AI (Artificial Intelligence) speakers adopt this method.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

[0007] In the first method described above, the information obtained as a response is limited to the database. Moreover, the relationship between user input and the response is not clear to the user, which hinders the development of dialogue with the user.

[0008] The second method described above has the problem that it is not possible to control the generated output as a response. Deep learning inherently has the problem that the process of generating a response from an input is not visible from the outside. Therefore, even if one tries to control the response, it is unclear how to do so. Furthermore, this method requires the collection of a large amount of dialogue data, and the scope of that data must be broad. It is generally known that collecting such data is extremely difficult.

[0009] Therefore, the second method described above had the problem that it was difficult to train the response device to generate responses that would develop the conversation in response to user input.

[0010] Furthermore, the third method has the constraint that responses remain within the scope of a prepared scenario. Trying to develop a dialogue within that constraint naturally has its limitations.

[0011] Therefore, the main objective of this invention is to provide a dialogue device capable of outputting responses that can develop a topic in response to user input, and a training device for training such a dialogue device. [Means for solving the problem]

[0012] A training device for a dialogue device according to the first aspect of the present invention includes a hypothetical input storage means that stores a plurality of hypothetical inputs, each of which is assumed to be an input to the dialogue device, and a causal relationship storage means that stores a plurality of causal relationship expressions, each of which includes a cause expression and an effect expression, and for each of the plurality of hypothetical inputs stored in the hypothetical input storage means, the hypothetical input and place The system includes a causal relationship expression extraction means for extracting a causal relationship expression having a fixed relationship from multiple causal relationship expressions; a training data creation means for taking the assumed input as input, creating a training data sample in which the causal relationship expression extracted by the causal relationship expression extraction means is used as the response, and storing it in a predetermined memory device; and a training means for training a dialogue device consisting of a neural network, designed to generate output sentences for natural language input sentences, using the training data sample stored in the training data creation means.

[0013] Preferably, the causal relationship expression extraction means includes a specific causal relationship expression extraction means for extracting causal relationships from multiple causal relationship expressions in which the noun phrase of the assumed input is used as the cause expression.

[0014] More preferably, the training device further includes a topic word model that, given a word, outputs the distribution probability of surrounding words for each word in a predetermined vocabulary, and a first training data sample adding means for identifying words with high distribution probabilities for each causal relationship representation of training data samples stored in a predetermined storage device, based on the output of the topic word model, and adding these to the input of the training data sample to generate a new training data sample and add it to the predetermined storage device.

[0015] More preferably, the training device further includes a second training data sample adding means for generating new training data samples and adding them to the predetermined storage device, based on the output of the topic word model, which extracts sentences from a predetermined corpus having distribution probabilities of surrounding words similar to the distribution probabilities of surrounding words of each causal relationship expression of the training data samples stored in a predetermined storage device, and adds these sentences to the input of the training data samples.

[0016] A computer program according to a second aspect of the present invention is a computer program that causes a computer to function as an assumption input storage means for storing a plurality of assumption inputs, each assumed to be an input to a dialogue device, and a causal relationship storage means for storing a plurality of causal relationship expressions, each of the plurality of causal relationship expressions including a cause expression and an effect expression, and the computer program further causes the computer to function as a causal relationship expression extraction means for extracting a causal relationship expression having a predetermined relationship with each of the plurality of assumption inputs stored in the assumption input storage means from the plurality of causal relationship expressions, and as a training data creation means for creating a training data sample that takes the assumption input as input and the causal relationship expression extracted by the causal relationship expression extraction means as the answer, and storing it in a predetermined storage device, and further causes the computer to function as a training means for training a dialogue device consisting of a neural network designed to generate output sentences for natural language input sentences using the training data samples stored in the training data creation means.

[0017] A dialogue device according to a third aspect of the present invention is a natural language dialogue device that includes a neural network designed to generate output sentences in response to natural language input sentences, wherein the neural network is trained so that the output sentences represent the latent consequences of the input sentences.

[0018] Preferably, the dialogue device further includes related expression adding means that, in response to being given an input sentence, adds related expressions, which are expressions containing words or sentences related to phrases in the input sentence, to the input sentence and inputs them to a neural network.

[0019] The dialogue device according to the fourth aspect of the present invention includes an utterance storage unit that stores past utterances of a user, a topic model that outputs a probability distribution of the occurrence of surrounding words for an input word, and a response generation unit that generates a response to a user utterance using the user utterances stored in the utterance storage unit and the topic model with the user utterance as an input.

[0020] The dialogue device according to the fifth aspect of the present invention includes an utterance storage unit that stores past utterances of a user, a topic model that outputs a probability distribution of the occurrence of surrounding words for an input word, a response generation unit that generates a response to the user's utterance with the user's utterance as an input, and a response generation adjustment unit that adjusts the generation of the response by the response generation unit based on the output of the topic model for the user's utterance.

[0021] The above and other objects, features, aspects, and advantages of the present invention will become apparent from the following detailed description of the present invention understood in connection with the accompanying drawings.

Brief Description of the Drawings

[0022] [Figure 1] FIG. 1 is a block diagram of a dialogue system for training a dialogue device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram of a training data creation unit that is a part of the dialogue device shown in FIG. 1. [Figure 3] FIG. 3 is a flowchart showing a control structure of a computer program that causes a computer to function as a training data creation device for the dialogue system shown in FIG. 1. [Figure 4] FIG. 4 is a block diagram of a dialogue system for training a dialogue device according to a second embodiment of the present invention. [Figure 5] FIG. 5 is a block diagram of a training data addition unit shown in FIG. 4. [Figure 6] FIG. 6 is a block diagram of a related expression search unit shown in FIG. 4. [Figure 7]Figure 7 is a flowchart showing the control structure of a computer program for causing a computer to function as a device for creating training data for a dialogue system according to the second embodiment. [Figure 8] Figure 8 is a flowchart showing the control structure of the routine that implements the related representation addition process shown in Figure 7. [Figure 9] Figure 9 is a block diagram of a dialogue device and a training device for training the dialogue device according to a third embodiment of the present invention. [Figure 10] Figure 10 is a schematic diagram showing an example of training data prepared by the training device according to the third embodiment. [Figure 11] Figure 11 is a flowchart showing the control structure of a computer program for causing a computer to function as a device for creating training data for a dialogue system according to the third embodiment. [Figure 12] Figure 12 is a flowchart showing the control structure of a computer routine that causes the computer to function in order to perform the related word addition process shown in Figure 11. [Figure 13] Figure 13 is a block diagram of a dialogue device according to the fourth embodiment of this invention. [Figure 14] Figure 14 is a block diagram of a dialogue device according to a fifth embodiment of this invention. [Figure 15] Figure 15 is an external view of a computer system that implements each of the above embodiments. [Figure 16] Figure 16 is a block diagram showing the hardware configuration of the computer system shown in Figure 15. [Modes for carrying out the invention]

[0023] In the following descriptions and drawings, identical parts are assigned the same reference number. Therefore, detailed descriptions of them will not be repeated.

[0024] First Embodiment 1. Configuration (1) Dialogue system 50 Figure 1 shows a block diagram of the configuration of a dialogue system 50 according to the first embodiment of the present invention. Referring to Figure 1, the dialogue system 50 includes a dialogue device 52 and a training data generation device 54 connected to the Internet 60 for generating training data for training the dialogue device 52 using causal relationship representations extracted from the Internet 60. The dialogue system 50 further includes a training device 56 for training the dialogue device 52 using the training data generated by the training data generation device 54.

[0025] The configuration of each element of the dialogue system 50 will be described below. It should be assumed that each word constituting the natural language sentence used in the following embodiment has been converted into a word vector beforehand. That is, each natural language sentence is represented as a sequence of word vectors consisting of word vectors.

[0026] (2) Dialogue device 52 The dialogue device 52 includes a response generation neural network 100, which is a neural network that receives user input 102 consisting of natural language sentences and generates a response sentence. The dialogue device 52 further includes a speech formatting unit 104 that formats the response sentence generated by the response generation neural network 100 so that it is in a form appropriate as a response to the user input 102, and outputs it as a response utterance 106.

[0027] The response generation neural network 100 is a so-called end-to-end type, and is pre-trained to generate natural language responses to natural language inputs. Training the response generation neural network 100 by the training device 56 corresponds to so-called fine tuning.

[0028] The response-generating neural network 100 can utilize a generative network consisting of a transformer encoder and a transformer decoder, or UniLM, which is a BERT model further pre-trained for generative purposes. However, it is not limited to these; the response-generating neural network 100 can be implemented using any generative network. Furthermore, if there is a large amount of data for generative training, implementation is possible even with just standard generative training. (For the BERT model, see the reference Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding.) (3) Training data generation device 54 The training data generation device 54 includes a causal relationship extraction unit 62 for extracting causal relationship expressions from the Internet 60 by known methods, and a causal relationship database 64 for storing the causal relationship expressions extracted by the causal relationship extraction unit 62. The training data generation device 54 further includes a chain causal relationship generation unit 66 that generates new candidate causal relationships in a chain by linking the cause part of a first causal relationship and the cause part of a second causal relationship when the consequence part of a causal relationship (first causal relationship) and the cause part of another causal relationship (second causal relationship) are semantically consistent among the causal relationship expressions stored in the causal relationship database 64, and further by linking the newly generated causal relationships with other causal relationships, and a generated causal relationship database 68 for storing the candidate causal relationships generated by the chain causal relationship generation unit 66. For the causal relationship extraction unit 62, for example, the causal relationship recognition method disclosed in Japanese Patent Publication No. 2018-60364 can be used to extract causal relationships. For example, the technology disclosed in Japanese Patent Publication No. 2015-121897 can be used as a method for generating new candidate causal relationships by linking causal relationships.

[0029] The training data generation device 54 further includes a chain causal relationship selection unit 70 that selects appropriate causal relationships from among the candidate causal relationships stored in the generated causal relationship DB 68, a chain causal relationship DB 72 for storing the causal relationships selected by the chain causal relationship selection unit 70, and an extended causal relationship DB 74 that integrates and stores the causal relationships stored in the causal relationship DB 64 and the causal relationships stored in the chain causal relationship DB 72.

[0030] The training data generation device 54 also includes an assumed input extraction unit 76 that extracts expressions from the Internet 60 that are expected to become user inputs 102 to the dialogue device 52, and an assumed input storage unit 78 for storing assumed inputs extracted by the assumed input extraction unit 76 and assumed inputs added manually from the console 80. The training data generation device 54 further includes a training data creation unit 82 that takes one of the assumed inputs stored in the assumed input storage unit 78 as input, creates a training data sample in which a causal relationship stored in the extended causal relationship DB 74 has a consequence part that can be an answer to the assumed input as the correct answer, and outputs it to the training device 56.

[0031] As will be described later, the training data creation unit 82 reads from the extended causal relationship DB 74 for each assumed input read from the assumed input storage unit 78, and creates training data samples by reading causal relationships in which the noun phrase contained in that assumed input is the cause.

[0032] Referring to Figure 2, the training data creation unit 82 includes an assumed input reading unit 150 that sequentially reads assumed inputs from the assumed input storage unit 78, and a noun phrase identification unit 152 that identifies noun phrases contained in the assumed inputs read by the assumed input reading unit 150. The training data creation unit 82 further uses the noun phrases identified by the noun phrase identification unit 152. cause The system includes a causal relationship search unit 154 for searching and retrieving all causal relationships included in the section from the extended causal relationship DB74, and a training data sample creation unit 156 for creating training data samples for each causal relationship retrieved by the causal relationship search unit 154, using assumed inputs as inputs and causal relationships as outputs, and storing them in the training data storage unit 84.

[0033] (4) Training device 56 The training device 56 includes a training data storage unit 84 for storing training data samples output by the training data creation unit 82, and a training unit 86 for training the response generation neural network 100 using the training data samples stored in the training data storage unit 84. The processing performed by the training unit 86 on the response generation neural network 100 is, as described above, fine-tuning of the response generation neural network 100. Specifically, for each training data sample, the training unit 86 provides the response generation neural network 100 with an assumed input and trains the response generation neural network 100 using the backpropagation method to generate an output that is the result of a causal relationship caused by that assumed input.

[0034] (5) Training data generation program Figure 3 shows a flowchart illustrating the control structure of a computer program for making the computer function as a training data generation device 54. Referring to Figure 3, the program includes, after the program is started, a step 200 that allocates and initializes memory, opens relevant files and connects to a database, and a step 202 that reads all assumed inputs from the assumed input storage unit 78 shown in Figure 1. The program further includes a step 204 that executes step 206 for each assumed input read in step 202, and a step 208 that, after step 204 is completed, performs a predetermined termination process to terminate the execution of the program.

[0035] Step 206 includes step 230, which identifies all noun phrases present in the expected input to be processed, and step 232, which performs step 234 for each of the noun phrases identified in step 230.

[0036] Step 234 includes step 260, which reads all causal relationships having the noun phrase being processed as the cause from the extended causal relationship DB74 shown in Figure 1, and step 262, which executes step 264, which creates training data samples consisting of combinations of each causal relationship read in step 260 with the assumed input being processed.

[0037] Specifically, step 264 involves creating training data samples where the expected input during processing is used as input and the causal relationship during processing is used as the response, and storing them in the training data storage unit 84 shown in Figure 1.

[0038] 2 operations The dialogue system 50, whose structure is described above, operates as follows:

[0039] (1) Creation of training data First, the training data generation device 54 operates as follows to create training data for the response generation neural network 100. First, the causal relationship extraction unit 62 extracts a large number of causal relationships from the internet 60 and stores them in the causal relationship DB 64. The causal relationship DB 64 stores these causal relationships in a format that allows for various searches. For example, in the technology disclosed in the above-mentioned Japanese Patent Application Publication No. 2018-60364, the cause and consequence parts of a causal relationship can be distinguished. If the causal relationship DB 64 is designed to store the cause and consequence parts of each causal relationship in separate columns, then, for example, only causal relationships that contain a specific word in the cause part can be easily extracted.

[0040] The causal relationship generation unit 66 generates all of the candidate causal relationships stored in the causal relationship DB 64 such that the consequence part of the first causal relationship has substantially the same meaning as the cause part of the second causal relationship. Further, by performing a similar process, new candidate causal relationships are generated by linking multiple causal relationships. In this embodiment, a large number of candidate causal relationships are generated by linking causal relationships up to a predetermined upper limit. All of these candidate causal relationships are stored in the generated causal relationship DB 68. The generated causal relationship DB 68 stores candidate causal relationships in the same format as the causal relationship DB 64, for example. In the case of the generated causal relationship DB 68, information that identifies the causal relationship from which the candidate causal relationship was generated may also be stored.

[0041] The causal relationship candidates generated by the chain causal relationship generation unit 66 are based solely on the fact that the consequence part of one causal relationship and the cause part of the other are semantically the same in two consecutive causal relationships within the chain. However, as disclosed in the above-mentioned Japanese Patent Publication No. 2015-121897, among the causal relationship candidates obtained through such a chain of causal relationships, there are some that do not represent a correct causal relationship as a whole. Therefore, the chain causal relationship selection unit 70 selects those that are considered correct from the causal relationship candidates stored in the generated causal relationship DB 68 and stores them in the chain causal relationship DB 72. The method for selecting causal relationship candidates here is the one disclosed in Japanese Patent Publication No. 2015-121897. It is also possible to fine-tune a pre-trained natural language model such as BERT and have it perform the selection of causal relationship candidates.

[0042] The Extended Causal Relationship DB74 stores causal relationships that integrate the Causal Relationship DB64 and the Chain Causal Relationship DB72. Specifically, the Extended Causal Relationship DB74 stores causal relationships extracted from the Internet 60 by the Causal Relationship Extraction Unit 62, and causal relationships generated from these by the Chain Causal Relationship Generation Unit 66 and the Chain Causal Relationship Selection Unit 70. The Extended Causal Relationship DB74 also stores causal relationships in the same format as the Causal Relationship DB64.

[0043] Meanwhile, the assumed input extraction unit 76, like the causal relationship extraction unit 62, extracts expressions from numerous web pages on the Internet 60 that are likely to be used as input to the response generation neural network 100, as assumed inputs for the response generation neural network 100. For example, candidates could include question sentences from numerous FAQ (Frequently Asked Questions) sites on the Internet 60, as well as question-format sentences posted on various information-providing sites. In addition to this, it is also possible to extract ordinary sentences and generate question sentences in which noun phrases within them become the answers. The assumed inputs extracted by the assumed input extraction unit 76 are stored in the assumed input storage unit 78.

[0044] Alternatively, the user may use console 80 to supplement the expected inputs. However, this supplementation is not mandatory.

[0045] Once the causal relationships are prepared in the extended causal relationship DB74 and the assumed inputs are prepared in the assumed input storage unit78 as described above, the training data creation unit82 creates the training data as follows. Specifically, the computer creates the training data by executing a program whose control structure is shown in Figure 3.

[0046] Referring to Figure 3, when this program is started, initial processing is first performed in step 200, and in step 202, all assumed inputs stored in the assumed input storage unit 78 are read and loaded into memory. In step 204, training data is created by executing step 206 for each of these assumed inputs as described below, until all assumed inputs are finished.

[0047] In step 206, all noun phrases in the expected input being processed are first identified (step 230), and in step 232, the process in step 234 is executed for each of those noun phrases.

[0048] In step 234, for the noun phrase being processed, all causal relationships that have that noun phrase as the cause are read from the extended causal relationship DB74 (step 260). Furthermore, in step 262, for each of these causal relationships, a training data sample is created with the assumed input being processed as the input and the consequence part of the causal relationship as the answer, and this is saved in the training data storage unit 84 in Figure 1. This process of step 264 is repeated until all causal relationships are completed.

[0049] Once the processing in step 232 is completed for all the noun phrases identified in step 230, step 206 for a given expected input is completed, and the processing in step 206 is executed again for the next expected input.

[0050] Once step 204 is completed for all the assumed inputs read in step 202, the creation of the training data is finished. The training data is stored in the training data storage unit 84 in Figure 1.

[0051] (2) Training Once the training data is prepared as described above, the training unit 86 of the training device 56 uses this training data prepared in the training data storage unit 84 to train the response generation neural network 100.

[0052] The training of the response-generating neural network 100 is carried out using the standard backpropagation method. That is, the response-generating neural network 100 is given a hypothetical input, and its output is obtained. At this time, the output of the response-generating neural network 100 is output sequentially in the form of word vectors. The parameters of the response-generating neural network 100 are modified so that the sequence of word vectors formed by these word vectors becomes the consequence (sequence of word vectors) of the causal relationship, which is the answer to the training data. Since this training is based on existing methods, the details will not be repeated here.

[0053] (3) Dialogue processing Once the response generation neural network 100 has been trained, the dialogue device 52 becomes available. When a user provides some input in natural language, this input is converted into a sequence of word vectors and provided to the response generation neural network 100 as user input 102. The response generation neural network 100 generates a response to this user input 102 and outputs it to the speech formatting unit 104. The speech formatting unit 104 formats the response provided by the response generation neural network 100 to make it appropriate as a response to user input 102 (for example, by adding some words to the beginning to match user input 102, repeating part of user input 102, or changing sentence endings to make them more conversational), and outputs it as a response utterance 106. This formatting may be done using a rule-based method, or a neural network trained using modified sentences and their endings as training data may be used.

[0054] Alternatively, the user may speak, and this speech may be recognized and input to the response generation neural network 100 as user input 102. In this case, the response utterance 106 may also be output as speech through speech synthesis.

[0055] 3. Effects Since the response generation neural network 100 is trained based on causal relationships, the response generation neural network 100 generates sentences related to some causal consequence of user input 102 in response to user input 102. The training data also includes causal relationships generated by linking causal relationships, and it is highly likely that the output of the response generation neural network 100 will not only be answers that can be directly derived from user input 102, but also sentences related to potential risks or opportunities derived from user input 102. As a result, the dialogue is more likely to develop compared to conventional dialogue systems where answers are searched within a certain framework.

[0056] In this embodiment, training data is created using causal relationships. As is done in conventional techniques, it is difficult to obtain large amounts of general dialogue data, but causal relationships can be obtained in large quantities from the internet using existing methods. Therefore, a large amount of training data can be prepared, and the response generation neural network 100 can be made more accurate. In other words, the response generation neural network 100 can generate responses based on causal relationships that have a high probability of further developing the dialogue in response to user input 102.

[0057] In conventional dialogue systems using neural networks, the neural network itself is a black box. Therefore, it is difficult to explain to the user the intention behind the responses output by the dialogue system. In contrast, with a neural network trained using causal relationships, such as the response-generating neural network 100 according to this embodiment, it can be explained that the network is informing the user of potential opportunities and risks derived from their utterances. Therefore, instead of viewing the dialogue system simply as a conversation partner, its applications can be expanded to include tools that develop the user's thinking or guide their actions.

[0058] Second second embodiment 1. Configuration (1) Dialogue system 300 Figure 4 shows the configuration of a dialogue system 300 according to a second embodiment of the present invention in block diagram form. Referring to Figure 4, this dialogue system 300 includes a training data generation device 54 and a training device 56 similar to those in the first embodiment. The dialogue system 300 further includes a training data expansion unit 58 connected to the training data storage unit 84 of the training device 56 and the assumed input storage unit 78 of the training data generation device 54, which creates a new training data sample for each training data sample by adding some word or phrase related to the assumed input as a topic to the assumed input and adds it to the training data storage unit 84, and a dialogue device 302 that uses a response generation neural network 340 trained by the training device 56 using this expanded training data.

[0059] Of these, the training data expansion unit 58 and the dialogue device 302 will be described in order below.

[0060] (2) Training data expansion unit 58 A. Composition The training data expansion unit 58 includes a topic word model 330, which is pre-prepared using corpus statistics, to output a surrounding word distribution vector whose elements are the probability of each word occurring around a given word. The training data expansion unit 58 further includes an association expression search unit 332 that, when given a word, identifies words with surrounding word distribution vectors similar to the surrounding word distribution vector for that word based on the output of the topic word model 330, and extracts assumed inputs from the assumed input storage unit 78 that include words with word vectors similar to the surrounding word distribution vector of that word. The training data expansion unit 58 further includes a training data addition unit 332 that, for each training data sample stored in the training data storage unit 84, extracts words included in the consequence part of its causal relationship expression and provides them to the association expression search unit 332, and in response adds the combination of words and assumed inputs output by the association expression search unit 332 to the assumed input of the training data sample to expand a new training data sample and add it to the training data storage unit 84. 4 include.

[0061] I. Training data addition unit 334 Referring to Figure 5, the configuration of the training data addition unit 334 is as follows. The training data addition unit 334 includes a training data reading unit 360 that reads out training data samples stored in the training data storage unit 84 one by one and extracts words that are included in the result section of the causal relationship therein, and a related expression query unit 362 that queries the related expression search unit 332 for related expressions for each of the words extracted by the training data reading unit 360.

[0062] The training data addition unit 334 further includes an association expression addition unit 366 for receiving association expressions output by the association expression search unit 332 in response to an inquiry from the association expression query unit 362, and adding predetermined combinations of these association expressions to the assumed input in the training data sample read by the training data reading unit 360 to generate a new training data sample. The training data addition unit 334 further includes a training data writing unit 36 ​​for adding and writing the new training data sample generated by the association expression addition unit 366 to the training data storage unit 84. 4 include.

[0063] Topic word model 330 As described above, the topic word model 330 is obtained by performing statistical processing on a predetermined corpus in advance, so that when a word is given, it outputs a marginal word distribution vector whose elements are the probability of each word appearing around that word (for example, before and after the word, or within two ranges both before and after the word). The marginal word distribution vector has a number of elements corresponding to a certain number of words selected from the target language. From this marginal word distribution vector, words that have a high probability of appearing around a given word can be identified. Furthermore, the marginal word distribution vectors of words that appear in similar situations are similar to each other. Therefore, words that appear in similar situations can also be estimated using this topic word model 330. Note that whether two marginal word distribution vectors are similar or not can be determined by their cosine similarity.

[0064] In addition to topic-word models that compute word embeddings like Word2Vec, it is also possible to use a topic-word model that fine-tunes a pre-trained BERT model to estimate the probability distribution of words appearing in the input sentence or in neighboring sentences.

[0065] E Related expression search unit 332 Referring to Figure 6, the related expression search unit 332, when given a word from the related expression query unit 362 shown in Figure 5, receives a surrounding word distribution vector for that word from the topic word model 330, and further includes a related assumption input search unit 400 for searching among the assumption inputs stored in the assumption input storage unit 78 for assumption inputs containing words with surrounding word distribution vectors similar to this surrounding word distribution vector, and retrieving them as related assumption inputs. The related expression search unit 332 further includes a related word search unit 402 for obtaining a surrounding word distribution vector for that word from the topic word model 330 when given a word from the related expression query unit 362, and selecting words that are highly likely to be used around the input word as related words from this surrounding word distribution vector. The related expression search unit 332 further includes a related expression expansion unit 404 for generating various combinations of related assumed input retrieved by the related assumed input search unit 400 and related words selected by the related word search unit 402, and a related expression output unit 406 for outputting each of the combinations generated by the related expression expansion unit 404 to the related expression addition unit 366 in Figure 5.

[0066] (3) Dialogue device 302 A. Composition Referring to Figure 4, the dialogue device 302 includes a response-generating neural network 340 having a similar configuration to the response-generating neural network 100 in Figure 1, but unlike the response-generating neural network 100, it includes a response-generating neural network 340 trained with training data expanded by the training data expansion unit 58. The dialogue device 302 further includes an information-adding device 33 that adds one or more words representing expressions previously uttered by the user in a dialogue or the topic of the dialogue with the user to the user input 102 to the dialogue device 302 and inputs them to the response-generating neural network 340. 8 include.

[0067] The information-adding device 338 includes a topic word model 344 similar to the topic word model 330 of the training data expansion unit 58, a user input storage unit 346 for storing past user utterances, and a word selection unit 348 that receives user input 102, extracts words contained in user input 102, and selects several words with a high probability of being related to those words by referring to the topic word model 344. The information-adding device 338 further includes an utterance selection unit 350 for selecting several utterances related to user input 102 from among the user inputs stored in the user input storage unit 346, and a word selection unit 348 and an utterance selection unit 350 for user input 102. ta utterance It includes an information addition unit 352 for selecting any combination of (including the case where nothing is selected) and adding it to the user input 102, and providing it as input to the response generation neural network 340.

[0068] Topic word model 344 Topic word model 344 is similar to topic word model 330 in Figure 4.

[0069] U. Word selection section 348 The word selection unit 348 extracts words from the user input 102 and obtains a surrounding word distribution vector for each of those words by referring to the topic word model 344. It then has the function of selecting a predetermined number of words that correspond to the element with the highest probability among those surrounding word distribution vectors.

[0070] E. User input storage unit 346 The user input storage unit 346 stores a predetermined number of the most recent user inputs entered by the user in the past. The time of input may also be added to the user input before storage.

[0071] O Speech selection unit 350 The speech selection unit 350 has a neural network that is pre-trained to receive a word vector sequence formed by combining user input 102 and past user inputs stored in the user input storage unit 346 via predetermined separation tokens, and to output a value indicating the degree to which the two are related. The unit has the function of selecting several past user utterances for which the value output by this neural network is the highest. If the input time of the user input is stored, the selection target may be limited to user inputs from a predetermined time before the present to the present, or the probability of selecting a user input may be inversely correlated with the elapsed time since the input was made.

[0072] Information addition section 352 The information addition unit 352 receives all the words selected by the word selection unit 348 and all the user utterances selected by the utterance selection unit 350, generates a set of all possible combinations of these, and randomly selects one combination from this set to add to the user input 102. The combinations may include combinations that do not contain any words or past inputs. Therefore, the output of the information addition unit 352 can take various forms, such as user input 102 only, user input 102 + a word, user input 102 + past input, user input 102 + two words, user input 102 + a word + past input, etc.

[0073] (4) Training data generation program A. Overall structure Figure 7 shows the control structure of the program for causing the computer to function as the training data generation device 54, training device 56, and training data expansion unit 58 in Figure 4.

[0074] Referring to Figure 7, this program differs from the first embodiment shown in Figure 3 in that it includes a step 450 between step 204 and step 208 in which a process is performed to add the associated representation of the user input 102.

[0075] I. Adding related expressions Figure 8 shows the control structure of the program routine that causes the computer to function to execute step 450 in Figure 7.

[0076] Referring to Figure 8, this routine includes a step 500 which performs predetermined initial processing when the routine starts executing, a step 502 which performs the processing of step 504 (described later) for each training data sample until the end of the training data sample, and a step 506 which, after the completion of step 502, performs predetermined termination processing to complete the execution of this routine and return control to the parent routine shown in Figure 7.

[0077] Step 504 includes step 530, which extracts words from the consequences of causal relationships contained in the training data sample to be processed, and step 532, which generates a set of word candidates to be attached to the assumed input by performing step 534 for each word until the words extracted in step 530 are finished. Step 504 further includes step 536, which, after step 532 is completed, generates all possible combinations of words contained in the set of words generated in step 532, and for each combination, generates a new training data sample by attaching that combination to the assumed input contained in the training data sample to be processed, and adds it to the training data.

[0078] Step 534 includes step 560, which calculates a peripheral word distribution vector for the word being processed using the topic word model 330, and step 562, which selects a predetermined number of words with the highest probabilities calculated in step 560. Step 534 further includes step 56, which selects a predetermined number of assumed inputs from the assumed input storage unit 78 in Figure 4 that have peripheral word distribution vectors similar to the one calculated in step 560. 4 include.

[0079] In this embodiment, when selecting words in step 562, only words with a probability above a threshold are selected, even if their probability is high. Similarly in step 564, only words whose similarity to the surrounding word distribution vector is above a predetermined level are selected. The threshold can be determined experimentally to be appropriate. Such constraints are not required.

[0080] 2 operations (1) Creation of training data A. Creation of basic training data The creation of basic training data is performed from steps 200 to 204 in Figure 7. In this case, the operation of this embodiment is the same as that of the first embodiment.

[0081] (i) Expansion of training data The expansion of the training data is performed in step 450 of Figure 7. to Referring to Figures 4 and 5, the training data reading unit 360 of the training data addition unit 334 reads training data one by one from the training data storage unit 84 (step 502 in Figure 8), extracts words from the causal relationship consequence section within the training data (step 530), and provides them to the related expression query unit 362. The related expression query unit 362 queries the related expression search unit 332 for related expressions for each of those words (step 532) (step 534).

[0082] Referring to Figure 6, when the related expression search unit 332 receives a word from the related expression query unit 362, the related assumed input search unit 400 calculates a surrounding word distribution vector for that word from the topic word model 330 (step 560). The related word search unit 402 selects a predetermined number of words that have a high probability of being used around the input word as related words from this surrounding word distribution vector (step 562). The related assumed input search unit 400 further searches the assumed inputs stored in the assumed input storage unit 78 for assumed inputs that contain words with a surrounding word distribution vector similar to this surrounding word distribution vector, and retrieves them as related assumed inputs (step 564). This process is performed for each word.

[0083] The related expression development unit 404 generates possible combinations of related assumption input retrieved by the related assumption input search unit 400 and related words selected by the related word search unit 402 (step 536). The related expression output unit 406 outputs each of the combinations generated by the related expression development unit 404 to the related expression addition unit 366 in Figure 5 as related expressions related to causal relationships.

[0084] Referring again to Figure 5, the association expression addition unit 366 receives the combination of association expressions output by the association expression search unit 332, adds these combinations to the assumed input in the training data sample being processed, and generates a new training data sample. The training data writing unit 364 adds and writes the new training data sample generated in this way by the association expression addition unit 366 to the training data storage unit 84 (step 538).

[0085] The above process will augment the training data.

[0086] (2) Training The training of the response-generating neural network 340 in this second embodiment is the same as that performed on the response-generating neural network 100 in the first embodiment. However, since the training data in the second embodiment is different from that in the first embodiment, the internal parameters of the response-generating neural network 340 at the end of training will be considerably different from the internal parameters of the response-generating neural network 100 in the first embodiment.

[0087] (3) Dialogue processing During dialogue processing, the dialogue device 302 operates as follows: First, when user input 102 is given to the dialogue device 302, the word selection unit 348 extracts words contained in the user input 102 and selects several words that have a high probability of being related to those words by referring to the topic word model 344. The utterance selection unit 350 selects several utterances related to user input 102 from past user inputs stored in the user input storage unit 346. The information addition unit 352 adds one of any combinations (including no selection) of words selected by the word selection unit 348 and user inputs selected by the utterance selection unit 350 to user input 102 by some method, for example, randomly, and provides it as input to the response generation neural network 340.

[0088] The internal operation of the response-generating neural network 340 after receiving this input is the same as that of the response-generating neural network 100 in the first embodiment. However, since the internal parameters of the two are different, even if the same user input 102 is given, the output of the response-generating neural network 340 is likely to differ from the output of the response-generating neural network 100. This possibility is further increased by adding words or other elements to the user input 102. In particular, in the case of the response-generating neural network 340, in generating a response, it is highly likely that the response will be generated not solely based on the user input 102, but based on a causal relationship in which the words contained in the user input 102 are included in the outcome, or a causal relationship in which the expression related to user input prior to user input 102 is included in the outcome. Therefore, while the dialogue with the user is continuous, it is possible to provide topics that the user had not anticipated by reflecting potential opportunities or risks based on a chain of causal relationships.

[0089] For example, let's assume user input 102 is "Artificial intelligence has advanced, hasn't it?". Let's further consider a scenario where a previous user input was "I'm worried about the elderly," and this previous user input is added to user input 102, resulting in "Artificial intelligence has advanced, hasn't it? + I'm worried about the elderly," which is then provided to the response generation neural network 340. In this case, for example, the output of the response generation neural network 340 could be "Let's use robots to provide care services and support the elderly," developing the topic from user input 102 and guiding the conversation in a direction the user hadn't considered at the time user input 102 was uttered.

[0090] 3. Effects As described above, this embodiment can generate responses to user input based on causal relationships. To generate such responses, it is necessary to train a neural network using a large number of causal relationship representations. However, unlike ordinary dialogue data, causal relationship representations can be easily collected in large quantities from the internet. Moreover, this amount is increasing day by day. Therefore, the accuracy of response generation by the neural network used in the dialogue can be easily improved. Furthermore, since the response is generated based on causal relationships, unlike conventional dialogue systems using neural networks, it can generate responses based on potential opportunities and risks that the user may not be aware of, reflecting causal relationships and chains of causal relationships. As a result, dialogue with the user can be developed in a more beneficial way compared to conventional systems.

[0091] Third Embodiment 1. Configuration Figure 9 is a block diagram showing the configuration of a dialogue system 600 according to a third embodiment of the present invention. Referring to Figure 9, the dialogue system 600 includes a training data generation device 610 for creating training data for the dialogue system, and a dialogue device 612 including a neural network. The dialogue system 600 further includes a training data expansion unit 58 similar to that shown in Figure 4, which creates and adds new training data samples to each training data sample created by the training data generation device 610, by adding some words related to the assumed input as a topic to the assumed input, and a training device 56 for training the dialogue device 612 using the training data generated by the training data generation device 610 and expanded by the training data expansion unit 58.

[0092] The difference between the training data generation device 610 and the training data generation device 54 shown in Figure 1 is that, after the chain causal relationship selection unit 70 in Figure 1, it includes a causal consequence chain generation unit 620 for generating a string (here called a causal consequence chain) by sequentially chaining together only the cause part of the leading causal relationship and the consequence parts of the causal relationships included in that chain causal relationship, selected by the chain causal relationship selection unit 70; it includes a causal consequence chain storage unit 622 for storing the causal consequence chains generated by the causal consequence chain generation unit 620, instead of the chain causal relationship DB 72 in Figure 1; and it includes a causal consequence chain DB 624 for storing individual causal relationships stored in the causal relationship DB 64 and causal consequence chains stored in the causal consequence chain storage unit 622, instead of the extended causal relationship DB 74 shown in Figure 1. A single causal relationship stored in the causal relationship DB64 can be considered a causal consequence chain consisting of only one causal consequence. The training data generation device 610 further differs from the training data generation device 54 shown in Figure 1 in that, instead of the training data creation unit 82 shown in Figure 1, it includes a training data creation unit 626 that creates training data for training the dialogue device 612 using the causal consequence chains stored in the causal consequence chain DB624 and the assumed inputs stored in the assumed input storage unit 78.

[0093] In this third embodiment, the user inputs a natural number N that specifies the number of causal consequence chains, which is used as the user input 630 to the dialogue device 612, which differs from the first and second embodiments.

[0094] The dialogue device 612 includes a response-generating neural network 632 trained by the training device 56, and an information-adding device 642, which includes a topic word model 344, a word selection unit 348, and an information-adding unit 352 within the information-adding device 338 shown in Figure 4. This device receives user input 630, extracts words representing the topic of conversation with the user, and adds any combination of these words to the user input 630 before inputting it to the response-generating neural network 632. Any combination of these added words is inserted between the user input 630 and a natural number N. That is, the input to the response-generating neural network 632 is in the form of "user input 630" + "combination of words added by the information-adding device 338" + "natural number N specifying the chain length". As will be described later, when the response-generating neural network 632 is given an input in the above format, it is trained to output a string consisting of N causal consequence chains, the last causal consequence chain containing one of the words added by the information-adding device 338.

[0095] The dialogue device 612 further includes a question generation unit 634 that generates a question sentence from the output result of a response generation neural network 632 to an input given from an information addition device 338, a response acquisition unit 636 that provides the question generated by the question generation unit 634 to an external question answering system 614 to obtain its response, and a speech formatting unit 638 that formats the response obtained by the response acquisition unit 636 into a dialogue sentence and outputs it as a response utterance 640.

[0096] In this embodiment (and in other embodiments as well), the processing in the question generation unit 634 involves adding an interrogative word to the beginning of the causal relationship expression output by the response generation neural network 632, and then formatting the end into a question in some way to create a question. These question sentences can be generated using a separately trained neural network or rule-based formatting means.

[0097] For example, let's assume that the causal relationship expression output by the response generation neural network 632 is "Go to Nara" → "Meet deer". From this obtained causal relationship expression, (1) by adding "How", we create the question "How is it good to go to Nara to meet deer?". (2) In addition to this result, by further adding the phrase "and good?" to the end, we create the question "How is it good to go to Nara to meet deer?".

[0098] The response acquisition unit 636 inputs one or both of these question sentences into the question answering system 614 and obtains the answer as output and the text from which the answer has been extracted. The speech formatting unit 638 ,child The system uses the response alone or in combination with the extracted text to format it appropriately for a user interaction, and outputs the final response utterance 640.

[0099] Here, the question-answering system 614 is envisioned as a system that performs a search process on an internal database or external information source (e.g., documents on the internet) based on the input, extracts descriptions that are related to the input question and are factual or well-founded, and outputs them as a response. An example of such a question-answering system is the question-answering system WISDOM X (https: / / www.wisdom-nict.jp / ) operated by the National Institute of Information and Communications Technology (NICT). By sequentially using such question-answering systems 614, it is possible to utilize supporting descriptions and texts in the response generation process, thereby giving basis to the response utterance 640 to the user. Note that it is also possible to use the output of a regular search engine, not just WISDOM X.

[0100] As described above, the input to the response generation neural network 632 is in the form of "user input 630" + "combination of utterances and words added by the information addition device 338" + "natural number N specifying the chain length". Therefore, the training data for the response generation neural network 632 is as shown in Figure 10.

[0101] Referring to Figure 10, each training sample that makes up this training data has "expected input + related words (group) + natural number N" as its input, and its output contains a causal consequence chain consisting of a number of causal consequences specified by the natural number N. In this case, the causal consequence at the end of the output of each training sample contains one of the words included in the related words (group).

[0102] Figure 11 is a flowchart showing the control structure of a computer program that causes the computer to function as the training data creation unit 626 shown in Figure 9. Referring to Figure 11, this program is similar to that of the second embodiment shown in Figure 7. However, the program shown in Figure 11 differs from that shown in Figure 7 in that it includes step 670 instead of step 206 in Figure 7, and step 672, which adds related words to the input instead of related expressions, instead of step 450 in Figure 7.

[0103] Step 670 is similar to step 206 in Figure 7. However, step 670 includes step 680 instead of step 234 in Figure 7. 06 and different.

[0104] Step 680 includes step 690, which reads out all causal consequence chains that have the noun phrase being processed as the cause (the beginning of the causal consequence chain), and step 692, which performs step 694 for each causal consequence chain read out in step 690.

[0105] Step 694 is for creating and saving training data samples for combinations of assumed inputs and causal outcome chains during processing, where the assumed input and the number of causal outcome chains are inputs, and the causal outcome chain is the response.

[0106] In other words, the process in step 680 yields a training data sample from which the "related words(groups)" have been removed from the training data sample shown in Figure 10. The "related words(groups)" are added in step 672.

[0107] Figure 12 shows the control structure of the program routine that implements step 672 in Figure 11. Referring to Figure 12, the program includes step 700 which performs predetermined initial processing, step 702 which performs step 704 for each training data sample, and step 706 which, after step 702 is completed, performs predetermined termination processing to terminate the execution of this routine.

[0108] Step 704 includes step 710, which extracts words from the last consequence of the causal chain of the training data sample being processed; step 711, which generates all combinations of the words extracted in step 710; and step 712, which performs step 714 for each combination generated in step 711.

[0109] Step 714 includes step 720, which inserts the combination of words being processed between the assumed input of the training data sample being processed and a natural number N indicating the number of chains, thereby creating a new sample, and step 722, which writes the new sample created in step 720 to the training data storage unit 84 shown in Figure 9, thereby ending the execution of step 714.

[0110] 2. Operation and Effects By having a computer run the program that shows the control structure in Figures 11 and 12, training data with the configuration shown in Figure 10 can be obtained. It is important to note that the words included in the "related words (group)" in the input section of the training data are included in the causal consequences at the end of the causal consequence chain in the output, and the number of causal consequences in the output section is the number specified by the natural number in the input section.

[0111] The response-generating neural network 632 shown in Figure 9 is trained using this training data. As a result, when the response-generating neural network 632 is given an input with the configuration shown as "Input" in Figure 10, the parameters of the response-generating neural network 632 are set so that it outputs a causal consequence chain that has a causal part related to the assumed input specified by the input, has the number of causal consequences specified by the input, and the last causal consequence has the related word specified by the input or a word close to the related word.

[0112] Therefore, when the user inputs user input 630 and a natural number N to the dialogue device 612 during actual inference, a word indicating the topic of the dialogue is selected as a related word for the user input 630 and added to the input. When this input is given to the response generation neural network 632, the response generation neural network 632 outputs a causal consequence chain with a specified number of causal consequences, the last causal consequence of which contains the word designated as a related word or a word close to it. Its cause is related to the input. This causal consequence chain is the "causal relationship expression" referred to in the first and second embodiments. From this causal relationship expression, the question generation unit 634 generates a question sentence. The response acquisition unit 636 gives the question sentence to the question answering system 614 and obtains its response. The speech formatting unit 638 formats the response into a form appropriate as a response to user input 630 and outputs it as a response utterance 640.

[0113] In this embodiment, user input + related words + natural number N are input to the response generation neural network 632, and multiple causal consequence chains consisting of N causal consequences are obtained. For example, consider the case where the user input is "Artificial intelligence will develop," and the related word "elderly people" is discovered. Assuming the values ​​of the natural number N are 1, 2, and 3, the following output is expected to be obtained.

[0114] N=1: Support the elderly.

[0115] N=2: Utilize robots to support the elderly.

[0116] N=3: Utilizing robots → enabling the provision of care services → supporting the elderly.

[0117] Therefore, the value of the natural number N has the effect of providing a more detailed understanding of the relationship between the input and the final outcome. In other words, the response utterance 640 includes a number of causal consequences specified by the user that intervene between the input and the output. Consequently, outcomes that the user did not anticipate when uttering the user input 630 are obtained, including the process leading up to them, which clarifies the thought process of the dialogue device 612 in the dialogue and provides an opportunity for the dialogue to develop.

[0118] Furthermore, during generation, the input to the response generation neural network 632 does not necessarily include additional words from the information addition device 642; these may be omitted.

[0119] Fourth Embodiment 1. Configuration Using the response generation neural network 632 trained in the third embodiment, the following embodiments are also possible. Figure 13 shows the configuration of the dialogue device 730 according to the fourth embodiment.

[0120] Referring to Figure 13, the dialogue device 730 includes the response generation neural network 632 described above, and an information addition device 642 that receives user input 740 and outputs the user input 740 with related words (groups) indicating a topic attached to it. The dialogue device 730 further includes a chain count addition unit 742 that adds an externally provided natural number N, which represents the number of causal consequences chained as described in the third embodiment, to the output of the information addition device 642 and inputs it to the response generation neural network 632. The dialogue device 730 further includes an upper limit storage unit 746 that stores an upper limit of the chain count, and a counting unit 744 that provides the natural number N to the chain count addition unit 742 in increments of 1 from 1 up to the upper limit stored in the upper limit storage unit 746, and sequentially outputs the output of the information addition device 642 with the natural number N attached.

[0121] The dialogue device 730 further includes an output storage unit 748 for storing a series of causal consequence sequences output successively from the response generation neural network 632 in response to these inputs, and a ranking unit 750 for ranking the series of causal consequence sequences stored in the output storage unit 748. The dialogue device 730 further includes an utterance selection unit 752 for selecting the causal consequence sequence ranked highest by the ranking unit 750 as an utterance candidate, and an utterance formatting unit 754 for formatting the causal consequence sequence selected by the utterance selection unit 752 to be suitable as a response to the user input 740 and outputting it as a response utterance 756.

[0122] For the ranking unit 750, a pre-trained neural network can be used. This neural network can be trained using sets of causal consequence chains for combinations of user input and related words, which can be manually evaluated.

[0123] 2 Effects This embodiment provides the following benefits. For example, assuming the input is "artificial intelligence develops" and the additional word "elderly people" is selected, the final outcome is "support the elderly," but a chain of multiple causal consequences can be considered in between. Moreover, the number of these multiple causal consequences is not known in advance. However, in this embodiment, a chain of causal consequences including the causal consequences of all numbers within the range of input natural numbers is generated, and the one with the highest rank can be output as response utterance 756. As a result, a causal consequence that the user had not considered at the time of user input 740, and that is appropriate as a response, can be selected.

[0124] Furthermore, the number of causal chains during inference does not need to be limited to the natural number N of the training data. For example, even if the maximum value of the natural number N in the training data is 10, there is no need to restrict N to 10 during inference. As long as a certain level of accuracy is achieved, it is possible to perform inference processing with a natural number N of 15 or higher.

[0125] Fifth embodiment 1. Configuration The fourth embodiment described above retrospectively evaluates multiple response causal consequences generated for multiple natural numbers N and selects the one with the highest score. However, this invention is not limited to such embodiments. It is also conceivable to perform a similar evaluation in advance and utilize the results in the dialogue system. The fifth embodiment is an example of such an example. Specifically, in this example, for each combination of a hypothetical input and its related word, multiple causal consequence chains consisting of 1 to N causal consequences are created in advance by a human or using the response generation neural network 632 of the fourth embodiment. The results are then ranked manually. Then, training data is created using the combination of the hypothetical input and its related word as input, and the natural number N obtained when the highest rank is obtained among the causal consequence chains from that combination as the ground truth data. Using this training data, a neural network (N-evaluation neural network) that outputs a natural number N when a user input is given is trained. Then, the input and related word are input to this trained N-evaluation neural network to obtain a natural number N, and "input + related word + N" is given to the response generation neural network to obtain an output.

[0126] Specifically, referring to Figure 14, the dialogue system 770 according to the fifth embodiment includes an information adding device 642 that adds relevant words to user input 780, and a chain count estimation unit 782 which includes the above-described N evaluation neural network that obtains the output of the information adding device 642 and outputs the optimal value of a natural number N for the combination of user input and relevant words included in the output. The dialogue system 770 further includes a chain count adding unit 784 that adds a natural number N, which is the chain count output by the chain count estimation unit 782, to the output of the information adding device 642.

[0127] The dialogue system 770 further includes a response generation neural network 632 that receives user input, related words, and a natural number N as input, which are the output of the chain count addition unit 784, and outputs a causal consequence chain containing a specified N causal consequences, and an utterance formatting unit 786 that formats the causal consequence chain output by the response generation neural network 632 so that it is appropriate as a response to the user input 780 and outputs a response utterance 788.

[0128] 2 Effects According to this fifth embodiment, the user simply inputs (speaks) user input 780, and the dialogue system 770 estimates related words, and the chain count estimation unit 782 estimates a natural number N and provides it to the response generation neural network 632. Therefore, the user does not need to think of related words or specify a natural number N. Furthermore, the response generation neural network 632 provides an appropriate response utterance 788, and the user can easily understand what corresponds to the thought process of the dialogue system 770 from the causal consequences in the response utterance 788. As a result, the dialogue with the dialogue system 770 can be developed in a meaningful way for the user.

[0129] In the descriptions of the third, fourth, and fifth embodiments so far, it has been assumed that the causal consequences at the end include the words designated as related words or words similar to them. However, this is not necessarily required, and learning may also be performed even if the expression representing the topic is not included in the chain of consequences.

[0130] 6. Implementation by computer Figure 15 is an external view of a computer system that implements each of the above embodiments. Figure 16 is a hardware block diagram of the computer system shown in Figure 15.

[0131] Referring to Figure 15, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972, all connected to the computer 970, for user interaction. Of course, these are just one example of a configuration for when user interaction is required, and any general hardware and software available for user interaction (e.g., touch panels, voice input, pointing devices in general) can be used.

[0132] Referring to Figure 16, the computer 970 includes, in addition to the DVD drive 1002, a CPU (Central Processing Unit) 990, a GPU (Graphics Processing Unit) 992, and a bus 1010 connected to the CPU 990, GPU 992, and DVD drive 1002. The computer 970 further includes a ROM (Read-Only Memory) 996 connected to the bus 1010 for storing the computer 970's boot-up program, etc., a RAM (Random Access Memory) 998 connected to the bus 1010 for storing program instructions, system programs, and work data, etc., and a non-volatile memory SSD (Solid State Drive) 1000 connected to the bus 1010. The SSD 1000 is for storing programs executed by the CPU 990 and GPU 992, as well as data used by programs executed by the CPU 990 and GPU 992. Computer 970 further includes a network interface 1008 that provides a connection to a network 986 enabling communication with other terminals, and a USB port 1006 that allows a USB (Universal Serial Bus) memory 984 to be attached and detached, and provides communication between the USB memory 984 and various parts within Computer 970.

[0133] Computer 970 further includes an audio interface 1004 connected to the microphone 982 and speaker 980 and bus 1010. The audio interface 1004 reads audio signals, video signals, and text data generated by the CPU 990 and stored in RAM 998 or SSD 1000 according to the instructions of the CPU 990, performs analog conversion and amplification processing to drive the speaker 980, and digitizes the analog audio signal from the microphone 982 and stores it in RAM 998 or SSD 1000 at any address specified by the CPU 990.

[0134] In the above embodiment, programs for implementing the dialogue systems 50, 300, 600, and 770, or parts thereof such as the training data generation devices 54 and 610, the training device 56, the dialogue devices 52, 302, 612, 730, and the training data expansion unit 58, as well as neural network parameters and neural network programs, training data, causal relationships, causal consequence chains, and assumed inputs, are all stored in storage media of external devices (not shown) connected via network I / F 1008 and network 986, for example, as shown in Figure 16: SSD 1000, RAM 998, DVD 978, USB memory 984, or network I / F 1008 and network 986. Typically, this data and parameters are written to the SSD 1000 from an external source and loaded into the RAM 998 when executed by the computer 970. In some cases, some of the code executed by the program may be loaded into the RAM 998 from outside the computer 970 via network 986 at the time of execution.

[0135] The computer programs for operating this computer system to realize the functions of the systems and their components in each of the embodiments described above are stored on a DVD 978 inserted into the DVD drive 1002 and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs are stored on a USB memory 984, the USB memory 984 is inserted into the USB port 1006, and the programs are transferred to the SSD 1000. Alternatively, these programs may be transmitted to a computer 970 via the network 986 and stored in the SSD 1000.

[0136] The program is loaded into RAM998 at runtime. Of course, the source program may be input using the keyboard 974, monitor 972, and mouse 976, and the compiled object program may be stored in SSD1000. In the case of a scripting language, the script entered using the keyboard 974, etc., may be stored in SSD1000. In the case of a program that runs on a virtual machine, a program that functions as a virtual machine must be installed on computer 970 in advance. Since training and testing neural networks involves a large amount of computation, it is preferable to implement each part of the embodiment of the present invention as an object program consisting of the computer's native code rather than a scripting language.

[0137] The CPU990 reads the program from RAM998 according to the address indicated by an internal register called the program counter (not shown), interprets the instructions, reads the data necessary for executing the instructions from RAM998, SSD1000, or other devices according to the address specified by the instructions, and executes the processing specified by the instructions. The CPU990 stores the execution result data at an address specified by the program, such as RAM998, SSD1000, or a register within the CPU990. At this time, the value of the program counter is also updated by the program. Computer programs may be loaded directly into RAM998 from DVD978, USB memory 984, or via a network. In addition, some tasks (mainly numerical calculations) within the program executed by the CPU990 are dispatched to the GPU992 according to the instructions included in the program or according to the analysis results when the CPU990 executes the instructions.

[0138] A program that implements the functions of the systems and their components according to each embodiment described above using the computer 970 includes a plurality of instructions written and arranged to operate the computer 970 to implement those functions. Some of the basic functions necessary to execute these instructions are provided by the operating system (OS) running on the computer 970, a third-party program, or modules of various toolkits installed on the computer 970. Therefore, this program does not necessarily have to include all the functions necessary to implement the system and method of this embodiment. This program only needs to include instructions that perform the operations of each of the above-described devices and their components by statically linking appropriate functions or functions of the "program library" in a controlled manner to obtain the desired result, or by dynamically calling them during program execution. The method of operating the computer 970 for this purpose is well known and will not be repeated here.

[0139] Furthermore, the GPU992 is capable of parallel processing, allowing it to execute large amounts of calculations associated with machine learning concurrently, in parallel, or in a pipelined manner. For example, parallel computation elements discovered in the program during compilation, or during program execution, are dispatched from the CPU990 to the GPU992 as needed, executed, and the results are returned to the CPU990 directly or via a predetermined address in RAM998, and assigned to a predetermined variable in the program.

[0140] 7. Variation In the above embodiment, causal relationships collected from the web are used as is, and only appropriate extended causal relationships are used. However, this invention is not limited to such embodiments. For example, the causal relationships used for training may be filtered in some way. For example, causal relationships with a positive or negative sentiment polarity may be selected. Only causal relationships whose affinity to a specific topic has been determined by a topic word model may be used. Alternatively, responses may be manually or automatically labeled to determine whether they are appropriate as dialogue responses, and only those deemed appropriate may be used.

[0141] The neural network used in the above embodiment is a pre-trained neural network (for example, the aforementioned BERT) that has been fine-tuned with training data prepared based on causal relationships. However, the neural network is not limited to this type; for example, a network that produces causal consequences by providing some kind of special type of input, such as GPT-3 described in the following literature, may also be used.

[0142] References: Tom B. Brown et al., Language Models are Few-Shot Learners, https: / / arxiv.org / pdf / 2005.14165.pdf Furthermore, when generating a response, for example, the user input is repeated, and sentence endings are also processed. ofThe system may transform and return responses according to a deep learning-based neural network or rules. It may also trace further causal relationships in response to user interjections in response to the dialogue system's replies and present the consequences in the form of a response. How to interpret these interjections can be determined by creating separate training data and using a neural network trained with deep learning. .versus If the user gives a negative response to the system's response, the system may ask a question to explain why, and then search for a causal relationship in which the words used in the response have a consequence, and present the cause. Similarly, the system may explain to the user the process by which the response was generated.

[0143] Conversely, if the user responds positively to the dialogue system's response, a causal relationship with the response's wording as the consequence can be searched for, and the response may be generated from the cause.

[0144] Furthermore, in the second to fifth embodiments described above, words or assumed inputs are added to user inputs 102, 630, 740, and 780. However, this invention is not limited to such embodiments. The combination of words to be added and assumed inputs may be converted into fixed-length feature vectors using an encoder consisting of a neural network, and these feature vectors, added to user inputs 102, 630, 740, and 780, may be used as input to response generation neural networks 340 and 632.

[0145] The embodiments disclosed herein are illustrative and not limited to those embodiments. The scope of the present invention is defined by each claim, with reference to the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wording contained herein. [Explanation of Symbols]

[0146] 50, 300, 600, 770 Dialogue Systems 52, 302, 612, 730 Dialogue device 54, 610 Training data generation device 56 Training equipment 58 Training Data Expansion Unit 60 Internet 62. Causal Relationship Extraction Unit 64 Causal Relationship Database 66. Chain Causal Relationship Generation Unit 68 Generating Causal Relationship Database 70. Chain of Causal Relationships Selection Unit 72 Chain Causal Relationship Database 74 Extended Causal Relationship Database 76. Assume Input Extraction Unit 78 Expected Input Storage Unit 80 Console 82,626 Training Data Creation Department 84 Training data storage unit 86 Training Department 100, 340, 632 Response Generating Neural Networks 102, 630, 740, 780 User input 104, 638, 754, 786 Speech Shaping Department 106, 640, 756, 788 Response utterances 150 Expected Input Reading Unit 152 Noun phrase identification part 154 Causal Relationship Search Department 156 Training Data Sample Creation Department 330, 344 Topic Word Models 332 Related Expression Search Section 334 Training Data Addition Section 338, 642 Information Addition Device 346 User input storage unit 348 Word Selection Section 350, 752 Speech Selection Section 352 Information Addition Unit 360 Training Data Readout Unit 362 Related Expressions Inquiry Department 364 Training data writing section 366 Related Expression Addition Section 400 Related Assumption Input Search Unit 402 Related Word Search Section 404 Related Expression Development Section 406 Related Expression Output Section 614 Question Answering System 620 Causal Consequence Chain Generation Unit 622 Causal Consequence Chain Memory Unit 624 Causal Consequence Chain DB 634 Question generation part 750 Ranking Section 782 Chain count estimation unit

Claims

1. A hypothetical input storage means that stores multiple hypothetical inputs, each of which is assumed to be an input to the dialogue device, It includes a causal relationship memory means for storing multiple causal relationship representations, Each of the aforementioned multiple causal relationship expressions includes a cause expression and a consequence expression, For each of the plurality of assumed inputs stored in the assumed input storage means, A means for extracting a causal relationship expression from a plurality of causal relationship expressions, in which the noun phrase of the assumed input is the cause expression, A training data creation means creates a training data sample in which the assumed input is used as the input unit and the consequence expression of the causal relationship expression extracted by the causal relationship expression extraction means is used as the response unit, and stores it in a predetermined storage device. A neural network training apparatus, comprising training means for training a neural network, which is designed to generate output sentences for natural language input sentences, using the training data samples stored in a predetermined memory device.

2. Furthermore, given a word, a pre-trained topic word model outputs the distribution probability of surrounding words for each word in a predetermined vocabulary, The training apparatus according to claim 1, further comprising: a first training data sample adding means for, for each of the causal relationship expressions of the training data samples stored in the predetermined storage device, identifying a predetermined number of words with the highest distribution probability for the words included in the causal relationship expression based on the output of the topic word model, and adding them to the input section of the training data sample to generate a new training data sample and add it to the predetermined storage device.

3. The training apparatus according to claim 2, wherein the first training data sample addition means includes a second training data sample addition means for identifying a predetermined number of words with the highest distribution probability for each of the causal relationship expressions of the training data samples stored in the predetermined storage device as related words based on the output of the topic word model, further extracting sentences containing the related words from the plurality of assumed inputs stored in the assumed input storage means, and adding them together with the related words to the input portion of the training data sample to generate a new training data sample and add it to the predetermined storage device.

4. Furthermore, given a word, a pre-trained topic word model outputs the distribution probability of surrounding words for each word in a predetermined vocabulary, A related word identification unit identifies, based on the output of the topic word model, a group of related words consisting of a predetermined number of words with the highest distribution probability for each of the training data samples stored in the predetermined storage device, It includes a causal consequence chain memory for storing multiple causal consequence chains, Each of the aforementioned multiple causal consequence chains is a string formed by sequentially chaining together the causal part of the first causal relationship and the consequence parts of the causal relationships included in the chain of causal relationships that begin with the first causal relationship. Furthermore, the causal consequence chain reading means is included for reading a causal consequence chain from the causal consequence chain storage device for each of the plurality of assumed inputs stored in the assumed input storage means, the causal consequence chain having the noun phrase of the assumed input as the cause part and any word included in the related word group as the consequence part at the end. The training apparatus according to claim 1, wherein the training data creation means includes a training data sample adding means for creating a training data sample for each of the causal consequence chains read by the causal consequence chain reading means for each of the plurality of assumed inputs, with the combination of the assumed input, the related word group, and a natural number N representing the length of the causal consequence chain as the input unit, and the causal consequence chain as the answer unit, and adding it to the predetermined storage device.

5. A natural language dialogue device, comprising a neural network designed to generate output sentences from natural language input sentences, A dialogue device wherein the neural network is trained by the training device described in claim 1 such that the output sentence represents a latent consequence of the input sentence.

6. A natural language dialogue device, comprising a neural network designed to generate output sentences from natural language input sentences, The neural network is trained by the training device described in claim 2. The dialogue device further includes related expression adding means that, in response to user input, adds related expressions, which are expressions containing words related to the phrases in the user input, to the user input and input them to the neural network. The aforementioned related expression addition means is A speech memory unit that stores the user's past utterances, A topic word model that outputs the probability distribution of the occurrence of surrounding words for an input word, A related word identification unit takes user utterances as input and uses the topic word model to identify related word groups that are associated with words contained in the user's past utterances stored in the utterance memory unit, A dialogue device comprising: related information adding means for adding related information consisting of any combination of words included in the group of related words identified by the related word identification unit to the user utterance as the related expression.

7. A natural language dialogue device, comprising a neural network designed to generate output sentences from natural language input sentences, The neural network is trained by the training device described in claim 3. The dialogue device further includes related expression adding means that, in response to a given user utterance, adds related expressions, which are expressions containing words or sentences related to phrases in the user utterance, to the user utterance and inputs them to the neural network. The aforementioned related expression addition means is A speech memory unit that stores the user's past utterances, A topic word model that outputs the probability distribution of the occurrence of surrounding words for an input word, A related word identification unit takes user utterances as input and uses the topic word model to identify related words that are associated with words contained in the user's past utterances stored in the utterance memory unit, A past utterance extraction unit retrieves utterances related to the user utterance from the user's past utterances stored in the utterance storage unit. A dialogue device including related information adding means for adding related information to a user utterance as a related expression, which consists of any combination of the related word identified by the related word identification unit and the user's past utterance extracted by the past utterance extraction unit.

8. A natural language dialogue device, comprising a neural network designed to generate an output sentence from a natural language input sentence in which natural numbers are appended at predetermined positions, The neural network is trained by the training device described in claim 4. A topic word model that outputs the probability distribution of the occurrence of surrounding words for an input word, A related word group generation unit generates a related word group using the topic word model by inputting the words contained in the input sentence into the topic word model in response to the input sentence, An input creation unit that combines the aforementioned input sentence, the words within the aforementioned related word group, and the aforementioned natural numbers and inputs them into the neural network, A question unit generates a question sentence from a causal relationship expression output by the neural network in response to input from the input generation unit, and inputs it into a question answering system to obtain an answer. A dialogue device comprising: an answer generation unit that generates an answer to an input sentence using the answer from the question answering system to a question sentence from the question unit.

Citation Information

Patent Citations

  • Training device for question answering system and computer program therefor

    JP2017049681A

  • Dialogue system, dialogue apparatus and computer program therefor

    JP2018156272A

  • Creation method of training data of question answering system and training method of question answering system

    JP2019133229A