Utterance filtering device, dialog system, context model learning data generation device and computer program

JP2024011901A5Pending Publication Date: 2025-06-23NAT INST OF INFORMATION & COMM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022114229
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Existing dialogue systems struggle to prevent the output of inappropriate expressions that may not contain problematic keywords but become offensive based on context, leading to potential criticism or issues.

Method used

An utterance filtering device that calculates the probability of words appearing in context using a context model trained on word vectors, determining whether to discard or approve utterances based on predetermined thresholds, considering the context and relationships between words.

Benefits of technology

Reduces the likelihood of outputting problematic utterances by evaluating the context and relationships of words, ensuring appropriate responses in dialogue systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an utterance filtering device capable of preventing output of a potentially problematic expression in an interactive system outputting an utterance in an interactive manner.SOLUTION: An utterance filtering device includes: a context model which is pre-learned to output a probability vector with an element for a probability of each word in a predetermined word group appeared in a context in which the utterance is placed when a word vector string representing the utterance is input; and a determination part 456 which determines whether the utterance should be discarded or approved in accordance with whether or not a value, which is determined as a prescribed function of the probability vector output by the context model in response to input of the word vector string representing the utterance to the context model, is greater than or equal to a threshold value.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a dialogue device, and more particularly to a technique for determining whether or not a system utterance generated by a dialogue device includes an inappropriate expression. [Background technology]

[0002] Systems in which users and systems communicate in some form of dialogue, such as search engines, question answering systems, and dialogue systems, are becoming more and more common. In such systems, it is desirable to ensure that system responses (hereinafter referred to as "system utterances") do not include inappropriate language.

[0003] A direct way to deal with this problem is to make a list of problematic keywords. Check the system utterance candidates from the beginning to see if they contain any of these keywords. If a system utterance candidate contains even one of these keywords, reject it and select the next system utterance candidate. If a system utterance candidate that does not contain any of the listed keywords is found, output that system utterance candidate.

[0004] Such a technology is disclosed in the below-mentioned Patent Document 1. In the technology disclosed in Patent Document 1, when dynamic content is displayed by a browser, the browser determines whether or not the dynamic content contains problematic expressions such as hate speech. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent Publication No. 2022-082538 Summary of the Invention [Problem to be solved by the invention]

[0006] The technology disclosed in Patent Document 1 is such that when a browser displays dynamic content, the browser transmits the dynamic content received from an application to a server that checks the content, and receives the check results from the server. As described above, the list of problematic keywords is used for the server's judgment.

[0007] The technology disclosed in Patent Document 1 is a judgment for the entire content. Therefore, if there is a problematic expression in the content, it is possible to stop the display of only that part or the entire content.

[0008] In contrast, the output of a dialogue system is generally one utterance. Therefore, if the technology disclosed in Patent Document 1 is applied to a dialogue system, if the utterance of the system contains a problematic keyword, the utterance will not be output, and if not, the utterance will be output.

[0009] However, in real speech, even if the speech itself does not contain problematic keywords, some speech may be problematic depending on the context. For example, after listing expressions such as "skin color" or "place of origin" as problematic expressions, a comment is made about the expression, or an utterance is made that implies malice. In this case, even if the comment itself is not malicious, or the expression itself cannot be said to be malicious, the output of the problematic expression itself may be problematic. For example, if such an expression is output on a site that provides a public service or a site operated by a company, there is a risk that the expression will be criticized by users even if it is an expression that should not be problematic if the context is examined. The output of a question-answering system, a dialogue system, etc. may only be a short expression, and a technology that checks the entire content and determines whether to output it, such as the system described in Patent Document 1, cannot prevent the output of a potentially problematic expression.

[0010] Therefore, an object of the present invention is to provide an utterance filtering device that prevents potentially problematic expressions from being output in an interactive system that outputs utterances in an interactive manner. [Means for solving the problem]

[0011] A speech filtering device according to a first aspect of the present invention includes a context model that has been trained in advance to output a probability vector whose elements are the probability that each word included in a predetermined word group will appear in the context in which the utterance is placed when a word vector sequence representing an utterance is input, and a determination means for inputting the word vector sequence representing the utterance into the context model and determining whether the utterance should be discarded or accepted depending on whether at least one element of the probability vector output by the context model in response to the input satisfies a predetermined condition.

[0012] Preferably, the determining means includes means for determining whether to discard or accept the utterance depending on whether a value determined as a predetermined function of at least one element of the probability vector is greater than or equal to a predetermined threshold.

[0013] A dialogue system according to a second aspect of the present invention includes a dialogue device, the above-mentioned speech filtering device coupled to the dialogue device so as to receive as input candidate utterances output by the dialogue device, and speech filtering means for filtering the utterances output by the dialogue device in accordance with the determination results by the speech filtering device.

[0014] A computer program according to a third aspect of the present invention causes a computer to function as: a context model that has been pre-trained to output a probability vector whose elements are the probability that each word included in a predetermined word group will appear in the context in which the utterance is placed when a word vector sequence representing an utterance is input; and a determination means for inputting the word vector sequence representing the utterance into the context model and determining whether the utterance should be discarded or approved depending on whether the probability of any word included in the predetermined word group is equal to or greater than a threshold value based on the probability vector output by the context model in response to the input.

[0015] A training data generation device according to a fourth aspect of the present invention includes: a context extraction means for extracting the context of each utterance stored in a corpus; a context vector generation means for generating a context vector indicating whether each of words included in a predetermined word group appears at least in the context; and a training data generation means for generating training data for each utterance stored in the corpus by combining the utterance as input and the context vector as output.

[0016] Preferably, the context extraction means includes a preceding and following utterance extraction means for extracting utterances before and after each utterance stored in the corpus as the context of the utterance.

[0017] More preferably, the context extraction means includes a subsequent utterance extraction means for extracting, as the context of each utterance stored in the corpus, an utterance immediately following that utterance.

[0018] More preferably, the corpus includes a plurality of causal relationships expressions each including a cause part and a result part, and the context extraction means includes result part extraction means for extracting, for each of the plurality of causal relationships expressions, the cause part of the causal relationship expression as an utterance and the result part of the causal relationship expression as the context of the utterance.

[0019] A computer program according to a fifth aspect of the present invention causes a computer to function as: context extraction means for extracting the context of each utterance stored in a corpus; context vector generation means for generating a context vector indicating whether each of words included in a predetermined word group appears at least in the context; training data generation means for generating training data for each utterance stored in the corpus by combining the utterance as input and the context vector as output; and learning means for training a context model consisting of a neural network using the training data generated by the training data generation means.

[0020] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief description of the drawings]

[0021] [Figure 1] FIG. 1 is a block diagram showing the configuration of a dialogue system according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is a flowchart showing a control structure of a computer program implementing the learning data generating unit shown in FIG. [Diagram 3] FIG. 3 is a flowchart showing a control structure of a computer program for implementing the steps shown in FIG. [Figure 4] FIG. 4 is a block diagram showing the configuration of the context model shown in FIG. [Diagram 5] FIG. 5 is a block diagram showing a mechanism for learning the context model shown in FIG. [Figure 6] FIG. 6 is a flowchart showing a control structure of a computer program implementing the interactive apparatus shown in FIG. [Figure 7] FIG. 7 is a flowchart showing a control structure of a computer program corresponding to FIG. 6 in accordance with the modification of the first embodiment. [Figure 8] FIG. 8 is a block diagram showing the configuration of a dialogue system according to the second embodiment of the present invention. [Figure 9] FIG. 9 is a flowchart showing a control structure of a computer program implementing the learning data generating unit shown in FIG. [Figure 10] FIG. 10 is a flowchart showing a control structure of a computer program for implementing part of the processing shown in FIG. [Figure 11] FIG. 11 is a block diagram showing the configuration of a dialogue system according to the third embodiment of the present invention. [Figure 12] FIG. 12 is a flowchart showing a control structure of a computer program implementing the dialogue system shown in FIG. [Figure 13] FIG. 13 is an external view of a computer for implementing each embodiment of the present invention. [Figure 14] FIG. 14 is a hardware block diagram of the computer system whose external appearance is shown in FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] In the following description and drawings, the same parts are given the same reference numbers, and therefore detailed descriptions thereof will not be repeated.

[0023] 1. First embodiment A.Configuration Referring to FIG. 1, a dialogue system 50 according to a first embodiment of the present invention includes a dialogue device 62, a context model 80 used in the dialogue device 62 when filtering candidates for system utterances, a passage DB (Database) 70 that stores a plurality of passages, and a context model learning system 60 for learning the context model 80 using each passage stored in the passage DB 70.

[0024] The dialogue device 62 includes a dialogue engine 84 for receiving an input utterance 82 and generating and outputting a plurality of response candidates as responses to the input utterance 82, and a filtering unit 86 for filtering the plurality of response candidates output by the dialogue engine 84 using a context model 80, and outputting as a system utterance 88 a response candidate that is determined by the context model 80 to be acceptable and is determined to be optimal as a response to the input utterance 82.

[0025] In this embodiment, the dialogue engine 84 has a function of selecting a plurality of sentences that are considered appropriate as responses to the input utterance 82 from among sentences collected from the Internet, calculating a score indicating the appropriateness of each as a response to the input utterance 82, and outputting a predetermined number of sentences with the highest scores as response candidates. For example, the dialogue engine 84 can be the dialogue system disclosed in JP 2019-197498 A. In the dialogue system described in the above document, candidates for the system utterance are selected from a large number of sentences collected in advance. In particular, the more sentences collected in advance, the higher the possibility of finding an appropriate response to the input utterance 82. Therefore, these many sentences are collected in advance from the Internet. As is well known, many sentences existing on the Internet can be problematic as expressions. Therefore, the problem is what kind of sentence should actually be selected as the system utterance.

[0026] The passage DB 70 stores a plurality of passages. Each of the plurality of passages includes a plurality of consecutive sentences that are part of a sentence. The number of sentences included in each passage ranges from 3 to 9, for example. In this embodiment, the number of sentences included in each passage stored in the passage DB 70 varies. As described above, all of these passages have been collected in advance from the Internet.

[0027] The context model learning system 60 includes a topic word list 74 that lists topic words prepared in advance, including expressions, keywords, concepts, etc. that may be problematic or indicate problems, and a learning data creation unit 72 for generating learning data for the context model 80 by using each of the topic words stored in the topic word list 74 based on each passage stored in the passage DB 70. In this embodiment, the topic word list 74 is assumed to be a file in which problematic keywords are separated by a predetermined delimiter and recorded in a computer-readable storage medium. The number of topic words is assumed to be N.

[0028] The context model training system 60 further includes a training data memory unit 76 for storing the training data generated by the training data creation unit 72, and a training unit 78 for performing training of the training unit 78 using the training data stored in the training data storage unit 76.

[0029] The learning data creation unit 72 shown in Fig. 1 is realized by computer hardware and a computer program executed by the computer hardware. With reference to Fig. 2, the computer program includes step 150, after startup, for performing initialization processing such as allocating and initializing a memory area to be used by the program, opening files to be used, reading initial parameters, and setting parameters for accessing a database, and step 152 for reading the topic word list 74 shown in Fig. 1 from a file, separating them at the positions indicated by delimiters, and expanding and storing them in memory as elements of an array T.

[0030] This program also sets the variable MAX T 1. In this embodiment, the subscripts of the array T start from 0. That is, the number of elements of the array T is the variable MAX T +1.

[0031] The program further includes step 158 of generating training data for context model 80 by executing the following step 160 for each passage stored in passage DB 70, and step 162 of storing the training model generated in step 158 in training data storage unit 76 and terminating execution of the program.

[0032] Step 160 includes a step 200 for dividing the passage to be processed into sentences and expanding each sentence into an array S, and a step 200 for expanding a variable MAX S Step 160 further includes step 202 of substituting the value of the maximum index of the array S into the iteration control variables j=1 to j=MAX. SThe method includes step 204 of executing the process of creating learning data in step 206 for each value of the variable j up to -1.

[0033] 3, step 206 shown in FIG. 2 includes step 250 of generating vector Z having N+1 elements, all of which are zero, step 252 of substituting a character string obtained by concatenating S[j-1], S[j], and S[j+1] into character string variable S3, and step 254 of repeatedly executing step 256 while incrementing variable i by 1 from i=0 to N-1. Vector Z is generated from element Z0 to element Z N As described above, N is the number of topic words listed in the topic word list 74 (see FIG. 1).

[0034] Step 256 includes step 300, which branches the flow of control according to whether the topic word to be processed, i.e., element T[i] with subscript = 0 of array T, exists in the string represented by string variable S3, and when the determination in step 300 is affirmative, the i-th element Z of vector Z is i and step 302 of assigning 1 to x = 1. If the determination in step 300 is negative and after step 302, step 256 ends.

[0035] Step 206 further includes step 258 of assigning the number of non-zero elements of vector Z to variable M after completion of step 254, and step 260 of branching the control flow according to whether the value of variable M is 0 or not. Step 206 further includes step 262 of assigning 1 to the N+1-th element of vector Z when the determination in step 260 is positive, step 264 of dividing vector Z by the value of variable M when the determination in step 260 is negative, and step 266 of adding, after steps 262 and 264, a record of the training data whose input is the j-th component of array S, i.e., S[j], and whose output is vector Z, to the training data and terminating step 206.

[0036] When the process of step 262 is executed, the N+1-th element Z N The value of only Z becomes 1, and all other elements Z k (k=0 to N-1) has a value of 0. When step 264 is executed, element Zk (k=0 to N-1) of the vector Z has a value of 1 / M if a topic word corresponding to the element exists in the string substituted for the string variable S3, and has a value of 0 otherwise. On the other hand, element Z N The value of is 1 if there is no topic word corresponding to that element in the string assigned to string variable S3, and 0 otherwise.

[0037] 4 shows a schematic configuration of the context model 80. Referring to FIG. 4, the context model 80 includes a neural network BERT (Bidirectional Encoder Representations from Transformers) 352 that receives as input an utterance 350 with a CLS token 340 indicating the beginning of the input at the beginning and a SEP token 342 indicating a sentence boundary at the end, and a fully connected layer 358 having N+1 outputs that is connected to receive as a vector the contents of a CLS-compatible layer 356, which is a transformer layer corresponding to the CLS token 340, in the final hidden layer 354 of the BERT 352. The context model 80 further includes a SoftMax layer 360 for performing a softMax operation on the N+1 outputs from the fully connected layer 358 and outputting a probability vector 362. In this embodiment, the BERT 352 is a pre-trained BERT Large It is.

[0038] Fig. 5 illustrates the relationship between the BERT 352 and training data during training of the BERT 352. As described above, the training data 400 includes a sentence (element S[j] at the time of creating the training data) as an input, and has a vector Z as an output (correct answer data).

[0039] During training, a CLS token 340 is added to the beginning of each sentence in training data 400 and a SEP token 342 is added to the end, and the sentences are input to BERT 352. In response to this input, a probability vector 362 is obtained at the output of SoftMax layer 360. Training of BERT 352 and fully connected layer 358 is performed by the backpropagation method using the error between each element of this probability vector 362 and the correct label vector 404 in training data 400.

[0040] Referring to Figure 6, the program for realizing the filtering unit 86 shown in Figure 1 includes a step 450 of inputting an input utterance 82 to a dialogue engine 84, and a step 452 of obtaining a system utterance candidate list output from the dialogue engine 84 in response to the processing in step 450.

[0041] This program further includes step 454 for executing step 456 for determining whether each candidate in the system utterance candidate list obtained in step 452 is suitable as a system utterance, approving and retaining it if suitable, and rejecting it if inappropriate, and after step 454 is completed, step 458 for modifying the approved candidates so that they become suitable as system utterances for the input utterance 82, re-scoring and re-ranking them, and outputting the system utterance candidate with the highest score as the system utterance 88 (Figure 1).

[0042] Step 456 includes step 480 of inputting the target system utterance candidates into a context model 80, step 482 of obtaining a probability vector 362 output from the context model 80 as a result of the processing in step 480, and step 484 of obtaining, from among the probability vectors obtained in step 482, the maximum value of elements corresponding to one or more words that have been previously designated as undesirable words.

[0043] Step 456 further includes step 486 of determining whether the value obtained in step 484 is greater than a predetermined threshold and branching the control flow according to the determination, step 488 of discarding the system utterance candidate to be processed and terminating step 456 if the determination in step 486 is positive, and step 490 of approving and leaving the system utterance candidate to be processed and terminating step 456 if the determination in step 486 is negative.

[0044] B. Operation The dialogue system 50 according to the first embodiment operates as follows. The operation of the dialogue system 50 includes a learning phase and a dialogue phase. First, the operation of the dialogue system 50 (context model learning system 60) in the learning phase will be described below. Then, the operation of the dialogue system 50 (dialogue device 62) in the dialogue phase will be described.

[0045] B1. Learning Phase In the learning phase, first, a passage DB 70 is prepared. In this embodiment, each passage to be stored in the passage DB 70 is collected from the Internet. Similarly, a topic word list 74 is prepared. The topic word list 74 is, for example, a list of words whose frequency of appearance in the group of passages stored in the passage DB 70 is higher than a predetermined threshold value. In other words, this list can be automatically extracted from the passage DB 70 or the like by specifying a threshold value. In this embodiment, the topic word list 74 is a file that stores character strings in which each word is divided by a predetermined delimiter.

[0046] The learning data creation unit 72 generates learning data from the passage DB 70 while referring to the topic word list 74 as follows.

[0047] 1, when the context model learning system 60 starts up, the learning data creation unit 72 initializes each part of the computer (step 150 in FIG. 2. Unless otherwise specified, the step numbers are those shown in FIG. 2). In this process, the learning data creation unit 72 sets parameters for accessing the passage DB 70 and opens the topic word list 74. The dialogue device 62 also secures storage areas for the arrays T and S, the variables S3 and M, the repetition control variables i and j, and the vector Z.

[0048] Next, the learning data creation unit 72 reads the topic word list 74, separates it with a predetermined delimiter, and stores the contents in each element of the array T (step 152). The learning data creation unit 72 further T The maximum value of the subscript of the array T is substituted into (step 154). The learning data creation unit 72 then connects to the passage DB 70 shown in FIG. 1 (step 156). In this embodiment, the subscript of the array T ranges from 0 to the variable MAX T The value is up to .

[0049] The learning data creation unit 72 further executes the following step 160 for each passage stored in the passage DB 70, thereby generating a record of the learning data (step 158).

[0050] In step 160, the learning data creation unit 72 first divides the passage to be processed into sentences, and stores each sentence in each element of the array S (step 200). S (step 202). The learning data creation unit 72 further assigns the value of the maximum subscript of the array S to the iteration control variables j=1 to j=MAX in step 204. S Step 206 is performed for each value of variable j up to -1 to create a new record of the training data.

[0051] 3, in step 206, the learning data creation unit 72 generates a vector Z whose elements are all zero (step 250 in FIG. 3). That is, in this step, the vector Z is initialized. Next, the learning data creation unit 72 assigns a character string obtained by concatenating S[j-1], S[j], and S[j+1] to a character string variable S3 (step 252 in FIG. 3). Furthermore, the learning data creation unit 72 repeatedly executes step 256 while incrementing the variable i by 1 from the repetition variable i=0 to N-1 (step 254 in FIG. 3).

[0052] In step 256, the learning data creation unit 72 determines whether the element T[i] of the array T to be processed is present in the character string represented by the character string variable S3 (step 300 in FIG. 3). When the determination in step 300 is positive, the learning data creation unit 72 performs the following: i (Step 302 in FIG. 3). If the determination in step 300 is negative, nothing is done.

[0053] The context model learning system 60 executes step 256 while repeatedly incrementing the variable i by 1 from 0 to N-1. If the element T[i] exists in the string represented by the string variable S3, the i-th element Z i has a value of 1, otherwise element Zi has a value of 0.

[0054] After completing step 254, the learning data creation unit 72 assigns the number of non-zero elements of the vector Z to the variable M (step 258 in FIG. 3). The learning data creation unit 72 judges whether the value of the variable M is 0 or not (step 260 in FIG. 3). If the judgment in step 260 is positive, that is, if there is no non-zero element among the elements of the vector Z, the learning data creation unit 72 assigns 1 to the N+1-th element of the vector Z (step 262 in FIG. 3). If the judgment in step 260 is negative, that is, if there is at least one non-zero element in the vector Z, the learning data creation unit 72 divides the vector Z by the value of the variable M (step 264 in FIG. 3).

[0055] The context model learning system 60 executes step 206 shown in FIG. 3 to learn the context of a passage for a certain value of the variable j (1≦j≦MAX S If at least one word in topic word list 74 is present in the string (the value of string variable S3) formed by combining the sentence indicated by (-1) with the sentences before and after it, then the value of the elements of vector Z corresponding to those words will be 1 / M, and the value of the other elements will be 0. If none of the words in topic word list 74 are present in the string represented by string variable S3, then the Nth element Z of vector Z will be N will have a value of 1 and all other elements will have a value of 0.

[0056] Thereafter, the learning data creation unit 72 generates a new record of learning data corresponding to element S[j] by combining element S[j] as input and vector Z as output, and adds it to the learning data storage unit 76 (step 266).

[0057] After completing the creation of the learning data, the learning unit 78 uses the learning data to learn the context model 80.

[0058] Learning of the context model 80 by the learning unit 78 will be described with reference to FIG. 5. As described above, the learning data 400 includes a sentence (element S[j] at the time of creating the learning data) as an input, and has a vector Z as an output (correct answer data). The learning unit 78 shown in FIG. 1 reads one record of the learning data 400, generates a learning utterance 402 by adding a CLS token 340 to the beginning of the sentence and a SEP token 342 to the end, and inputs it to the BERT 352. The BERT 352 performs an operation on this input and changes the internal state of each hidden layer. The fully connected layer 358 receives the output vector of the CLS-compatible layer 356, which is the final hidden layer of the BERT 352, and inputs N+1 outputs to the SoftMax layer 360. The output of each position of the fully connected layer 358 is a numerical value representing the probability that the learning utterance 402 is related to the word corresponding to that position among the words listed in the topic word list 74. The correct label vector 404 performs a softMax operation on these N+1 numerical values, and outputs a probability vector 362 consisting of N+1 elements P(0) to P(N).

[0059] The learning unit 78 uses the error between this probability vector 362 and each element of the ground truth label vector 404 corresponding to the training utterance 402 to learn the parameters of the BERT 352 and the fully connected layer 358 by the error backpropagation method. In practice, the learning unit 78 repeats the above-mentioned process for each mini-batch selected from the training data until a predetermined termination condition is met. In this embodiment, this learning is performed by minimizing the value of the loss function L shown below.

[0060]

number

[0061] B2. Dialogue Phase 1, a user inputs an input utterance 82 to a dialogue engine 84. In response to the input utterance 82, the dialogue engine 84 selects a number of system utterance candidates that are considered appropriate as responses to the input utterance 82 from a large number of sentences previously collected from the Internet. The input utterance 82 calculates a score for each of the multiple system utterance candidates using a predetermined scoring method, and ranks the system utterance candidates based on the scores. The input utterance 82 provides a predetermined number of top system utterance candidates based on this ranking to a filtering unit 86.

[0062] In this embodiment, the filtering unit 86 inputs each system utterance candidate received from the dialogue engine 84 into the context model 80 and obtains a probability vector 362 as its output. The filtering unit 86 determines whether the probability value of an element of the probability vector 362 that is predefined as being unsuitable as a system utterance is greater than a predefined threshold (step 486). If the determination is positive, the filtering unit 86 discards the system utterance candidate (step 488). If the determination is negative, the filtering unit 86 accepts and leaves the system utterance candidate (step 490).

[0063] The filtering unit 86 modifies the remaining system utterance candidates in this way to make them suitable as responses to the input utterance 82. The filtering unit 86 rescores the modified system utterance candidates, and outputs the system utterance candidate with the highest score as the system utterance 88.

[0064] As described above, according to this embodiment, a system utterance in a dialogue is selected by considering not only the text of the system utterance candidate itself but also the possibility of words appearing in the context. A system utterance is usually one sentence, and there is actually no context before or after it. Therefore, it is difficult to determine whether or not the utterance may cause a problem from the system utterance alone. However, according to this embodiment, the system utterance is selected using information on the relationship that the system utterance may have with the context before and after it, so that the probability of some problem occurring by outputting the system utterance can be reduced.

[0065] C. Modifications In the first embodiment, as shown in steps 484 to 490 in FIG. 6, whether to discard or retain a candidate is determined according to whether the maximum value of the value of a specified element in the output probability vector is greater than a threshold value. That is, the value of the element in the output probability vector is used as it is for the judgment. However, this invention is not limited to such an embodiment. When judging whether the probability value of an element previously determined as being unsuitable for system utterance is a predetermined threshold value or not, the judgment may be made using not only one element of the probability vector but multiple elements. When making the judgment using multiple elements, it is possible to make the judgment based on the value of a logical expression of a condition for multiple elements, such as making a positive judgment when either the values ​​of two elements are both less than a predetermined threshold value or the other element is greater than the predetermined threshold value, or more generally, the judgment may be made using a value obtained by substituting one or multiple elements of the probability vector into a predetermined function. The following describes such a modified example.

[0066] Fig. 7 shows a control structure of a program for implementing the process corresponding to the process shown in Fig. 6 in a modified example of the first embodiment. This program differs from that shown in Fig. 6 in that it includes step 500 for executing step 502 for each candidate instead of step 454 in Fig. 6.

[0067] 7, step 502 includes steps 480 and 482, which are the same as those shown in FIG 6, step 510 for performing a predetermined operation between the elements of the output vector, and step 512 for branching the control flow depending on whether the result of the operation in step 510 is 1 or not. If the determination in step 512 is positive, i.e., the result of the logical operation in step 510 is 1, then the candidate being processed is discarded in step 488. If the determination in step 512 is negative, then the candidate being processed is accepted and retained in step 490.

[0068] In this embodiment, the calculation in step 510 is realized by previously setting up logic according to the conditions that the elements of the output probability vector should satisfy. i Then, a i represents the probability that the i-th word in the topic word list appears in the vicinity of a system utterance candidate. Therefore, by performing a predetermined logical operation on multiple elements of this output probability vector, a composite condition regarding whether the target system utterance candidate should be discarded or retained can be determined.

[0069] For example, for the condition "If the probability that the i1th word and the i2th word in the topic word list appear simultaneously around the system utterance candidate is higher than a threshold, discard the system utterance candidate," i1 *a i2 >If it exceeds the threshold, discard the system utterance candidate.

[0070] That is, this modification can also provide the same effects as the first embodiment. In the modification, more complex conditions can be set than in the first embodiment, so that the intention of the system developer can be more clearly reflected in the operation of the dialogue system.

[0071] In the first embodiment, the output probability vector is normalized by the SoftMAX function so that the sum of all element values ​​is 1. However, when performing the above-mentioned calculation, the output vector of the BERT before input to the SoftMAX function may be used as it is if the threshold value can be appropriately adjusted. Also, the first embodiment and the above-mentioned modified example may be combined.

[0072] 2. Second embodiment A.Configuration In the first embodiment, for each passage stored in the passage DB 70, as shown in Fig. 1, a context model 80 is trained using a target sentence, the sentence immediately preceding it, and the sentence immediately following it as context. However, in the second embodiment, a context model is trained using only the expression following the target expression as the context of the target expression.

[0073] The second embodiment also differs from the first embodiment in that learning data for a context model is created in such a way that the relationship between a target expression and the immediately following expression, which is its context, constitutes a causal relationship.

[0074] Referring to FIG. 8, a dialogue system 550 according to the second embodiment includes a context model 580, a context model learning system 560, and a dialogue device 562 that uses the learned context model 580 to filter system utterances and output system utterances 584 in response to input utterances 82.

[0075] The context model learning system 560 includes a corpus 570 that stores a large number of expressions collected from the Internet, a causal relationship extraction unit 572 for extracting sentences or expressions expressing causal relationships from the corpus 570, and a causal relationship corpus 574 for storing the extracted causal relationships by the causal relationship extraction unit 572.

[0076] A causal relationship is a phrase pair including a cause phrase, which is an expression expressing the cause of a causal relationship, and a result phrase, which is an expression expressing the result of the causal relationship. In this embodiment, training data for the context model 580 is generated by using a result phrase corresponding to a cause phrase as a context for the cause phrase.

[0077] The context model training system 560 further includes a topic word list 74, a training data creation unit 576 for creating each record of training data using each phrase pair stored in the causal corpus 574 while referring to the topic word list 74, and a training data storage unit 578 for storing each record of the training data created by the training data creation unit 576.

[0078] The context model training system 560 further includes a training unit 78 for training the context model 580 with training data stored in a training data storage unit 578 .

[0079] As in the first embodiment, the dialogue device 562 includes a dialogue engine 84 for receiving an input utterance 82 and outputting a plurality of system utterance candidates, and a filtering unit 582 for filtering the plurality of response candidates output by the dialogue engine 84 using a context model 580, and outputting as a system utterance 584 a response candidate that is determined by the context model 580 to be acceptable and optimal as a response to the input utterance 82.

[0080] For the process of extracting causal relationships from a corpus including a large number of documents, such as the causal relationship extraction unit 572, the technology disclosed in JP 2018-60364 A can be applied.

[0081] Referring to FIG. 9, a program executed by a computer to realize the context model learning system 560 shown in FIG. 8 includes step 620 for performing initialization immediately after startup, and step 152 for reading the topic word list 74 shown in FIG. 8 from a file, separating them at the points indicated by delimiters, and expanding and storing them in memory as elements of an array T.

[0082] This program also sets the variable MAX T 8 , step 622 of connecting to the causal relation corpus 574 shown in FIG. 8 , step 624 of creating training data by executing step 626 for each causal relation stored in the causal relation corpus 574, and step 628 of storing the training data created in step 624 in the training data storage unit 578 shown in FIG. 8 and terminating the process.

[0083] 10, step 626 shown in Fig. 9 has a control structure similar to that of the program for realizing step 206 of the first embodiment shown in Fig. 3. Step 626 differs from step 206 in that, instead of step 252 in Fig. 3, step 626 includes step 650 for substituting a result phrase of the causal relationship to be processed into a character string variable S3. Step 626 further differs from step 206 in that, instead of step 266 in Fig. 3, step 626 includes step 654 for adding a record of the learning data, the input of which is the cause phrase of the causal relationship to be processed and the output of which is the vector Z, to the learning data, and then terminating step 626.

[0084] B. Operation The dialogue system 550 shown in Fig. 8 according to the second embodiment operates as follows. The operation of the dialogue system 550 includes a learning phase and a dialogue phase. Of these, the configuration of the dialogue device 562 in the dialogue phase is the same as that of the dialogue device 62 in the first embodiment, except that the context model used is different, and the operation is also the same. Therefore, the operation of the dialogue system 550 (context model learning system 560) in the learning phase will be described below.

[0085] B1. Learning Phase Prior to the learning phase, a large amount of text is stored in the corpus 570. These texts may be collected, for example, from the Internet. A causal relation extraction unit 572 extracts causal relations from these large amounts of text and stores them in a causal relation corpus 574.

[0086] A training data creation unit 576 creates training data using each causal relationship stored in the causal relationship corpus 574 while referring to the topic word list 74 , and stores the data in a training data storage unit 578 .

[0087] 8, when the context model learning system 560 is started, the learning data creation unit 576 initializes each part of the computer (step 620 in FIG. 9. Unless otherwise specified, the step numbers are those shown in FIG. 9). In this process, the learning data creation unit 576 sets parameters for accessing the causal relationship corpus 574 and opens the topic word list 74. The learning data creation unit 576 also secures storage areas for the arrays T and S, the variables S3 and M, the repetition control variables i and j, and the vector Z.

[0088] Next, the learning data creation unit 576 reads the topic word list 74, separates it using a predetermined delimiter, and stores the contents in each element of the array T (step 152). The learning data creation unit 576 further T The maximum value of the subscript of the array T is substituted into (step 154). The learning data creation unit 576 then connects to the causal relation corpus 574 shown in FIG. 8 (step 622). In this embodiment as well, the subscript of the array T is changed from 0 to the variable MAX T The value is up to .

[0089] The training data creation unit 576 further performs the following step 626 for each causal relation stored in the causal relation corpus 574 to generate a record of the training data (step 624).

[0090] 10, in step 626, the learning data creation unit 576 generates a vector Z whose elements are all zero (step 250 in FIG. 10). That is, in this step, the vector Z is initialized. Next, the learning data creation unit 576 assigns the character string of the result phrase of the causal relationship to be processed to the character string variable S3 (step 650 in FIG. 10). Furthermore, the learning data creation unit 576 repeatedly executes step 256 while incrementing the repetition variable i by 1 from i=0 to N-1 (step 652 in FIG. 10).

[0091] In step 256, the learning data creation unit 576 determines whether the element T[i] of the array T to be processed is present in the character string represented by the character string variable S3 (step 300 in FIG. 10). When the determination in step 300 is positive, the learning data creation unit 576 performs the following: i (step 302). If the determination in step 300 is negative, learning data creation unit 576 does nothing.

[0092] The learning data creation unit 576 executes step 256 while incrementing the variable i by 1 from the repetition variable i=0 to N-1. If the element T[i] exists in the character string represented by the character string variable S3, this process i has a value of 1, otherwise element Z i The value of will be 0.

[0093] After completing step 254, the learning data creation unit 576 assigns the number of non-zero elements among the elements of vector Z to the variable M (step 258 in FIG. 10). The learning data creation unit 576 judges whether the value of variable M is 0 or not (step 260). If the judgment in step 260 is positive, that is, if there is no non-zero element among the elements of vector Z, the learning data creation unit 576 assigns the (N+1)th element Z of vector Z to N(step 262 in FIG. 10). If the determination in step 260 is negative, that is, if there is at least one non-zero element in vector Z, learning data creation unit 576 divides vector Z by the value of variable M (step 264 in FIG. 10). That is, each element of vector Z is divided by the value of variable M.

[0094] 10, the learning data creation unit 576 executes step 626, which is shown in FIG. 10, to obtain a vector Z in which, if any one of the words in the topic word list 74 is present in the result phrase of a causal relationship, the values ​​of the elements of vector Z corresponding to those words are 1 / M, and the values ​​of the other elements are 0. If no word in the topic word list 74 is present in the string represented by the string variable S3, the Nth element Z of vector Z is N will have a value of 1 and all other elements will have a value of 0.

[0095] Thereafter, the learning data creation unit 576 generates a new record of learning data corresponding to the causal relationship to be processed by combining the cause phrase of the causal relationship to be processed as input and the vector Z as output, and adds it to the learning data storage unit 578 shown in FIG. 8 (step 654).

[0096] The dialogue device 562 uses the training data thus created to train the context model 580. The processing by the training unit 78 is the same as that by the training unit 78 shown in FIG. 1, except for the training data used.

[0097] B2. Dialogue Phase The dialogue processing by the dialogue device 562 in the second embodiment is no different from that of the filtering unit 86 in the first embodiment, except that the dialogue device 562 uses a context model 580 learned by the method described above instead of the context model 80 used in the first embodiment.

[0098] Thus, according to the second embodiment, a large number of causal relationships are prepared in advance, and the result phrase of each causal relationship is regarded as the context of the cause phrase, and training data is prepared in the same manner as in the first embodiment. By using this training data to train the context model 580, as in the first embodiment, the validity of the system utterance is determined by considering not only the text of the system utterance candidate itself, but also the possibility of words appearing in the context. A system utterance in a dialogue is usually one sentence, and there is actually no context before or after it. Therefore, it is difficult to determine whether or not the utterance is an utterance that may cause a problem from the system utterance alone. However, according to this embodiment, the system utterance is selected using information on the relationship that the system utterance may have with the context before and after it, so that the probability of some problem occurring by outputting the system utterance can be reduced.

[0099] 3. Third embodiment A.Configuration In the above first and second embodiments, when a system utterance candidate is input, basically only the output of the context model for the system utterance candidate is used to determine whether to discard or retain the system utterance candidate. However, the present invention is not limited to such embodiments. In the third embodiment, the similarity between the vector output by the context model for the system utterance candidate and a plurality of control vectors prepared in advance is examined, and the system utterance candidate is discarded when the similarity satisfies a certain condition.

[0100] 11 shows a block diagram of a dialogue system 700 according to the third embodiment of the present invention. Referring to FIG. 11, the dialogue system 700 includes a dialogue engine 84 and a context model 80 similar to those used in the first embodiment, and a filtering unit 712 that checks the cosine similarity between an output probability vector output by the context model 80 for a system utterance candidate output by the dialogue engine 84 and a plurality of control vectors prepared in advance, and keeps the system utterance candidate if the number of control vectors with cosine similarity equal to or greater than a predetermined threshold is less than the threshold, and discards the system utterance candidate if not, and outputs a system utterance 714 based on a final scoring. It is assumed that the context model 80 has been trained according to the method described in the description of the first embodiment.

[0101] The dialogue system 700 further includes a filtering vector generating unit 710 that generates and stores in advance a contrast vector that is used by the filtering unit 712 for filtering.

[0102] More specifically, the filtering vector generation unit 710 includes a filtering expression storage unit 720 for storing a plurality of expressions that are considered to be likely to have undesirable expressions around them, a contrast vector generation unit 722 for generating a contrast vector consisting of an output probability vector of the context model 80 for each expression by inputting each expression stored in the filtering expression storage unit 720 to the context model 80, and a contrast vector storage unit 724 for storing the contrast vector generated by the contrast vector generation unit 722. The contrast vector storage unit 724 is connected to the filtering unit 712 so as to be accessible from the filtering unit 712.

[0103] This embodiment is based on the discovery that when there is a high similarity between an output probability vector obtained from an expression that has a high probability of having undesirable expressions appearing around it and an output probability vector obtained from a system utterance candidate, there is a high probability that undesirable expressions will appear around that system utterance candidate. In other words, the idea that it is undesirable to use such a system utterance candidate as the output of a dialogue system would not have been possible without such a discovery.

[0104] Figure 12 is a flowchart showing a control structure of a computer program for implementing filtering unit 712 shown in Figure 11. Referring to Figure 12, this program includes steps 450 and 452 similar to those shown in Figure 6, and a step 800 of executing step 802 for each system utterance candidate.

[0105] Step 802 includes steps 480 and 482 similar to those shown in Fig. 6, and step 820 following step 482, in which 0 is substituted for a variable representing a counter. This counter is used in the following processing to count the number of filtering expressions whose similarity to the probability vector obtained from the system utterance candidates is equal to or greater than a threshold value.

[0106] Step 802 further includes step 824 of incrementing a counter by 1 for each control vector if it is similar to the probability vector obtained from the system utterance candidate, step 826 of branching the control flow according to whether the counter value is less than a second threshold value after the process of step 822 is completed, and step 830 of retaining the system utterance candidate as the target when the determination in step 826 is positive, and discarding the system utterance candidate when the determination in step 826 is negative. Step 802 ends after steps 828 and 830.

[0107] Step 824 includes step 840 of calculating a cosine similarity between the target vector and a probability vector obtained from a system utterance candidate, step 842 of branching the control flow according to whether the cosine similarity calculated in step 840 is equal to or greater than a first threshold, and step 844 of incrementing a counter value by 1 and terminating execution of step 824 when the determination in step 842 is negative. When the determination in step 842 is negative, execution of step 824 is terminated without incrementing the counter.

[0108] It is desirable to determine the value of the first threshold value through experiments. The second threshold value may be any value greater than or equal to 1, but it is considered desirable to typically set the second threshold value to 1. However, since the value of the second threshold value also depends on the type of expression used for filtering, it is considered desirable to determine the value of the second threshold value through experiments.

[0109] B. Operation The dialogue system 700 according to the third embodiment has three operation phases. The first is a learning phase of the dialogue system 700. The second is a generation phase of a comparison vector. The third is a dialogue phase in which the filtering unit 712 is used. Of these, the learning phase is as described in relation to the first embodiment. Therefore, the generation phase of a comparison vector and the dialogue phase will be described in this order.

[0110] B1. Control vector generation phase 11, expressions that are highly likely to have undesirable expressions around them are collected in advance as filtering expressions and stored in a filtering expression storage unit 720. A contrast vector generation unit 722 provides each of these filtering expressions to the context model 80, obtains a probability vector output by the context model 80 in response thereto, and stores the probability vector as a contrast vector in a contrast vector storage unit 724. In this way, when contrast vectors are generated for all filtering expressions stored in the filtering expression storage unit 720 and stored in the contrast vector storage unit 724, the contrast vector generation phase is completed.

[0111] Of course, in this embodiment, a control vector may be generated from a filtering expression newly found after the filtering unit 712 is operated and added to the control vector storage unit 724 .

[0112] B2. Dialogue Phase The dialogue engine 84 generates a plurality of system utterance candidates for the input utterance 82 (step 450 in FIG. 12) and provides the candidates as a system utterance candidate list to the filtering unit 712 (step 452).

[0113] The filtering unit 712 performs the following process (step 802) for each of these system utterance candidates (step 800). The filtering unit 712 first inputs each system utterance candidate into the context model 80 (step 480) to obtain its output probability vector (step 482). The filtering unit 712 assigns 0 to a variable representing a counter (step 820), and performs the process shown in step 824 for each contrast vector (step 824).

[0114] In step 824, the filtering unit 712 calculates the cosine similarity between the system utterance candidate being processed and the comparison vector being processed (step 840), and determines whether the value is equal to or greater than a first threshold (step 842). If the cosine similarity is equal to or greater than the first threshold, the counter is incremented by 1 in step 844, and processing of the next comparison vector proceeds. If the cosine similarity is less than the first threshold, nothing is done, and processing of the next comparison vector proceeds.

[0115] When the process of step 824 is completed for all the comparison vectors in this manner, the counter stores the number of comparison vectors whose cosine similarity with the system utterance candidate being processed is equal to or greater than the first threshold value.

[0116] The filtering unit 712 further determines whether the counter value is less than a second threshold value (step 826). If the counter value is less than the second threshold value, the filtering unit 712 leaves the system utterance candidate being processed (step 828) and starts processing the next system utterance candidate. If the counter value is equal to or greater than the second threshold value, the filtering unit 712 discards the system utterance candidate being processed (step 830) and starts processing the next system utterance candidate.

[0117] In this way, after determining whether to discard or retain all system utterance candidates, the filtering unit 712 performs a re-ranking process on the remaining system utterance candidates and outputs the system utterance candidate with the highest score as the system utterance 714 (FIG. 11).

[0118] As described above, in the dialogue system 700 according to this embodiment, instead of using only the value of the probability vector output by the context model 80, the similarity between each of a plurality of prepared contrast vectors and a system utterance candidate is calculated. If the number of contrast vectors with high calculated similarity is equal to or greater than a predetermined number (second threshold value), the system utterance candidate is discarded, and other system utterance candidates are retained. The second threshold value may be a number equal to or greater than 1, and may be simply set to 1.

[0119] As described above, in the third embodiment, the same context model as in the first and second embodiments is used, but the filtering method used is different from that in the first and second embodiments. The third embodiment can also obtain the same effects as the first and second embodiments.

[0120] In the third embodiment, vector similarity is used to compare the control vector with the system utterance candidate. However, the present invention is not limited to such an embodiment. Any value that is a measure of the similarity between two vectors may be used. For example, after normalizing the two vectors, the two vectors may be regarded as position vectors, and the distance between the tips of the two vectors may be used as the measure of similarity. Alternatively, the sum of squared errors between corresponding elements after normalizing the vectors may be used as the measure of similarity.

[0121] 4. Computer implementation Fig. 13 is an external view of an example of a computer system for implementing each of the above-described embodiments, and Fig. 14 is a block diagram showing an example of a hardware configuration of the computer system shown in Fig. 13.

[0122] 13, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972 for interacting with a user, all of which are connected to the computer 970. Of course, these are just one example of a configuration for when user interaction is required, and any general hardware and software available for user interaction (e.g., a touch panel, voice input, or a general pointing device) can be used.

[0123] 14, the computer 970 includes, in addition to the DVD drive 1002, a CPU (Central Processing Unit) 990, a GPU (Graphics Processing Unit) 992, a bus 1010 connected to the CPU 990, the GPU 992, and the DVD drive 1002, a ROM (Read-Only Memory) 996 connected to the bus 1010 and storing a boot-up program of the computer 970, a RAM (Random Access Memory) 998 connected to the bus 1010 and storing instructions constituting a program, a system program, working data, and the like, and a SSD (Solid State Drive) 1000 which is a non-volatile memory connected to the bus 1010. The SSD 1000 is for storing programs executed by the CPU 990 and the GPU 992, and data used by the programs executed by the CPU 990 and the GPU 992, and the like. The computer 970 further includes a network I / F (Interface) 1008 that provides a connection to a network 986 enabling communication with other terminals, and a USB port 1006 to which a USB (Universal Serial Bus) memory 984 can be attached / detached and that provides communication between the USB memory 984 and each component within the computer 970.

[0124] The computer 970 further includes an audio I / F 1004 that is connected to a microphone 982, a speaker 980, and a bus 1010, and that reads out audio signals, video signals, and text data generated by the CPU 990 and stored in the RAM 998 or SSD 1000 in accordance with instructions from the CPU 990, performs analog conversion and amplification processing to drive the speaker 980, and digitizes the analog audio signal from the microphone 982 and stores it in any address of the RAM 998 or SSD 1000 specified by the CPU 990.

[0125] In the above embodiment, the programs for implementing the various parts of the dialogue system 50 shown in Fig. 1 and the dialogue system 550 shown in Fig. 8, the neural network parameters, the neural network programs, etc. are all stored in, for example, the SSD 1000, RAM 998, DVD 978, or USB memory 984 shown in Fig. 14, or in a storage medium of an external device (not shown) connected via the network I / F 1008 and the network 986. Typically, these data, parameters, etc. are written into the SSD 1000 from the outside, for example, and loaded into the RAM 998 when executed by the computer 970.

[0126] 1 and 8, respectively, and computer programs for operating the computer system to realize the functions of the dialogue systems 50 and 550 and their respective components are stored in a DVD 978 inserted in a DVD drive 1002, and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs are stored in a USB memory 984, and the USB memory 984 is inserted in a USB port 1006 and the programs are transferred to the SSD 1000. Alternatively, the programs may be transmitted to the computer 970 via a network 986 and stored in the SSD 1000.

[0127] Of course, a source program may be input using the keyboard 974, monitor 972, and mouse 976, and the compiled object program may be stored in the SSD 1000. In the case of a script language, a script input using the keyboard 974 or the like may be stored in the SSD 1000. In the case of a program that runs on a virtual machine, a program that functions as a virtual machine must be installed in the computer 970 in advance. Since training and testing of a neural network involves a large amount of calculation, it is preferable to realize each part of the embodiment of the present invention as an object program consisting of native computer code, rather than a script language, particularly for the program portion that is the entity that performs numerical calculations.

[0128] The program is loaded into the RAM 998 when it is executed. The CPU 990 reads the program from the RAM 998 according to an address indicated by a register (not shown) called a program counter inside the CPU 990, interprets the instruction, reads data required for executing the instruction from the RAM 998, the SSD 1000, or other devices according to an address specified by the instruction, and executes the process specified by the instruction. The CPU 990 stores the execution result data in an address specified by the program, such as the RAM 998, the SSD 1000, or a register in the CPU 990. At this time, the value of the program counter is also updated by the program. The computer program may be directly loaded into the RAM 998 from the DVD 978, the USB memory 984, or via a network. Note that some tasks (mainly numerical calculations) among the programs executed by the CPU 990 are dispatched to the GPU 992 according to instructions included in the program or according to the analysis results when the CPU 990 executes the instructions.

[0129] The program for implementing the functions of each unit according to the above-described embodiment in cooperation with the computer 970 includes a plurality of instructions written and arranged to operate the computer 970 to implement those functions. Some of the basic functions required to execute those instructions are provided by an operating system (OS) or a third-party program running on the computer 970, or by modules of various tool kits installed on the computer 970. Thus, the program does not necessarily include all of the functions required to implement the system and method of the embodiment. The program may include only instructions to perform the operations of each of the above-described devices and their components by statically linking appropriate functions or functions of a "programming tool kit" in a controlled manner to obtain the desired results, or by dynamically linking those functions when the program is executed. The method of operating the computer 970 for that purpose is well known, and will not be repeated here.

[0130] The GPU 992 is capable of parallel processing, and can simultaneously execute a large amount of calculations associated with machine learning in parallel or in a pipelined manner. For example, parallel calculation elements found in a program when the program is compiled, or parallel calculation elements found when the program is executed, are dispatched from the CPU 990 to the GPU 992 as needed, and executed, and the results are returned to the CPU 990 directly or via a predetermined address in the RAM 998, and assigned to a predetermined variable in the program.

[0131] 4. Variations In the above embodiment, the topic word list 74 is a list of words whose frequency of occurrence in a group of passages or the like is higher than a threshold value. However, the present invention is not limited to such an embodiment. For example, a predetermined number of words whose frequency of occurrence in a group of passages or the like is high may be listed. Instead of such a method, the topic word list 74 may be created by extracting words contained in expressions to be noted that have been collected manually in advance. Alternatively, the topic word list 74 may be a union or intersection of words whose frequency of occurrence in a group of passages or the like is higher than a threshold value, or a predetermined number of words whose frequency of occurrence is high, and a list of words to be noted that has been manually created in advance.

[0132] Furthermore, in the above embodiment, there is no particular restriction on the types of words, such as the parts of speech. However, the present invention is not limited to such an embodiment. Words may be restricted by specific parts of speech (e.g., verbs, adjectives, and nouns), or may be restricted to so-called content words only. Also, so-called phrases may be added to the topic word list 74, in addition to words.

[0133] In the above embodiment, BERT is used as the context model, but the present invention is not limited to such an embodiment, and a model based on an architecture other than BERT may be used as the context model.

[0134] The above embodiment relates to a dialogue system. However, the present invention is not limited to such an embodiment. The present invention can be applied to any system that performs dialogue-based communication between a person and some other system, such as a question-answering system, a dialogue-based task-oriented system, or a system that responds to a message from a user.

[0135] In the first embodiment, the passages used to create the training data are not particularly limited. However, as in the second embodiment, good results are obtained by creating the training data from causal relationships. Therefore, in the first embodiment, the training data may be created using passages that include specific expressions such as causal relationships.

[0136] Also, the second embodiment uses a causal relationship. A causal relationship is a combination of a cause phrase and a result phrase. When a result phrase of a causal relationship is similar to a cause phrase of another causal relationship, the two causal relationships can be linked together. Such a chain of causal relationships results in two result phrases from the cause phrase of the first causal relationship. Similarly, three or more result phrases can be associated with the first cause phrase. Using such a relationship, learning data can be created using not only one result phrase but two or more chained result phrases as the context in the second embodiment.

[0137] The embodiments disclosed herein are merely illustrative, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims of the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the scope and meaning equivalent to the words described therein. [Explanation of symbols]

[0138] 50, 550, 700 Dialogue System 60, 560 Context Model Learning System 62, 562 Interactive device 70 Passage DB 72,576 Learning Data Creation Department 74 Topic Word List 76,578 Learning data storage unit 78 Learning Department 80, 580 Context Model 82 Input utterance 84 Dialogue Engine 86, 582, 712 Filtering section 88, 584, 714 System Speech 340 CLS Tokens 342 SEP tokens 350 utterances 352 BERT 354 Final hidden layer 356 CLS compatible layer 358 fully connected layer 360 SoftMax layer 362 Probability Vector 400 training data 402 Training utterances 404 Correct label vector 570 Corpus 572 Causal Relation Extraction Unit 574 Causal Corpus 710 Filtering Vector Generation Unit 722 Control Vector Generator

Claims

1. a pre-trained context model that, when a word vector sequence representing an utterance is input, outputs a probability vector whose elements are the probability that each word included in a predetermined word group appears in the context in which the utterance is placed; and a determination means for inputting a word vector sequence representing an utterance into the context model, and determining whether the utterance should be discarded or accepted according to whether at least one element of the probability vector output by the context model in response to the input satisfies a predetermined condition.

2. 2. The speech filtering device according to claim 1, wherein the determining means includes means for determining whether the utterance should be discarded or accepted according to whether a value determined as a predetermined function of at least one element of the probability vector is equal to or greater than a predetermined threshold.

3. A dialogue device; 2. The speech filtering device of claim 1, coupled to the dialogue device to receive as input candidate utterances output by the dialogue device; and a speech filtering means for filtering the speech output by the dialogue device in accordance with a result of the determination by the speech filtering device.

4. Computer, a pre-trained context model that, when a word vector sequence representing an utterance is input, outputs a probability vector whose elements are the probability that each word included in a predetermined word group appears in the context in which the utterance is placed; A computer program that functions as a determination means for inputting a word vector sequence representing an utterance into the context model, and determining whether the utterance should be discarded or accepted depending on whether the probability of any word included in a predetermined word group is equal to or greater than a threshold based on the probability vector output by the context model in response to the input.

5. a context extraction means for extracting, for each utterance stored in the corpus, a context of the utterance; a context vector generating means for generating a context vector indicating whether each of the words included in a predetermined word group appears in at least the context; and a training data generation means for generating training data for each utterance stored in a corpus by combining the utterance as an input and the context vector as an output.

6. the corpus includes a plurality of causal relationship expressions, each of which includes a cause part and an effect part; 6. The training data generating device according to claim 5, wherein the context extraction means includes a result part extraction means for extracting, for each of the plurality of causal relationship expressions, the cause part of the causal relationship expression as the utterance and the result part of the causal relationship expression as the context of the utterance.