Utterance dictionary learning device, utterance sentence generation device, and utterance dictionary learning program

By employing a dynamic speech dictionary learning and generation system, communication robots can expand their vocabulary and generate diverse utterances, addressing the limitations of pre-registered speech and enhancing user interaction.

JP7682678B2Active Publication Date: 2025-05-26NIPPON HOSO KYOKAI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021064728
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-06
Publication Date
2025-05-26
Estimated Expiration
2041-04-06

AI Technical Summary

Technical Problem

Communication robots are limited to speaking only pre-registered words or sentences, leading to user boredom and short usage periods, as existing methods struggle to generate appropriate utterances using unknown words.

Method used

A speech dictionary learning device and a speech sentence generation device that learn and update dictionaries dynamically, allowing communication robots to gradually increase their vocabulary by storing input sentence histories, learning sentence models, and registering unknown words with approximate word estimates.

Benefits of technology

Enables communication robots to generate a wider range of utterances, providing a sense of growth and personality variation, thereby extending user engagement and usage periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682678000001
    Figure 0007682678000001
  • Figure 0007682678000002
    Figure 0007682678000002
  • Figure 0007682678000003
    Figure 0007682678000003
Patent Text Reader

Abstract

To provide an utterance dictionary learning device and an utterance sentence generation device which can gradually increase the number of words which can be uttered by a communication robot.SOLUTION: An utterance dictionary learning device 100 includes: an inputted sentence record C which stores inputted sentences I which have been inputted in the past; a sentence model learning unit 5 which learns a feature vector of a word which appears in the sentences stored in the inputted sentence record C in a frequency equal to or higher than a certain frequency, via a sentence model D for estimating a word in a sentence, and registers a known word whose feature vector has been learned in a known word dictionary A; and an unknown word dictionary updating unit 3 which registers, in an unknown word dictionary B, an unknown word which appears in an inputted sentence I and also has not been registered in the known word dictionary A, and the sentence model learning unit 5 repeats learning of the sentence model D at a certain timing and deletes a known word which has been registered in the known word dictionary A from the unknown word dictionary B.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a speech sentence generation device for a communication robot to start speaking based on a program being watched.

Background Art

[0002] Conventionally, technologies related to communication robots that watch videos such as TV programs with people have been proposed. For example, in Patent Document 1, by using social media comments related to a video and generating a speech sentence according to the internal personality, emotional state, etc. of the robot and operating the robot, a technology has been proposed to realize an empathic action as if the robot is watching the video with the user. Also, in Patent Document 2, a robot control technology has been proposed in which, during TV watching, the robot responds to command dialogues such as channel switching from a person and whispers while facing the direction of the TV, so that the robot performs an operation as if it is autonomously watching the TV.

[0003] In such a communication robot, not only passive actions such as responding and empathizing to a person's call are required, but also active actions such as providing an opportunity to talk to a person are required. As one of the active actions, there is a function in which the robot automatically generates a speech sentence and talks to a person. As a method of generating a speech sentence based on a TV program being watched, for example, in Patent Document 3, a method has been proposed in which keywords are extracted from the subtitle information of the TV program being watched, and a speech sentence is generated based on these keywords. In this method, the speech sentence is generated by inserting the extracted keywords into a stored template sentence. However, since the extracted keywords are assumed to be those registered in a dictionary in advance, there is a problem that only sentences based on the determined keywords can be generated.

[0004] In addition, as a method for automatically generating text from images, many techniques called visual captioning have been developed. These are methods for automatically generating a caption from an input image by constructing a model that generates text from an image through deep learning using a large amount of learning data of captions paired with images. For example, in the method of Non-Patent Document 1, a template sentence is generated from an image, and the template sentence is filled with detected objects in the image to generate a sentence regarding the detected objects. However, in the visual captioning technology based on deep learning, there is a problem that the types (number of classes) of detected objects are determined during learning, and the types of keywords used in the uttered sentence are limited. In addition, to add a new class of detected objects, a large amount of learning data is required, so it cannot be easily increased. Furthermore, since the detected objects in the image are abstract class names such as "dog" and "cat", it is difficult to generate a sentence including various words such as proper nouns.

[0005] Therefore, in order to automatically generate an uttered sentence using words other than pre-determined words (words registered in a dictionary) (unknown words), for example, the method of Patent Document 4 has been proposed. In this method, in a sentence including an unknown word, an approximate word fitting the unknown word part is estimated from the words before and after the unknown word in a dictionary, a sentence is generated using this estimated word, and finally the estimated word is replaced with the unknown word to automatically generate a sentence including the unknown word. Although this method can generate a sentence using an unknown word, the estimation accuracy depends on the sentences used for learning.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Non-Patent Literature

[0007]

Non-Patent Literature 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] As a problem of communication robots, there are cases where users get bored with the operations of the robots and the usage period becomes short. One of the factors that cause people to get bored with robots is that the robot manufacturer can only perform pre-determined operations. For example, regarding the speech of the robot, as described above, it could only speak using pre-registered words or sentences. According to the method of Patent Document 4, although a sentence using unknown words can be generated, since the estimation accuracy of approximate words depends on the sentences used for learning, appropriate utterance sentences could not often be generated.

[0009] An object of the present invention is to provide a speech dictionary learning device and a speech sentence generation device capable of gradually increasing the words that a communication robot can speak.

Means for Solving the Problems

[0010] The utterance dictionary learning device according to the present invention stores the input sentence history storing the input sentences input in the past, and the feature vectors of the words that appear in the sentences stored in the input sentence history with a frequency of a predetermined value or more, and learns them by a sentence model for predicting the words in the sentence. A sentence model learning unit that registers the learned known words of the feature vectors in a known word dictionary, and an unknown word dictionary update unit that registers an unknown word that appears in the input sentence and is not registered in the known word dictionary in the unknown word dictionary. The sentence model learning unit repeats the learning of the sentence model at a predetermined timing, and deletes the known words registered in the known word dictionary from the unknown word dictionary.

[0011] The utterance dictionary learning device may include a relearning request unit that requests relearning from the sentence model learning unit based on the number of appearances of the unknown word.

[0012] The relearning request unit may request relearning from the sentence model learning unit when a predetermined number or more of the unknown words with a predetermined number of appearances or more are registered in the unknown word dictionary.

[0013] The utterance dictionary learning device includes an unknown word prediction unit that estimates an approximate word of the unknown word based on a prediction result based on the sentence model of the words that appear in the input sentence, and the unknown word dictionary update unit may register the unknown word and the approximate word in the unknown word dictionary in association with each other.

[0014] The utterance sentence generation device according to the present invention is a device that refers to the known word dictionary and the unknown word dictionary learned by the utterance dictionary learning device and generates an utterance sentence, and includes an utterance sentence generation unit that generates an utterance sentence by combining the known words included in the input sentence and registered in the known word dictionary or the unknown words registered in the unknown word dictionary in a template sentence.

[0015] The utterance sentence generation device according to the present invention is a device that generates an utterance sentence by referring to the known word dictionary and the unknown word dictionary learned by the utterance dictionary learning device, and includes a template creation unit that registers an input sentence including the known word registered in the known word dictionary as a template sentence, a template selection unit that selects the template sentence for the known word included in the input sentence or the unknown word registered in the unknown word dictionary, and an utterance sentence generation unit that generates an utterance sentence by combining the known word or the unknown word with the selected template sentence. The template selection unit selects a synonym whose feature vector is similar to the known word, and selects the template sentence corresponding to the synonym. For the unknown word, the template sentence corresponding to the approximate word associated with the unknown word in the unknown word dictionary is selected.

[0016] Among the plurality of words estimated by the unknown word prediction unit for the unknown word, the word with a high estimation frequency may be prioritized as the approximate word.

[0017] The utterance dictionary learning program according to the present invention is for causing a computer to function as the utterance dictionary learning device.

Effects of the Invention

[0018] According to the present invention, the words that a communication robot can speak can be gradually increased.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0020] Hereinafter, an example of an embodiment of the present invention will be described. The uttered sentence generation device in the present embodiment is incorporated into or connected to a device that speaks to humans such as a robot or an anthropomorphic agent, and stores a dictionary of known words, a dictionary of words that can be spoken, and a template sentence for generating an uttered sentence step by step from the input sentence such as the subtitle sentence of a TV program. Thereby, for example, a communication robot feels growth as the number of words and sentences that can be spoken increases while watching TV with the user, and realizes an utterance that makes the user feel a personality such that the uttered sentence varies depending on the user of the robot.

[0021] FIG. 1 is a block diagram showing the functional configuration of the uttered sentence generation device 200 in the present embodiment. The utterance sentence generation device 200 generates an utterance sentence U from the input sentence I. In the present embodiment, it will be described that the subtitle sentence of the TV program being watched is sequentially input to the utterance sentence generation device 200 as the input sentence I.

[0022] The utterance sentence generation device 200 includes an utterance dictionary learning device 100 inside, and generates the utterance sentence U using the words output from the utterance dictionary learning device 100. Here, the utterance sentence generation device 200 is an information processing device (computer) including a control unit, a storage unit, and various input / output devices, and the control unit executes the software stored in the storage unit to realize various functions described later. Note that the utterance dictionary learning device 100 may be configured as an information processing device having an interface with the utterance sentence generation device 200, but its functions and database may be integrated into the control unit and storage unit of the utterance sentence generation device 200.

[0023] The utterance dictionary learning device 100 includes, in the control unit, a known word extraction unit 1, an unknown word extraction unit 2, an unknown word dictionary update unit 3, a sentence storage unit 4, a sentence model learning unit 5, an unknown word prediction unit 6, and a re-learning request unit 7, and accumulates data of a known word dictionary A, an unknown word dictionary B, an input sentence history C, and a sentence model D in the storage unit. In addition to the utterance dictionary learning device 100, the utterance sentence generation device 200 includes a template creation unit 9, an utterance sentence generation unit 10, and a template selection unit 11 in the control unit, and accumulates data of a template sentence E in the storage unit.

[0024] The known word dictionary A is a database that stores words (known words) that the utterance dictionary learning device 100 has learned and can be used to speak in generating an utterance sentence. In the present embodiment, the words stored in the known word dictionary A are nouns, but the words are not limited to nouns and may include other parts of speech such as verbs, adjectives, and adverbs.

[0025] FIG. 2 is a diagram showing an example of words stored in the known word dictionary A in the present embodiment. In this example, 13 words, namely, "Akita dog", "Shiba dog", "Tokyo Station", "Kyoto Station", "Japanese buckwheat noodles", "ramen", "udon", "spaghetti", "calico cat", "tabby cat", "Setagaya", "UK", and "France", are registered as known words.

[0026] The unknown word dictionary B is a database that stores words (nouns) included in the input sentence I that are not included in the known word dictionary A, that is, words that are not learned but known (unknown words). In addition, for each registered unknown word in the unknown word dictionary B, an approximate word, which is the result of estimating a known word close to this unknown word from the known word dictionary A, is stored. The estimation of the approximate word is performed by the unknown word prediction unit 6 described later each time this unknown word appears in the input sentence I, and the estimated approximate word is appended and associated with the corresponding unknown word in the unknown word dictionary B each time it is estimated.

[0027] Figure 3 is a diagram showing an example of data stored in the unknown word dictionary B in this embodiment. In this example, 4 words, namely, "Napolitana", "Kai dog", "Hanamaki", and "Munchkin", are registered as unknown words. In addition, for "Napolitana", 3 words, namely, "udon", "UK", and "ramen", registered in the known word dictionary A (Figure 2), are estimated as approximate words. Similarly, for "Kai dog", "Akita dog" and "calico cat" are registered as approximate words, for "Hanamaki", "Setagaya" and "udon" are registered as approximate words, and for "Munchkin", "tabby cat" is registered as an approximate word.

[0028] Note that in the example of Figure 3, since the approximate words are appended each time they are estimated, for example, the approximate word "ramen" for the unknown word "Napolitana" is registered repeatedly. The recording format in the unknown word dictionary B is not limited to this. For example, the approximate words may be recorded together with the number of times.

[0029] The known word extraction unit 1 extracts known words included in the known word dictionary A from the input sentence I and provides them to the utterance sentence generation unit 10. Further, the known word extraction unit 1 also extracts words (nouns) not included in the known word dictionary A and provides them to the unknown word dictionary update unit 3. Furthermore, the known word extraction unit 1 provides the input sentence I to the template creation unit 9 together with the extracted known words in order to create a template based on the input sentence I.

[0030] FIG. 4 is a diagram showing the processing flow by the known word extraction unit 1 in the present embodiment. This process starts when an input sentence I such as a subtitle sentence is input.

[0031] In step S11, the known word extraction unit 1 divides the input sentence I into words by morphological analysis. For example, when the input sentence I is "A cute calico cat.", a result of word division is obtained together with the part-of-speech, such as "cute (adjective) / calico cat (noun) / desu (auxiliary verb) / ne (particle)".

[0032] In step S12, the known word extraction unit 1 determines whether there is a word whose part-of-speech is a noun and which is registered in the known word dictionary A in the divided sentence. If this determination is YES, the process proceeds to step S11, and if the determination is NO, the process proceeds to step S15.

[0033] In step S13, the known word extraction unit 1 notifies the detected word (known word) and the divided input sentence I to the template creation unit 9.

[0034] In step S14, the known word extraction unit 1 notifies the detected word (known word) to the utterance sentence generation unit 10.

[0035] In step S15, the known word extraction unit 1 determines whether there is a noun other than the word registered in the known word dictionary A in the input sentence I. If this determination is YES, the process proceeds to step S16, and if the determination is NO, the process ends.

[0036] In step S16, the unknown word extraction unit 1 notifies the detected word (unknown word), together with the word-segmented input sentence I, to the unknown word dictionary update unit 3. For example, when the input sentence I is "I want to eat a tart.", the input sentence I is word-segmented into "tart (noun) / を (particle) / eat (verb) / たい (auxiliary verb)", and "tart (noun)" is notified to the unknown word dictionary update unit 3 as an unknown word.

[0037] The unknown word extraction unit 2 extracts the unknown words registered in the unknown word dictionary B from the input sentence I and provides them to the utterance sentence generation unit 10 together with any of the associated approximate words. Here, the selection of the approximate word to be provided may be random. Since multiple estimated approximate words are registered in the unknown word dictionary B either redundantly or along with the number of times, these redundant approximate words have a higher probability of being selected than other approximate words. Alternatively, approximate words with a higher number of estimations may be set with a higher priority and selected with a probability according to the priority.

[0038] FIG. 5 is a diagram showing the processing flow by the unknown word extraction unit 2 in the present embodiment. This process starts when an input sentence I such as a subtitle sentence is input.

[0039] In step S21, the unknown word extraction unit 2 word-segments the input sentence I by morphological analysis.

[0040] In step S22, the unknown word extraction unit 2 determines whether there is an unknown word registered in the unknown word dictionary B in the word-segmented sentence. If this determination is YES, the process proceeds to step S23, and if the determination is NO, the process ends.

[0041] In step S23, the unknown word extraction unit 2 randomly selects one of the approximate words registered in the unknown word dictionary B and notifies the unknown word detected in step S22 and the selected approximate word to the utterance sentence generation unit 10.

[0042] For example, when the input sentence I is "The spaghetti in this store is extremely delicious.", the word "spaghetti" registered in the unknown word dictionary B (Figure 3) is detected as an unknown word. Also, in the unknown word dictionary B, as approximate words for "spaghetti", "udon", "ramen", and "England" are registered, and among them, "ramen" has been estimated as an approximate word twice. The unknown word extraction unit 2 randomly selects one of these approximate words. For example, if "ramen" is selected, the unknown word extraction unit 2 notifies the utterance sentence generation unit 10 of the approximate word "ramen" together with the unknown word "spaghetti".

[0043] The unknown word dictionary update unit 3 registers the word notified as an unknown word from the known word extraction unit 1 in the unknown word dictionary B, and additionally records the estimation result of the approximate word for this unknown word in the unknown word dictionary B.

[0044] Figure 6 is a diagram showing the processing flow by the unknown word dictionary update unit 3 in this embodiment. This processing starts when the input sentence I and the unknown word X are notified from the known word extraction unit 1.

[0045] In step S31, the unknown word dictionary update unit 3 determines whether the unknown word X notified from the known word extraction unit 1 is already registered in the unknown word dictionary B. If this determination is YES, the processing proceeds to step S33; if the determination is NO, the processing proceeds to step S32.

[0046] In step S32, the unknown word dictionary update unit 3 registers the unknown word X that is not registered in the unknown word dictionary B.

[0047] In step S33, the unknown word dictionary update unit 3 obtains the estimation result of the approximate word for the unknown word X from the unknown word prediction unit 6.

[0048] In step S34, the unknown word dictionary update unit 3 registers the obtained approximate word as the approximate word for the unknown word X in the unknown word dictionary B. For example, when "tart" (noun) / "wo" (particle) / "taberu" (verb) / "tai" (auxiliary verb) is notified as the input sentence I that has been word-segmented from the known word extraction unit 1 and "tart" is notified as an unknown word, the unknown word dictionary update unit 3 searches for "tart" among the words registered in the unknown word dictionary B (Fig. 3). Since it is not registered, it is newly registered as an unknown word. Also, when "udon" is estimated as an approximate word for "tart" from the input sentence I by the unknown word prediction unit 6, the unknown word dictionary update unit 3 adds "udon" as an approximate word for "tart" in the unknown word dictionary B.

[0049] The sentence storage unit 4 stores the input input sentence I as the input sentence history C in order to learn or relearn the sentence model D.

[0050] The sentence model learning unit 5 learns or relearns the sentence model D using the sentences input in the past stored in the input sentence history C at a predetermined timing such as an instruction input or a fixed period, or when there is a request from the relearning request unit 7, and also executes the update process of the known word dictionary A and the unknown word dictionary B.

[0051] In this embodiment, as an example of the learning method of the sentence model D, Word2vec shown in the following Non-Patent Document 2 is used. Non-Patent Document 2: T. Mikolov, K. Chen, G. Corrado, J. Dean: "Efficient Estimation of Word Representations in Vector Space," Proceedings of Workshop at ICLR, arXiv: 1301.3781 (2013).

[0052] Word2vec is a neural network that uses word-segmented sentences and learns so that the target word is output from the surrounding words of the target word. By using a large number of input sentences, the weights in the learned neural network can be used as the feature vectors of words, and it is empirically known that the feature vectors of words that are semantically approximate become close vector values.

[0053] FIG. 7 is a diagram illustrating a method of learning the sentence model D in the present embodiment. Here, an example of learning by Word2vec is shown in the case where the target word is "ramen" in the sentence "delicious / looks like / is / ramen / is / right" which is split into words.

[0054] This example shows the case where the number of vocabulary words learned by Word2vec is V and the dimension number of the feature vector of the word is N. By inputting the surrounding words excluding the target word and learning so that the word "ramen" is output, the N-dimensional feature vector in the network is learned. Here, W N×V and W' V×N represent the weights of the network, and the sum of the N-dimensional vectors W wi ·x obtained from the one-hot vector x N×V corresponding to each input word wi wi becomes the feature vector of the target word. The sentence model learning unit 5 performs learning by Word2vec using the input sentence history C, and stores W N×V and W' V×N as the sentence model D.

[0055] In order to learn the feature vector of a word by Word2vec, a certain amount of learning sentences are required for the word. Therefore, for words with a low appearance frequency, a highly accurate feature vector cannot be obtained, and the vocabulary number V of the words to be learned is determined after counting the appearance frequency in advance. In the present embodiment, among the V vocabularies used for learning, noun words are registered in the known word dictionary A.

[0056] FIG. 8 is a diagram showing the flow of processing by the sentence model learning unit 5 in the present embodiment. In step S51, the sentence model learning unit 5 counts the number of occurrences of each word in the sentences of the input sentence history C, and determines the words that appear K times or more as the words to be learned. Here, the words that appear K or more times may be the words that exist K or more times in the entire input sentence history C, but are not limited to this. For example, words that appear K or more times within a certain period may be selected.

[0057] In step S52, the sentence model learning unit 5 performs Word2vec learning with the number V of words determined in step S51 as the dimensions of the input and output, and the learning results W N×V and W’ V×N are stored as the sentence model D.

[0058] In step S53, the sentence model learning unit 5 registers the nouns among the vocabulary used for learning as known words in the known word dictionary A.

[0059] In step S54, the sentence model learning unit 5 deletes the words used for learning, that is, the words registered in the known word dictionary A, from among the words registered in the unknown word dictionary B. As a result, a part of the unknown words in the unknown word dictionary B is transferred to the known word dictionary A by learning.

[0060] The unknown word prediction unit 6 estimates one word from among the known words registered in the known word dictionary A as an approximate word for the specified unknown word X and the input sentence I in which the unknown word X is included. As shown in FIG. 7, since the sentence model D can predict the word with the highest appearance probability from the surrounding words, the unknown word prediction unit 6 predicts the word that fits the position of the unknown word X from the surrounding words other than the unknown word X in the input sentence I in which the unknown word X is included, and estimates an approximate word from among the known word dictionary A.

[0061] The relearning request unit 7 requests the sentence model learning unit 5 to relearn the sentence model D. As described above, a certain degree of appearance frequency (K or more) is required for the target word for learning. For this reason, the relearning request unit 7 monitors the number of estimation results of approximate words for each unknown word registered in the unknown word dictionary B, and requests relearning when the number of unknown words with K or more estimated words reaches a certain number L or more.

[0062] For example, when K = 50 and L = 100, if all the approximate words registered in association with 100 unknown words are in a state where there are 50 or more of them, the relearning request unit 7 will notify the sentence model learning unit 5 of the request for relearning. In this embodiment, although the relearning by the sentence model learning unit 5 is automatically requested from the relearning request unit 7, the user may request it at an arbitrary timing.

[0063] The template creation unit 9 creates a template sentence with the known word part filled in from the word-segmented sentence notified from the known word extraction unit 1 and the known words included in this sentence, and stores it in a database (template sentence E).

[0064] FIG. 9 is a diagram showing an example of the template sentence E in this embodiment. Here, the words (known words) enclosed in brackets [] indicate replaceable nouns. For example, in the template sentence created from the past input sentence I "Even when becoming an adult, Akita dogs are still popular, aren't they?" in the first line, it is shown that the replaceable words are [adult] [Akita dog] [popularity]. Similarly, in the second line, a template sentence in which [Shiba dog] [abroad] [popularity] can be replaced is created from the past input sentence I "Is a Shiba dog popular even abroad?"

[0065] In this way, the template creation unit 9 creates a template sentence from the word-segmented input sentence I and the known words notified from the known word extraction unit 1, and stores it in the template sentence E. For example, when the known word extraction unit 1 notifies the input sentence "cute (adjective) / calico cat (noun) / is (auxiliary verb) / right (particle)" and the known word "calico cat", the template creation unit 9 creates a template sentence "cute [calico cat] is right." and adds it to the template sentence E.

[0066] In the present embodiment, all the input sentences I notified from the known word extraction unit 1 are used as template sentences, but it is not limited to this. For example, only sentences with a predetermined number of characters or less may be used as template sentences. Further, only sentences containing emotions, such as verb phrases like "want to go" and "want to eat", or adjectives like "beautiful" and "big", may be stored as template sentences, so that the robot may be configured to make utterances containing emotions.

[0067] The utterance sentence generation unit 10 generates an utterance sentence U from the known word Z notified from the known word extraction unit 1, or the unknown word X and the approximate word Y notified from the unknown word extraction unit 2.

[0068] FIG. 10 is a diagram showing the processing flow by the utterance sentence generation unit 10 in the present embodiment. This processing is started when the known word Z is notified from the known word extraction unit 1, or when the unknown word X and the approximate word Y are notified from the unknown word extraction unit 2.

[0069] In step S101, the utterance sentence generation unit 10 determines whether the known word Z is notified or the unknown word X is notified. If the known word Z is notified, the process proceeds to step S102, and if the unknown word X is notified, the process proceeds to step S104.

[0070] In step S102, in order to generate an utterance sentence U using the known word Z notified from the known word extraction unit 1, the utterance sentence generation unit 10 acquires the template sentence selected by the template selection unit 11. The method of selecting the template sentence by the template selection unit 11 will be described later, but a template sentence including a synonym of the known word Z is selected.

[0071] In step S103, the utterance sentence generation unit 10 replaces the synonym part in the acquired template sentence with the known word Z to generate the utterance sentence U.

[0072] In step S104, in order to generate the utterance sentence U using the unknown word X notified by the unknown word extraction unit 2, the utterance sentence generation unit 10 acquires a template sentence for the approximate word Y selected by the template selection unit 11.

[0073] In step S105, the utterance sentence generation unit 10 generates the utterance sentence U by replacing the approximate word part in the acquired template sentence with the unknown word X.

[0074] In this way, the utterance sentence generation unit 10 generates the utterance sentence U using the template sentence for the synonym of the known word Z or the approximate word Y of the unknown word X. Note that since the accuracy (certainty) of the approximate word Y of the unknown word X is lower than that of the synonym of the known word Z, the utterance sentence using the unknown word X may be generated not from the template sentence for the approximate word Y, but using a pre-registered general template sentence such as "What is ~?", "Do you know ~?", etc.

[0075] The template selection unit 11 selects one from the template sentences stored in the template sentence E from the specified word. First, the template selection unit 11 obtains a word similar to the specified word (known word Z or unknown word X). Here, the word similar to the unknown word X is the approximate word Y selected by the unknown word extraction unit 2 corresponding to the unknown word X. The word similar to the known word Z is obtained as follows based on the similarity of the learned word feature vectors.

[0076] FIG. 11 is a diagram showing an image of the feature vectors of the words included in the known word dictionary A learned as the sentence model D in the present embodiment. As shown in FIG. 11, it is empirically known that in the learning result by Word2vec, the feature vectors of words with similar meanings are distributed at positions with close distances.

[0077] In FIG. 11, it shows that the feature vectors of "ramen", "udon", "Japanese soba", "spaghetti" stored in the known word dictionary A, the feature vectors of "Akita dog", "Shiba Inu", "calico cat", "tiger cat", the feature vectors of "Tokyo Station", "Kyoto Station", and the feature vectors of "Setagaya", "UK", "France" are distributed close to each other respectively. On the other hand, the distance between the feature vectors of "udon" and "tiger cat" is far. In this case, for example, template sentences containing words (synonyms) such as "calico cat", "Akita dog", "Shiba Inu" whose feature vectors are similar to that of "tiger cat" are likely to be usable even if these words are replaced with "tiger cat". Therefore, the template selection unit 11 selects a template sentence containing any synonym (for example, "calico cat") for the known word "tiger cat".

[0078] FIG. 12 is a diagram showing the flow of processing by the template selection unit 11 in this embodiment. This processing starts when the known word Z is input.

[0079] In step S111, the template selection unit 11 extracts, from the vocabulary of the sentence model D, other words whose feature vectors are similar to the known word Z, for example, words with the top predetermined number (for example, 5) of similarity degrees, as candidates for synonyms. For the similarity degree of the feature vectors, for example, the cosine similarity (cos(a, b) = a·b / |a||b|) between two vectors a and b can be used. Although it is assumed that words with higher similarity degrees are extracted as synonyms, in addition to this, words with a similarity degree of a predetermined value or more may also be extracted. Also, the similarity degree is not limited to the cosine similarity.

[0080] In step S112, the template selection unit 11 randomly selects one of the extracted words as a synonym.

[0081] In step S113, the template selection unit 11 randomly selects one template sentence containing the selected synonym from among the template sentences stored in the template sentence E, and outputs the selected template sentence and the selected synonym.

[0082] According to the present embodiment, the utterance dictionary learning device 100 includes a known word dictionary A that stores words that can be used as words for a robot to speak by learning feature vectors, and an unknown word dictionary B that stores words that are not included in the known word dictionary A but exist in past input sentences and have been heard. Using the input sentence I input daily, the words included in these dictionaries are updated step by step. Thereby, the utterance dictionary learning device 100 can gradually increase the words that the communication robot can speak, and can give a "sense of growth" to the communication robot that performs utterances using these words.

[0083] The utterance dictionary learning device 100 can update the sentence model D, the known word dictionary A, and the unknown word dictionary B at an appropriate timing by performing re-learning based on the number of appearances of unknown words. Specifically, when a predetermined number or more of unknown words with a predetermined number or more of appearances are registered in the unknown word dictionary B, re-learning is performed, so that the frequency of re-learning of the known word dictionary A for moving the words in the unknown word dictionary B, which usually takes time, to the known word dictionary A can be reduced.

[0084] The utterance dictionary learning device 100 estimates approximate words included in the known word dictionary A for unknown words based on the prediction results based on the sentence model D and registers them in the unknown word dictionary B. Thereby, the utterance dictionary learning device 100 can provide the unknown word dictionary B so that even before re-learning, an utterance sentence using an unknown word can be generated using the approximate word as a clue.

[0085] The utterance sentence generation device 200 can grow the utterance function of the communication robot using an increasing vocabulary by referring to these known word dictionary A and unknown word dictionary B.

[0086] In addition, the utterance sentence generation device 200 registers the input sentence I including known words as a template sentence, so that different template sentences are accumulated based on different input sentences I for each user, such as the TV programs watched. As a result, for each communication robot, utterances with "personality" that change with "growth" become possible.

[0087] The utterance sentence generation device 200 preferentially selects, as approximate words, words with a high estimated frequency by the sentence model D for unknown words before they are registered in the known word dictionary A. Thereby, the utterance sentence generation device 200 can select a template sentence using highly reliable approximate words and generate an utterance sentence even for unknown words for which the learning of the feature vector has not been performed.

[0088] As described above, the embodiments of the present invention have been described, but the present invention is not limited to the above-described embodiments. In addition, the effects described in the above embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments.

[0089] In this embodiment, mainly the configurations and operations of the utterance dictionary learning device 100 and the utterance sentence generation device 200 have been described, but the present invention is not limited to this, and it may be configured as a method or a program for causing a communication robot to utter, including each component.

[0090] Furthermore, a program for realizing the functions of the utterance dictionary learning device 100 and the utterance sentence generation device 200 may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to be realized.

[0091] The "computer system" referred to here is assumed to include hardware such as an OS and peripheral devices. In addition, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built in a computer system.

[0092] Furthermore, the "computer-readable recording medium" may include those that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and those that hold a program for a certain period of time, such as volatile memory inside a computer system serving as a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and furthermore, it may be capable of realizing the aforementioned functions in combination with a program already recorded in a computer system.

Explanation of Signs

[0093] A Known Word Dictionary B Unknown Word Dictionary C Input Sentence History D Sentence Model E Template Sentence 1 Known Word Extraction Unit 2 Unknown Word Extraction Unit 3 Unknown Word Dictionary Update Unit 4 Sentence Storage Unit 5 Sentence Model Learning Unit 6 Unknown Word Prediction Unit 7 Relearning Request Unit 9 Template Creation Unit 10 Spoken Sentence Generation Unit 11 Template Selection Unit 100 Spoken Word Dictionary Learning Device 200 Spoken Sentence Generation Device

Claims

1. An input sentence history that stores input sentences input in the past, and A sentence model learning unit that learns the feature vectors of words that appear in the sentences stored in the input sentence history at a frequency equal to or higher than a predetermined value using a sentence model for predicting words in a sentence, and registers the learned known words with the feature vectors in a known word dictionary; An unknown word dictionary update unit that registers an unknown word that appears in the input sentence and is not registered in the known word dictionary in the unknown word dictionary; An unknown word prediction unit that estimates an approximate word for the unknown word based on a prediction result of the sentence model for words that appear in the input sentence, and The sentence model learning unit repeatedly learns the sentence model at a predetermined timing, and deletes the known words newly registered in the known word dictionary from among the words registered in the unknown word dictionary. The unknown word dictionary update unit can generate a spoken sentence by registering the unknown word and the approximate word in the unknown word dictionary in association with each other and replacing the approximate word with the unknown word using a template corresponding to the approximate word. A spoken dictionary learning device.

2. The spoken dictionary learning device according to claim 1, further comprising a re-learning request unit that requests re-learning to the sentence model learning unit when a predetermined number or more of unknown words for which the number of estimation results of the approximate words registered in the unknown word dictionary is equal to or greater than a predetermined value are registered.

3. A spoken sentence generation device that refers to the known word dictionary and the unknown word dictionary learned by the spoken dictionary learning device according to claim 1 or claim 2 and generates a spoken sentence, A spoken sentence generation unit that generates a spoken sentence by combining the known words registered in the known word dictionary or the unknown words registered in the unknown word dictionary included in the input sentence with a template sentence. A spoken sentence generation device.

4. A spoken sentence generation device that refers to the known word dictionary and the unknown word dictionary learned by the spoken dictionary learning device according to claim 1 or claim 2 and generates a spoken sentence, A template creation unit that registers an input sentence including the known words registered in the known word dictionary as a template sentence; A template selection unit that selects the template sentence for the known words registered in the known word dictionary or the unknown words registered in the unknown word dictionary included in the input sentence; A spoken sentence generation unit that generates a spoken sentence by combining the known word or the unknown word with the selected template sentence, and The template selection unit For the known word, select a synonym for which the feature vector is similar, and select the template sentence corresponding to the synonym. An utterance sentence generation device that selects a template sentence corresponding to the approximate word associated with the unknown word in the unknown word dictionary for the unknown word. **Claim 5** The utterance sentence generation device according to claim 4, wherein among the plurality of words estimated by the unknown word prediction unit for the unknown word, the word with a higher estimated frequency is prioritized as the approximate word. **Claim 6** An utterance dictionary learning program for causing a computer to function as the utterance dictionary learning device according to claim 1 or claim 2.

Citation Information

Patent Citations

  • Word vector construction method and device, medium and electronic equipment

    CN110321552A

  • Control circuit of ac motor

    JP1986022792A

  • Control device, control method, and control program

    JP2018180472A

  • Speech generation device, speech generation method and speech generation program

    JP2018190077A

  • Sentence generation device, sentence generation method, and sentence generation program

    JP2019185400A