Voice processing device, toy, and computer program
The voice processing device addresses processor limitations by using pre-stored voice information to facilitate conversations with reduced processing load and varied responses, enhancing user interaction.
Patent Information
- Application Number
- PCT/JP2025/022582
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-23
- Publication Date
- 2026-01-02
AI Technical Summary
Existing simulated conversation systems are limited by the number of processable input words due to processor performance constraints, necessitating a technology to facilitate conversations even when input word processing is restricted.
A voice processing device that includes voice acquisition, processing, output, and storage means, utilizing pre-stored voice information to determine and output predetermined input words, reducing the need for continuous speech recognition and minimizing processing load.
Enables pseudo-conversations by processing limited input words efficiently, reducing power consumption and maintaining user engagement through varied responses.
Smart Images

Figure JP2025022582_02012026_PF_FP_ABST
Abstract
Description
Speech processing device, toy, and computer program
[0001] The present invention relates to a voice processing device, a toy, and a computer program.
[0002] Patent Document 1 discloses an invention of a simulated conversation system, in which a question to start a conversation is output from a speaker 30, and when a voice in response to the output question is input, the voice (input word) is recognized and a registered word corresponding to the voice is determined. In accordance with the determined registered word, a backchannel and the next question are determined from a group of backchannels and a group of next questions, and the voices are synthesized and output from the speaker, continuing the conversation.
[0003] Japanese Patent Application Laid-Open No. 2002-169591
[0004] In simulated conversation systems such as those described above, the user is not particularly aware of an upper limit on the number of input words that can be spoken. However, if there are constraints on the processing of input words, such as when the processor used has low performance, the number of input words that can be processed is limited. Therefore, there is a need for technology that can establish a simulated conversation even when the number of input words that can be processed is limited.
[0005] Therefore, a technology is provided that makes it possible to establish a pseudo-conversation even when the input words that can be processed are limited.
[0006] One form for solving the above problem is a voice processing device comprising: a voice acquisition means for acquiring voice uttered by a user; a processing means for processing the acquired voice; a voice output means for outputting voice in accordance with the result of processing by the processing means; and a storage means for storing voice information of a plurality of predetermined input words and output voice information by the voice output means associated with each of the plurality of predetermined input words, wherein the processing means starts the processing in accordance with the acquisition of the voice by the voice acquisition means, and in the processing, makes a first determination to determine whether the voice acquired by the voice acquisition means corresponds to any of the voice information of the plurality of predetermined input words, and in accordance with determining in the first determination that the voice corresponds to first voice information among the voice information of the plurality of predetermined input words, outputs first output voice information stored in the storage means in association with the first voice information from the voice output means.
[0007] To enable a pseudo conversation to be established even when processable input words are limited.
[0008] A diagram showing an example of the configuration of a voice processing system 10 corresponding to an embodiment. A flowchart showing an example of processing corresponding to an embodiment. A flowchart showing an example of voice registration processing corresponding to an embodiment. A diagram showing an example of the data configuration of registered voice corresponding to an embodiment. A flowchart showing an example of first mode processing corresponding to an embodiment. A diagram for explaining first mode processing corresponding to an embodiment. A diagram showing an example of the configuration of a data table corresponding to the first mode corresponding to an embodiment. A flowchart showing an example of second mode processing corresponding to an embodiment. A diagram showing an example of the configuration of a data table corresponding to the second mode corresponding to an embodiment. A flowchart showing an example of third mode processing corresponding to an embodiment. A diagram showing an example of the configuration of a data table corresponding to the third mode corresponding to an embodiment. A diagram for explaining third mode processing corresponding to an embodiment.
[0009] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant explanations will be omitted.
[0010] First, the configuration of a voice processing system corresponding to this embodiment will be described. Fig. 1 is a diagram showing an example of the configuration of a voice processing system 10 corresponding to this embodiment. The voice processing system 10 is composed of a voice processing device 100 and a controller 110. The voice processing device 100 is composed of a processing unit 101, a microphone 102, a speaker 103, a notification unit 104, an infrared receiving unit 105, an operation unit 106, and a storage unit 107. The controller 110 transmits commands via infrared rays in response to operations by a user of the voice processing system 10.
[0011] The system configuration is not limited to that shown in Figure 1, and the audio processing device 100 may be further connected to an external device such as an external server or a cloud server via a wired or wireless connection, and may be configured to be able to obtain update data for updating output audio data, etc. from the external device.
[0012] The processing unit 101 of the audio processing device 100 controls the operation of the entire device. The processing unit 101 can be configured with a CPU or the like. The microphone 102 is an audio input device for acquiring audio utterances from the user. Audio information acquired by the microphone 102 is output to the processing unit 101 and processed. The speaker 103 is an audio output device that outputs audio such as messages, alarm sounds, and notification sounds to the user under the control of the processing unit 101. The notification unit 104 is a notification device that provides non-audio notifications to the user under the control of the processing unit 101, and can output, for example, light, video, vibration, etc. The notification unit 104 may perform notification by, for example, controlling the lighting of an LED, or by displaying a video on a display such as an LCD, or by using a vibrator to perform notification by vibration. These can also be used in any combination.
[0013] The infrared receiving unit 105 receives command information transmitted via infrared rays from the controller 110 and outputs it to the processing unit 101. The operation unit 106 is an operation mechanism for user operation, and includes, for example, a power button, operation buttons, switches, and the like.
[0014] The storage unit 107 is a storage device for managing and storing information and programs necessary for processing by the processing unit 101, and includes RAM, ROM, EEPROM, etc. The storage unit 107 stores a processing program for the processing unit 101 to execute processing corresponding to this embodiment. When the processing unit 101 is realized by a CPU, for example, it can control each unit of the voice processing device 100 using RAM as a work memory in accordance with a program stored in ROM. The storage unit 107 also stores user voice data (voice information, hereinafter the same) acquired by the microphone 102, output voice data (output voice information, hereinafter the same) output from the speaker 103, alarm sound data, notification sound data, etc. In this embodiment, the storage capacity of the storage unit 107 can be reduced by implementing the voice recognition process in a simplified manner.
[0015] The controller 110 is configured to have buttons, switches, a touch panel, etc., and when any button, etc. is operated, it transmits a corresponding command signal via infrared rays. The infrared signal is received and processed by the infrared receiving unit 105 of the audio processing device 100.
[0016] The voice processing device 100 may be configured to have, for example, the external shape of a toy. The toy may be, for example, a model having a doll-like (robot or human) appearance. The model may be, for example, a plastic model that has been assembled and painted. Alternatively, the model may be a completed model such as a figurine (action figure) with movable joints. The model may not only have a doll-like appearance, but may also have any shape, such as a general vehicle, racing vehicle, military vehicle, aircraft, ship, animal, or virtual life form.
[0017] <Overall Processing in Audio Processing Device 100> Next, an example of processing executed by the audio processing device 100 according to this embodiment will be described with reference to Fig. 2. At least a part of the processing corresponding to this flowchart is realized by the CPU of the processing unit 101 of the audio processing device 100 executing a program stored in the storage unit 107.
[0018] First, in S201, the processing unit 101 executes a voice registration process. This process involves having the user speak a predetermined input word (such as a line, a message, or a sentence) and registering the user's spoken voice as voice data (e.g., voice waveform data) in the storage unit 107. In the voice recognition process of this embodiment, the user's spoken voice is not converted into text, but is instead saved as voice data in the voice registration process. The saved voice data is then compared with the input voice data to determine which of the predetermined input words the user uttered. The predetermined input words are prepared in advance, and the user can register them by reading them out loud in order. The voice registration process can be executed in response to the operation of a button on the controller to instruct the start of the voice registration process, or can be executed when the voice processing system 10 is first started up. Details of S201 will be described with reference to FIG. 3A.
[0019] In the following S202, the infrared receiving unit 105 of the audio processing device 100 receives a command from the controller. The received command information is provided to the processing unit 101 for processing. Specifically, in S203, the processing unit 101 determines the type of the received command information and determines which of a plurality of modes has been instructed to execute. Here, three modes, a first mode, a second mode, and a third mode, will be described, but it is also possible to determine two or four or more modes.
[0020] If it is determined as a result of the determination in S203 that the first mode has been requested, the process proceeds to S204, where first mode processing is executed. Details of the first mode processing will be described with reference to FIG. 4A. Furthermore, if it is determined as a result of the determination in S203 that the second mode has been requested, the process proceeds to S205, where second mode processing is executed. Details of the second mode processing will be described with reference to FIG. 5. Furthermore, if it is determined as a result of the determination in S203 that the third mode has been requested, the process proceeds to S206, where third mode processing is executed. Details of the third mode processing will be described with reference to FIG. 6A. The processing in the audio processing device 100 is performed in the above manner.
[0021] <Voice Registration Processing> Next, the voice registration processing in S201 will be described in detail. Fig. 3A is a flowchart showing an example of the voice registration processing.
[0022] In the speech registration process, in S301, the processing unit 101 resets the value of a parameter n, which indicates the number of an input word to be registered, to 1. In the following S302, the processing unit 101 acquires, via the microphone 102, the speech of a user reading a predetermined input word n. In the following S303, the processing unit 101 divides the speech data acquired from the microphone 102 into syllables and stores them in the storage unit 107. For example, if the predetermined input is "chocolate," the processing unit 101 divides the speech data into "cho," "ko," "ree," and "to," and stores them in the storage unit 107. Note that the syllable division shown here is merely an example, and syllable division can be performed in any manner as long as it is possible to compare the registered speech data with newly input speech data and determine whether they match or mismatch.
[0023] In the next S304, the processing unit 101 determines whether the number of the registered input word has reached a predetermined number N. If N has not yet been reached ("NO" in S304), the processing proceeds to S305, where the input word number n is incremented by one, and the processing returns to S302 to continue the above-described processing. On the other hand, if N has been reached ("YES" in S304), the processing ends.
[0024] By the above voice registration process, it is possible to register the user's voice data for a predetermined number N of predetermined input words. By using the registered voice data in subsequent processing, when the user utters one of the predetermined input words, it becomes possible to reliably determine which word has been uttered.
[0025] An example of the input words registered at this time is shown in FIG. 3B . FIG. 3B shows an example in which 10 input words (e.g., input word 1 to input word 10) are divided into syllables and registered in the storage unit 107. In FIG. 3B , syllable information for the registered input words is shown row by row, and the information for each row is alternatively displayed as a combination of the input word number and the syllable numbers that make up the word. For example, 1_1 represents information about the first (first) syllable of input word 1, and 8_4 represents information about the fourth syllable of input word 8. Here, input word 1, registered at the top of the table in FIG. 3B , has the input sound divided into four syllables, 1_1 to 1_4. Specifically, if the word "partner" (aibou) is spoken and registered by a user, "partner" is pronounced as "aibou" (aibou) and divided into the syllables "a," "i," "bo," and "u." Therefore, the syllable information of "a" is registered as 1_1, "i" as 1_2, "bo" as 1_3, and "u" as 1_4. Similarly, if the phrase "Switch to manual" is pronounced as "Manyuarunikirikaero," it is divided into the syllables "ma," "nyu," "a," "ru," "ni," "ki," "ri," "ka," "e," and "ro" and registered. By using the voice data divided into these syllables, input voice data uttered by the user and acquired by microphone 102 is compared from the beginning with the registered voice data similarly divided into syllables and stored in storage unit 107, and it is possible to determine which input word matches.
[0026] <First Mode Processing> Next, the processing content of the first mode in which the above determination is made and audio output is performed will be described with reference to Figures 4A and 4B. Figure 4A is a flowchart corresponding to an example of the processing in the first mode, and Figure 4B is a diagram for explaining the round-robin determination processing corresponding to the first mode.
[0027] 4A, in S401, voice uttered by the user is acquired via the microphone 102. The voice data from the microphone 102 is transmitted to the processing unit 101 and processed therein. Specifically, in S402, it is determined whether the voice intensity is greater than a predetermined threshold Th1. The processing unit 101 determines whether the amplitude value of the voice data acquired from the microphone 102 is greater than the threshold Th1, and if it is greater than the threshold Th1 ("YES" in S402), the process proceeds to S403. On the other hand, if it is equal to or less than the threshold Th1 ("NO" in S402), the process returns to S401. In the process of S402, only voices generated at a volume equal to or greater than a certain volume are processed in order to prevent malfunction due to erroneous input.
[0028] Next, in S403, it is determined whether the length of the input voice is longer than a predetermined threshold value Th2. Th2 can be set to, for example, one second. The processing unit 101 determines whether the length of the voice data acquired from the microphone 102 is longer than the threshold value Th2, and if it is longer than the threshold value Th2 ("YES" in S403), the processing proceeds to S404. On the other hand, if it is equal to or less than the threshold value Th2 ("NO" in S403), the processing returns to S401. In order to prevent malfunction due to erroneous input, the processing in S403 also processes only voices generated with a certain length or longer.
[0029] In S403, if a certain amount of audio is acquired, the process proceeds to S404. Here, the length of the audio to be processed may be limited. For example, the audio data may be limited to a length not exceeding a threshold value Th3. Here, Th3 may be set to, for example, 3 seconds.
[0030] In the next step S404, the processing unit 101 performs a process of dividing the acquired voice data into syllables, and in step S405, the processing unit 101 performs a brute force test (first test) on the input voice data divided into syllables against the registered voices of the user registered in the storage unit 107. A specific example of this brute force test process will be described with reference to FIG. 4B. The process in step S404 can be performed when the user has finished speaking. The end of the user's speaking can be determined in the same manner as described later in relation to step S505 in FIG. 5A.
[0031] In Figure 4B, the input speech data is divided into syllables and numbered from 1 to 5, but there may be other syllables after these. The input speech data is compared starting with the first syllable numbered 1. The registered speech data to be compared is the same as that shown in Figure 3B. The registered speech data is compared sequentially with the registered speech data of the syllable at position 1 of the registered speech data, and if there is a match, a value of 1 is assigned, and if there is a mismatch, a value of 0 is assigned. This is performed for each syllable in the input speech data.
[0032] In the subsequent process of S406, the processing unit 101 selects the registered voice data with the highest judgment score in the judgment process of S405. Here, the registered voice data for which all syllables in the registered voice data are judged to be 1, or the registered voice data for which the number of syllables is the most and which contains 1 at a certain rate or more, can be judged to have the highest score. In this case, registered voice data whose number of syllables highly differs from that of the input voice data may be excluded in advance from the selection candidates. For example, if the number of syllables in the input voice data is four and the number of syllables in the registered voice data is five or more, at least one zero will be included, but data containing a certain number of zeros or more may be excluded.
[0033] In the following S407, the processing unit 101 selects the voice data of the first line stored in association with the input word corresponding to the registered voice selected in S406, and outputs the voice from the speaker 103. The first line stored in association with the input word can be identified by referring to a data table registered in the storage unit 107. FIG. 4C shows an example of the data structure of a data table for registering output voice data corresponding to this embodiment. In the data table 400 shown in FIG. 4C, voice data of a first line 402 and a second line 403 (first output voice information) are registered in association with a predetermined input word 401. The predetermined input word 401 corresponds to the registered voice data that was the target of determination in S406. In FIG. 4C, two types of lines are registered in association with each input word registered in the input word 401, but the number of registered lines may be three or more. In this case, the additional lines may be lines different from the first and second lines registered in association with each other. These voice data may be updated as appropriate by connecting to an external device.
[0034] In S407, when the output of the voice data of the first line is completed, the process proceeds to S408. In S408, voice uttered by the user is acquired via the microphone 102. The voice data from the microphone 102 is transmitted to the processing unit 101, and the processing unit 101 performs a second determination process. As the second determination, first in S409, the processing unit 101 determines whether the voice intensity is greater than a predetermined threshold Th1. If the amplitude value of the voice data acquired from the microphone 102 is greater than the threshold Th1 ("YES" in S409), the process proceeds to S410. On the other hand, if the amplitude value is equal to or less than the threshold Th1 ("NO" in S409), the process returns to S408. In the process of S409, only voices generated at a volume equal to or greater than a certain volume are processed in order to prevent malfunction due to erroneous input.
[0035] Next, in S410, it is determined whether the length of the input voice is longer than a predetermined threshold value Th2. Th2 can be set to, for example, one second. If the length of the voice data acquired from the microphone 102 is longer than the threshold value Th2 ("YES" in S410), the process proceeds to S411. On the other hand, if the length is equal to or less than the threshold value Th2 ("NO" in S410), the process returns to S408. In order to prevent malfunctions due to erroneous input, the process of S410 also processes only voices generated with a certain length or longer.
[0036] If it is determined in the second determinations of S409 and S410 that a voice input from the user has been made (or that a voice has been detected or acquired), in S411 the processing unit 101 selects the voice data of the second line registered in association with the input word selected in S406, outputs the voice data from the speaker 103, and ends this processing. Note that the processing of S411 can be performed when the user has finished speaking. The end of the user's speaking can be determined in the same manner as described below in relation to S505 in FIG. 5A.
[0037] The above describes the process of outputting up to the second line, but even after the audio data of the second line has been output, the same process as from S408 to S411 may be repeated to output audio data of additional lines registered in data table 400 in association with the input word selected in S406.
[0038] In the above example, the voice output is performed using the voice data of the lines registered in association with the predetermined words, but the voice output may be performed by selecting one of the lines not associated with the predetermined words. The selection method may be selection according to a predetermined order or random selection. In this case, the same line may be repeatedly output.
[0039] In the processing of the first mode described above, after the first determination determines which of the predetermined input words the voice initially uttered by the user corresponds to, the second determination can output the second line and subsequent lines without determining (or specifying) the content of the user's utterance in response to the acquisition (or detection) of the voice, thereby reducing the processing load and enabling a reduction in power consumption of the voice processing device 100. When the voice processing device 100 is battery-powered, reducing processing power is extremely important because it also has a significant impact on the battery life.
[0040] <Second Mode Processing> In the first mode described above, the voice processing device 100 responds to speech from the user, but in the second mode, the voice processing device 100 starts outputting voice and the user responds. The processing content of the second mode will be described below with reference to Figures 5A and 5B. Figure 5A is a flowchart corresponding to an example of processing in the second mode, and Figure 5B shows an example of the data configuration of a data table 500 corresponding to the second mode.
[0041] In S501 of Fig. 5A, the processing unit 101 selects one of the voice data of the third lines stored in the storage unit 107 and outputs the voice from the speaker 103. Here, Fig. 5B shows an example of the data configuration of a data table that registers voice data corresponding to the second mode. In the data table 500 of Fig. 5B, voice data of the third line 501 and voice data of the fourth line 502 (second output voice information) are registered in association with each other. In the data table 500, a plurality of combinations of the third line and the fourth line are registered, and in S501, one of these combinations may be selected randomly or in the order of registration. These voice data may be updated as needed by connecting to an external device.
[0042] After the audio output in S501 is completed, in the following S502, the processing unit 101 acquires the audio uttered by the user via the microphone 102. The audio data from the microphone 102 is transmitted to the processing unit 101 and processed by the processing unit 101. In the following S503, the processing unit 101 determines whether the audio intensity is greater than a predetermined threshold Th1. If the amplitude value of the audio data acquired from the microphone 102 is greater than the threshold Th1 ("YES" in S503), the processing proceeds to S504. On the other hand, if the amplitude value is equal to or less than the threshold Th1 ("NO" in S503), the processing returns to S502. In the processing of S503, only audio generated at a volume above a certain level is processed to prevent malfunction due to erroneous input.
[0043] Next, in S504, the processing unit 101 determines whether the length of the input voice is longer than a predetermined threshold value Th2. Th2 can be set to, for example, one second. If the length of the voice data acquired from the microphone 102 is longer than the threshold value Th2 ("YES" in S504), the process proceeds to S505. On the other hand, if the length is equal to or less than the threshold value Th2 ("NO" in S504), the process returns to S502. In order to prevent malfunctions due to erroneous input, the process of S504 also processes only voices generated with a certain length or more.
[0044] In the following S505, the processing unit 101 determines whether the user's speech has ended. For example, if the amplitude value of the voice data acquired from the microphone 102 reaches the noise level and it is determined that the speech has ended ("YES" in S505), the processing proceeds to S506. On the other hand, if the amplitude value is higher than the noise level and it is determined that speech is still occurring ("NO" in S505), the processing returns to S502. This determination is made in order to output the next line after the user's speech has ended. If the next line is output in the middle of the user's speech, the user may not be able to hear it or the conversation may sound unnatural, so the processing unit waits for the user's speech to end before proceeding to the next process.
[0045] In the following S506, the processing unit 101 selects the voice data of the fourth line registered in the data table 500 in association with the third line output in S501, outputs the voice data from the speaker 103, and ends this processing.
[0046] The above describes the processing up to the fourth line, but even after the audio data of the fourth line has been output, the same processing as that from S502 to S506 may be repeated to output the audio data of lines that are associated with the third line selected in S501 and registered in the data table 500. Also, the above describes the case where lines that are associated with each other and registered in the data table 500 are output in sequence, but lines that are not associated with each other may be randomly selected from a plurality of lines and output in sequence. In this case, the same line may be output repeatedly.
[0047] <Third Mode Processing> In the first and second modes described above, speech data registered in association with each other are output in sequence. In contrast, in the third mode, multiple pieces of speech data are registered in association with a specific input word, and one piece is selected and output each time, enabling different responses to the same speech input. The processing content of the third mode will be described below with reference to Figures 6A, 6B, and 6C. Figure 6A is a flowchart corresponding to an example of processing in the third mode, Figure 6B shows an example of the data configuration of a data table 600 corresponding to the third mode, and Figure 6C is a conceptual diagram of processing corresponding to the third mode.
[0048] 6A, the processing unit 101 acquires the voice uttered by the user via the microphone 102. The voice data from the microphone 102 is transmitted to the processing unit 101 and processed by the processing unit 101. In the following S602, the processing unit 101 determines whether the voice intensity is greater than a predetermined threshold Th1. If the amplitude value of the voice data acquired from the microphone 102 is greater than the threshold Th1 ("YES" in S602), the processing proceeds to S603. On the other hand, if the amplitude value is equal to or less than the threshold Th1 ("NO" in S602), the processing returns to S601. In the processing of S602, only voices generated at a volume equal to or greater than a certain volume are processed to prevent malfunction due to erroneous input.
[0049] Next, in S603, the processing unit 101 determines whether the length of the input voice is longer than a predetermined threshold value Th2. Th2 can be set to, for example, one second. If the length of the voice data acquired from the microphone 102 is longer than the threshold value Th2 ("YES" in S603), the process proceeds to S604. On the other hand, if the length is equal to or less than the threshold value Th2 ("NO" in S603), the process returns to S601. In order to prevent malfunctions due to erroneous input, the process of S603 also processes only voices generated with a certain length or more.
[0050] In the next step S604, the processing unit 101 determines whether the input speech corresponds to a predetermined input word or line. This determination can be made in the same manner as in the processes of steps S404 to S406 in Fig. 4. In the next step S605, the processing unit 101 selects a fifth line stored in the storage unit 107 in association with the predetermined line determined in step S604, outputs the selected line from the speaker 103, and ends the process. At this time, since multiple fifth lines are registered in association with the predetermined line, any one of them is selected.
[0051] 6B shows an example of the data configuration of a data table 600 for registering voice data for a fifth line. In the data table 600, an input word 601 and a fifth line 602 are registered. The input word 601 may be the same as the input word 401 in FIG. 4C , or an additional input word may be used with some commonality. Looking at input word 1 in FIG. 6B , for example, voice data for multiple lines is registered, such as a fifth line 11, a fifth line 12, a fifth line 13, and a fifth line 14. Therefore, when input word 1 is identified as the line uttered by the user in S604, multiple output candidates will exist.
[0052] More specifically, Figure 6C shows an example of a response to a voice input of "partner." "Partner" is a line registered as an input word 601 in Figure 6B, and "Yes, what's wrong?", "What's wrong?", "Where are you now?", etc. are registered as fifth lines 602. When it is determined in S604 that the input voice is "partner," in S605, the fifth lines registered as output voices are output in order from top to bottom. When the last registered line, "Am I that noisy?", is output, the process returns to the top and "Yes, what's wrong?" is selected. Note that the selection may be in the order of registration or random.
[0053] In this way, in the third mode, the voice data registered as the fifth line is output in sequence or randomly, so that a different response is returned each time the user tweets a specific line, thereby improving preference and interest.
[0054] In Fig. 6A, this process ends after the fifth line is output, but the embodiment is not limited to this, and the process from S408 onwards in Fig. 4A may be combined. This makes it possible to randomize the response to a predetermined input word and to continue the dialogue, thereby further improving the user's preference and interest.
[0055] As described above, the present embodiment provides a technology that enables a dialogue with a user while reducing power consumption by reducing the number of words that require speech recognition (identifying the content of a user's utterance based on a comparison with registered speech data) in the speech processing device 100, thereby reducing the processing load. A dialogue with a user may be performed by speech from the user or by speech from the speech processing device 100. In particular, when a dialogue is initiated by a speech from the speech processing device 100, speech recognition is not required, thereby enabling a dialogue with a particularly reduced processing load. Even when a dialogue is initiated by a user's speech, speech recognition is performed only on the initial utterance, and the dialogue can be continued thereafter without speech recognition, thereby reducing the processing load.
[0056] Summary of the Embodiments The above embodiments disclose at least the following voice processing device, toy, and computer program: (1) A voice processing device comprising: voice acquisition means for acquiring voice uttered by a user; processing means for processing the acquired voice; voice output means for outputting voice in accordance with a result of processing by the processing means; and storage means for storing voice information of a plurality of predetermined input words and output voice information by the voice output means associated with each of the plurality of predetermined input words, wherein the processing means starts the processing in accordance with the acquisition of the voice by the voice acquisition means, makes a first determination in the processing to determine whether the voice acquired by the voice acquisition means corresponds to any of the voice information of the plurality of predetermined input words, and outputs first output voice information stored in the storage means in association with the first voice information from the voice output means in accordance with the first determination that the voice corresponds to first voice information among the voice information of the plurality of predetermined input words. (2) The storage means stores a plurality of pieces of first output audio information in association with the first audio information, and the processing means outputs one of the plurality of pieces of first output audio information from the audio output means in response to determining in the first determination that the piece of first output audio information corresponds to the first audio information. (3) The first output audio information includes audio information of a plurality of lines, and when the processing means outputs the first output audio information from the audio output means, the processing means outputs audio information of a first line among the plurality of lines, and after the audio information of the first line is output, performs a second determination to determine whether audio has been acquired by the audio acquisition means, The voice processing device according to claim 1, wherein when it is determined in the second determination that the voice has been acquired, the voice output means outputs voice information of a second line of the plurality of lines that is different from the first line. (4) The voice processing device according to claim 3, wherein the voice information of the plurality of lines included in the first output voice information further includes voice information of at least one additional line that is different from the first line and the second line, and that is output after the voice information of the second line has been output.(5) The audio processing device according to (1), wherein the first output audio information includes audio information of a first line, and the processing means, when causing the audio output means to output the first output audio information, causes the audio information of the first line to be output, and after the audio information of the first line is output, makes a second determination to determine whether audio has been acquired by the audio acquisition means, and if it is determined in the second determination that the audio has been acquired, causes the audio output means to output audio information of any line from the audio information of a plurality of lines stored in the storage means. (6) The audio processing device according to any one of (3) to (5), wherein, in the second determination, the processing means determines that the audio has been acquired in response to the audio being detected by the audio acquisition means. (7) The audio processing device according to any one of (3) to (6), wherein, in the second determination, the processing means determines whether the audio has been acquired, but does not make a determination on the content of the acquired audio. (8) The speech processing device according to any one of (3) to (7), wherein the second determination includes determining an end of the speech acquired by the speech acquisition means, and the processing means causes the speech output means to output speech in response to the second determination determining that the speech has ended. (9) The speech processing device according to any one of (3) to (8), wherein the storage means further stores a plurality of second output speech information without associating it with speech information of the predetermined input word, and the second output speech information includes speech information of a plurality of lines, and the processing means selects one from the plurality of second output speech information, causes the speech output means to output speech information of a third line from the plurality of lines included in the selected second output speech information, performs the second determination after the speech information of the third line has been output, and causes the speech output means to output speech information of a fourth line different from the third line from the plurality of lines when it is determined in the second determination that the speech has been acquired. (10) The speech processing device according to any one of (1) to (9), wherein the plurality of predetermined input words are different input words.(11) The speech processing device according to any one of (1) to (10), wherein the storage means stores speech information of the plurality of predetermined input words in syllable units for each spoken input word, and the processing means, in the first determination, determines whether the speech acquired by the speech acquisition means matches the speech information for each syllable unit of the plurality of predetermined input words. (12) The speech processing device according to (11), wherein the processing means performs the determination of whether the speech matches the speech information for each syllable unit of the plurality of predetermined input words, starting with the speech information of the first syllable of each input word. (13) The speech processing device according to (12), wherein the processing means performs the determination of whether the speech matches the speech information for each syllable unit of the plurality of predetermined input words in a round-robin manner for the plurality of predetermined input words. (14) The speech processing device according to any one of (1) to (13), wherein the processing means starts the processing in response to the speech acquisition means acquiring speech for a length of at least a predetermined time. (15) The voice processing device according to (14), wherein the predetermined time is one second. (16) The voice processing device according to (14) or (15), wherein the processing means starts the processing in response to the voice acquisition means acquiring voice with a predetermined sound pressure or higher. (17) The voice processing device according to any one of (1) to 16, wherein the voice information of the plurality of predetermined input words is stored in the storage means by the voice acquisition means acquiring voices uttered by the user for each of the plurality of predetermined input words. (18) A toy including the voice processing device according to any one of (1) to (17). (19) A computer program for causing a computer to function as the processing means of the voice processing device according to any one of (1) to (17).
[0057] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention.
Claims
1. A speech processing device comprising: a speech acquisition means for acquiring speech uttered by a user; a processing means for processing the acquired speech; a speech output means for outputting speech in accordance with the results of processing by the processing means; and a storage means for storing speech information of a plurality of predetermined input words and output speech information by the speech output means associated with each of the plurality of predetermined input words, wherein the processing means: starts the processing in accordance with the acquisition of the speech by the speech acquisition means; in the processing, makes a first determination to determine whether the speech acquired by the speech acquisition means corresponds to any of the speech information of the plurality of predetermined input words; and in accordance with the first determination that the speech corresponds to first speech information of the speech information of the plurality of predetermined input words, outputs first output speech information stored in the storage means in association with the first speech information from the speech output means.
2. The audio processing device described in claim 1, wherein the storage means stores a plurality of pieces of first output audio information in association with the first audio information, and the processing means outputs one of the plurality of pieces of first output audio information from the audio output means in response to determining in the first determination that the first output audio information corresponds to the first audio information.
3. The audio processing device of claim 2, wherein the first output audio information includes audio information of a plurality of lines, and the processing means, when outputting the first output audio information from the audio output means, outputs audio information of a first line among the plurality of lines, performs a second determination to determine whether audio has been acquired by the audio acquisition means after the audio information of the first line has been output, and if it is determined in the second determination that the audio has been acquired, outputs audio information of a second line different from the first line among the plurality of lines from the audio output means.
4. The audio processing device described in claim 3, wherein the audio information of the multiple lines included in the first output audio information further includes audio information of at least one additional line different from the first line and the second line, which is output after the audio information of the second line is output.
5. The audio processing device of claim 1, wherein the first output audio information includes audio information of a first line, and the processing means, when outputting the first output audio information from the audio output means, outputs the audio information of the first line, and after the audio information of the first line has been output, makes a second determination to determine whether audio has been acquired by the audio acquisition means, and when it is determined in the second determination that the audio has been acquired, outputs audio information of any one of the multiple lines of audio information stored in the storage means from the audio output means.
6. The audio processing device according to claim 3, wherein the processing means determines that the audio has been acquired in response to the audio being detected by the audio acquisition means in the second determination.
7. The audio processing device according to claim 6, wherein in the second determination, the processing means determines whether the audio has been acquired, but does not make a determination about the content of the acquired audio.
8. The audio processing device according to claim 7, wherein the second determination includes determining that the audio acquired by the audio acquisition means has ended, and the processing means causes the audio output means to output audio in response to the second determination determining that the audio has ended.
9. The voice processing device according to claim 8, wherein the storage means further stores a plurality of second output voice information without associating it with the voice information of the specified input word, the second output voice information including voice information of a plurality of lines, the processing means selects one of the plurality of second output voice information, causes the voice output means to output voice information of a third line among the plurality of lines included in the selected second output voice information, performs the second judgment after the voice information of the third line has been output, and when it is determined in the second judgment that the voice has been acquired, causes the voice output means to output voice information of a fourth line different from the third line among the plurality of lines.
10. The speech processing device according to claim 9, wherein the plurality of predetermined input words are different input words.
11. A speech processing device as described in claim 10, wherein the storage means stores speech information of the plurality of predetermined input words in units of syllables for each spoken input word, and the processing means, in the first determination, determines whether the speech acquired by the speech acquisition means matches the speech information in units of syllables for each of the plurality of predetermined input words.
12. A speech processing device according to claim 11, wherein said processing means performs said determination of whether said speech matches said syllable-by-syllable speech information of each of said plurality of predetermined input words, starting with the speech information of the first syllable of each input word.
13. The speech processing device according to claim 12, wherein said processing means performs said determination of whether said speech matches said syllable-by-syllable speech information of each of said plurality of predetermined input words in a brute-force manner for said plurality of predetermined input words.
14. The audio processing device according to claim 13, wherein said processing means starts said processing in response to said audio acquisition means acquiring audio having a length equal to or longer than a predetermined time.
15. The audio processing device according to claim 14, wherein the predetermined time is one second.
16. The audio processing device according to claim 14, wherein said processing means starts said processing in response to said audio acquisition means acquiring audio with a predetermined sound pressure or higher.
17. A speech processing device according to claim 16, wherein the speech information of the plurality of predetermined input words is stored in the storage means by the speech acquisition means acquiring speech uttered by the user for each of the plurality of predetermined input words.
18. A toy equipped with a sound processing device according to any one of claims 1 to 17.
19. A computer program for causing a computer to function as the processing means of the speech processing device according to any one of claims 1 to 17.
Citation Information
Patent Citations
Voice conversation device
JP1995210193A
Voice recognition and generation apparatuses, toy with the apparatus, and recording medium having voice recognition and generation control program recorded therein
JP1998179941A
Robot controller and method for robot control and recording medium
JP2002116792A
Simulated conversation system and information storage medium
JP2002169591A