Robot, speech synthesis program, and speech output method
The robot system addresses the issue of excessive linguistic information in voice communication by generating and outputting modified phonological information, promoting user attachment and comfort through nonverbal communication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GROOVE X INC
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-13
AI Technical Summary
Existing voice communication by robots can diminish therapeutic effects due to excessive linguistic information, making the robot's voice seem persuasive and explanatory, and verbal communication is not always necessary for providing comfort.
A robot system that includes a phonology acquisition unit, a phonology generation unit, and a speech synthesis unit to generate and output speech with modified phonological information, reducing linguistic information and enabling imperfect mimicry.
Promotes user attachment and comfort by engaging in nonverbal communication through imperfect mimicry, enhancing the robot's endearing qualities and maintaining conversation without relying on exact linguistic repetition.
Smart Images

Figure 2026064239000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robot that outputs voice, a voice synthesis program, and a voice output method.
Background Art
[0002] When a robot outputs voice in response to an appeal from a user (e.g., a call, contact, etc.) or internal parameters (e.g., emotion parameters, etc.), the user can get the feeling that the robot has intentions and can develop an attachment to the robot.
[0003] Voice includes para-linguistic information in addition to linguistic information. Linguistic information is the information of phonemes representing concepts, and para-linguistic information is non-linguistic information such as voice quality, rhythm (pitch, intonation, rhythm, pause, etc.) of voice. It is known that users can obtain a healing effect by performing non-linguistic communication such as animal therapy. However, voice communication also includes not only linguistic communication by linguistic information but also non-linguistic communication by para-linguistic information. By effectively utilizing this non-linguistic communication in the voice output of a robot, healing can be given to the user (see, for example, Patent Document 1).
[0004] On the other hand, when a robot expresses some concept (emotion, intention, meaning, etc.) by linguistic information in voice, the linguistic communication between the robot and the user is enriched, and the user comes to have an attachment to the robot.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, when a robot communicates with a user through voice output, if the linguistic communication contains too much clear linguistic information, the user may perceive the robot's voice as persuasive and explanatory, diminishing the therapeutic effect of nonverbal communication.
[0007] Furthermore, in voice communication between robots, verbal communication is not always necessary; by engaging in conversations that do not rely on language, robots can provide comfort to users who are watching.
[0008] Therefore, the present invention aims to promote the formation of attachment between users and robots in voice communication with others through the output of voices by robots. [Means for solving the problem]
[0009] A robot according to one aspect of the present invention includes a phonology acquisition unit that acquires first phonological information consisting of a plurality of phonemes, a phonology generation unit that generates second phonological information different from the first phonological information based on at least some of the phonemes included in the first phonological information, a speech synthesis unit that synthesizes speech according to the second phonological information, and a speech output unit that outputs the speech.
[0010] Furthermore, in one aspect of the present invention, a speech synthesis program causes a robot's computer to function as a phonology acquisition unit that acquires first phonological information consisting of a plurality of phonemes, a phonology generation unit that generates second phonological information different from the first phonological information based on at least some of the phonemes included in the first phonological information, and a speech synthesis unit that synthesizes speech according to the second phonological information.
[0011] Furthermore, one embodiment of the present invention is a voice output method for a robot, comprising: a phonology acquisition step of acquiring first phonological information consisting of a plurality of phonemes; a phonology generation step of generating second phonological information different from the first phonological information based on at least some of the phonemes included in the first phonological information; a voice synthesis step of synthesizing speech according to the second phonological information; and a voice output step of outputting the speech. [Effects of the Invention]
[0012] According to the present invention, a phonological generation unit generates second phonological information based on at least some of the phonologies contained in the acquired first phonological information. A speech synthesis unit synthesizes speech according to such second phonological information. This makes it possible to promote the formation of attachment between the user and the robot in voice communication with others through the robot's speech output. [Brief explanation of the drawing]
[0013] The aforementioned objectives, as well as other objectives, features, and advantages, will become even clearer from the preferred embodiments described below and the accompanying drawings.
[0014] [Figure 1A] Figure 1A is a front view of a robot according to an embodiment of the present invention. [Figure 1B] Figure 1B is a side view of a robot according to an embodiment of the present invention. [Figure 2] Figure 2 is a schematic cross-sectional view showing the structure of a robot according to an embodiment of the present invention. [Figure 3] Figure 3 shows the hardware configuration of a robot according to an embodiment of the present invention. [Figure 4] Figure 4 is a block diagram showing the configuration for outputting sound in a robot according to an embodiment of the present invention. [Figure 5] Figure 5 is a block diagram showing in detail the configuration of the string input unit, sensing unit, and acquisition unit according to an embodiment of the present invention. [Figure 6]FIG. 6 is an example of a prosody-emotion table that defines the relationship between prosody and emotion parameters in an embodiment of the present invention. [Figure 7] FIG. 7 is a block diagram showing in detail the configuration of a generation unit, a voice synthesis unit, and an output unit in an embodiment of the present invention. [Figure 8A] FIG. 8A is a diagram showing an example of a prosody curve used by the voice synthesis unit in an embodiment of the present invention. [Figure 8B] FIG. 8B is a diagram showing an example of a prosody curve used by the voice synthesis unit in an embodiment of the present invention. [Figure 8C] FIG. 8C is a diagram showing an example of a prosody curve used by the voice synthesis unit in an embodiment of the present invention. [Figure 8D] FIG. 8D is a diagram showing an example of a prosody curve used by the voice synthesis unit in an embodiment of the present invention. [Figure 9] FIG. 9 is a diagram showing an example of two prosody curves connected by the voice synthesis unit in an embodiment of the present invention.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, embodiments of the present invention will be described. Note that the embodiments described below are examples of implementing the present invention, and the present invention is not limited to the specific configurations described below. In implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.
[0016] [[ID=:30]]A robot according to an embodiment of the present invention includes an acquisition unit that acquires first prosody information composed of a plurality of prosodies, a generation unit that generates second prosody information different from the first prosody information based on at least a part of the prosodies included in the first prosody information, a voice synthesis unit that synthesizes voice according to the second prosody information, and an output unit that outputs the voice.
[0017] In this configuration, the robot first outputs speech by synthesizing speech according to phonological information, rather than by playing pre-prepared sound sources. The robot then generates second phonological information that is different from the first phonological information, based on at least some of the phonological information acquired in the first phonological information, and the speech synthesis unit synthesizes speech according to the second phonological information thus generated. This allows the robot to generate second phonological information with some phonological modifications, even when mimicking the first phonological information acquired by speech sensing. This enables imperfect mimicry (speech imitation), increasing the robot's cuteness and promoting user attachment to the robot. Furthermore, when robots converse with each other, they acquire first phonological information from the other robot's speech and synthesize and output speech according to the second phonological information, which is different from that. By having both robots converse in this way, they can continue the conversation, which in turn promotes user attachment to the robot.
[0018] The phonological generation unit may generate the second phonological information having less linguistic information than the first phonological information.
[0019] This configuration reduces the amount of linguistic information contained in the acquired first phonological information to generate the second phonological information, thereby enabling speech communication at the level of an infant, for example, due to immature language abilities. The method for reducing the amount of linguistic information contained in the first phonological information may be, for example, the deletion, modification, or addition of some characters or phonemes to the phonomes of the first phonological information.
[0020] The robot may further include a sensing unit that senses the external environment and generates an input signal, and the phoneme acquisition unit may acquire the first phoneme information based on the input signal.
[0021] The sensing unit may be a microphone that senses sound and generates an audio signal as the input signal, and the phonological acquisition unit may determine the language information based on the audio signal and acquire the first phonological information including the language information.
[0022] The phonological acquisition unit may perform speech recognition on the speech signal and acquire the first phonological information having the recognized speech as linguistic information.
[0023] This configuration allows the robot to perform an imperfect mimicry, where it incompletely imitates and repeats the sounds it hears. For example, if a user says "orange" to the robot, the robot acquires first phonological information containing the linguistic information "orange." The robot then generates second phonological information containing the linguistic information "two oranges," which is "orange" with some consonants swapped, and outputs this as speech. As a result, the user understands that the robot is trying to parrot back "orange," while also finding the imperfect mimicry endearing.
[0024] The phonological acquisition unit may perform speech recognition on the speech signal and acquire the first phonological information having the response to the recognized speech as linguistic information.
[0025] This configuration allows the robot to engage in conversations by responding to heard sounds with incomplete linguistic expressions, enabling the user to understand the robot's responses while also increasing its endearing qualities. For example, if a user asks the robot, "What should we do?", and the robot receives first phonological information containing the linguistic information "hug," it will generate second phonological information, "dako," by removing the geminate consonant in "dakko," and output it as speech. This allows the user to understand that the robot is requesting a hug, while also finding its incomplete linguistic expression endearing.
[0026] The sensing unit may be a camera that senses incident light and generates an image signal as the input signal, and the phonological acquisition unit may determine the language information based on the image signal and acquire first phonological information having the language information.
[0027] The phonological acquisition unit may perform character recognition on the image signal and acquire the first phonological information which includes the recognized characters as linguistic information.
[0028] This configuration allows the robot to pronounce the characters it sees and recognizes not exactly as they are, but as incomplete linguistic expressions. Users can understand that the robot is trying to read the characters it sees, and it also enhances the robot's endearing quality. For example, if the robot recognizes characters from an image signal and acquires first phonological information containing the linguistic information "tokei" (clock), it will generate second phonological information, "toke" (toke), by deleting some characters from "tokei," and output it as speech. This allows the user to understand that the robot is trying to read the word "tokei," while also finding its incomplete linguistic expression endearing.
[0029] The phonological acquisition unit may perform object recognition on the image signal and acquire the first phonological information having linguistic information representing the recognized object.
[0030] This configuration allows the robot to express the recognized object not directly, but with incomplete linguistic information. This enables the user to understand that the robot is trying to represent the recognized object, and also enhances the robot's endearing qualities. For example, if the robot recognizes a clock by performing object recognition on an image signal and obtains first phonological information containing the linguistic information "tokei" (clock), it will generate second phonological information containing the linguistic information "toke" (toke), which is a partial deletion of "tokei," and output this as speech. This allows the user to understand that the robot has recognized a clock, while also finding its incomplete linguistic expression endearing.
[0031] The phonological generation unit may identify emotion parameters corresponding to at least some of the phonologies of the first phonological information and generate the second phonological information based on the identified emotion parameters.
[0032] This configuration allows the robot to generate second phonological information based on emotional parameters corresponding to the phonological information, rather than the linguistic information of the first phonological information it acquires, thereby enabling nonverbal communication. In this nonverbal communication, the first and second phonological information may be meaningless sequences of phonological information with little linguistic information, such as onomatopoeia (e.g., "ooh-ooh").
[0033] The phonological generation unit may generate second phonological information having emotion parameters close to the emotion parameters.
[0034] The robot may further include a table defining the relationship between phonomes and emotion parameters, and the phonome generation unit may refer to the table to identify emotion parameters corresponding to at least some of the phonomes of the first phonome information.
[0035] The robot may further include a table defining the relationship between phonology and emotion parameters, and the phonology generation unit may generate the second phonological information by referring to the table.
[0036] The robot may further include a microphone that senses sound and generates an audio signal, and the phonological acquisition unit may acquire first phonological information by performing speech recognition on the audio signal.
[0037] The phonological generation unit may generate the second phonological information consisting of a predetermined number of syllables or less (for example, two syllables), regardless of the number of syllables in the first phonological information.
[0038] Furthermore, a speech synthesis program according to one aspect of the present invention is executed on a robot's computer, causing the robot's computer to function as a phonology acquisition unit that acquires first phonological information consisting of a plurality of phonemes, a phonology generation unit that generates second phonological information different from the first phonological information based on at least some of the phonemes included in the first phonological information, and a speech synthesis unit that synthesizes speech according to the second phonological information.
[0039] Furthermore, one embodiment of the present invention is a speech output method for a robot, comprising: a phonology acquisition step of acquiring first phonological information consisting of a plurality of phonemes; a phonology generation step of generating second phonological information different from the first phonological information based on at least some of the phonemes included in the first phonological information; and a speech synthesis step of synthesizing speech according to the second phonological information. The step includes an audio output step that outputs the aforementioned audio.
[0040] The robot of this embodiment will be described below with reference to the drawings.
[0041] Figure 1A is a front view of the robot, and Figure 1B is a side view of the robot. In this embodiment, robot 100 is an autonomous robot that determines its actions, gestures, and voice based on the external environment and internal state. The external environment is detected by a group of sensors including a camera, microphone, accelerometer, and touch sensor. The internal state is quantified as various parameters that express the emotions of robot 100.
[0042] As a parameter for expressing emotions, robot 100 has, for example, an affinity parameter for each user. When robot 100 performs an action that shows affection towards the user, such as picking them up or talking to them, the sensor group detects this action and increases the affinity level with that user. On the other hand, robot 100 decreases its affinity level with users who do not interact with it, users who act rudely, or users it does not encounter often.
[0043] The body 104 of the robot 100 has an overall rounded shape and includes an outer skin made of a soft, elastic material such as urethane, rubber, resin, or fiber. The weight of the robot 100 is 15 kg or less, preferably 10 kg or less, and more preferably 5 kg or less. The height of the robot 100 is 1.2 m or less, preferably 0.7 m or less. In particular, it is desirable to make the robot small and lightweight by making it weigh about 5 kg or less and keeping the height about 0.7 m or less, so that users, including children and the elderly, can easily carry the robot 100.
[0044] Robot 100 is equipped with three wheels for three-wheeled movement. As shown in the figure, robot 100 includes a pair of front wheels 102 (left wheel 102a, right wheel 102b) and one rear wheel 103. The front wheels 102 are drive wheels, and the rear wheel 103 is a driven wheel. The front wheels 102 do not have a steering mechanism, but the rotational speed and direction of the left wheel 102a and the right wheel 102b can be controlled individually.
[0045] The rear wheels 103 are so-called omni wheels or casters, and are rotatable to move the robot 100 forward, backward, left, and right. By increasing the forward rotation speed of the right wheel 102b compared to the left wheel 102a (including cases where the left wheel 102a is stopped or rotates in the backward direction), the robot 100 can turn left or rotate counterclockwise. Also, by increasing the forward rotation speed of the left wheel 102a compared to the right wheel 102b (including cases where the right wheel 102b is stopped or rotates in the backward direction), the robot 100 can turn right or rotate clockwise.
[0046] The front wheels 102 and rear wheels 103 can be completely retracted into the body 104 by the drive mechanism. Even when the robot is moving, most of each wheel is hidden within the body 104, but when each wheel is completely retracted into the body 104, the robot 100 becomes immobile. That is, as the wheels are retracted, the body 104 lowers and the robot 100 sits on the floor surface F. In this seated state, the flat seating surface 108 (installation bottom surface) formed on the bottom of the body 104 comes into contact with the floor surface F, allowing the robot 100 to maintain a stable seated position.
[0047] The robot 100 has two hands 105. The robot 100 is capable of actions such as raising, shaking, and vibrating the hands 105. The two hands 105 can be controlled individually.
[0048] The eye 106 is capable of displaying images using a display device consisting of elements such as liquid crystal elements or organic EL elements. The robot 100 has a microphone capable of identifying the direction of the sound source, an ultrasonic sensor, and an odor sensor. It is equipped with various sensors such as a sensor, distance sensor, and acceleration sensor. Furthermore, the robot 100 has a built-in speaker and can output sound. A capacitive touch sensor is installed on the body 104 of the robot 100. The touch sensor allows the robot 100 to detect user touch.
[0049] A horn 109 is attached to the head of robot 100. A panoramic camera is mounted on the horn 109, allowing for simultaneous imaging of the entire upper area of robot 100.
[0050] Figure 2 is a schematic cross-sectional view showing the structure of the robot 100. As shown in Figure 2, the body 104 of the robot 100 includes a base frame 308, a main frame 310, a pair of resin wheel covers 312, and an outer shell 314. The base frame 308 is made of metal and forms the axis of the body 104 and supports the internal structure. The base frame 308 is constructed by connecting an upper plate 332 and a lower plate 334 vertically with a plurality of side plates 336. Sufficient spacing is provided between the plurality of side plates 336 to allow for ventilation. Inside the base frame 308 are a battery 117, a control circuit 342, and various actuators.
[0051] The main frame 310 is made of resin and includes a head frame 316 and a torso frame 318. The head frame 316 is hollow and hemispherical and forms the head skeleton of the robot 100. The torso frame 318 consists of a neck frame 3181, a chest frame 3182, and an abdominal frame 3183, and as a whole has a stepped cylindrical shape and forms the torso skeleton of the robot 100. The torso frame 318 is fixed integrally with the base frame 308. The head frame 316 is assembled to the upper end (neck frame 3181) of the torso frame 318 so as to be displaceable relative to it.
[0052] The head frame 316 is provided with three axes: a yaw axis 320, a pitch axis 322, and a roll axis 324, and actuators 326 for rotationally driving each axis. The actuators 326 include multiple servo motors for individually driving each axis. The yaw axis 320 is driven for head-turning motion, the pitch axis 322 is driven for nodding motion, and the roll axis 324 is driven for head tilting motion.
[0053] A plate 325 is fixed to the upper part of the head frame 316 to support the yaw axis 320. The plate 325 has multiple ventilation holes 327 to ensure airflow between the upper and lower parts.
[0054] A metal base plate 328 is provided to support the head frame 316 and its internal mechanism from below. The base plate 328 is connected to plate 325 via cross link 329 (pantograph mechanism), while also being connected to upper plate 332 (base frame 308) via joint 330.
[0055] The fuselage frame 318 houses the base frame 308 and the wheel drive mechanism 370. The wheel drive mechanism 370 includes a rotating shaft 378 and an actuator 379. The lower half of the fuselage frame 318 (abdominal frame 3183) is narrowed to form a housing space Sp for the front wheels 102 between it and the wheel cover 312.
[0056] The outer skin 314 covers the main frame 310 and the pair of hands 105 from the outside. The outer skin 314 has a thickness that a person can feel as elastic, and is formed by using a soft, stretchable material such as urethane sponge as the base material and wrapping it with a smooth-to-the-touch fabric such as polyester. As a result, when a user hugs the robot 100, they feel a moderate softness and can engage in natural physical contact as a person would with a pet. The upper end of the outer skin 314 is protected from the outside air. An opening 309 is provided for introducing [something].
[0057] Figure 3 shows the hardware configuration of robot 100. Robot 100 includes a display device 110, an internal sensor 111, a speaker 112, a communication unit 113, a storage device 114, a processor 115, a drive mechanism 116, and a battery 117 within its housing 101. The drive mechanism 116 includes the wheel drive mechanism 370 described above. The processor 115 and the storage device 114 are included in the control circuit 342.
[0058] Each unit is connected to the others by power lines 120 and signal lines 122. Battery 117 supplies power to each unit via power lines 120. Each unit sends and receives control signals via signal lines 122. Battery 117 is, for example, a lithium-ion rechargeable battery and is the power source for robot 100.
[0059] The drive mechanism 116 is an actuator that controls the internal mechanisms. The drive mechanism 116 has the function of moving and changing the direction of the robot 100 by driving the front wheels 102 and rear wheels 103. The drive mechanism 116 also controls the hand 105 via wire 118 to perform actions such as raising the hand 105, swinging the hand 105, and driving the hand 105. The drive mechanism 116 also has the function of controlling the head to change the direction of the head.
[0060] The internal sensor 111 is a collection of various sensors built into the robot 100. Examples of internal sensors 111 include a camera (360-degree camera), a microphone, a distance sensor (infrared sensor), a thermal sensor, a touch sensor, an acceleration sensor, and an odor sensor. The speaker 112 outputs sound.
[0061] The communication unit 113 is a communication module that performs wireless communication with various external devices such as servers, external sensors, other robots, and user-held portable devices. The storage device 114 consists of non-volatile memory and volatile memory and stores various programs, including the speech synthesis program described later, and various setting information.
[0062] The display device 110 is installed at the eye position of the robot 100 and has the function of displaying an image of the eye. The display device 110 displays an image of the robot 100's eye by combining eye parts such as the pupil and eyelid. If external light shines into the eye, a catchlight may be displayed at a position corresponding to the location of the external light source.
[0063] Figure 4 is a block diagram showing the configuration for outputting voice in robot 100. Robot 100 includes an emotion generation unit 51, a sensing unit 52, a phonology acquisition unit 53, a phonology generation unit 54, a speech synthesis unit 55, and a voice output unit 56. The emotion generation unit 51, the phonology acquisition unit 53, the phonology generation unit 54, and the speech synthesis unit 55 are realized by a computer executing the speech synthesis program of this embodiment.
[0064] The emotion generation unit 51 determines the emotions of the robot 100. The emotions of the robot 100 are expressed by multiple emotion parameters. The emotion generation unit 51 determines the emotions of the robot 100 according to predetermined rules, based on the external environment and internal parameters sensed by the sensing unit 52.
[0065] The sensing unit 52 corresponds to the internal sensors 111 mentioned above and includes a camera (360-degree camera), microphone, distance sensor (infrared sensor), thermal sensor, touch sensor, acceleration sensor, odor sensor, etc. The sensing unit 52 senses the environment outside the robot 100 and generates an input signal.
[0066] The phonological acquisition unit 53 acquires phonological information based on emotion parameters input from the emotion generation unit 51 or input signals input from the sensing unit 52. Phonological information is generally information about a sequence of multiple phonemes arranged in order, but it may also consist of a single phoneme (one syllable). Phonemes can be represented, for example, in the case of Japanese, by kana; in the case of English, by phonetic symbols; and in the case of Chinese, by pinyin. The method of acquiring phonological information in the phonological acquisition unit 53 will be described in detail later.
[0067] The phonological generation unit 54 generates phonological information different from the phonological information acquired by the phonological acquisition unit 53, based on at least some of the phonological information acquired by the phonological acquisition unit 53. Hereinafter, the phonological information acquired by the phonological acquisition unit 53 will be referred to as "first phonological information," and the phonological information generated by the phonological generation unit 54 will be referred to as "second phonological information." The second phonological information is different from the first phonological information, but is generated based on at least some of the phonological information of the first phonological information. In this embodiment, even if the first phonological information input from the phonological acquisition unit 53 has three or more syllables, the phonological generation unit 54 generates 2-syllable phonological information as the second phonological information. Typically, for example, if the first phonological information consists of three syllables, the phonological generation unit 54 deletes one of them, and uses only the remaining two syllables as the second phonological information. The method of generating the second phonological information in the phonological generation unit 54 will be described in detail later.
[0068] The speech synthesis unit 55 synthesizes speech according to the second phonological information generated by the phonological generation unit 54. The speech synthesis unit 55 can be configured as a synthesizer. The speech synthesis unit 55 stores parameters for speech synthesis corresponding to each phonological, and when the second phonological information is provided, it determines the parameters for outputting the corresponding phonological speech and synthesizes speech. The speech synthesis in the speech synthesis unit 55 will be described in detail later.
[0069] The audio output unit 56 corresponds to the speaker 112 and outputs the audio synthesized by the audio synthesis unit 55.
[0070] As described above, the robot 100 of this embodiment is equipped with a speech synthesis unit 55 that synthesizes speech, so it can synthesize and output any desired speech. Therefore, it is not limited to outputting only fixed speech, as in the case of playing pre-prepared speech files, but can output speech according to second speech information generated based on first speech information. As a result, the user can perceive the robot 100's speech as being lifelike.
[0071] Furthermore, the robot 100 of this embodiment does not synthesize speech using the acquired first phonological information as is, but rather generates second phonological information based on at least some of the phonologies of the first phonological information, and synthesizes speech according to the second phonological information.Here, if the first phonological information contains linguistic information, generating second phonological information using some of the phonologies of the first phonological information reduces the amount of linguistic information contained in the first phonological information.
[0072] This allows for the synthesis of speech with some phonemes modified, even when mimicking speech recognized by speech recognition. This enables imperfect mimicry (speech imitation), increasing the robot's cuteness. Furthermore, when robots converse with each other, they can recognize the other robot's speech and synthesize speech with different phoneme sequences while utilizing at least some of the recognized phonemes. By having both robots converse, they can maintain a conversation (without repeating the same speech). In this specification, the linguistic information contained in phonological information consisting of multiple phonemes (phoneme sequences) refers to the linguistic meaning represented by that phoneme sequence. For example, phoneme sequences that do not represent a specific meaning, such as onomatopoeia, are understood to have no linguistic information, or to have an extremely low amount of linguistic information.
[0073] Next, the acquisition of the first phonological information in the phonological acquisition unit 53 will be described in detail. Figure 5 is a block diagram showing in detail the configuration of the emotion generation unit 51, sensing unit 52, and phonological acquisition unit 53 of the robot 100 shown in Figure 4. In the example in Figure 5, the sensing unit 52 includes a microphone 521 and a camera 522. The phonological acquisition unit 53 includes a speech recognition unit 531, a character recognition unit 532, an object recognition unit 533, an emotion acquisition unit 534, a response generation unit 535, and a phonological information acquisition unit 536.
[0074] As described above, the emotion generation unit 51 determines the emotion of the robot 100 according to predetermined rules based on the external environment and internal parameters sensed by the sensing unit 52, and outputs emotion parameters to the phonology acquisition unit 53. The microphone 521 senses sound as the external environment, generates an audio signal as an input signal, and outputs it to the phonology acquisition unit 53. The camera 522 senses incident light as the external environment, generates an image signal as an input signal, and outputs it to the phonology acquisition unit 53.
[0075] The speech recognition unit 531 performs speech recognition on the audio signal obtained by sensing sound with the microphone 521 to acquire a string of characters. The speech recognition unit 531 outputs the string of characters obtained by speech recognition to the response generation unit 535 and the phonological information acquisition unit 536. Any existing speech recognition engine can be used for this speech recognition. In general speech recognition engines, after recognizing phoneme sequences from the input audio signal, natural language processing such as morphological analysis is performed on those phoneme sequences to obtain a string of characters that contains linguistic information. In this embodiment, the string of characters from which linguistic information has been obtained by natural language processing is output to the response generation unit 535 and the phonological information acquisition unit 536. This string of characters contains phonological information (i.e., phoneme sequences) and linguistic information (i.e., information obtained by natural language processing).
[0076] The response generation unit 535 generates a response to the speech recognized by the speech recognition unit 531 and outputs the string of this response to the phonological information acquisition unit 536. Any existing dialogue engine can be used to generate this response. This dialogue engine may generate a response to the recognized speech using a machine learning model that has learned responses to input strings.
[0077] The character recognition unit 532 acquires a string of characters by performing character recognition on the image signal obtained from the camera 522 capturing the area around the robot 100, and outputs it to the phonological information acquisition unit 536. Any existing character recognition engine can be used for this character recognition. The character recognition engine can perform character recognition using a machine learning model such as a neural network. The character recognition engine may recognize each character of the string of characters independently from the input image signal. Alternatively, the character recognition engine may obtain a string of characters containing linguistic information by performing natural language processing on the strings after recognizing them from the input image signal.
[0078] The object recognition unit 533 performs object recognition on the image signal obtained by the camera 522 capturing images of the area around the robot 100. Any existing object recognition engine can be used for this object recognition. The object recognition engine recognizes objects in the image and assigns a label indicating the name of the object. Machine learning models such as neural networks can also be used for the object recognition engine. This object recognition also includes person recognition, which involves recognizing the faces of people in the image to identify users. In the case of person recognition, the user's name is obtained as a label as a result of face recognition. The object recognition unit 533 outputs the string of labels obtained through recognition to the phonological information acquisition unit 536.
[0079] The emotion acquisition unit 534 acquires emotion parameters from the emotion generation unit 51 and, by referring to the phonology-emotion table, determines the two-syllable phonology that is closest to the acquired emotion parameters.
[0080] Figure 6 shows an example of a phonome-emotion table that defines the relationship between phonomes and emotion parameters. As shown in Figure 6, each phonome has four emotion parameters defined: "calm," "anger," "joy," and "sarrow." Each emotion parameter takes a value between 0 and 100.
[0081] The emotion acquisition unit 534 determines the two-syllable phoneme closest to the acquired emotion parameter by selecting from the phoneme-emotion table the two-syllable phoneme having the emotion parameter whose sum of differences with each acquired emotion parameter is smallest. The method of determining the phoneme based on the emotion parameter is not limited to this; for example, the emotion acquisition unit 534 may select the phoneme whose sum of differences with the largest values of some (e.g., two) of the acquired emotion parameters is smallest.
[0082] The phonological information acquisition unit 536 acquires strings input from the speech recognition unit 531, the response generation unit 535, the character recognition unit 532, and the object recognition unit 533, and converts these strings into first phonological information. In the case of Japanese, the phonological information acquisition unit 536 acquires strings that include kanji characters or strings that consist only of kana characters. In the case of English, the phonological information acquisition unit 536 acquires strings consisting of one or more words expressed in the alphabet. In the case of Chinese, the phonological information acquisition unit 536 acquires strings consisting of multiple kanji characters. Furthermore, if the phonological information acquisition unit 536 acquires a phonological sequence from the emotion acquisition unit 534, it uses this phonological sequence as the first phonological information.
[0083] Here, phonological information consists of phonemes, which are the phonetic units in each language. As mentioned above, in the case of Japanese, phonological information can be represented by kana. In the case of English, phonological information can be represented by phonetic symbols. In the case of Chinese, phonological information can be represented by pinyin. In the case of Japanese, the phonological information acquisition unit 536 obtains first phonological information by referring to a dictionary that defines the relationship between kanji and their readings, replacing the kanji with kana, and arranging all the kana. In the case of English, the phonological information acquisition unit 536 obtains first phonological information by referring to a dictionary that defines the relationship between words and phonetic symbols, and replacing each word in the string with a phonetic symbol. In the case of Chinese, the phonological information acquisition unit 536 obtains first phonological information by referring to a dictionary that defines the relationship between each kanji and its pinyin, and replacing the kanji with its pinyin. The phonological information acquisition unit 536 outputs the obtained first phonological information to the phonological generation unit 54.
[0084] Figure 7 is a block diagram showing in detail the configurations of the phonology generation unit 54, speech synthesis unit 55, and speech output unit 56 of the robot 100 shown in Figure 4. The phonology generation unit 54 comprises an onomatopoeia generation unit 541, a language information generation unit 542, and a phonology information generation unit 543. The onomatopoeia generation unit 541 identifies emotion parameters corresponding to at least some of the phonoes in the first phonological information by referring to a phonology-emotion table. The onomatopoeia generation unit 541 determines a phono based on the identified emotion parameters and outputs the determined phono to the phonology information generation unit 543. Specifically, the onomatopoeia generation unit 541 in this embodiment determines a phono having emotion parameters close to the emotion parameters of the phono in the first phonological information.
[0085] Specifically, if the first phonological information includes a single-syllable phonological element, the onomatopoeia generation unit 541 refers to the phonological-emotion table to identify the emotion with the highest value among the emotion parameters of that phonological element. Then, the onomatopoeia generation unit 541 determines two other phonological elements that have the same emotion parameter value as that emotion. For example, if the first phonological information is only the single syllable "a", the onomatopoeia generation unit 541 refers to the four emotion parameters for the syllable "a" in the table. Of the four emotion parameters for "a", the one with the highest value is the "joy" parameter, which has a value of 50. Therefore, the onomatopoeia generation unit 541 determines Search for other phonemes with a "joy" parameter of 50, and determine phonemes such as "ru" and "ni".
[0086] If the first phonological information includes a two-syllable phono, the onomatopoeia generation unit 541 determines a two-syllable phono corresponding to the two-syllable phono in the first phonological information in the same manner as described above. If the first phonological information has three or more syllables, the onomatopoeia generation unit 541 arbitrarily or based on predetermined rules selects a two-syllable phono from the three or more syllable phonos. Then, for each selected phono, the onomatopoeia generation unit 541 determines a corresponding two-syllable phono in the same manner as described above. The number of syllables may be a predetermined number or less instead of two syllables.
[0087] The language information generation unit 542 generates a string with less linguistic information than the input first phonological information and outputs it to the phonological information generation unit 543. The language information generation unit 542 reduces the amount of linguistic information by partially deleting, partially changing, or partially adding characters or phonemes to the string of first phonological information. Whether to partially delete, partially change, or partially add, and which characters or phonemes to delete, change, or add, can be determined arbitrarily or based on predetermined rules.
[0088] The language information generation unit 542 may, for example, generate the string "toke" by deleting one character from "tokei" when the first phoneme information "tokei" is input. The language information generation unit 542 may generate the string "nikan" by swapping some of the consonants of "mikan" when the first phoneme information "mikan" is input. The language information generation unit 542 may generate the string "oayou" by deleting some of the consonants of "ohayou" when the first phoneme information "ohayou" is input. The language information generation unit 542 may generate the string "tukei" by adding a diphthong to "tokei" when the first phoneme information "tokei" is input. The language information generation unit 542 may generate the string "dako" by deleting the geminate consonant from "dakko" when the first phoneme information "dakko" is input. The strings "toke," "nikan," "oayou," "tukei," and "dako" generated by the language information generation unit 542 are similar to "tokei," "mikan," "ohayou," "tokei," and "dakko," respectively, but they are not exact matches, which means that the amount of information in their language information is reduced. The language information generation unit 542 may further reduce the language information by using a combination of methods such as partially deleting, partially changing, partially adding characters or phonemes, or rearranging the order of phonemes. Partial changes to characters or phonemes may be made by changing them to similar phonologies in another language.
[0089] The methods for reducing the amount of information in linguistic information are not limited to those described above. Reducing the number of phonemes, eliminating linguistic meaning, making words incomplete, or making some phonemes difficult to hear all reduce the amount of information in linguistic information. Alternatively, the types of phonemes that can be used may be limited, and each phoneme in the first phonological information may be replaced with one of the limited phonemes to generate the second phonological information. Alternatively, the second phonological information may be generated by deleting all phonemes in the first phonological information that are not usable.
[0090] In this way, by reducing the amount of linguistic information in the first phonological information that contains linguistic information and generating second phonological information, second phonological information similar to the linguistic information of the first phonological information is generated. Therefore, by having robot 100 synthesize and output speech according to this second phonological information, the user can guess, and is inclined to guess, what robot 100 is trying to say. That is, by robot 100 deliberately using childish language, it can make the user think, "The robot seems to want to say something, it wants to communicate something." In turn, it can lead the user to unconsciously come to understand robot 100, or to develop curiosity about robot 100. This allows users to hold objects or draw attention to Robot 100. This can have a psychological effect, keeping users engaged and gradually leading them to develop an attachment to Robot 100.
[0091] If robot 100 were to synthesize and output speech using the first phonological information containing linguistic information as is, for example, if robot 100 clearly pronounced "tokei" (clock), the user would simply recognize that it was saying "tokei" and would not pay any further attention to robot 100. On the other hand, if robot 100 reduced the amount of linguistic information and pronounced the incomplete "toke" (toke), the user might become aware of robot 100 and wonder if it was trying to say "tokei." Furthermore, if the user finds that incompleteness endearing, it could promote the formation of affection for robot 100 in the user.
[0092] In the above, an example of generating characters with 2 to 4 syllables was described in order to explain the generation of strings with reduced information content by the language information generation unit 542. As described above, the phonological generation unit 54 generates second phonological information containing a 2-syllable phonology. The language information generation unit 542 makes the generated second phonological information 2 syllables by partially deleting or partially adding characters or phonemes. By similar processing, it is possible to generate second phonological information with a predetermined number of syllables or less.
[0093] The onomatopoeia generation unit 541 determines the syllables in the manner described above, thereby generating second phonological information that has phonological expressions similar to the emotion represented by the phonological expression of the first phonological information. In this case, however, since linguistic information is not considered when generating the second phonological information, second phonological information consisting of two meaningless syllables is generated.
[0094] Furthermore, since the language information generation unit 542 generates a string with reduced information content from the first phonological information as described above, it is possible to generate a second phonological information that incompletely represents the first phonological information.
[0095] The phonological information generation unit 543 generates phonological information for the phonological sequence determined by the onomatopoeia generation unit 541, or generates phonological information for the string generated by the language information generation unit 542, and outputs it to the speech synthesis unit 55 as second phonological information.
[0096] The speech synthesis unit 55 synthesizes speech based on information other than phonological information. For example, the prosody (stress, length, pitch, etc.) of the synthesized speech may be determined based on information other than the second phonological information. Specifically, the speech synthesis unit 55 stores four types of prosodic curves as prosodic patterns, and determines the prosody of each syllable by applying one of these prosodic patterns to each syllable of the generated speech.
[0097] Figures 8A to 8D show four types of prosodic curves. The speech synthesis unit 55 determines the prosodicity of each syllable by assigning one of these prosodic curves to each syllable. The speech synthesis unit 55 selects the prosodic curve to assign according to the phonology (pronunciation) of the syllable. The prosodic curves assigned to each phonology are predetermined and stored in the speech synthesis unit 55 as a phonology-prosodic curve table. The prosodic curve in Figure 8A is an example of a prosodic curve assigned to the phonology "a". The prosodic curve in Figure 8B is an example of a prosodic curve assigned to the phonology "i". The speech synthesis unit 55 determines the prosodicity of each syllable by referring to this phonology-prosodic curve table.
[0098] Figure 9 shows a diagram of the prosody of two syllables. When the speech synthesis unit 55 determines the prosody of two consecutive syllables using a prosodic curve, it smoothly connects the prosodic curves of the two consecutive syllables, as shown in Figure 9. In the example in Figure 9, the prosodic curve in Figure 8A and the prosodic curve in Figure 8C are connected.
[0099] The speech synthesis unit 55 has a virtual vocal organ. Generally, the vocalization process is common to all organisms that have vocal organs. For example, in humans, the vocalization process involves air guided from the lungs and abdomen through the trachea, vibrating the vocal cords to produce sound, which then resonates in the oral cavity and nasal cavity, resulting in a louder sound. The shape of the mouth and tongue then changes, producing a variety of voices. Individual differences in voice arise from various factors such as body size, lung capacity, vocal cords, trachea length, oral cavity size, nasal cavity size, tooth alignment, and tongue movement. Furthermore, even in the same person, the condition of the trachea and vocal cords changes depending on their physical condition, causing their voice to change. Due to this vocalization process, each person has a different voice quality, and their voice also changes depending on their physical condition and internal state, such as emotions.
[0100] In another embodiment, the speech synthesis unit 55 generates sound by simulating the vocalization process in a virtual vocal organ based on this vocalization process. In other words, the speech synthesis unit 55 is a virtual vocal organ (hereinafter referred to as "virtual vocal organ"), and generates voice using a virtual vocal organ implemented in software. For example, the virtual vocal organ may have a structure that mimics the vocal organs of a human, or it may have a structure that mimics the vocal organs of an animal such as a dog or cat. By having a virtual vocal organ, it is possible to generate individual-specific voices even if the basic structure of the vocal organs is the same, by changing the size of the trachea in the virtual vocal organ, adjusting the tension of the vocal cords, or changing the size of the oral cavity for each individual. The parameters for generating sound do not simply include direct parameters for generating sound with a synthesizer, but also include values that specify the structural characteristics of each organ in the virtual vocal organ as parameters (hereinafter referred to as "static parameters"). Using these static parameters, the vocalization process is simulated and a voice is generated.
[0101] For example, humans can produce a variety of voices. High-pitched voices, low-pitched voices, singing along to melodies, laughing, shouting—they can produce virtually any voice as long as the structure of their vocal organs allows. This is because the shape and state of each organ that makes up the vocal organs change, and these can be consciously altered or unconsciously altered in response to emotions and stimuli. The speech synthesis unit 55 also has parameters (hereinafter referred to as "dynamic parameters") for the state of these organs that change in conjunction with the external environment and internal state, and performs simulations by changing these dynamic parameters in conjunction with the external environment and internal state.
[0102] Generally, stretching the vocal cords produces a higher pitch, while relaxing them causes them to contract, resulting in a lower pitch. For example, an organ that mimics the vocal cords has a static parameter called the degree of vocal cord stretching (hereinafter referred to as "tension"), and by adjusting the tension, it is possible to produce high-pitched or low-pitched voices. This makes it possible to create robots 100 with high-pitched voices and robots 100 with low-pitched voices. Also, just as a person's voice can become high-pitched when they are nervous, similarly, by changing the tension of the vocal cords as a dynamic parameter in conjunction with the tension state of robot 100, it is possible to make robot 100's voice higher when it is nervous. For example, when robot 100 recognizes a stranger or is suddenly lowered from being held, etc., when the internal parameter indicating the state of nervousness swings to a tense value, the tension of the vocal cords can be increased in conjunction with this to produce a high-pitched voice. In this way, by associating the internal state of robot 100 with the organs involved in the vocalization process, and adjusting the parameters of the related organs according to the internal state, it is possible to change the voice according to the internal state.
[0103] Here, static and dynamic parameters are parameters that describe the morphological state of each organ over time. The virtual vocal organs are simulated based on these parameters.
[0104] Furthermore, by generating voices based on simulations, only voices based on the structural constraints of the vocal organs are generated. In other words, since voices that are impossible for living beings are not generated, it is possible to generate voices that sound natural. By performing simulations and generating voices... Furthermore, it can not only pronounce similar syllables, but also generate voices that are influenced by the internal state of robot 100.
[0105] Robot 100 keeps its sensor group, including microphone 521 and camera 522, constantly operational, and its emotion generation unit 51 is also constantly operational. In this state, when a user speaks to robot 100, the above processing begins when robot 100's microphone 521 senses the sound and outputs an audio signal to phoneme acquisition unit 53. The above processing also begins when camera 522 captures the user's face and outputs an image signal to phoneme acquisition unit 53. The above processing also begins when camera 522 captures text and outputs an image signal to phoneme acquisition unit 53. Furthermore, the above processing begins when emotion generation unit 51 generates emotion parameters based on the external environment and internal parameters and outputs them to phoneme acquisition unit 53. Note that the detection results of the external environment by sensing unit 52 do not all trigger the generation of speech; it is determined according to the internal state of robot 100 at that time.
[0106] In the above embodiment, the phonological acquisition unit 53 received a string containing language information from the speech recognition unit 531 to the phonological information acquisition unit 536. However, instead, the phonological sequence recognized by the speech recognition unit 531 may be directly input to the phonological information acquisition unit 536, and the phonological information acquisition unit 536 may use the input phonological sequence as the first phonological information. In other words, natural language processing in the speech recognition unit 531 may not be necessary.
[0107] Furthermore, although the above embodiment illustrates a configuration in which the sensing unit 52 includes a microphone 521 and a camera 522, for example, if a thermosensor is used as the sensing unit 52, the sensing unit 52 may detect temperature and the phoneme acquisition unit 53 may acquire first phoneme information such as "cold" or "hot" according to the detected temperature. Similarly, if an odor sensor is used as the sensing unit 52, the sensing unit 52 may detect odor and the phoneme acquisition unit 53 may acquire first phoneme information such as "smelly" according to the detected odor.
[0108] Furthermore, in the above embodiment, the onomatopoeia generation unit 541 determined other phonemes that share the largest emotion parameter among the emotion parameters corresponding to the phonemes of the first phoneme information as phonemes with similar emotion parameters. However, the method for determining other phonemes is not limited to this. For example, phonemes having multiple emotion parameters where the differences between each of the multiple emotion parameters corresponding to the phonemes of the first phoneme information are small (for example, the sum of the differences is small) may be determined as phonemes with similar emotion parameters. Also, the onomatopoeia generation unit 541 may determine phonemes whose emotion parameters are significantly different from those corresponding to the phonemes of the first phoneme information. For example, for phonemes with a strong emotion parameter of "anger," a phoneme with a strong emotion parameter of "sadness" may be determined.
[0109] The robot 100 of this embodiment enables, for example, the following effects. Specifically, in the robot 100 of this embodiment, when the phonology acquisition unit 53 acquires first phonological information containing a three-syllable phonology through speech recognition, character recognition, object recognition, etc., the phonology generation unit 54 deletes one syllable from those three syllables and generates second phonological information consisting of a two-syllable phonology. As a result, the robot 100 imitates and outputs the heard sound with fewer syllables, making it possible to create an effect that resembles an infant with limited language ability imperfectly imitating and outputting the heard sound.
[0110] Furthermore, in the robot 100 of this embodiment, when the phonology acquisition unit 53 recognizes a two-syllable speech output from another robot and acquires first phonology information, the phonology generation unit 54 determines a phonology having emotional parameters that are close to or far from the emotional parameters corresponding to those two-syllable phonologies and generates second phonology information. Therefore, when such robots 100 meet... By having them talk, it becomes possible to create the illusion that the robots are having a conversation, influenced by each other's emotions.
[0111] The following describes various modifications of the robot 100 described above. The phonology acquisition unit 53 recognizes the pitch of the audio signal input from the microphone 521, and the speech synthesis unit 55 may synthesize a speech with the same pitch as the input audio signal. For example, if an audio signal of 440 Hz is input from the microphone 521, the speech synthesis unit 55 may synthesize a speech of the same 440 Hz. Alternatively, the speech synthesis unit 55 may synthesize a speech that matches the pitch of the input audio signal to a predetermined scale. For example, if an audio of 438 Hz is input from the microphone 521, the speech synthesis unit 55 may synthesize a speech of 440 Hz.
[0112] Furthermore, the phonology acquisition unit 53 may recognize changes in the pitch of the sound input from the microphone 521, and the speech synthesis unit 55 may synthesize a sound with the same pitch changes as the input sound signal. This makes it possible to create the effect that the robot 100 is mimicking and producing the melody of the sound it hears.
[0113] Furthermore, the sensing unit 52 may be equipped with a torque sensor for the front wheel 102, and the voice synthesis unit 55 may generate voice according to the value of this torque sensor. For example, when the robot 100 is unable to move in the direction of travel due to an obstacle and the torque of the front wheel increases, the voice synthesis unit 55 may synthesize a straining voice such as "hmmm".
[0114] Furthermore, in person recognition by the object recognition unit 533, if a person's face is suddenly recognized in the image at a predetermined size, the speech synthesis unit 55 may synthesize a laughing sound. Alternatively, if a person's face is suddenly recognized in the image at a predetermined size, the emotion generation unit 51 may generate an emotion parameter for "joy" and output it to the phonology acquisition unit 53, which may then perform the above-mentioned processing to acquire the first phonological information and generate the second phonological information to synthesize the sound.
[0115] Furthermore, in the above embodiment, the phoneme acquisition unit 53 acquired first phoneme information representing recognized characters and recognized objects from the image captured by the camera 522. However, when an object is recognized from the image, the phoneme acquisition unit 53 may generate a string of characters to speak to the object and acquire first phoneme information. For example, when the phoneme acquisition unit 53 recognizes a person through object recognition, it may acquire first phoneme information such as "dakko" (hug), which means "hug". Alternatively, when the phoneme acquisition unit 53 recognizes an object from the image, it may generate a string of related words associated with the object and acquire first phoneme information. For example, when an airplane is recognized through object recognition, the phoneme acquisition unit 53 may acquire first phoneme information such as the onomatopoeia "boom" associated with the airplane.
[0116] Furthermore, if a request is not fulfilled after outputting a request voice, the speech synthesis unit 55 may synthesize a voice with different volume, speaking speed, etc. For example, if the voice synthesis unit 55 synthesizes and outputs the voice "dako" as a request to be held, but the child is not held, the speech synthesis unit 55 may generate a voice with stronger emphasis, such as "Dako!"
[0117] Furthermore, the emotion generation unit 51 may generate the emotion of "joy" if, after outputting sound from the sound output unit 56, the speech recognition unit 531 recognizes a sound with the same phonemes as that sound. This makes it possible to create an effect where the robot 100 reacts with joy when the user imitates the robot 100's speech. In addition, after outputting sound from the sound output unit 56, the robot 100 may detect the user's reaction and assign a score to the outputted sound for learning purposes. For example, if the object recognition unit 533 detects a smile from an image after outputting sound, the robot 100 may assign a high score to that sound for learning purposes. The robot 100 may, for example, prioritize synthesizing and outputting sounds with high scores.
[0118] Furthermore, if the object recognition unit 533 recognizes an object and the speech recognition unit 531 recognizes a sound at the same time, the recognized object and the recognized sound are associated and learned, and then the object is recognized. In this case, the phonological acquisition unit 53 may acquire the first phonological information of the associated sound. For example, if the object recognition unit 533 recognizes a cup and the speech recognition unit 531 recognizes the sound "cup", this combination is learned, and then when the object recognition unit 533 recognizes a cup, the phonological acquisition unit 53 may acquire the first phonological information "cup". This allows the user to teach the robot 100 the names of objects, and enables the robot 100 to learn the names of objects taught by the user.
[0119] Furthermore, by repeating the learning process, the amount of information lost between the first and second phonological information can be reduced. For example, when the first phonological information, "otousan," is acquired during learning, initially, some of the phonomes in the first phonological information are deleted, their order is changed, and the second phonological information, "uo," is generated by placing the non-adjacent "u" and "o" in order. As learning progresses, the amount of information lost can be gradually reduced, for example, by generating the second phonological information, "tosa," by placing the non-adjacent "to" and "sa" in order, while deleting some of the phonomes, and finally, by placing the adjacent "o" and "to" in order, which are characteristic sounds (for example, phonomes with a strong accent), resulting in "oto."
[0120] Furthermore, the audio output unit 56 may adjust the volume of the output audio according to the volume of the sound sensed by the microphone 521. For example, if the volume of the sound sensed by the microphone 521 is high, the volume of the output audio may be increased. In addition, the audio output unit 56 may adjust the volume of the output audio according to the volume of the sound recognized as noise by the speech recognition unit 531. That is, in a noisy environment, the volume of the output audio may be increased.
[0121] Furthermore, while the above embodiment describes how robot 100 can continue a conversation with other robots 100, each robot 100 may also have the following additional functions in order to have conversations with other robots 100.
[0122] The emotion generation unit 51 may develop a story in a conversation between the robots 100 and generate emotions according to this story. The robot 100 then outputs speech that expresses the emotion using the functions of the phonology acquisition unit 53 or the speech output unit 56 described above. A machine learning model such as a neural network may also be used in the development of the story in the emotion generation unit 51.
[0123] The speech synthesis unit 55 may synthesize speech so as to harmonize its pitch with the speech of other robots 100 input from the microphone 521. This makes it possible to create the effect of multiple robots 100 singing in a chorus. Alternatively, by deliberately using a different pitch from the speech of other robots 100, it is possible to create the effect of tone-deafness.
[0124] Furthermore, the speech synthesis unit 55 may synthesize speech at pitches not typically used by humans. While the pitch of a normal human voice is at most around 500Hz, the robot 100 outputs speech at a higher pitch (for example, around 800Hz). Other robots 100 can recognize that the speech is coming from another robot 100 based solely on the pitch information. For example, when robots 100 are playing chase, they need to recognize the other robot's call and direction. If the input pitch is within a predetermined range, they can recognize that it is coming from the other robot 100 (meaning something like "Come here"). Additionally, combining the pitch with patterns (such as changes in the pitch curve) can further improve recognition accuracy. Also, while recognizing solely by pitch might pick up sounds like ambulance sirens, conversely, high-pitched sounds can be picked up. The unconditional response to something can also be used as an expression of animalistic behavior.
[0125] Furthermore, the phonology acquisition unit 53 acquires first phonological information based on input signals from the sensing unit 52 and emotion parameters from the emotion generation unit 51. The phonology acquisition unit 53 may also acquire information on volume, pitch, and timbre, which are elements that constitute sound, based on input signals, emotion parameters, or other information. In this case, the phonology generation unit 54 may also determine the volume, pitch, and timbre of the speech to be synthesized by the speech synthesis unit 55 based on the volume, pitch, and timbre information acquired by the phonology acquisition unit 53 and output it to the speech synthesis unit 55. In addition, the phonology acquisition unit 53 may also acquire the length (speech rate) of each phonology, and the phonology generation unit 54 may determine the speech rate of the speech to be output by the speech output unit 56 based on the acquired speech rate. Furthermore, the phonology acquisition unit 53 may also acquire language-specific characteristics as elements that constitute sound.
[0126] Furthermore, the phonology acquisition unit 53 may have a function to determine whether or not there is a melody (i.e., whether or not the input sound is a song or melody) based on the audio signal input from the microphone 521. In this case, the phonology acquisition unit 53 specifically assigns a score according to the change in pitch at predetermined intervals and determines whether or not there is a melody (i.e., whether or not a song is being sung) based on the score. If the phonology acquisition unit 53 determines that there is a melody in the input audio signal, the speech synthesis unit 55 may determine the length and pitch of each phonology of the speech to be synthesized so as to mimic the recognized melody. Also, if the phonology acquisition unit 53 determines that there is a melody in the input audio signal, the phonology generation unit 54 generates second phonological information using predetermined phonologies. The speech synthesis unit 55 may then determine the length and pitch of each phonology of the speech to be synthesized so as to mimic the recognized melody. This makes it possible to create an effect that sounds like humming.
[0127] Furthermore, the phonology acquisition unit 53 may acquire strings of characters in languages other than Japanese based on the input signal from the sensing unit 52. That is, the speech recognition unit 531 may recognize speech in a language other than Japanese and generate strings of characters in that language, the character recognition unit 532 may recognize characters in a language other than Japanese and generate strings of characters in that language, and the object recognition unit 533 may recognize an object and generate strings of characters in a language other than Japanese that represent that object.
[0128] Furthermore, when the robot 100 has made a predetermined number of mimic responses (for example, 5 times), it may return the previous predetermined number of mimic responses (for example, 4 times) in succession. As mentioned above, the robot 100 outputs two-syllable sounds, but if only two-syllable mimics are repeated, the user may get bored. Therefore, at predetermined intervals, the robot may combine and generate sounds that it has previously mimicked and spoken. This is expected to have the effect of making the user feel as if the robot 100 is trying to say something.
[0129] To this end, the robot 100 includes a memory unit that stores second phoneme information generated as a mimic, a counting unit that counts the number of mimics, and a determination unit that determines whether the number of mimics has reached a predetermined number (for example, 5 times). When the determination unit determines that the number of mimics has reached the predetermined number, the speech synthesis unit 35 reads the mimics stored in the memory unit and synthesizes speech by concatenating them. [Industrial applicability]
[0130] This invention is useful as a robot that can output voice, as it can facilitate the formation of user attachment to the robot during voice communication with others through the robot's voice output.
Claims
1. A phonological acquisition unit that acquires first phonological information consisting of multiple phonemes, A phonological generation unit that generates a second phonological information different from the first phonological information based on at least some of the phonologies included in the first phonological information, A speech synthesis unit that synthesizes speech according to the second phonological information, The audio output unit outputs the aforementioned sound, A robot equipped with [the following features].
2. The robot according to claim 1, wherein the phonological generation unit generates the second phonological information having less information than the linguistic information contained in the first phonological information.
3. It further includes a sensing unit that senses the external environment and generates an input signal, The robot according to claim 2, wherein the phoneme acquisition unit acquires the first phoneme information based on the input signal.
4. The sensing unit is a microphone that senses sound and generates an audio signal as the input signal. The robot according to claim 3, wherein the phonological acquisition unit determines the language information based on the speech signal and acquires the first phonological information including the language information.
5. The robot according to claim 4, wherein the phonological acquisition unit performs speech recognition on the speech signal and acquires the first phonological information having the recognized speech as linguistic information.
6. The robot according to claim 4, wherein the phonological acquisition unit performs speech recognition on the speech signal and acquires the first phonological information having a response to the recognized speech as linguistic information.
7. The sensing unit is a camera that senses incident light and generates an image signal as the input signal. The robot according to claim 3, wherein the phonological acquisition unit determines the language information based on the image signal and acquires first phonological information having the language information.
8. The robot according to claim 7, wherein the phoneme acquisition unit performs character recognition on the image signal and acquires the first phoneme information having the recognized characters as language information.
9. The robot according to claim 7, wherein the phoneme acquisition unit performs object recognition on the image signal and acquires the first phoneme information having linguistic information representing the recognized object.
10. The robot according to claim 1, wherein the phonological generation unit identifies emotion parameters corresponding to at least some of the phonologies of the first phonological information and generates the second phonological information based on the identified emotion parameters.
11. The robot according to claim 10, wherein the phonological generation unit generates second phonological information having emotion parameters close to the emotion parameters.
12. It also includes a table that defines the relationship between phonology and emotion parameters, The robot according to claim 10 or 11, wherein the phonological generation unit refers to the table to identify emotion parameters corresponding to at least some of the phonological information of the first phonological information.
13. It also includes a table that defines the relationship between phonology and emotion parameters. The robot according to claim 10 or 11, wherein the phonological generation unit generates the second phonological information by referring to the table.
14. The sensing unit is a microphone that senses sound and generates an audio signal as the input signal. The robot according to any one of claims 10 to 13, wherein the phonological acquisition unit acquires first phonological information by performing speech recognition on the speech signal.
15. The robot according to any one of claims 1 to 14, wherein the phonological generation unit generates the second phonological information consisting of a predetermined number of syllables or less, regardless of the number of syllables in the first phonological information.
16. The robot according to claim 15, wherein the phonological generation unit generates the second phonological information consisting of two syllables as the second phonological information consisting of a predetermined number of syllables or less.
17. The robot's computer, A phonological acquisition unit that acquires first phonological information consisting of multiple phonemes. A phonological generation unit that generates a second phonological information different from the first phonological information based on at least some of the phonologies included in the first phonological information, and A speech synthesis unit that synthesizes speech according to the second phonological information, A speech synthesis program that functions as such.
18. A method for outputting voice in a robot, A phonological acquisition step that acquires first phonological information consisting of multiple phonemes, A phonological generation step that generates a second phonological information different from the first phonological information based on at least some of the phonologies included in the first phonological information, A speech synthesis step that synthesizes speech according to the second phonological information, The audio output step includes outputting the aforementioned audio, A method of outputting audio, including the audio output method.
Citation Information
Patent Citations
Voice synthesis method and program
JP2018128690A