Robot and speech generation program
The robot's real-time voice generation based on sensor information and personality development addresses the issue of robots losing their lifelike quality, enhancing interaction and attachment through responsive and unique audio output.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GROOVE X INC
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-21
AI Technical Summary
Robots that output pre-recorded voices lose the illusion of being a living entity when interacted with over a long period, leading to a diminished user attachment.
A robot equipped with a voice generation unit that generates audio in real-time based on sensor information, allowing it to produce sounds that reflect external stimuli and internal states, including varying volume and pitch, and develop a unique personality through interaction.
Enhances the robot's lifelike interaction by generating responsive and distinctive voices, fostering a stronger user attachment and perception of the robot as a living being.
Smart Images

Figure 2026067976000001_ABST
Abstract
Description
Cross-reference to Related Applications
[0001] In this application, the benefits of Patent Application No. 2018-161616 filed in Japan on August 30, 2018 and Patent Application No. 2018-161617 filed in Japan on August 30, 2018 are claimed, and the contents of said applications are hereby incorporated by reference.
Technical Field
[0002] At least one embodiment relates to a robot that outputs voice and a voice generation program for generating the voice output by the robot.
Background Art
[0003] Conventionally, robots that output voice have been known (see, for example, Japanese Patent Application Laid-Open No. 2010-94799). Such robots are equipped with sensors, and when the robot receives some external stimulus, the sensor detects it and outputs a voice corresponding to the external stimulus. Alternatively, such a robot outputs voice according to internal information processing. Thereby, the user can obtain a feeling that the robot is a living thing.
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, when a robot outputs voice and plays back a fixed voice prepared in advance, if the user comes into contact with such a robot over a long period of time, the feeling that the robot is a living thing will be lost, and it will be difficult to form an attachment to the robot.
[0005] At least one embodiment aims to provide a robot that allows the user to feel more like a living thing in view of the above background.
Means for Solving the Problems
[0006] At least one embodiment is a robot comprising a voice generation unit that generates voice, and a generation The robot has a configuration that includes an audio output unit that outputs the aforementioned audio. The robot does not output pre-prepared audio, but rather outputs audio that it generates itself. [Brief explanation of the drawing]
[0007] [Figure 1A] Figure 1A is a front view of the robot. [Figure 1B] Figure 1B is a side view of the robot. [Figure 2] Figure 2 is a schematic cross-sectional view showing the structure of the robot. [Figure 3A] Figure 3A is a front view of the robot's neck. [Figure 3B] Figure 3B is a perspective view of the robot's neck, seen from the front and above. [Figure 3C] Figure 3C is a cross-sectional view of the robot's neck (AA). [Figure 3D] Figure 3D is a perspective view of the robot's neck, seen from diagonally above. [Figure 4] Figure 4 shows the hardware configuration of the robot. [Figure 5] Figure 5 is a block diagram showing the configuration for outputting sound in a robot. [Figure 6] Figure 6 is a table showing the relationship between intonation patterns and indices. [Figure 7] Figure 7 is a table showing the relationship between accent patterns and indices. [Figure 8] Figure 8 is a table showing the relationship between the duration pattern and the index. [Figure 9] Figure 9 is a table showing the relationship between vibrato patterns and indices. [Figure 10] Figure 10 is a block diagram showing the configuration of multiple robots capable of synchronized speech. [Figure 11]Figure 11 shows an example of a screen displaying the robot's status in a robot control application. [Figure 12A] Figure 12A shows an example of an app screen related to audio settings. [Figure 12B] Figure 12B shows an example of an application screen that appears in response to the operation of the application screen shown in Figure 12A. [Figure 12C] Figure 12C shows an example of an application screen that appears in response to the operation of the application screen shown in Figure 12A. [Figure 12D] Figure 12D shows an example of an app screen that appears when a user customizes the voice output from the robot. [Modes for carrying out the invention]
[0008] The embodiments described below are examples, and the present invention is not limited to the specific configurations described below. In implementing the present invention, specific configurations may be adopted as appropriate depending on the embodiment.
[0009] The robot of this embodiment has a configuration comprising a voice generation unit that generates sound and a voice output unit that outputs the generated sound.
[0010] This configuration allows the robot to output sounds it generates itself, rather than pre-recorded sounds. This enables it to generate and output sounds corresponding to sensor information, or even generate and output sounds unique to the robot, allowing users to perceive the robot as more lifelike. The timing of sound generation and sound output do not necessarily have to be different. In other words, the sound generated by the sound generation unit can be stored, and the stored sound can be output when certain conditions are met.
[0011] The robot according to the embodiment may include a sensor that detects a physical quantity and outputs sensor information, and the voice generation unit may generate a voice based on the sensor information.
[0012] With this configuration, since the voice is generated based on the sensor information, not only the external stimulus is detected as a physical quantity and the voice is output, but also the nature (e.g., magnitude) of the stimulus can be used to generate a corresponding voice.
[0013] The robot according to the embodiment may be such that the voice generation unit generates a voice when predetermined sensor information is continuously input over a predetermined period of time.
[0014] With this configuration, rather than simply reflecting the sensor information as it is in voice generation, flexible voice generation with respect to the sensor information becomes possible.
[0015] The robot according to the embodiment may include a plurality of sensors that detect a physical quantity and output sensor information, and the voice generation unit may generate a voice based on the sensor information of the plurality of sensors.
[0016] With this configuration, rather than simply generating a voice based on one piece of sensor information, generating a voice based on a plurality of pieces of sensor information enables more flexible voice generation.
[0017] The robot according to the embodiment may include a plurality of sensors that detect a physical quantity and output sensor information, and an interpretation unit that interprets the semantic situation in which the robot is placed based on the sensor information, and the voice generation unit may generate a voice based on the semantic situation interpreted by the interpretation unit.
[0018] With this configuration, not only is the sensor information simply reflected in voice generation, but the semantic situation is interpreted from the sensor information and a voice is generated based on the interpreted semantic situation, so that a voice showing a more lifelike reaction can be generated.
[0019] In the robot of this embodiment, the voice generation unit may generate a voice when the interpretation unit interprets that the robot is being held.
[0020] In the robot of this embodiment, the voice generation unit may generate voice that reflexively reflects the sensor information.
[0021] This configuration allows for the generation of sounds that reflexively reflect sensor information. For example, if the sensor is an impact-sensing sensor and the sound output unit outputs a sound in response to the impact, the sound generation unit can output a loud sound when the impact is large and a quiet sound when the impact is small. This makes it possible to output sounds of varying volume according to the stimulus, like a living organism. More specifically, if the robot is to output the sound "ouch" when hit, it is possible to create a robot that says "ouch" softly when hit lightly and loudly when hit hard.
[0022] In the robot of the embodiment, the voice generation unit may generate voice with a volume based on the sensor information.
[0023] This configuration allows, for example, if the sensor is an accelerometer, to generate a loud sound when a large acceleration is detected.
[0024] In the robot of the embodiment, the sensor may be an acceleration sensor that detects acceleration as the physical quantity and outputs acceleration as sensor information, and the sound generation unit may generate sound whose volume changes in accordance with the change in acceleration.
[0025] This configuration allows, for example, a robot to be vibrated, which then outputs sound whose volume increases or decreases according to the vibration period.
[0026] In the robot of this embodiment, the voice generation unit may generate voices of pitch based on the sensor information.
[0027] In the robot of this embodiment, the sensor may be an acceleration sensor that detects acceleration as the physical quantity and outputs acceleration as sensor information, and the sound generation unit may generate sound whose pitch changes in accordance with the change in acceleration.
[0028] This configuration allows, for example, a robot to vibrate, and in response to that vibration, it can output a vibrato-laden sound.
[0029] The program of this embodiment is a program for generating sound to be output from a robot, and causes a computer to perform a sound generation step of generating sound and a sound output step of outputting the generated sound.
[0030] Furthermore, the robot in this embodiment is equipped with a voice output unit that outputs a distinctive voice. This configuration allows the robot to emit a unique voice, making it easier for the user to perceive the robot as a living being. This, in turn, promotes the user's attachment to the robot.
[0031] In this embodiment, a distinctive voice refers to, for example, a voice that is distinguishable from other individuals and has identity within the same individual. Distinguishing voice means, for example, that even when multiple robots output voices with the same content, their spectral and prosodic features differ from one individual to another. Identical voice means, for example, that even when the same robot outputs voices with different content, the voice is recognized as belonging to the same individual.
[0032] The robot of this embodiment may further include a personality-forming unit that creates personality, and the voice output unit may output voice corresponding to the formed personality. With this configuration, personality is not fixedly given, but is formed during the process of use.
[0033] The robot of this embodiment may further include a voice generation unit that generates voices corresponding to the formed personality, and the voice output unit may output the generated voices. With this configuration, since the voice generation unit generates the voices, voices corresponding to the personality can be easily output.
[0034] In the robot of this embodiment, the voice generation unit may include a standard voice determination unit that determines a standard voice, and a voice adjustment unit that adjusts the determined standard voice to produce a unique voice. This configuration makes it easy to generate a unique voice.
[0035] The robot of this embodiment may further include a growth management unit that manages the growth of the robot, and the personality formation unit may form the personality according to the growth of the robot. With this configuration, personality is formed according to the growth of the robot.
[0036] The robot of this embodiment may further include an instruction receiving unit that receives instructions from a user, and the personality forming unit may form the personality based on the received instructions. With this configuration, the robot's personality can be formed based on user instructions.
[0037] The robot of this embodiment may further include a microphone that converts sound into electrical signals, and the personality-forming unit may form the personality based on the electrical signals. Based on the sound waves received, individuality is formed.
[0038] The robot in this embodiment may be equipped with a positioning device for measuring its position, and the personality-forming unit may form the personality based on the measured position. With this configuration, the personality is formed according to the robot's position.
[0039] In the robot of this embodiment, the personality-forming unit may determine the personality randomly.
[0040] Furthermore, in another embodiment of the robot, a voice generation unit is provided that generates sound by simulating the vocalization mechanism of a predetermined vocal organ, and a voice output unit outputs the generated sound. With this configuration, a unique voice can be generated and output by simulating the vocalization mechanism of the vocal organ.
[0041] The robot of this embodiment may further include sensors for acquiring external environmental information, and the voice generation unit may change the parameters used in the simulation based on the environmental information obtained from the sensors.
[0042] In the robot of this embodiment, the voice generation unit may change the parameters in conjunction with the environmental information obtained from the sensor.
[0043] The robot of the embodiment may further include an internal state management unit that changes its internal state based on environmental information obtained from the sensor, and the voice generation unit may change the parameters in conjunction with the internal state.
[0044] In the robot of this embodiment, the vocal organ may have a vocal cord organ that mimics the vocal cords, and parameters related to the vocal cord organ may be changed in conjunction with changes in the internal state.
[0045] In the robot of this embodiment, the vocal organ may have multiple organs, and the voice generation unit may use parameters that indicate the morphological state of each organ over time in its simulation.
[0046] The robot of the embodiment may further include a microphone that inputs sound output from another robot, and a comparison unit that compares the sound output from the other robot with its own sound, and the sound generation unit may change the parameters indicating the shape state so that the sound of the other robot and its own sound are different.
[0047] The robot of the embodiment further includes a voice generation condition recognition unit that recognizes voice conditions including a voice output start timing and at least some voice parameters, the voice generation unit may generate voice that matches a set of at least some voice parameters included in the voice conditions, and the voice output unit may output the voice generated by the voice generation unit at the voice output start timing included in the voice conditions.
[0048] In the robot of the embodiment, the voice generation unit may generate voices that match at least some of the voice parameters included in the voice conditions, as well as voices that are appropriate to the individuality of the robot.
[0049] In the robot of this embodiment, the voice generation condition recognition unit may recognize a first voice condition, which is its own voice condition, via communication, and the first voice condition may be shown to other robots. The conditions may be identical to at least a part of the second phonetic condition, which is a phonetic condition.
[0050] In the robot of this embodiment, some of the voice parameters may include a parameter indicating pitch.
[0051] In the robot of the embodiment, the first pitch included in the first sound condition may have a predetermined relationship with the second pitch included in the second sound condition.
[0052] In the robot of the embodiment, the first output start timing included in the first sound condition may be the same timing as the second output start timing included in the second sound condition, and the interval, which is the relative relationship between the first pitch included in the first sound condition and the second pitch included in the second sound condition, may be a consonant interval.
[0053] In the robot of the embodiment, the voice condition may include a condition indicating the length of the voice content, and the robot may further include a standard voice determination unit that randomly determines voice content that matches the length of the voice content.
[0054] In the robot of the embodiment, the voice condition may include a condition indicating the length of the voice content, and the robot may further include a standard voice determination unit that determines the voice content that matches the length of the voice content based on previously collected voices.
[0055] The program of the embodiment is a program for generating voice output from a robot, and causes a computer to perform the following steps: a personality formation step of forming the personality of the robot; a voice generation step of generating voice corresponding to the formed personality; and a step of outputting the generated voice.
[0056] Figure 1A is a front view of the robot, and Figure 1B is a side view of the robot. In this embodiment, robot 100 is an autonomous robot that determines its actions, gestures, and voice based on external environmental information and internal state. External environmental information is detected by a group of sensors including a camera, microphone, accelerometer, and touch sensor. The internal state is quantified as various parameters that express the emotions of robot 100.
[0057] As a parameter for expressing emotions, robot 100 has, for example, an intimacy parameter for each user. When robot 100 recognizes actions that show affection towards the user, such as picking them up or talking to them, the intimacy level with that user increases. On the other hand, the intimacy level will be lower for users who do not interact with robot 100, users who act rudely, or users who are rarely encountered.
[0058] The body 104 of the robot 100 has an overall rounded shape and includes an outer skin made of a soft, elastic material such as urethane, rubber, resin, or fiber. The weight of the robot 100 is 15 kg or less, preferably 10 kg or less, and more preferably 5 kg. The height of the robot 100 is 1.2 m or less, preferably 0.7 m or less. In particular, it is desirable to make the robot small and lightweight by making it weigh about 5 kg or less and keeping the height about 0.7 m or less, so that users, including children, can easily hold the robot.
[0059] Robot 100 is equipped with three wheels for three-wheeled movement. As shown in the figure, robot 100 includes a pair of front wheels 102 (left wheel 102a, right wheel 102b) and one rear wheel 103. The front wheels 102 are drive wheels, and the rear wheel 103 is a driven wheel. The front wheels 102 do not have a steering mechanism, but their rotational speed and direction can be controlled individually.
[0060] The rear wheels 103 are so-called omni-wheels or casters, and are rotatable to allow the robot 100 to move forward, backward, left, and right. By increasing the rotation speed of the right wheel 102b compared to the left wheel 102a, the robot 100 can turn left or rotate counterclockwise. Similarly, by increasing the rotation speed of the left wheel 102a compared to the right wheel 102b, the robot 100 can turn right or rotate clockwise.
[0061] The front wheels 102 and rear wheels 103 can be fully retracted into the body 104 by the drive mechanism (including the rotation mechanism and linkage mechanism). Even when moving, most of each wheel is hidden within the body 104, but when each wheel is fully retracted into the body 104, the robot 100 becomes immobile. That is, the retraction of the wheels causes the body 104 to lower and sit on the floor surface F. In this seated state, the flat seating surface 108 (ground contact bottom surface) formed on the bottom of the body 104 comes into contact with the floor surface F.
[0062] The robot 100 has two hands 105. The robot 100 is capable of actions such as raising, shaking, and vibrating the hands 105. The two hands 105 can be controlled individually.
[0063] The eye 106 is capable of displaying images using a liquid crystal element or an organic EL element. The robot 100 is equipped with various sensors, including a microphone capable of identifying the direction of a sound source, an ultrasonic sensor, a smell sensor, a distance sensor, and an acceleration sensor. The robot 100 also has a built-in speaker and can emit simple sounds of about 1 to 3 syllables. A capacitive touch sensor is installed on the body 104 of the robot 100. The touch sensor allows the robot 100 to detect user touches.
[0064] A horn 109 is attached to the head of robot 100. A 360-degree camera is attached to the horn 109, allowing it to capture images of the entire upper part of robot 100 at once.
[0065] Figure 2 is a schematic cross-sectional view showing the structure of the robot 100. As shown in Figure 2, the body 104 of the robot 100 includes a base frame 308, a main frame 310, a pair of resin wheel covers 312, and an outer shell 314. The base frame 308 is made of metal and forms the axis of the body 104 and supports the internal structure. The base frame 308 is constructed by connecting an upper plate 332 and a lower plate 334 vertically with a plurality of side plates 336. Sufficient spacing is provided between the plurality of side plates 336 to allow for ventilation. Inside the base frame 308 are a battery 117, a control circuit 342, and various actuators.
[0066] The main frame 310 is made of resin and includes a head frame 316 and a torso frame 318. The head frame 316 is hollow and hemispherical and forms the head skeleton of the robot 100. The torso frame 318 consists of a neck frame 3181, a chest frame 3182, and an abdominal frame 3183, and as a whole has a stepped cylindrical shape and forms the torso skeleton of the robot 100. The torso frame 318 is fixed integrally with the base frame 308. The head frame 316 is assembled to the upper end (neck frame 3181) of the torso frame 318 so as to be displaceable relative to it.
[0067] The head frame 316 is provided with three axes: a yaw axis 320, a pitch axis 322, and a roll axis 324, and actuators 326 for rotationally driving each axis. The actuators 326 include multiple servo motors for individually driving each axis. The yaw axis 320 is driven for head swinging motion, the pitch axis 322 is driven for nodding motion, and the roll axis 324 is driven for head tilting motion.
[0068] A plate 325 supporting the yaw axis 320 is fixed to the upper part of the head frame 316. Multiple ventilation holes 327 are formed in the plate 325 to ensure ventilation between the upper and lower parts.
[0069] A metal base plate 328 is provided to support the head frame 316 and its internal mechanism from below. The base plate 328 is connected to plate 325 via cross link 329 (pantograph mechanism), while also being connected to upper plate 332 (base frame 308) via joint 330.
[0070] The fuselage frame 318 houses the base frame 308 and the wheel drive mechanism 370. The wheel drive mechanism 370 includes a rotating shaft 378 and an actuator 379. The lower half of the fuselage frame 318 (abdominal frame 3813) is made narrow to form a housing space Sp for the front wheels 102 between it and the wheel cover 312.
[0071] The outer shell 314 is made of urethane rubber and covers the main frame 310 and wheel cover 312 from the outside. The handle 105 is integrally molded with the outer shell 314. An opening 390 for introducing outside air is provided at the upper end of the outer shell 314.
[0072] Figure 3A is a front view of the neck of robot 100. Figure 3B is a perspective view of the neck from the front and above. Figure 3C is a cross-sectional view of the neck at AA. Figure 3D is a perspective view of the neck from diagonally above. The neck of robot 100 is made up of a neck frame 3181 on which various components, including a circuit board, are mounted. A speaker 112 is provided on the neck frame 3181.
[0073] The speaker 112 is mounted facing upward on the front side of the neck frame 3181. That is, the diaphragm 1121 of the speaker 112 is mounted horizontally. A horn 1122 extending upward and forward is formed on the upper part of the diaphragm 1121, and the tip of the horn 1122 is open toward the front. The open surface of the horn 1122 corresponds to the position of the robot 100's mouth. Furthermore, the area of the open surface of the horn 1122 is formed to be approximately equal to the area of the diaphragm 1121. By providing the horn 1122, flexibility in the placement of the speaker 112 can be increased.
[0074] In this configuration, the sound waves generated by the vibration of the diaphragm 1121 and emitted upward are redirected forward by the horn 1122 and output. Therefore, to the user, it sounds as if the sound is coming from the mouth of the robot 100. In particular, when the robot 100 emits a low volume sound, it can be more clearly perceived that the sound is being emitted from the mouth. It is conceivable that the user may bring their ear closer to the mouth of the robot 100 to hear the sound clearly.
[0075] Figure 4 shows the hardware configuration of robot 100. Robot 100 includes a display device 110, an internal sensor 111, a speaker 112, a communication unit 113, a storage device 114, a processor 115, a drive mechanism 116, and a battery 117 within its housing 101. The drive mechanism 116 includes the wheel drive mechanism 370 described above. The processor 115 and the storage device 114 are included in the control circuit 342.
[0076] Each unit is connected to the others by power lines 120 and signal lines 122. Battery 117 supplies power to each unit via power lines 120. Each unit sends and receives control signals via signal lines 122. Battery 117 is, for example, a lithium-ion rechargeable battery and is the power source for robot 100.
[0077] The drive mechanism 116 is an actuator that controls the internal mechanisms. The drive mechanism 116 has the function of moving and changing the direction of the robot 100 by driving the front wheels 102 and rear wheels 103. The drive mechanism 116 also controls the hand 105 via wire 118 to perform actions such as raising the hand 105, shaking the hand 105, and vibrating the hand 105. The drive mechanism 116 also has the function of controlling the head to change the direction of the head.
[0078] The internal sensor 111 is a collection of various sensors built into the robot 100. Examples of internal sensors 111 include a camera (spherical camera), microphone, distance sensor (infrared sensor), thermal sensor, touch sensor, acceleration sensor, odor sensor, etc. The speaker 112 outputs sound.
[0079] The communication unit 113 is a communication module that performs wireless communication with various external devices such as servers, external sensors, other robots, and user-held mobile devices. The storage device 114 consists of non-volatile memory and volatile memory and stores various programs, including the voice generation program described later, and various setting information. The drive mechanism 116 is an actuator that controls the internal mechanism.
[0080] The display device 110 is installed at the position of the robot's eyes and has the function of displaying an image of the eyes. The display device 110 displays an image of the robot's eyes by combining eye parts such as pupils and eyelids. If external light shines into the eyes, a catchlight may be displayed at a position corresponding to the location of the external light source.
[0081] The processor 115 has the function of operating the robot 100 by controlling the drive mechanism 116, speaker 112, display device 110, etc., based on sensor information acquired by the internal sensor 111 and various information acquired through the communication unit 113. The robot 100 also has a clock (not shown) that manages the current date and time. The current date and time information is provided to each unit as needed.
[0082] Figure 5 is a block diagram showing the configuration for outputting sound in robot 100. Robot 100 includes an internal sensor 111, a positioning device 131, an instruction receiving unit 132, a growth management unit 133, a situation interpretation unit 134, a personality formation unit 135, a voice generation unit 136, and a speaker 112 as a voice output unit.
[0083] The voice generation unit 136 comprises a voice content database 1361, a standard voice determination unit 1362, a voice synthesis unit 1363, and a voice adjustment unit 1364. The growth management unit 133, situation interpretation unit 134, personality formation unit 135, standard voice determination unit 1362, voice synthesis unit 1363, and voice adjustment unit 1364 are software modules realized when the processor 115 executes the voice generation program of the embodiment.
[0084] Furthermore, the voice content database 1361 is composed of the storage device 114. The instruction receiving unit 132 receives instructions via communication, and the communication unit 113 is responsible for this. In this embodiment, the instruction receiving unit 132 specifically receives instructions from the user regarding the formation of personality in the personality formation unit 135.
[0085] The internal sensor 111 detects various physical quantities in the external environment of the robot 100 (i.e., external environmental information) and outputs sensor information indicating the environmental information (i.e., sensor detection values). The internal sensor 111 includes a touch sensor 1111, an acceleration sensor 1112, a camera 1113, and a microphone 1114. In Figure 5, the above-mentioned sensors are shown as sensors related to audio output in this embodiment, but audio may be output based on sensor information from other sensors as described above.
[0086] Furthermore, although only one touch sensor 1111 is shown in Figure 5, touch sensors 1111 may be provided on the back of the robot 100's head, face, right hand, left hand, abdomen, back, etc. The touch sensor 1111 is a capacitive touch sensor that detects when a user touches the corresponding part of the robot 100 and outputs sensor information indicating that contact has occurred.
[0087] Furthermore, although only one accelerometer 1112 is shown in Figure 5, it may include three accelerometers that detect acceleration in the vertical, horizontal, and forward / backward directions, respectively. These three accelerometers 1112 output acceleration in the vertical, horizontal, and forward / backward directions as sensor information. Since the accelerometer 1112 also detects gravitational acceleration, the posture (orientation) of the robot 100 when it is stationary and the direction of movement when the robot 100 is moving can be determined based on the accelerations of the three accelerometers 1112 in mutually orthogonal axial directions.
[0088] As described above, camera 1113 is mounted on the horn 109 and captures the entire upper area of robot 100 at once. Camera 1113 outputs the image obtained from the capture as sensor information. Microphone 1114 converts sound into an electrical signal and outputs this electrical signal as sensor information.
[0089] The situation interpretation unit 134 interprets the semantic situation in which the robot 100 is located based on sensor information from various sensors 1111 to 1114. To this end, the situation interpretation unit 134 stores sensor information output from the internal sensor 111 over a certain period of time.
[0090] For example, if the situation interpretation unit 134 detects that a touch is being made using the touch sensor 1111, and the acceleration sensor 1112 detects that the robot 100 has moved upward, and there is a gradual change in acceleration thereafter, the situation interpretation unit 134 interprets that the robot 100 is being held by the user.
[0091] In addition, the situation interpretation unit 134 can interpret, based on the sensor information from the touch sensor 1111, that the robot is being stroked by a user, and based on the sensor information from the microphone 1114, that the robot is being spoken to. Thus, semantic situation interpretation means, for example, not simply handling sensor information as is, but appropriately using various sensor information according to the posture, situation, and state of the robot 100 to be judged, thereby identifying the posture of the robot 100, identifying the situation in which the robot 100 is placed, and determining the state of the robot 100. The situation interpretation unit 134 outputs the interpreted content as an event so that it can be used in subsequent processing.
[0092] The situation interpretation unit 134 stores candidate semantic situations to be interpreted. Based on multiple sensor information, the situation interpretation unit 134 estimates a semantic situation from among several pre-prepared candidates. For this estimation, various sensor information may be used as input, and a lookup table, a decision tree, a support vector machine (SVM), a neural network, or other methods may be used.
[0093] Although not shown in Figure 5, the semantic situation interpreted by the situation interpretation unit 134 is also reflected in actions or gestures of the robot 100 other than voice. In other words, the robot 100 interprets the semantic situation from the physical quantities of the external environment detected by the internal sensor 111 and performs a reaction to the external environment. For example, if it is interpreted that the robot is being held, it will perform a control such as closing its eyes as a reaction. The voice output described in this embodiment is also one of these reactions to the external environment.
[0094] The growth management unit 133 manages the growth of the robot 100. Based on sensor information from the internal sensor 111, the robot 100 interprets the semantic situation in which it is placed and grows according to the content and number of experiences in performing reactions. This "growth" is expressed by growth parameters.
[0095] The growth management unit 133 updates and stores these growth parameters. The growth management unit 133 may manage multiple growth parameters. For example, the growth management unit 133 may manage growth parameters that represent emotional growth and physical growth, respectively. Physical growth refers to, for example, the speed at which movement occurs. For example, the unit may initially not output the maximum possible speed, and then increase the output speed as the user grows. The growth management unit 133 also stores the date and time when the power was turned on and manages the elapsed time from the power-on date and time to the present. The growth management unit 133 manages growth parameters in relation to the elapsed time. For example, growth parameters that represent emotional growth and physical growth, respectively, may be managed by the growth management unit 133.
[0096] The personality formation unit 135 forms the personality of the robot 100. The personality of the robot 100 is expressed by at least one personality parameter. The personality formation unit 135 forms the personality based on the situation (experience) interpreted by the situation interpretation unit 134 and the growth parameters managed by the growth management unit 133. To this end, the personality formation unit 135 accumulates the semantic situation interpreted by the situation interpretation unit 134 over a certain period of time.
[0097] In this embodiment, the robot 100 initially has no personality, i.e., at the time of power-on, and the personality parameters are the same for all robots 100. The robot 100 forms personality based on the semantic situation interpreted by the situation interpretation unit 134, and fixes the formed personality according to the growth parameters. Specifically, the personality formation unit 135 gradually changes the personality parameters from their initial values based on the accumulated semantic situation, and as the growth parameters are updated (grow), the changes in the personality parameters are reduced, and finally the personality parameters are fixed.
[0098] Here, "individuality" in this embodiment means, for example, having distinguishability from other individuals and having identity as an individual. That is, even if multiple individuals interpret the same semantic situation based on sensor information, if those multiple individuals react differently, then those multiple individuals (robots 100) can be said to have distinguishability. Furthermore, if there is commonality among multiple types of reactions in the same individual, then it can be said to have identity. However, regarding the requirement of distinguishability, it is permissible for there to be combinations of multiple individuals that have the same individuality with a sufficiently small probability.
[0099] The personality formation unit 135 updates and stores personality parameters that represent personality. The personality formation unit 135 may handle multiple types of personality parameters. In this embodiment, the personality formed by the personality formation unit 135 includes "voice quality". Other personality traits may include temperament (e.g., lonely, active, short-tempered, easygoing), physical abilities (e.g., maximum movement speed), etc.
[0100] Even if there is only one type of personality parameter, the personality represented by that parameter may be meaningless. Furthermore, personality parameters may be continuous, or multiple types may be provided as candidates, and the personality formation unit 135 may form a personality by selecting from among the candidates. For example, if personality is represented by one type of personality parameter, dozens or even hundreds of candidate personality parameters may be provided. This number of types is sufficient to achieve identifiability (i.e., the possibility of different individuals having the same personality when compared can be sufficiently reduced).
[0101] The personality formation unit 135 may form personality based on the position of the robot 100 as determined by the positioning device 131. For example, for the personality trait of "voice quality," the accent of the region may be used as the personality trait according to the robot 100's location (region). The personality formation unit 135 may also form (set) personality based on instructions from the instruction receiving unit 132.
[0102] The standard voice determination unit 1362 decides to generate voice and determines the content of the voice to be generated. Candidate voice content to be generated is stored as standard voice in the voice content database 1361. The standard voice determination unit 1362 determines the content of the voice to be output by selecting a standard voice from the voice content database 1361.
[0103] The standard voice determination unit 1362 determines the output and content of a voice according to the external environment and / or internal state. The voice generation unit 136 may generate a voice consciously or reflexively. Conscious voice generation means, for example, that the standard voice determination unit 1362 determines the output and content of a voice based on the internal state of the robot 100 and the semantic situation interpreted by the situation interpretation unit 134. For example, if the situation interpretation unit 134 interprets that the robot is being held, the voice generation unit 136 outputs a voice that expresses a feeling of happiness.
[0104] The standard voice determination unit 1362 consciously generates voices in response to internal states such as emotions, which change in response to the external environment obtained as multiple sensor values. For example, when a user speaks to it, the standard voice determination unit 1362 generates a voice in response. In addition, the standard voice determination unit 1362 generates voices triggered by changes in emotions (internal states) such as happiness, sadness, or fear, for example, when the user wants to get the user's attention or wants to express happiness with voice in addition to limb movements.
[0105] Reflexive speech generation refers to, for example, determining the generation and content of speech based on sensor information from the internal sensor 111. In conscious speech generation, the semantic situation is interpreted from the sensor information, or an internal state such as emotion changes, and speech is generated in response to such changes in semantic situation or internal state. In contrast, in reflexive speech generation, the sensor information is directly reflected in the generation of speech. An internal state management unit may be provided that changes the internal state based on sensor information indicating environmental information obtained from the internal sensor 111.
[0106] For example, in response to a large acceleration, the voice generation unit 136 outputs a sound that expresses a surprised reaction. Alternatively, it may output a predetermined sound in response to an acceleration of a predetermined value or higher that continues for a predetermined time or longer. A large acceleration would occur, for example, if the robot 100 is hit or collides with something. An acceleration of a predetermined value or higher that continues for a predetermined time or longer would occur, for example, if the robot is swung around violently or falls from a height. Furthermore, generating sound that directly reflects sensor information such as the sound pressure of a sound detected by a microphone or the intensity (brightness) of light detected by an illuminance sensor also constitutes reflexive voice generation.
[0107] In this type of reflexive voice generation, the voice is generated to directly reflect the sensor information, resulting in minimal delay and enabling the generation of voice that corresponds to the stimulus received by the robot 100. For this type of reflexive voice generation, the sensor value of each sensor may be used as a trigger when a predetermined condition is met (for example, when it exceeds a predetermined value).
[0108] Furthermore, the standard voice determination unit 1362 may decide to output voice in response to an action, such as an action or gesture, when the action is performed based on internal information processing, rather than as a reaction to the external environment. The robot 100 may determine a standard voice by deciding to output a corresponding voice when it tenses up or when it releases tension. Conversely, when the standard voice determination unit 1362 decides to output a voice, it may move its hands or head in accordance with the output of that voice.
[0109] Furthermore, the standard voice determination unit 1362 may decide to output a voice based on the combination of the semantic situation interpreted by the situation interpretation unit 134 and the sensor information, and determine the corresponding standard voice. For example, if the standard voice determination unit 1362 interprets that the child is being held on someone's lap, it may decide to output a corresponding voice when the child is being rocked up and down for a certain period of time. The rocking state can be determined by focusing on the temporal change in the value of the acceleration sensor. The waveform of the sensor value may be pattern-matched, or the determination may be based on machine learning.
[0110] In this way, by considering sensor values over a certain period, it is possible to change the time it takes to generate a voice based on, for example, the level of intimacy with the user holding the baby on their lap. For example, if there is past experience where generating a voice in the same situation pleased the user, the voice can be generated in a shorter time based on that experience. In this way, the time from when the external environment or internal state changes until the voice is generated can be shortened.
[0111] The voice content database 1361 stores standard voices that are consciously output in response to semantic situations interpreted by the internal state and situation interpretation unit 134, and standard voices that are reflexively output in response to sensor information from the internal sensor 111. These voices are simple voices of about 1 to 3 syllables. For example, the voice content database 1361 stores the standard voice "Ahhh" which expresses a pleasant feeling in response to the situation of being held, and the standard voice "It hurts" which is uttered reflexively in response to a large acceleration. Thus, the content of the voices may be interjections, sounds like humming, nouns, or adjectives.
[0112] In this embodiment, the audio content database 1361 does not store standard audio as audio data such as wav or mp3, but rather stores a set of parameters for generating audio. This set of parameters is output to the audio synthesis unit 1363, which will be described later. The audio synthesis unit 1363 uses the set of parameters to adjust the synthesizer and generate audio. Alternatively, instead of this embodiment, the audio content database 1361 may be used to store basic audio data in a format such as wav beforehand, and the audio adjustment unit 1364 may be used to make adjustments to it.
[0113] The speech synthesis unit 1363 is composed of a synthesizer and is implemented, for example, by software. The speech synthesis unit 1363 reads the set of standard speech parameters determined by the standard speech determination unit 1362 from the speech content database 1361 and synthesizes speech using the read set of parameters.
[0114] The voice adjustment unit 1364 adjusts the voice synthesized by the voice synthesis unit 1363 based on personality parameters stored in the personality formation unit 135. In particular, the voice adjustment unit 1364 adjusts the voice so that it can be recognized as being emitted by the same individual, regardless of the standard voice. The voice adjustment unit 1364 also changes the tone of voice according to the level of familiarity with the user it is communicating with.
[0115] The voice adjustment unit 1364 performs voice conversion, which adjusts non-verbal information such as voice quality and prosody without changing the linguistic (phonological) information contained in the standard voices, in order to adjust multiple standard voices according to individuality parameters.
[0116] Individuality in speech depends on the characteristics that appear in the spectrum and prosody of the speech. In living organisms, spectral characteristics are determined by the characteristics of the individual's articulatory organs, i.e., physical characteristics such as the shape of the vocal cords and vocal tract, and mainly manifest as differences in the individual's voice quality. On the other hand, prosodic characteristics manifest as differences in intonation, accent of each syllable, duration of each syllable, and vibrato of each syllable. Therefore, in order to achieve voice quality conversion, the speech adjustment unit 1364 converts the spectral and prosodic characteristics of the standard speech according to the individuality formed in the individuality formation unit 135.
[0117] First, let's explain the spectral features. In voice quality conversion from a standard voice to the voice of the individual in question, let xt be the spectral features of the standard voice at time t (e.g., Mel-cepstrum coefficient vector or line spectral frequency vector, etc.), and let yt be the spectral features of the individual's voice converted from xt. Then, voice quality conversion focusing on the conversion of spectral features is as follows. That is, the conversion function yt = Fs(xt) that converts the spectral features of the standard voice to spectral features according to the individual's personality is expressed by equation (1) below.
number
[0118] The personality formation unit 135 initially stores initial values for voice quality parameters, and gradually changes these initial values as the robot grows. After a certain period of time, the amount of change is gradually reduced. In other words, the personality formation unit 135 changes the voice quality parameters over time. This allows us to represent how the voice quality changes as the robot grows and how it stabilizes at a certain point. Furthermore, during this voice quality transformation process, the personality formation unit 135 creates differences between the voice quality of one robot and those of other robots. Hereinafter, the period until the voice quality stabilizes will be referred to as the "transformation period". The initial value of the transformation matrix Ai is the identity matrix, the initial value of the bias bi is the zero vector, and the initial value of the weight coefficient wi is the unit vector.
[0119] The conversion in equation (1) in the voice adjustment unit 1364 makes it possible to change voice quality, such as high / low pitch, and filtered voices (clear voice, hoarse voice, etc.). This voice quality conversion makes it possible to generate unique voices.
[0120] Next, we will explain the prosodic features. There are various methods for transforming prosodic features according to the individual characteristics of each person. In this embodiment, prosodic features are represented by intonation, accent of each syllable, duration of each syllable, index of vibrato of each syllable, and volume, speech rate level, and pitch compression level.
[0121] Figure 6 is a table showing the relationship between intonation patterns and indices. For intonation, there are intonation patterns consisting of low, medium, and high combinations for one-syllable, two-syllable, and three-syllable words, and each pattern for each number of syllables is assigned an index Ii1, Ii2, and Ii3.
[0122] Figure 7 is a table showing the relationship between accent patterns and the index. Regarding accent, it also shows weak, medium, and strong accents for one-syllable, two-syllable, and three-syllable words. Accent patterns consisting of combinations of these elements are provided, and each pattern for each number of syllables is assigned an index Ia1, Ia2, and Ia3.
[0123] Figure 8 is a table showing the relationship between duration patterns and indices. For duration, there are duration patterns consisting of short, medium, and long combinations for one-syllable, two-syllable, and three-syllable cases, and each pattern for each number of syllables is assigned an index Il1, Il2, and Il3.
[0124] Figure 9 is a table showing the relationship between vibrato patterns and indices. For vibrato, there are vibrato patterns available for one-syllable, two-syllable, and three-syllable words, each consisting of combinations of vibrato with and without it, and each pattern for each number of syllables is assigned indices Iv1, Iv2, and Iv3.
[0125] In this embodiment, volume V, speech rate level S, and pitch compression level C are further provided as prosodic features. Volume V is the loudness (volume) of the voice. Speech rate level S is the level at which the pronunciation time of the standard voice is compressed. Pitch compression level C is the level at which the pitch difference of the standard voice is reduced. In this way, by changing the combination of prosodic parameters such as intonation, accent, duration, vibrato, volume, speech rate level, and pitch compression level for each individual, it is possible to generate unique voices.
[0126] Other methods for transforming prosodic features include vector quantization-based methods, methods that simply match the average values of the fundamental frequency F0 and speech rate to the average values of the individual, methods that linearly transform the fundamental frequency F0 while considering variance, and methods based on prosody generation using Hidden Markov Model (HMM) speech synthesis. It is also possible to transform both spectral and prosodic features based on HMM speech synthesis and speaker adaptation.
[0127] The individuality formation unit 135 stores the above-mentioned voice quality parameters (Ai, bi, wi) and prosodic parameters (Ii, Ia1, Ia2, Ia3, Il1, Il2, Il3, Iv1, Iv2, Iv3, V, S, C) as individuality parameters related to voice (hereinafter referred to as "individual voice parameters"). The voice adjustment unit 1364 reads the individuality voice parameters from the individuality formation unit 135 and uses them to convert the standard voice determined by the standard voice determination unit 1362. The individuality formation unit 135 may also store different prosodic parameters for each standard voice stored in the voice content database 1361. By changing these voice quality parameters and prosodic parameters for each individual, individual voices can be generated.
[0128] The method for determining the individual speech parameters in the personality formation unit 135 will be explained. The prosodic parameters may be determined based on the individual's personality. As mentioned above, the personality in the personality formation unit 135 changes with experience and growth, so the prosodic parameters may also change during the personality formation process. When determining the corresponding prosodic parameters from the personality, a conversion function may be used, or a learning model learned by machine learning may be used.
[0129] Furthermore, the prosodic parameters may be determined based on the individual's region. In this case, as described above, the individuality formation unit 135 acquires location information from the positioning device 131, and determines the prosodic parameters to correspond to the relevant region based on this location information. This makes it possible to adjust the voice according to the region, such as a Kansai accent or a Tohoku accent. When determining prosodic parameters from location information, a lookup table may be used.
[0130] Furthermore, the unique voice parameters may be determined based on instructions received by the instruction receiving unit 132. For example, the personality formation unit 135 may receive instructions from the instruction receiving unit 132 specifying gender. If accepted, the voice quality parameters and prosodic parameters will be determined to correspond to the specified gender.
[0131] To issue these instructions, an information terminal (e.g., a smartphone or personal computer) with a control application installed may be used. This information terminal receives instructions using the control application as a user interface and transmits the received instructions to the robot 100 via a network or relay server as needed. The instruction receiving unit 132 of the robot 100 can receive the instructions transmitted in this manner.
[0132] The above describes how the voice adjustment unit 1364 adjusts the standard voice based on the unique voice parameters stored in the personality formation unit 135. The voice adjustment unit 1364 further adjusts the prosody based on sensor information from the internal sensor 111. For example, if the acceleration from the acceleration sensor 1112 vibrates at a frequency higher than a predetermined frequency, the voice adjustment unit 1364 vibratos the voice in accordance with the vibration. In other words, it periodically changes the pitch. It may also increase the volume or raise the pitch depending on the magnitude of the acceleration from the acceleration sensor 1112.
[0133] Adjusting the voice in this way to directly reflect sensor information can be likened to adjusting the voice to reflexively reflect stimuli from the external environment, for example, in the case of a living organism. In other words, the voice adjustment unit 1364 adjusts the voice based on individual characteristics, and also adjusts the voice as a reflex to stimuli when a predetermined external environment stimulus is applied.
[0134] The speech generation unit 136 generates individualistic sounds, reflexive sounds, and conscious sounds. In other words, even if the sounds generated by the speech generation unit 136 are the same sound, the basis for generation, i.e., the triggers and parameters for sound generation, are different. It is also possible to provide separate speech generation units for generating reflexive sounds and for generating conscious sounds.
[0135] <Generating unique voices> The voice generation unit 136 generates a distinctive voice with discriminative properties. If there are multiple robots 100 in a household, the voice of each robot 100 is generated so that they do not all sound the same. To achieve this, robot 100 captures the voices of other robots with a microphone 1114 and generates a voice that is different from them. The voice generation unit 136 has a comparison unit (not shown) that captures the voices of other robots with the microphone 1114 and compares them with its own voice to determine whether or not there is a difference in voice. If the comparison unit determines that there is no difference from other voices, i.e., that there is no discriminative property, the personality formation unit 135 changes the distinctive voice parameters. The voice parameters are changed in response to the voices of other robots only for robots in the transformation period; for robots that have passed the transformation period, the voice parameters are not changed. In this way, by listening to the voices of other robots and creating a different voice, the voices are the same when the power is turned on, but individual differences become clear after a certain period of time has passed.
[0136] <Reflexive voice generation> Humans sometimes unconsciously and reflexively produce sounds when subjected to certain external stimuli. For example, when experiencing pain or surprise. The reflexive sound generation in the sound generation unit 136 mimics such unconscious vocalizations. The sound generation unit 136 generates a reflexive sound when certain sensor information changes rapidly. To achieve this, sensor values from the accelerometer 1112, microphone 1114, and touch sensor 1111 are used as triggers. Furthermore, sensor values are also used to determine the volume of the generated sound.
[0137] For example, if the voice generation unit 136 is making a sound and is suddenly lifted, it will stop the current sound and generate a surprised sound like "Wow!". Another example is that when the voice generation unit 136 is suddenly lifted, it increases the volume of the current sound in conjunction with the magnitude of the acceleration. In this way, in reflexive voice generation, the sensor value is directly linked to the sound. Specifically, if the user is rhythmically making sounds like "run, run, run, run," and is suddenly lifted around the third "run," the volume of "run" will suddenly increase, the rhythm will break, and the user will shout.
[0138] <Conscious voice generation> Conscious voice generation produces sounds that express emotions. Just as the atmosphere changes when the background music changes in a movie or drama, sound in Robot 100 can be considered a part of the overall presentation. In Robot 100, emotions change like waves. That is, emotions constantly change in response to stimuli from the external environment. Because emotions change like waves, the parameter value indicating emotion will reach its peak in a given situation and then gradually decrease over time.
[0139] For example, if robot 100 is surrounded by users and detects many smiles and laughter through the camera 1113 and microphone 1114, a wave of "joy" emotion will rise within robot 100. At that moment, the voice generation unit 136 selects parameters for joy from the voice content database 1361 and generates speech. For example, if robot 100 is being held on someone's lap and rocked up and down, robot 100 interprets this and experiences a wave of "happiness," generating speech that expresses happiness. Furthermore, because the acceleration sensor value shows a periodic fluctuation, the robot begins to add vibrato to its speech that matches the period.
[0140] Sensor information from internal sensors 111, such as the acceleration sensor 1112, not only triggers reflexive voice generation but also influences the voice even while it is speaking. This influence may be based on quantitative conditions, such as the sensor information exceeding a predetermined value, or it may vary qualitatively depending on the internal state, such as emotions at the time. In this way, by generating voice in the voice generation unit 136 based on sensor information, the robot 100 does not always generate the same voice, but rather generates voice by comprehensively reflecting the sensor information and internal state.
[0141] Speaker 112 outputs sound adjusted by the sound adjustment unit 1364.
[0142] As described above, the robot 100 of this embodiment does not output pre-prepared sound sources as they are, but rather generates and outputs sound using the sound generation unit 136, thus enabling more flexible sound output.
[0143] Specifically, robot 100 can output a unique voice that is distinctive and identical. This allows users to distinguish and recognize their own robot 100 from other robots 100 by listening to its voice, effectively promoting the formation of attachment between users and robot 100. Furthermore, robot 100 can generate voices based on sensor information. This allows it to output reflexively adjusted voices, such as vibrating in response to vibrations.
[0144] In the above embodiment, the robots 100 initially lack individuality (in particular, they do not possess identifiability), and the spectral and prosodic characteristics of the voice are the same for all robots 100, with individuality being formed during the process of use. However, instead, each robot may initially have a different individuality.
[0145] Furthermore, in the above embodiment, the user could specify the robot's personality through the control application. However, instead of this, or in addition to this, the system may be configured so that the user can also instruct the robot 100 to cancel its personality and return it to its initial value through the control application.
[0146] Alternatively, the robot 100 may initially lack a defined personality, but when it is first powered on, the personality formation unit 135 may randomly determine a personality. Furthermore, the personality of the robot 100 may be visualized through a control application. In this case, the robot 100 transmits personality parameters through the communication unit 113, and the personality parameters are received and displayed on an information terminal where the control application is installed.
[0147] Furthermore, in the above embodiment, the personality formation unit 135 formed personality based on the situation interpreted by the situation interpretation unit 134 and the growth parameters managed by the growth management unit 133. However, instead of this, or in addition to this, the personality formation unit 135 may analyze the user's voice detected by the microphone 1114, acquire the spectral and prosodic features of the user's voice, and determine unique voice parameters to approximate the user's voice. This makes it possible to create the effect that the robot 100's voice approximates the user's voice.
[0148] Furthermore, the robot 100 may learn the user's reaction when it outputs sound and reflect it in the personality formation in the personality formation unit 135. For example, the user's reaction can be detected by performing image recognition on the image from the camera 1113 to detect that the user is smiling, or by detecting that the user is stroking the robot 100 based on the sensor information from the touch sensor 1111.
[0149] Furthermore, in the above embodiment, the standard voice determination unit 1362 determines the content of the voice to be output, and then the voice adjustment unit 1364 adjusts the standard voice before outputting it from the speaker 112. However, instead, the standard voice stored in the voice content database 1361 may be pre-adjusted based on the personality formed by the personality formation unit 135 and stored in the voice content database 1361. In other words, voice generation may be performed in advance, rather than immediately before voice output.
[0150] In this case, when the standard voice determination unit 1362 determines whether to output voice and the content of that voice, it may output an adjusted voice corresponding to that content from the speaker 112. In this case as well, the voice adjustment unit 1364 may perform reflective adjustment of the voice based on sensor information.
[0151] Furthermore, in the above embodiment, the growth management unit 133, situation interpretation unit 134, personality formation unit 135, and voice generation unit 136 were all provided on the robot 100, but some or all of these may be provided on a device separate from the robot 100 that can communicate with the robot 100. Such a device may communicate with the robot 100 via short-range communication such as Wi-Fi (registered trademark), or via a wide-area network such as the Internet.
[0152] <Robot with virtual vocal organs> Generally, the vocalization process is common to all organisms with vocal organs. For example, in humans, the vocalization process involves air from the lungs and abdomen passing through the trachea, vibrating the vocal cords to produce sound, which then resonates in the oral and nasal cavities, resulting in a louder sound. The shape of the mouth and tongue then changes, producing a variety of voices. Individual differences in voice arise from various factors such as body size, lung capacity, vocal cords, trachea length, oral cavity size, nasal cavity size, tooth alignment, and tongue movement. Furthermore, even in the same person, the condition of the trachea and vocal cords changes depending on their physical condition, which in turn changes their voice. Due to this vocalization process, each person has a different voice quality, and their voice also changes depending on their internal state, such as their physical condition and emotions.
[0153] In another embodiment, the speech synthesis unit 1363 generates speech by simulating the speech process in a virtual vocal organ based on this speech process. In other words, the speech synthesis unit 1363 is a virtual vocal organ (hereinafter referred to as "virtual vocal organ"), and generates voice using a virtual vocal organ implemented in software. For example, the virtual vocal organ may have a structure that mimics the vocal organs of a human, or it may have a structure that mimics the vocal organs of an animal such as a dog or cat. By having a virtual vocal organ, it is possible to generate individual-specific voices even if the basic structure of the vocal organs is the same, by changing the size of the trachea in the virtual vocal organ, adjusting the tension of the vocal cords, or changing the size of the oral cavity for each individual. The set of parameters for generating speech held in the speech content database 1361 does not simply include direct parameters for generating sound with a synthesizer, but also includes values that specify the structural characteristics of each organ in the virtual vocal organ as parameters (hereinafter referred to as "static parameters"). Using these static parameters, the speech process is simulated and a voice is generated.
[0154] For example, humans can produce a variety of voices. High-pitched voices, low-pitched voices, singing along to melodies, laughing, shouting—they can produce virtually any voice as long as the structure of their vocal organs allows. This is because the shape and state of each organ that makes up the vocal organs change, and these can be consciously altered or unconsciously altered in response to emotions and stimuli. The speech synthesis unit 1363 also has parameters (hereinafter referred to as "dynamic parameters") for the state of these organs that change in conjunction with the external environment and internal state, and performs simulations by changing these dynamic parameters in conjunction with the external environment and internal state.
[0155] Generally, stretching the vocal cords produces a higher pitch, while relaxing them causes them to contract, resulting in a lower pitch. For example, an organ that mimics the vocal cords has a static parameter called the degree of vocal cord stretching (hereinafter referred to as "tension"), and by adjusting the tension, it is possible to produce high-pitched or low-pitched voices. This makes it possible to create robots 100 with high-pitched voices and robots 100 with low-pitched voices. Also, just as a person's voice may become high-pitched when they are nervous, similarly, by changing the tension of the vocal cords as a dynamic parameter in conjunction with the tension state of robot 100, it is possible to make robot 100's voice higher when it is nervous. For example, when robot 100 recognizes a stranger or is suddenly lowered from being held, etc., when the internal parameter indicating the state of nervousness swings to a tense value, the tension of the vocal cords is increased in conjunction with this, allowing it to produce a high-pitched voice. In this way, by associating the internal state of robot 100 with the organs involved in the vocalization process, and adjusting the parameters of the related organs according to the internal state, it is possible to change the voice according to the internal state.
[0156] Here, static and dynamic parameters are parameters that describe the morphological state of each organ over time. The virtual vocal organs are simulated based on these parameters.
[0157] Furthermore, by generating voices based on simulations, only voices that are based on the structural constraints of the vocal organs are generated. In other words, voices that are impossible for living beings are not generated, so it is possible to generate voices that sound natural and lifelike.
[0158] <Synchronized speech by multiple robots> Figure 10 is a block diagram showing the configuration of multiple robots capable of synchronous speech. In the example in Figure 10, robot 100A and robot 100B perform synchronous speech. Robots 100A and 100B have the same configuration. Robots 100A and 100B are equipped with instruction receiving units 132A, 132B, personality formation units 135A, 135B, voice generation units 136A, 136B, and speakers 112A, 112B, similar to the embodiment described above. Robots 100A and 100B are also equipped with an internal sensor 111, a positioning device 131, a growth management unit 133, and a situation interpretation unit 134, similar to the embodiments described above, although these are omitted from the illustration in Figure 10.
[0159] As described above, the instruction receiving units 132A and 132B correspond to the communication unit 113 (see Figure 4) that performs wireless communication. In this embodiment, the instruction receiving units 132A and 132B can communicate wirelessly with each other. In this embodiment, in order to achieve synchronized speech between the two robots 100A and 100B, one robot determines the content of the speech generated by the standard speech determination unit 1362, and also determines speech conditions consisting of a set of speech output start timings for itself and the other robot and at least some speech parameters, and the other robot outputs speech according to the speech conditions determined by the other robot. In this embodiment, an example is described in which robot 100A determines the speech conditions and robot 100B outputs speech according to those speech conditions.
[0160] The standard voice determination unit 1362A of robot 100A determines the content of the voice generated by robot 100A, and also determines voice conditions including the output start timing and at least some voice parameters for itself and robot 100B. That is, robot 100A determines voice conditions (second voice conditions) including the output start timing (second output start timing) for robot 100A and voice conditions (first voice conditions) including the output start timing (first output start timing) for robot B. The instruction receiving unit 132A of robot 100A transmits the first voice conditions to robot 100B.
[0161] The instruction receiving unit 132B of robot 100B receives a first voice condition from robot 100A, and the standard voice determination unit 1362B of robot 100B recognizes at least some of the voice parameters included in the received first voice condition. In addition, the synchronization control unit 1365B of robot 100B recognizes the first output start timing included in the received first voice condition. The standard voice determination unit 1362B and the synchronization control unit 1365B correspond to the voice generation condition recognition unit.
[0162] The voice adjustment unit 1364A of robot 100A generates a voice that matches at least some of the voice parameters included in the second voice condition determined by the standard voice determination unit 1362A. The voice adjustment unit 1364B of robot 100B generates a voice that matches at least some of the set of voice parameters included in the first voice condition recognized by the standard voice determination unit 1362B.
[0163] The synchronous control unit 1365A of robot 100A outputs the sound generated by the sound adjustment unit 1364 to the speaker 112A according to the second output start timing included in the second sound condition determined by the standard sound determination unit 1362A. The synchronous control unit 1365B of robot 100B outputs the sound generated by the sound adjustment unit 1364 to the speaker 112B according to the first output start timing included in the first sound condition.
[0164] Some of the speech parameters determined by the standard speech determination unit 1362A as speech conditions are, for example, tempo expressed in BPM, rhythm, pitch, length of speech content (e.g., number of syllables), timbre, volume, or the time-series change pattern of at least one of these elements. The speech adjustment units 1364A and 1364B, respectively, convert the spectral and prosodic features of the standard speech according to the individuality formed by the individuality formation unit 135, similar to the embodiment described above, thereby adjusting the standard speech determined by the standard speech determination units 1362A and 1362B according to the individuality parameters formed by the individuality formation units 135A and 135B, and performing voice quality conversion that adjusts non-verbal information such as voice quality and prosody without changing the linguistic (phonological) information contained in the standard speech. At this time, the speech adjustment units 1364A and 1364B convert the speech parameters specified in the speech conditions (e.g., tempo, rhythm, pitch, number of syllables, timbre, volume, or the time-series change patterns of these, etc.) For ), adjust according to the voice conditions, and adjust other voice parameters according to the personality parameters.
[0165] Specifically, the standard voice determination unit 1362A can determine the output start timing for robot 100A and robot 100B so that they output voices at the same time. This makes it possible to make robot 100A and robot 100B speak at the same time.
[0166] Alternatively, the standard voice determination unit 1362A may determine the respective output start timings (first output start timing and second output start timing) for robot 100A and robot 100B so that robots 100A and 100B output voice at predetermined time intervals. For example, the voice output start timings for each robot may be determined so that the other robot outputs voice when the voice output of the other robot ends.
[0167] Furthermore, the standard voice determination unit 1362A can, specifically, determine the pitch (second pitch) as a part of the voice parameters of robot 100A and the pitch (first pitch) as a part of the voice parameters of robot 100B, respectively, so that robot 100A and robot B output sounds of the same pitch. Alternatively, the standard voice determination unit 1362A may determine the first and second pitches so that robot 100A and robot 100B output sounds of different pitches. In this case, the first and second pitches may have a predetermined relationship.
[0168] For example, the speech parameters may be determined such that the ratio of the second pitch (frequency) to the first pitch (frequency) falls within a predetermined range. For example, the intervals (ratios of pitches) may have a relationship that results in a consonant interval. By outputting two consonant sounds at the same time, a harmony can be created. The consonant interval may be an imperfect consonant, a perfect consonant, or an absolute consonant. Furthermore, if the immaturity of robots 100A and 100B is to be expressed, dissonant intervals may be used.
[0169] For example, the speech parameters may be determined such that the second pitch (frequency) is higher or lower than the first pitch (frequency) by a predetermined interval (e.g., a third) or a predetermined frequency.
[0170] Alternatively, the pitch may not be specified in the voice parameters, and robots 100A and 100B may each determine the pitch and generate the sound. In this case, for example, the following processing may be performed. For example, suppose the same tempo is specified by the voice parameters. Before outputting the sound, robot 100A generates the sound to match the specified tempo. Robot 100A transmits the time-series change of the pitch to be output to robot 100B via communication. Robot 100B may generate the sound it outputs so that the ratio of the pitch to be output at the same timing to the pitch of the sound generated by robot 100A (pitch) falls within a predetermined range, and so that it matches the specified tempo. In this case, robot 100B may generate a list of pitches to be output at the same timing based on the time-series change of the pitch of the sound generated by robot 100A, so that the ratio of the pitches to be output at the same timing falls within a predetermined range, and select the pitch to be output at the time from the list.
[0171] Furthermore, robots 100A and 100B independently generate speech to match the speech parameters, share the generated speech via communication before outputting it, and determine whether the ratio of the pitches of the speech falls within a predetermined range at each timing of the time-series changes in the speech. If it does not fall within the predetermined range, the ratio of the pitches falls within the predetermined range. One pitch may be corrected to be higher or lower by a predetermined interval (e.g., one octave) or a predetermined frequency so that it is included. The same process may be performed by changing the condition that the ratio of pitches falls within a predetermined range to whether the difference in frequencies falls within a predetermined range. The predetermined range may be, for example, a range specified by both a lower limit and an upper limit, a range specified by only a lower limit or an upper limit, a continuous range, or an intermittent range.
[0172] Furthermore, if some of the speech parameters determined by the standard speech determination unit 1362A of robot 100A are syllables, the number of syllables of robot 100A and robot 100B may be the same. Also, if some of the speech parameters are syllables, the content of the speech may be determined randomly according to the number of syllables.
[0173] The standard voice determination unit 1362A determines the content of the voice to be spoken by at least robot 100A. The standard voice determination unit 1362A may also determine the content of the voice to be spoken by robot 100B as some voice parameters. In this case, the content of the voice of robot 100A and the content of the voice of robot 100B may be the same. In this case, robot 100A may determine the content of the voice randomly. Similarly, when robot 100B determines the content of the voice, robot 100B may also determine the content of the voice randomly.
[0174] Alternatively, robots 100A and 100B may perform a predetermined task in sync by having them output sounds at different times and specifying the content of those sounds. For example, by determining the sound conditions such that robot 100A outputs the sound "Janken" and robot 100B outputs the sound "Pon" at the moment robot 100A finishes outputting that sound, robots 100A and 100B can complete the Janken call together.
[0175] If the standard voice determination unit 1362A does not determine the content of the robot 100B's voice, the robot 100B's standard voice determination unit 1362B determines the content of the voice itself. In this case, if the number of syllables is included as some of the voice parameters, the standard voice determination unit 1362B may determine the content of the voice that matches the number of syllables according to previously collected voice (for example, the user's voice). That is, the content of the voice may be learned from voices picked up by the microphone in order to reproduce the content of voices that the user frequently uses, or a part thereof. In addition, if the user's voice satisfies predetermined conditions (such as the pitch being within a certain range, or a case where there is a high probability that the user is singing), the standard voice determination unit 1362A or the standard voice determination unit 1362B may determine the content of the voice in order to reproduce the content of voices that have been collected, or a part thereof.
[0176] In addition to reproducing the content of the voice, the personality formation unit 135 may form a personality based on the pitch and intonation of the collected user's voice to reproduce that pitch and intonation, and then generate or correct the voice according to that personality.
[0177] In the above explanation, robot 100A determined the voice conditions, and robot 100B generated and output voice according to the voice conditions determined by robot 100A. However, it is also possible for robot 100B to determine the voice conditions, in which case robot 100A generates and outputs voice according to the voice conditions determined by robot 100B.
[0178] In the example above, when there are multiple robots, one robot determines the voice conditions and transmits them to the other robots. Alternatively, a control device that allows multiple robots to communicate may determine the voice conditions for each robot and transmit them to each robot. In this case, the control device determines and transmits common voice conditions for multiple robots. Alternatively, different voice conditions may be determined for each robot, and these conditions may be sent individually.
[0179] <User-defined personalization> As described above, the personality formation unit 135 may form (set) a personality for voice based on instructions from the instruction receiving unit 132, and below, an example of setting personality parameters based on user instructions will be described. In this case, an information terminal (e.g., a smartphone or personal computer) with a control application installed may be used as the device that gives instructions to the instruction receiving unit 132. This information terminal receives instructions using the control application as a user interface and transmits the received instructions to the robot 100 via a network or relay server as needed. The instruction receiving unit 132 of the robot 100 can receive instructions that have been transmitted in this manner.
[0180] Figures 11, 12A to 12C, and 13 are examples of screens (hereinafter referred to as "app screens") of the control application related to audio settings displayed on the terminal device 204 of this embodiment. Note that the configuration of the app images and the method of setting audio using the app screen shown in Figures 11, 12A to 12C, and 13 are examples only and are not limited thereto.
[0181] Figure 11 shows an example of a screen displaying the robot's status in a robot control application. This control application can display the status of robot A and robot B by selecting tags. The various statuses of robot 100 selected by tags are indicated by icons. The status of robot 100 shown on application screen 280 includes the robot's voice. Icon 291 indicates the voice status of robot 100, and in the example in Figure 11, it is shown that "gentle voice" is selected as the voice of robot 100. The user can adjust the voice settings by pressing icon 291.
[0182] Figure 12A shows an example of the app screen when setting voice. When the user presses button 281 labeled "Select Voice," personality parameters are automatically generated to produce multiple (four in this embodiment) unique voices at random. All generated personality parameters do not overlap with personality parameters set on other robots 100. Once button 281 is pressed by the user and multiple (four in this embodiment) voices are generated, app screen 280A changes to app screen 280B shown in Figure 10B. Personality parameters will be described later.
[0183] Furthermore, when the user presses the button 282 labeled "Current Voice" on the app screen 280A, the user can hear the voice generated based on the personality parameters currently set for the robot 100. This confirmation voice is output from the robot 100's speaker 112, which is an output device that the user can perceive.
[0184] The app screen 280B shown in Figure 12B displays voice selection buttons 283A to 283D, which allow the user to select one of several generated personality parameters. When the user presses one of the voice selection buttons 283A to 283D to make a selection, the corresponding voice is output. This allows the user to confirm their preferred voice. When one of the voice selection buttons 283A to 283D is selected, and the user presses the button 284 labeled "Confirm," the personality parameter selected by the user is set for the robot.
[0185] Furthermore, a button for regenerating personality parameters (hereinafter referred to as the "regenerate button") may be displayed on the app screen 280B. The user will press the regenerate button if the personality parameters that produce the voice the user prefers have not been generated. When the regenerate button is pressed, the personality parameters are newly generated, and the newly generated personality parameters are associated with the voice selection buttons 283A to 283D.
[0186] The audio selection buttons 283A through 283D are lined up with multiple vertical bar objects that move in a way that reflects the sound elements, which are the individual parameters of each bar. These vertical bar objects dynamically and continuously change in length. The change in length of a vertical bar object has continuity with the change in length of an adjacent vertical bar object (such as a similar change with a time delay), and in this way, the multiple parallel vertical bar objects express the characteristics of sound (waveform).
[0187] The speed at which the waves of each vertical bar object pass, that is, the time it takes for one change to be reflected in an adjacent vertical bar object, represents the speed of the audio. The base length of the vertical bar object, that is, the length of the vertical bar object when it is not changing, represents the pitch of the audio. The dispersion of the waves of the vertical bar object, that is, the amount of change in the vertical bar object, represents the pitch range of the audio. The color of the vertical bar object represents the brightness of the audio. The decay and bounce of the vertical bar object represent the lip reflection of the audio. In other words, the decay of the amount of change in the vertical bar object is important; a large decay means that the time it takes for a wave to subside after it occurs is shorter, and a small decay means that the time it takes for a wave to subside after it occurs is longer. The thickness of the lines of the vertical bar object represents the tract length of the audio.
[0188] Figure 12C shows the application screen 280C, which is displayed when the history button 285 on the application screen 280A is pressed. The application screen 280C displays a list of voice selection buttons 286 for selecting from multiple personality parameters previously set by the user. In other words, the information terminal on which the control application is installed stores a history of personality parameters previously set by the user. For example, the maximum number of past personality parameters that can be selected is predetermined. Furthermore, this history of personality parameters may be stored in the cloud.
[0189] The voice selection button 286 displays the date on which the personality parameter was set for robot 100. This allows the user to recognize the differences in the voices corresponding to the voice selection button 286. Furthermore, by scrolling on the information terminal screen, the voice selection buttons 286 corresponding to previously set personality parameters that were not previously displayed will be revealed. When the user presses the button 287 labeled "Confirm," the personality parameter corresponding to the voice selection button 286 selected by the user is set for robot 100.
[0190] Furthermore, when a user selects a personality parameter via the app screen 280C, it is determined whether or not that voice feature data is not currently set on any other robots in the robot group that includes the robot 100. In other words, even if a personality parameter is stored, if it is currently set by another robot, it cannot be set on the robot 100.
[0191] Figure 12D shows the application screen 280D displayed on the information terminal when a user customizes the voice output from the robot. The application screen 280D displays personality parameters (combinations of values for multiple parameters) for generating the voice, which the user can select. For example, the user selects the parameter values by moving the slider bars 288A to 288F corresponding to each parameter left or right. In other words, the user manually generates the voice to be set for the robot 100 according to their own preferences, rather than having it automatically generated.
[0192] In the example in Figure 12D, the user-selectable parameters include speed, pitch, pitch range, brightness, lip vibration, and vocal cord length. This is treated as a single parameter set, and any voice parameter set that is not currently set for another robot can be set as the voice for robot 100. This parameter set is the same as the personality parameter generated when the user presses button 281 on the application screen 280A mentioned above.
[0193] Speed is the speaking speed per unit of sound. In language, a unit of sound is a syllable. The higher this value, the faster the speaking speed. Pitch is the average pitch. The higher this value, the higher the pitch. Pitch range is the range of pitches that can be pronounced. The higher this value, the wider the range of pitches. Brightness is a parameter that indicates the brightness of the voice (sound). The brightness of the sound can be changed by altering some of the frequency components of the sound being pronounced (for example, harmonic components). The higher this value, the more likely the voice (sound) is to be perceived as brighter. Lip vibration is the degree of lip vibration in a pronunciation structure that mimics the human vocal structure (mouth). The higher this value, the greater the sound reflectivity within the human vocal structure. Vocal cord length is a parameter that indicates the length of the vocal cords in a pronunciation structure that mimics the human vocal structure (mouth). The higher this value, the more low-frequency components there are in the sound, resulting in a more mature-sounding voice.
[0194] Furthermore, when the user presses button 289 labeled "Listen with robot," the voice generated with the selected personality parameters (parameter set) is output from the speaker 112 of robot 100. When the press operation of the confirmation button 290 is detected, the selected parameter set is set for the corresponding robot 100. Note that personality parameters generated manually are also voices that do not overlap with personality parameters set on other robots 100. Specifically, the information terminal sends the personality parameters (parameter set) selected by the user to the server that manages the voice settings, and the server determines whether the personality parameters overlap with personality parameters set on other robots 100 and sends the determination result to the upper terminal. If a personality parameter (parameter set) that overlaps with another robot 100 is selected, the information terminal disables the confirmation button 290 and displays a message on the touch panel display 210 indicating that it is in use by another robot. In this case, the server or information terminal may generate personality parameters similar to the personality parameters selected by the user, and a button to play the generated personality parameters may be displayed on the touch panel display.
[0195] Furthermore, voice customization using the app screen 280D may be permitted for users who meet certain conditions. These conditions include, for example, users who have used the robot 100 for a predetermined period, users who have earned a predetermined number of points, or users who have been charged a predetermined amount. [Industrial applicability]
[0196] In at least one embodiment, the robot emits a distinctive voice, which has the effect of making it easier for the user to feel as if the robot is a living being, making it useful as a voice-emitting robot. [Explanation of Symbols]
[0197] 100 robots 102 wheels 104 Body 105 hands 106th 108 seating surface 109 Horns 110 Display device 111 Internal Sensor 1111 Touch Sensor 1112 Accelerometer 1113 Camera 1114 Microphone 112 speakers 113 Communications Department 114 Storage device 115 processors 116 Drive Unit 117 Battery 118 wires 120 Power line 122 signal line 131 Positioning device 132 Instruction Reception Department 133 Growth Management Department 134 Situation Interpretation Section 135 Personality Development Department 136 Voice generation unit 1361 Audio Content Database 1362 Standard Speech Determination Unit 1363 Speech Synthesis Unit 1364 Audio Adjustment Unit
Claims
1. A voice generation unit that generates sound, An audio output unit that outputs the generated audio, A robot equipped with [this feature].
2. Equipped with a sensor that detects physical quantities and outputs sensor information, The voice generation unit generates voice based on the sensor information, as described in claim 1.
3. The robot according to claim 2, wherein the voice generation unit generates voice when predetermined sensor information is continuously input over a predetermined period of time.
4. Equipped with multiple sensors that detect physical quantities and output sensor information, The voice generation unit generates voice based on sensor information from the plurality of sensors, as described in claim 1.
5. Multiple sensors that detect physical quantities and output sensor information, An interpretation unit that interprets the semantic situation in which the robot is located based on the aforementioned sensor information, Equipped with, The robot according to claim 1, wherein the voice generation unit generates voice based on the semantic situation interpreted by the interpretation unit.
6. The robot according to claim 5, wherein the voice generation unit generates a voice when the interpretation unit interprets that the robot is being held.
7. The robot according to claim 2, wherein the voice generation unit generates voice that reflects the sensor information reflectively.
8. The voice generation unit generates voice with a volume based on the sensor information, according to claim 7.
9. The aforementioned sensor is an acceleration sensor that detects acceleration as the physical quantity and outputs acceleration as sensor information. The robot according to claim 7, wherein the voice generation unit generates sound whose volume changes in accordance with the change in acceleration.
10. The robot according to claim 7, wherein the voice generation unit generates voice of pitch based on the sensor information.
11. The aforementioned sensor is an acceleration sensor that detects acceleration as the physical quantity and outputs acceleration as sensor information. The robot according to claim 7, wherein the voice generation unit generates sound whose pitch changes in accordance with the change in acceleration.
12. It is further equipped with a personality-forming section that creates individuality, The aforementioned audio output unit outputs audio corresponding to the formed individuality, The robot further comprises a voice generation unit that generates voice corresponding to the formed personality, The robot according to claim 1, wherein the voice output unit outputs the generated voice.
13. The aforementioned speech generation unit, A standard voice determination unit that determines the standard voice, A voice adjustment unit that adjusts the determined standard voice to produce a distinctive voice, The robot according to claim 12, comprising:
14. The robot further comprises a growth management unit for managing the robot's growth, The robot according to claim 12 or 13, wherein the personality-forming unit forms the personality in accordance with the growth of the robot.
15. It also includes an instruction receiving unit that receives instructions from the user, The robot according to claim 12 or 13, wherein the personality forming unit forms the personality based on the received instructions.
16. It also features a microphone that converts sound into electrical signals, The robot according to claim 12 or 13, wherein the personality-forming unit forms the personality based on the electrical signal.
17. Equipped with a positioning device to measure location, The robot according to claim 12 or 13, wherein the personality-forming unit forms the personality based on the measured position.
18. The robot according to any one of claims 1 to 17, wherein the voice generation unit comprises a first voice generation unit that generates voice corresponding to the output value of an internal sensor, and a second voice generation unit that interprets the meaning of the output value of the internal sensor and generates voice according to the interpretation of the meaning.
19. It includes a voice generation condition recognition unit that recognizes voice conditions, including the voice output start timing and at least some voice parameters, The voice generation unit generates voice that matches at least some of the voice parameters included in the voice conditions, The robot according to any one of claims 1 to 17, wherein the voice output unit outputs the voice generated by the voice generation unit at the timing of the start of voice output included in the voice conditions.
20. The robot according to claim 19, wherein the voice generation unit generates voice that matches at least some of the voice parameters included in the voice conditions and that is appropriate to the robot's personality.
21. The voice generation condition recognition unit recognizes a first voice condition, which is its own voice condition, via communication. The robot according to claim 19 or 20, wherein the first voice condition is the same as at least a part of the second voice condition, which is a voice condition shown to another robot.
22. The robot according to claim 21, wherein some of the aforementioned audio parameters include a parameter indicating pitch.
23. The robot according to claim 22, wherein the first pitch included in the first sound condition has a predetermined relationship with the second pitch included in the second sound condition.
24. The first output start timing included in the first audio condition is included in the second audio condition The robot according to claim 23, wherein the timing of the start of the second output is the same as the timing of the start of the second output, and the interval which is the relative relationship between the first pitch included in the first sound condition and the second pitch included in the second sound condition is a consonant interval.
25. The aforementioned audio conditions include conditions indicating the length of the audio content, The robot according to any one of claims 19 to 24, further comprising a standard voice determination unit that randomly determines voice content that matches the length of the aforementioned voice content.
26. The aforementioned audio conditions include conditions indicating the length of the audio content, The robot according to any one of claims 19 to 24, further comprising a standard voice determination unit that determines voice content that matches the length of the aforementioned voice content based on previously collected voices.
27. A voice generation program for generating voice output from a robot, which is used on a computer. A speech generation step that generates speech, Audio output step to output the generated audio, A voice generation program that executes the command.