Robot, learning device, control method, and program
By integrating the acquisition of captured images and biological information in the robot and combining the behavioral decision model of machine learning, the limitations of timing design in the prior art are solved, and the effect of effectively inducing users to touch each other and providing healing in daily scenarios is achieved.
Patent Information
- Application Number
- CN202380069514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, robots based on body temperature cycles have limitations in timing concepts, and it is difficult to measure the timing of communication with users without stressing the user in daily scenarios, resulting in the inability to effectively promote contact with users, thereby unable to provide healing.
A robot is designed to determine the behavior that induces users to touch by obtaining the user's captured images and biological information, and using the state observation unit and the learning unit to generate a learning model based on machine learning to generate behavioral value.
It is achieved to induce users to contact at appropriate timing, promote contact between users and robots, and thus provide healing effects.
Smart Images

Figure CN119998015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robot, a learning device, a control method and a program. Background Art
[0002] Conventionally, there is a known robot that assumes contact with a user. Also, there is a known robot that provides healing by contact with a user.
[0003] Patent Document 1 discloses that a robot that measures a user's body temperature and determines a physical condition based on the body temperature cycle refers to a woman's menstrual cycle and changes its behavior pattern when a timing that requires attention comes.
[0004] <Prior Art Literature>
[0005] <Patent Documents>
[0006] Patent Document 1: Japanese Patent Application Publication No. 2009-104878 Summary of the invention
[0007] <Problems to be Solved by the Invention>
[0008] However, in Patent Document 1, the timing that should be taken into consideration assumes the user's body temperature cycle, which has limitations. In various daily scenarios, it is difficult to measure the timing of communication with the user based on the body temperature cycle without causing pressure on the user. If the communication with the user is not performed at the right time, the contact with the user cannot be promoted, and thus the healing cannot be provided.
[0009] The technology of the present invention aims to induce contact with a user at an appropriate timing.
[0010] <Methods used to solve the problem>
[0011] One embodiment of the present invention is a robot comprising: an acquisition unit that acquires a photographed image of a user and at least one of the biological information of the user; and a behavior control unit that, based on the photographed image and at least one of the biological information, instructs the execution of a given behavior that induces contact with the user according to the state of the user.
[0012] Another embodiment of the present invention is a learning device that is connected to a robot in a communicative manner, and the learning device comprises: a state observation unit that observes the state of the user based on at least one of a photographed image of the user and biological information of the user; and a learning unit that generates a learning model that inputs the state of the user and outputs the value of the robot's behavior through machine learning.
[0013] <Effects of the Invention>
[0014] According to one aspect of the present invention, it is possible to induce contact with a user at an appropriate timing. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a perspective view of a robot according to one embodiment.
[0016] Figure 2 This is a side view of a robot according to one embodiment.
[0017] Figure 3 This is a cross-sectional view of the robot along the III-III cutting line of one embodiment.
[0018] Figure 4 It is a diagram showing the structure of a life sensor according to one embodiment.
[0019] Figure 5 This is a block diagram showing a hardware configuration of a control unit according to one embodiment.
[0020] Figure 6 This is a block diagram showing a functional structure of a control unit according to one embodiment.
[0021] Figure 7 This is a block diagram showing a hardware configuration of an estimating unit (learning device) according to one embodiment.
[0022] Figure 8 This is a block diagram showing a functional configuration of an estimating unit (learning device) according to one embodiment.
[0023] Fig. 9 This is a schematic diagram of a neuron learning model according to one embodiment.
[0024] Fig.10 This is a schematic diagram of a learning model of a neural network according to one embodiment.
[0025] Fig.11 This is a flowchart showing the processing of the control unit according to one embodiment.
[0026] Fig.12 This is a flowchart showing the processing of the estimating unit (learning device) according to one embodiment.
[0027] Fig.13 It is a block diagram showing a functional configuration of an estimating unit (learning device) according to a modified example.
[0028] Fig.14 This is a flowchart showing the processing of the estimating unit (learning device) according to the modification. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In each of the drawings, the same components are denoted by the same reference numerals, and repeated descriptions are appropriately omitted.
[0030] The embodiments shown below illustrate robots for embodying the technical concept of the present invention, but the present invention is not limited to the embodiments shown below. The sizes, materials, shapes and relative arrangements of the components described below are intended to be illustrative unless otherwise specified, and are not intended to limit the scope of the present invention to these. In addition, the sizes and positional relationships of the components shown in the drawings may be exaggerated to make the description clear.
[0031] <Overall Configuration Example of Robot 100>
[0032] Reference Figures 1 to 3 , a structure of a robot 100 according to one embodiment is described. Figure 1 It is a perspective view illustrating a robot 100 according to an embodiment. Figure 2 is a side view of the robot 100 . Figure 3 It is along Figure 2 A cross-sectional view taken along the III-III cutting line.
[0033] The robot 100 is a robot having an outer casing 10 and can be driven by supplied power. The robot 100 illustrated in the present embodiment is a communication robot of a doll type imitating a bear. The robot 100 is made in a size and weight suitable for a user to hold. Here, the user refers to the user of the robot 100. Representative examples of users include single people living alone, elderly people whose children are already independent, and frail elderly people who are the objects of home medical treatment. In addition, in addition to the user of the robot 100, the user may also include the manager of the robot 100 and other people who only come into contact with the robot 100.
[0034] The outer casing 10 has flexibility. For example, the outer casing 10 includes a soft raw material that gives a good touch when the user of the robot 100 touches the robot 100. As the raw material of the outer casing 10, a raw material including an organic material such as polyurethane foam, rubber, resin, fiber, etc. can be used. The outer casing 10 is preferably composed of an outer casing such as a polyurethane foam material having heat insulation properties, and a soft cloth covering the outer surface of the outer casing.
[0035] As an example, the robot 100 includes a body 1, a head 2, an arm 3, and a leg 4. The head 2 includes a right eye 2a, a left eye 2b, a mouth 2c, a right cheek 2d, and a left cheek 2e. The arm 3 includes a right arm 3a and a left arm 3b, and the leg 4 includes a right leg 4a and a left leg 4b. Here, the body 1 corresponds to the robot body. The head 2, the arm 3, and the leg 4 correspond to the driving bodies connected to the robot body in a manner that allows relative displacement.
[0036] In the present embodiment, the arm 3 is configured to be displaceable relative to the body 1. For example, when the robot 100 is hugged by the user, the right arm 3a and the left arm 3b are displaced to contact the user's head, body, etc. in a manner of hugging the user. Through this action, the user feels close to the robot 100, so that the contact between the user and the robot 100 can be promoted. In addition, the so-called contact with the user refers to the action (contact action) of the user and the robot 100 touching each other, such as rubbing, patting (touching) and hugging (embracing).
[0037] The body 1, the head 2, the arms 3, and the legs 4 are all covered by the outer casing 10. The outer casing in the body 1 is integrated with the outer casing in the arms 3, and the outer casing in the head 2 and the legs 4 is separated from the outer casing in the body 1 and the arms 3. However, it is not limited to these structures, and for example, only the parts of the robot 100 that are easily contacted by the user may be covered by the outer casing 10. In addition, at least one of the outer casing 10 in each of the body 1, the head 2, the arms 3, and the legs 4 may be separated from the other outer casings. In addition, the parts of the head 2, the arms 3, and the legs 4 that do not displace may not include components such as sensors on their inner sides and may be composed only of the outer casing 10.
[0038] The robot 100 has a camera 11, a tactile sensor 12, a control unit 13, a vital sensor 14, a battery 15, a first electrostatic capacitance sensor 21, and a second electrostatic capacitance sensor 31 inside the outer casing 10. In addition, the robot 100 has a camera 11, a tactile sensor 12, a control unit 13, a vital sensor 14, and a battery 15 inside the outer casing 10 in the body 1. Furthermore, the robot 100 has a first electrostatic capacitance sensor 21 inside the outer casing 10 in the head 2, and has a second electrostatic capacitance sensor 31 inside the outer casing 10 in the arm 3.
[0039] In addition, the robot 100 has a display 24, a speaker 25, and a light 26 inside the exterior member 10 in the head 2. Furthermore, the robot 100 has a display 24 inside the exterior member 10 in the right eye 2a and the left eye 2b. In addition, the robot 100 has a speaker 25 inside the exterior member 10 in the mouth 2c, and a light 26 inside the exterior member 10 in the right cheek 2d and the left cheek 2e.
[0040] In more detail, Figure 3As shown, the robot 100 has a body frame 16 and a body mounting platform 17 inside the exterior member 10 in the body 1. In addition, the robot 100 has a head frame 22 and a head mounting platform 23 inside the exterior member 10 in the head 2. Furthermore, the robot 100 has a right arm frame 32a and a right arm mounting platform 33 inside the exterior member 10 in the right arm 3a, and a left arm frame 32b inside the exterior member 10 in the left arm 3b. In addition, the robot 100 has a right leg frame 42a inside the exterior member 10 in the right leg 4a, and a left leg frame 42b inside the exterior member 10 in the left leg 4b.
[0041] The body frame 16, the head frame 22, the right arm frame 32a, the left arm frame 32b, the right leg frame 42a, and the left leg frame 42b are structures formed by combining a plurality of columnar members. The body mounting platform 17, the head mounting platform 23, and the right arm mounting platform 33 are plate-like members having a mounting surface. The body mounting platform 17 is fixed to the body frame 16, the head mounting platform 23 is fixed to the head frame 22, and the right arm mounting platform 33 is fixed to the right arm frame 32a. In addition, the body frame 16, the head frame 22, the right arm frame 32a, the left arm frame 32b, the right leg frame 42a, and the left leg frame 42b may also be formed in a box shape including a plurality of plate-like members.
[0042] The right arm frame 32a is connected to the body frame 16 via the right arm connection mechanism 34a, and is driven by the right arm servo motor 35a to be relatively displaced with respect to the body frame 16. The right arm frame 32a is displaced, so that the right arm 3a is relatively displaced with respect to the body 1. The right arm connection mechanism 34a preferably has a speed reducer that increases the output torque of the right arm servo motor 35a, for example.
[0043] In this embodiment, the right arm frame 32a is composed of a multi-joint robot including a plurality of frame members and a plurality of connection mechanisms. For example, the right arm frame 32a includes a right shoulder frame F1a, a right upper arm frame F2a, a right elbow frame F3a, and a right forearm frame F4a. The trunk frame 16, the right shoulder frame F1a, the right upper arm frame F2a, the right elbow frame F3a, and the right forearm frame F4a are connected to each other via the connection mechanisms.
[0044] The right arm servo motor 35a is a general term for a plurality of servo motors. For example, the right arm servo motor 35a includes a right shoulder servo motor M1a, a right upper arm servo motor M2a, a right elbow servo motor M3a, and a right forearm servo motor M4a. The right shoulder servo motor M1a rotates the right shoulder frame F1a around a rotation axis that is perpendicular to the trunk frame 16. The right upper arm servo motor M2a rotates the right upper arm frame F2a around a rotation axis that is perpendicular to the rotation axis of the right shoulder frame F1a. The right elbow servo motor M3a rotates the right elbow frame F3a around a rotation axis that is perpendicular to the rotation axis of the right upper arm frame F2a. The right forearm servo motor M4a rotates the right forearm frame F4a around a rotation axis that is perpendicular to the rotation axis of the right elbow frame F3a.
[0045] The left arm frame 32b is connected to the body frame 16 via the left arm connection mechanism 34b, and is driven by the left arm servo motor 35b to be relatively displaced with respect to the body frame 16. The displacement of the left arm frame 32b causes the left arm 3b to be relatively displaced with respect to the body 1. The left arm connection mechanism 34b preferably has, for example, a speed reducer that increases the output torque of the left arm servo motor 35b.
[0046] In this embodiment, the left arm frame 32b is composed of a multi-joint robot including a plurality of frame members and a plurality of connection mechanisms. For example, the left arm frame 32b includes a left shoulder frame F1b, a left upper arm frame F2b, a left elbow frame F3b, and a left forearm frame F4b. The trunk frame 16, the left shoulder frame F1b, the left upper arm frame F2b, the left elbow frame F3b, and the left forearm frame F4b are connected to each other via the connection mechanisms.
[0047] The left arm servo motor 35b is a general term for a plurality of servo motors. For example, the left arm servo motor 35b includes a left shoulder servo motor M1b, a left upper arm servo motor M2b, a left elbow servo motor M3b, and a left forearm servo motor M4b. The left shoulder servo motor M1b rotates the left shoulder frame F1b around a rotation axis that is perpendicular to the trunk frame 16. The left upper arm servo motor M2b rotates the left upper arm frame F2b around a rotation axis that is perpendicular to the rotation axis of the left shoulder frame F1b. The left elbow servo motor M3b rotates the left elbow frame F3b around a rotation axis that is perpendicular to the rotation axis of the left upper arm frame F2b. The left forearm servo motor M4b rotates the left forearm frame F4b around a rotation axis that is perpendicular to the rotation axis of the left elbow frame F3b.
[0048] Since the arm 3 has a four-axis joint, the robot 100 can perform more realistic movements. For example, the robot 100 can induce a user to touch the user by turning the arm 3 and "opening the hand" when the user is indifferent and does not move. In addition, the robot 100 can induce a user to touch the user by shaking the arm 3 and "fearing" when the user walks with anger.
[0049] The head frame 22 is connected to the body frame 16 via the head connection mechanism 27, and can be relatively displaced with respect to the body frame 16 by being driven by the head servo motor 35c. The head 2 is relatively displaced with respect to the body 1 by the displacement of the head frame 22. The head connection mechanism 27 preferably has a speed reducer that increases the output torque of the head servo motor 35c, for example.
[0050] In this embodiment, the head frame 22 includes a neck frame F1c and a face frame F2c. The body frame 16, the neck frame F1c, and the face frame F2c are connected to each other via a connection mechanism.
[0051] The head servo motor 35c is a general term for a plurality of servo motors. For example, the head servo motor 35c includes a neck servo motor M1c and a face servo motor M2c. The neck servo motor M1c rotates the neck frame F1c around a rotation axis perpendicular to the body frame 16. The face servo motor M2c rotates the face frame F2c around a rotation axis perpendicular to the rotation axis of the neck frame F1c.
[0052] Since the head 2 has a two-axis joint, the robot 100 can achieve more realistic movements. For example, if the robot 100 turns its head 2 and "looks up" (looks at) a user who is operating a mobile phone with disgust, it can express its concern for the user and induce contact with the user.
[0053] The right leg frame 42a is connected to the trunk frame 16 via the right leg connection mechanism 44a, and has a right leg wheel 41a on the bottom side. In order to stabilize the posture of the robot 100, the robot 100 preferably has two right leg wheels 41a in the front-to-back direction of the right leg frame 42a. The right leg wheel 41a is driven by the right leg servo motor 35d so as to be able to rotate around a rotation axis perpendicular to the front-to-back direction of the right leg frame 42a. The robot 100 becomes able to travel by rotating the right leg wheel 41a. The right leg connection mechanism 44a, for example, preferably has a reducer that increases the output torque of the right leg servo motor 35d.
[0054] The left leg frame 42b is connected to the trunk frame 16 via the left leg connection mechanism 44b, and has a left leg wheel 41b on the bottom side. In order to stabilize the posture of the robot 100, the robot 100 preferably has two left leg wheels 41b in the front-to-back direction of the left leg frame 42b. The left leg wheel 41b is driven by the left leg servo motor 35e and can rotate around a rotation axis perpendicular to the front-to-back direction of the left leg frame 42b. The robot 100 becomes able to travel by rotating the left leg wheel 41b. The left leg connection mechanism 44b, for example, preferably has a reducer that increases the output torque of the left leg servo motor 35e.
[0055] In this embodiment, the robot 100 moves forward or backward by turning the right leg wheel 41a and the left leg wheel 41b forward or backward at the same time. The robot 100 turns right or left by braking one of the right leg wheel 41a and the left leg wheel 41b and turning the other forward or backward.
[0056] Thus, the robot 100 can realize more realistic actions through the legs 4. For example, the robot 100 can perform an action to induce contact with a user by turning the legs 4 to "approach" a user who is standing still with a sad or surprised emotion.
[0057] The camera 11 is fixed to the body frame 16. The tactile sensor 12, the control unit 13, the vital sensor 14, and the battery 15 are fixed to the body mounting platform 17. The control unit 13 and the battery 15 are fixed to the side of the body mounting platform 17 opposite to the side to which the tactile sensor 12 and the vital sensor 14 are fixed. In addition, the arrangement of the control unit 13 and the battery 15 here is determined according to the specific conditions of the space that can be arranged on the body mounting platform 17, and is not limited to the above arrangement. However, when the battery 15 is fixed to the side of the body mounting platform 17 opposite to the side to which the tactile sensor 12 and the vital sensor 14 are fixed, the center of gravity of the robot 100 is lowered because the battery 15 is heavier than other components. When the center of gravity of the robot 100 is lower, at least one of the position and posture of the robot 100 is stabilized, and at least one of charging and replacing the battery 15 becomes easy, so it is preferable.
[0058] The first electrostatic capacitance sensor 21 is fixed to the head mounting platform 23, and the second electrostatic capacitance sensor 31 is fixed to the right arm mounting platform 33. The display 24 includes a right eye display 24a and a left eye display 24b. The right eye display 24a, the left eye display 24b and the speaker 25 are fixed to the head frame 22. The light 26 includes a right cheek light 26a and a left cheek light 26b. The right cheek light 26a and the left cheek light 26b are fixed to the head frame 22.
[0059] In addition, the camera 11, the tactile sensor 12, the control unit 13, the life sensor 14, the battery 15, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, etc. can be fixed by screw members or adhesive members, etc. In addition, the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, the left cheek light 26b, etc. can also be fixed by screw members or adhesive members, etc.
[0060] The materials of the body frame 16, the body mounting platform 17, the head frame 22, the head mounting platform 23, the right arm frame 32a, the right arm mounting platform 33 and the left arm frame 32b are not particularly limited, and resin materials or metal materials can be used. However, from the viewpoint of ensuring the strength during driving, metal materials such as aluminum are preferably used for the body frame 16, the right arm frame 32a and the left arm frame 32b. On the other hand, as long as the strength can be ensured, in order to make the robot 100 lightweight, resin materials are preferably used for the materials of these parts. The materials of the body mounting platform 17, the head frame 22, the head mounting platform 23, the right arm mounting platform 33 and the left arm frame 32b are also not particularly limited, and resin materials or metal materials can be used, but from the viewpoint of making the robot 100 lightweight, resin materials are preferably used.
[0061] The control unit 13 is connected to the camera 11, the tactile sensor 12, the vital sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the right arm servo motor 35a, and the left arm servo motor 35b respectively by wire or wireless so as to enable communication. In addition, the control unit 13 is also connected to the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e respectively by wire or wireless so as to enable communication. Furthermore, the control unit 13 is also connected to the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, and the left cheek light 26b respectively by wire or wireless so as to enable communication.
[0062] The camera 11 is an image sensor that outputs a captured image of the robot 100's surroundings to the control unit 13. In the present embodiment, the camera 11 is an example of a capturing unit that captures a user. The camera 11 includes a lens and an image capturing element that captures an image formed by the lens. The image capturing element may use a CCD (Charge Coupled Device) or a CMOS (Complementary Metal-Oxide Semiconductor). The captured image may be either a still image or a dynamic image.
[0063] In addition, the camera 11 is preferably composed of a TOF (Time Of Flight) camera that outputs a distance image around the robot 100 to the control unit 13. Therefore, the captured image output from the camera 11 sometimes includes a three-dimensional captured image (distance image) in addition to the two-dimensional captured image, or includes a three-dimensional captured image (distance image) instead of a two-dimensional camera image. The captured image is used to detect the presence or approach of a user, detect the distance from the robot 100 to the user, authenticate the user, or infer the user's emotions or behaviors. The captured image is an example of a captured image of the user. In addition, the robot 100 may also have, in addition to the camera 11, a human sensor such as an ultrasonic sensor, an infrared sensor, a millimeter wave radar, or a LiDAR (light Detection And Raging).
[0064] The tactile sensor 12 is a sensor element that detects information sensed by the tactile sense possessed by a human hand or the like, converts the information into a tactile signal as an electrical signal, and outputs the information to the control unit 13. For example, the tactile sensor 12 converts information on pressure or vibration generated by contact between the user and the robot 100 into a tactile signal through a piezoelectric element, and outputs the tactile signal to the control unit 13. The tactile signal output from the tactile sensor 12 is used to detect contact or presence of the user 200 relative to the robot 100.
[0065] The life sensor 14 is an example of an electromagnetic wave sensor that uses electromagnetic waves to obtain biological information of the user. Figure 4 Details will be given separately.
[0066] The first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are sensor elements that output electrostatic capacitance signals to the control unit 13 for detecting the contact or proximity of the user with the robot 100 based on the change in electrostatic capacitance. From the viewpoint of stabilizing the exterior member 10, the first electrostatic capacitance sensor 21 is preferably a rigid sensor without flexibility. Since the arm 3 is a part that the user easily touches, from the viewpoint of making the touch feel good, the second electrostatic capacitance sensor 31 is preferably a sensor that includes a conductive wire and has flexibility. The electrostatic capacitance signals output from the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are used to detect the proximity or presence of the user relative to the robot 100.
[0067] The right eye display 24a and the left eye display 24b are display modules that display character strings or images such as characters, numbers, and symbols according to instructions from the control unit 13. The right eye display 24a and the left eye display 24b are composed of, for example, liquid crystal display modules. The character strings or images displayed on the right eye display 24a and the left eye display 24b are used for the emotional expression of the robot 100. For example, for a user sitting with a happy emotion, the robot 100 displays a "smiling" image on the right eye display 24a and the left eye display 24b to empathize with the happiness, thereby being able to suggestively induce contact with the user.
[0068] The speaker 25 is a speaker unit that amplifies the sound signal from the control unit 13 and outputs the sound. The sound output from the speaker 25 is the speech or call of the robot 100, and is used to express the emotion of the robot 100. For example, the robot 100 outputs a sound "making a (caring) sound" from the speaker 25 to a user who is doing housework with a sad mood, thereby inducing a behavior of contacting the user.
[0069] The right cheek light 26a and the left cheek light 26b are light modules that flash or change color according to the on / off signal from the control unit 13. The right cheek light 26a and the left cheek light 26b are composed of, for example, LED (Light Emitting Diode) light modules. The flashing or color change of the right cheek light 26a and the left cheek light 26b is used for the emotional expression of the robot 100. For example, the robot 100 can express empathy for a user who is sitting with a sad mood by flashing the right cheek light 26a and the left cheek light 26b in blue, thereby inducing a behavior of touching the user.
[0070] The battery 15 is a power source that supplies power to the camera 11, the tactile sensor 12, the control unit 13, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the right arm servo motor 35a, and the left arm servo motor 35b. In addition, the battery 15 also supplies power to the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e. Furthermore, the battery 15 also supplies power to the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, and the left cheek light 26b. The battery 15 can use various secondary batteries such as lithium ion batteries and lithium polymer batteries.
[0071] In addition, various sensors such as the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 in the robot 100 are not essential components. The robot 100 only needs to have at least the camera 11, the life sensor 14 and the tactile sensor 12. Their installation positions can also be changed appropriately. Furthermore, various sensors such as the camera 11, the life sensor 14 and the tactile sensor 12 can also be configured on the outside of the robot 100 to send necessary information to the robot 100 or an external device via wireless. For example, a learning device composed of a PC (Personal Computer) or a server is an example of an external device.
[0072] The robot 100 does not necessarily need to include the control unit 13 inside the exterior member 10, and the control unit 13 may communicate with each device via wireless from outside the exterior member 10. The battery 15 may supply power to each component from outside the exterior member 10.
[0073] In this embodiment, the structure in which the head 2, the arm 3, and the leg 4 are all displaceable is illustrated, but the present invention is not limited thereto, and at least one of the head 2, the arm 3, and the leg 4 may be displaceable. In addition, the arm 3 is composed of a 4-axis multi-joint robot arm, but it may actually be composed of a 6-axis multi-joint robot arm. Furthermore, the arm 3 is preferably capable of connecting an end effector such as a hand. In addition, the leg 4 is composed of a wheel system, but it may also be composed of a track system or a leg system.
[0074] The structure and shape of the robot 100 are not limited to those illustrated in this embodiment, and can be appropriately changed according to the user's preference, the usage form of the robot 100, etc. For example, the robot 100 may not be in the form of a bear but in the form of a mechanical arm of an industrial robot, etc., or in the form of a humanoid puppet. In addition, the robot 100 may be in the form of a mobile device such as a drone or a vehicle having at least one of an arm, a display, a speaker, and a light.
[0075] <Configuration example of the life sensor 14>
[0076] Figure 4 14 is a diagram illustrating a configuration of a life sensor 14. The life sensor 14 is a microwave Doppler sensor having a microwave transmitting unit 141 and a microwave receiving unit 142. Microwaves are an example of electromagnetic waves.
[0077] The life sensor 14 transmits a transmission wave Ms as a microwave from the inside of the outer casing 10 of the robot 100 toward the user 200 through the microwave transmitting unit 141. In addition, the life sensor 14 receives a reflected wave Mr resulting from the transmission wave Ms being reflected by the user 200 through the microwave receiving unit 142.
[0078] The life sensor 14 detects the minute displacements on the body surface caused by the heart beats of the user 200 in a non-contact manner by using the Doppler effect based on the difference between the frequencies of the transmission wave Ms and the reflection wave Mr. The life sensor 14 can obtain information such as the heartbeat, respiration, pulse wave, blood pressure, etc. as biological information of the user 200 based on the detected minute displacements, and output the obtained biological information to the control unit 13.
[0079] However, the life sensor 14 is not limited to a microwave Doppler sensor, and may be a life sensor that detects micro-displacements generated on the body surface by using changes in the coupling between the human body and the antenna, or may be a life sensor that uses electromagnetic waves other than microwaves such as near-infrared light. In addition, the life sensor 14 may also be a millimeter wave radar, a microwave radar, etc. Furthermore, the life sensor 14 preferably has a non-contact thermometer that detects infrared rays emitted from the user 200 in addition to the Doppler sensor. In this case, the life sensor 14 detects biological information of the user 200 including information related to at least one of the heartbeat (pulse), respiration, blood pressure, and body temperature.
[0080] In this embodiment, the life sensor 14 is provided inside the outer casing 10, so that the user 200 cannot visually recognize the life sensor 14. Thus, the user 200 can suppress the resistance to the detection of biological information, and can smoothly obtain biological information. In addition, the life sensor 14 can obtain biological information in a non-contact manner, so unlike a contact sensor that requires the user to contact the same place for a certain period of time, the biological information can be obtained even if the user moves to some extent.
[0081] Furthermore, by promoting contact between the user 200 and the robot 100 through the hugging action of the robot 100, the robot 100 is hugged by the user 200 and can acquire biological information while in contact or close to the user 200. Thus, the robot 100 can acquire highly reliable biological information with suppressed noise.
[0082] <Configuration Example of Control Unit 13>
[0083] (Hardware Configuration Example)
[0084] Figure 51 is a block diagram showing the hardware structure of the control unit 13. The control unit 13 is constructed by a computer and has a CPU (Central Processing Unit) 131, a ROM (Read Only Memory) 132, and a RAM (Random Access Memory) 133. In addition, the control unit 13 has a HDD / SSD (Hard Disk Drive / Solid State Drive) 134, a device connection I / F (Interface) 135, and a communication I / F 136. These are connected via a system bus A so as to be able to communicate with each other.
[0085] The CPU 131 performs control processing including various calculation processing. The ROM 132 stores programs such as IPL (Initial Program Loader) for driving the CPU 131. The RAM 133 is used as a work area for the CPU 131. The HDD / SSD 134 stores various information such as programs, camera images obtained by the camera 11, biological information obtained by the life sensor 14, and detection information obtained by various sensors such as tactile signals obtained by the tactile sensor 12.
[0086] The device connection I / F 135 is an interface for connecting the control unit 13 to various external devices. The external devices here include the camera 11, the tactile sensor 12, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the servo motor 35, and the battery 15. In addition, the external devices also include the display 24, the speaker 25, and the light 26.
[0087] Here, the servo motor 35 is a generic term for the right arm servo motor 35a, the left arm servo motor 35b, the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e. In addition, the display 24 is a generic term for the right eye display 24a and the left eye display 24b. Furthermore, the light 26 is a generic term for the right cheek light 26a and the left cheek light 26b.
[0088] The communication I / F 136 is an interface for communicating with an external device via a communication network, etc. For example, the control unit 13 is connected to the Internet via the communication I / F 136 and communicates with an external device via the Internet.
[0089] In addition, at least a part of the functions implemented by the CPU 131 may be implemented by an electric circuit or an electronic circuit.
[0090] (Functional configuration example)
[0091] Figure 6 1 is a block diagram showing the functional structure of the control unit 13. The control unit 13 includes an acquisition unit 101, a communication control unit 102, a storage unit 103, an authentication unit 104, a registration unit 105, a start control unit 106, a motor control unit 107, and an output unit 108. Furthermore, the control unit 13 includes a detection unit 110, an estimation unit 111, and a behavior control unit 112.
[0092] The control unit 13 can realize the functions of the acquisition unit 101 and the output unit 108 through the device connection I / F 135 and the like, and can realize the function of the communication control unit 102 through the communication I / F 136 and the like. In addition, the control unit 13 can realize the functions of the storage unit 103 and the registration unit 105 through the non-volatile memory such as the HDD / SSD 134. Furthermore, the functions of the authentication unit 104, the start control unit 106, and the motor control unit 107 can be realized by the processor such as the CPU 131 executing the processing specified by the program stored in the non-volatile memory such as the ROM 132.
[0093] In addition, the functions of the detection unit 110, the estimation unit 111, and the behavior control unit 112 can be realized by executing the processing specified by the program stored in the non-volatile memory such as the ROM 132 by the processor such as the CPU 131. In addition, part of the above functions of the control unit 13 can also be realized by an external device such as a PC or a server, or can be realized by distributed processing between the control unit 13 and the external device. For example, the estimation unit 111 can also be configured as a learning device connected to the robot 100 in a manner that allows communication.
[0094] The acquisition unit 101 acquires the captured image Im of the user 200 from the camera 11 by controlling the communication between the control unit 13 and the camera 11. In addition, the acquisition unit 101 acquires the tactile signal S from the tactile sensor 12 by controlling the communication between the control unit 13 and the tactile sensor 12. Furthermore, the acquisition unit 101 acquires the biological information B of the user 200 from the life sensor 14 by controlling the communication between the control unit 13 and the life sensor 14.
[0095] The acquisition unit 101 acquires the first capacitance signal C1 from the first capacitance sensor 21 by controlling the communication between the control unit 13 and the first capacitance sensor 21. The acquisition unit 101 acquires the second capacitance signal C2 from the second capacitance sensor 31 by controlling the communication between the control unit 13 and the second capacitance sensor 31.
[0096] The communication control unit 102 controls communication with an external device via a communication network, etc. For example, the communication control unit 102 can send a captured image Im obtained by the camera 11, biological information B obtained by the life sensor 14, a tactile signal S obtained by the tactile sensor 12, etc. to an external device (e.g., a learning device described later) via the communication network.
[0097] The storage unit 103 stores the biological information B acquired by the life sensor 14. The storage unit 103 continuously stores the acquired biological information B while the acquisition unit 101 acquires the biological information B from the life sensor 14. In addition, the storage unit 103 can also store information obtained from the captured image Im acquired by the camera 11, the tactile signal S from the tactile sensor 12, the first electrostatic capacitance signal C1 from the first electrostatic capacitance sensor 21, and the second electrostatic capacitance signal C2 from the second electrostatic capacitance sensor 31.
[0098] The authentication unit 104 performs personal authentication on the user 200 based on the captured image Im of the user 200 acquired by the camera 11. For example, the authentication unit 104 performs facial authentication based on the captured image Im including the face of the user 200 captured by the camera 11, referring to the registration information 109 of the facial image pre-registered in the registration unit 105. In this way, the user 200 who is currently in contact with or close to the robot 100 can be associated with the pre-registered personal information, and the biometric information B acquired by the life sensor 14 can be associated with the personal information. In addition, the control unit 13 may control to stop the acquisition of the biometric information by the life sensor 14 when the facial image included in the captured image Im is not registered in the registration unit 105.
[0099] The start control unit 106 causes the life sensor 14 to start acquiring the biological information B. For example, when the detection unit 110 detects that the user 200 is in contact with or close to the robot 100, the start control unit 106 turns on a switch that supplies power from the battery 15 to the life sensor 14. Thus, the start control unit 106 causes the life sensor 14 to start acquiring the biological information B.
[0100] The detection unit 110 detects the presence or approach of the user 200 around the robot 100 based on the captured image Im obtained by the camera 11 or the like. The detection unit 110 preferably detects the distance from the robot 100 to the user 200 based on the captured image Im (distance image) obtained by the camera 11. In addition, the detection unit 110 may detect the approach or presence of the user 200 relative to the robot 100 based on the first electrostatic capacitance signal C1 or the second electrostatic capacitance signal C2. Furthermore, the detection unit 110 detects the contact or presence of the user 200 relative to the robot 100 based on the tactile signal S from the tactile sensor 12.
[0101] The estimation unit 111 estimates a given behavior an (n is the identification number of the behavior a) of the robot 100 that matches the state of the user 200 based on the camera image Im of the user 200 and the biological information B of the user 200. In the present embodiment, the estimation unit 111 performs reinforcement learning to estimate the behavior at (t is the time) of the robot 100 that matches the state st (t is the time) of the user 200. However, the estimation unit 111 may also perform other machine learning such as supervised learning, semi-supervised learning, or unsupervised learning to estimate the given behavior at of the robot 100 that matches the state st of the user 200.
[0102] Regarding the configuration of the estimation unit 111 that performs reinforcement learning, refer to Figure 8 In addition, regarding the structure of the estimation unit 111 for supervised learning, please refer to Fig.13 Further, when learning converges, the inference unit 111 may use the completed learning model (behavior value table or neural network in this embodiment) to infer the behavior at of the robot 100 that matches the state st of the user 200. In this case, the inference unit 111 infers the given behavior at of the robot 100 that matches the state st of the user 200 through a given logic or a given algorithm.
[0103] The behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to perform the behavior at of the robot 100. In addition, the behavior control unit 112 instructs the execution of the behavior at of inducing contact with the user 200 according to the state st of the user 200. The robot 100 performs the behavior of inducing contact with the user 200 at an appropriate timing, thereby promoting contact with the user 200, thereby providing healing to the user 200.
[0104] The behavior at inducing contact with the user 200 is, for example, "opening hands" or "waving hands" to explicitly induce contact with the user 200. The behavior at inducing contact with the user 200 may also be, for example, "fearful" or "singing" to implicitly induce contact with the user 200.
[0105] The storage unit 103 stores information related to the behavior an of the robot 100 that has been defined in advance. The information related to the behavior an of the robot 100 is managed, for example, by a table of a database. Table 1 below is an example of a behavior table TB1 related to the behavior an of the robot 100. The behavior table TB1 contains a behavior ID (Identity Document) for identifying the behavior an of the robot 100, the behavior content of the robot 100, the instruction content of the behavior an, the behavior time per cycle, and a usage example.
[0106]
Table 1
[0107]
[0108] The symbols in the instruction contents shown in Table 1 represent the symbols of the control objects. In addition, the so-called teaching instruction is an action instruction that is taught in advance using a teaching method such as offline teaching, online teaching, or direct teaching. In addition, the so-called tracking instruction is an action instruction that tracks the position and posture of the user 200 based on various sensor information such as the captured image Im (range image).
[0109] The motor control unit 107 controls the driving of the servo motor 35 according to the execution instruction of the behavior at of the robot 100 from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "open the hand", the motor control unit 107 executes the action instruction of "open the hand" taught in advance.
[0110] The output unit 108 controls the communication between the control unit 13 and the display 24 according to the execution instruction of the behavior at from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "smiling", the output unit 108 outputs the image data of the smiling face to the right eye display 24a and the left eye display 24b.
[0111] In addition, the output unit 108 controls the communication between the control unit 13 and the speaker 25 according to the execution instruction of the behavior at from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "making a (greeting) sound", the output unit 108 outputs a greeting sound output signal to the speaker 25.
[0112] Furthermore, the output unit 108 controls the communication between the control unit 13 and the lights 26 according to the execution instruction of the behavior at from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "blinking cheeks", the output unit 108 outputs on-off signals to the switch elements of the right cheek light 26a and the left cheek light 26b.
[0113] <Configuration of the Estimation Unit 111>
[0114] (Hardware Configuration Example)
[0115] Figure 7 2 is a block diagram showing the hardware configuration of the estimation unit 111 . Figure 7 Shows the Figure 6 The estimation unit 111 shown in FIG. 1 is configured as an example of a learning device 300 connected to the robot 100 in a communicable manner. However, the function of the estimation unit 111 may also be as follows: Figure 6 As shown, it is arranged inside the robot 100.
[0116] The estimation unit 111 is constructed by a computer and has a CPU 301, a ROM 302, a RAM 303, a HDD / SSD 304, a device connection I / F 305, and a communication I / F 306. These are connected via a system bus A' so as to be able to communicate with each other. In addition, in order to improve the learning processing capability of the computer, the learning device 300 is preferably composed of a PC cluster having a GPU (Graphics Processing Unit) or a plurality of computers.
[0117] The CPU 301 performs control processing including various calculation processing. The ROM 302 stores programs such as IPL (Initial Program Loader) for driving the CPU 301. The RAM 303 is used as a work area for the CPU 301. The HDD / SSD 304 stores various information such as programs, the camera image Im obtained by the camera 11, the biological information B obtained by the life sensor 14, or the detection information based on various sensors.
[0118] The device connection I / F 305 is an interface for connecting the estimation unit 111 to various external devices. The external devices here are the camera 11, the tactile sensor 12, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, etc. However, the estimation unit 111 can also obtain detection information of these various sensors from the robot 100 via the communication I / F 306 described later.
[0119] The communication I / F 306 is an interface for communicating with an external device such as the robot 100 via a communication network, etc. For example, the estimation unit 111 is connected to the Internet via the communication I / F 136 and communicates with the external device via the Internet. In addition, the estimation unit 111 directly communicates wirelessly with the external device via the communication I / F 306.
[0120] In addition, at least a part of the functions implemented by the CPU 301 may be implemented by an electric circuit or an electronic circuit.
[0121] (Functional configuration example)
[0122] Figure 8 1 is a block diagram showing the functional structure of the estimation unit 111. The estimation unit 111 includes a state observation unit 121, a behavior determination unit 122, a result acquisition unit 123, a learning unit 124, a communication control unit 125, and a storage unit 126. In addition, when the function of the estimation unit 111 is set inside the robot 100, the communication control unit 125 and the storage unit 126 become unnecessary. In addition, when the estimation unit 111 estimates the behavior at of the robot 100 adapted to the state st of the user 200 using a learned learning model LM or based on a given algorithm, the result acquisition unit 123 and the learning unit 124 become unnecessary.
[0123] The various functions of the state observation unit 121, the behavior determination unit 122, the result acquisition unit 123, and the learning unit 124 can be realized by a processor such as the CPU 301 executing processing specified by a program stored in a non-volatile memory such as the ROM 302. In addition, the function of the communication control unit 125 can be realized by the communication I / F 306, etc. Furthermore, the function of the storage unit 126 can be realized by a non-volatile memory such as the HDD / SDD 304.
[0124] The estimation unit 111 of this embodiment performs reinforcement learning to estimate the behavior at of the robot 100 that is suitable for the state st of the user 200. As a reinforcement learning algorithm, any one of Q learning, Sarsa, Monte Carlo method, and deep reinforcement learning (reinforcement learning using DQN (Deep-Q-Network)) can be used. In the following, Q learning and deep reinforcement learning are exemplified for explanation.
[0125] The state observation unit 121 performs various processes for observing the state st of the user 200. The state observation unit 121 observes the state st of the user 200 based on at least one of the captured image Im of the user 200 and the biological information B of the user 200. The captured image Im includes at least one of a facial image and a full-body image of the user 200, and the biological information B includes information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. For example, the biological information B includes at least one of the heart rate [bpm], the respiration rate [times / minute], the blood pressure value [mmHg], and the body temperature [°C] of the user 200.
[0126] The state st of the user 200 includes at least one of the state of emotion classified based on the facial image of the user 200 and information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. In addition, the state st of the user 200 preferably also includes the state of behavior classified based on the movement of the skeleton estimated from the full-body image of the user 200. That is, the state st of the user 200 is a given state classified based on the combination of the emotion and behavior of the user 200 estimated from at least one of the captured image Im and the biological information B. In addition, the state st of the user 200 may be only the state of emotion of the user 200 estimated from the captured image Im and the biological information B, or may be only the state of behavior of the user 200 estimated from the captured image Im of the user 200.
[0127] The state observation unit 121 has an emotion behavior estimation unit 151. The emotion behavior estimation unit 151 is responsible for a process in the state observation unit 121. In addition, the function of the emotion behavior estimation unit 151 can also be assumed by other external devices connected to the learning device 300 in a communicable manner. The emotion behavior estimation unit 151 estimates the emotion of the user 200 based on the facial image of the user 200 and at least one of the information related to at least one of the heartbeat, breathing, blood pressure and body temperature of the user 200. The emotion of the user 200 is classified into a given state based on the facial image of the user 200 and the information related to at least one of the heartbeat, breathing, blood pressure and body temperature of the user 200. For example, the emotion of the user 200 is preferably classified into at least one of "indifferent", "happy", "sad", "disgusted", "fear", "surprised" and "angry". The emotion of the user 200 can also be a combination of "fear" and "surprised", etc.
[0128] The emotion behavior estimation unit 151 estimates the emotion of the user 200 using a learned learning model or performing machine learning. For example, the emotion behavior estimation unit 151 uses training data to perform deep learning on a learning model of a neural network that inputs a facial image, heart rate, and blood pressure value of the user 200 and outputs emotions such as "happiness", "sadness", and "disgust" of the user 200. Thus, if a new facial image, heart rate, and blood pressure value of the user 200 with a smiling face is input to the learning model of the neural network and analyzed by AI (artificial intelligence), the emotion of the user 200 can be classified as "happiness".
[0129] In addition, the emotional behavior estimation unit 151 estimates the behavior of the user 200 based on the movement of the skeleton estimated from the full-body image of the user 200. For example, the emotional behavior estimation unit 151 learns the position of each joint from the full-body image of the user 200, estimates the skeleton of the user 200 from the position of each joint, and estimates the behavior of the user 200 based on the movement of the estimated skeleton. The movement of the skeleton includes basic movements such as "standing", "sitting", "squatting", "walking", "right arm forward", "left arm forward", "shaking head", etc. In addition, the behavior of the user 200 includes a combination and sequence of one or more basic movements of the skeleton. For example, the behavior of the user 200 is classified into at least one of "no movement", "walking", "sitting", "operating a mobile phone", "cooking", "typing the keyboard", "watching TV", and "sleeping".
[0130] The emotional behavior estimation unit 151 estimates the behavior of the user 200 using a learned learning model or while performing machine learning. For example, the emotional behavior estimation unit 151 uses training data to perform deep learning on a learning model of a first neural network that inputs a full-body image of the user 200 and outputs the positions of each joint of the user 200. Secondly, the emotional behavior estimation unit 151 uses training data to perform deep learning on a learning model of a second neural network that inputs information related to changes in the skeleton of the user 200 (changes in the positions of each joint) and outputs the basic movements of the skeleton of the user 200. Then, the emotional behavior estimation unit 151 estimates the behavior of the user 200 based on the combination and sequence of the basic movements of the skeleton. Thus, if a new full-body image of the user 200 operating a mobile phone is input into the learning model of the neural network and AI analysis is performed, the behavior of the user 200 can be classified as "operating a mobile phone."
[0131] The storage unit 126 stores information related to the state sn (n is the identification number of the state s) of the user 200. The information related to the state sn of the user 200 is managed, for example, by a table of a database. Table 2 below is an example of a state table TB2 related to the state sn of the user 200. The state table TB2 contains a state ID for identifying the state sn of the user 200, and the state content of the user 200. The state sn of the user 200 exists in a number corresponding to the number of combinations of emotions and behaviors of the user 200 that have been defined in advance. In addition, in a case where the function of the inference unit 111 is set inside the robot 100, the state table TB2 is stored by the storage unit 103 of the robot 100.
[0132]
Table 2
[0133] Status ID Status content s0 Emotion: Flat, Behavior: Walking s1 Emotion: Happiness, Behavior: Walking s2 Emotion: Sad, Behavior: Walking s3 Emotion: Disgust, Behavior: Walking s4 Emotion: Fear, Behavior: Walking s5 Emotion: Surprised, Behavior: Walking s6 Emotion: Angry, Behavior: Walking s7 Emotion: Flat, Behavior: Sitting s8 Emotion: Happiness, Behavior: Sitting s9 Emotion: Sad, Behavior: Sitting s ···
[0134] For example, when the state st of the user 200 is "emotion: calm, behavior: typing on the keyboard", it is observed that the user 200 is in a state of being calm but busy with work, etc. Therefore, if the robot 100 performs a certain behavior at, it may cause stress to the user 200. When the user 200 has negative emotions such as "a little busy now" or "depressed", the user 200 will be annoyed with the robot 100, and the robot 100 will no longer provide healing for the user 200.
[0135] On the other hand, when the state st of the user 200 is "emotion: happy, behavior: sitting", it is observed that the user 200 is in a relatively unbusy state with positive emotions. Therefore, if the robot 100 performs the behavior at that induces contact with the user 200, it will not cause stress to the user 200, and the possibility of being able to contact with the user 200 will increase. If the chance of being able to contact with the user 200 increases, the robot 100 can provide healing for the user 200. In this way, it can be inferred that there is a certain correlation between the state st of the user 200 and the value Q of the behavior at of the robot 100.
[0136] The behavior determination unit 122 determines the behavior at of the robot 100 relative to the state st of the user 200 based on the value Q of the behavior at-1 (t-1 is the previous time). The storage unit 126 stores the behavior value table TB3, which indicates the value Q of the robot's behavior an relative to the state sn of the user 200. The following Table 3 is an example of the behavior value table TB3 at a certain time t.
[0137]
Table 3
[0138] a0 a1 a2 a3 a4 a5 a6 a7 a8 a9 a10 an s0 +1 +1 +1 +2 +1 +1 +1 0 +2 +2 +1 … s1 +2 +1 +2 +3 +2 +1 +1 +1 +3 +3 +2 … s2 +1 0 0 +1 0 0 0 0 +1 +2 0 … s3 0 0 0 0 0 0 0 0 +1 0 0 … s4 -1 0 -2 0 +1 -1 0 0 +1 0 0 … s5 -2 +1 -1 -1 +2 -1 +1 0 +1 -1 -1 … s6 -1 +1 -2 -2 +2 -1 +1 +1 +1 -2 -1 … s7 0 0 +1 +1 0 +1 0 0 +1 +2 +1 … s8 +1 0 +1 +2 +1 0 0 0 +2 +2 +1 … s9 0 0 0 0 0 0 0 0 +1 0 0 … s … … … … … … … … … … … …
[0139] In the initial state of the behavior value table TB3 (time t=0, etc.), the value Q of the robot 100's behavior at relative to the state st of the user 200 is unknown. Therefore, the behavior determination unit 122 preferably initializes the value Q of all behaviors an with a random number and selects one behavior at from the given behaviors an.
[0140] In addition, if the behavior decision unit 122 only continuously selects the behavior at with the highest value Q for learning, it will not transition to a state st+1 (t+1 is the next time) that has not been experienced yet. Therefore, the behavior decision unit 122 preferably uses the ε-greedy method, etc., to select the behavior at with the highest value Q with probability 1-ε, and selects one behavior at from all behaviors an with probability ε.
[0141] For example, when the state st of the user 200 is "emotion: sad, behavior: walking", the behavior decision unit 122 selects the behavior at of "dancing" with the highest value Q with a probability of 0.9 (ε=0.1). As a result, the possibility of inducing contact with the user 200 is increased. In addition, the behavior decision unit 122 selects any behavior at from all behaviors an with a probability of 0.1. As a result, the user 200 will feel that the robot 100 has selected the behavior at under free will and will not be annoyed by the robot 100. In addition, when there is no highest value Q but there are multiple identical values Q, the behavior decision unit 122 selects any behavior at from the highest parallel values Q with a random number.
[0142] The communication control unit 125 sends an execution instruction of the behavior at determined by the behavior determination unit 122 to the robot 100. The robot 100 receives the execution instruction of the behavior at through the communication control unit 102. Then, the behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to execute the behavior at of the robot 100. Thus, the robot 100 executes the behavior at of inducing contact with the user 200 according to the state st of the user 200.
[0143] The result acquisition unit 123 acquires information related to the result of contact with the user 200 as the result of the behavior at of the robot 100. The information related to the result of contact with the user 200 preferably includes information related to the proximity of the user 200, information related to the emotion of the user 200, and information related to the time of contact with the user.
[0144] The information related to the approach of the user 200 includes at least the presence or absence of the approach of the user 200 (whether the user 200 approaches the robot 100). In addition, the information related to the emotion of the user 200 includes at least the positive or negative emotion level of the user 200. Furthermore, the time of contact with the user 200 includes at least the duration of contact with the user 200.
[0145] The result acquisition unit 123 includes a proximity information acquisition unit 152, an emotion level estimation unit 153, and a contact time acquisition unit 154. In addition, the function of at least one of the proximity information acquisition unit 152, the emotion level estimation unit 153, and the contact time acquisition unit 154 may be performed by another external device connected to the learning device 300 in a communicable manner. In addition, the function of the emotion level estimation unit 153 may also be performed by the emotional behavior estimation unit 151.
[0146] The approach information acquisition unit 152 acquires information related to the approach of the user 200 (whether the user 200 is approaching) based on various sensor information such as the captured image Im (distance image) and the first electrostatic capacitance signal C1 or the second electrostatic capacitance signal C2. For example, when the distance from the robot 100 to the user 200 is greater than a given threshold value (for example, greater than 1 meter), the approach information acquisition unit 152 acquires information indicating that the user 200 is not approaching. In addition, when the distance from the robot 100 to the user 200 is less than a given threshold value (for example, less than 1 meter), the approach information acquisition unit 152 acquires information indicating that the user 200 is approaching. Furthermore, it is preferable that the approach information acquisition unit 152 also acquires the approach speed of the user 200 relative to the robot 100.
[0147] The emotion level estimation unit 153 estimates the emotion level of the user 200 based on the facial image of the user 200 and at least one of the information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. For example, the emotion level estimation unit 153 estimates the emotion level to be “neutral” when the emotion of the user 200 is classified as “neutral”, and estimates the emotion level to be “very positive” when the emotion is classified as “happy”. In addition, the emotion level estimation unit 153 estimates the emotion level to be “negative” when the emotion of the user 200 is classified as “sad”, and estimates the emotion level to be “very negative” when the emotion is classified as “disgusted”.
[0148] The contact time acquisition unit 154 acquires information related to the time of contact with the user 200 based on various sensor information such as the tactile signal S. For example, the contact time acquisition unit 154 calculates the total time of all the on periods until the tactile signal S is completely turned off in a given period after the robot 100 performs the behavior at, thereby acquiring the duration of contact with the user 200.
[0149] Through the above, the result acquisition unit 123 acquires information related to the result of the contact with the user 200 (such as the presence or absence of the user 200, the emotional level of the user 200, and the duration of the contact with the user 200) as the result of the behavior at of the robot 100.
[0150] The learning unit 124 generates a learning model LM that inputs the state st (e.g., a combination of emotions and behaviors) of the user 200 and outputs the value Q (st, at) of the behavior at of the robot 100 through reinforcement learning. In addition, the learning unit 124 updates the learning model LM based on the result of the contact with the user 200. In the present embodiment, the learning unit 124 obtains a reward r corresponding to the behavior at of the robot 100 based on the result of the contact with the user 200, and updates the value Q (behavior value table TB3) of the behavior at corresponding to the state st of the user 200 based on the reward r.
[0151] The learning unit 124 includes a reward acquisition unit 155 and a value update unit 156. The reward acquisition unit 155 acquires a reward r corresponding to the behavior at of the robot 100 based on the result of the contact with the user 200 (e.g., the proximity of the user 200, the emotional level of the user 200, and the duration of the contact with the user 200).
[0152] The storage unit 126 stores information related to a given report rn (n is an identification number of the report r) based on the result of contact with the user 200. The information related to the report rn is managed, for example, by a table in a database. The following Table 4 is an example of a report table TB4 representing a given report rn. The report table TB4 contains a report ID for identifying the report rn, the result of contact with the user 200, and the report rn based on the result of contact with the user 200. The number of reports rn corresponds to the number of results of contact with the user 200 that have been defined in advance.
[0153]
Table 4
[0154] Report ID The result of contact with users Rewards r0 Approach: None, Emotional level: *, Contact duration: * 0 r1 Approach: Yes, Emotional level: Normal, Contact duration: 1s-3s +1 r2 Approach: Yes, Emotional level: Positive, Contact duration: 1s-3s +2 r3 Approach: Yes, Emotional level: Very positive, Contact duration: 1s-3s +3 r4 Approach: Yes, Emotional level: Negative, Contact duration: 1s-3s -1 r5 Approach: Yes, Emotional level: Very negative, Contact duration: 1s-3s -2 r6 Approach: Yes, Emotional level: Normal, Contact duration: 3s-5s +2 r7 Approach: Yes, Emotional level: Positive, Contact duration: 3s-5s +3 r8 Approach: Yes, Emotional level: Very positive, Contact duration: 3s-5s +5 r9 Approach: Yes, Emotional level: Negative, Contact duration: 3s-5s +2 rn ··· ···
[0155] When the user 200 approaches the robot 100 (approach presence or absence: yes) and the time of contact with the user 200 is longer than a given threshold (e.g., longer than 1 second), the reward acquisition unit 155 determines that contact with the user 200 is established. In this case, the reward acquisition unit 155 acquires a given reward rn corresponding to the emotion level of the user 200 and the duration of contact with the user 200. In addition, although not shown in Table 4, the reward acquisition unit 155 may also acquire a given reward rn corresponding to the approach speed of the user 200 in addition to the emotion level of the user 200 and the duration of contact with the user 200.
[0156] When the emotion level of the user 200 is positive, the reward rn may be defined as a positive reward rn as the duration of contact with the user 200 is longer. When the structure of the robot 100 is consistent with the intention of the user 200, it is expected to increase the frequency of contact with the user 200. In addition, when the emotion level of the user 200 is negative, the reward rn may be defined as a negative reward rn as the duration of contact with the user 200 is longer. When the structure of the robot 100 is consistent with the intention of the user 200, it is expected to prevent the frequency of contact with the user 200 from decreasing.
[0157] On the other hand, when the user 200 does not approach the robot 100 (approach presence or absence: no), or when the duration of contact with the user 200 is less than a given time, the reward acquisition unit 155 determines that contact with the user 200 is not established. In this case, the reward acquisition unit 155 acquires a reward rn of zero.
[0158] The value updating unit 156 updates the value Q of the behavior at of the robot 100 relative to the state st of the user 200 based on the given reward rn. In Q-learning, the value Q is updated by the following formula 1.
[0159] [Formula 1]
[0160] Q(st,at)=Q(st,at)+α(r+γmax Q(st+1,at+1)-Q(st,at)) ...Equation 1
[0161] In Formula 1, st is the state of the user 200 at a certain time t, and at is the behavior of the robot 100 at a certain time t. Through the behavior at of the robot 100, the state of the user 200 becomes st+1 (t+1 is the next time). r is the reward obtained by the change in the state of the user 200. In addition, the item with max is the item obtained by multiplying the value Q of the behavior at+1 with the highest value Q known at this time when the state is st+1 by the discount rate γ (0<γ≦1). In addition, α is a learning coefficient (0<α≦1) used to adjust the learning speed.
[0162] Formula 1 represents the following method: based on the reward r returned as a result of the behavior at of the robot 100, the value Q(st,at) of the behavior at relative to the state st of the user 200 is updated. When the value Q of a certain behavior at of the robot 100 relative to a certain state st of the user 200 is smaller than the total value of its reward r and the discounted value Q of the best behavior at+1 relative to the next state st+1, the value Q(st,at) is increased. On the contrary, when the value Q(st,at) of the behavior at of the robot 100 relative to the state st of the user 200 is larger than the total value of its reward r and the discounted value Q of the best behavior at+1 relative to the next state st+1, the value Q(st,at) is reduced. Therefore, Formula 1 makes the value Q of a certain behavior at in a certain state st close to the total value of the reward r as a result and the discounted value Q of the best behavior at+1 relative to the next state st+1.
[0163] The value updating unit 156 updates the value Q(sn, an) of the behavior value table TB3 by using equation 1. Then, the state observation unit 121 observes the state st+1 of the next user 200. The behavior determination unit 122 determines the behavior at+1 of the robot 100 corresponding to the state st+1 of the next user 200 based on the value Q of the behavior at (in this example, the behavior value table TB3) by using the ε-greedy method or the like.
[0164] The communication control unit 125 sends an execution instruction of the determined behavior at+1 to the robot 100. The robot 100 receives the instruction of the behavior at+1 from the learning device 300 via the communication control unit 102. Then, the behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to execute the behavior at+1 of the robot 100. As a result, the robot 100 executes the behavior at+1 that matches the state st+1 of the user 200.
[0165] Here, regarding the method of expressing the value Q(st, at) on a computer, there is the following method: as described above, for all combinations of the states sn of the users 200 and the behaviors an of the robots 100, the value Q(st, at) is saved in advance in the form of a behavior value table TB3. In addition, there is the following method: prepare a behavior value function that is approximate to the behavior value table TB3. The latter method can be achieved by adjusting the parameters of the approximate function using a method such as a probabilistic gradient descent method. For example, as an approximate function, the learning unit 124 preferably generates a learning model of a neural network (DQN) that inputs the state st (a combination of emotions and behaviors) of the user 200 and outputs the value Q of the behavior at of the robot 100 through deep reinforcement learning.
[0166] The following is an explanation of deep reinforcement learning, but before that, a description of neural networks is given. Fig. 9 is a diagram schematically showing a learning model of a neuron. Fig.10 It is a schematic representation of Fig. 9 The neural network is composed of a three-layer neural network composed of a combination of neurons as shown in FIG. Fig. 9 The neuron (single-layer perceptron) model shown is composed of a computing device, a memory, etc.
[0167] like Fig. 9 As shown, the neuron output is related to multiple inputs x (in Fig. 9 As an example, the output (result) y is relative to the input x1 to input x3. For each input x (x1, x2, x3), a weight w (w1, w2, w3) corresponding to the input x is assigned. As a result, the neuron outputs the output y expressed by the following formula 2. In addition, the input x, output y and weight w are all vectors. In addition, in the formula 2 described later, θ is the deviation rate and fk is the activation function.
[0168] [Formula 2]
[0169]
[0170] Fig.10 It is shown in Fig. 9 The three-layer neural network is composed of the neurons shown in Fig.10 As shown, multiple inputs x (here, as an example, inputs x1 to x3) are input from the left side of the neural network, and results y (here, as an example, outputs y1 to y3) are output from the right side. Specifically, inputs x1, x2, and x3 are input to the three neurons N11 to N13 with corresponding weights assigned to them. The weights assigned to these inputs are collectively referred to as W1.
[0171] Neurons N11~N13 output z11~z13 respectively. Fig.10 In the above example, z11 to z13 are collectively referred to as feature vector Z1, which can be regarded as a vector that extracts the feature value of the input vector. This feature vector Z1 is a feature vector between weight W1 and weight W2. z11 to z13 are respectively assigned corresponding weights to the two neurons N21 and N22 and input. The weights assigned to these feature vectors are collectively referred to as W2.
[0172] Neurons N21 and N22 output z21 and z22 respectively. Fig.10 In the above example, z21 and z22 are collectively referred to as feature vector Z2. This feature vector Z2 is a feature vector between weight W2 and weight W3. z21 and z22 are input to three neurons N31 to N33 with corresponding weights assigned to them. The weights assigned to these feature vectors are collectively referred to as W3.
[0173] Finally, neurons N31 to N33 output outputs y1 to y3, respectively. In the operation of the neural network, there is a learning mode for learning weights W1 to W3 of the neural network, and an estimation mode for estimating outputs y1 to y3 based on inputs x1 to x3. For example, in the learning mode, the weights W1 to W3 are learned using a learning data set, and the behavior at of the robot 100 is determined using the parameters in the estimation mode. In addition, although it is written as "estimation" for convenience, it goes without saying that various tasks such as detection and classification can be performed.
[0174] In addition, weights W1 to W3 can be learned by backpropagation. Error information enters from the right side of the neural network and flows to the left side. Backpropagation is a method of adjusting (learning) the weights of each neuron in a way that reduces the difference (error) between the output y when input x is input and the true output y (label data).
[0175] Such a neural network can also be further increased in layers to perform deep learning. In addition, a convolutional neural network (CNN) that extracts input features in stages and a computing device for a neural network that classifies or regresses outputs can also be automatically obtained based only on training data.
[0176] In the above-mentioned behavior value table TB3, when the number of states sn of the user 200 and the number of behaviors an of the robot 100 are large, the memory space of the behavior value table TB3 becomes too large. Therefore, by using a neural network (DQN) to perform function approximation on the behavior value table TB3, the increase of the memory space can be prevented.
[0177] Refer again Figure 8 , the structure of the estimation unit 111 for deep reinforcement learning will be described. The learning unit 124 includes a target network TN (value Q (st, at) | θ - ) and a Q network QN (value Q(st,at)|θ). The two networks are stored in the storage unit 126. The two networks have the same structure, but the parameters θ (equivalent to the above-mentioned weights) are different. The inputs of the two networks are both the state st of the user 200, and the outputs are both the value Q(st,at) of the behavior at of the robot 100.
[0178] The state observation unit 121 observes the state st of the user 200 and outputs it to the behavior decision unit 122 and the learning unit 124. The behavior decision unit 122 inputs the state st of the user 200 into the target network TN and determines the value Q(st, at|θ) of the behavior at output from the target network TN. - ), the behavior at of the robot 100 is determined by the ε-greedy method or the like. The communication control unit 125 sends an execution instruction of the determined behavior at to the robot 100, and the robot 100 executes the behavior at according to the execution instruction of the behavior at.
[0179] The result acquisition unit 123 acquires the result of the contact with the user 200 (the presence or absence of the user 200, the emotional level of the user 200, and the duration of the contact with the user 200) as the result of the behavior at of the robot 100, and outputs it to the learning unit 124. The learning unit 124 acquires the reward r based on the result of the contact with the user 200. In addition, the state observation unit 121 observes the state st+1 of the next user 200, and outputs it to the behavior determination unit 122 and the learning unit 124.
[0180] The learning unit 124 stores the experience et (<st, at, st+1, r>) of the robot 100 in the storage unit 126 as an experience buffer. Here, st is the state of the user 200, at is the behavior of the robot 100, st+1 is the state of the next user 200, and r is the reward. In addition, the learning unit 124 preferably clips the reward r within the range of -1 to +1 so as not to overreact to abnormal values, etc. (so-called reward clipping).
[0181] The learning unit 124 periodically obtains arbitrary experience et from the storage unit 103 (Experience Buffer) and makes the Q network QN learn it. For example, the learning unit 124 obtains experience (B = e0 ~ en) for small batch learning B from the storage unit 126. Then, the learning unit 124 updates the parameter θ of the Q network QN in a way that minimizes the TD (Temporal Difference) error L (θ) shown in the following formula 3 (so-called experience replay).
[0182] [Formula 3]
[0183]
[0184] Next, the learning unit 124 reflects the parameters θ of the Q network QN in the target network TN at arbitrary intervals. The learning unit 124 may periodically copy all the parameters θ of the Q network QN to the target network TN, or may reflect the parameters θ of the Q network QN little by little each time the parameters θ of the Q network QN are updated.
[0185] The behavior decision unit 122 inputs the state st+1 of the next user 200 to the target network TN. Then, the behavior decision unit 122 determines the value Q(st+1, at+1|θ) of the behavior at+1 output from the target network TN. -), the behavior at+1 of the robot 100 is determined by the ε-greedy method, etc. The communication control unit 125 sends an execution instruction of the determined behavior at+1 to the robot 100, and the robot 100 executes the behavior at+1 according to the execution instruction of the behavior at+1.
[0186] Through the above, the estimating unit 111 can perform deep reinforcement learning and estimate the behavior at of the robot 100 that is suitable for the state st of the user 200 .
[0187] <Processing Example of Control Unit 13>
[0188] Fig.11 It is a flowchart which illustrates the processing of the control unit 13. Fig.11 The following process is shown: the control unit 13 instructs the execution of the action at of inducing contact with the user 200 according to the state st of the user 200.
[0189] First, in step S10, the control unit 13 obtains at least one of the captured image Im of the user 200 and the biometric information B of the user 200 through the acquisition unit 101. The captured image Im includes at least one of the face image and the whole body image of the user 200, and the biometric information B includes information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200.
[0190] In step S10, power is supplied from the battery 15 to the camera 11, the tactile sensor 12, the first capacitance sensor 21, the second capacitance sensor 31, and the life sensor 14. However, in order to reduce power consumption of the battery 15, power may not be supplied to the servo motor 35, the display 24, the speaker 25, and the lamp 26.
[0191] Next, in step S11, the control unit 13 estimates the given behavior at of the robot 100 that is adapted to the state st of the user 200 based on at least one of the captured image Im and the biological information B through the estimation unit 111. The control unit 13 preferably estimates the given behavior at of the robot 100 that is adapted to the state st of the user 200 through the estimation unit 111 while performing reinforcement learning or using a learned learning model. In addition, the processing in step S11 is not necessary processing by the control unit 13, and may also be performed in an external device (a learning device described later) that is connected to the robot 100 in a communicative manner. For detailed processing in step S11, refer to Fig.12 Details will be given separately.
[0192] Then, in step S12, the control unit 13 instructs the execution of the behavior at of inducing contact with the user 200 according to the state st of the user 200 through the behavior control unit 112. After the robot 100 executes the behavior at, the control unit 13 repeatedly performs the processing of steps S10 to S12, thereby inducing contact with the user 200, thereby providing healing to the user 200.
[0193] As described above, the control unit 13 performs processing to instruct execution of a behavior of inducing contact with the user 200 according to the state st of the user 200. In addition, in order to improve the learning processing capability or to suppress the power consumption of the battery 15, when the learning device 300 connected to the robot 100 in a communicable manner assumes the function of the estimation unit 111, the processing of step S11 is performed by the learning device 300.
[0194] In addition, Fig.11 When the process shown is started, the servo motor 35, the display 24, the speaker 25 and the lamp 26 may be in a standby state (sleep state) with a suppressed power supply. That is, the control unit 13 preferably restores various devices from the standby state with a suppressed power supply as needed, thereby suppressing the power consumption of the battery 15.
[0195] <Processing by the Estimation Unit 111>
[0196] Fig.12 It is a flowchart showing the processing of the estimation unit 111 (for example, the learning device 300 ). Fig.12 The following process is shown: the estimating unit 111 performs reinforcement learning to estimate a given behavior at of the robot 100 that is suitable for the state st of the user 200 . Fig.12 The steps shown are Fig.11 The detailed processing of step S11 is shown.
[0197] First, in step S20, the estimation unit 111 observes the state st of the user 200 through the state observation unit 121 based on at least one of the captured image Im and the biological information B. The estimation unit 111 estimates the emotion of the user 200 based on the facial image of the user 200 and at least one of the information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. In addition, the estimation unit 111 estimates the behavior of the user 200 based on the movement of the skeleton estimated from the whole body image of the user 200. Then, the estimation unit 111 preferably observes the given state st of the user 200 classified according to the combination of the emotion and behavior of the user 200.
[0198] Next, in step S21, the estimation unit 111 determines the behavior at of the robot 100 that is adapted to the state st of the user 200 based on the value Q of the behavior at-1 (the above-mentioned behavior value table TB3 or the learning model LM such as DQN) through the behavior determination unit 122. The estimation unit 111 outputs the execution instruction of the behavior at to the behavior control unit 112 or sends it to the robot 100 via the communication control unit 125, so that the robot 100 executes the behavior at (step S12).
[0199] In addition, step S20 and step S21 are an estimation phase for estimating the behavior at of the robot 100 that matches the state of the user 200 , and the other steps are a learning phase.
[0200] In step S22, the estimation unit 111 obtains information related to the result of contact with the user 200 as the result of the behavior at of the robot 100 through the result acquisition unit 123. The result of contact with the user 200 preferably includes, for example, the presence or absence of the user 200 approaching, the emotional level of the user 200, and the duration of contact with the user 200.
[0201] In step S23 , the estimation unit 111 obtains the reward r corresponding to the behavior at of the robot 100 based on the result of the contact with the user 200 through the reward acquisition unit 155 .
[0202] In step S24 , the estimating unit 111 updates the value Q of the behavior at of the robot 100 relative to the state st of the user 200 based on the reward r through the value updating unit 156 .
[0203] After learning, the process returns to step S20, and the estimation unit 111 observes the state st+1 of the next user 200 through the state observation unit 121. Then, in step S21, the estimation unit 111 determines the behavior at+1 of the robot 100 that is adapted to the state st+1 of the next user 200 based on the value Q of the updated behavior at through the behavior determination unit 122. Then, the robot 100 executes the next behavior at+1 (step S12).
[0204] In addition, after step S24, the following step may be provided: the estimation unit 111 determines whether the value Q of the behavior at has converged (i.e., whether the learning has converged) through the value updating unit 156. When the estimation unit 111 determines that the learning has converged, the learning phase may not be executed in the subsequent processing. That is, the estimation unit 111 only executes the estimation phase, and estimates the behavior at+n (t+n is the time after n times) of the robot 100 that is adapted to the state st+n of the user 200 using the learned learning model LM (behavior value table TB3 or DQN, etc.).
[0205] <Configuration of Estimation Unit 111 of Modification Example>
[0206] Fig.13 1 is a block diagram showing the functional configuration of the estimation unit 111 according to the modified example. Figure 8 The difference in the functional structure of the estimation unit 111 shown in the figure is that it performs supervised learning to estimate the behavior at of the robot 100 that is suitable for the state st of the user 200. That is, the learning unit 124 includes a training data recording unit 157, an error calculation unit 158, and a learning model updating unit 159. Figure 8 The differences in the configuration of the estimating unit 111 shown in FIG. 1 will be described.
[0207] The function of the training data recording unit 157 can be realized by a nonvolatile memory such as HDD / SSD 304. In addition, the functions of the error calculation unit 158 and the learning model update unit 159 can be realized by a processor such as CPU 301 executing processing specified by a program stored in a nonvolatile memory such as ROM 302.
[0208] The learning unit 124 can use a decision tree (regression tree), a neural network, or a logistic regression, etc., as a learning model LM for supervised learning. Hereinafter, an example of a learning model LM of a neural network that generates a value Q of the behavior at of the robot 100 and inputs the state st of the user 200 through supervised learning by the learning unit 124 will be described.
[0209] The training data recording unit 157 stores training data obtained in the past by other robots 100 or simulations, for example. The training data is result (label) data including the state st-n (tn is the time n times ago) of the user 200, the behavior at-n of the robot 100, and the value Q (equivalent to a label) of the behavior at-n. The estimation unit 111 receives training data from other robots 100 or other external devices via the communication control unit 125, etc. In addition, the estimation unit 111 may store the experience experienced by the robot 100 itself as training data.
[0210] The error calculation unit 158 first obtains the training data from the training data recording unit 157, and calculates the error L of the value Q of the behavior at based on the training data. For example, when the contact with the user 200 is actually established, the error calculation unit 158 regards that there is an error of -log(Q(st,at)) and calculates the error L. In addition, when the contact with the user 200 is not actually established, the error calculation unit 158 regards that there is an error of -log(1-Q(st,at)) and calculates the error L.
[0211] The learning model updating unit 159 updates the parameters (the weights, etc.) of the learning model LM of the neural network in a manner that minimizes the error L. In updating the learning model LM, the error back propagation method (Backpropagation) can be used. Thus, the learning unit 124 generates a learning model LM that has been learned to a certain level through training data.
[0212] Then, the estimating unit 111 estimates the behavior at that matches the actual state st of the user 200 using the learning model LM generated by supervised learning. Then, the robot 100 performs the behavior at that induces contact with the user 200 according to the state st of the user 200 .
[0213] More specifically, the state observation unit 121 observes the state st of the user 200, and the behavior determination unit 122 uses the learning model LM to determine a given behavior at that is suitable for the state st of the user 200. Then, the communication control unit 125 sends an execution instruction of the behavior at to the robot 100, and the robot 100 executes the behavior at according to the received execution instruction of the behavior at.
[0214] The result acquisition unit 123 acquires the result of the contact with the user 200 as the result of the behavior at of the robot 100. The error calculation unit 158 calculates the error of the value Q of the behavior at based on the result of the contact with the user 200, and the learning model updating unit 159 further updates the learning model LM of the neural network so as to minimize the error L. Then, the behavior determination unit 122 determines the behavior at+1 of the robot 100 that is suitable for the next state st+1 of the user 200 using the learning model LM.
[0215] As described above, the estimation unit 111 can estimate the behavior of the robot 100 that is suitable for the state st of the user 200 using the learning model LM learned to a certain level through supervised learning. For example, even if the robot 100 fails and is replaced with a robot 100 of the same model, the replaced robot 100 can learn from past experience based on the training data of the failure, and thus immediately perform the behavior at that is suitable for the state st of the emotion of the user 200. In addition, the robot 100 can also perform the behavior at that is suitable for the state st of the user 200 to a certain level for the user 200 that it has contacted for the first time.
[0216] <Processing of the Estimation Unit 111 of Modification Example>
[0217] Fig.14 This is a flowchart showing the processing of the estimating unit 111 (learning device 300 ) according to the modification. Fig.14The following process is shown: the estimating unit 111 performs supervised learning to estimate the behavior at of the robot 100 that is suitable for the state st of the user 200 . Fig.14 The steps shown are Fig.11 The detailed processing of step S11 is shown.
[0218] First, in step S30 , the estimation unit 111 obtains the training data from the training data recording unit 157 via the error calculation unit 158 , and calculates the error L of the value Q of the behavior at of the robot 100 based on the training data.
[0219] Next, in step S31, the estimation unit 111 updates the parameters (the weights, etc., described above) of the learning model LM of the neural network through the learning model updating unit 159 in a manner that minimizes the error L. Thus, the estimation unit 111 can estimate the behavior at that matches the actual state st of the user 200 using the learning model LM learned to a certain level through the training data.
[0220] Then, in step S32, the estimation unit 111 observes the actual state st of the user 200 through the state observation unit 121 based on at least one of the captured image Im and the biological information B. The estimation unit 111 estimates the emotion of the user 200 based on the facial image of the user 200 and at least one of the information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. In addition, the estimation unit 111 estimates the behavior of the user 200 based on the movement of the skeleton estimated from the whole body image of the user 200. Then, the estimation unit 111 preferably observes the given state st of the user 200 classified according to the combination of the emotion and behavior of the user 200.
[0221] Next, in step S33, the estimation unit 111 determines the behavior at of the robot 100 that matches the state st of the user 200 based on the value Q of the behavior at-1 (learning model LM such as a neural network) through the behavior determination unit 122. The estimation unit 111 outputs the execution instruction of the behavior at to the behavior control unit 112 or sends it to the robot 100 via the communication control unit 125. As a result, the robot 100 executes the behavior at corresponding to the state st of the user 200 (step S12).
[0222] In addition, step S32 and step S33 are an estimation phase for estimating the behavior at of the robot 100 that matches the state of the user 200, and the other steps are a learning phase.
[0223] In step S34, the estimation unit 111 obtains information related to the result of contact with the user 200 as the result of the behavior at of the robot 100 through the result acquisition unit 123. The information related to the result of contact with the user 200 preferably includes the presence or absence of the user 200 approaching, the emotional level of the user 200, and the duration of contact with the user 200.
[0224] Next, returning to step S30 , the estimation unit 111 calculates the error L of the value Q of the action at based on the result of the contact with the user 200 through the error calculation unit 158 .
[0225] Next, in step S31 , the estimation unit 111 updates the parameters (weights, etc.) of the learning model LM based on the error L through the learning model updating unit 159 .
[0226] After learning, in step S32, the estimation unit 111 observes the state st+1 of the next user 200 through the state observation unit 121. Then, in step S33, the estimation unit 111 determines the behavior at+1 of the robot 100 that is adapted to the state st+1 of the next user 200 based on the value Q of the updated behavior at through the behavior determination unit 122. Then, the robot 100 executes the next behavior at+1 (step S12).
[0227] In addition, after step S31, the following step may be provided: the estimation unit 111 determines whether the value Q of the behavior at has converged (i.e., whether the learning has converged) through the learning model updating unit 159. When the estimation unit 111 determines that the learning has converged, the learning phase may not be performed in the subsequent processing. That is, the estimation unit 111 only performs the estimation phase, and estimates the behavior at+n of the robot 100 that is adapted to the state st+n (t+n is the time after n times) of the user 200 using the learning model LM (neural network, etc.) that has been learned.
[0228] <Function and Effect of the Present Embodiment>
[0229] As described above, the robot 100 estimates a given behavior at that matches the state st of the user 200 based on at least one of the captured image Im of the user 200 and the biological information B of the user 200. Then, the robot 100 performs the behavior at that induces contact with the user 200 according to the state st of the user 200. Therefore, the robot 100 can induce contact with the user 200 at an appropriate timing. Furthermore, the contact with the user 200 can be continuously established, and healing can be continuously provided to the user 200.
[0230] In addition, the captured image Im includes at least one of a face image and a whole body image of the user 200, and the biological information B includes information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. Compared with the case where the user's state st is observed only based on the captured image Im or only based on the biological information B, the state st of the user 200 can be observed with high accuracy.
[0231] It is particularly preferred that the state st of the user 200 includes the user's emotion classified based on the face image of the user 200 and at least one of the information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200. Even if the user 200 is smiling temporarily, it is not necessarily in a state of having a positive emotion. Therefore, by adding not only the face image of the user 200 but also the information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200, the state st of the user 200 can be observed with high accuracy.
[0232] Furthermore, the state st of the user 200 preferably includes the behavior of the user 200 classified based on the motion of the skeleton estimated from the full-body image of the user 200. Even if the user 200 is in a positive mood, it is possible that the user 200 is busy with housework or work, etc. Therefore, by observing the state st of the user 200 based not only on the emotion of the user 200 but also on the behavior of the user 200, the state st of the user 200 can be observed with high accuracy.
[0233] Furthermore, the state st of the user is a given state sn classified according to a combination of the emotion and behavior of the user 200 estimated from at least one of the captured image Im and the biometric information B. Thus, the state sn of the user 200 is classified into dozens to hundreds of types, thereby suppressing an increase in the memory space of the behavior value table TB3.
[0234] In addition, the robot 100 generates a learning model LM that inputs the state st of the user 200 and outputs the value Q of the behavior at of the robot 100 through machine learning (reinforcement learning or supervised learning, etc.). Therefore, the robot 100 can learn the behavior at that matches the state st of the user 200 and can perform the behavior at that induces contact with the user 200 at an appropriate timing. Furthermore, the contact with the user 200 can be continuously established, and the user 200 can be continuously provided with healing.
[0235] Furthermore, the robot 100 further includes a result acquisition unit 123 that acquires information related to the result of the contact with the user 200, and the learning unit 124 updates the learning model LM based on the result of the contact with the user 200. Therefore, the robot 100 can continuously learn the behavior at that matches the state st of the user 200, and can execute the behavior at that induces the contact with the user 200 at an appropriate timing. In addition, the learning model LM can improve the reliability of the learning ability of the robot 100 by using the behavior value table TB3 or the neural network (DQN) whose effect has been proven.
[0236] In addition, the information related to the result of the contact with the user 200 includes the presence or absence of the user 200 approaching, the emotional level of the user 200, and the duration of the contact with the user 200. Therefore, the robot 100 can determine whether the contact with the user 200 is established based on the presence or absence of the user 200 approaching and the duration of the contact with the user 200. Then, the learning unit 124 obtains the reward r corresponding to the behavior at of the robot 100 based on the result of the contact with the user 200, and updates the value Q of the behavior at corresponding to the state st of the user based on the reward r. Therefore, the robot 100 can learn the behavior at that is suitable for the state st of the user 200 according to the emotional level of the user 200 and the duration of the contact with the user 200.
[0237] In addition, the robot 100 can also use the learning model LM learned to a certain level through supervised learning to infer the behavior at of the robot 100 that is suitable for the state st of the user 200. As a result, for example, even if the robot 100 fails and is replaced with a robot 100 of the same model, the replaced robot 100 can learn past experience based on the training data, so that it can immediately perform the behavior at that is suitable for the state st of the user 200. In addition, the robot 100 can perform the behavior at that is suitable for the state st of the user 200 at a certain level even for the user 200 that it contacts for the first time.
[0238] Furthermore, the behavior an of the robot 100 that induces contact with the user 200 includes not only the behavior of explicitly inducing contact, but also the behavior of causing the user 200 to pay attention to the movement of the robot 100, or the behavior of implicitly inducing contact. Therefore, regardless of whether the user 200 is in a state st with negative emotions or in a busy state st, it is not easy to cause stress to the user 200. Instead, the user 200 can have a positive impression of the robot 100, such as being cute or interesting.
[0239] The functions of the estimation unit 111 of the robot 100 described above can also be provided in a learning device 300 connected to the robot 100 in a communicable manner so as to be distributed. Thus, the learning processing capability of the computer can be improved. In addition, through the distributed processing based on the learning device 300, it is possible to obtain technical effects such as reducing the power consumption of the battery 15 of the robot 100, reducing the number of charging times, and reducing the weight of the battery.
[0240] As mentioned above, although the preferred embodiment was described in detail, it is not limited to the said embodiment, Various deformation|transformation and substitution can be added to the said embodiment without departing from the scope described in a claim.
[0241] In addition, the numbers such as ordinal numbers and quantities used in the description of the above-mentioned embodiments are all illustrative for the purpose of specifically describing the technology of the present invention, and the present invention is not limited to the illustrative numbers. In addition, the connection relationship between the constituent elements is the connection relationship illustrative for the purpose of specifically describing the technology of the present invention, and the connection relationship for realizing the functions of the present invention is not limited thereto.
[0242] The robot involved in this embodiment is particularly suitable for the following purposes: promoting the secretion of oxytocin and providing healing (sense of security or self-affirmation) for single people living alone, elderly people whose children have become independent, and frail elderly people who are the objects of home medical treatment. However, it is not limited to the above purposes and can be used to provide healing for various users.
[0243] The embodiments of the present invention are as follows, for example.
[0244] <1> A robot comprising: an acquisition unit that acquires at least one of a captured image of a user and biological information of the user; and a behavior control unit that, based on at least one of the captured image and the biological information and in accordance with the state of the user, instructs the execution of a given behavior that induces contact with the user.
[0245] <2> The robot as described in <1> above also has an estimating unit for estimating the given behavior that is compatible with the state of the user, and the estimating unit has: a state observation unit that observes the state of the user based on at least one of the captured image and the biological information; and a behavior determination unit that determines the given behavior that is compatible with the state of the user based on the value of the given behavior.
[0246] <3> A robot as described in <1> or <2> above, wherein the captured image includes at least one of a facial image and a full-body image of the user, and the biological information includes information related to at least one of the user's heartbeat, respiration, blood pressure and body temperature.
[0247] <4> A robot as described in any one of <1> to <3> above, wherein the user's state includes the user's emotion classified based on the user's facial image and at least one of information related to at least one of the user's heartbeat, respiration, blood pressure and body temperature.
[0248] <5> The robot according to any one of <1> to <4> above, wherein the state of the user includes a behavior of the user classified based on a motion of a skeleton estimated from a full-body image of the user.
[0249] <6> The robot according to any one of <1> to <5> above, wherein the state of the user is a given state classified based on a combination of an emotion and a behavior of the user estimated from at least one of the captured image and the biological information.
[0250] <7> The robot according to <2> above, wherein the estimating unit includes a learning unit configured to generate a learning model that inputs the state of the user and outputs the value of the robot's behavior through machine learning.
[0251] <8> A robot as described in <7> above, wherein the estimation unit further includes a result acquisition unit, the result acquisition unit acquires information related to the result of contact with the user as a result of the robot's behavior, and the learning unit updates the learning model based on the result of contact with the user.
[0252] <9> A robot as described in <7> or <8> above, wherein the learning model is a behavior value table or a neural network.
[0253] <10> A robot as described in <8> above, wherein information related to the result of contact with the user includes the presence or absence of the user, the user's emotional level, and the duration of contact with the user, and the learning unit obtains a reward relative to the behavior of the robot based on the result of contact with the user, and updates the value of the behavior relative to the state of the user based on the reward.
[0254] <11> A learning device connected to a robot in a manner capable of communicating, comprising: a state observation unit, which observes the state of the user based on at least one of a photographed image of the user and biological information of the user; and a learning unit, which generates a learning model through machine learning that inputs the state of the user and outputs the value of the robot's behavior.
[0255] <12> The learning device as described in <11> above further comprises: a behavior determination unit, which determines the behavior of the robot adapted to the state of the user based on the value of the behavior; and a communication control unit, which sends an execution instruction of the behavior to the robot.
[0256] <13> A learning device as described in <11> or <12> above, further comprising a result acquisition unit, which acquires information related to the result of the contact with the user as a result of the behavior of the robot, and the learning unit updates the learning model based on the result of the contact.
[0257] <14> A control method, which is a control method for a robot, wherein the robot executes the following steps: a step of obtaining a photographic image of a user and at least one of the biological information of the user; and a step of inducing a given behavior of contacting the user according to the state of the user based on the photographic image and at least one of the biological information.
[0258] <15> A program that causes a computer controlling a robot to execute the following steps: a step of obtaining a photographic image of a user and at least one of the biological information of the user; and a step of instructing the execution of a given behavior of inducing contact with the user based on the state of the user and the photographic image and at least one of the biological information.
[0259] This application claims priority based on Japanese Patent Application No. 2022-156758 filed with the Japan Patent Office on September 29, 2022, and incorporates all the contents of the above-mentioned Japanese patent application.
[0260] Explanation of symbols
[0261] 1: Main body
[0262] 2: Head
[0263] 2a: Right eye
[0264] 2b: Left eye
[0265] 2c: Mouth
[0266] 2d: right cheek
[0267] 2e: Left cheek
[0268] 3: Arm
[0269] 3a: Right arm
[0270] 3b: Left arm
[0271] 4: Legs
[0272] 4a: Right leg
[0273] 4b: Left leg
[0274] 10: Exterior components
[0275] 11: Camera
[0276] 12: Tactile sensor
[0277] 13: Control Department
[0278] 14: Life sensor (electromagnetic wave sensor)
[0279] 141: Microwave Transmitter
[0280] 142: Microwave receiving unit
[0281] 15: Battery
[0282] 16: Body frame
[0283] 17: Body loading platform
[0284] 21: First electrostatic capacitance sensor
[0285] 22: Head frame
[0286] 23: Head loading platform
[0287] 24: Display
[0288] 24a: Right eye display
[0289] 24b: Left eye display
[0290] 25: Speaker
[0291] 26: Lights
[0292] 26a: Right cheek light
[0293] 26b: Left cheek light
[0294] 27: Head connection mechanism
[0295] 31: Second electrostatic capacitance sensor
[0296] 32a: Right arm frame
[0297] 32b: Left arm frame
[0298] 33: Right arm support platform
[0299] 34a: Right arm connection mechanism
[0300] 34b: Left arm connection mechanism
[0301] 35: Servo motor
[0302] 35a: Right arm servo motor
[0303] 35b: Left arm servo motor
[0304] 35c: Head servo motor
[0305] 35d: Right leg servo motor
[0306] 35e: Left leg servo motor
[0307] 41a: Right leg wheel
[0308] 41b: Left leg wheel
[0309] 42a: Right leg frame
[0310] 42b: Left leg frame
[0311] 44a: Right leg connection mechanism
[0312] 44b: Left leg connection mechanism
[0313] 100: Robot
[0314] 101: Acquisition
[0315] 102: Communication control unit
[0316] 103: Preservation Department
[0317] 104: Certification Department
[0318] 105: Registration Department
[0319] 106: Start control department
[0320] 107: Motor control unit
[0321] 108: Output unit
[0322] 109: Registration information
[0323] 110: Inspection Department
[0324] 111: Presumption Department
[0325] 112: Behavior Control Department
[0326] 121: Status Observation Department
[0327] 122: Behavior Decision Department
[0328] 123: Result acquisition unit
[0329] 124: Learning Department
[0330] 125: Communication control unit
[0331] 126: Preservation Department
[0332] 131: CPU
[0333] 132: ROM
[0334] 133: RAM
[0335] 134: HDD / SSD
[0336] 135: Device connection I / F
[0337] 136: Communication I / F
[0338] 151: Emotional Behavior Inference Department
[0339] 152: Proximity information acquisition unit
[0340] 153: Emotional level estimation
[0341] 154: Contact time acquisition unit
[0342] 155: Reward Acquisition Department
[0343] 156: Value Update Department
[0344] 157: Training Data Recording Department
[0345] 158: Error calculation unit
[0346] 159: Learning Model Update Department
[0347] 200: User
[0348] 300: Learning device
[0349] 301: CPU
[0350] 302: ROM
[0351] 303: RAM
[0352] 304: HDD / SSD
[0353] 305: Device connection I / F
[0354] 306: Communication I / F
[0355] A, A': system bus
[0356] B: Biological information
[0357] C1: First electrostatic capacitance signal
[0358] C2: Second electrostatic capacitance signal
[0359] F1a: Right shoulder frame
[0360] F2a: Right upper arm frame
[0361] F3a: Right elbow frame
[0362] F4a: Right forearm frame
[0363] F1b: Left shoulder frame
[0364] F2b: Left upper arm frame
[0365] F3b: Left elbow frame
[0366] F4b: Left forearm frame
[0367] F1c: Neck frame
[0368] F2c: Face framing
[0369] Im: Take an image
[0370] L: Error
[0371] LM: Learning Model
[0372] Ms: Emission wave
[0373] Mr: Reflection wave
[0374] M1a: Right shoulder servo motor
[0375] M2a: Right upper arm servo motor
[0376] M3a: Right elbow servo motor
[0377] M4a: Right forearm servo motor
[0378] M1b: Left shoulder servo motor
[0379] M2b: Left upper arm servo motor
[0380] M3b: Left elbow servo motor
[0381] M4b: Left forearm servo motor
[0382] M1c: Neck servo motor
[0383] M2c: Face servo motor
[0384] Q: Value
[0385] S: Tactile signal
[0386] s: Status
[0387] a: behavior
[0388] r: return.
Claims
1. A robot comprising: an acquisition unit that acquires at least one of an image of a user and biological information of the user; and A behavior control unit instructs, based on at least one of the captured image and the biological information, to induce execution of a given behavior of contact with the user in accordance with the state of the user.
2. The robot according to claim 1, wherein: The robot further includes an estimating unit for estimating the given behavior suitable for the state of the user. The estimating unit comprises: a state observation unit that observes the state of the user based on at least one of the captured image and the biological information; and A behavior determination unit determines the given behavior suitable for the state of the user based on the value of the given behavior.
3. The robot according to claim 1 or 2, wherein: The captured image includes at least one of a facial image and a full-body image of the user, The biological information includes information related to at least one of a heartbeat, a respiration, a blood pressure, and a body temperature of the user.
4. The robot according to claim 1 or 2, wherein: The state of the user includes a state of emotion of the user classified based on at least one of a facial image of the user and information related to at least one of a heartbeat, a respiration, a blood pressure, and a body temperature of the user.
5. The robot according to claim 1 or 2, wherein: The state of the user includes a state of the user's behavior classified based on a motion of a skeleton estimated from a full-body image of the user.
6. The robot according to claim 1 or 2, wherein: The state of the user is a given state classified according to a combination of an emotion and a behavior of the user estimated from at least one of the captured image and the biometric information.
7. The robot according to claim 2, wherein: The estimating unit includes a learning unit that generates a learning model that inputs the state of the user and outputs the value of the behavior of the robot through machine learning.
8. The robot according to claim 7, wherein: The estimating unit further includes a result acquiring unit that acquires information related to a result of the establishment of the user contact as a result of the behavior of the robot. The learning unit updates the learning model based on a result of the establishment of contact with the user.
9. The robot according to claim 7, wherein: The learning model is a behavior value table or a neural network.
10. The robot according to claim 8, wherein: The information related to the result of the contact with the user includes whether the user is close, the emotional level of the user, and the duration of the contact with the user. The learning unit obtains a reward corresponding to the behavior of the robot based on a result of the contact with the user, and updates a value of the behavior corresponding to the state of the user based on the reward.
11. A learning device connected to a robot in a communicative manner, comprising: a state observation unit that observes the state of the user based on at least one of a captured image of the user and biological information of the user; and A learning unit generates a learning model that inputs the state of the user and outputs the value of the behavior of the robot through machine learning.
12. The learning device according to claim 11, further comprising: a behavior determination unit that determines the behavior of the robot adapted to the state of the user based on the value of the behavior; and A communication control unit sends an execution instruction of the behavior to the robot.
13. The learning device according to claim 11 or 12, wherein: The learning device further includes a result acquisition unit that acquires information related to a result of the user contact as a result of the behavior of the robot. The learning unit updates the learning model based on the establishment result of the contact.
14. A control method is a control method for a robot, wherein the robot performs the following steps: A step of acquiring at least one of an image of a user and biological information of the user; and Based on at least one of the captured image and the biological information, a step of inducing a given behavior of contact with the user is performed according to the state of the user.
15. A program that causes a computer controlling a robot to execute the following steps: A step of acquiring at least one of an image of a user and biological information of the user; and A step of instructing, based on at least one of the captured image and the biological information, inducing execution of a given behavior of contact with the user in accordance with the state of the user.
Citation Information
Patent Citations
Method for manufacturing fuel cell separator, and method for manufacturing fuel cell
JP2009104878A
Analyzer, analysis method, and computer program
JP2022156758A