Robot, learning device, control method, and program
By integrating the acquisition unit and the behavior control unit in the robot system, using machine learning to infer the user's contact mode status, the problem of robots in the prior art is difficult to accurately match the contact modes required by users, and the effect of reducing discomfort and promoting long-term contact is achieved.
Patent Information
- Application Number
- CN202380069507.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, robots adjust their behavior according to the user's contact method, and it is difficult to accurately match the contact method requested by the user, resulting in discomfort and affecting the long-term contact and healing effect.
A robot system is designed, including a acquisition unit and a behavior control unit, by obtaining information related to user contact, generate a learning model based on machine learning, inferring the user's contact mode status, and instructing appropriate behavior to reduce discomfort.
By accurately matching users' contact methods, it reduces discomfort, improves user acceptance, and promotes long-term contact and healing effects.
Smart Images

Figure CN119947803A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robot, a learning device, a control method and a program. Background Art
[0002] Conventionally, there is a known robot that is designed to come into contact with a user. Also, there is a known robot that provides healing by coming into contact with a user.
[0003] Patent Document 1 discloses that the contact of a user with the body surface of a robot is detected, whether the action is comfortable is determined based on the contact position and contact strength, and the behavior of the robot is changed based on the determination result.
[0004] <Prior Art Literature>
[0005] <Patent Documents>
[0006] Patent Document 1: Japanese Patent Application Publication No. 2019-72495 Summary of the invention
[0007] <Problems to be Solved by the Invention>
[0008] However, in Patent Document 1, since the robot's behavior changes depending on whether the user's contact is a comfortable action for the robot, it is difficult to perform a behavior that matches the contact method requested by the user. If the robot performs a behavior that does not match the user's contact method, it will become an uncomfortable communication for the user. It may even prevent the user from long-term contact, and thus fail to provide healing.
[0009] The purpose of the technology of the present invention is to reduce behaviors that are inconsistent with the user's touch method.
[0010] <Methods used to solve the problem>
[0011] One embodiment of the present invention is a robot comprising: an acquisition unit that acquires information related to a user's contact with the robot; and a behavior control unit that, based on the information related to the contact and in accordance with the state of the user's contact method, instructs the execution of a given behavior that induces a good response from the user.
[0012] Another embodiment of the present invention is a learning device that is connected to a robot in a communicative manner, and the learning device comprises: a state observation unit that observes the state of the user's contact mode based on information related to the user's contact with the robot; and a learning unit that generates a learning model through machine learning that inputs the state of the user's contact mode and outputs the value of the robot's behavior.
[0013] <Effects of the Invention>
[0014] According to one aspect of the present invention, it is possible to reduce behaviors that are inconsistent with the user's touch method. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a perspective view of a robot according to one embodiment.
[0016] Figure 2 This is a side view of a robot according to one embodiment.
[0017] Figure 3 This is a cross-sectional view of the robot along the III-III cutting line of one embodiment.
[0018] Figure 4 It is a diagram showing the structure of a life sensor according to one embodiment.
[0019] Figure 5 This is a block diagram showing a hardware configuration of a control unit according to one embodiment.
[0020] Figure 6 This is a block diagram showing a functional structure of a control unit according to one embodiment.
[0021] Figure 7 This is a block diagram showing a hardware configuration of an estimating unit (learning device) according to one embodiment.
[0022] Figure 8 This is a block diagram showing a functional configuration of an estimating unit (learning device) according to one embodiment.
[0023] Fig. 9 This is a schematic diagram of a neuron learning model according to one embodiment.
[0024] Fig.10 This is a schematic diagram of a learning model of a neural network according to one embodiment.
[0025] Fig.11 This is a flowchart showing the processing of the control unit according to one embodiment.
[0026] Fig.12 This is a flowchart showing the processing of the estimating unit (learning device) according to one embodiment.
[0027] Fig.13 It is a block diagram showing a functional configuration of an estimating unit (learning device) according to a modified example.
[0028] Fig.14 This is a flowchart showing the processing of the estimating unit (learning device) according to the modification. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In each of the drawings, the same components are denoted by the same reference numerals, and repeated descriptions are appropriately omitted.
[0030] The embodiments shown below illustrate robots for embodying the technical concept of the present invention, but the present invention is not limited to the embodiments shown below. The sizes, materials, shapes and relative arrangements of the components described below are intended to be illustrative unless otherwise specified, and are not intended to limit the scope of the present invention to these. In addition, the sizes and positional relationships of the components shown in the drawings may be exaggerated to make the description clear.
[0031] <Overall Configuration Example of Robot 100>
[0032] Reference Figures 1 to 3 , a structure of a robot 100 according to one embodiment is described. Figure 1 It is a perspective view illustrating a robot 100 according to an embodiment. Figure 2 is a side view of the robot 100 . Figure 3 It is along Figure 2 A cross-sectional view taken along the III-III cutting line.
[0033] The robot 100 is a robot having an outer casing 10 and can be driven by supplied power. The robot 100 illustrated in the present embodiment is a communication robot of a doll type imitating a bear. The robot 100 is made in a size and weight suitable for a user to hold. Here, the user refers to the user of the robot 100. Representative examples of users include single people living alone, elderly people whose children are already independent, and frail elderly people who are the objects of home medical treatment. In addition, in addition to the user of the robot 100, the user may also include the manager of the robot 100 and other people who only come into contact with the robot 100.
[0034] The outer casing 10 has flexibility. For example, the outer casing 10 includes a soft raw material that feels good when the user of the robot 100 touches the robot 100. As the raw material of the outer casing 10, a raw material including an organic material such as polyurethane foam, rubber, resin, fiber, etc. can be used. The outer casing 10 is preferably composed of an outer casing such as a polyurethane foam material having heat insulation properties, and a soft cloth covering the outer surface of the outer casing.
[0035] As an example, the robot 100 includes a body 1, a head 2, an arm 3, and a leg 4. The head 2 includes a right eye 2a, a left eye 2b, a mouth 2c, a right cheek 2d, and a left cheek 2e. The arm 3 includes a right arm 3a and a left arm 3b, and the leg 4 includes a right leg 4a and a left leg 4b. Here, the body 1 corresponds to the robot body. The head 2, the arm 3, and the leg 4 correspond to the driving bodies connected to the robot body in a relatively displaceable manner.
[0036] In the present embodiment, the arm 3 is configured to be displaceable relative to the trunk 1. For example, when the robot 100 is hugged by the user, the right arm 3a and the left arm 3b are displaced to contact the user's head, body, etc. in a manner of hugging the user. Through this action, the user feels close to the robot 100, so that the contact between the user and the robot 100 can be promoted. In addition, the so-called contact with the user refers to the action (contact action) of the user and the robot 100 touching each other, such as rubbing, patting (touching) and hugging (hugging).
[0037] The body 1, the head 2, the arms 3, and the legs 4 are all covered by the outer casing 10. The outer casing in the body 1 is integrated with the outer casing in the arms 3, and the outer casing in the head 2 and the legs 4 is separated from the outer casing in the body 1 and the arms 3. However, it is not limited to these structures, and for example, only the parts of the robot 100 that are easily contacted by the user may be covered by the outer casing 10. In addition, at least one of the outer casings 10 in each of the body 1, the head 2, the arms 3, and the legs 4 may be separated from the other outer casings. In addition, the parts of the head 2, the arms 3, and the legs 4 that do not displace may not include components such as sensors on their inner sides and may be composed only of the outer casing 10.
[0038] The robot 100 has a camera 11, a tactile sensor 12, a control unit 13, a vital sensor 14, a battery 15, a first electrostatic capacitance sensor 21, and a second electrostatic capacitance sensor 31 inside the outer casing 10. In addition, the robot 100 has a camera 11, a tactile sensor 12, a control unit 13, a vital sensor 14, and a battery 15 inside the outer casing 10 in the body 1. Furthermore, the robot 100 has a first electrostatic capacitance sensor 21 inside the outer casing 10 in the head 2, and has a second electrostatic capacitance sensor 31 inside the outer casing 10 in the arm 3.
[0039] In addition, the robot 100 has a display 24, a speaker 25, and a light 26 inside the exterior member 10 in the head 2. Furthermore, the robot 100 has a display 24 inside the exterior member 10 in the right eye 2a and the left eye 2b. In addition, the robot 100 has a speaker 25 inside the exterior member 10 in the mouth 2c, and a light 26 inside the exterior member 10 in the right cheek 2d and the left cheek 2e.
[0040] In more detail, Figure 3As shown, the robot 100 has a body frame 16 and a body mounting platform 17 inside the exterior member 10 in the body 1. In addition, the robot 100 has a head frame 22 and a head mounting platform 23 inside the exterior member 10 in the head 2. Furthermore, the robot 100 has a right arm frame 32a and a right arm mounting platform 33 inside the exterior member 10 in the right arm 3a, and a left arm frame 32b inside the exterior member 10 in the left arm 3b. In addition, the robot 100 has a right leg frame 42a inside the exterior member 10 in the right leg 4a, and a left leg frame 42b inside the exterior member 10 in the left leg 4b.
[0041] The trunk frame 16, the head frame 22, the right arm frame 32a, the left arm frame 32b, the right leg frame 42a, and the left leg frame 42b are structures formed by combining a plurality of columnar members. The trunk mounting platform 17, the head mounting platform 23, and the right arm mounting platform 33 are plate-like members having a mounting surface. The trunk mounting platform 17 is fixed to the trunk frame 16, the head mounting platform 23 is fixed to the head frame 22, and the right arm mounting platform 33 is fixed to the right arm frame 32a. In addition, the trunk frame 16, the head frame 22, the right arm frame 32a, the left arm frame 32b, the right leg frame 42a, and the left leg frame 42b may also be formed in a box shape including a plurality of plate-like members.
[0042] The right arm frame 32a is connected to the body frame 16 via the right arm connection mechanism 34a, and is driven by the right arm servo motor 35a to be relatively displaced with respect to the body frame 16. The right arm frame 32a is displaced, so that the right arm 3a is relatively displaced with respect to the body 1. The right arm connection mechanism 34a preferably has a speed reducer that increases the output torque of the right arm servo motor 35a, for example.
[0043] In this embodiment, the right arm frame 32a is composed of a multi-joint robot including a plurality of frame members and a plurality of connection mechanisms. For example, the right arm frame 32a includes a right shoulder frame F1a, a right upper arm frame F2a, a right elbow frame F3a, and a right forearm frame F4a. The trunk frame 16, the right shoulder frame F1a, the right upper arm frame F2a, the right elbow frame F3a, and the right forearm frame F4a are connected to each other via connection mechanisms.
[0044] The right arm servo motor 35a is a general term for a plurality of servo motors. For example, the right arm servo motor 35a includes a right shoulder servo motor M1a, a right upper arm servo motor M2a, a right elbow servo motor M3a, and a right forearm servo motor M4a. The right shoulder servo motor M1a rotates the right shoulder frame F1a around a rotation axis that is perpendicular to the trunk frame 16. The right upper arm servo motor M2a rotates the right upper arm frame F2a around a rotation axis that is perpendicular to the rotation axis of the right shoulder frame F1a. The right elbow servo motor M3a rotates the right elbow frame F3a around a rotation axis that is perpendicular to the rotation axis of the right upper arm frame F2a. The right forearm servo motor M4a rotates the right forearm frame F4a around a rotation axis that is perpendicular to the rotation axis of the right elbow frame F3a.
[0045] The left arm frame 32b is connected to the body frame 16 via the left arm connection mechanism 34b, and is driven by the left arm servo motor 35b to be relatively displaced with respect to the body frame 16. The displacement of the left arm frame 32b causes the left arm 3b to be relatively displaced with respect to the body 1. The left arm connection mechanism 34b preferably has a speed reducer that increases the output torque of the left arm servo motor 35b, for example.
[0046] In this embodiment, the left arm frame 32b is composed of a multi-joint robot including a plurality of frame members and a plurality of connection mechanisms. For example, the left arm frame 32b includes a left shoulder frame F1b, a left upper arm frame F2b, a left elbow frame F3b, and a left forearm frame F4b. The trunk frame 16, the left shoulder frame F1b, the left upper arm frame F2b, the left elbow frame F3b, and the left forearm frame F4b are connected to each other via connection mechanisms.
[0047] The left arm servo motor 35b is a general term for a plurality of servo motors. For example, the left arm servo motor 35b includes a left shoulder servo motor M1b, a left upper arm servo motor M2b, a left elbow servo motor M3b, and a left forearm servo motor M4b. The left shoulder servo motor M1b rotates the left shoulder frame F1b around a rotation axis that is perpendicular to the trunk frame 16. The left upper arm servo motor M2b rotates the left upper arm frame F2b around a rotation axis that is perpendicular to the rotation axis of the left shoulder frame F1b. The left elbow servo motor M3b rotates the left elbow frame F3b around a rotation axis that is perpendicular to the rotation axis of the left upper arm frame F2b. The left forearm servo motor M4b rotates the left forearm frame F4b around a rotation axis that is perpendicular to the rotation axis of the left elbow frame F3b.
[0048] Since the arm 3 has a four-axis joint, the robot 100 can achieve more realistic movements. For example, the robot 100 can perform actions that match the user's contact situation by turning the arm 3 to gently "hug" the user who gently hugs (embraces) the robot for a long time. In addition, the robot 100 can perform actions that match the user's busy situation by turning the arm 3 to "rub" the user for a short time when the user caresses the robot for a short time.
[0049] The head frame 22 is connected to the body frame 16 via the head connection mechanism 27, and can be relatively displaced with respect to the body frame 16 by being driven by the head servo motor 35c. The head 2 is relatively displaced with respect to the body 1 by the displacement of the head frame 22. The head connection mechanism 27 preferably has a speed reducer that increases the output torque of the head servo motor 35c, for example.
[0050] In this embodiment, the head frame 22 includes a neck frame F1c and a face frame F2c. The body frame 16, the neck frame F1c, and the face frame F2c are connected to each other via a connection mechanism.
[0051] The head servo motor 35c is a general term for a plurality of servo motors. For example, the head servo motor 35c includes a neck servo motor M1c and a face servo motor M2c. The neck servo motor M1c rotates the neck frame F1c around a rotation axis perpendicular to the body frame 16. The face servo motor M2c rotates the face frame F2c around a rotation axis perpendicular to the rotation axis of the neck frame F1c.
[0052] The robot 100 can achieve more realistic movements by having a two-axis joint in the head 2. For example, the robot 100 can perform actions that match the user's touch by turning the head 2 and "looking up" (looking at) a user who hugs the robot for a short time.
[0053] The right leg frame 42a is connected to the trunk frame 16 via the right leg connection mechanism 44a, and has a right leg wheel 41a on the bottom side. In order to stabilize the posture of the robot 100, the robot 100 preferably has two right leg wheels 41a in the front-to-back direction of the right leg frame 42a. The right leg wheel 41a is driven by the right leg servo motor 35d and can rotate around a rotation axis perpendicular to the front-to-back direction of the right leg frame 42a. The robot 100 becomes able to travel by rotating the right leg wheel 41a. The right leg connection mechanism 44a, for example, preferably has a reducer that increases the output torque of the right leg servo motor 35d.
[0054] The left leg frame 42b is connected to the trunk frame 16 via the left leg connection mechanism 44b, and has a left leg wheel 41b on the bottom side. In order to stabilize the posture of the robot 100, the robot 100 preferably has two left leg wheels 41b in the front-to-back direction of the left leg frame 42b. The left leg wheel 41b is driven by the left leg servo motor 35e and can rotate around a rotation axis perpendicular to the front-to-back direction of the left leg frame 42b. The robot 100 becomes able to travel by rotating the left leg wheel 41b. The left leg connection mechanism 44b, for example, preferably has a reducer that increases the output torque of the left leg servo motor 35e.
[0055] In this embodiment, the robot 100 moves forward or backward by turning the right leg wheel 41a and the left leg wheel 41b forward or backward at the same time. The robot 100 turns right or left by braking one of the right leg wheel 41a and the left leg wheel 41b and turning the other forward or backward.
[0056] Thus, the robot 100 can realize more realistic actions through the legs 4. For example, the robot 100 can perform actions that match the user's contact situation by rotating the legs 4 to "lean" on the user who touches the robot for a long time.
[0057] The camera 11 is fixed to the body frame 16. The tactile sensor 12, the control unit 13, the vital sensor 14, and the battery 15 are fixed to the body mounting platform 17. The control unit 13 and the battery 15 are fixed to the side of the body mounting platform 17 opposite to the side to which the tactile sensor 12 and the vital sensor 14 are fixed. In addition, the arrangement of the control unit 13 and the battery 15 here is determined according to the specific conditions of the space that can be arranged on the body mounting platform 17, and is not limited to the above arrangement. However, when the battery 15 is fixed to the side of the body mounting platform 17 opposite to the side to which the tactile sensor 12 and the vital sensor 14 are fixed, the center of gravity of the robot 100 is lowered because the battery 15 is heavier than other components. When the center of gravity of the robot 100 is lower, at least one of the position and posture of the robot 100 is stabilized, and at least one of charging and replacing the battery 15 becomes easy, so it is preferable.
[0058] The first electrostatic capacitance sensor 21 is fixed to the head mounting platform 23, and the second electrostatic capacitance sensor 31 is fixed to the right arm mounting platform 33. The display 24 includes a right eye display 24a and a left eye display 24b. The right eye display 24a, the left eye display 24b and the speaker 25 are fixed to the head frame 22. The light 26 includes a right cheek light 26a and a left cheek light 26b. The right cheek light 26a and the left cheek light 26b are fixed to the head frame 22.
[0059] In addition, the camera 11, the tactile sensor 12, the control unit 13, the life sensor 14, the battery 15, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, etc. can be fixed by screw members or adhesive members, etc. In addition, the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, the left cheek light 26b, etc. can also be fixed by screw members or adhesive members, etc.
[0060] As the materials of the body frame 16, the body mounting platform 17, the head frame 22, the head mounting platform 23, the right arm frame 32a, the right arm mounting platform 33 and the left arm frame 32b, there is no particular restriction, and resin materials or metal materials can be used. However, from the perspective of ensuring the strength during driving, the body frame 16, the right arm frame 32a and the left arm frame 32b are preferably made of metal materials such as aluminum. On the other hand, only the strength can be ensured. In order to make the robot 100 lightweight, the materials of these parts are preferably made of resin materials. As the materials of the body mounting platform 17, the head frame 22, the head mounting platform 23, the right arm mounting platform 33 and the left arm frame 32b, there is no particular restriction, and resin materials or metal materials can be used, but from the perspective of making the robot 100 lightweight, it is preferably used.
[0061] The control unit 13 is connected to the camera 11, the tactile sensor 12, the vital sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the right arm servo motor 35a, and the left arm servo motor 35b respectively by wire or wireless so as to enable communication. In addition, the control unit 13 is also connected to the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e respectively by wire or wireless so as to enable communication. Furthermore, the control unit 13 is also connected to the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, and the left cheek light 26b respectively by wire or wireless so as to enable communication.
[0062] The camera 11 is an image sensor that outputs a captured image of the robot 100's surroundings to the control unit 13. In the present embodiment, the camera 11 is an example of a capturing unit that captures a user. The camera 11 includes a lens and an image capturing element that captures an image formed by the lens. The image capturing element may use a CCD (Charge Coupled Device) or a CMOS (Complementary Metal-Oxide Semiconductor). The captured image may be either a still image or a dynamic image.
[0063] In addition, the camera 11 is preferably composed of a TOF (Time Of Flight) camera that outputs a distance image of the robot 100's surroundings to the control unit 13. Therefore, the captured image output from the camera 11 sometimes includes a three-dimensional captured image (distance image) in addition to the two-dimensional captured image, or includes a three-dimensional captured image (distance image) instead of a two-dimensional camera image. The captured image is used to detect the presence or approach of a user, detect the distance from the robot 100 to the user, authenticate the user, or estimate the user's emotions. The captured image is an example of a facial image of the user. In addition, the robot 100 may also have, in addition to the camera 11, a human sensor such as an ultrasonic sensor, an infrared sensor, a millimeter wave radar, or a LiDAR (light Detection And Raging).
[0064] The tactile sensor 12 is a sensor element that detects information perceived by the sense of touch possessed by human hands, etc., converts the information into a tactile signal as an electrical signal, and outputs the information to the control unit 13. From the viewpoint of making the touch of the robot 100 good, the tactile sensor 12 is preferably configured so as not to affect the touch of the user. For example, the tactile sensor 12 is preferably configured to be flexible and to be able to deform following the deformation of the outer casing 10 of the robot 100. For example, the tactile sensor 12 converts information on pressure or vibration generated by the user's contact with the robot 100 into a tactile signal through a piezoelectric element and outputs the tactile signal to the control unit 13. The tactile signal output from the tactile sensor 12 is an example of information related to the user's contact with the robot 100.
[0065] The life sensor 14 is an example of an electromagnetic wave sensor that uses electromagnetic waves to obtain biological information of the user. Figure 4 Details will be given separately.
[0066] The first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are sensor elements that output electrostatic capacitance signals to the control unit 13 for detecting the contact or proximity of the user with the robot 100 based on the change in electrostatic capacitance. From the viewpoint of making the touch of the robot 100 good, the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are preferably configured so as not to affect the touch of the user. For example, the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are configured as flexible sensors including conductive threads or the like fixed to the exterior member 10 in a mesh shape. In addition, the electrostatic capacitance signals output from the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are an example of information related to the contact of the user with the robot 100.
[0067] In addition, "information related to the user's contact" includes information related to at least one of the contact position, contact range, contact time, and contact strength detected from information of a sensor that detects the user's contact, such as a tactile signal or an electrostatic capacitance signal. The information related to the contact position, contact range, contact time, and contact strength is defined as follows.
[0068] Information related to the "contact part" is, for example, information about the body part of the robot 100 that is touched by the user in one user contact. For example, when the user touches the head 2 of the robot 100, the "contact part" is the head 2, and when the user hugs the robot 100, the "contact part" is the torso 1 and the arm 3. In addition, the "contact part" sometimes changes over time in one user contact, and in this case, it can be the body part of the robot 100 with the largest number of contact parts. In addition, the so-called "one user contact" refers to the duration of the contact from the time the user touches the robot 100 to the time when a given time (for example, 1 second) has passed. For example, even if the user touches the robot 100 twice in succession, if the first touch and the second touch are not separated by more than a given time (for example, more than 1 second), it is also considered as one contact.
[0069] In addition, the information related to the "contact part" is preferably not a large category such as the trunk 1, head 2, arm 3 or leg 4, but a medium category or a small category thereof. For example, the "contact part" is a medium category or a small category of the trunk 1 such as the chest, abdomen, back and waist, and is a medium category or a small category of the head 2 such as the eyes, nose, mouth, jaw, cheek, forehead, top of the head, side of the head and back of the head. In addition, for example, the "contact part" is a medium category or a small category of the arm 3 such as the shoulder, upper arm, forearm and hand of the right arm 3a or the left arm 3b, and is a medium category or a small category of the leg 4 such as the thigh, knee, calf and foot of the right leg 4a or the left leg 4b. Therefore, the tactile sensor 12, the first electrostatic capacitance sensor 21 and the second electrostatic capacitance sensor 31 are preferably capable of detecting a large category, a medium category or a small category of the body part.
[0070] Information related to the “contact range” is, for example, information on at least one of the surface region and the surface area of the robot 100 touched by the user in one user contact. For example, when the user hugs the robot 100, the “contact range” is a surface area including connection information of the body parts of the robot 100 touched by the user, such as the chest (torso 1) -> the right upper arm and the left upper arm (arm 3) -> the back (torso 1). In addition, when the user hugs the robot 100, the “contact range” is the surface area obtained by adding up the surface areas of the robot 100 touched by the user in the chest (torso 1), the right upper arm (arm 3), the left upper arm (arm 3), and the back (torso 1). In addition, the “contact range” may change over time in one user contact. In this case, it may be the contact range of the robot 100 with the largest contact range in one user contact.
[0071] Information related to "contact time" is, for example, information on the duration of one user contact. "Contact time" does not refer to the duration of the user continuously touching a body part of the robot 100. For example, when the user caresses the arm 3 of the robot 100 and then hugs the robot, it refers to the duration from when the user caresses the robot 100 to when the hug ends. In addition, the information related to "contact time" may also include at least one of the user's contact speed and contact acceleration relative to the robot 100 (hereinafter referred to as the user's "contact change rate"). Furthermore, the information related to "contact time" may also include the frequency of the user's contact methods, such as the number of times the user hugs per day or the number of times the user caresses per day.
[0072] Information related to the "contact strength" is, for example, information on the strength of the force applied by the user to the surface of the robot 100 during one user contact. For example, when the user hugs the robot 100, the "contact strength" is the maximum value or average value of the strength of the force applied by the user to the chest (body part 1), the right upper arm (arm 3), the left upper arm (arm 3), and the back (body part 1). In addition, the "contact strength" sometimes changes over time during one user contact. In this case, it can be the maximum value or average value of the strength of the force applied by the user during one user contact.
[0073] The right eye display 24a and the left eye display 24b are display modules that display character strings or images such as characters, numbers, and symbols according to instructions from the control unit 13. The right eye display 24a and the left eye display 24b are composed of, for example, liquid crystal display modules. The character strings or images displayed on the right eye display 24a and the left eye display 24b are used to express the emotions of the robot 100. For example, the robot 100 displays an image of "changing expression" (opening or closing eyes) on the right eye display 24a and the left eye display 24b for a user who hugs the robot tightly for a short time, thereby enabling the robot 100 to perform actions that match the user's touch method.
[0074] The speaker 25 is a speaker unit that amplifies the sound signal from the control unit 13 and outputs the sound. The sound output from the speaker 25 is the speech or call of the robot 100, which is used to express the emotion of the robot 100. For example, the robot 100 "makes a sound" (greet) from the speaker 25 to the user who touches the robot once, thereby being able to perform a behavior that matches the user's touch method.
[0075] The right cheek light 26a and the left cheek light 26b are light modules that flash or change color according to the on / off signal from the control unit 13. The right cheek light 26a and the left cheek light 26b are composed of, for example, LED (Light Emitting Diode) light modules. The flashing or color change of the right cheek light 26a and the left cheek light 26b is used to express the emotions of the robot 100. For example, the robot 100 can perform actions that match the user's touch method by flashing the right cheek light 26a and the left cheek light 26b in red for a user who hugs the robot.
[0076] The battery 15 is a power source that supplies power to the camera 11, the tactile sensor 12, the control unit 13, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the right arm servo motor 35a, and the left arm servo motor 35b. In addition, the battery 15 also supplies power to the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e. Furthermore, the battery 15 also supplies power to the right eye display 24a, the left eye display 24b, the speaker 25, the right cheek light 26a, and the left cheek light 26b. The battery 15 can use various secondary batteries such as lithium ion batteries and lithium polymer batteries.
[0077] In addition, the installation positions of the tactile sensor 12, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the camera 11, and the life sensor 14 in the robot 100 can be changed appropriately. Furthermore, various sensors such as the camera 11 and the life sensor 14 can also be arranged outside the robot 100 to send necessary information to the robot 100 or an external device via wireless. For example, a learning device composed of a PC (Personal Computer) or a server is an example of an external device.
[0078] The robot 100 does not necessarily need to include the control unit 13 inside the exterior member 10, and the control unit 13 may communicate with each device via wireless from outside the exterior member 10. The battery 15 may supply power to each component from outside the exterior member 10.
[0079] In this embodiment, the structure in which the head 2, the arm 3, and the leg 4 are all displaceable is illustrated, but the present invention is not limited thereto, and at least one of the head 2, the arm 3, and the leg 4 may be displaceable. In addition, the arm 3 is composed of a 4-axis multi-joint robot, but may also be composed of a 6-axis multi-joint robot. Furthermore, the arm 3 is preferably capable of connecting an end effector such as a hand. In addition, the leg 4 is composed of a wheel system, but may also be composed of a track system or a leg system.
[0080] The structure and shape of the robot 100 are not limited to those illustrated in this embodiment, and can be appropriately changed according to the user's preference or the usage form of the robot 100. For example, the robot 100 may not be in the form of a bear but in the form of a mechanical arm of an industrial robot or the like, or in the form of a humanoid puppet. In addition, the robot 100 may be in the form of a mobile device such as a drone or a vehicle having at least one of an arm, a display, a speaker, and a light.
[0081] <Configuration example of the life sensor 14>
[0082] Figure 4 14 is a diagram illustrating a configuration of a life sensor 14. The life sensor 14 is a microwave Doppler sensor having a microwave transmitting unit 141 and a microwave receiving unit 142. Microwaves are an example of electromagnetic waves.
[0083] The life sensor 14 transmits a transmission wave Ms as a microwave from the inside of the outer casing 10 of the robot 100 to the user 200 through the microwave transmitting unit 141 . In addition, the life sensor 14 receives a reflected wave Mr resulting from the transmission wave Ms being reflected by the user 200 through the microwave receiving unit 142 .
[0084] The life sensor 14 detects the minute displacements on the body surface caused by the heart beats of the user 200, etc., in a non-contact manner by using the Doppler effect based on the difference between the frequencies of the transmitted wave Ms and the reflected wave Mr. The life sensor 14 can obtain information such as the heartbeat, respiration, pulse wave, blood pressure, etc., as biological information of the user 200, based on the detected minute displacements, and output the obtained biological information to the control unit 13.
[0085] However, the life sensor 14 is not limited to a microwave Doppler sensor, and may be a life sensor that detects micro-displacements generated on the body surface by using the coupling change between the human body and the antenna, or may be a life sensor that uses electromagnetic waves other than microwaves such as near-infrared light. In addition, the life sensor 14 may also be a millimeter wave radar, a microwave radar, etc. Furthermore, the life sensor 14 preferably has a non-contact thermometer that detects infrared rays emitted from the user 200 in addition to the Doppler sensor. In this case, the life sensor 14 detects biological information of the user 200 including information related to at least one of the heartbeat (pulse), respiration, blood pressure, and body temperature.
[0086] In this embodiment, the life sensor 14 is provided inside the outer casing 10, so the user 200 cannot visually recognize the life sensor 14. Thus, the user 200 can suppress the resistance to the detection of biological information, and can smoothly obtain biological information. In addition, since the life sensor 14 can obtain biological information in a non-contact manner, it is different from a contact sensor that requires the user to contact the same place for a certain period of time, even if the user moves to some extent, the biological information can be obtained.
[0087] Furthermore, by embracing the robot 100, the user 200 and the robot 100 are encouraged to come into contact with each other, so that the robot 100 is embraced by the user 200 and can obtain biological information while in contact or close to the user 200. Thus, the robot 100 can obtain biological information with high reliability while suppressing noise.
[0088] <Configuration Example of Control Unit 13>
[0089] (Hardware Configuration Example)
[0090] Figure 51 is a block diagram showing the hardware structure of the control unit 13. The control unit 13 is constructed by a computer and has a CPU (Central Processing Unit) 131, a ROM (Read Only Memory) 132, and a RAM (Random Access Memory) 133. In addition, the control unit 13 has a HDD / SSD (Hard Disk Drive / Solid State Drive) 134, a device connection I / F (Interface) 135, and a communication I / F 136. They are connected via a system bus A so as to be able to communicate with each other.
[0091] The CPU 131 performs control processing including various calculation processing. The ROM 132 stores programs such as IPL (Initial Program Loader) for driving the CPU 131. The RAM 133 is used as a work area for the CPU 131. The HDD / SSD 134 stores various information such as programs, and detection information obtained by various sensors such as the captured images obtained by the camera 11 and the biological information obtained by the life sensor 14.
[0092] The device connection I / F 135 is an interface for connecting the control unit 13 to various external devices. The external devices here include the camera 11, the tactile sensor 12, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, the servo motor 35, and the battery 15. In addition, the external devices also include the display 24, the speaker 25, and the light 26.
[0093] Here, the servo motor 35 is a generic term for the right arm servo motor 35a, the left arm servo motor 35b, the head servo motor 35c, the right leg servo motor 35d, and the left leg servo motor 35e. In addition, the display 24 is a generic term for the right eye display 24a and the left eye display 24b. Furthermore, the light 26 is a generic term for the right cheek light 26a and the left cheek light 26b.
[0094] The communication I / F 136 is an interface for communicating with an external device via a communication network, etc. For example, the control unit 13 is connected to the Internet via the communication I / F 136 and communicates with an external device via the Internet.
[0095] In addition, at least a part of the functions implemented by the CPU 131 may be implemented by an electric circuit or an electronic circuit.
[0096] (Functional configuration example)
[0097] Figure 6 1 is a block diagram showing the functional structure of the control unit 13. The control unit 13 includes an acquisition unit 101, a communication control unit 102, a storage unit 103, an authentication unit 104, a registration unit 105, a start control unit 106, a motor control unit 107, and an output unit 108. Furthermore, the control unit 13 includes a detection unit 110, an estimation unit 111, and a behavior control unit 112.
[0098] The control unit 13 can realize the functions of the acquisition unit 101 and the output unit 108 through the device connection I / F 135 and the like, and can realize the function of the communication control unit 102 through the communication I / F 136 and the like. In addition, the control unit 13 can realize the functions of the storage unit 103 and the registration unit 105 through the non-volatile memory such as the HDD / SSD 134. Furthermore, the functions of the authentication unit 104, the start control unit 106, and the motor control unit 107 can be realized by the processor such as the CPU 131 executing the processing specified by the program stored in the non-volatile memory such as the ROM 132.
[0099] In addition, the functions of the detection unit 110, the estimation unit 111, and the behavior control unit 112 can be realized by the CPU 131 or other processors executing the processing specified by the program stored in the non-volatile memory such as the ROM 132. In addition, part of the above functions of the control unit 13 can also be realized by an external device such as a PC or a server, or can be realized by distributed processing between the control unit 13 and the external device. For example, the estimation unit 111 can also be configured as a learning device connected to the robot 100 in a communicable manner.
[0100] The acquisition unit 101 acquires the captured image Im of the user 200 from the camera 11 by controlling the communication between the control unit 13 and the camera 11. In addition, the acquisition unit 101 acquires the tactile signal S from the tactile sensor 12 by controlling the communication between the control unit 13 and the tactile sensor 12. Furthermore, the acquisition unit 101 acquires the biological information B of the user 200 from the life sensor 14 by controlling the communication between the control unit 13 and the life sensor 14.
[0101] The acquisition unit 101 acquires the first capacitance signal C1 from the first capacitance sensor 21 by controlling the communication between the control unit 13 and the first capacitance sensor 21. The acquisition unit 101 acquires the second capacitance signal C2 from the second capacitance sensor 31 by controlling the communication between the control unit 13 and the second capacitance sensor 31.
[0102] The communication control unit 102 controls communication with an external device via a communication network, etc. For example, the communication control unit 102 can send a captured image Im obtained by the camera 11, biological information B obtained by the life sensor 14, a tactile signal S obtained by the tactile sensor 12, etc. to an external device (e.g., a learning device described below) via a communication network.
[0103] The storage unit 103 stores the biological information B acquired by the life sensor 14. The storage unit 103 continuously stores the acquired biological information B while the acquisition unit 101 acquires the biological information B from the life sensor 14. In addition, the storage unit 103 can also store information obtained from the captured image Im acquired by the camera 11, the tactile signal S from the tactile sensor 12, the first electrostatic capacitance signal C1 from the first electrostatic capacitance sensor 21, and the second electrostatic capacitance signal C2 from the second electrostatic capacitance sensor 31.
[0104] The authentication unit 104 performs personal authentication on the user 200 based on the captured image Im of the user 200 acquired by the camera 11. For example, the authentication unit 104 performs facial authentication based on the captured image Im including the face of the user 200 captured by the camera 11, referring to the registration information 109 of the facial image pre-registered in the registration unit 105. In this way, the user 200 who is currently in contact with or close to the robot 100 can be associated with the pre-registered personal information, and the biometric information B acquired by the life sensor 14 can be associated with the personal information. In addition, the control unit 13 may control to stop the acquisition of the biometric information by the life sensor 14 when the facial image included in the captured image Im is not registered in the registration unit 105.
[0105] The start control unit 106 causes the life sensor 14 to start acquiring the biological information B. For example, when the detection unit 110 detects that the user 200 is in contact with or close to the robot 100, the start control unit 106 turns on a switch that supplies power from the battery 15 to the life sensor 14. Thus, the start control unit 106 causes the life sensor 14 to start acquiring the biological information B.
[0106] The detection unit 110 detects the presence or approach of the user 200 around the robot 100 based on the captured image Im obtained by the camera 11 or the like. The detection unit 110 preferably detects the distance from the robot 100 to the user 200 based on the captured image Im (distance image) obtained by the camera 11. In addition, the detection unit 110 detects information related to the contact of the user 200 with respect to the robot 100 based on the tactile signal S from the tactile sensor 12. Furthermore, the detection unit 110 may also detect information related to the contact of the user 200 with respect to the robot 100 based on the first electrostatic capacitance signal C1 or the second electrostatic capacitance signal C2.
[0107] The estimation unit 111 estimates a given behavior an (n is the identification number of the behavior a) of the robot 100 that matches the state st (t is the time) of the contact mode of the user 200 based on the information related to the contact mode of the user 200 with respect to the robot 100. In the present embodiment, the estimation unit 111 performs reinforcement learning to estimate a given behavior at (t is the time) of the robot 100 that matches the state st of the contact mode of the user 200. However, the estimation unit 111 may also perform other machine learning such as supervised learning, semi-supervised learning, or unsupervised learning to estimate a given behavior at of the robot 100 that matches the state st of the contact mode of the user 200.
[0108] Regarding the configuration of the estimation unit 111 that performs reinforcement learning, refer to Figure 8 In addition, regarding the structure of the estimation unit 111 for supervised learning, please refer to Fig.13 Further, when learning converges, the inference unit 111 may use the learned learning model (in this embodiment, the learned behavior value table or the learned neural network) to infer the behavior at of the robot 100 that matches the state st of the contact mode of the user 200. In this case, the inference unit 111 infers the given behavior at of the robot 100 that matches the state st of the contact mode of the user 200 through a given logic or a given algorithm.
[0109] The behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to perform the behavior at of the robot 100. In addition, the behavior control unit 112 instructs the execution of a given behavior at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200. By the robot 100 performing the given behavior at that matches the contact mode of the user 200, it is possible to reduce behaviors that make the user 200 feel uncomfortable. Furthermore, it is possible to make the user 200 accept the long-term contact with the robot 100, thereby providing healing for the user 200.
[0110] In addition, the given behavior at of the robot 100 is preferably a behavior that imitates the contact method of the user 200, for example, as a reaction to being hugged for a long time by the user 200, the robot 100 “hugs” the user for a long time, etc. In addition, the given behavior at of the robot 100 is preferably a behavior that imitates at least one of the duration of contact performed by the user 200, the contact change rate (contact speed or contact acceleration) of the user 200, and the contact strength of the user 200.
[0111] In addition, the given behavior at of the robot 100 is preferably a behavior that matches the busy state of the user 200, such as "changing expression" (making cheeks flicker or smiling) for a short time as a reaction to being touched by the user 200 for a short time. Furthermore, the given behavior at of the robot 100 is preferably a behavior that matches the way the user 200 touches, such as "making a sound" (making a greeting sound) for a short time as a reaction to being gently touched on the head by the user 200 for a short time. In addition, the given behavior at of the robot 100 is preferably a behavior that matches the playful way or the playful way of the user 200, such as "fearing" or "changing expression" (shedding tears, opening eyes or closing eyes) as a reaction to being slapped by the user 200.
[0112] The storage unit 103 stores information related to the behavior an of the robot 100 that has been defined in advance. The information related to the behavior an of the robot 100 is managed, for example, by a table of a database. Table 1A and Table 1B below are examples of a behavior table TB1 related to the behavior an of the robot 100. The behavior table TB1 contains a behavior ID (Identity Document, identification number) for identifying the behavior an of the robot 100, the behavior content of the robot 100, the instruction content of the behavior an, the behavior time per cycle, and a usage example.
[0113]
Table 1A
[0114]
[0115]
Table 1B
[0116]
[0117] The symbols in the instruction contents shown in Table 1A and Table 1B represent the symbols of the control objects. In addition, the so-called teaching instruction refers to an action instruction that is taught in advance using a teaching method such as offline teaching, online teaching, or direct teaching. In addition, the so-called tracking instruction refers to an action instruction that tracks the position and posture of the user 200 based on various sensor information such as the captured image Im (range image).
[0118] The motor control unit 107 controls the driving of the servo motor 35 according to the instruction content of the behavior at of the robot 100 from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "hug", the motor control unit 107 controls the position and posture of the robot 100 relative to the position and posture of the user 200 by the tracking instruction, and then executes the action instruction of "hug" that has been taught in advance.
[0119] The output unit 108 controls the communication between the control unit 13 and the display 24 according to the execution instruction of the behavior at of the robot 100 from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "smiling", the output unit 108 outputs a display instruction of a smiling face image to the right eye display 24a and the left eye display 24b.
[0120] In addition, the output unit 108 controls the communication between the control unit 13 and the speaker 25 according to the execution instruction of the behavior at of the robot 100 from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "making a (greeting) sound", the output unit 108 outputs a greeting sound output signal to the speaker 25.
[0121] Furthermore, the output unit 108 controls the communication between the control unit 13 and the lights 26 according to the execution instruction of the behavior at of the robot 100 from the behavior control unit 112. When the behavior content of the robot 100 is, for example, "blinking cheeks", the output unit 108 outputs an on-off signal to the switch elements of the right cheek light 26a and the left cheek light 26b.
[0122] <Configuration of the Estimation Unit 111>
[0123] (Hardware Configuration Example)
[0124] Figure 7 2 is a block diagram showing the hardware configuration of the estimation unit 111 . Figure 7 Shows the Figure 6 The estimation unit 111 shown in the figure is configured as an example of a learning device 300 connected to the robot 100 in a communicable manner. However, the function of the estimation unit 111 may also be as follows Figure 6 As shown, it is arranged inside the robot 100.
[0125] The estimation unit 111 is constructed by a computer and has a CPU 301, a ROM 302, a RAM 303, a HDD / SSD 304, a device connection I / F 305, and a communication I / F 306. These are connected via a system bus A' so as to be able to communicate with each other. In addition, in order to improve the learning processing capability of the computer, the learning device 300 is preferably composed of a PC cluster having a GPU (Graphics Processing Unit) or a plurality of computers.
[0126] The CPU 301 performs control processing including various calculation processing. The ROM 302 stores programs such as IPL (Initial Program Loader) for driving the CPU 301. The RAM 303 is used as a work area for the CPU 301. The HDD / SSD 304 stores various information such as programs, as well as the camera image Im obtained by the camera 11, the biological information B obtained by the life sensor 14, or the detection information obtained by various other sensors.
[0127] The device connection I / F 305 is an interface for connecting the estimation unit 111 to various external devices. The external devices here are the camera 11, the tactile sensor 12, the life sensor 14, the first electrostatic capacitance sensor 21, the second electrostatic capacitance sensor 31, etc. However, the estimation unit 111 can also obtain detection information of these various sensors from the robot 100 via the communication I / F 306 described below.
[0128] The communication I / F 306 is an interface for communicating with an external device such as the robot 100 via a communication network, etc. For example, the estimation unit 111 is connected to the Internet via the communication I / F 136 and communicates with the external device via the Internet. In addition, the estimation unit 111 directly communicates wirelessly with the external device via the communication I / F 306.
[0129] In addition, at least a part of the functions implemented by the CPU 301 may be implemented by an electric circuit or an electronic circuit.
[0130] (Functional configuration example)
[0131] Figure 81 is a block diagram showing the functional structure of the estimation unit 111. The estimation unit 111 includes a state observation unit 121, a behavior determination unit 122, a result acquisition unit 123, a learning unit 124, a communication control unit 125, and a storage unit 126. In addition, when the function of the estimation unit 111 is set inside the robot 100, the communication control unit 125 and the storage unit 126 become unnecessary. In addition, when the estimation unit 111 estimates the behavior at of the robot 100 adapted to the state st of the contact mode of the user 200 using a learned learning model LM or based on a given algorithm, the result acquisition unit 123 and the learning unit 124 become unnecessary.
[0132] The various functions of the state observation unit 121, the behavior determination unit 122, the result acquisition unit 123, and the learning unit 124 can be realized by a processor such as the CPU 301 executing processing specified by a program stored in a non-volatile memory such as the ROM 302. In addition, the function of the communication control unit 125 can be realized by the communication I / F 306, etc. Furthermore, the function of the storage unit 126 can be realized by a non-volatile memory such as the HDD / SDD 304.
[0133] The estimation unit 111 of this embodiment performs reinforcement learning to estimate the behavior at of the robot 100 that matches the state st of the contact mode of the user 200. As a reinforcement learning algorithm, any one of Q learning, Sarsa, Monte Carlo method, and deep reinforcement learning (reinforcement learning using DQN (Deep-Q-Network)) can be used. The following is an example of Q learning and deep reinforcement learning for explanation.
[0134] The state observation unit 121 performs various processes for observing the state st of the user 200. The state observation unit 121 observes the state st of the contact mode of the user 200 based on the information related to the contact of the user 200 with respect to the robot 100. The information related to the contact of the user 200 in this example is the tactile signal S, but it may also be the first electrostatic capacitance signal S1 or the second electrostatic capacitance signal S2.
[0135] The state observation unit 121 includes a contact mode estimation unit 151. The contact mode estimation unit 151 performs a process in the state observation unit 121. The function of the contact mode estimation unit 151 may also be performed by other external devices connected to the learning device 300 in a communicable manner. The contact mode estimation unit 151 estimates the state st of the contact mode of the user 200 based on the information related to the contact with the user 200. The contact mode of the user 200 is classified into a given state based on the information related to at least one of the contact part, contact range, contact time and contact strength of the robot 100. For example, the contact mode of the user 200 is classified into a given state sn including any one of "hugging", "patting", "stroking", "slapping", "pushing away", "shaking hands", "rubbing the face", "stroking" and "holding hands". In addition, the contact mode of the user 200 may be substantially the same as the behavior content of the robot 100 shown in Table 1A and Table 1B.
[0136] The contact mode estimation unit 151 uses a learned learning model or performs machine learning while estimating the contact mode. For example, the contact mode estimation unit 151 uses training data to perform deep learning on a learning model of a neural network that inputs information related to contact with the user 200 (information related to the contact position, contact range, and contact time) and outputs the contact mode of the user 200. Therefore, if information related to new contact when being hugged by the user 200 is input into the learning model of the neural network and AI (artificial intelligence) analysis is performed, the contact mode of the user 200 can be classified as "hugging".
[0137] The storage unit 126 stores information related to the state sn (n is the identification number of the state s) of the contact mode of the user 200. The information related to the state sn of the contact mode is managed, for example, by a table of a database. Table 2 below is an example of a state table TB2 related to the state sn of the contact mode of the user 200. The state table TB2 contains a state ID for identifying the state sn of the contact mode of the user 200, and the state content of the contact mode of the user 200. The state sn of the contact mode of the user 200 only has the number of types of contact modes of the user 200 that have been defined in advance. In addition, in the case where the function of the inference unit 111 is set inside the robot 100, the state table TB2 is saved by the storage unit 103 of the robot 100.
[0138]
Table 2
[0139] Status ID Status content s0 Contact method: hug, contact time: 1s~3s, contact strength: 1N~3N s1 Contact method: hug, contact time: 3s~5s, contact strength: 1N~3N s2 Contact method: hug, contact time: more than 5s, contact strength: 1N~3N s3 Contact method: tapping, contact time: 1s~3s, contact strength: 1N~3N s4 Contact method: tapping, contact time: 3s~5s, contact strength: 1N~3N s5 Contact method: tapping, contact time: more than 5s, contact strength: 1N~3N s6 Contact mode: rubbing, contact time: 1s~3s, contact strength: 1N~3N s ···
[0140] In addition, the state content of the contact mode of the user 200 includes the contact mode, contact time and contact strength of the user 200, but it may also include the contact change rate of the user 200 relative to the robot 100 (contact speed and contact acceleration, etc.), and the frequency of the user's contact mode (the number of times the user hugs per day, etc.).
[0141] When the state st of the contact mode of the user 200 is, for example, "contact mode: hug, contact time: 1s to 3s, contact strength: 1N to 3N", since it is a relatively short hug, it can be observed that the user 200 expects a relatively short contact state. Therefore, if the robot 100 performs the behavior at of "hugging" for a relatively long time, it will become a behavior that makes the user 200 uncomfortable. If the user 200 feels that he has become the passive party, he will eventually get tired of the robot 100 and stop interacting with it. On the other hand, if the robot 100 performs the behavior at of "hugging" for a relatively short time, the possibility of inducing a good response from the user 200 increases.
[0142] Furthermore, when the state st of the contact mode of the user 200 is "contact mode: hug, contact time: 3s to 5s, contact strength: 3N to 5N", since it is a relatively strong hug, it can be observed that the user 200 expects a relatively strong contact state. Therefore, if the robot 100 performs a relatively gentle "hug" behavior at, for example, it may not match the contact scenario desired by the user 200. If it does not match the contact scenario of the user 200, the user 200 will feel that they are not compatible with the robot 100, and will feel annoyed and stay away. On the other hand, if the robot 100 performs a behavior at that matches the contact mode of the user 200, such as a relatively strong "hug" or "break free", the possibility of inducing a good response from the user 200 increases. In this way, it can be inferred that there is a certain correlation between the state st of the contact mode of the user 200 and the value Q of the behavior at of the robot 100.
[0143] The behavior determination unit 122 determines the behavior at of the robot 100 relative to the state st of the contact mode of the user 200 based on the value Q of the behavior at-1 (t-1 is the previous time). The storage unit 126 stores the behavior value table TB3, which indicates the value Q of the robot's behavior an relative to the state sn of the contact mode of the user 200. The following Table 3 is an example of the behavior value table TB3 at a certain time t.
[0144]
Table 3
[0145]
[0146] In the initial state of the behavior value table TB3 (time t=0, etc.), the value Q of the robot 100's behavior at relative to the state st of the contact mode of the user 200 is unknown. Therefore, the behavior determination unit 122 preferably initializes the value Q of all behaviors an with a random number and selects one behavior at from the given behaviors an.
[0147] In addition, if the behavior decision unit 122 only continuously selects the behavior at with the highest value Q for learning, it will not change to the state st+1 (t+1 is the next time) of the contact mode that has not been experienced yet. Therefore, the behavior decision unit 122 preferably uses the ε-greedy method, etc., to select the behavior at with the highest value Q with probability 1-ε, and selects one behavior at from all behaviors an with probability ε.
[0148] For example, when the state st of the contact mode of the user 200 is "contact mode: hug, contact time: more than 5s, contact strength: 1N to 3N", the behavior decision unit 122 selects the behavior a0 of "hug" with the highest value Q with a probability of 0.9 (ε=0.1). As a result, the possibility of the user 200 giving a good response increases. In addition, the behavior decision unit 122 selects an arbitrary behavior at from all behaviors an with a probability of 0.1. As a result, the user 200 will feel that the robot 100 has selected the behavior at of its own will, and will not be annoyed by the robot 100. In addition, when there is no highest value Q but there are multiple identical values Q, the behavior decision unit 122 selects any one of the behaviors at from the highest parallel values Q with a random number.
[0149] The communication control unit 125 sends the execution instruction of the behavior at determined by the behavior determination unit 122 to the robot 100. The robot 100 receives the execution instruction of the behavior at through the communication control unit 102. Then, the behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to execute the behavior at of the robot 100. Thus, the robot 100 executes the behavior at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200.
[0150] The result acquisition unit 123 acquires information related to the good or bad result of the user 200's response as the result of the behavior at of the robot 100. The information related to the good or bad result of the user 200's response preferably includes information related to the emotion of the user 200. The information related to the emotion of the user 200 includes at least the positive or negative emotion level of the user 200.
[0151] The result acquisition unit 123 includes an emotion level estimation unit 152. The function of the emotion level estimation unit 152 may also be performed by another external device connected to the learning device 300 in a communicable manner. The emotion level estimation unit 152 estimates the emotion level of the user 200 based on the captured image Im (facial image) of the user 200 and information related to at least one of the heartbeat, respiration, blood pressure, and body temperature of the user 200.
[0152] For example, the emotion of the user 200 is classified into a given state based on the facial image of the user 200 and at least one of the heart rate [bpm], breathing rate [times / minute], blood pressure value [mmHg], and body temperature [°C] of the user 200. For example, the emotion of the user 200 is preferably classified into at least one of "neutral", "happy", "sad", "disgusted", "fearful", "surprised", and "angry". The emotion of the user 200 may also be a combination of "fearful" and "surprised", for example.
[0153] The emotion level estimation unit 152 estimates emotions using a learned learning model or while performing machine learning. For example, the emotion level estimation unit 152 uses training data to perform deep learning on a learning model of a neural network that inputs a facial image, heart rate, and blood pressure value of the user 200 and outputs emotions of the user 200 including "neutral", "happy", "sad", and "disgusted". Thus, when a new facial image, heart rate, and blood pressure value of the user 200 with a smiling face is input to the learning model of the neural network and AI (artificial intelligence) analysis is performed, the emotion of the user 200 can be classified as "happy".
[0154] Then, the emotion level estimation unit 152 estimates the emotion level to be “neutral” when the emotion of the user 200 is classified as “neutral”, and estimates the emotion level to be “very positive” when the emotion of the user 200 is classified as “happy”. Furthermore, the emotion level estimation unit 152 estimates the emotion level to be “negative” when the emotion of the user 200 is classified as “sad”, and estimates the emotion level to be “very negative” when the emotion of the user 200 is classified as “disgusted”.
[0155] Through the above, the result acquisition unit 123 acquires the good or bad result of the user 200's reaction (for example, the emotional level of the user 200) as the result of the behavior at of the robot 100. In particular, when the emotional level estimation unit 152 estimates that the emotional level of the user 200 is "positive", the result acquisition unit 123 acquires a result indicating "a good reaction of the user 200".
[0156] The learning unit 124 generates a learning model LM that inputs the state st of the contact mode of the user 200 and outputs the value Q(st,at) of the behavior at of the robot 100 through reinforcement learning. In addition, the learning unit 124 updates the learning model LM based on the good or bad result of the reaction of the user 200. In this embodiment, the learning unit 124 obtains the reward r corresponding to the behavior at of the robot 100 based on the good or bad result of the reaction of the user 200, and updates the value Q (behavior value table TB3) of the behavior at corresponding to the state st of the contact mode of the user 200 based on the reward r.
[0157] The learning unit 124 includes a reward acquisition unit 155 and a value update unit 156. The reward acquisition unit 155 acquires a reward r corresponding to the behavior at of the robot 100 based on the good or bad result of the reaction of the user 200 (for example, the emotional level of the user 200).
[0158] The storage unit 126 stores information related to a given report rn (n is an identification number of the report r) based on the good or bad result of the response of the user 200. The information related to the report rn is managed, for example, by a table of a database. Table 4 below is an example of a report table TB4 showing a given report rn. The report table TB4 has a report ID for identifying the report rn, the good or bad result of the response of the user 200, and the report rn based on the good or bad result of the response of the user 200. The number of reports rn corresponds to the number of good or bad results of the response of the user 200 defined in advance.
[0159]
Table 4
[0160] Report ID Good or bad results of user response Rewards r0 Emotional level: Neutral 0 r1 Sentiment Level: Slightly Positive +1 r2 Mood Level: Positive +2 r3 Mood Level: Very Positive +3 r4 Mood Level: Slightly Negative -1 r5 Emotional level: Negative -2 r6 Mood Level: Very Negative -3 rn ··· ···
[0161] The reward acquisition unit 155 acquires a given reward rn corresponding to the emotion level of the user 200. With respect to the reward rn, it is preferred that the more "positive" the emotion level of the user 200 is, the more positive the reward rn is defined. In addition, with respect to the reward rn, it is preferred that the more "negative" the emotion level of the user 200 is, the more negative the reward rn is defined. On the other hand, when the emotion level of the user 200 is "neutral", the reward acquisition unit 155 acquires a reward rn of zero.
[0162] The value updating unit 156 updates the value Q of the behavior at of the robot 100 with respect to the state st of the contact pattern of the user 200 based on the given reward rn. In Q learning, the value Q is updated by the following formula 1.
[0163] [Formula 1]
[0164] Q(st, at)=Q(st, at)+α(r+γmaxQ((st+1, at+1)-Q(st, at))…Equation 1
[0165] In Formula 1, st is the state of the contact mode of the user 200 at a certain time t, and at is the behavior of the robot 100 at a certain time t. Through the behavior at of the robot 100, the state of the contact mode of the user 200 becomes st+1 (t+1 is the next time). r is the reward obtained by the change in the state of the contact mode. In addition, the item with max is the item obtained by multiplying the value Q when the behavior at+1 with the highest value Q known at this time is selected when the state of the contact mode is st+1 by the discount rate γ (0<γ≦1). In addition, α is a learning coefficient (0<γ≦1) used to adjust the learning speed.
[0166] Formula 1 represents the following method: based on the reward r returned as a result of the behavior at of the robot 100, the value Q(st,at) of the behavior at relative to the state st of the contact mode of the user 200 is updated. When the value Q of a certain behavior at of the robot 100 relative to the state st of a certain contact mode of the user 200 is smaller than the sum of its reward r and the discounted value Q of the best behavior at+1 relative to the state st+1 of the next contact mode, the value Q(st,at) is increased. On the contrary, when the value of the behavior at of the robot 100 relative to the state st of the contact mode of the user 200 is larger than the sum of its reward r and the discounted value Q of the best behavior at+1 relative to the state st+1 of the next contact mode, the value Q(st,at) is reduced. Therefore, Formula 1 makes the value Q of a certain behavior at under a certain contact mode state st close to the sum of the reward r as a result and the discounted value Q of the best behavior at+1 relative to the state st+1 of the next contact mode.
[0167] The value updating unit 156 updates the value Q(sn, an) of the behavior value table TB3 by using equation 1. Then, the state observation unit 121 observes the state st+1 of the contact mode of the next user 200. The behavior determination unit 122 determines the behavior at+1 of the robot 100 corresponding to the state st+1 of the contact mode of the next user 200 based on the value Q of the behavior at (in this example, the behavior value table TB3) by using the ε-greedy method or the like.
[0168] The communication control unit 125 sends the execution instruction of the determined behavior at+1 to the robot 100. The robot 100 receives the execution instruction of the behavior at+1 from the learning device 300 via the communication control unit 102. Then, the behavior control unit 112 instructs the motor control unit 107 or the output unit 108 to execute the behavior at+1 of the robot 100. As a result, the robot 100 executes the behavior at+1 that matches the state st+1 of the contact mode of the user 200.
[0169] Here, regarding the method of expressing the value Q(st, at) on a computer, there is the following method: as described above, for all combinations of the contact states sn of the users 200 and the behaviors an of the robots 100, the value Q(st, at) is saved in advance as a behavior value table TB3. In addition, there is the following method: prepare a behavior value function that is approximate to the behavior value table TB3. The latter method can be achieved by adjusting the parameters of the approximate function using a method such as a probabilistic gradient descent method. For example, as an approximate function, the learning unit 124 preferably generates a learning model of a neural network (DQN) that inputs the state st of the contact state of the user 200 and outputs the value Q of the behavior at of the robot 100 through deep reinforcement learning.
[0170] Below, we will explain deep reinforcement learning, but first we will explain neural networks. Fig. 9 is a diagram schematically showing a learning model of a neuron. Fig.10 It is a schematic representation of Fig. 9 A diagram of a learning model of a three-layer neural network composed of a combination of neurons shown in FIG. A neural network is composed of, for example, simulated Fig. 9 The model of the neuron (simple perceptron) shown is composed of a computing device, a memory, etc.
[0171] like Fig. 9 As shown, the neuron output is related to multiple inputs x (in Fig. 9 As an example, the output (result) y is relative to the input x1 to input x3. For each input x (x1, x2, x3), a weight w (w1, w2, w3) corresponding to the input x is assigned. As a result, the neuron outputs the output y expressed by the following formula 2. In addition, the input x, output y and weight w are all vectors. In the following formula 2, θ is the bias and fk is the activation function.
[0172] [Formula 2]
[0173]
[0174] Fig.10 Shows the Fig. 9 The three-layer neural network is composed of the neurons shown in Fig.10 As shown, multiple inputs x are input from the left side of the neural network (here, as an example, inputs x1 to x3), and results y are output from the right side (here, as an example, outputs y1 to y3). Specifically, inputs x1, x2, and x3 are input to the three neurons N11 to N13 with corresponding weights assigned to them. The weights assigned to these inputs are collectively referred to as W1.
[0175] Neurons N11~N13 output z11~z13 respectively. Fig.10In the above example, z11 to z13 are collectively referred to as feature vectors Z1, which can be regarded as vectors that extract the feature value of the input vector. This feature vector Z1 is a feature vector between weights W1 and W2. z11 to z13 are input to two neurons N21 and N22 with corresponding weights assigned to them. The weights assigned to these feature vectors are collectively referred to as W2.
[0176] Neurons N21 and N22 output z21 and z22 respectively. Fig.10 In the above example, z21 and z22 are collectively referred to as feature vector Z2. This feature vector Z2 is a feature vector between weight W2 and weight W3. z21 and z22 are input to three neurons N31 to N33 with corresponding weights assigned to them. The weights assigned to these feature vectors are collectively referred to as W3.
[0177] Finally, neurons N31 to N33 output output y1 to output y3, respectively. In the operation of the neural network, there is a learning mode for learning weights W1 to W3 of the neural network, and an estimation mode for estimating outputs y1 to y3 based on inputs x1 to x3. For example, in the learning mode, the weights W1 to W3 are learned using a learning data set, and the behavior at of the robot 100 is determined using the parameters in the estimation mode. In addition, although it is recorded as "estimation" for the sake of convenience, it goes without saying that various tasks such as detection and classification can be performed.
[0178] In addition, weights W1 to W3 can be learned by backpropagation. Error information enters from the right side of the neural network and flows to the left side. Backpropagation is a method of adjusting (learning) each weight in a way that reduces the difference (error) between the output y when input x is input and the true output y (label data) for each neuron.
[0179] Such a neural network can also be more than three layers, and further layers can be added to perform deep learning. In addition, a convolutional neural network (CNN) that extracts input features in stages and a computing device for a neural network that classifies or regresses outputs can also be automatically obtained based only on training data.
[0180] In the above-mentioned behavior value table TB3, when the number of states sn of the contact mode of the user 200 and the number of behaviors an of the robot 100 are large, the memory space of the behavior value table TB3 may become too large. Therefore, by using a neural network (DQN) to perform function approximation on the behavior value table TB3, the increase of the memory space can be prevented.
[0181] Refer again Figure 8, the structure of the estimation unit 111 for performing deep reinforcement learning is described. The learning unit 124 includes a target network TN (value Q (st, at) | θ - ) and a Q network QN (value Q(st,at)|θ). The two networks are stored in the storage unit 126. The two networks have the same structure, but the parameters θ (equivalent to the above-mentioned weights) are different. The inputs of the two networks are both the state st of the contact mode of the user 200, and the outputs are both the value Q(st,at) of the behavior at of the robot 100.
[0182] The state observation unit 121 observes the state st of the contact mode of the user 200 and outputs it to the behavior determination unit 122 and the learning unit 124. The behavior determination unit 122 inputs the state st of the contact mode of the user 200 into the target network TN, and based on the value Q(st,at|θ - ), the behavior at of the robot 100 is determined by the ε-greedy method, etc. The communication control unit 125 sends an execution instruction of the determined behavior at to the robot 100, and the robot 100 executes the behavior at according to the execution instruction of the behavior at.
[0183] The result acquisition unit 123 acquires the good or bad result of the user 200's reaction (for example, the emotional level of the user 200) as the result of the robot 100's behavior at, and outputs it to the learning unit 124. The learning unit 124 acquires the reward r based on the good or bad result of the user 200's reaction. In addition, the state observation unit 121 observes the state st+1 of the next contact method of the user 200, and outputs it to the behavior determination unit 122 and the learning unit 124.
[0184] The learning unit 124 stores the experience et (<st, at, st+1, r>) of the robot 100 in the storage unit 126 as an experience buffer. Here, st is the state of the contact mode of the user 200, at is the behavior of the robot 100, st+1 is the state of the contact mode of the next user 200, and r is the reward. In addition, the learning unit 124 preferably clips the reward r within the range of -1 to +1 so as not to overreact to abnormal values or the like (so-called reward clipping).
[0185] The learning unit 124 periodically obtains arbitrary experience et from the storage unit 103 (Experience Buffer) and learns the Q network QN. For example, the learning unit 124 obtains experience (B = e0 to en) for small batch learning B from the storage unit 126. Then, the learning unit 124 updates the parameter θ of the Q network QN in a manner that minimizes the TD (Temporal Difference) error L(θ) shown in the following formula 3 (so-called experience replay).
[0186] [Formula 3]
[0187]
[0188] Next, the learning unit 124 reflects the parameters θ of the Q network QN in the target network TN at arbitrary intervals. The learning unit 124 may periodically copy all the parameters θ of the Q network QN to the target network TN, or may reflect the parameters θ of the Q network QN little by little each time the parameters θ of the Q network QN are updated.
[0189] The behavior decision unit 122 inputs the state st+1 of the contact mode of the next user 200 to the target network TN. Then, the behavior decision unit 122 determines the value Q (st+1, at+1|θ) of the behavior at+1 output from the target network TN. - ), the behavior at+1 of the robot 100 is determined by the ε-greedy method, etc. The communication control unit 125 sends an execution instruction of the determined behavior at+1 to the robot 100, and the robot 100 executes the behavior at+1 according to the execution instruction of the behavior at+1.
[0190] Through the above processing, the estimating unit 111 can perform deep reinforcement learning and estimate the behavior at of the robot 100 that is suitable for the state st of the contact method of the user 200.
[0191] <Processing Example of Control Unit 13>
[0192] Fig.11 It is a flowchart which illustrates the processing of the control unit 13. Fig.11 The following process is shown: the control unit 13 instructs the execution of the action at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200.
[0193] First, in step S10, the control unit 13 obtains information related to the contact of the user 200 with respect to the robot 100 through the obtaining unit 101. The information related to the contact of the user 200 is, for example, at least one of the tactile signal S, the first electrostatic capacitance signal C1, and the second electrostatic capacitance signal C2.
[0194] In step S10, power is supplied from the battery 15 to the tactile sensor 12, the first capacitance sensor 21, and the second capacitance sensor 31. However, in order to reduce power consumption of the battery 15, power may not be supplied to the camera 11, the vital sensor 14, the servo motor 35, the display 24, the speaker 25, and the light 26.
[0195] Next, in step S11, the control unit 13 estimates the given behavior at of the robot 100 that is compatible with the state st of the contact mode with the user 200 through the estimation unit 111 based on the information related to the contact with the user 200. The control unit 13 preferably estimates the given behavior at of the robot 100 that is compatible with the state st of the contact mode with the user 200 through the estimation unit 111 while performing reinforcement learning or using a learned learning model. In addition, the processing in step S11 is not necessary processing of the control unit 13, and can be executed in an external device (the learning device described below) connected to the robot 100 in a communicative manner. For detailed processing in step S11, refer to Fig.12 Details will be given separately.
[0196] Then, in step S12, the control unit 13 instructs the execution of the behavior at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200 through the behavior control unit 112. After the robot 100 executes the behavior at, the control unit 13 repeatedly performs the processing of steps S10 to S12, thereby accumulating behaviors at that do not make the user 200 feel uncomfortable. The accumulation of natural behaviors at will give the impression of a good partner, and establishing communication becomes natural. Furthermore, the user 200 can be made to accept long-term contact, and can provide healing for the user 200.
[0197] As described above, the control unit 13 performs processing for performing the behavior at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200. In addition, in order to improve the learning processing capability or suppress the power consumption of the battery 15, when the learning device 300 connected to the robot 100 in a communicable manner assumes the function of the estimation unit 111, the processing of step S11 is performed by the learning device 300.
[0198] In addition, Fig.11 When the process shown is started, the servo motor 35, the display 24, the speaker 25 and the lamp 26 may be in a standby state (sleep state) with a suppressed power supply. That is, the control unit 13 preferably restores various devices from the standby state with a suppressed power supply as needed, thereby suppressing the power consumption of the battery 15.
[0199] <Processing by the Estimation Unit 111>
[0200] Fig.12 It is a flowchart showing the processing of the estimation unit 111 (for example, the learning device 300 ). Fig.12 The following process is shown: the estimating unit 111 performs reinforcement learning to estimate a given behavior at of the robot 100 that is suitable for the state st of the contact pattern of the user 200 . Fig.12 The steps shown are Fig.11 The detailed processing of step S11 is shown.
[0201] First, in step S20, the estimation unit 111 observes the state st of the contact mode of the user 200 through the state observation unit 121 based on the information related to the contact of the user 200 with respect to the robot 100. The estimation unit 111 preferably estimates the contact mode of the user 200 based on the information related to at least one of the contact position, contact range, contact time, and contact strength of the user 200, and observes the state st of the contact mode.
[0202] Next, in step S21, the estimation unit 111 determines the behavior at of the robot 100 that is adapted to the state st of the contact mode of the user 200 based on the value Q of the behavior at-1 (the above-mentioned behavior value table TB3 or the learning model LM such as DQN) through the behavior determination unit 122. The estimation unit 111 outputs the execution instruction of the behavior at to the behavior control unit 112 or sends it to the robot 100 via the communication control unit 125, so that the robot 100 executes the behavior at (step S12).
[0203] In addition, step S20 and step S21 are an estimation phase for estimating the behavior at of the robot 100 that matches the state st of the contact pattern of the user 200, and the other steps are a learning phase.
[0204] In step S22 , the estimation unit 111 obtains information on the good or bad result of the user's 200 response (eg, the emotion level of the user 200 ) as the result of the behavior at of the robot 100 through the result acquisition unit 123 .
[0205] In step S23 , the estimation unit 111 obtains the reward r corresponding to the behavior at of the robot 100 based on the good or bad result of the response of the user 200 through the reward acquisition unit 155 .
[0206] In step S24 , the estimating unit 111 updates the value Q of the behavior at of the robot 100 with respect to the state st of the contact pattern of the user 200 based on the response r through the value updating unit 156 .
[0207] After learning, the process returns to step S20, and the estimation unit 111 observes the state st+1 of the contact mode of the next user 200 through the state observation unit 121. Then, in step S21, the estimation unit 111 determines the behavior at+1 of the robot 100 that is adapted to the state st+1 of the contact mode of the next user 200 based on the value Q of the updated behavior at through the behavior determination unit 122. Then, the robot 100 executes the next behavior at+1 (step S12).
[0208] In addition, after step S24, the following steps may be provided: the estimation unit 111 determines whether the value Q of the behavior at has converged (i.e., whether the learning has converged) through the value updating unit 156. When the estimation unit 111 determines that the learning has converged, the learning phase may not be executed in the subsequent processing. That is, the estimation unit 111 only executes the estimation phase, and estimates the behavior at+n (t+n is the time after n times) of the robot 100 that is adapted to the state st+n of the contact mode of the user 200 using the learned learning model LM (behavior value table TB3 or DQN, etc.).
[0209] <Configuration of Estimation Unit 111 of Modification Example>
[0210] Fig.13 1 is a block diagram showing the functional configuration of the estimation unit 111 according to the modified example. Figure 8 The difference in the functional structure of the estimation unit 111 shown in the figure is that it performs supervised learning to estimate the behavior at of the robot 100 that is suitable for the state st of the contact mode of the user 200. That is, the learning unit 124 has a training data recording unit 157, an error calculation unit 158, and a learning model updating unit 159. Figure 8 The differences in the structure of the estimating unit 111 shown will be described.
[0211] The function of the training data recording unit 157 can be realized by a nonvolatile memory such as HDD / SSD 304. In addition, the functions of the error calculation unit 158 and the learning model update unit 159 can be realized by a processor such as CPU 301 executing processing specified by a program stored in a nonvolatile memory such as ROM 302.
[0212] The learning unit 124 can use a decision tree (regression tree), a neural network, or logistic regression as a learning model LM for supervised learning. Hereinafter, an example of a learning model LM of a neural network that generates a state st of the contact mode of the user 200 and outputs a value Q of the behavior at of the robot 100 by the learning unit 124 through supervised learning is described.
[0213] The training data recording unit 157 stores training data obtained in the past, for example, by other robots 100 or simulation. The training data is data with results (labels), including the state st-n (t-n is the time n times ago) of the contact mode of the user 200, the behavior at-n of the robot 100, and the value Q (equivalent to a label) of the behavior at-n. The estimation unit 111 obtains training data from other robots 100 or other external devices via the communication control unit 125, etc. In addition, the estimation unit 111 can also store the experience experienced by the robot 100 itself as training data.
[0214] The error calculation unit 158 first obtains the training data from the training data recording unit 157, and calculates the error L of the value Q of the behavior at based on the training data. For example, when the reaction of the user 200 is actually good, the error calculation unit 158 regards that there is an error of -log(Q(st,at)) and calculates the error L. In addition, when the reaction of the user 200 is actually not good, the error calculation unit 158 regards that there is an error of -log(1-Q(st,at)) and calculates the error L.
[0215] The learning model updating unit 159 updates the parameters (the weights, etc.) of the learning model LM of the neural network in a manner that minimizes the error L. In updating the learning model LM, the error back propagation method (Backpropagation) can be used. Thus, the learning unit 124 generates a learning model LM that has been learned to a certain level through training data.
[0216] Then, the estimation unit 111 estimates the behavior at that matches the actual state st of the contact pattern of the user 200 using the learning model LM generated by supervised learning. Then, the robot 100 executes the behavior at that induces a good response from the user 200 according to the state st of the contact pattern of the user 200.
[0217] More specifically, the state observation unit 121 observes the state st of the contact mode of the user 200, and the behavior determination unit 122 uses the learning model LM to determine a given behavior at that matches the state st of the contact mode of the user 200. Then, the communication control unit 125 sends an execution instruction of the behavior at to the robot 100, and the robot 100 executes the received execution instruction of the behavior at.
[0218] The result acquisition unit 123 acquires the good or bad result of the user 200's reaction as the result of the robot 100's behavior at. The error calculation unit 158 calculates the error of the value Q of the behavior at based on the good or bad result of the user 200's reaction, and the learning model updating unit 159 further updates the learning model LM of the neural network in a manner to minimize the error L. Then, the behavior determination unit 122 determines the behavior at+1 of the robot 100 that is suitable for the next state st+1 of the contact method of the user 200 using the learning model LM.
[0219] As described above, the estimation unit 111 estimates the behavior at of the robot 100 that is suitable for the state st of the contact mode of the user 200 using the learning model LM learned to a certain level through supervised learning. Thus, for example, even if the robot 100 fails and is replaced with a robot 100 of the same model, the replaced robot 100 can learn from past experience based on the training data of the failed robot 100 and immediately perform the behavior at that is suitable for the state st of the contact mode of the user 200. In addition, the robot 100 can also perform the behavior at that is suitable for the state st of the contact mode of the user 200 at a certain level for the user 200 that the robot 100 contacts for the first time.
[0220] <Processing of the Estimation Unit 111 of Modification Example>
[0221] Fig.14 This is a flowchart showing the processing of the estimating unit 111 (learning device 300 ) according to the modification. Fig.14 The following process is shown: the estimation unit 111 performs supervised learning to estimate the behavior at of the robot 100 that is suitable for the state st of the contact pattern of the user 200. Fig.14 The steps shown are Fig.11 The detailed processing of step S11 is shown.
[0222] First, in step S30 , the estimation unit 111 obtains the training data from the training data recording unit 157 via the error calculation unit 158 , and calculates the error L of the value Q of the behavior at of the robot 100 based on the training data.
[0223] Next, in step S31, the estimation unit 111 updates the parameters (the weights, etc., described above) of the learning model LM of the neural network through the learning model updating unit 159 in a manner that minimizes the error L. Thus, the estimation unit 111 can estimate the behavior at that matches the state st of the actual contact method of the user 200 using the learning model LM learned to a certain level through the training data.
[0224] Then, in step S32, the estimation unit 111 observes the actual state st of the contact mode of the user 200 through the state observation unit 121 based on the information related to the contact of the user 200 with respect to the robot 100. The estimation unit 111 preferably estimates the contact mode of the user 200 based on the information related to at least one of the contact position, contact range, contact time, and contact strength of the user 200, and observes the state st of the contact mode.
[0225] Next, in step S33, the estimation unit 111 determines the behavior at of the robot 100 that matches the state st of the contact mode of the user 200 based on the value Q of the behavior at-1 (learning model LM such as a neural network) through the behavior determination unit 122. The estimation unit 111 outputs the execution instruction of the behavior at to the behavior control unit 112 or sends it to the robot 100 via the communication control unit 125. As a result, the robot 100 executes the behavior at corresponding to the state st of the contact mode of the user 200 (step S12).
[0226] In addition, step S32 and step S33 are an estimation phase for estimating the behavior at of the robot 100 that matches the state of the contact method of the user 200, and the other steps are a learning phase.
[0227] In step S34 , the estimation unit 111 obtains information on the good or bad result of the reaction of the user 200 (the emotional level of the user 200 ) as the result of the behavior at of the robot 100 through the result acquisition unit 123 .
[0228] Next, the process returns to step S30 , and the estimation unit 111 calculates the error L of the value Q of the action at based on the good or bad result of the response of the user 200 through the error calculation unit 158 .
[0229] Then, in step S31 , the estimation unit 111 updates the parameters (weights, etc.) of the learning model LM based on the error L through the learning model updating unit 159 .
[0230] After learning, in step S32, the estimation unit 111 observes the state st+1 of the contact mode of the next user 200 through the state observation unit 121. Then, in step S33, the estimation unit 111 determines the behavior at+1 of the robot 100 adapted to the state st+1 of the contact mode of the next user 200 based on the value Q of the updated behavior at through the behavior determination unit 122. Then, the robot 100 executes the next behavior at+1 (step S12).
[0231] In addition, after step S31, the following step may be provided: the estimation unit 111 determines whether the value Q of the behavior at has converged (i.e., whether the learning has converged) through the learning model updating unit 159. When the estimation unit 111 determines that the learning has converged, the learning phase may not be executed in the subsequent processing. That is, the estimation unit 111 only executes the estimation phase, and estimates the behavior at+n of the robot 100 that is adapted to the state st+n (t+n is the time after n times) of the contact mode of the user 200 using the learning model LM (neural network, etc.) that has been learned.
[0232] <Function and Effect of the Present Embodiment>
[0233] As described above, the robot 100 performs the behavior at that induces a good reaction of the user 200 according to the state st of the contact mode of the user 200 based on the information related to the contact of the user 200 with respect to the robot 100. Therefore, it is possible to reduce the behavior that is inconsistent with the contact mode of the user 200. The accumulation of natural behaviors will give the impression of a good partner, and it becomes natural to establish communication. Furthermore, it is possible to make the user 200 accept long-term contact, and it is possible to provide healing for the user 200.
[0234] In addition, the information related to the contact of the user 200 includes information related to at least one of the contact position, contact range, contact time, and contact strength of the robot 100. Even when the user 200 hugs the robot 100, it is not necessarily in a state where the user can touch the robot 100 leisurely, and the user 200 may be busy. In addition, the user 200 may want to hug the robot 100 strongly. Therefore, by adding not only the information related to the contact position and contact range but also the information related to the contact time and contact strength, the state st of the contact mode of the user 200 can be observed with high accuracy.
[0235] In addition, the sensor that detects information related to the contact with the user 200 is preferably a sensor that does not affect the touch of the user 200, that is, a sensor with good touch, such as the tactile sensor 12, the first electrostatic capacitance sensor 21, and the second electrostatic capacitance sensor 31. For example, it is preferably a flexible structure that follows the contact change of the outer casing 10 of the robot 100. As a result, the user 200 will be willing to touch the robot 100, and the frequency or number of times the user 200 touches the robot 100 can be increased.
[0236] In addition, the state st of the contact mode of the user 200 is a given state sn classified according to a combination of the contact mode estimated based on the information related to the contact with the user 200, the contact time with the contact mode, and at least one of the contact strength. For example, the state st of the contact mode is any given state sn of "hugging", "patting", "stroking", "slapping", "pushing away", "shaking hands", "rubbing the face", "stroking", and "holding hands". As a result, the state st of the contact mode of the user 200 is classified into dozens to hundreds of types, so that the increase in the memory space of the behavior value table TB3 can be suppressed.
[0237] Furthermore, the given behavior an of the robot 100 is a behavior that imitates the contact method of the user 200. In addition, the given behavior an is a behavior that imitates at least one of the duration of contact performed by the user 200, the contact change rate of the user 200, and the contact strength. That is, the robot 100 performs a behavior of the state st that imitates the contact method of the user 200, thereby increasing the possibility of inducing a good response from the user 200.
[0238] In addition, the robot 100 generates a learning model LM that inputs the state st of the contact mode of the user 200 and outputs the value Q of the behavior at of the robot 100 through machine learning (reinforcement learning or supervised learning, etc.). Therefore, the robot 100 can learn the behavior at that matches the contact mode of the user 200, thereby reducing the discomfort of communication. Furthermore, continuous contact with the user 200 can be achieved, and healing can be provided to the user 200.
[0239] Furthermore, the robot 100 also includes a result acquisition unit 123, which acquires information related to the good or bad results of the user 200's response, and the learning unit 124 updates the learning model LM based on the good or bad results of the user 200's response. Therefore, the robot 100 can continuously learn the behavior at that matches the state st of the contact method of the user 200, so that it can continuously perform the behavior that matches the contact method of the user 200, and can reduce the discomfort of communication. In addition, the learning model LM can improve the reliability of the learning ability of the robot 100 by using the behavior value table TB3 or the neural network (DQN) whose effect has been proven.
[0240] In addition, the information related to the good or bad result of the reaction of the user 200 includes the emotional level of the user 200 estimated based on the facial image and biological information of the user 200. Then, the learning unit 124 obtains the reward r corresponding to the behavior at of the robot 100 based on the good or bad result (emotional level) of the reaction of the user 200, and updates the value Q of the behavior at corresponding to the state st of the contact method of the user 200 based on the reward r. Therefore, the robot 100 can learn the behavior at that is suitable for the state st of the contact method of the user 200 according to the emotional level of the user 200.
[0241] In addition, the robot 100 can also use the learning model LM learned to a certain level through supervised learning to infer the behavior at of the robot 100 that is suitable for the state st of the contact mode of the user 200. As a result, for example, even if the robot 100 fails and is replaced with a robot 100 of the same model, the replaced robot 100 can learn past experience based on the training data and immediately perform the behavior at that is suitable for the state st of the contact mode of the user 200. In addition, the robot 100 can also perform the behavior at that is suitable for the state st of the contact mode of the user 200 at a certain level for the user 200 with whom it has first contact.
[0242] The functions of the estimation unit 111 of the robot 100 described above can also be provided in a learning device 300 connected to the robot 100 in a communicable manner so as to be distributed. Thus, the learning processing capability of the computer can be improved. In addition, through the distributed processing of the learning device 300, technical effects such as reduction in power consumption of the battery 15 of the robot 100, reduction in the number of charging times, and reduction in battery weight can be obtained.
[0243] As mentioned above, although the preferred embodiment was described in detail, it is not limited to the said embodiment, Various deformation|transformation and substitution can be added to the said embodiment without departing from the scope described in a claim.
[0244] In addition, the numbers such as ordinal numbers and quantities used in the description of the above-mentioned embodiments are all illustrative for the purpose of specifically describing the technology of the present invention, and the present invention is not limited to the illustrative numbers. In addition, the connection relationship between the constituent elements is the connection relationship illustrative for the purpose of specifically describing the technology of the present invention, and the connection relationship for realizing the functions of the present invention is not limited thereto.
[0245] The robot involved in this embodiment is particularly suitable for the following purposes: promoting the secretion of oxytocin and providing healing (sense of security or self-affirmation) for people living alone, elderly people whose children have become independent, and frail elderly people who are the objects of home medical treatment. However, it is not limited to this purpose and can be used to provide healing for various users.
[0246] The embodiments of the present invention are as follows, for example.
[0247] <1> A robot comprising:
[0248] an acquisition unit that acquires information related to the user's contact with the robot; and
[0249] A behavior control unit instructs execution of a given behavior that induces a good response from the user according to the state of the contact mode of the user based on the information related to the contact.
[0250] <2> The robot according to the above <1> also has an inference unit, which infers the given behavior that is compatible with the state of the contact mode of the user based on information related to the contact, and the inference unit has: a state observation unit, which observes the state of the contact mode of the user based on the information related to the contact; and a behavior determination unit, which determines the given behavior that is compatible with the state of the contact mode based on the value of the behavior.
[0251] <3> The robot according to <1> or <2>, wherein the information related to the contact includes information related to at least one of a contact location, a contact range, a contact time, and a contact strength.
[0252] <4> The robot according to any one of <1> to <3>, further comprising a sensor for detecting information related to the contact, wherein the sensor is a sensor that does not affect a sense of touch of the user.
[0253] <5> A robot according to any one of <1> to <4> above, wherein the state of the contact mode is a given state classified based on a combination of the contact mode inferred based on information related to the contact, and information related to at least one of the contact time and contact strength of the contact mode.
[0254] <6> The robot according to any one of <1> to <5> above, wherein the given behavior is a behavior that imitates the contact manner of the user.
[0255] <7> The robot according to any one of <1> to <6> above, wherein the given behavior is a behavior that imitates at least one of the duration of contact performed by the user, the contact change rate of the user, and the contact strength of the user.
[0256] <8> The robot according to <2> above, wherein the estimating unit includes a learning unit that generates a learning model that inputs the state of the contact method of the user and outputs the value of the behavior of the robot through machine learning.
[0257] <9> According to the robot described in <8> above, the inference unit also has a result acquisition unit, which acquires information related to the good or bad results of the user's reaction as the result of the robot's behavior, and the learning unit updates the learning model based on the good or bad results of the reaction.
[0258] <10> The robot according to <8> or <9> above, wherein the learning model is a behavior value table or a neural network.
[0259] <11> According to the robot described in <9> above, the acquisition unit also acquires the facial image and biological information of the user, and the information related to the good or bad result of the reaction includes the emotional level of the user estimated based on the facial image and the biological information, and the learning unit obtains the reward corresponding to the behavior of the robot based on the good or bad result of the reaction, and based on the reward, updates the value of the behavior relative to the state of the contact method of the user.
[0260] <12> A learning device connected to a robot in a communicative manner, the learning device comprising: a state observation unit, which observes the state of the user's contact mode based on information related to the user's contact with the robot; and a learning unit, which generates a learning model through machine learning that inputs the state of the user's contact mode and outputs the value of the robot's behavior.
[0261] <13> The learning device according to <12> above comprises: a behavior determination unit, which determines the behavior of the robot to be adapted to the state of the contact mode based on the value of the behavior; and a communication control unit, which sends an instruction for executing the behavior to the robot.
[0262] <14> A control method, which is a control method for a robot, wherein the robot executes the following steps: a step of obtaining information related to contact of a user relative to the robot; and a step of executing a given behavior that induces a good response from the user based on the information related to the contact and in accordance with the state of the user's contact method.
[0263] <15> A program that causes a computer controlling a robot to execute the following steps: obtaining information related to a user's contact with the robot; and based on the information related to the contact, instructing the execution of a given behavior that induces a good response from the user according to the state of the user's contact method.
[0264] This application is based on Japanese Patent Application No. 2022-156765 filed with the Japan Patent Office on September 29, 2022, claims priority, and incorporates all the contents of the Japanese patent application.
[0265] Explanation of symbols
[0266] 1: Body
[0267] 2: Head
[0268] 2a: Right eye
[0269] 2b: Left eye
[0270] 2c: Mouth
[0271] 2d: right cheek
[0272] 2e: Left cheek
[0273] 3: Arm
[0274] 3a: Right arm
[0275] 3b: Left arm
[0276] 4: Legs
[0277] 4a: Right leg
[0278] 4b: Left leg
[0279] 10: Exterior components
[0280] 11: Camera
[0281] 12: Tactile sensor
[0282] 13: Control Department
[0283] 14: Life sensor (electromagnetic wave sensor)
[0284] 141: Microwave Transmitter
[0285] 142: Microwave receiving unit
[0286] 15: Battery
[0287] 16: Body frame
[0288] 17: Body loading platform
[0289] 21: First electrostatic capacitance sensor
[0290] 22: Head frame
[0291] 23: Head loading platform
[0292] 24: Display
[0293] 24a: Right eye display
[0294] 24b: Left eye display
[0295] 25: Speaker
[0296] 26: Lights
[0297] 26a: Right cheek light
[0298] 26b: Left cheek light
[0299] 27: Head connection mechanism
[0300] 31: Second electrostatic capacitance sensor
[0301] 32a: Right arm frame
[0302] 32b: Left arm frame
[0303] 33: Right arm support platform
[0304] 34a: Right arm connection mechanism
[0305] 34b: Left arm connection mechanism
[0306] 35: Servo motor
[0307] 35a: Right arm servo motor
[0308] 35b: Left arm servo motor
[0309] 35c: Head servo motor
[0310] 35d: Right leg servo motor
[0311] 35e: Left leg servo motor
[0312] 41a: Right leg wheel
[0313] 41b: Left leg wheel
[0314] 42a: Right leg frame
[0315] 42b: Left leg frame
[0316] 44a: Right leg connection mechanism
[0317] 44b: Left leg connection mechanism
[0318] 100: Robot
[0319] 101: Acquisition
[0320] 102: Communication control unit
[0321] 103: Preservation Department
[0322] 104: Certification Department
[0323] 105: Registration Department
[0324] 106: Start control department
[0325] 107: Motor control unit
[0326] 108: Output unit
[0327] 109: Registration information
[0328] 110: Inspection Department
[0329] 111: Presumption Department
[0330] 112: Behavior Control Department
[0331] 121: Status Observation Department
[0332] 122: Behavior Decision Department
[0333] 123: Result acquisition unit
[0334] 124: Learning Department
[0335] 125: Communication control unit
[0336] 126: Preservation Department
[0337] 131: CPU
[0338] 132: ROM
[0339] 133: RAM
[0340] 134: HDD / SSD
[0341] 135: Device connection I / F
[0342] 136: Communication I / F
[0343] 151: Contact mode estimation unit
[0344] 152: Emotional level estimation
[0345] 155: Reward Acquisition Department
[0346] 156: Value Update Department
[0347] 157: Training Data Recording Department
[0348] 158: Error calculation unit
[0349] 159: Learning Model Update Department
[0350] 200: User
[0351] 300: Learning device
[0352] 301: CPU
[0353] 302: ROM
[0354] 303: RAM
[0355] 304: HDD / SSD
[0356] 305: Device connection I / F
[0357] 306: Communication I / F
[0358] A, A': system bus
[0359] B: Biological information
[0360] C1: First electrostatic capacitance signal
[0361] C2: Second electrostatic capacitance signal
[0362] F1a: Right shoulder frame
[0363] F2a: Right upper arm frame
[0364] F3a: Right elbow frame
[0365] F4a: Right forearm frame
[0366] F1b: Left shoulder frame
[0367] F2b: Left upper arm frame
[0368] F3b: Left elbow frame
[0369] F4b: Left forearm frame
[0370] F1c: Neck frame
[0371] F2c: Face framing
[0372] Im: Take an image
[0373] L: Error
[0374] LM: Learning Model
[0375] Mr: Reflection wave
[0376] Ms: Emission wave
[0377] M1a: Right shoulder servo motor
[0378] M2a: Right upper arm servo motor
[0379] M3a: Right elbow servo motor
[0380] M4a: Right forearm servo motor
[0381] M1b: Left shoulder servo motor
[0382] M2b: Left upper arm servo motor
[0383] M3b: Left elbow servo motor
[0384] M4b: Left forearm servo motor
[0385] M1c: Neck servo motor
[0386] M2c: Face servo motor
[0387] Q: Value
[0388] S: Tactile signal
[0389] s: Status
[0390] a: behavior
[0391] r: return.
Claims
1. A robot comprising: an acquisition unit that acquires information related to the user's contact with the robot; and A behavior control unit instructs execution of a given behavior that induces a good response from the user in accordance with the state of the contact manner of the user based on the information related to the contact.
2. The robot according to claim 1, wherein: The robot further includes an estimating unit for estimating the given behavior that matches the state of the contact mode of the user. The estimating unit comprises: a state observation unit that observes the state of the contact mode of the user based on the information related to the contact; and A behavior determination unit determines the given behavior that is suitable for the state of the contact mode based on the value of the behavior.
3. The robot according to claim 1 or 2, wherein: The information related to the contact includes information related to at least one of a contact location, a contact range, a contact time, and a contact strength.
4. The robot according to claim 1 or 2, wherein: The robot further includes a sensor for detecting information related to the contact, and the sensor is configured not to affect the user's sense of touch.
5. The robot according to claim 1 or 2, wherein: The state of the contact pattern is a given state classified according to a combination of the contact pattern estimated based on the information related to the contact and information related to at least one of a contact time and a contact strength of the contact pattern.
6. The robot according to claim 1 or 2, wherein the given behavior is a behavior that imitates the contact manner of the user.
7. The robot according to claim 1 or 2, wherein: The given behavior is a behavior that imitates at least one of a duration of contact performed by the user, a contact change rate of the user, and a contact intensity of the user.
8. The robot according to claim 2, wherein: The estimating unit includes a learning unit that generates a learning model that inputs the state of the contact mode of the user and outputs the value of the behavior of the robot through machine learning.
9. The robot according to claim 8, wherein: The estimating unit further includes a result acquiring unit that acquires information related to a good or bad result of the user's response as a result of the behavior of the robot. The learning unit updates the learning model based on the good or bad result of the response.
10. The robot according to claim 8, wherein: The learning model is a behavior value table or a neural network.
11. The robot according to claim 9, wherein: The acquisition unit further acquires the user's facial image and biometric information. The information related to the good or bad result of the reaction includes the emotional level of the user estimated based on the facial image and the biological information, The learning department, Based on the good or bad results of the reaction, a reward corresponding to the behavior of the robot is obtained, And based on the reward, the value of the behavior relative to the state of the contact mode of the user is updated.
12. A learning device connected to a robot in a communicative manner, comprising: a state observation unit that observes the state of the user's contact pattern based on information related to the user's contact with the robot; and A learning unit generates a learning model that inputs the state of the contact mode of the user and outputs the value of the behavior of the robot through machine learning.
13. The learning device according to claim 12, further comprising: a behavior determination unit that determines the behavior of the robot adapted to the state of the contact mode based on the value of the behavior; and A communication control unit sends an instruction for executing the behavior to the robot.
14. A control method is a control method for a robot, wherein the robot performs the following steps: The step of obtaining information related to the user's contact with the robot; and Based on the information related to the contact, a step of performing a given action to induce a good response of the user is performed according to the state of the contact mode of the user.
15. A program that causes a computer controlling a robot to execute the following steps: The step of obtaining information related to the user's contact with the robot; and Based on the information related to the contact, the step of instructing the execution of a given action that induces a good response from the user is performed according to the state of the user's contact mode.
Citation Information
Patent Citations
Autonomous travel robot understanding physical contact
JP2019072495A
High place work vehicle
JP2022156765A