Emotional expression estimation system, emotional expression estimation device, and emotional expression estimation method
The emotion expression estimation system trains devices to evoke and adjust specific emotions, enhancing emotional understanding in children with developmental disorders by mirroring and teaching appropriate emotional responses.
Patent Information
- Application Number
- JP2023216714
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing systems lack the ability to train devices to express specific emotions, particularly for individuals with developmental disorders such as autism, which hinders their understanding of emotional cues.
An emotion expression estimation system that includes a sensor to acquire user information, an execution unit to evoke specific emotions through stories, and a comparison unit to adjust the user's apparent emotion to match the desired emotion, with optional change and teaching units to facilitate emotional mirroring.
The system effectively trains devices to express specific emotions, improving the emotional expression and understanding of children with developmental disorders by adjusting tone, facial expressions, and movements to enhance emotional mirroring.
Smart Images

Figure 2025099785000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an emotion expression estimation system, an emotion expression estimation device, and an emotion expression estimation method.
Background Art
[0002] A system for evaluating and providing feedback on a person's developmental state has been proposed (see, for example, Patent Document 1). In the technique described in Patent Document 1, a processor is caused to receive an input regarding an individual related to a behavioral disorder, developmental delay, or neurological disorder. Then, the processor uses a trained classifier module of a computer program trained using data from a plurality of individuals with a behavioral disorder, developmental delay, or neurological disorder to determine that the individual has signs of the presence of a behavioral disorder, developmental delay, or neurological disorder. Then, the processor uses a machine learning model generated by the computer program to determine that the behavioral disorder, developmental delay, or neurological disorder with signs of the presence is improved by digital therapy that promotes social reciprocity. Further, the processor is caused to provide digital therapy that promotes social reciprocity.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, there is a proposal for a program for diagnosing ADHD and an approach thereto, but there is no specific training for a robot or the like to express specific emotions.
[0005] The present invention has been made in view of the above problems, and an object thereof is to provide an emotion expression estimation system, an emotion expression estimation device, and an emotion expression estimation method that can train a device to express a specific emotion.
Means for Solving the Problems
[0006] (1) To achieve the above object, an emotion expression estimation system according to an aspect of the present invention includes a sensor (e.g., an environmental sensor, an acquisition unit, an emotion estimation unit) that acquires information of a user, an execution unit that causes the user to listen to a story that evokes a specific emotion, and a comparison unit that compares the apparent emotion of the user with the specific emotion using the information of the user acquired by the sensor when the execution unit causes the user to listen to the story. The comparison unit is an emotion expression estimation system that outputs an instruction to bring the apparent emotion of the user closer to the specific emotion based on the comparison result.
[0007] (2) In the emotion expression estimation system according to an aspect of the present invention described in (1) above, a change unit that changes the way of speaking so as to prompt the apparent emotion of the user to approach the specific emotion may be further provided.
[0008] (3) In the emotion expression estimation system according to an aspect of the present invention described in (1) or (2) above, an emotion teaching unit that prompts the apparent emotion of the user to approach the specific emotion may be further provided.
[0009] (4) In the emotion expression estimation system according to an aspect of the present invention among any one of (1) to (4) above, the specific emotion may be classified into "excitement", "love", "fear", "anger", "sadness", "joy" in psychology.
[0010] (5) In the emotion expression estimation system according to an aspect of the present invention described in (2) above, the change unit may change at least one of the speed of the way of speaking, the tone of speaking, and the type of the story.
[0011] (6) In the emotion expression estimation system according to one aspect of the present invention described in (3) above, when the emotion teaching unit does not understand the emotion of the story provided by the user, it may talk to attract children so as to have the correct emotion, and when the user understands the emotion of the story provided, it may make a speech to praise that the user can correctly understand the emotion.
[0012] (7) In the emotion expression estimation system according to one aspect of the present invention among (1) to (4) above, the sensor includes a photographing device and a sound collecting unit, and acquires information of the user by estimating the emotion of the user using at least one of an image photographed by the photographing device and an acoustic signal collected by the sound collecting unit.
[0013] (8) To achieve the above object, an emotion expression estimation device according to one aspect of the present invention includes a sensor that acquires information of a user, an execution unit that makes the user listen to a story that evokes a specific emotion, and a comparison unit that compares the apparent emotion of the user with the specific emotion using the information of the user acquired by the sensor when the execution unit makes the user listen to the story. The comparison unit is an emotion expression estimation device that outputs an instruction to bring the apparent emotion of the user closer to the specific emotion based on the comparison result.
[0014] (9) To achieve the above object, an emotion expression estimation method according to one aspect of the present invention includes a sensor that acquires information of a user, an execution unit that makes the user listen to a story that evokes a specific emotion, a comparison unit that compares the apparent emotion of the user with the specific emotion using the information of the user acquired by the sensor when the execution unit makes the user listen to the story, and the comparison unit that outputs an instruction to bring the apparent emotion of the user closer to the specific emotion based on the comparison result.
Effect of the Invention
[0015] According to (1) to (9) above, the device can be trained to express specific emotions. Further, according to (1), (8), and (9) above, the emotion expression estimation device can be trained using the comparison results.
Brief Description of Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0017] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scales of each member are appropriately changed in order to make each member recognizable in size. In all the drawings for explaining the embodiments, those having the same function are denoted by the same reference numerals, and repeated explanations are omitted. In addition, "based on XX" as used in this application means "at least based on XX", and includes cases where it is based on XX in addition to other elements. Also, "based on XX" is not limited to the case of directly using XX, but also includes cases where it is based on something obtained by performing operations or processing on XX. "XX" is any element (for example, any information).
[0018] [Overview] One of the problems of children with autism is that they cannot understand differences in body language, facial expressions, and tone of voice. For this reason, training in the meaning of facial expressions has been pursued by experts for children so that they can recognize emotional cues from others through self-awareness. In this embodiment, a robot capable of communicating with children reads aloud a story (for example, a paper puppet show) that evokes specific emotions in a rich emotional manner. In this embodiment, the act of uttering the lines and narration of the story, etc. is referred to as "reading aloud". Also, "specific emotions" are classified, for example, into "excitement", "love", "fear", "anger", "sadness", and "joy" in psychology.
[0019] When reading aloud, the robot observes the children. The robot tries to notice how the child's emotion mirroring ability correlates with the actual emotion of the story in the process. In this embodiment, through the robot, the child is enabled to maximally reflect the evidence of the emotion of the story. And in this embodiment, the robot is enabled to adjust its expressive power (tone of voice, facial expression, movement, etc.) to maximize the child's understanding of emotions.
[0020] FIG. 1 is a diagram showing a configuration example of an emotion expression estimation system according to this embodiment. The emotion expression estimation system 1 includes, for example, a robot 2 (emotion expression estimation device), an environmental sensor 3, and an image display device 4. In the following embodiments, the robot 2 will be described as an example of the emotion expression estimation device, but it is not limited to this. The emotion expression estimation device may be composed of a robot 2 and an information processing device, etc.
[0021] While the child Hu (user) communicates with the robot 2, the child Hu views, for example, a puppet show provided by the robot 2 on the image display device 4.
[0022] The robot 2 is a robot capable of communicating with the child Hu. The robot 2 includes, for example, a main body, a display unit capable of displaying expressions of eyes and mouth, a boom capable of operating the eye part, an audio output unit for outputting an audio signal, a sound collection unit, a camera, a control unit for controlling each functional unit, and the like. The robot 2 provides, for example, a puppet show. During the reading aloud, the robot 2 acquires the detection information detected by the environment sensor 3. The robot 2 observes the child Hu using the acquired information. The robot 2 estimates how the child Hu's emotion mirroring ability correlates with the actual emotion of the story using the observed result information. Note that the reading aloud by the robot may be performed by the method of Japanese Patent Application No. 2021-130727, for example. The configuration, operation, processing, etc. of the robot 2 will be described later.
[0023] The environment sensor 3 performs, for example, photographing of an area including the face of the child Hu, detection of an audio signal, etc. Note that the environment sensor 3 may detect the line of sight of the child Hu, etc. The environment sensor 3 may include, for example, a stereo camera or an RGBD camera capable of obtaining depth information, and a sound collection unit.
[0024] The image display device 4 displays an image signal in the following cases where an image is displayed for display, such as a puppet show. Note that the image display device 4 may be provided with a speaker for outputting an acoustic signal such as a sound effect or BGM (BackGround Music).
[0025] [Example of the external shape of the robot] Next, an example of the external shape of the robot 2 will be described. Figure 2 is a diagram showing an external appearance example of the robot according to the present embodiment. In FIG. 2, the front view g101 and the side view g102 are diagrams showing an external appearance example of the robot 2 according to the embodiment. The robot 2 includes, for example, three display units 111 (eye display unit 111a, eye display unit 111b, mouth display unit 111c). Also, in the example of FIG. 2, the imaging unit 102a is attached above the eye display unit 111a, and the imaging unit 102b is attached above the eye display unit 111b. The eye display units 111a and 111b correspond to human eyes and present images and image information corresponding to human eyes. The screen sizes of the eye display units 111a and 111b are, for example, 3 inches. For example, the audio output unit 112, which is a speaker, is attached near the mouth display unit 111c that displays an image corresponding to a human mouth on the housing 120. The mouth display unit 111c is composed of, for example, a plurality of LEDs (light-emitting diodes), each LED can be address-specified, and can be individually driven to turn on and off. The sound collection unit 103 is attached to the housing 120.
[0026] Also, the robot 2 includes a boom 121. The boom 121 is movably attached to the housing 120 via a movable part 131. A horizontal bar 122 is rotatably attached to the boom 121 via a movable part 132. The eye display unit 111a is rotatably attached to the horizontal bar 122 via a movable part 133, and the eye display unit 111b is rotatably attached to the horizontal bar 122 via a movable part 134. Note that the external appearance of the robot 2 shown in FIG. 2 is an example and is not limited thereto.
[0027] [Examples of robot emotional expressions and operation examples] Figure 3 is a diagram showing an example of presenting the emotional expression of the robot according to the present embodiment in animation. As shown in FIG. 3, the robot 2 presents an emotional expression by changing the animation displayed on the display unit 111 (eye display unit 111a, eye display unit 111b, mouth display unit 111c). Animation examples of each of the symbols g11 to g17 are "Angry", "Ecstatic", "Disinterested", "Confused", "Blushing", "Sad", and "Sympathetic".
[0028] Note that the emotional expressions and animations shown in FIG. 3 are just examples and are not limited to this. Emotional expressions may exist other than in FIG. 3, and the animations of each emotional expression may be different from those in FIG. 3. Also, when each emotional expression is made, the angle and position of the display unit 111 are changed as in FIG. 3, the angle of the boom 121 is changed, or an audio signal is output together.
[0029] As described with reference to FIGS. 2 and 3, the robot 2 of the present embodiment can richly display images of eyes and a mouth on the display unit 111, can richly output an audio signal from the audio output unit, and can communicate with the child Hu by moving the boom 121 and the display unit 111. Such rich expressions of the robot 2 can provide an excellent learning mechanism for children to understand expressions and emotions both in terms of vocalization and visual animations.
[0030] In the present embodiment, such a robot 2 reads aloud, for example, in accordance with a paper puppet show. The words contain emotional nuances and can be easily curated. Moreover, children generally love stories very much. And in the present embodiment, when reading aloud, the robot 2 observes the expression and speech of the child Hu and determines whether the emotion of the child Hu is the correct emotion corresponding to the story. And when the emotion of the child Hu is not the correct emotion corresponding to the story, the robot 2 changes the speed, tone, etc. of the reading aloud so that the child Hu can maximally reflect the emotional support of the story.
[0031] [Configuration Example of Each Device of the Emotional Expression Estimation System] Next, a configuration example of each device of the emotion expression estimation system will be described. FIG. 4 is a diagram showing a configuration example of each device of the emotion expression estimation system according to the present embodiment. The robot 2 includes, for example, a photographing unit 102 (sensor), a sound collection unit 103 (sensor), an execution unit 104, a modification unit 105, an acquisition unit 106 (sensor), an emotion estimation unit 107 (sensor), an emotion model 108, a comparison unit 109, an emotion teaching unit 110, a display unit 111, a voice output unit 112, a drive unit 113, a control unit 114, and a storage unit 115. In FIG. 4, the housing 120, boom 121, horizontal bar 122, movable parts 131, 132, and 134, etc. described with reference to FIG. 2 are omitted for illustration. Also, the robot 2 may not include the photographing unit 102 and the sound collection unit 103, and in that case, it may acquire the image and the audio signal acquired by the environment sensor 3. The environment sensor 3 includes, for example, a photographing unit 301, a sound collection unit 302, and a communication unit 303.
[0032] (Environment sensor) The photographing unit 301 is, for example, a stereo camera or an RGBD camera capable of obtaining depth information. The photographing unit 301 outputs the captured image to the robot 2 via the communication unit 303.
[0033] The sound collection unit 302 is a microphone, and may be a microphone array including a plurality of microphones. The sound collection unit 302 outputs the collected audio signal to the robot 2 via the communication unit 303.
[0034] The communication unit 303 outputs the image and the acoustic signal to the robot 2. Note that the transmission timing is, for example, at regular intervals. The captured image signal includes an image of the operator's face. The acoustic signal includes game sounds and the operator's speech.
[0035] (Robot) During the game, the imaging unit 102, for example, captures an image of the game board to capture the progress of the game. The imaging unit 102 may be, for example, an RGB (Red, Green, Blue) camera, or may be an RGBD camera that can also acquire depth information D. Note that the imaging unit 102 captures an image of the child's face during communication with the child.
[0036] During the game, the sound collection unit 103, for example, collects game sound effects, the voices of children, etc. The sound collection unit 103 is a microphone, or may be a microphone array including a plurality of microphones.
[0037] The execution unit 104, for example, executes, based on a program stored in the storage unit 115, an electronic puppet show that evokes a specific emotion. In this embodiment, when reading a puppet show, the "correct" mirror emotion of a child who does not have attention deficit / hyperactivity disorder (ADHD) is defined as the "specific emotion". The execution unit 104 outputs the image data of the electronic puppet show to the image display device 4 via a wired or wireless network NW. Further, the execution unit 104 outputs text data or audio data associated with the image data of the electronic puppet show to the audio output unit 112. Note that parameter information related to reading, such as the speed of reading, the tone of the voice being read, the type of story being read, and the movement of the robot 2, is added to the text data or audio data. Note that the parameter information represents the expressiveness during the reading.
[0038] Based on the result estimated by the emotion estimation unit 107, the modification unit 105 changes the way of speaking to prompt the user's facial emotion to approach a specific emotion.
[0039] The acquisition unit 106 acquires the image and acoustic signal output by the environment sensor 3.
[0040] The emotion estimation unit 107 performs well-known image recognition processing on the image acquired by the acquisition unit 106 to estimate the operator's expression. The emotion estimation unit 107 performs well-known speech recognition processing on the acoustic signal acquired by the acquisition unit 106 to extract and recognize the operator's speech signal. The emotion estimation unit 107 estimates the emotion mirroring of the user based on the recognized expression and speech of the operator (Estimation of emotion mirroring). Note that the method for acquiring emotions may be performed, for example, by the method described in Japanese Patent Application Laid-Open No. 2023-026244. The emotion estimation unit 107 generates images of eyes and mouth to be displayed on the display unit 111 based on the estimation result. The emotion estimation unit 107 generates a speech signal to be output to the speech output unit 112 based on the estimation result.
[0041] The emotion model 108 (Emotion model) models and stores the "correct" mirror image emotions (verification of the child's emotions and the story's emotions) of children who are not inattentive or hyperactive when reading a picture book aloud. Note that the emotion model 108 may be placed, for example, on the cloud.
[0042] The comparison unit 109 compares the apparent emotion of the user estimated by the emotion estimation unit 107 after the story of the execution unit 104 with the specific emotion stored in the emotion model 108.
[0043] Based on the comparison between the estimated value of the emotion mirroring of child Hu and the selected emotion, the emotion instruction unit 110 causes the corresponding expression to be displayed on the display unit 111 or outputs a speech signal from the speech output unit 112 to speak.
[0044] The display unit 111 displays the images of eyes and mouth generated by the emotion estimation unit 107.
[0045] The speech output unit 112 converts the text data of the story output by the execution unit 104 into a speech signal and outputs it. Alternatively, the speech output unit 112 outputs to the speech signal of the story output by the execution unit 104. Further, the speech output unit 112 outputs the speech signal generated by the emotion estimation unit 107.
[0046] The drive unit 113 includes, for example, a drive circuit, an actuator, and sensors such as an encoder. The drive unit 113 drives and operates the actuators attached to the movable parts 131 to 134, the boom 121, and the horizontal bar 122 according to the control of the control unit 114.
[0047] The control unit 114 generates a control signal for driving the drive unit 113 using the information indicating emotions and the like output by the emotion teaching unit 110.
[0048] The storage unit 115 stores information related to the paper theater, programs used for controlling the robot 2, threshold values, predetermined values, mathematical formulas used in processing, and the like.
[0049] [Paper theater data] Next, an example of the paper theater data stored in the storage unit 115 will be described. FIG. 5 is a diagram showing an example of the paper theater data stored in the storage unit according to the present embodiment. As shown in FIG. 5, the storage unit 115 stores the image data in association with the "scenario, dialogue data" for each scene. Note that the storage unit 115 may also store the "sound effect data" in association with the image data. Also, the "scenario, dialogue data" is text data or voice data. Note that the storage unit 115 stores data such as that shown in FIG. 5 for each story.
[0050] [Emotion model] Next, examples of the input / output data during learning and use of the emotion model 108 will be described. FIG. 6 is a diagram showing examples of the input / output data during learning and use of the emotion model according to the present embodiment. Reference numeral g201 indicates an example of the input / output data of the emotion model 108 during learning. Reference numeral g202 indicates an example of the input / output data of the emotion model 108 during use.
[0051] During learning, as with reference sign g201, in the emotion model 108, an image of the face of child Hu during reading aloud, an audio signal uttered by child Hu during reading aloud, and information indicating the correct "emotion" which is teacher data are input, and information indicating "emotion" is output. The robot 2 repeats learning until the difference between the information indicating the correct "emotion" which is teacher data and the information indicating the "emotion" of the output reaches a predetermined value, for example.
[0052] During use, as with reference sign g202, in the learned emotion model 108, an image of the face of child Hu during reading aloud and an audio signal uttered by child Hu during reading aloud are input, and information indicating "emotion" is output.
[0053] Note that the example described with reference to FIG. 6 is just an example and is not limited thereto. For example, the input to the emotion model 108 may be a feature amount of an image of the face of child Hu during reading aloud, or may be a feature amount of an audio signal uttered by child Hu during reading aloud.
[0054] [Example of processing content, example of processing procedure] Next, an example of the processing content and an example of the processing procedure performed by the robot 2 will be described. FIG. 7 is a diagram showing an example of the processing content and an example of the processing procedure performed by the robot according to the present embodiment. FIG. 8 is a flowchart of an example of the processing procedure performed by the robot according to the present embodiment.
[0055] (Step S1) The robot 2 selects an emotion based on, for example, information on child Hu given in advance, and starts a puppet show including emotional elements (reference sign g105). Note that the robot 2 sets parameter information related to reading aloud (reference signs g101 to g104) based on, for example, information on child Hu given in advance or based on initial values, and starts reading aloud.
[0056] (Step S2) The robot 2 observes using the result detected by the environment sensor 3 of the reaction of child Hu (reference sign g106). Note that the information detected by the environment sensor 3 preferably reflects the emotion underlying the story in the reaction.
[0057] (Step S3) The emotion estimation unit 107 of the robot 2 observes the reaction (such as voice and facial features) of the child Hu using the information acquired from the environment sensor 3 to confirm how emotion mirroring functions (reference sign g107).
[0058] (Step S4) The comparison unit 109 of the robot 2 compares the estimated value of the child's emotion mirroring with the selected emotion using the pre-trained emotion model 108 (reference sign g108) (reference sign g109). The comparison unit 109 extracts a reward (reference sign g110) based on the comparison result. Note that the reward may be, for example, to increase the value of the reward when the emotion of the child Hu can be acquired and the acquired emotion is within or close to a predetermined range of the emotion mirroring model (emotion model 108). Also, the reward may be, for example, to decrease the value of the reward when the emotion cannot be acquired or when the acquired emotion is outside the predetermined range of the emotion mirroring model. Note that the reward may be information indicating the comparison result, information indicating that the emotion is within or close to a predetermined range of the emotion mirroring model, information indicating that the emotion cannot be acquired, information indicating that the acquired emotion is outside the predetermined range of the emotion mirroring model, etc. The change unit 105 adjusts the parameter information (reference signs g101 to g104) regarding reading aloud according to the ability of the child Hu based on the comparison result and human-centered reinforcement learning (reference signs g111, g112).
[0059] (Step S5) The emotion teaching unit 110 of the robot 2 speaks to the child while making corresponding facial expressions (displayed on the display unit 111) and movements (of the robot 2) based on the comparison between the estimated value of the child Hu's emotion mirroring and the selected emotion. In other words, the emotion estimation unit 107 and the comparison unit 109 determine (reference sign g111) whether they understand the emotion provided by the child Hu based on the estimation result and the comparison result. If they do not understand the emotion provided by the child Hu, the emotion teaching unit 110 talks to the child in a way and with facial expressions that attract the child so as to achieve the "correct emotion". If they understand the emotion provided by the child Hu, the emotion teaching unit 110, for example, makes a statement praising the child for correctly understanding the emotion. In this way, in the present embodiment, the "selected emotion" encourages the child to do their best to reflect the emotion expressed in the storytelling (reading aloud).
[0060] (Step S6) The modification unit 105 of the robot 2 continues the reading aloud while adjusting the emotional undertone and expressiveness of the selected emotion and changing the narration. Note that the modification unit 105 may change the story to be read aloud and start reading the story from the beginning based on the estimation result and the comparison result.
[0061] (Step S7) The robot 2 returns the process to Step S2 and repeats until the child's emotion mirroring for the selected emotion is satisfied based on the estimation.
[0062] (Step S8) The robot 2 selects another emotion, repeats the above steps, and adjusts the emotional expression of the narration so that the emotion mirroring of the child Hu's emotions and facial expressions is maximized through human-centered reinforcement learning and encouraging speech. Note that maximizing the emotion mirroring of the child Hu's emotions and facial expressions means, for example, making the difference between the "correct" emotion stored in the emotion model 108 and the emotion of the child Hu fall within a predetermined range.
[0063] Note that the processing procedures and the like described with reference to FIGS. 7 and 8 are merely examples and are not limited thereto. For example, some processes may be performed simultaneously, and other processes may be added.
[0064] As described above, in this embodiment, for example, the "correct" mirrored emotion (mirrored emotion; the confirmation of the child's emotion and the emotion in the story) of a child who does not have ADHD is modeled by the emotion model 108 and used to compare with the actual emotion of an ADHD child during online use. The robot 2 uses the emotion model 108 in this way to detect the inconsistency or correctness of the emotion expressed by the child Hu, and adjusts the expressiveness (such as tone of voice, facial expression, movement, etc.) to teach the correct emotion. Then, the robot 2 can personalize the story according to various understanding levels and individual needs in terms of understanding the facial expressions and emotions of children with autism by estimating the emotional mirroring of the child as a reward signal. Thereby, the robot 2 can monitor the progress of the child regarding correctly identifying the emotional implications included in the story, for example, for evaluation by medical experts.
[0065] As described above, in this embodiment, first, the robot 2 selects a specific emotion from the stories prepared in advance for such purposes (for example, a specific emotion is set for each story), and reads aloud for emotion training. Next, the robot 2 observes the reaction of the child Hu (checks the facial expression or detects the emotion) using the emotion model 108. If the emotion expressed in the robot 2's way of speaking does not reach the child Hu well, the robot 2 teaches while prompting, for example, "This story is sad, so being sad is like this." After that, the robot adjusts the expressiveness by the speed of the storytelling speech, the tone of the voice, etc., to show sadness (or any emotion). When the child Hu correctly reflects that emotion, the robot 2 praises the child Hu by saying, for example, "You did a great job," and then performs a happy routine. Robot 2 can continue the reading aloud and focus on emotions that are difficult for children to mirror. Furthermore, Robot 2 learns from the rewards provided by the estimation of the emotional mirroring of child Hu, and adjusts the expression of the emotional reading aloud so that the emotional mirroring of child Hu's emotions / facial expressions is maximized.
[0066] Thus, according to this embodiment, Robot 2 can be trained to improve the child's emotional expression. As a result, according to this embodiment, the child's emotional expression can be improved by the reading aloud of Robot 2.
[0067] In the embodiment, a child with ADHD is described as an example of the user, but the user is not limited to this. The user may be, for example, a person who cannot express emotions richly, or not only a child but also an adult. Also, in the embodiment, an example of one user is described, but there may be a plurality of users. In this case, for example, Robot 2 may correspond to a plurality of users in turn.
[0068] Also, in the above-described embodiment, an example in which Robot 2 reads a story to child Hu while displaying the paper puppet on the image display device 4 has been described, but it is not limited to this. For example, another person may actually turn the pages or operate the paper puppet. Or, like a radio drama, Robot 2 may perform the reading aloud using only sound effects and dialogue without using images.
[0069] Robot 2 may obtain a policy on which emotions to train from a person related to child Hu (such as a guardian, teacher, etc.) in advance and select a story to read aloud. Or, Robot 2 may obtain a policy on which emotions to train through several small readings aloud or conversations and select a story to read aloud.
[0070] Note that a program for realizing all or part of the functions of the robot 2 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform all or part of the processing performed by the robot 2. Here, the "computer system" shall include hardware such as an OS and peripheral devices. Also, the "computer system" shall include a WWW system equipped with a homepage providing environment (or display environment). Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, magneto-optical disk, ROM, CD-ROM, etc., and a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" shall also include a volatile memory (RAM) inside a computer system that becomes a server or a client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and that holds the program for a certain period of time. Alternatively, some or all of these components may be realized by hardware (including a circuit part; circuitry) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), SOC (System On Chip), or may be realized by the cooperation of software and hardware.
[0071] Further, the above program may be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by a transmission wave in the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium having a function of transmitting information, such as a network (communication network) like the Internet or a communication line (communication wire) like a telephone line. Also, the above program may be for realizing a part of the functions described above. Furthermore, it may be a so-called difference file (difference program) that can realize the functions described above in combination with a program already recorded in a computer system.
[0072] As described above, the embodiments for implementing the present invention have been described using the embodiments. However, the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the gist of the present invention.
Explanation of Reference Numerals
[0073] 1... Emotion Expression Estimation System, 2... Robot, 3... Environment Sensor, 4... Image Display Device, 102... Photographing Unit, 103... Sound Collection Unit, 104... Execution Unit, 105... Change Unit, 106... Acquisition Unit, 107... Emotion Estimation Unit, 108... Emotion Model, 109... Comparison Unit, 110... Emotion Instruction Unit, 111... Display Unit, 112... Voice Output Unit, 113... Driving Unit, 114... Control Unit, 115... Storage Unit, 120... Housing, 121... Boom, 122... Horizontal Bar, 131... Movable Part, 132... Movable Part, 133... Movable Part, 134... Movable Part, 301... Photographing Unit, 302... Sound Collection Unit, 303... Communication Unit
Claims
1. A sensor that acquires user information, An execution unit that tells a story that evokes a specific emotion, A comparison unit that compares the user's facial emotion with the specific emotion using the user information acquired by the sensor when the execution unit tells the story, Comprising, Based on the comparison result, the comparison unit outputs an instruction to bring the user's facial emotion closer to the specific emotion. An emotion expression estimation system.
2. A modification unit that changes the way of speaking to encourage the user's facial emotion to approach the specific emotion, The emotion expression estimation system according to claim 1, further comprising.
3. An emotion teaching unit that encourages the user's facial emotion to approach the specific emotion, The emotion expression estimation system according to claim 1 or claim 2, further comprising.
4. The specific emotion is classified into "excitement", "love", "fear", "anger", "sadness", "joy" in psychology, The emotion expression estimation system according to claim 1 or claim 2.
5. The modification unit changes at least one of the speed of speaking, the tone of speaking, and the type of story, The emotion expression estimation system according to claim 2.
6. The emotion teaching unit, When the user does not understand the emotion of the story provided, talk to attract the child so as to have the correct emotion, When the user understands the emotion of the story provided, make a statement praising that the emotion is correctly understood, The emotion expression estimation system according to claim 3.
7. The sensor, Comprises a photographing device and a sound collecting unit, Using at least one of the image captured by the photographing device and the acoustic signal collected by the sound collecting unit to estimate the user's emotion, thereby acquiring the user's information, The emotion expression estimation system according to claim 1 or claim 2.
8. A sensor that acquires user information, An execution unit that tells a story that evokes a specific emotion, A comparison unit that compares the user's facial emotion with the specific emotion using the user information acquired by the sensor when the execution unit tells the story, Comprising, Based on the comparison result, the comparison unit outputs an instruction to bring the user's facial emotion closer to the specific emotion. An emotion expression estimation device.
9. The sensor acquires user information, The execution unit tells a story that evokes a specific emotion, The comparison unit compares the apparent emotion of the user with the specific emotion using the information of the user acquired by the sensor when the execution unit makes the user listen to the story. Based on the result of the comparison, the comparison unit outputs an instruction to bring the apparent emotion of the user closer to the specific emotion. Emotion expression estimation method.
Citation Information
Patent Citations
Personalized digital therapeutic methods and devices
JP2022527946A