Information processing system, information processing device, and information processing method
The information processing system enhances attention in individuals with ADHD by using a virtual co-operator robot to engage in games, estimating and improving concentration through interactive gameplay.
Patent Information
- Application Number
- JP2023216727
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-12-22
AI Technical Summary
Existing developmental disorder support systems fail to improve attention in individuals with disorders such as ADHD, despite aggregating and sharing various types of information.
An information processing system that includes an acquisition unit, a policy unit, an output unit, a discriminator, and a reward network to enhance attention by interacting with individuals through a virtual co-operator, such as a robot, using games to estimate and improve concentration and interest.
The system effectively improves attention and concentration in individuals with ADHD by providing interactive and engaging gameplay experiences, utilizing a robot to monitor and respond to the individual's attention levels and provide feedback.
Smart Images

Figure 2025099794000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system, an information processing apparatus, and an information processing method.
Background Art
[0002] Attention-deficit / hyperactivity disorder (ADHD) usually presents symptoms such as inattention (e.g., lack of concentration) and hyperactivity / impulsivity (e.g., restlessness, inability to wait one's turn). This disorder poses a problem for children, parents, and teachers because it makes it difficult for children to concentrate and affects their performance in daily life at home and school.
[0003] For such disorders, a developmental disorder support system has been proposed (see, for example, Patent Document 1). In such a developmental disorder support system, supportee identification information of a supportee and a plurality of types of test information regarding the supportee are associated in advance. Authority information for associating the supportee identification information with the supporter identification information of the supporter is also stored in the support server. Then, the support server receives test information regarding a supportee who is a person with a developmental disorder. When the support server receives a request for acquisition of test information, the requester of the test information becomes the target of the test information. On the condition that the requester has access authority regarding the supportee, the test information is transmitted to the requester. This enables the aggregation of various types of information regarding persons with developmental disorders and allows the relevant parties to smoothly share the information while taking privacy into consideration.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the prior art's developmental disorder support system, even though various types of information regarding individuals with developmental disorders could be aggregated and shared, it was not possible to improve the attention of individuals with developmental disorders.
[0006] The present invention has been made in view of the above problems, and an object thereof is to provide an information processing system, an information processing apparatus, and an information processing method capable of improving attention.
Means for Solving the Problems
[0007] (1) To achieve the above object, an information processing system according to an aspect of the present invention includes an acquisition unit that acquires the situation of an operator (e.g., a child), a policy unit (Generator (policy)) that has information on a virtual co-operator (robot) that plays a game with rules together with the operator, an output unit (speaker, display unit) that can output an approach of the virtual co-operator to the operator, a discriminator (Discriminator) trained using the activities of the operator and the virtual co-operator, and a reward network (Human reward net) that predicts the reward of the operator. The policy unit learns a policy for operating the game of the virtual co-operator from the reward provided by the reward network and the discriminator, and the output unit performs an output that prompts the interest of the operator based on the information acquired by the acquisition unit and the information of the policy unit.
[0008] (2) In the information processing system according to an aspect of the present invention described in (1) above, the output unit may perform an operation of the game that prompts the interest of the operator based on the information acquired by the acquisition unit and the information of the policy unit.
[0009] (3) In the information processing system according to one aspect of the present invention described in (1) or (2) above, the information acquired by the acquisition unit includes at least one of an image including the face of the operator and an audio signal of the operator. Using the image including the face of the operator, estimate the line-of-sight direction of the operator, estimate the time during which the line-of-sight direction remains using the image including the face of the operator, estimate the expression of the operator using the image including the face of the operator, and estimate the degree to which the operator shows interest in the game using the audio signal of the operator. By performing at least one of the above estimations, an attention estimation unit for estimating the attention of the operator may be further provided.
[0010] (4) In the information processing system according to one aspect of the present invention according to any one of (1) to (3) above, the discriminator compares, for each progress of the game, the state of the game of the operator and the activity that is a combination of actions with respect to the game with the state of the game of the virtual co-operator and the activity that is a combination of actions with respect to the game, and may obtain the reward of the discriminator.
[0011] (5) In the information processing system according to one aspect of the present invention described in (3) above, when the result estimated by the attention estimation unit indicates that the attention of the operator is sustained or improved, the output unit outputs a statement praising the operator. When the result estimated by the attention estimation unit indicates that the attention of the operator is not sustained, the output unit outputs a statement prompting the operator to participate further in the game. When the time during which the operator has not participated in the game continues for a predetermined time, the output unit may output a statement prompting the operator to participate in the game.
[0012] (6) In the information processing system according to one aspect of the present invention according to any one of (1) to (4) above, during learning, the state of the game and the actions of the operator may be weighted to train the reward network.
[0013] (7) In the information processing system according to any one of aspects (1) to (4) and (6) of the present invention, an environmental sensor including a photographing device and a sound collection unit is further provided, and the acquisition unit may acquire at least one of an image photographed by the environmental sensor and an acoustic signal collected by the sound collection unit to acquire the situation of the operator.
[0014] (8) In the information processing system according to any one of aspects (1) to (4), (6), and (7) of the present invention, the virtual co-operator is a robot, the robot includes a voice output unit and a display unit, the output unit is at least one of the voice output unit and the display unit, and the robot may output an approach to the operator of the virtual co-operator by performing at least one of speaking from the voice output unit and displaying an image expressing an expression on the display unit to the operator.
[0015] (9) To achieve the above object, an information processing apparatus according to an aspect of the present invention includes an acquisition unit that acquires the situation of an operator, a policy unit that has information on a virtual co-operator who plays a game with rules together with the operator, an output unit that generates an approach to the operator of the virtual co-operator and can output it to the virtual co-operator, a discriminator trained using the activities of the operator and the activities of the virtual co-operator, and a reward network that predicts the reward of the operator. The policy unit learns a policy for operating the game of the virtual co-operator from the reward provided by the reward network and the discriminator, and the output unit generates an output that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit and outputs it to the virtual co-operator.
[0016] (10) To achieve the above object, an information processing method according to an aspect of the present invention is as follows in an information processing apparatus. An acquisition unit acquires the situation of an operator. When a policy unit has information of a virtual co-operator who plays a game with rules together with the operator, an output unit generates an approach of the virtual co-operator to the operator and outputs it to the virtual co-operator. A discriminator trained using the activities of the operator and the activities of the virtual co-operator compares the activities of the operator and the activities of the virtual co-operator to obtain a reward. A reward network predicts the reward of the operator. The policy unit learns a policy for operating the game of the virtual co-operator from the rewards provided by the reward network and the discriminator. The output unit generates an output that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit, and outputs it to the virtual co-operator. This is the information processing method.
Advantages of the Invention
[0017] (1) - (10) can improve attention.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scales of the respective members are appropriately changed in order to make the respective members recognizable in size. In all the drawings for explaining the embodiments, those having the same function are denoted by the same reference numerals, and repeated explanations are omitted. In addition, “based on XX” as used in the present application means “based on at least XX”, and includes cases where it is based on another element in addition to XX. Further, “based on XX” is not limited to the case where XX is directly used, and includes cases where it is based on something obtained by performing an operation or processing on XX. “XX” is an arbitrary element (for example, arbitrary information).
[0020] [Configuration Example of Information Processing System] FIG. 1 is a diagram showing a configuration example of an information processing system according to the present embodiment. The information processing system 1 includes, for example, a robot 2 (virtual co-operator), an environmental sensor 3, an image display device 4, and an information processing device 5.
[0021] The child Hu (operator) operates, for example, a game provided by the information processing device 5 by operating an operation unit 9 (for example, a game controller) while communicating with the robot 2.
[0022] Robot 2 is a robot capable of communicating with child Hu. Robot 2 includes, for example, a main body, a display unit capable of displaying expressions of eyes and mouth, a boom capable of operating the eye part, an audio output unit for outputting an audio signal, a sound collection unit, a camera, a control unit for controlling each functional unit, etc. Robot 2 monitors the attention and concentration of child Hu using the information output by information processing device 5. When child Hu is losing concentration, Robot 2 uses the information output by information processing device 5 to arouse attention and assist in maintaining concentration. Note that the configuration, operation, processing, etc. of Robot 2 will be described later.
[0023] Environment sensor 3 performs, for example, detection of the line of sight of child Hu, detection of the face direction, detection of an audio signal, etc. Environment sensor 3 includes, for example, a stereo camera or an RGBD camera capable of obtaining depth information, and a sound collection unit.
[0024] Image display device 4 displays an image signal in the case of a game or the like where an image is displayed for display. Note that image display device 4 may be provided with a speaker for outputting an acoustic signal of the game.
[0025] Information processing device 5 provides, for example, a game. Information processing device 5 may be provided with a game device. Information processing device 5 acquires the operation result of child Hu operating operation unit 9. Information processing device 5 acquires the detection information detected by environment sensor 3. Information processing device 5 uses the acquired information to perform the processing and judgment described later, and outputs the processed and judged information to Robot 2. Note that information processing device 5 and Robot 2 are connected to each other by wire or wirelessly. Note that the configuration, operation, processing, etc. of information processing device 5 will be described later.
[0026] In this way, in the present embodiment, Robot 2 provides an attractive interaction for child Hu with ADHD, for example, both by vocalization and visual animation, and arouses attention. Robot 2 trains and maintains the attention of child Hu, for example, through a game. Note that in the embodiment, a child with ADHD is described as an example of the person to be supported.
[0027] [Example of the external shape of the robot] Next, an example of the external shape of the robot 2 will be described. FIG. 2 is a diagram showing an example of the external shape of the robot according to the present embodiment. In FIG. 2, the front view g101 and the side view g102 are diagrams showing an example of the external shape of the robot 2 according to the embodiment. The robot 2 includes, for example, three display units 111 (eye display unit 111a, eye display unit 111b, mouth display unit 111c). Also, in the example of FIG. 2, the imaging unit 102a is attached above the eye display unit 111a, and the imaging unit 102b is attached above the eye display unit 111b. The eye display units 111a and 111b correspond to human eyes and present images and image information corresponding to human eyes. The screen size of the eye display units 111a and 111b is, for example, 3 inches. For example, the audio output unit 112, which is a speaker, is attached in the vicinity of the mouth display unit 111c that displays an image corresponding to the human mouth of the housing 120. The mouth display unit 111c is composed of, for example, a plurality of LEDs (light emitting diodes), each LED can be address-specified, and can be individually driven to turn on and off. The sound collection unit 103 is attached to the housing 120.
[0028] Also, the robot 2 includes a boom 121. The boom 121 is movably attached to the housing 120 via a movable part 131. A horizontal bar 122 is rotatably attached to the boom 121 via a movable part 132. The eye display unit 111a is rotatably attached to the horizontal bar 122 via a movable part 133, and the eye display unit 111b is rotatably attached to the horizontal bar 122 via a movable part 134. Note that the external shape of the robot 2 shown in FIG. 2 is an example and is not limited thereto.
[0029] [Examples of emotional expressions and actions of the robot] FIG. 3 is a diagram showing an example of presenting the emotional expression of the robot according to the present embodiment in animation. As shown in FIG. 3, the robot 2 presents an emotional expression by changing the animation displayed on the display unit 111 (eye display unit 111a, eye display unit 111b, mouth display unit 111c). The animation examples of symbols g11 to g17 are "Angry", "Ecstatic", "Disinterested", "Confused", "Blushing", "Sad", and "Sympathetic", respectively.
[0030] Note that the emotional expressions and animations shown in FIG. 3 are just examples and are not limited to this. The emotional expressions may exist other than in FIG. 3, and the animations of each emotional expression may be different from those in FIG. 3. Also, when each emotional expression is made, the angle and position of the display unit 111 are changed as in FIG. 3, the angle of the boom 121 is changed, or an audio signal is output together.
[0031] [Configuration Examples of Each Device of the Information Processing System] Next, configuration examples of each device of the information processing system will be described. FIG. 4 is a diagram showing configuration examples of each device of the information processing system according to the present embodiment. The robot 2 includes, for example, a photographing unit 102, a sound collection unit 103, a generation unit 110, a display unit 111 (output unit), an audio output unit 112 (output unit), a drive unit 113, a control unit 114, a storage unit 115, and a communication unit 116. In FIG. 4, the housing 120, boom 121, horizontal bar 122, movable parts 131, 132, and 134, etc. described with reference to FIG. 2 are omitted. Also, the robot 2 may not include the photographing unit 102 and the sound collection unit 103, and in that case, the robot 2 may acquire the image and audio signal acquired by the environmental sensor 3. The environmental sensor 3 includes, for example, a photographing unit 301, a sound collection unit 302, and a communication unit 303. The information processing device 5 includes, for example, an execution unit 501, an acquisition unit 502, an audio-visual processing unit 503, an attention estimation unit 504, a policy unit 505, an output unit 506, a discriminator 507, a human reward network 508, a learning unit 509, a storage unit 510, a communication unit 511, and a preprocessing unit 512.
[0032] (Environmental Sensor) The imaging unit 301 is, for example, a stereo camera or an RGBD camera capable of obtaining depth information. The imaging unit 301 outputs the captured image to the information processing device 5 via the communication unit 303.
[0033] The sound collection unit 302 is a microphone, and may be a microphone array including a plurality of microphones. The sound collection unit 302 outputs the collected voice signal to the information processing device 5 via the communication unit 303.
[0034] The communication unit 303 outputs the image and the acoustic signal to the information processing device 5. The timing of transmission is, for example, at regular intervals. The captured image signal includes an image of the face of the child Hu. The acoustic signal includes, for example, game sounds (such as sound effects) and the speech of the child Hu.
[0035] Note that there may be two or more environmental sensors 3. In that case, for example, the first environmental sensor 3-1 detects information about the child Hu, and the second environmental sensor 3-2 detects the area of the game board on which the game is projected and the positional relationship between the game board and the child. The second environmental sensor 3-2 is installed at a position where it can capture both the game board and the child Hu, for example, on the ceiling or wall.
[0036] (Information Processing Device) The execution unit 501 executes a game, for example, based on a program stored in the storage unit 510. During the game, the execution unit 501 outputs the image information of the game to the image display device 4. The output information may include acoustic information. Also, during the game execution, the execution unit 501 acquires information indicating the operation result of the child Hu operating the operation unit 9. Note that during the game, the robot 2 may be configured to acquire an operation instruction for the output game.
[0037] The preprocessing unit 512 performs pre-processing on the game state of the child Hu output by the execution unit 501. The preprocessing unit 512 performs pre-processing on the game state of the robot 2 output by the execution unit 501. Note that the preprocessing on the game state of the child Hu is a process of extracting the position and changes of game items as a result of the operation by the child Hu. Also, the preprocessing on the game state of the robot is a process of extracting the position and changes of game items as a result of the operation by the child Hu. Note that the execution unit 501 may include the preprocessing unit 512, or may perform the processing performed by the preprocessing unit 512.
[0038] The acquisition unit 502 acquires the images and acoustic signals output by the environment sensor 3 to acquire and monitor the situation of the child Hu.
[0039] The audio-visual processing unit 503 performs well-known image recognition processing on the images acquired by the acquisition unit 502 to estimate the orientation, expression, and line-of-sight direction of the face of the child Hu. The audio-visual processing unit 503 performs well-known speech recognition processing on the acoustic signals acquired by the acquisition unit 502 to extract and recognize the voice signals of the child Hu.
[0040] The attention estimation unit 504 uses the information processed by the audio-visual processing unit 503 to estimate the degree to which the face direction of the child Hu is directed, for example, towards the game board (that is, within the attention area) as the attention level. Note that the attention estimation unit 504 has previously acquired the coordinates of the area of the game board where the game is projected based on the images captured by the environment sensor 3. For line-of-sight detection, for example, the method described in Japanese Patent Application Laid-Open No. 2023-026245 may be used.
[0041] The policy unit 505 holds information about the robot 2, which is a virtual co-operator that plays games with the child Hu. Also, the policy unit 505 uses the rewards and the like provided by the reward network 508 and the discriminator 507 to learn the policy for the robot 2 to operate the game. The policy unit 505 uses the rewards and the like provided by the reward network 508 and the discriminator 507 to generate approach instructions for the child Hu. Note that the learning method of the policy unit 505 will be described later.
[0042] The output unit 506 outputs the approach instructions for the child Hu generated by the policy unit 505 to the robot 2.
[0043] The discriminator 507 is trained using the activities of the child Hu and the robot 2. The discriminator 507 compares the trajectory of the child Hu's actions (Human trajectory) with the trajectory of the robot's actions (Robot trajectory) and outputs the comparison result. Note that the discriminator 507 outputs the difference, for example, every predetermined time, every screen, and for each operation on the game. Note that the learning method of the discriminator 507 will be described later.
[0044] The reward network 508 predicts the reward for the child Hu (human reward) based on the state output by the execution unit 501 and the actions of the child Hu. Note that the reward for the child Hu is, for example, information based on a state in which the child Hu can be made (or has been made) to concentrate more and be interested in the game, for example. For example, if the expression of the child Hu changes from expressionless to smiling, laughing, or the number of utterances increases, it is estimated that the concentration has improved, that is, the value of the reward is large. Conversely, if the expression of the child Hu remains expressionless, changes from smiling to expressionless, or the number of utterances decreases, it is estimated that the concentration has decreased, that is, the value of the reward is small. Note that the learning method of the reward network 508 will be described later.
[0045] The learning unit 509 may pre-train the policy unit 505, the discriminator 507, and the reward network 508. Alternatively, the learning unit 509 may train the policy unit 505, the discriminator 507, and the reward network 508 even during the game.
[0046] The memory unit 510 stores programs used by each functional unit of the information processing apparatus 5, predetermined values, threshold values, identification information for identifying the environmental sensor 3, identification information for identifying the robot 2, mathematical formulas used in processing, and the like.
[0047] The communication unit 511 transmits and receives information to and from the robot 2.
[0048] (Robot) The robot 2 may include a part of the functions of the information processing apparatus 5, for example, the policy unit 505 and the like. Alternatively, the robot 2 may include all of the functions of the information processing apparatus 5.
[0049] During the game, the imaging unit 102, for example, images the game board to image the progress of the game. The imaging unit 102 may be, for example, an RGB (red, green, blue) camera, or an RGBD camera that can also acquire depth information D. Note that the imaging unit 102 images the child's face during communication with the child.
[0050] During the game, the sound collection unit 103, for example, collects game sound effects, the voices of children, and the like. The sound collection unit 103 is a microphone, and may be a microphone array including a plurality of microphones.
[0051] During the game, the generation unit 110 generates an image to be displayed on the display unit 111 and a voice signal to be output from the voice output unit 112 based on the information output by the information processing device 5. Note that during communication with a child, the generation unit 110 performs well-known image processing on the captured image of the child's face to estimate the expression and emotion of the child's face, and generates an image to be displayed on the display unit 111 based on the estimated result. Further, during communication with a child, the generation unit 110 performs well-known voice recognition processing on the collected acoustic signal to estimate the emotion of the child, and generates a voice signal to be output to the voice output unit 112 based on the estimated result. Note that the method for obtaining emotions may be performed, for example, by the method described in Japanese Patent Application Laid-Open No. 2023-026244. The generation unit 110 outputs information indicating the estimated emotions and the like to the control unit 114.
[0052] The display unit 111 displays the image generated by the generation unit 110.
[0053] The voice output unit 112 outputs the voice signal generated by the generation unit 110.
[0054] The drive unit 113 includes, for example, a drive circuit, an actuator, and sensors such as an encoder. The drive unit 113 drives and operates the actuators attached to the movable parts 131 to 134, the boom 121, and the horizontal bar 122 in accordance with the control of the control unit 114.
[0055] The control unit 114 generates a control signal for driving the drive unit 113 using the information indicating emotions and the like output by the generation unit 110.
[0056] The storage unit 115 stores programs, threshold values, predetermined values, mathematical formulas used in processing, identification information of the information processing device 5, etc. used for controlling the robot 2.
[0057] The communication unit 116 transmits and receives information to and from the information processing device 5.
[0058] [Attention arousal, attention estimation] Next, the arousal of concentration and the estimation of concentration for the child's concentration will be described. FIG. 5 is a diagram showing an example of a state in which a robot and a child are playing a game. In the example of FIG. 5, the robot 2 and the child Hu are playing a video game with a predetermined rule (for example, an air hockey game). Reference numeral g21 is an example of a game image. The region of reference numeral g11 is an example of a region (attention region) for detecting the line-of-sight direction and facial expression of the child Hu. The region of reference numeral g12 is an example of a region (tracking region) for tracking the line-of-sight direction and expression even when the line-of-sight direction of the child Hu deviates from the attention region g11.
[0059] In the present embodiment, the environmental sensor 3 tracks the line-of-sight direction and fixation period of the child Hu, estimates the attention of the child Hu, and causes the child Hu to participate in the game in order to improve the concentration of the child Hu. The attention of the child Hu can be generally divided into the following three states. The robot 2 always (or every predetermined time) monitors it and uses it as a judgment material.
[0060] I. First state; The range of focus is constant (Constant focus) When the face direction of the child Hu is directed onto the game board (that is, within the attention region), the attention estimation unit 504 of the information processing device 5 outputs the estimation result information to the robot 2. Based on the acquired estimation result information, the robot 2 utters encouraging words to strengthen the "concentrated" attention state of the child Hu. The utterance content is, for example, "You're doing great! You're gradually understanding how to play the game" (an utterance praising the operator). Note that the robot 2 can also express rich facial expressions and daily movements through the display on the display unit 111 (eye display unit 111a, eye display unit 111b, mouth display unit 111c) and the movement of the boom 121 (for example, "happy" or "excited").
[0061] II. Second state; Partial loss of concentration (Losing focus (partial)) When the attention of child Hu moves outside the game board (outside area g11 and within area g12), and the fixation time in that area exceeds a threshold value (for example, 5 seconds), the attention estimation unit 504 outputs estimation result information, which is information indicating a partial loss of concentration in robot 2. Based on the acquired estimation result information, robot 2 gently prompts child Hu's attention, for example. For example, robot 2 says to child Hu, "Ah, if you don't keep up, I'll start winning!" (utterance to prompt the operator to further participate in the game), and creates a routine with rich expressions.
[0062] III. The third state; Losing focus (full) When the concentration of child Hu completely goes outside the screen (outside area g12), and the fixation time in this area and the time when the child is not operating during the game exceed a threshold value (for example, 5 seconds), the attention estimation unit 504 outputs estimation result information, which is information indicating a complete loss of concentration in robot 2. Based on the acquired estimation result information, robot 2 talks to child Hu, for example, "Shall we play again?" and moves its expression and emotion accordingly.
[0063] [Processing method] Next, an overview of the processing performed by robot 2 and information processing device 5 will be described. FIG. 6 is a diagram showing an overview of the processing performed by the robot and information processing device according to this embodiment. The information processing device 5 starts the game. Reference numeral g101 is an example of an image of the game. Note that the game is one that child Hu knows the rules of. For example, before starting the game, robot 2 may communicate with child Hu through conversation to find out whether child Hu knows the rules. For example, robot 2 may ask questions such as "What kind of games do you like?" to detect games that child Hu knows the rules of. Note that the game can also be said to be a task for enhancing concentration.
[0064] Then, the robot 2 may output the detected detection result to the information processing device 5. The information processing device 5 may select a game based on the information acquired from the robot 2. Note that the robot 2 is a virtual co-operator and, depending on the game, plays with or competes against the child Hu alternately. Note that the robot 2 outputs an operation instruction for the game as a command to the information processing device 5 without operating the operation unit, for example. Note that the robot 2 may be provided with a hand having an arm or finger portions, etc., and in that case, the hand may be used to operate the operation unit for the robot 2.
[0065] The child Hu plays a game. By observing this, the robot 2 learns the way to play from the child Hu and learns the rules of the game. Also, while learning the way to play together with the child Hu, the robot 2 further monitors the child Hu's concentration (reference sign g121) and gives feedback to the child to maintain the child Hu's concentration (reference sign g122). Note that the monitoring of the concentration is performed by the attention estimation unit 504 of the information processing device 5 as described above (reference sign g131). The utterances are, as described above, utterances to make the child Hu interested in the game, praising utterances, encouraging utterances, utterances to improve concentration, etc.
[0066] Note that the robot 2 imitates the play strategy (rules) of the game of the child Hu. Thereby, the robot 2 (including the information processing device 5) can give a sense of satisfaction that the child Hu is not bored with the game and that the rules can be taught to the robot 2.
[0067] The attention estimation unit 504 estimates the attention of child Hu by tracking the direction and time in which child Hu gazes at the game board, the expression of child Hu, and the vocal expression of child Hu (reference sign g114). That is, the information processing apparatus 5 determines the concentration state of child Hu (reference sign g111), performs feedback (reference sign g112), learns using the GAILHF method (reference sign g113), and outputs the result to the robot 2. The GAILHF method is a reinforcement learning approach that combines the ideas of imitation learning and generative adversarial networks (GANs). This is a method used to learn a policy from expert demonstrations without requiring an explicit reward function, enabling learning even without an actual reward.
[0068] Accordingly, according to the present embodiment, through attractive and competitive interaction with the robot 2 based on attention estimation, the concentration of child Hu can be improved. As a result, according to the present embodiment, child Hu can be trained to pay more attention in order to win the game, or child Hu can teach the robot 2 the rules of the game.
[0069] [Example of processing procedure] Next, an example of the processing procedure performed by the information processing system 1 will be described. FIG. 7 is a diagram showing an example of the processing procedure of the information processing system according to the present embodiment. FIG. 8 is a flowchart of an example of the processing procedure of the information processing system according to the present embodiment.
[0070] In FIG. 7, reference sign g201 corresponds to the game execution unit 501. Reference sign g205 corresponds to the reward network 508. Reference sign g207 corresponds to the policy unit 505. Reference sign g209 corresponds to the discriminator 507. Reference signs g202 and g203 correspond to the preprocessing unit 512.
[0071] Reference sign g204 is a set {τ1^, τ2^, …, τ n ^} of trajectories of the state of child Hu at each predetermined time or for each screen. Reference sign g208 is a set {τ1, τ2, …, τn} is. Also, s is the game state, a is the action of child Hu or the robot, and the superscript ^ indicates the estimated result.
[0072] (Step S1) The preprocessing unit 512 preprocesses a set of a series of states s and actions a of child Hu during the game play (reference sign g202). The preprocessing unit 512 stores the preprocessed set (s’^, a^) in the storage unit 510 as the human trajectory (reference sign g204). The preprocessing unit 512 preprocesses a set of a series of states s and actions a of the robot 2 during the game play (reference sign g203). The preprocessing unit 512 stores the preprocessed set (s’, a) in the storage unit 510 as the trajectory of the robot 2 (reference sign g204). Note that Game env (reference sign g201) outputs a series of states s of child Hu during the game play and a series of states s of the robot 2 during the game play to the preprocessing unit 512 and the Human reward ner (reference sign g209).
[0073] (Step S2) The learning unit 509 learns the discriminator 507 using both the trajectory of the robot 2 and the human trajectory. Subsequently, the learning unit 509 uses the discriminator 507 to give a reward and learns a policy part (i.e., the policy of the robot 2) for the robot 2 to play the game by inverse reinforcement learning.
[0074] (Step S3) The information processing system 1 can give positive or negative feedback to child Hu based on the evaluation of the performance of the game play from the robot 2 (reference sign g212). This feedback is used together with the discriminator 507 to provide a reward (in this way, the robot 2 can learn the game).
[0075] (Step S4) In order not to burden the child Hu (the operator) by providing feedback and to train the reward network 508 that predicts the reward of the child Hu, the learning unit 509 collects a small amount of feedback (reference sign g207). For example, the learning unit 509 may weight a small value for the reward from the child Hu and add it (reference sign g207) to other rewards.
[0076] (Step S5) The robot 2 learns the game rules and the like from the rewards provided by both the reward network 508 and the discriminator 507 (g206, g207, g208, g209).
[0077] (Step S6) The attention estimation unit 504 estimates the degree of attention of the child Hu during the game play by tracking the line-of-sight direction, time, expression, and voice tone of the child Hu during the game play (reference sign g210). Subsequently, the robot 2 determines the concentration state of the child Hu from the estimated degree of attention of the child Hu (reference sign g211), and makes a corresponding utterance to attract the child (reference sign g212) (involvement in the degree of attention of the child Hu). In addition, when the robot 2 gains or loses points, the robot 2 controls to further draw out the interest of the child Hu by showing expressions of joy or sadness and uttering words expressing emotions (involvement in the degree of attention of the child Hu).
[0078] The robot 2 repeats Steps S2 to S6 to play, learn, and communicate with the child Hu, and improves its concentration through attractive and competitive communication with the child Hu. In this case, the child Hu can be trained to pay more attention to win the game, or can teach the robot 2 the game.
[0079] [Policy Unit, Reward Network, Discriminator] Next, an example of the learning method and an example of the input / output during use of the policy unit 505, the reward network 508, and the discriminator 507 will be described. FIG. 9 is a diagram showing an input example during the learning of the policy unit, the reward network, and the discriminator according to the present embodiment, and an input / output example during use.
[0080] During learning, like in reference sign g301, the policy unit 505 receives as input the preprocessing result for the game state of robot 2 output by the preprocessing unit 512, the reward, and the teacher data (correct data), and outputs an action instruction for robot 2 to perform on child Hu. Note that the reward is information obtained by adding the output of the reward network 508 and the output of the discriminator 507. During use, like in reference sign g301, the policy unit 505 receives as input the preprocessing result for the game state of robot 2 output by the preprocessing unit 512 and the reward, and outputs an action instruction for robot 2 to perform on child Hu.
[0081] During learning, like in reference sign g311, the reward network 508 receives as input the game state output by the execution unit 501, the action a^ of child Hu, and the teacher data (correct data), and outputs the estimated reward for child Hu. During use, like in reference sign g311, the reward network 508 receives as input the game state output by the execution unit 501 and the action a^ of child Hu, and outputs the estimated reward for child Hu.
[0082] During learning, like in reference sign g321, the discriminator 507 receives as input the action trajectory of child Hu, the action trajectory of robot 2, and the teacher data (correct data), and outputs information indicating the discrimination result based on the difference between the action trajectory of child Hu and the action trajectory of robot 2. During use, like in reference sign g321, the discriminator 507 receives as input the action trajectory of child Hu and the action trajectory of robot 2, and outputs information indicating the discrimination result based on the difference between the action trajectory of child Hu and the action trajectory of robot 2.
[0083] Note that the learning may be performed in advance as described above, or may be performed during the game.
[0084] As described above, in this embodiment, the attention of child Hu can be trained and maintained as follows. I. Child Hu can teach robot 2 a game, and robot 2 can learn how to play the game with child Hu. II. While child Hu is teaching the robot 2 a game, the robot 2 can monitor child Hu's attention and concentration. III. When child Hu is about to lose concentration, the robot 2 can draw attention through speech or the like and help maintain concentration. IV. The robot 2 can also monitor the metrics of their interaction in the form of measurement values that are useful for, for example, medical experts to evaluate the child's performance, such as the reference sign g131 in FIG. 6.
[0085] Note that in the above example, the person to be supported is not limited to a child with ADHD, and may be an elderly person, a person with low attention, etc. Also, in the embodiment, an example of one person to be supported is described, but there may be a plurality of them. In that case, the information processing system 1 may, for example, observe the expressions, gazes, and speech of each child and make corresponding speeches or the like for each child. Note that when there are a plurality of children, since the progress of the game and the way of teaching the rules may differ for each child, the reward network 508 and the discriminator 507 may be provided for each child.
[0086] [First Modification Example] Note that in the above example, an example has been described in which the operator, child Hu, knows the rules of the game to some extent and aims to improve concentration and the like by proceeding with the game and teaching the robot 2 the rules of the game, but it is not limited to this. For example, depending on child Hu, there may be cases where the game cannot be advanced well. In such a case, the robot 2 and the information processing device 5 may show child Hu how to play the game to promote understanding of the rules. Alternatively, when child Hu is not good at the game and cannot proceed easily, the information processing device 5 may lower the difficulty level of the game or change to another game. Also, the robot 2 may play against child Hu according to the level of the game of child Hu so that child Hu's concentration does not decrease and the interest can be improved and sustained.
[0087] In addition, the speech that the robot 2 makes to the child Hu can be either positive or negative. Such classification is performed by the processing of the discriminator 507. The discriminator 507 may, for example, perform discrimination with weighting on either the trajectory of the child Hu's actions or the trajectory of the robot 2's actions.
[0088] In this way, in the present embodiment, the robot 2 and the information processing device 5 may select, change, or adjust the battle level of the game so that the concentration of the child Hu does not decrease and the interest can be improved and sustained.
[0089] [Second Modification Example] Note that the speech of the robot 2 described with reference to FIG. 5 is an example and is not limited thereto. Also, the game is not limited to a so-called TV game displayed on a monitor or the like, and may be, for example, a board game or a card game that does not use a display unit. Also, the game may be a battle type or may be one that the child Hu plays alone. Or, for example, it may be a game that the child Hu plays while wearing a VR (Virtual Reality) goggle or the like.
[0090] When the game is a board game or a card game, the execution unit 501 stores in advance the rules of the game and the battle patterns. Note that the battle patterns may be placed on the cloud or may be acquired via a network. Also, in such a case, it is preferable that the robot 2 is provided with at least one set of arms and hands. Then, the execution unit 501 causes the robot 2 to perform, for example, the initial arrangement of the board game or the card game. Note that the initial arrangement of the board game or the card game may be performed by a person related to the operator (for example, parents, siblings, friends, teachers). In such a case, the execution unit 501 may, for example, grasp the situation of the game and the positions of each piece and card based on the image captured by the environment sensor 3. Then, the execution unit 501 may output the state of the game instead of the Game env (reference sign g201) in FIG. 5 based on the grasped information.
[0091] [Third Modification Example] In the above example, for instance, an example of a game where the child Hu and the robot 2 play alternately or in a battle was described, but it is not limited to this. For example, the information processing device 5 may become the parent of the game, play the game with the child Hu, and the robot 2 or the information processing device 5 may observe the state and actions of the child Hu at that time. Then, based on the observed results, the robot 2 may talk to the child Hu to improve concentration or praise the child Hu.
[0092] [Fourth Modification Example] In the above example, an example of improving or sustaining concentration by having the child Hu play a game was described, but such a tool is not limited to games. A game only needs to have a predetermined rule. For example, it may be building blocks, puzzles, other types of play, or a battle-type game. Even for such play, based on the observed results, the robot 2 may talk to the child Hu to improve concentration or praise the child Hu. In this way, the methods of the present embodiment and each modification example can be applied to various types of play including games.
[0093] Note that each of the above modification examples is just an example and is not limited to this. For example, other processes may be added to the above modification examples, or conversely, the processes may be reduced, or the modification examples may be combined with the above embodiments.
[0094] Note that a program for realizing all or part of the functions of the robot 2 and the information processing device 5 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform part or all of the processing performed by the robot 2 and the information processing device 5. Here, the "computer system" shall include hardware such as an OS and peripheral devices. Also, the "computer system" shall include a WWW system equipped with a homepage providing environment (or display environment). Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" shall include a volatile memory (RAM) inside a computer system that becomes a server or a client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and that holds the program for a certain period of time. Alternatively, some or all of these components may be realized by hardware (including a circuit part; circuitry) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), SOC (System On Chip), or may be realized by the cooperation of software and hardware.
[0095] In addition, the above program may be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by a transmission wave in the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium having a function of transmitting information, such as a network (communication network) like the Internet or a communication line (communication wire) like a telephone line. Further, the above program may be for realizing a part of the functions described above. Furthermore, it may be a so-called difference file (difference program) that can realize the functions described above in combination with a program already recorded in a computer system.
[0096] As described above, the embodiments for carrying out the present invention have been described using the embodiments. However, the present invention is not limited to such embodiments at all, and various modifications and substitutions can be made without departing from the gist of the present invention.
Explanation of Reference Numerals
[0097] 1... Information processing system, 2... Robot, 3... Environment sensor, 4... Image display device, 5... Information processing device, 9... Operation unit, 102... Photographing unit, 103... Sound collection unit, 110... Generation unit, 111... Display unit, 112... Voice output unit, 113... Driving unit, 114... Control unit, 115... Storage unit, 116... Communication unit, 120... Housing, 121... Boom, 122... Horizontal bar, 131... Movable part, 132... Movable part, 133... Movable part, 134... Movable part, 301... Photographing unit, 302... Sound collection unit, 303... Communication unit, 501... Execution unit, 502... Acquisition unit, 503... Audio-visual image processing unit, 504... Attention estimation unit, 505... Policy unit, 506... Output unit, 507... Identifier, 508... Reward network, 509... Learning unit, 510... Storage unit, 511... Communication unit
Claims
1. An acquisition unit that acquires the situation of the operator; A policy unit having information of a virtual co-operator who plays a game with rules together with the operator; An output unit capable of outputting an approach of the virtual co-operator to the operator; A discriminator trained using the activities of the operator and the activities of the virtual co-operator; A reward network that predicts the reward of the operator; Comprising: The policy unit learns a policy for operating the game of the virtual co-operator from the rewards provided by the reward network and the discriminator; The output unit outputs an output that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit; An information processing system.
2. The output unit performs an operation of the game that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit; The information processing system according to claim 1.
3. The information acquired by the acquisition unit includes at least one of an image including the face of the operator and an audio signal of the operator; An attention estimation unit that estimates the attention of the operator by performing at least one of estimating the line-of-sight direction of the operator using an image including the face of the operator, estimating the time during which the line-of-sight direction remains using an image including the face of the operator, estimating the expression of the operator using an image including the face of the operator, and estimating the degree of interest of the operator in the game using the audio signal of the operator; The information processing system according to claim 1 or claim 2, further comprising:
4. The discriminator: The activity that is a combination of the state of the game of the operator and the actions for the game; The activity that is a combination of the state of the game of the virtual co-operator and the actions for the game; Compare for each progress of the game and obtain the reward of the discriminator; The information processing system according to claim 1 or claim 2.
5. The output unit: When the result estimated by the attention estimation unit indicates that the attention of the operator is sustained or improved, outputs a statement praising the operator; When the result estimated by the attention estimation unit indicates that the attention of the operator is not sustained, outputs a statement prompting the operator to participate further in the game; When the time during which the operator does not participate in the game continues for a predetermined time, outputs a statement prompting the operator to participate in the game; The information processing system according to claim 3.
6. During learning, the state of the game and the actions of the operator are weighted to train the reward network. The information processing system according to claim 1 or claim 2.
7. Further comprising an environmental sensor including a photographing device and a sound collection unit, wherein the acquisition unit acquires at least one of the image photographed by the environmental sensor and the acoustic signal collected, and acquires the situation of the operator. The information processing system according to claim 1 or claim 2.
8. The virtual co-operator is a robot, the robot includes a voice output unit and a display unit, the output unit is at least one of the voice output unit and the display unit, the robot outputs an approach to the operator of the virtual co-operator by performing at least one of speaking from the voice output unit and displaying an image expressing an expression on the display unit to the operator. The information processing system according to claim 1 or claim 2.
9. An acquisition unit that acquires the situation of the operator; A policy unit having information on a virtual co-operator who plays a game with rules together with the operator; An output unit that generates an approach to the operator of the virtual co-operator and can output it to the virtual co-operator; A discriminator trained using the activities of the operator and the activities of the virtual co-operator; A reward network that predicts the reward of the operator; comprising the policy unit learns a policy for operating the game of the virtual co-operator from the rewards provided by the reward network and the discriminator, the output unit generates an output that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit, and outputs it to the virtual co-operator. An information processing device.
10. In an information processing device, the acquisition unit acquires the situation of the operator, when the policy unit has information on a virtual co-operator who plays a game with rules together with the operator, the output unit generates an approach to the operator of the virtual co-operator and outputs it to the virtual co-operator, the discriminator trained using the activities of the operator and the activities of the virtual co-operator compares the activities of the operator and the activities of the virtual co-operator to obtain a reward, the reward network predicts the reward of the operator. The policy unit learns a policy for operating the game of the virtual co-operator from the reward provided by the reward network and the identifier, The output unit generates an output that prompts the interest of the operator from the information based on the information acquired by the acquisition unit and the information of the policy unit, and outputs the output to the virtual co-operator. An information processing method.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2017225489A
Game operation learning program, game program, game play program, and game operation learning method
JP2020166528A
Behavioral control devices, behavioral control method, and program
JP2022029599A
Collaborative Adversarial Imitation Learning for Dynamic Treatment
JP2022542283A
Systems and methods to measure, predict and optimize brain function
WO2023239647A2