Learning device, learning method and program

The learning device enhances state estimation by using TGN models to process audiovisual data for robot interaction, addressing limitations of conventional methods by incorporating gaze, body language, and position information.

JP2025137268APending Publication Date: 2025-09-19HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024036377
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Conventional technologies struggle to accurately estimate a person's state by accounting for internal and external stimuli and rely solely on voice data, which is unreliable due to false facial expressions and limited applicability to brain wave processing.

Method used

A learning device that utilizes a detection unit to acquire audiovisual data, a modeling unit to create modeling data representing attention dynamics, and a model to control robot behavior based on Temporal Graph Network (TGN) models, incorporating gaze, body language, and position information to enhance state estimation.

Benefits of technology

The solution enables more practical estimation of a person's state by integrating multimodal sensory data, improving robustness and accuracy in real-world interactions with robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137268000001_ABST
    Figure 2025137268000001_ABST
Patent Text Reader

Abstract

To provide a learning device, learning method and program capable of estimating a person's state that can be implemented more practically than before.SOLUTION: A learning device includes: a detection unit for acquiring a plurality of sensing data, the information relating to states of a plurality of people based on audiovisual information in an environment where a plurality of people exist; a modeling unit that uses the acquired plurality of sensing data to create a plurality of modeling data, the information relating to attention of the plurality of people; and a model that inputs the plurality of modeling data and outputs information for controlling behavior of a robot that communicates with the plurality of people in the environment where the plurality of people exist.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a learning method, and a program. [Background technology]

[0002] For example, when a robot or the like communicates with a person, a technology has been proposed that estimates the state of the person and engages in a dialogue, etc. Methods for estimating the state of the person include, for example, physiological responses using electroencephalograms or the like, voice data (see, for example, Patent Document 1), and visual data such as facial appearance. Note that the state of a person tends to be influenced by internal stimuli (experiences, memories, etc.) and external stimuli (other people, objects, the environment, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-191521 Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional technologies have difficulty responding to the influence of internal and external stimuli, can only use voice data while a person is speaking, and are unreliable because people may falsely express their facial expressions, making them impractical for implementing processing of brain waves, etc.

[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a learning device, a learning method, and a program that can estimate a person's state in a more practical manner than conventional methods. [Means for solving the problem]

[0006] (1) In order to achieve the above-mentioned object, a learning device according to one aspect of the present invention is a learning device that includes: a detection unit that acquires, in an environment where a plurality of people are present, a plurality of sensing data that is information about the states of the plurality of people based on audiovisual information; a modeling unit that uses the acquired plurality of sensing data to create a plurality of modeling data that is information about the attention of the plurality of people; and a model that inputs the plurality of modeling data and outputs information to control the behavior of a robot that is in an environment where a plurality of people are present and communicates with the plurality of people.

[0007] (2) In a learning device according to one aspect of (1) above, the plurality of sensing data may be at least two of information indicating whether each of the plurality of people is speaking, body language information of the upper body of each of the plurality of people, position information of each of the plurality of people, and information indicating the gaze direction of each of the plurality of people.

[0008] (3) In a learning device according to one aspect of (1) or (2) above, the modeling unit may input information from the detection unit that can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the model, and output a node embedding representation to the model.

[0009] (4) In a learning device according to any one of the above (1) to (3), the modeling unit may be a model based on a TGN (Temporal Graph Network) model, and may input the modeling data output by the modeling unit for one time or multiple consecutive times, convert it into a TGN node-embedded representation, input information based on the converted node-embedded representation to a Gaussian model, and input the output from the Gaussian model to a decoder to create a behavioral command for the robot.

[0010] (5) In order to achieve the above object, a learning method according to one aspect of the present invention is a learning method for a learning device having a detection unit, a modeling unit, and a model, in which the detection unit acquires a plurality of sensing data, which is information regarding the state of a plurality of people based on audiovisual information, in an environment where a plurality of people are present, the modeling unit uses the acquired plurality of sensing data to create a plurality of modeling data, which is information regarding the attention of the plurality of people, and the model inputs the plurality of modeling data and outputs information to control the behavior of a robot that is in the environment where a plurality of people are present and communicates with the plurality of people.

[0011] (6) In order to achieve the above object, one aspect of the present invention provides a program that causes a computer having a learning device with a model to acquire, in an environment where a plurality of people are present, a plurality of sensing data that is information about the states of a plurality of people based on audiovisual information, use the acquired plurality of sensing data to create a plurality of modeling data that is information about the attention of the plurality of people, and input the plurality of modeling data to output information that controls the behavior of a robot that is in an environment where a plurality of people are present and communicates with the plurality of people. [Effects of the Invention]

[0012] According to the above (1) to (6), it is possible to estimate a person's state in a more practical manner than before. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a diagram for explaining an overview of a system configuration and an overview of processing in an embodiment. [Figure 2] 1 is a diagram illustrating an example of the configuration of a communication system according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of processing of a TGN model. [Figure 4] FIG. 1 shows the operational flow of the improved TGN used for learning memory-related modules. [Figure 5] FIG. 2 is a diagram for explaining each symbol in equation (3). [Figure 6] FIG. 10 is a diagram showing an example of processing in the case of attention (gazing) between C and H. [Figure 7] FIG. 10 is a diagram showing an example of a graph representing interactions between the gazes of multiple people. [Figure 8] FIG. 10 is a diagram showing an example of converting each piece of estimated information into text. [Figure 9] FIG. 10 is a diagram for explaining generation of a context-embedded representation. [Figure 10] 10A and 10B are diagrams illustrating an example of data used during learning and a learning method, and an example of data used during use and an estimation method in an embodiment. [Figure 11] 10 is a flowchart of a processing procedure when learning a model according to an embodiment. [Figure 12] 10 is a flowchart of a processing procedure when using a model according to an embodiment. [Figure 13] FIG. 10 is a diagram showing examples of evaluation results of the first method and the second method. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).

[0015] [overview] FIG. 1 is a diagram illustrating an overview of the system configuration and processing in this embodiment. As shown in FIG. 1, there are multiple people in an environment, engaged in conversation, etc. The environment also includes a robot 3 that can communicate with the people. Furthermore, a sensor 2 that senses the state of the people is installed in the environment. A control device 4 controls communication between the robot 3 and the people based on detection data detected by the sensor 2. The control device 4 trains a model to output behavioral commands to the robot 3 using the detection data. Note that there are multiple types of sensed data, such as facial images, audio signals, and human movements. Furthermore, there are multiple types of modeling data generated using the detection data. In this embodiment, the control device 4 estimates the state of multiple people using non-verbal cues and multimodal temporal audiovisual observation information of the environmental context.

[0016] [Example of communication system configuration] Next, an example of the configuration of the communication system 1 will be described. 2 is a diagram showing an example of the configuration of a communication system according to this embodiment. As shown in FIG. 2, the communication system 1 includes, for example, a sensor 2 (detection unit), a robot 3, and a control device 4.

[0017] The sensor 2 includes, for example, a sound collector 21 (detection unit), an imaging device 22 (detection unit), and a communication unit 23. The robot 3 includes, for example, a control unit 31 and a communication unit 32.

[0018] The control device 4 includes, for example, a learning device 5, a first generation unit 6, and a second generation unit 7. The learning device 5 includes, for example, a detection unit 51 (detection unit), a modeling unit 52, a state model 53, and a sociological model 54 (social dynamics model) (also simply referred to as a "model"). The detection unit 51 includes, for example, a first feature information detection unit 511, a second feature information detection unit 512, a third feature information detection unit 513, and a fourth feature information detection unit 514. The modeling section 52 includes, for example, a first modeling device 521, a second modeling device 522, a third modeling device 523, and a fourth modeling device 524.

[0019] (sensor) The sensor 2 is an environmental sensor or the like that detects the state (sensing data) of each person. Note that the sensor 2 may include sensors other than the sound pickup device 21 and the image pickup device 22.

[0020] The sound collector 21 is, for example, a microphone array made up of multiple microphones. The sound collector 21 collects human voice signals (sensing data) and may output the collected voice signals in analog form, or may convert the analog signals into digital signals and output them. When outputting digital signals, the multiple microphones are synchronized.

[0021] The image capturing device 22 is, for example, an RGB (red-green-blue)-D (depth) camera that can also acquire depth information. The image capturing device 22 may be an RGB camera and a distance sensor. The image capturing device 22 captures the state of a person (sensing data).

[0022] The communication unit 23 outputs the sound signal (sensing data) collected by the sound collector 21 and the image data (sensing data) captured by the image capturing device 22 to the control device 4. The sensor 2 and the control device 4 are connected via a wired or wireless network or the like.

[0023] (robot) The robot 3 may be, for example, a humanoid robot, a robot capable of walking on two legs, or a robot with a head and a body. The robot 3 may be equipped with, for example, an image display unit, a sound pickup unit, a photographing device, an audio output device, a power supply unit, etc. The robot 3 may also be equipped with a sensor 2. The robot 3 may also be equipped with a control device 4. The robot 3 may receive operation instructions, for example, via remote control. The robot 3 may also be any robot, including, for example, a desktop robot, a mobile robot, an industrial robot arm, or an intelligent agent such as an avatar in any embodiment in a virtual space. The robot 3 may also be a human model such as an avatar, or a robot with multiple degrees of freedom, including but not limited to key points on the hand joints, head, and face.

[0024] The control unit 31 controls the operation of the robot 3 in accordance with the action command output by the control device 4.

[0025] The communication unit 32 acquires behavioral commands output by the control device 4. The communication unit 32 outputs information acquired by the robot 3 (for example, detection data from sensors provided at the joints, captured image data, etc.) to the control device 4. The robot 3 and the control device 4 are connected via a wired or wireless network or the like.

[0026] (Control device) The learning device 5 acquires the audio signal output by the sensor 2, the photographic data of the time series (t-2, t-1, t), etc., and detects a plurality of pieces of state information of the person from the acquired data. The learning device 5 generates a plurality of pieces of modeling data using the plurality of pieces of state information of the detected person (detection data). The learning device 5 trains each of the modelers (first modeler 521 to fourth modeler 524) included in the modeling unit 52 using the generated plurality of modeling data.

[0027] The detection unit 51 acquires audio signals, photographic data, etc. output by the sensor 2, and outputs low-level sensing data from the acquired data, including but not limited to face and body key points, head position, etc., to be used in the modeling unit 52 and sociological model 54. The sensing data is information about the state of multiple people based on audiovisual information.

[0028] The first feature information detection unit 511 inputs the image data to a convolutional neural network (CNN) that has been trained using, for example, training data, and detects or extracts facial feature information, attention feature information based on facial direction, gaze, etc. as first feature information. The first feature information detection unit 511 outputs the detected first feature information for consecutive times (timesteps) t-2, t-1, and t to the first modeler 521, the second modeler 522, the third modeler 523, the fourth modeler 524, and the sociological model 54. Note that well-known image processing may be performed on the first image data to detect or extract facial features, including but not limited to facial keypoints, a cut-out human face and head, CNN feature information corresponding to the face or head, and notable features, and use these as the first feature information.

[0029] The second feature information detection unit 512 performs, for example, well-known image processing on the image data to detect or extract, as second feature information, main features of the human body, such as but not limited to key points of the human body, segmented human bodies and their corresponding CNN features, human body parts, and the skeleton of a human pose. The second feature information detection unit 512 detects or extracts, as second feature information, information on the human body's physical attention (for example, body orientation, posture, hand and arm movements during conversation, etc.). The second feature information detection unit 512 outputs the detected second feature information for consecutive times t-2, t-1, and t to the first modeler 521, the second modeler 522, the third modeler 523, the fourth modeler 524, and the sociological model 54.

[0030] The third feature information detection unit 513 performs well-known speech recognition processing (e.g., sound source separation, sound source direction estimation, noise suppression, etc.) on the speech signal to detect or extract acoustic feature information, such as, but not limited to, Mel-frequency cepstrum (MFCC) representing intonation, pitch, etc., prosodic features, as third feature information. The third feature information detection unit 513 outputs the detected third feature information for a period similar to consecutive times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54. In this embodiment, paralanguage refers to information conveyed from a speaker to a listener that is outside the context of language, excluding body language such as gestures, and is, for example, the strength and weakness of the voice when speaking, pitch, intonation, etc.

[0031] The fourth feature information detection unit 514 performs, for example, well-known image processing on the image data to detect or extract true feature information of the environment, feature information of objects in the environment, etc. as fourth feature information. The fourth feature information detection unit 514 outputs the detected fourth feature information for times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54.

[0032] The modeling unit 52 generates multiple modeling data to be used for training the sociological model 54 using multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time step t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information. The modeling data is information related to the attention of multiple people. That is, the modeling unit 52 uses low-level information from the detection unit 51 to generate high-level modeling data, including but not limited to gaze behavior, body language, and human-object interactions. The output of the modeling unit 52 is, for example, data that can be used to establish edges and nodes in a graph representation and their features. Each modeler (the first modeler 521 to the fourth modeler 524) is, for example, a gaze targeting model, a joint attention model, a body language recognition model, etc. The nodes are not limited to people, but may also be objects (people, objects, etc.) that people pay attention to.

[0033] The first modeling device 521 generates modeling data of the gaze trajectory of each person using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the degree of attention to a common target by most people in the scene. Note that the first modeling device 521 can also output the gaze target of everyone in the scene.

[0034] The second modeling device 522 generates modeling data of body language or behavior for each person using multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for a single time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the physical behavior (e.g., walking, sitting, etc.) of everyone in the scene.

[0035] The third modeler 523 converts the paralanguage of each person into high-level auditory behavior using one or more pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. This generates modeling data, which is output to the sociological model 54 as environmental information and instructs auditory behavior such as sound source direction estimation, sound source position estimation, and sound source estimation.

[0036] The fourth modeler 524 generates modeling data of the environment and context using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. Then, the fourth modeler 524 outputs the generated modeling data to the sociological model 54 as environmental information not limited to object recognition and how humans interact with these objects, i.e., human-object interactions.

[0037] The state model 53 is a model for inferring human internal states, such as emotions, intentions, or internal states, such as positive and negative, positive or negative, including but not limited to valence and arousal. Such internal states can also be represented by physiological responses, such as information about electroencephalograms, such as electroencephalogram frequency bands, and information about electrocardiograms, including but not limited to heart rate.

[0038] The sociological model 54 is, for example, a model based on social interaction dynamics, such as a model based on the TGN (Temporal Graph Network) model (see, for example, Reference 1). The sociological model 54 is a trained model that receives environmental information from the sound pickup device 21, the camera device 22, the detection unit 51, and the modeling unit 52. This information can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the TGN model. These inputs can also be encoded as features to better represent multimodal temporal interaction dynamics. Encoding can be performed using different methods, such as a text encoder or a visual encoder. The output of the sociological model 54 is a node embedding representation and feature, or embedding representation or feature, of the TGN model. During training, the sociological model 54 is trained using data stored in a database, for example. For example, during training, the sociological model 54 uses link prediction to learn social interaction dynamics. Furthermore, the training may be performed using privileged learning (PL), etc. The sociological model 54 may be an LCBM model for representing the dynamics of social interactions. In this case, the sociological model 54 fine-tunes a large language model (LLM) (see, for example, Reference 2) using context extracted from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the text or visual encoder of the modeling unit 52. The fine-tuned LLM model is a large context and behavioral model (LCBM). To generate behaviors such as movements and utterances, the output from the LCBM is passed to, for example, a Gaussian mixture model and decoded to generate behavior commands for controlling the different degrees of freedom of the robot 3.

[0039] Reference 1; Emanuele Rossi, Ben Chamberlain, et al., “Temporal Graph Networks for Deep Learning on Dynamic Graphs”, arXiv:2006.10637 [cs.LG], 2020 Reference 2;Ashmit Khandelwal, Aditya Agrawal, et al., “Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior”, arXiv:2309.00359 [cs.CL], 2023

[0040] The first generation unit 6 generates nonverbal cues (NOCs) based on the output of the sociological model 54. The nonverbal cues (NOCs) will be described later. The first generation unit 6 inputs the output from the large-scale language model of the sociological model 54 to a Gaussian model (e.g., a Gaussian mixture model) and the linguistic information and embeddings of a TGN model or LCBM representing social interaction dynamics. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder to create nonverbal cue behavioral commands for the robot 3.

[0041] The second generation unit 7 generates speech information (text, audio signals) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the generated speech information to a Gaussian model (e.g., a Gaussian mixture model), inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 from the speech information.

[0042] In the example shown in FIG. 2, the detection unit 51 is provided with four feature information detection units (511 to 514), but this is not limiting. The number of feature information detection units may be two or more, and may be five or more. Furthermore, the modeling unit 52 is provided with four modeling devices (first modeling device 521 to fourth modeling device 524), but this is not limiting. The number of modeling devices may be two or more, and may be five or more. In the example shown in FIG. 2, the first generation unit 6 and the second generation unit 7 are provided, but the present invention is not limited to this and it is sufficient that at least one of them is provided.

[0043] With this configuration, in this embodiment, link prediction is used to train a TGN model (sociological model 54) to learn social interaction dynamics. In addition, in this embodiment, node embeddings and features representing social interaction dynamics are extracted to infer human states. These embeddings / features can also be used for other tasks, such as next speaker prediction and group state estimation.

[0044] [Nonverbal Cues (NOC)] Next, we will explain non-verbal cues (NOCs). In order to infer human states in real environments using audiovisual data, it is necessary to understand and represent the dynamics of social interactions using multimodal temporal observations, not limited to non-verbal behaviors and vocal activities, as human states tend to depend on environmental context. Here, non-verbal behavior refers to gestures (body movements), posture, movements (nodding and general behavior), facial expressions, tone of voice (volume, speed, pauses), etc. When a robot communicates with a person, for example, rather than simply outputting an automatically generated voice signal, a communication robot has been proposed that responds to the voice signal by adjusting the tone and strength of the voice, and even by displaying images of the eyes and mouth on a display unit to speak with facial expressions (for example, JP 2022-6610 A). Furthermore, it has been proposed to improve communication with the user by moving the robot's head or other parts to make the user feel as if the robot is drooping or happy, or by having a robot with arms and hands make gestures.

[0045] For this reason, in this embodiment, verbal information and non-verbal information (non-verbal cues, non-verbal behavior) are generated as behavioral commands for the robot 3. Note that in this embodiment, the non-verbal cues include gaze interaction, body language, vocal activities such as paralanguage, speech characteristics, facial expressions, etc.

[0046] [Internal and external stimuli] Next, we explain internal and external stimuli. Internal stimuli are represented using temporal information (i.e., TGN memory). External stimuli are represented by multimodal inputs from the sound pickup device 21, the camera device 22, the detection unit 51, and the modeling unit 52 to the sociological model 54, enabling the representation of social interaction dynamics using the TGN. The sociological model 54 then converts these inputs into TGN node embeddings as outputs. These node embeddings can be used for downstream supervised learning tasks, such as, but not limited to, human state inference, next speaker prediction, and action generation, as shown in Figure 2.

[0047] (TGN model) Here, an outline of the TGN model will be explained (see, for example, Reference 1). First, an example of calculations that the TGN performs on a batch of time-stamped interactions will be explained. Figure 3 shows an example of processing in the TGN model. The first module g11 generates an embedding emb using the time graph and node memory (step S1). The second module g12 uses the embedding emb to predict batch interactions (step S2). The third module g13 uses the embedding emb to calculate the probability of each edge and calculate the loss (step S3). The fourth module g14 to the sixth module g16 are storage modules. The interaction of the processes in steps S1 to S3 is used to update each storage module. However, this configuration and processing is a simplified operation flow, and since the fourth module g14 to the sixth module g16 do not receive gradients, it will prevent learning of all modules.

[0048] Next, we will explain a TGN model (hereinafter referred to as an improved TGN model) that solves the above problems (see, for example, Reference 1). Figure 4 shows the operation flow of the improved TGN used for learning memory-related modules. The first module (Raw Message Store) g21 stores the messages of interactions previously processed by the model, i.e., the raw information {rm1(t1), rm2(t2), rm3(t3), rm4(t4)} needed to calculate the inputs to the message function (called raw messages). This allows the improved model to postpone the update of the storage module brought about by interactions to a later batch.

[0049] The first module g21, the second module g22, and the third module g23 update their storage modules using messages calculated from raw messages stored in the previous batch (steps S21 to S23). The fourth module g24 calculates the embedding emb using the updated information (arrow g28) of the storage module (step S24). The fifth module g25 uses the embedding emb to calculate the probability of each edge and calculate the loss (step S25). Note that this probability of each edge represents the probability that it is the target of interest.

[0050] The output msg of the first module g21 is expressed by the following equation (1).

[0051]

number

[0052] In equation (1), s i(t-1) is the memory data of past i, and s j(t-1) is the memory data of past j. rm i,j(t) is the characteristic of the last interaction between i and j, and t enc is the linear transformation of time between the current time t and the update of past memory. Furthermore, the output s of the third module g23 i(t) The relationship between the output mem of the second module g22 and the output mem of the second module g22 is given by the following equation (2).

[0053]

number

[0054] In equation (2), mem is the memory update model (GRU), and m (i,j) is the message generated at the time of the last update, and s i(t-1) is the previous state in memory. The fourth module g24 generates a transformation that represents the state of the node for the entire graph. Here, the state of the node is updated (node ​​embedding) h i ' is expressed by the following equation (3).

[0055]

number

[0056] In equation (3), αij is the attention mechanism as shown in Figure 5, symbol g43, W is the learnable weight matrix as shown in Figure 5, symbol g42, and h j is the target node (neighborhood) as shown by symbol g41 in FIG. 5. Also, σ is the activation function. FIG. 5 is a diagram for explaining each symbol in equation (3). Also, Σ(·) in equation (3) is h1'*α as shown by symbol g43 in FIG. 11、0 h2'*α 12 and h3'*α 13 The output of each node is added together. Note that h1' is an embedding node for 1, e.g., [0.234, 0.412, ..., 0.98] (1,100) is. Furthermore, α in Eq. (3) ij is used as follows:

[0057]

number

[0058] α c,I,j is the attention (gaze) that node i should pay to node j, c is the multi-head attention (set to 1 in Eq. (3) for simplicity), and l is the layer (starting from 0). c,i (l) is the feature information of the source node as shown in Figure 6, and k c,j (l) is the characteristic information of the destination node as shown in Figure 6, and e c,i,j is edge feature information between the source node and the destination node as shown in Fig. 6. Fig. 6 is a diagram showing a processing example in the case of attention (gazing) between C and H. q in the numerator on the right side of equation (4) c,i (l) corresponds to symbol g51 in Figure 6, and k c,j (l) +e c,i,j corresponds to symbol g52 in Figure 6. The numerator of equation (4) is the attention of only one neighbor, as symbol g53 in Figure 6, and the denominator of equation (4) is the operation of finding the same thing for all neighbors, and the meaning of equation (4) is the normalization of attention based on others.

[0059] This allows the calculation of memory-related modules to directly affect the loss (steps S25 and S26) and receive the gradient. Finally, the sixth module stores the raw messages of this batch of interactions in the first module g21 for use in future batches (step S26).

[0060] In an embodiment, such an improved TGN model is used to detect, for example, gaze interactions between multiple people. For example, suppose three people (A, B, and C) are having a conversation. Assume that A is speaking at time t1, and C is speaking at time t2. The gaze interactions between multiple people in this situation can be represented by a network such as that shown in FIG. 7.

[0061] Figure 7 shows an example of a graph representing the interaction of the gazes of multiple people. In Figure 7, each node represents a different person. The tip of the edge arrow indicates the object that each person is looking at at each time. In the example of Figure 7, for example, at time t1, it is estimated from the graph representation that A is speaking, and at time t2, it is estimated from the graph representation that B and C are speaking. In the graph representation, multiple interactions can be estimated using environmental information such as the relationship between A, B, and C (for example, A is a teacher and B and C are students).

[0062] In this embodiment, such environmental information is converted into text as shown in Fig. 8. Fig. 8 is an example of converting each piece of estimated information into text. The information indicated by reference symbol g101 is a diagram showing an example of information generated for converting each piece of estimated information into text based on, for example, speaking information (talking information), the relative seating position of multiple people, and the role of each person among the multiple people. The information indicated by reference symbol g102 is an example of a summary of the information indicated by reference symbol g101. The text indicated by reference symbol g103 is an example of the information indicated by reference symbol g102 converted into text according to predetermined rules.

[0063] Note that the configurations and processes described using Figures 4 and 5 are merely examples and are not limited to these. For example, data input to the TGN model may be encoded using a Privileged Encoder, Wild Encoder, or the like, and then embedded in a node. Alternatively, audio signals may be processed by an acoustic encoder, image data may be processed by an image encoder, and the results of processing by the acoustic encoder and image decoder may be encoded using a Wild Encoder or the like, embedded in a node, and then input to the TGN model. Furthermore, encoding may be performed using different methods, such as a text encoder or a visual encoder.

[0064] Next, in this embodiment, as shown in Fig. 9, text g111 is generated by verbalizing information about the environment generated by compiling information for converting each piece of information into text, and the resulting text is input to a language model g112 to output a fixed-size context-embedded representation g113. Fig. 9 is a diagram for explaining the generation of a context-embedded representation. Note that the language model is based on, for example, BERT (Bidirectional Encoder Representations from Transformers), but is not limited to this and may be another language model.

[0065] Next, in this embodiment, the fixed-size context embedded representation thus generated is used to train, for example, an improved TGN model, which is a graph representation, as shown in Fig. 10. Fig. 10 is a diagram showing an example of data used during training and a training method, and an example of data used during use and an estimation method, in this embodiment.

[0066] The box g120 in the upper row shows the data used for model training and the graph representation model (TGN model) to be trained. The table g121 is mutual data of the context embedded representation for each time generated as described with reference to FIG. 9, and the person's identification information (Id (Gazing Subject)) and gaze direction identification information (Gaze Id (Gazed Subject)) for each time, and is data used when training a model. The identification information is assigned by, for example, the modeling unit 52. Note that the identification information may also be assigned by the detection unit 51. The symbol g122 is a model, for example, an improved TGN model.

[0067] The box g140 in the lower row shows the data, processing, etc. at the time of use. The table g141 is the mutual data of the person's identification information (Id (Gazing Subject)) and gaze direction identification information (Gaze Id (Gazed Subject)) for each time, and is the data input to the trained model. Reference symbol g142 is a diagram showing an example of a trained graph representation. The trained model receives time-varying mutual data such as the table in reference symbol g141 and outputs information representing the probability of each edge, i.e., the target of interest. The table g143 shows the probability of each edge output by the model and state information in which the probability of each edge is set to 1 if it is equal to or greater than a threshold (for example, 0.5) and set to 0 if it is less than the threshold.

[0068] That is, when in use, information indicating the focus of each person at each time is input to the representation of the interaction dynamics, which is such a trained model, and the focus based on the gaze is estimated, for example, based on the probability of each edge. Note that the probability of each edge is calculated by the fifth module g25 in the improved TGN model. In this embodiment, for example, the first modeling device 521 has the above-mentioned improved TGN model.

[0069] As described above, this embodiment has the following configuration and processing. I. We have designed the system to infer human states (emotions, values, arousal values, etc.) using multimodal temporal and non-verbal cues. II. Internal stimuli are represented using time information. III. External stimuli are represented by multimodal inputs from the TGN model. IV. Non-verbal cues included gaze interactions, body language, vocal activities such as paralanguage, speech characteristics, and facial expressions.

[0070] As a result, the configuration and method of this embodiment are practical for implementation in real environments because the sociological model 54 relies on audiovisual data. Furthermore, because the configuration and method of this embodiment rely on multimodal input, the robustness of the sociological model 54 can be improved compared to conventional methods.

[0071] In this embodiment, the output of the sociological model 54 is used to generate verbal and non-verbal behavioral commands by the generation units (first generation unit 6, second generation unit 7), which then control the robot 3. As a result, according to this embodiment, there is no need to wear any special device, and it is possible to understand the state (interaction) of multiple people using the visual data detected by the sensor 2. Furthermore, according to this embodiment, it is possible to make the robot 3 behave (speak, gesture, show facial expressions, etc.) using the sociological model 54 and state model 53 for inferring a person's state and non-verbal behavior, thereby realizing empathetic interaction with people in a real environment.

[0072] With this configuration and processing, the configuration and method of this embodiment are practical for implementation in a real environment because the model relies on audiovisual data. Furthermore, the configuration and method of this embodiment rely on multimodal input, which improves the robustness of the model compared to conventional methods.

[0073] [Processing Procedure] Next, an example of a processing procedure performed by the control device 4 will be described. First, an example of the procedure for training the model (modeling unit 52, sociological model 54) will be described. Fig. 11 is a flowchart of the processing procedure for training the model according to this embodiment.

[0074] (Step S101) The detection unit 51 acquires the detection data detected by the sensor 2.

[0075] (Step S102) The detection unit 51 acquires relationship data (for example, information about relationships between people).

[0076] (Step S103) The detection unit 51 detects the feature information (first feature information to fourth feature information) from the detection data and the relational data.

[0077] (Step S104) The modeling unit 52 converts the information about the environment included in the feature information (information based on the relational data) into text.

[0078] (Step S105) The modeling unit 52 receives inputs of environmental information from the sound pickup device 21, the image capture device 22, the detection unit 51, and the modeling unit 52, which can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the TGN model. These inputs can also be encoded as features to better represent multimodal temporal interaction dynamics. The output is node embeddings of the TGN model. During training, the sociological model 54 is trained for link prediction using data stored in a database, for example. This is a type of supervised learning to represent social interaction dynamics using the output of the sociological model 54, i.e., node embeddings.

[0079] (Step S106) The modeling unit 52 uses the generated node embedding representation to train another supervised learning model, which is, for example, a human state estimation result, a next speaker estimation result, etc. Note that the learning model is a model provided in each of the modeling units 52 itself.

[0080] In this embodiment, the above steps S104 to S106 are referred to as a modeling data generation process.

[0081] Next, an example of a procedure for making an estimation using the trained model (the modeling unit 52 and the sociological model 54) will be described. Fig. 12 is a flowchart of the processing procedure when using the model according to this embodiment.

[0082] (Step S201) The detection unit 51 acquires the detection data detected by the sensor 2.

[0083] (Step S202) The detection unit 51 detects feature information (first feature information to fourth feature information) from the detection data.

[0084] (Step S203) The modeling unit 52 assigns identification information to a plurality of people.

[0085] (Step S204) The modeling unit 52 converts environmental information (information based on relational data) included in the feature information into text. The modeling unit 52 then receives input of environmental information from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the modeling unit 52, which can be used to construct a graph consisting of nodes and edges. The modeling unit 52 then outputs the node embedding data of the TGN model to the sociological model 54.

[0086] (Step S205) The sociological model 54 receives information output from the first modeling device 521 to the fourth modeling device 524 for one time or a plurality of consecutive times, and converts the information into node embeddings of the TGN.

[0087] (Step S206) The first generation unit 6 generates a non-verbal cue (NOC) using the output of the sociological model 54. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the non-verbal cue. Furthermore, the second generation unit 7 generates speech information (text, audio signal) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the speech information.

[0088] (Step S207) The first generating unit 6 and the second generating unit 7 each output the generated action instructions to the robot 3 to control the action of the robot 3.

[0089] 11 and 12 are merely examples, and are not limiting. The control device 4 may, for example, perform several processes in parallel.

[0090] In the above-described embodiment, an example has been described in which each modeling device of the modeling unit 52 is a TGN model, but this is not limiting. The modeling devices (first modeling device 521 to fourth modeling device 524) may be other models, and the models of each modeling device may be different. Furthermore, although the sociological model 54 is an example of a TGN model, the sociological model 54 may be another model such as an LCBM.

[0091] [Evaluation results] Next, an example of the evaluation results will be described. Figure 13 shows example evaluation results for the first and second methods. Note that the first method is a process that uses the output of BERT. The second method is a process in which, for example, a text encoder is replaced with a visual encoder. The horizontal axis D1 to D8 in Figure 13 indicates the attributes of the subjects, and the vertical axis indicates the score (100% is the best score). Note that the score indicates whether the result of estimating a person's state (attention position) is correct or not.

[0092] As shown in FIG. 13, the evaluation results were good according to the configuration and processing method of this embodiment. Furthermore, the second method resulted in a greater overall improvement in performance than the first method, with the greatest improvement occurring in the unsupervised sessions and the highest performance occurring in the trained sessions.

[0093] In addition, a program for implementing some or all of the functions of the control device 4 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the control device 4. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. Alternatively, some or all of these components may be realized by LSI (Large Scale Integration) hardware (including circuitry) such as an ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), or SOC (System On Chip), or may be realized by a combination of software and hardware.

[0094] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0095] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0096] 1...communication system, 2...sensor, 3...robot, 4...control device, 5...learning device, 6...first generation unit, 7...second generation unit, 21...sound pickup device, 22...image capture device, 23...communication unit, 31...control unit, 32...communication unit, 51...detection unit, 52...modeling unit, 53...state model, 54...sociological model, 511...first feature information detection unit, 512...second feature information detection unit, 513...third feature information detection unit, 514...fourth feature information detection unit, 521...first modeling device, 522...second modeling device, 523...third modeling device, 524...fourth modeling device, g21...first module, g22...second module, g23...third module, g24...fourth module, g25...fifth module, g26...sixth module

Claims

1. a detection unit that acquires a plurality of sensing data pieces that are information about the states of a plurality of people based on audiovisual information in an environment where a plurality of people are present; a modeling unit that uses the acquired plurality of sensing data to create a plurality of modeling data that are information relating to the plurality of people's attentions; a model that receives a plurality of the modeling data as input and outputs information for controlling the behavior of a robot that is in an environment where the plurality of people are present and communicates with the plurality of people; A learning device equipped with

2. The plurality of sensing data At least two of information indicating whether each of the plurality of people is speaking, body language information of the upper body of each of the plurality of people, position information of each of the plurality of people, and information indicating a line of sight of each of the plurality of people, The learning device according to claim 1 .

3. The modeling unit inputting information from the detector that can be used to construct a graph of nodes and edges to represent the social interaction dynamics of the model, and outputting a node embedding representation to the model; The learning device according to claim 1 or 2.

4. The model is This model is based on the TGN (Temporal Graph Network) model, inputting the modeling data output by the modeling unit for one time or a plurality of consecutive times, and converting the modeling data into a node-embedded representation of a TGN; inputting information based on the converted node embedding representation into a Gaussian model, and inputting the output from the Gaussian model into a decoder to generate a behavioral command for the robot; The learning device according to claim 1 or 2.

5. A learning method for a learning device having a detection unit, a modeling unit, and a model, comprising: a detection unit that acquires a plurality of sensing data pieces that are information about the states of a plurality of people based on audiovisual information in an environment where a plurality of people are present; a modeling unit, using the acquired plurality of pieces of sensing data, to create a plurality of pieces of modeling data which are information relating to the plurality of pieces of attention of the person; a model inputting a plurality of the modeling data and outputting information for controlling the behavior of a robot that is in an environment where the plurality of people are present and communicates with the plurality of people; How to learn.

6. A learning computer with a model In an environment where a plurality of people are present, a plurality of sensing data are acquired which are information relating to the states of the plurality of people based on audiovisual information; creating a plurality of modeling data, which is information relating to the attention of a plurality of people, using the plurality of acquired sensing data; inputting a plurality of the modeling data and outputting information for controlling the behavior of a robot that is in an environment where the plurality of people are present and communicates with the plurality of people; program.

Citation Information

Patent Citations

  • Communication robot, input sound determination method and input sound determination program

    JP2019191521A