Learning apparatus, learning method, and program

WO2025187803A8PCT designated stage Publication Date: 2025-10-02HONDA MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/008333
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2025-03-06
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Conventional technologies struggle to accurately estimate a person's state by considering both internal and external stimuli, are limited to voice data during speech, and are unreliable due to false facial expressions, making them impractical for processing brain waves.

Method used

A learning device and method that utilizes a detection unit to acquire audiovisual information, a modeling unit to create modeling data, and a model to control robot behavior, incorporating a TGN model for social interaction dynamics, and generating non-verbal cues and speech information based on multimodal inputs.

Benefits of technology

Enables more practical estimation of a person's state by considering internal and external stimuli, improving robustness and reliability through multimodal audiovisual data processing, allowing robots to engage in empathic interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025008333_02102025_PF_FP_ABST
    Figure JP2025008333_02102025_PF_FP_ABST
Patent Text Reader

Abstract

This learning apparatus comprises: a detection unit for, in an environment in which a plurality of persons are present, acquiring a plurality of pieces of sensing data, which are information pertaining to the state of the plurality of persons as based on audiovisual information; a modeling unit for using the acquired plurality of pieces of sensing data to create a plurality of pieces of modeling data, which are information pertaining to the attention of the plurality of persons; and a model for inputting the plurality of pieces of modeling data and outputting information for controlling the behavior of a robot that communicates with the plurality of persons in the environment in which the plurality of persons are present.
Need to check novelty before this filing date? Find Prior Art

Description

Learning devices, learning methods, and programs

[0001] This application claims priority to Japanese Patent Application No. 2024-036377, filed March 8, 2024, the contents of which are incorporated herein by reference.

[0002] For example, when a robot or the like communicates with a person, a technology has been proposed that estimates the state of the person and engages in a dialogue, etc. Methods for estimating the state of the person include, for example, physiological responses using electroencephalograms or the like, voice data (see, for example, Patent Document 1), and visual data such as facial appearance. Note that the state of a person tends to be influenced by internal stimuli (experiences, memories, etc.) and external stimuli (other people, objects, the environment, etc.).

[0003] Japanese Patent Application Laid-Open No. 2019-191521

[0004] However, conventional technologies have difficulty responding to the influence of internal and external stimuli, can only use voice data while a person is speaking, and are unreliable because people may falsely express their facial expressions, making them impractical for implementing processing of brain waves, etc.

[0005] The aspects of the present invention have been made in consideration of the above-mentioned problems, and aim to provide a learning device, a learning method, and a program that can estimate a person's state in a more practical manner than conventional methods.

[0006] In order to solve the above problems, the present invention employs the following aspects: (1) A learning device according to one aspect of the present invention is a learning device including: a detection unit that acquires, in an environment where a plurality of people are present, a plurality of sensing data that is information relating to the states of a plurality of people based on audiovisual information, a modeling unit that uses the acquired plurality of sensing data to create a plurality of modeling data that is information relating to the attention of the plurality of people, and a model that inputs the plurality of modeling data and outputs information to control the behavior of a robot that is in the environment where a plurality of people are present and communicates with the plurality of people.

[0007] (2) In the above aspect (1), the plurality of sensing data may be at least two of information indicating whether each of the plurality of people is speaking, body language information of the upper body of each of the plurality of people, position information of each of the plurality of people, and information indicating the line of sight of each of the plurality of people.

[0008] (3) In the above aspect (1) or (2), the modeling unit may input information from the detection unit that can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the model, and output a node embedding representation to the model.

[0009] (4) In any one of the above aspects (1) to (3), the modeling unit may be a model based on a TGN (Temporal Graph Network) model, and may input the modeling data output by the modeling unit for one time or multiple consecutive times, convert it into a TGN node-embedded representation, input information based on the converted node-embedded representation to a Gaussian model, and input the output from the Gaussian model to a decoder to create a behavioral command for the robot.

[0010] (5) A learning method according to one aspect of the present invention is a learning method for a learning device having a detection unit, a modeling unit, and a model, in which the detection unit acquires a plurality of sensing data in an environment where a plurality of people are present, the sensing data being information regarding the state of the plurality of people based on audiovisual information, the modeling unit uses the acquired sensing data to create a plurality of modeling data being information regarding the attention of the plurality of people, and the model inputs the modeling data and outputs information to control the behavior of a robot that is in the environment where a plurality of people are present and communicates with the plurality of people.

[0011] (6) A program according to one aspect of the present invention is a program that causes a computer having a learning device with a model to acquire, in an environment where a plurality of people are present, a plurality of sensing data that is information regarding the state of a plurality of people based on audiovisual information, use the acquired plurality of sensing data to create a plurality of modeling data that is information regarding the attention of the plurality of people, input the plurality of modeling data, and output information that controls the behavior of a robot that is in an environment where a plurality of people are present and communicates with the plurality of people.

[0012] According to the aspects of the present invention, it is possible to estimate a person's state in a manner that is more practical to implement than conventional methods.

[0013] 1 is a diagram for explaining an overview of the system configuration and an overview of the processing in an embodiment. FIG. 2 is a diagram showing an example of the configuration of a communication system according to an embodiment. FIG. 3 is a diagram showing an example of processing of a TGN model. FIG. 4 is a diagram showing the operation flow of an improved TGN used for learning a memory-related module. FIG. 5 is a diagram for explaining each symbol in equation (3). FIG. 6 is a diagram showing an example of processing in the case of attention (gazing) between C and H. FIG. 7 is a diagram showing an example of a graphical representation of the interaction of the gazes of multiple people. FIG. 8 is a diagram showing an example of converting each estimated piece of information into text. FIG. 9 is a diagram for explaining the generation of a context-embedded representation. FIG. 10 is a diagram showing an example of data used during learning and a learning method in an embodiment, and an example of data used during use and an estimation method. FIG. 11 is a flowchart of the processing procedure when learning a model according to an embodiment. FIG. 12 is a flowchart of the processing procedure when using a model according to an embodiment. FIG. 13 is a diagram showing example evaluation results of a first method and a second method.

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component has been appropriately changed so that each component can be recognized. In all drawings used to explain the embodiments, components having the same function are designated by the same reference numerals, and repeated explanations will be omitted. In addition, "based on XX" in this application means "based on at least XX" and includes cases where the component is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where the component is based on XX after calculation or processing. "XX" is any element (for example, any information).

[0015] [Overview] FIG. 1 is a diagram illustrating an overview of the system configuration and processing in this embodiment. As shown in FIG. 1, multiple people are in an environment, engaging in conversation, etc. The environment also includes a robot 3 that can communicate with the people. Furthermore, a sensor 2 that senses the state of the people is installed in the environment. A control device 4 controls communication between the robot 3 and the people based on detection data detected by the sensor 2. The control device 4 trains a model to output behavioral commands to the robot 3 using the detection data. The sensed data includes multiple types of data, such as facial images, audio signals, and human movements. Furthermore, multiple types of modeling data are generated using the detection data. In this embodiment, the control device 4 estimates the state of multiple people using non-verbal cues and multimodal temporal audiovisual observation information of the environmental context.

[0016] [Configuration Example of Communication System] Next, a configuration example of the communication system 1 will be described. Fig. 2 is a diagram showing a configuration example of the communication system according to this embodiment. As shown in Fig. 2, the communication system 1 includes, for example, a sensor 2 (detection unit), a robot 3, and a control device 4.

[0017] The sensor 2 includes, for example, a sound pickup device 21 (detection unit), an imaging device 22 (detection unit), and a communication unit 23. The robot 3 includes, for example, a control unit 31 and a communication unit 32.

[0018] The control device 4 includes, for example, a learning device 5, a first generating unit 6, and a second generating unit 7. The learning device 5 includes, for example, a detecting unit 51 (detecting unit), a modeling unit 52, a state model 53, and a sociological model 54 (also simply referred to as a "model"). The detecting unit 51 includes, for example, a first feature information detecting unit 511, a second feature information detecting unit 512, a third feature information detecting unit 513, and a fourth feature information detecting unit 514. The modeling unit 52 includes, for example, a first modeling unit 521, a second modeling unit 522, a third modeling unit 523, and a fourth modeling unit 524.

[0019] (Sensor) The sensor 2 is an environmental sensor or the like that detects the state (sensing data) of each person. Note that the sensor 2 may include sensors other than the sound pickup device 21 and the image pickup device 22.

[0020] The sound collector 21 is, for example, a microphone array composed of multiple microphones. The sound collector 21 collects human voice signals (sensing data) and may output the collected voice signals in analog form, or may convert the analog signals into digital signals and output them. When outputting digital signals, the multiple microphones are synchronized.

[0021] The image capturing device 22 is, for example, an RGB (red-green-blue)-D (depth) camera that can also acquire depth information. The image capturing device 22 may also be an RGB camera and a distance sensor. The image capturing device 22 captures the state of a person (sensing data).

[0022] The communication unit 23 outputs the sound signal (sensing data) collected by the sound collector 21 and the image data (sensing data) captured by the image capturing device 22 to the control device 4. The sensor 2 and the control device 4 are connected via a wired or wireless network or the like.

[0023] (Robot) The robot 3 may be, for example, a humanoid robot, a robot capable of walking on two legs, or a robot with a head and a body. The robot 3 may be, for example, equipped with an image display unit, a sound pickup unit, a photographing device, an audio output device, a power supply unit, etc. The robot 3 may also be equipped with a sensor 2. The robot 3 may also be equipped with a control device 4. The robot 3 may receive operational instructions, for example, via remote control. The robot 3 may also be any robot, including, for example, a desktop robot, a mobile robot, an industrial robot arm, or an intelligent agent such as an avatar in any embodiment in a virtual space. The robot 3 may also be a human model such as an avatar, or a robot with multiple degrees of freedom, including but not limited to key points such as hand joints, head, and face.

[0024] The control unit 31 controls the operation of the robot 3 in accordance with the action command output by the control device 4 .

[0025] The communication unit 32 acquires behavioral commands output by the control device 4. The communication unit 32 outputs information acquired by the robot 3 (for example, detection data from sensors provided at the joints, captured image data, etc.) to the control device 4. The robot 3 and the control device 4 are connected via a wired or wireless network or the like.

[0026] (Control Device) The learning device 5 acquires audio signals output by the sensor 2, image capture data of a time series (t-2, t-1, t), etc., and detects multiple pieces of status information of a person from the acquired data. The learning device 5 generates multiple pieces of modeling data using the multiple pieces of status information of the detected person (detection data). The learning device 5 trains each of the modelers (first modeler 521 to fourth modeler 524) included in the modeling unit 52 using the generated multiple pieces of modeling data.

[0027] The detection unit 51 acquires audio signals, photographic data, etc. output by the sensor 2, and outputs low-level sensing data from the acquired data, including but not limited to face and body key points, head position, etc., to be used in the modeling unit 52 and sociological model 54. The sensing data is information relating to the state of multiple people based on audiovisual information.

[0028] The first feature information detection unit 511 inputs image data to a convolutional neural network (CNN) that has been trained using, for example, training data, and detects or extracts facial feature information, attention feature information based on facial direction, gaze, etc., as first feature information. The first feature information detection unit 511 outputs the detected first feature information for consecutive times (timesteps) t-2, t-1, and t to the first modeler 521, the second modeler 522, the third modeler 523, the fourth modeler 524, and the sociological model 54. Note that well-known image processing may be performed on the first image data to detect or extract facial features, including but not limited to facial key points, a cut-out human face and head, CNN feature information corresponding to the face or head, and notable features, and use these as the first feature information.

[0029] The second feature information detection unit 512 performs, for example, well-known image processing on the image data to detect or extract, as second feature information, main features of the human body, such as but not limited to key points of the human body, segmented human bodies and their corresponding CNN features, human body parts, and the skeleton of a human pose. The second feature information detection unit 512 detects or extracts, as second feature information, information on the human body's physical attention (e.g., body orientation, posture, hand and arm movements during conversation, etc.). The second feature information detection unit 512 outputs the detected second feature information for consecutive times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54.

[0030] The third feature information detection unit 513 performs well-known speech recognition processing (e.g., sound source separation, sound source direction estimation, noise suppression, etc.) on the speech signal to detect or extract acoustic feature information, such as Mel-frequency cepstrum (MFCC) and prosodic features, which represent intonation, pitch, etc., as the third feature information. The third feature information detection unit 513 outputs the detected third feature information for a period similar to consecutive times t-2, t-1, and t to the first modeler 521, the second modeler 522, the third modeler 523, the fourth modeler 524, and the sociological model 54. In this embodiment, paralanguage refers to information conveyed from a speaker to a listener that is outside the context of language, excluding body language such as gestures, and includes, for example, the strength and weakness of the voice when speaking, pitch, intonation, etc.

[0031] The fourth feature information detection unit 514 performs, for example, well-known image processing on the image data to detect or extract true feature information of the environment, feature information of objects in the environment, etc. as fourth feature information. The fourth feature information detection unit 514 outputs the detected fourth feature information for times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54.

[0032] The modeling unit 52 uses the feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time step t or multiple times t-2, t-1, and t to generate multiple modeling data to be used for training the sociological model 54, and outputs the generated modeling data to the sociological model 54 as environmental information. The modeling data is information related to the attention of multiple people. That is, the modeling unit 52 uses low-level information from the detection unit 51 to generate high-level modeling data, including but not limited to gaze behavior, body language, and human-object interactions. The output of the modeling unit 52 is, for example, data that can be used to establish edges and nodes in a graph representation and their features. Each modeler (the first modeler 521 to the fourth modeler 524) may be, for example, a gaze targeting model, a joint attention model, a body language recognition model, or the like. The nodes are not limited to people, but may also be objects (people, objects, etc.) that attract people's attention.

[0033] The first modeling device 521 generates modeling data of the gaze trajectory of each person using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection device 51 for one time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the degree of attention paid to a common target by most people in the scene. Note that the first modeling device 521 can also output the gaze target of everyone in the scene.

[0034] The second modeling device 522 generates modeling data of body language or behavior for each person using the multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for a single time t, or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the physical behavior (e.g., walking, sitting, etc.) of all people in the scene.

[0035] The third modeler 523 converts the paralanguage of each person into high-level auditory behavior using one or more pieces of feature information (first to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. This generates modeling data, which is output to the sociological model 54 as environmental information and instructs auditory behavior such as sound source direction estimation, sound source position estimation, and sound source estimation.

[0036] The fourth modeler 524 generates modeling data of the environment and context using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. The fourth modeler 524 then outputs the generated modeling data to the sociological model 54 as environmental information not limited to object recognition and how humans interact with these objects, i.e., as human-object interactions.

[0037] The state model 53 is a model for inferring human internal states, such as emotions, intentions, or internal states, such as positive and negative, positive or negative, including but not limited to valence and arousal, which may also be represented by physiological responses, such as electroencephalographic information, such as electroencephalographic frequency bands, and electrocardiographic information, including but not limited to heart rate characteristics.

[0038] The sociological model 54 is, for example, a model based on social interaction dynamics, such as a model based on the TGN (Temporal Graph Network) model (see, for example, Reference 1). The sociological model 54 is a trained model that receives environmental information from the sound pickup device 21, the image capture device 22, the detection unit 51, and the modeling unit 52. The trained model can be used to construct a graph of nodes and edges to represent the social interaction dynamics of the TGN model. These inputs can also be encoded as features to better represent multimodal temporal interaction dynamics. The encoding can be performed using different methods, such as a text encoder or a visual encoder. The output of the sociological model 54 is a node embedding representation and features of the TGN model, or an embedding representation or feature. During training, the sociological model 54 is trained using data stored in a database, for example. For example, during training, the sociological model 54 uses link prediction to learn social interaction dynamics. Alternatively, the learning may be privileged learning (PL) or the like. The sociological model 54 may be an LCBM model for representing the dynamics of social interactions. In this case, the sociological model 54 fine-tunes a large language model (LLM) (see, for example, Reference 2) using context extracted from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the text or visual encoder of the modeling unit 52. The fine-tuned LLM model is a large context and behavioral model (LCBM). To generate behaviors such as movements and utterances, the output from the LCBM is passed to, for example, a Gaussian mixture model and decoded to generate behavior commands for controlling different degrees of freedom of the robot 3.

[0039] Reference 1; Emanuele Rossi, Ben Chamberlain, et al., “Temporal Graph Networks for Deep Learning on Dynamic Graphs”, arXiv:2006.10637 [cs.LG], 2020 Reference 2; Ashmit Khandelwal, Aditya Agrawal, et al., “Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior”, arXiv:2309.00359 [cs.CL], 2023

[0040] The first generation unit 6 generates nonverbal cues (NOCs) based on the output of the sociological model 54. The nonverbal cues (NOCs) will be described later. The first generation unit 6 receives the output from the large-scale language model of the sociological model 54 as input to a Gaussian model (e.g., a Gaussian mixture model), and inputs the linguistic information and embeddings of a TGN model or LCBM representing social interaction dynamics. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder to generate nonverbal cue behavioral commands for the robot 3.

[0041] The second generation unit 7 generates speech information (text, audio signals) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the generated speech information to a Gaussian model (e.g., a Gaussian mixture model), inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the speech information.

[0042] In the example shown in FIG. 2, the detection unit 51 is provided with four feature information detection units (511 to 514), but this is not limiting. The number of feature information detection units may be two or more, and may be five or more. In addition, the modeling unit 52 is provided with four modeling devices (first modeling device 521 to fourth modeling device 524), but this is not limiting. The number of modeling devices may be two or more, and may be five or more. In addition, in the example shown in FIG. 2, the detection unit 51 is provided with a first generation unit 6 and a second generation unit 7, but this is not limiting, and it is sufficient that at least one is provided.

[0043] With this configuration, in this embodiment, link prediction is used to train a TGN model (sociological model 54) to learn social interaction dynamics. In addition, in this embodiment, node embeddings and features representing social interaction dynamics are extracted to infer human states. These embeddings / features can be used for other tasks, such as next speaker prediction, group state estimation, etc.

[0044] [Nonverbal Cues (NOC)] Next, we will explain nonverbal cues (NOC). In order to infer human states in real environments using audiovisual data, it is necessary to understand and represent the dynamics of social interactions using multimodal temporal observations, not just nonverbal behavior and vocal activity, because human states tend to depend on environmental context. Here, nonverbal behavior refers to gestures (body movements), posture, movement (nodding and general behavior), facial expressions, and tone of voice (volume, speed, and pauses). When a robot communicates with a person, for example, rather than simply outputting an automatically generated voice signal, a communication robot has been proposed that responds to the voice signal with voice tone and volume, and even displays images of eyes and mouths on a display unit to accompany the speech (e.g., JP 2022-6610 A). Furthermore, it has been proposed to improve communication with users by moving the robot's head or other parts to make the user feel depressed or happy, or by having a robot with arms and hands perform gestures.

[0045] For this reason, in this embodiment, verbal information and non-verbal information (non-verbal cues, non-verbal behavior) are generated as behavioral commands for the robot 3. Note that in this embodiment, the non-verbal cues include gaze interaction, body language, vocal activities such as paralanguage, speech characteristics, facial expressions, and the like.

[0046] [Internal and External Stimuli] Next, we will explain internal stimuli and external stimuli. Internal stimuli are represented using temporal information (i.e., TGN memory). External stimuli are represented by multimodal inputs from the sound pickup device 21, the image capture device 22, the detection unit 51, and the modeling unit 52 to the sociological model 54, enabling the representation of TGN, i.e., social interaction dynamics using TGN. The sociological model 54 then converts these inputs into TGN node embeddings as outputs. These node embeddings can be used for downstream supervised learning tasks, such as human state inference, next speaker prediction, and action generation, as shown in Figure 2.

[0047] (TGN Model) Here, an overview of the TGN model will be described (see, for example, Reference 1). First, an example of the calculations that the TGN performs on a batch of time-stamped interactions will be described. FIG. 3 illustrates an example of the processing of the TGN model. The first module g11 generates an embedding emb using a time graph and node memory (step S1). The second module g12 predicts batch interactions using the embedding emb (step S2). The third module g13 uses the embedding emb to calculate the probability of each edge and calculate the loss (step S3). The fourth module g14 to the sixth module g16 are storage modules. The interaction of the processing in steps S1 to S3 is used to update each storage module. However, this configuration and processing is a simplified operation flow, and the fourth module g14 to the sixth module g16 do not receive gradients, which prevents learning of all modules.

[0048] Next, we will explain a TGN model (hereinafter also referred to as an improved TGN model) that solves the above problems (see, for example, Reference 1). Figure 4 shows the operation flow of the improved TGN model used for learning memory-related modules. The first module (Raw Message Store) g21 stores raw information {rm 1 (t 1 ), rm 2 (t 2 ), rm 3 (t 3 ), rm 4 (t 4 )}, which allows the improved model to postpone updates to the storage module caused by interactions until a later batch.

[0049] The first module g21, the second module g22, and the third module g23 update the storage module using messages calculated from raw messages stored in the previous batch (steps S21-S23). The fourth module g24 calculates the embedding emb using the updated information in the storage module (arrow g28) (step S24). The fifth module g25 uses the embedding emb to calculate the probability of each edge and calculate the loss (step S25). Note that this probability of each edge represents the probability that it is the target of interest.

[0050] The output msg of the first module g21 is expressed by the following equation (1).

[0051]

[0052] In formula (1), s i(t-1) is the stored data of the past i, and s j(t-1) is the stored data of the past j. i,j(t) is the characteristic of the last interaction between i and j, and t enc is a linear transformation of time between the current time t and the update of past memory. Furthermore, the output s of the third module g23 i(t)The relationship between the output mem of the second module g22 and the output mem of the second module g22 is expressed by the following equation (2).

[0053]

[0054] In formula (2), mem is the Memory update model (GRU), and m (i,j) is the message generated at the time of the last update, and s i(t-1) is the previous state in memory. The fourth module g24 generates a transformation that represents the state of the node for the entire graph. Here, the node state is updated (node ​​embedding) h i ' is expressed by the following equation (3).

[0055]

[0056] In formula (3), α ij is the attention mechanism as shown by symbol g43 in FIG. 5, W is the learnable weight matrix as shown by symbol g42 in FIG. 5, and h j is the target node (neighborhood) as shown by symbol g41 in FIG. 5. Also, σ is the activation function. FIG. 5 is a diagram for explaining each symbol in equation (3). Also, Σ(·) in equation (3) is the activation function of h as shown by symbol g43 in FIG. 1 '*α 11、0 h 2 '*α 12 and h 3 '*α 13 The output of each is added together. 1 ' is the embedding node for 1, e.g., [0.234, 0.412, ..., 0.98] (1,100) Furthermore, α in equation (3) ij is used as in the following equation (4).

[0057]

[0058] α c,I,j is the attention that node i should pay to node j, c is the multi-head attention (set to 1 in Equation (3) for simplicity), and l is the layer (starting from 0). c,i (l) is the characteristic information of the source node as shown in FIG.c,j (l) is the characteristic information of the converted node as shown in FIG. c,i,j is edge feature information between the source node and the destination node as shown in FIG. 6. FIG. 6 is a diagram showing a processing example in the case of attention (gaze) between C and H. The numerator q of the right side of Equation (4) c,i (l) corresponds to symbol g51 in FIG. 6, and k c,j (l) +e c,i,j corresponds to symbol g52 in Figure 6. The numerator of equation (4) is the attention of only one neighbor, as shown by symbol g53 in Figure 6, and the denominator of equation (4) is the operation of finding the same thing for all neighbors, and the implication of equation (4) is the normalization of attention based on others.

[0059] This allows the calculation of the memory-related modules to directly affect the loss (steps S25 and S26) and receive the gradient. Finally, the sixth module stores the raw messages of this batch of interactions in the first module g21 for use in future batches (step S26).

[0060] In an embodiment, such an improved TGN model is used to detect, for example, gaze interactions between multiple people. For example, assume that three people (A, B, and C) are having a conversation. Assume that A is speaking at time t1 and C is speaking at time t2. The gaze interactions between multiple people in this situation can be represented by a network such as that shown in FIG. 7.

[0061] FIG. 7 is a diagram showing an example of a graph representation of gaze interactions between multiple people. In FIG. 7, each node represents a different person. The tip of the edge arrow indicates the object that each person is looking at at each time. In the example of FIG. 7, for example, it is estimated from the graph representation that A is speaking at time t1, and that B and C are speaking at time t2. In the graph representation, multiple interactions can be estimated using environmental information such as the relationship between A, B, and C (for example, A is a teacher and B and C are students).

[0062] In this embodiment, such environmental information is converted into text as shown in FIG. 8. FIG. 8 is an example of converting each piece of estimated information into text. The information indicated by reference symbol g101 is a diagram showing an example of generating information for converting each piece of estimated information into text based on, for example, speaking information (talking information), the relative seating positions of multiple people, and the roles of each person among the multiple people. The information indicated by reference symbol g102 is an example of a summary of the information indicated by reference symbol g101. The text indicated by reference symbol g103 is an example of converting the information indicated by reference symbol g102 into text according to predetermined rules.

[0063] Note that the configurations and processes described using Figures 4 and 5 are merely examples and are not limited to these. For example, data input to the TGN model may be encoded using a privileged encoder, a wild encoder, or the like, and then embedded in a node. Alternatively, an audio signal may be processed by an audio encoder, image data may be processed by an image encoder, and the results of processing by the audio encoder and the image decoder may be encoded using a wild encoder or the like, embedded in a node, and then input to the TGN model. Furthermore, encoding may be performed using a different method, such as a text encoder or a visual encoder.

[0064] Next, in this embodiment, as shown in Fig. 9, a text g111 that verbalizes the environmental information generated by aggregating the information for converting each piece of information into text is input to a language model g112, and a fixed-size context-embedded representation g113 is output. Fig. 9 is a diagram for explaining the generation of a context-embedded representation. Note that the language model is based on, for example, BERT (Bidirectional Encoder Representations from Transformers), but is not limited to this and other language models may be used.

[0065] Next, in this embodiment, the fixed-size context embedded representation thus generated is used to train a graph representation, for example, an improved TGN model, as shown in Fig. 10. Fig. 10 is a diagram showing an example of data used during training and a training method, and an example of data used during use and an estimation method, in this embodiment.

[0066] The box g120 in the upper row indicates data used for model training and the graph representation model (TGN model) to be trained. The table indicated by reference symbol g121 is the context embedded representation for each time generated as described with reference to FIG. 9, and mutual data of the person's identification information (Id (Gazing Subject)) and gaze direction identification information (Gaze Id (Gazed Subject)) for each time, and is data used when training the model. The identification information is assigned by, for example, the modeling unit 52. Note that the identification information may also be assigned by the detection unit 51. Reference symbol g122 indicates a model, for example, an improved TGN model.

[0067] The box g140 in the lower row shows data, processing, etc. during use. The table with symbol g141 is mutual data between the person's identification information (Id (Gazing Subject)) at each time and the gaze identification information (Gaze Id (Gazed Subject)) at each time, and is data input to the trained model. Symbol g142 is a diagram showing an example of a trained graph representation. The trained model receives input of mutual data at each time such as the table with symbol g141, and outputs the probability of each edge, i.e., information representing the object being watched. The table with symbol g143 shows the probability of each edge output by the model and state information in which the probability of each edge is set to 1 if it is equal to or greater than a threshold (e.g., 0.5), and 0 if it is less than the threshold.

[0068] That is, when in use, information indicating the focus of each person at each time is input to the representation of the interaction dynamics, which is such a trained model, and the focus based on the gaze is estimated, for example, based on the probability of each edge. Note that the probability of each edge is calculated by the fifth module g25 in the improved TGN model. In this embodiment, for example, the first modeling device 521 has the above-mentioned improved TGN model.

[0069] As described above, this embodiment has the following configuration and processing: I. Human state inference (emotions, values, arousal levels, etc.) is performed using multimodal temporal and non-verbal cues. II. Internal stimuli are represented using temporal information. III. External stimuli are represented by multimodal input from the TGN model. IV. Non-verbal cues include gaze interactions, body language, vocal activities such as paralanguage, speech features, facial expressions, etc.

[0070] As a result, the configuration and method of this embodiment are practical for implementation in a real environment because the sociological model 54 relies on audiovisual data. Furthermore, the configuration and method of this embodiment rely on multimodal input, which allows the sociological model 54 to be more robust than before.

[0071] In this embodiment, the generation units (first generation unit 6, second generation unit 7) generate verbal and non-verbal behavior commands using the output of the sociological model 54, and control the robot 3. As a result, according to this embodiment, it is not necessary to wear any special equipment, and it is possible to understand the states (interactions) of multiple people using visual data detected by the sensor 2. Furthermore, according to this embodiment, it is possible to make the robot 3 behave (speak, gesture, show facial expressions, etc.) using the sociological model 54 and state model 53 for inferring people's states and non-verbal behavior, thereby realizing empathic interaction with people in a real environment.

[0072] With this configuration and processing, the configuration and method of this embodiment are practical for implementation in a real environment because the model relies on audiovisual data. Furthermore, the configuration and method of this embodiment rely on multimodal input, which improves the robustness of the model compared to conventional methods.

[0073] [Processing Procedure] Next, an example of a processing procedure performed by the control device 4 will be described. First, an example of a procedure for training a model (modeling unit 52, sociological model 54) will be described. Fig. 11 is a flowchart of the processing procedure for training a model according to this embodiment.

[0074] (Step S101 ) The detection unit 51 acquires the detection data detected by the sensor 2 .

[0075] (Step S102) The detection unit 51 acquires relationship data (for example, information on relationships between people).

[0076] (Step S103) The detection unit 51 detects feature information (first feature information to fourth feature information) from the detection data and the relational data.

[0077] (Step S104) The modeling unit 52 converts the information about the environment included in the feature information (information based on the relational data) into text.

[0078] (Step S105) The modeling unit 52 receives inputs of environmental information from the sound pickup device 21, the image capture device 22, the detection unit 51, and the modeling unit 52, which can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the TGN model. These inputs can also be encoded as features to better represent multimodal temporal interaction dynamics. The output is the node embedding of the TGN model. During training, the sociological model 54 is trained for link prediction using data stored in a database, for example. This is a type of supervised learning to represent social interaction dynamics using the output of the sociological model 54, i.e., the node embedding.

[0079] (Step S106) The modeling unit 52 uses the generated node embedding representation to train another supervised learning model, which is, for example, a human state estimation result, a next speaker estimation result, etc. Note that the learning model is a model provided in each of the modeling units 52 itself.

[0080] In this embodiment, the above steps S104 to S106 are referred to as a modeling data generation process.

[0081] Next, a description will be given of an example of a procedure for making an estimation using the trained model (the modeling unit 52 and the sociological model 54). Fig. 12 is a flowchart of the processing procedure when using the model according to this embodiment.

[0082] (Step S201) The detection unit 51 acquires the detection data detected by the sensor 2.

[0083] (Step S202) The detection unit 51 detects feature information (first feature information to fourth feature information) from the detection data.

[0084] (Step S203) The modeling unit 52 assigns identification information to a plurality of people.

[0085] (Step S204) The modeling unit 52 converts environmental information (information based on relational data) contained in the feature information into text. The modeling unit 52 then receives environmental information from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the modeling unit 52 that can be used to construct a graph consisting of nodes and edges. The modeling unit 52 then outputs the node embedding data of the TGN model to the sociological model 54.

[0086] (Step S205) The sociological model 54 receives information output from the first modeler 521 to the fourth modeler 524 for one time or a series of consecutive times, and converts the information into node embedding for the TGN.

[0087] (Step S206) The first generation unit 6 generates a non-verbal cue (NOC) using the output of the sociological model 54. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the non-verbal cue. Furthermore, the second generation unit 7 generates speech information (text, audio signal) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the speech information.

[0088] (Step S207) The first generation unit 6 and the second generation unit 7 each output the generated action instructions to the robot 3 to control the action of the robot 3.

[0089] 11 and 12 are merely examples, and are not intended to be limiting. The control device 4 may, for example, perform several processes in parallel.

[0090] In the above embodiment, an example has been described in which each modeler of the modeling unit 52 is a TGN model, but this is not limiting. The modelers (first modeler 521 to fourth modeler 524) may be other models, and the models of each modeler may be different. Furthermore, an example has been described in which the sociological model 54 is a TGN model, but the sociological model 54 may be another model, such as an LCBM.

[0091] [Evaluation Results] Next, examples of evaluation results will be described. FIG. 13 is a diagram showing examples of evaluation results for the first and second methods. Note that the first method is processing that uses the output of BERT. The second method is processing in which, for example, a text encoder is replaced with a visual encoder. The horizontal axis D1 to D8 in FIG. 13 indicates the attributes of the subjects, and the vertical axis indicates the score (100% is the best score). Note that the score indicates whether the result of estimating the person's state (focus position) is correct or not.

[0092] As shown in Figure 13, the evaluation results were good using the configuration and processing method of this embodiment. Furthermore, when the second method was used, the overall performance was even better than that of the first method. The unsupervised session showed the greatest improvement. The trained session also showed the highest performance.

[0093] In addition, a program for implementing some or all of the functions of the control device 4 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the control device 4. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Furthermore, the term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. Alternatively, some or all of these components may be realized by LSI (Large Scale Integration) hardware (including circuitry) such as an ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), or SOC (System On Chip), or may be realized by a combination of software and hardware.

[0094] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0095] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention.

[0096] 1...Communication system, 2...Sensor, 3...Robot, 4...Control device, 5...Learning device, 6...First generation unit, 7...Second generation unit, 21...Sound collector, 22...Image capture device, 23...Communication unit, 31...Control unit, 32...Communication unit, 51...Detection unit, 52...Modeling unit, 53...State model, 54...Sociological model, 511...First feature information detection unit, 512...Second feature information detection unit, 513...Third feature information detection unit, 514...Fourth feature information detection unit, 521...First modeling device, 522...Second modeling device, 523...Third modeling device, 524...Fourth modeling device, g21...First module, g22...Second module, g23...Third module, g24...Fourth module, g25...Fifth module, g26...Sixth module

Claims

1. A learning device comprising: a detection unit that acquires, in an environment where a plurality of people are present, a plurality of sensing data that is information relating to the states of a plurality of people based on audiovisual information; a modeling unit that uses the acquired plurality of sensing data to create a plurality of modeling data that is information relating to the attention of the plurality of people; and a model that inputs the plurality of modeling data and outputs information to control the behavior of a robot that is in an environment where a plurality of people are present and communicates with the plurality of people.

2. The learning device described in claim 1, wherein the plurality of sensing data are at least two of information indicating whether each of the plurality of people is speaking, body language information of the upper body of each of the plurality of people, position information of each of the plurality of people, and information indicating the gaze direction of each of the plurality of people.

3. The learning device of claim 1 or claim 2, wherein the modeling unit inputs information from the detection unit that can be used to construct a graph consisting of nodes and edges to represent the social interaction dynamics of the model, and outputs a node embedding representation to the model.

4. The learning device according to claim 1 or claim 2, wherein the model is based on a TGN (Temporal Graph Network) model, and inputs the modeling data output by the modeling unit for one time or multiple consecutive times, converts it into a TGN node-embedded representation, inputs information based on the converted node-embedded representation into a Gaussian model, and inputs the output from the Gaussian model into a decoder to create behavioral commands for the robot.

5. A learning method for a learning device having a detection unit, a modeling unit, and a model, wherein the detection unit acquires a plurality of sensing data in an environment where a plurality of people are present, the sensing data being information relating to the states of the plurality of people based on audiovisual information; the modeling unit uses the acquired sensing data to create a plurality of modeling data being information relating to the attention of the plurality of people; and the model inputs the modeling data and outputs information to control the behavior of a robot that is in the environment where a plurality of people are present and communicates with the plurality of people.

6. A program that causes a computer having a learning device with a model to acquire, in an environment where a plurality of people are present, a plurality of sensing data that is information about the states of a plurality of people based on audiovisual information, use the acquired plurality of sensing data to create a plurality of modeling data that is information about the attention of the plurality of people, and input the plurality of modeling data to output information that controls the behavior of a robot that is in an environment where a plurality of people are present and communicates with the plurality of people.