Visual behavior estimation apparatus, method for estimating visual behavior, and program

The integration of multimodal data using a transformer encoder and LSTM network addresses the limitation of conventional methods by accurately estimating visual behavior, accounting for internal and external stimuli, thereby improving prediction accuracy in real-world social interactions.

JP2025137269APending Publication Date: 2025-09-19HONDA MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024036378
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Conventional visual behavior estimation techniques fail to consider the influence of internal and external stimuli on human visual behavior, leading to incomplete and inaccurate estimates.

Method used

A visual behavior estimation device and method that integrates multimodal sensing data, including audio and visual inputs, using a transformer encoder and LSTM network to infer attention areas, accounting for temporal dynamics and social interaction contexts.

Benefits of technology

Enables accurate estimation of visual behavior by considering internal and external stimuli, enhancing the robustness and accuracy of visual behavior prediction in real-world environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137269000001_ABST
    Figure 2025137269000001_ABST
Patent Text Reader

Abstract

To provide a visual behavior estimation apparatus, a method for estimating visual behavior, and a program capable of estimating visual behavior by taking into account effects of internal and external stimuli.SOLUTION: A visual behavior estimation apparatus includes: a detection unit configured to acquire plural pieces of sensing data; and an estimation unit configured to concatenate the plural pieces of sensing data, which are multimodal data, combine the concatenated sensing data with learnable tokens, integrate the combined data at different times, and estimate locations that a plurality of persons are paying attention to using the integrated data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a visual behavior estimation device, a visual behavior estimation method, and a program. [Background technology]

[0002] Technologies have been developed to estimate what people are looking at. These technologies use gaze models. Attention models, such as gaze direction models, gaze target models, and joint attention models, are usually developed using single inputs, such as images. Recent research on "interactive joint attention estimation" uses multimodal inputs from visual data for a volleyball scenario (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Chihiro Nakatani, Hiroaki Kawashima, et al., “Interaction-aware Joint Attention Estimation Using People Attributes”, arXiv:2308.05382 [cs.CV], p10224-10233, 2023 Summary of the Invention [Problem to be solved by the invention]

[0004] Visual behavior is a primary human sense. It is complex and depends on internal stimuli (experiences, memories, emotional states, intentions, etc.) and external stimuli (other people's actions and states, objects, the environment, etc.). However, conventional techniques have not been able to take into account the influence of such internal and external stimuli when estimating visual behavior.

[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a visual behavior estimation device, a visual behavior estimation method, and a program that can estimate visual behavior by taking into account the influence of internal and external stimuli. [Means for solving the problem]

[0006] (1) In order to achieve the above object, a visual behavior inference device according to one aspect of the present invention is a visual behavior inference device including: a detection unit that acquires multiple pieces of sensing data; and a modeling unit that extracts gaze information from an estimation unit that connects the multiple pieces of sensing data that are multimodal data, combines the connected pieces of sensing data with a learnable token, integrates the connected data from different times, and uses the integrated data to infer where multiple people are focusing their attention; weights the gaze information with environmental situation information; and identifies areas of attention within the environment.

[0007] (2) In the visual behavior estimation device according to one aspect of (1) above, the estimation unit may input the concatenated multiple pieces of sensing data to a fully connected layer, input the concatenated multiple pieces of sensing data and a learnable token to a transformer encoder to combine them, integrate the combined data at different times using a long short-term memory (LSTM) network to take temporal dynamics into consideration, pass the data integrated by the long short-term memory network through a fully connected layer, and then estimate the locations where the multiple people are paying attention using an activation function.

[0008] (3) In the visual behavior estimation device according to one aspect of (1) or (2) above, the plurality of sensing data may be at least two of information indicating whether a person is speaking, body language information of the person, and position information of either the person or the object.

[0009] (4) In order to achieve the above object, a visual behavior estimation method according to one aspect of the present invention is a visual behavior estimation method in which a detection unit acquires multiple pieces of sensing data, and an estimation unit concatenates the multiple pieces of sensing data that are multimodal data, combines the concatenated multiple pieces of sensing data with a learnable token, integrates the combined data from different times, and uses the integrated data to estimate areas where multiple people are paying attention.

[0010] (5) To achieve the above object, a program according to one aspect of the present invention is a program that causes a computer of a visual behavior estimation device to acquire multiple pieces of sensing data, concatenate the multiple pieces of sensing data that are multimodal data, combine the concatenated multiple pieces of sensing data with a learnable token, integrate the combined data from different times, and use the integrated data to estimate areas where multiple people are paying attention. [Effects of the Invention]

[0011] According to the above (1) to (5), visual behavior can be estimated taking into consideration the influence of internal and external stimuli. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram for explaining an overview of a system configuration and an overview of processing in an embodiment. [Figure 2] 1 is a diagram illustrating an example of the configuration of a communication system according to an embodiment. [Figure 3] 1A and 1B are diagrams illustrating an example of a configuration of a model and an example of processing according to an embodiment. [Figure 4] 10A and 10B are diagrams for explaining a detection process of a point of interest according to an embodiment. [Figure 5] FIG. 15 is a block diagram showing an example of the overall configuration of the block diagram of FIG. 14. [Figure 6] 10 is a flowchart of a processing procedure when learning a model according to an embodiment. [Figure 7] 10 is a flowchart of a processing procedure when using a model according to an embodiment. [Figure 8] FIG. 10 shows the effect of interaction length on training loss and test AUC in different sessions. [Figure 9] FIG. 10 is a diagram showing an example of an evaluation result. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).

[0014] [overview] FIG. 1 is a diagram illustrating an overview of the system configuration and processing in this embodiment. As shown in FIG. 1, there are multiple people in an environment, engaged in conversation, etc. The environment also includes a robot 3 that can communicate with the people. Furthermore, a sensor 2 that senses the state of the people is installed in the environment. A control device 4 controls communication between the robot 3 and the people based on detection data detected by the sensor 2. The control device 4 trains a model to output behavioral commands to the robot 3 using the detection data. Note that there are multiple types of sensed data, such as facial images, audio signals, and human movements. Furthermore, there are multiple types of modeling data generated using the detection data. In this embodiment, the control device 4 estimates the state of multiple people using non-verbal cues and multimodal temporal audiovisual observation information of the environmental context.

[0015] [Example of communication system configuration] Next, an example of the configuration of the communication system 1 will be described. 2 is a diagram showing an example of the configuration of a communication system according to this embodiment. As shown in FIG. 2, the communication system 1 includes, for example, a sensor 2 (detection unit), a robot 3, and a control device 4.

[0016] The sensor 2 includes, for example, a sound collector 21 (detection unit), an imaging device 22 (detection unit), and a communication unit 23. The robot 3 includes, for example, a control unit 31 and a communication unit 32.

[0017] The control device 4 includes, for example, a learning device 5, a first generation unit 6, and a second generation unit 7. The learning device 5 includes, for example, a detection unit 51 (detection unit), a modeling unit 52 (estimation unit), a state model 53, and a sociological model 54 (social dynamics model) (also simply referred to as "model") (estimation unit). The detection unit 51 includes, for example, a first feature information detection unit 511, a second feature information detection unit 512, a third feature information detection unit 513, and a fourth feature information detection unit 514. The modeling section 52 includes, for example, a first modeling device 521, a second modeling device 522, a third modeling device 523, and a fourth modeling device 524. The visual behavior estimation device 50 includes, for example, a detection unit 51 (first feature information detection unit 511 to fourth feature information detection unit 514) and a first modeling device 521.

[0018] (sensor) The sensor 2 is an environmental sensor or the like that detects the state (sensing data) of each person. Note that the sensor 2 may include sensors other than the sound pickup device 21 and the image pickup device 22.

[0019] The sound collector 21 is, for example, a microphone array made up of multiple microphones. The sound collector 21 collects human voice signals (sensing data) and may output the collected voice signals in analog form, or may convert the analog signals into digital signals and output them. When outputting digital signals, the multiple microphones are synchronized.

[0020] The image capturing device 22 is, for example, an RGB (red-green-blue)-D (depth) camera that can also acquire depth information, etc. Note that the image capturing device 22 may be an RGB camera and a distance sensor, etc.

[0021] The communication unit 23 outputs the audio signal (sensing data) collected by the sound collector 21 and the image data (sensing data) captured by the photographing device 22 to the control device 4. The sensor 2 and the control device 4 are connected via a wired or wireless network or the like. The photographing device 22 captures the state of a person (sensing data).

[0022] (robot) The robot 3 may be, for example, a humanoid robot, a robot capable of walking on two legs, or a robot with a head and a body. The robot 3 may be equipped with, for example, an image display unit, a sound pickup unit, a photographing device, an audio output device, a power supply unit, etc. The robot 3 may also be equipped with a sensor 2. The robot 3 may also be equipped with a control device 4. The robot 3 may receive operation instructions, for example, via remote control. The robot 3 may also be any robot, including, for example, a desktop robot, a mobile robot, an industrial robot arm, or an intelligent agent such as an avatar in any embodiment in a virtual space. The robot 3 may also be a human model such as an avatar, or a robot with multiple degrees of freedom, including but not limited to key points on the hand joints, head, and face.

[0023] The control unit 31 controls the operation of the robot 3 in accordance with the action command output by the control device 4.

[0024] The communication unit 32 acquires behavioral commands output by the control device 4. The communication unit 32 outputs information acquired by the robot 3 (for example, detection data from sensors provided at the joints, captured image data, etc.) to the control device 4. The robot 3 and the control device 4 are connected via a wired or wireless network or the like.

[0025] (Control device) The learning device 5 acquires the audio signal output by the sensor 2, the photographic data of the time series (t-2, t-1, t), etc., and detects a plurality of pieces of state information of the person from the acquired data. The learning device 5 generates a plurality of pieces of modeling data using the plurality of pieces of state information of the detected person (detection data). The learning device 5 trains each of the modelers (first modeler 521 to fourth modeler 524) included in the modeling unit 52 using the generated plurality of modeling data.

[0026] The detection unit 51 acquires audio signals, photographic data, etc. output by the sensor 2, and outputs low-level sensing data from the acquired data, including but not limited to face and body key points, head position, etc., to be used in the modeling unit 52 and sociological model 54. The sensing data is information about the state of multiple people based on audiovisual information.

[0027] The first feature information detection unit 511 inputs the image data to a convolutional neural network (CNN) that has been trained using, for example, training data, and detects or extracts facial feature information, attention feature information based on facial direction, gaze, etc. as first feature information. The first feature information detection unit 511 outputs the detected first feature information for consecutive times (timesteps) t-2, t-1, and t to the first modeler 521, the second modeler 522, the third modeler 523, the fourth modeler 524, and the sociological model 54. Note that well-known image processing may be performed on the first image data to detect or extract facial features, including but not limited to facial keypoints, a cut-out human face and head, CNN feature information corresponding to the face or head, and notable features, and use these as the first feature information.

[0028] The second feature information detection unit 512 performs, for example, well-known image processing on the image data to detect or extract, as second feature information, main features of the human body, such as but not limited to key points of the human body, segmented human bodies and their corresponding CNN features, human body parts, and the skeleton of a human pose. The second feature information detection unit 512 detects or extracts, as second feature information, information on the human body's physical attention (for example, body orientation, posture, hand and arm movements during conversation, etc.). The second feature information detection unit 512 outputs the detected second feature information to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54.

[0029] The third feature information detection unit 513 performs well-known speech recognition processing (e.g., sound source separation, sound source direction estimation, noise suppression, etc.) on the speech signal to detect or extract acoustic feature information, such as, but not limited to, Mel-frequency cepstrum (MFCC) representing intonation, pitch, etc., prosodic features, as third feature information. The third feature information detection unit 513 outputs the detected third feature information for a period similar to consecutive times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54. In this embodiment, paralanguage refers to information conveyed from a speaker to a listener that is outside the context of language, excluding body language such as gestures, and is, for example, the strength and weakness of the voice when speaking, pitch, intonation, etc.

[0030] The fourth feature information detection unit 514 performs, for example, well-known image processing on the image data to detect or extract true feature information of the environment, feature information of objects in the environment, etc. as fourth feature information. The fourth feature information detection unit 514 outputs the detected fourth feature information for times t-2, t-1, and t to the first modeling device 521, the second modeling device 522, the third modeling device 523, the fourth modeling device 524, and the sociological model 54.

[0031] The modeling unit 52 generates a plurality of modeling data to be used for training the sociological model 54 using a plurality of pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information. The modeling data is information related to the attention of multiple people. That is, the modeling unit 52 uses low-level information from the detection unit 51 to generate high-level modeling data that is not limited to gaze behavior, body language, human-object interaction, etc. Each modeler (first modeler 521 to fourth modeler 524) is, for example, a gaze targeting model, a joint attention model, a body language recognition model, etc.

[0032] The first modeling device 521 generates modeling data of the gaze trajectory of each person using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the degree of attention to a common target by most people in the scene. Note that the first modeling device 521 can also output the gaze (or attention) target of everyone in the scene.

[0033] The second modeling device 522 generates modeling data of body language or behavior for each person using multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for a single time t or multiple times t-2, t-1, and t, and outputs the generated modeling data to the sociological model 54 as environmental information indicating the physical behavior (e.g., walking, sitting, etc.) of everyone in the scene.

[0034] The third modeler 523 converts the paralanguage of each person into high-level auditory behavior using one or more pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. This generates modeling data, which is output to the sociological model 54 as environmental information and instructs auditory behavior such as sound source direction estimation, sound source position estimation, and sound source estimation.

[0035] The fourth modeler 524 generates modeling data of the environment and context using the singular points or multiple pieces of feature information (first feature information to fourth feature information) detected by the detection unit 51 for one time t or multiple times t-2, t-1, and t. Then, the fourth modeler 524 outputs the generated modeling data to the sociological model 54 as environmental information not limited to object recognition and how humans interact with these objects, i.e., as human-object interactions.

[0036] The state model 53 is a model for inferring human internal states, such as emotions, intentions, or internal states, such as positive and negative, positive or negative, including but not limited to valence and arousal. Such internal states can also be represented by physiological responses, such as information about electroencephalograms, such as electroencephalogram frequency bands, and information about electrocardiograms, including but not limited to heart rate.

[0037] The sociological model 54 is, for example, a model based on social interaction dynamics, such as a Large Content Behavior Model (LCBM). The sociological model 54 fine-tunes a Large Language Model (LLM) (see, for example, Reference 1) using context extracted from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the text or visual encoders of the modeling unit 52. The fine-tuned LLM model is a Large Context and Behavioral Model (LCBM). To generate behaviors, such as movements and utterances, the output from the LCBM is passed to, for example, a Gaussian Mixture Model and decoded to generate behavior commands for controlling the different degrees of freedom of the robot 3.

[0038] The first generation unit 6 generates nonverbal cues (NOCs) based on the output of the sociological model 54. The nonverbal cues (NOCs) will be described later. The first generation unit 6 inputs the output (e.g., linguistic information) from the large-scale language model of the sociological model 54 to a Gaussian model (e.g., a Gaussian mixture model) as linguistic information and embeddings of an LCBM representing social interaction dynamics. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder to generate nonverbal cue behavioral commands for the robot 3.

[0039] The second generation unit 7 generates speech information (text, audio signals) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the generated speech information to a Gaussian model (e.g., a Gaussian mixture model), inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 from the speech information.

[0040] 2, the detection unit 51 is provided with four feature information detection units (511 to 514), but this is not limiting. The number of feature information detection units may be two or more, and may be five or more. Furthermore, the modeling unit 52 is provided with four modeling devices (first modeling device 521 to fourth modeling device 524), but this is not limiting. The number of modeling devices may be two or more, and may be five or more. In the example shown in FIG. 2, the first generation unit 6 and the second generation unit 7 are provided, but the present invention is not limited to this and it is sufficient that at least one of them is provided.

[0041] In this embodiment, with this configuration, audiovisual data is used to train the modeling unit 52 and sociological model 54. During use, information based on the audiovisual data is input to the model (modeling unit 52, sociological model 54) to estimate interactions between multiple people.

[0042] [Nonverbal Cues (NOC)] Next, we will explain non-verbal cues (NOCs). In order to infer human states in real environments using audiovisual data, it is necessary to understand and represent the dynamics of social interactions using multimodal temporal observations, not limited to non-verbal behaviors and vocal activities, as human states tend to depend on environmental context. Here, non-verbal behavior refers to gestures (body movements), posture, movements (nodding and general behavior), facial expressions, tone of voice (volume, speed, pauses), etc. When a robot communicates with a person, for example, rather than simply outputting an automatically generated voice signal, a communication robot has been proposed that responds to the voice signal by adjusting the tone and strength of the voice, and even by displaying images of the eyes and mouth on a display unit to speak with facial expressions (for example, JP 2022-6610 A). Furthermore, it has been proposed to improve communication with the user by moving the robot's head or other parts to make the user feel as if the robot is drooping or happy, or by having a robot with arms and hands make gestures.

[0043] For this reason, in this embodiment, verbal information and non-verbal information (non-verbal cues, non-verbal behavior) are generated as behavioral commands for the robot 3. Note that in this embodiment, the non-verbal cues include gaze interaction, body language, vocal activities such as paralanguage, speech characteristics, facial expressions, etc.

[0044] [Model] Next, processing using the sociological model 54 will be described. LCBM takes a similar approach to recent models such as the Bootstrapping Language-Image Pre-training (BLIP) model, the Large Language-and-Vision Assistant (Llava) model, and VideoLlaMa (a multi-modal framework that empowers Large Language Models) to understand both image and text content, using a Visual Encoder (EVA-CLIP) to encode images and a Large Scale Language Model (LLM)541 to encode text (see, e.g., Reference 1).

[0045] 3 is a diagram illustrating an example of the configuration and processing of a model according to this embodiment. As shown in FIG. 4, the sociological model 54 includes, for example, a conversion unit 541, an integration unit 542, a large-scale language model 543, and an output unit 544. The sociological model 54 obtains environmental information from the modeling unit 52 . The conversion unit 541 converts environmental information (for example, image data) obtained as non-linguistic information into linguistic information, and uses the environmental information obtained as linguistic information as it is. The integration unit 542 integrates the obtained linguistic information and inputs the integrated information to, for example, a large-scale language model (LLM) 543. The output unit 544 outputs the language information, which is the output from the large-scale language model, to the first generation unit 6 and the second generation unit 7.

[0046] Reference 1;Ashmit Khandelwal, Aditya Agrawal, et al., “Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior”, arXiv:2309.00359 [cs.CL], 2023

[0047] In this embodiment, the output of the sociological model 54 is used to generate verbal and non-verbal behavioral commands by the generation units (first generation unit 6, second generation unit 7), which then control the robot 3. As a result, according to this embodiment, there is no need to wear any special equipment, and it is possible to understand the state (interaction) of multiple people using the visual data detected by the sensor 2. Furthermore, according to this embodiment, it is possible to make the robot 3 behave (speak, gesture, show facial expressions, etc.) using the sociological model 54 and database (state model 53) for inferring people's states and non-verbal behavior, thereby realizing empathetic interaction with people in a real environment.

[0048] [Detection of points of interest] Next, the process of detecting a point of interest will be described. 4 is a diagram for explaining the process of detecting a portion of interest according to this embodiment. Note that the blocks in FIG. 4 correspond to, for example, the detection unit 51 and the first modeling device 521 of the learning device 5.

[0049] The input data is information t indicating whether each of the persons 1 to Np is speaking. 1 ~t Np and the body language information of each person 1 to Np. 1 ~a Np and the location information of each person 1 to Np (Location) 1 ~l Np The input data includes information g indicating the gaze direction of each person. 1 ~g NpThe input information may include at least two of information indicating whether each of the plurality of people is speaking, body language information of each of the plurality of people, position information of each of the plurality of people, and information indicating the line of sight of each of the plurality of people.

[0050] The integration unit 551 integrates the input data of each of the people 1 to Np, and calculates the n-th (n is an integer between 1 and Np) pixel value H n Outputs (552).

[0051] Next, the feature extraction unit 553 extracts the n-th pixel value H n Extract feature (554) from (552) and extract feature F n (554) is input to the transformer encoder 556.

[0052] Next, the joint attention token J (555) is input to the transformer encoder 556.

[0053] Next, the transformer encoder 556 processes multiple inputs in parallel to generate features F for each person. JA 1~Np and joint attention feature information J JA By this process, the interaction generates joint attention feature information J using the features of each person and the features of the joint attention token (e.g., J) that can be learned. JA is embedded in.

[0054] The joint attention estimation unit 559 calculates the joint attention feature information J JA , and a pixel value H of the heat map that indicates a high temperature at the point of interest estimated by the model of the first modeler 521 is obtained. JAを The output is the gaze target (or attention target) (joint attention) X=f1. The gaze target can be, for example, a person speaking, nearby voices, body language, human state, physiological reactions, etc. When in use, the pixel values ​​H of the heat map output by the joint attention estimation unit 559 are JAをis the output of the first modeler, for example.

[0055] During training, the learner 5 estimates the pixel value H JA n and the pixel value G of the nth heatmap of the ground truth data JA n The sum of the squared differences between JA The learning device 5 performs learning until this loss function becomes a minimum or falls within a predetermined value.

[0056]

number

[0057] FIG. 5 is a block diagram showing an example of the overall configuration of the block diagram of FIG. The dashed line rectangle 600 is the processing block at time t-2. The dashed line rectangle 620 is the processing block at time t-1, with the processing and block configuration being the same as at time t-2, and only a portion is shown. The dashed line rectangle 630 is the processing block at time t, with the processing and block configuration being the same as at time t-2, and only a portion is shown.

[0058] The integration unit 601 combines the input data for each of the people 1 to Np and calculates the nth (n is an integer between 1 and Np) pixel value H n Output.

[0059] The FC Layer 602 (Fully Connected Layer) corresponds to the feature extraction unit 553 in FIG. 4, and extracts the n-th pixel value H n Extract features from the extracted feature F n Output.

[0060] At time t-2, the transformer encoder 604 converts the feature F n and the joint attention token J (603) are input, and multiple inputs are processed in parallel to generate the features F JA1~Np and joint attention feature information J JA_(t-2) Similarly, at time t-1, the transformer encoder 604 processes multiple inputs in parallel to generate the features F JA 1~Np and joint attention feature information J JA_(t-1) At time t, the transformer encoder 604 processes multiple inputs in parallel to generate the features F JA 1~Np and joint attention feature information J JA_(t) Outputs (631).

[0061] The LSTM640 (Long Short-Term Memory network) corresponds to the sociological model 54. The LSTM640 contains the joint attention feature information J at time t-2. JA_(t-2) (606), joint attention feature information J at time t-1 JA_(t-1) (621), joint attention feature information J at time t JA_(t) (631) is input. The LSTM 640 and the FC Layer 650 use the joint attention feature information J JA The integrated information is output to Sigmoid660.

[0062] For example, Sigmoid660 is a sigmoid function and an activation function, which is a pixel value H of the heat map that shows high temperatures at points of interest. JA Output.

[0063] The learner 5 estimates the pixel value H JA n and the pixel value G of the nth heatmap of the ground truth data JA n The sum of the squared differences between JA It is calculated as follows.

[0064] With this configuration and processing, the configuration and method of this embodiment are practical for implementation in a real environment because the model relies on audiovisual data. Furthermore, the configuration and method of this embodiment rely on multimodal input, which improves the robustness of the model compared to conventional methods.

[0065] 4 and 5 are merely examples, and are not limiting. For example, other components may be provided, and other processes may be performed.

[0066] As mentioned above, the input of each modality is represented as a different attribute of the subject, and in the model of this embodiment, it is input to the transformer encoder. This integration allows the present embodiment to provide an integrated understanding of the diverse cues present in social interactions and their potential influence on the formation of shared attentional states.

[0067] In this embodiment, the binary classification of spoken information (speaking or not) is input to a transformer encoder to capture the participation state during the group discussion. The gaze angle is expressed as a vector, including both pitch and yaw. These vectorized gaze angles serve as dynamic inputs, providing information about the focus of each participant's gaze direction during the discussion. Body language is also encoded as vectors, capturing multiple different classes of body language. This vectorized representation allows for a nuanced understanding of participants' non-verbal cues and engagement levels.

[0068] Head position is represented by the center coordinates (x, y) of a bounding box around each participant's head. This position information is input into a transformer encoder to contribute to the spatial context of the group discussion.

[0069] The transformer-encoder processes these diverse attributes and harmonizes them into a shared latent space. By embedding information from each modality in the encoder, the model achieves a unified representation of the subject's attentional cues. The attention mechanism within the transformer automatically adjusts the importance of each attribute, recognizing their unique contribution to the joint attention dynamics. The trainable token J is used as one of the inputs of the transformer encoder, which aggregates information from all subjects to take advantage of the context of the entire group discussion.

[0070] This integrated approach recognizes the distinct characteristics of each modality, allowing the present embodiment to synthesize these inputs, whether capturing speaking state, gaze direction, subtle body language, or spatial head position, to comprehensively understand and estimate joint attention during complex group discussions.

[0071] In this embodiment, we also extend the framework of the joint attention model by incorporating an LSTM layer to model ongoing temporal awareness in a multi-person group discussion, as shown in Figures 4 and 5. In this embodiment, the inclusion of an LSTM allows the model to consider information from three consecutive frames (e.g., t-2, t-1, t), providing more dynamic and contextual predictions. The output from the transformer encoder represents the fused information from various modalities for each subject and serves as the input to the subsequent LSTM module. A multi-layer LSTM module, consisting of, for example, three layers, processes the input sequence, allowing the model to capture complex patterns of joint attention across three consecutive frames. Note that while the example shown in Figure 5 uses three consecutive sequences, the sequence used may be four or more consecutive sequences.

[0072] Then, as mentioned above, following the LSTM, the model passes through three fully connected layers, with the final layer employing sigmoid activation to generate a heatmap representing the predicted joint attention distribution for the current frame. Note that this model is trained using supervised learning using the ground-truth joint attention heatmap. In this embodiment, the incorporation of multi-layer LSTMs allows the model to enhance its temporal modeling capabilities and provide a more sophisticated understanding of the temporal dynamics inherent in joint attention during group discussions. Furthermore, in this embodiment, we introduce a mean squared error loss function to facilitate the alignment of the predicted heatmap with the true annotations and to guide the learning process. Note that the loss L JA is the estimated heatmap of the model H JA n and the ground truth heatmap G JA n It is calculated by evaluating the mean squared error between

[0073] As described above, in this embodiment, the following is included to perform "joint attention estimation taking interaction into consideration." I. Audio input (i.e., audio activity) was used. II. Multimodal visual input allows us to sense non-verbal cues relevant to multi-party interactions. III. Time dependency is included in the model (modeling part 52, sociological model 54). The model then uses multimodal audiovisual input and human state (potentially derived from audiovisual observation) to output a single joint attention output, which may be the most salient location in the scene.

[0074] In addition, in this embodiment, the following configuration is provided and processing is performed. I. The internal stimuli were represented using the temporal information of the LSTM layer of the model. II. The external stimulus was represented by multimodal input derived from the audiovisual input of all subjects in the frame. III. Non-verbal cues include gaze interactions, body language, vocal activities such as paralanguage, speech characteristics, facial expressions, etc. Human physiological responses may also be used to infer joint attention, complementing non-verbal cues.

[0075] Furthermore, the configuration and method of this embodiment are practical for implementation in a real environment because the sociological model 54 relies on audiovisual data. Furthermore, the configuration and method of this embodiment rely on multimodal input, making it possible to improve the robustness of the sociological model 54 compared to conventional methods. The sociological model 54 is a model that represents the dynamics of social interactions. For example, this is an LCBM model. The joint attention model is one of the models of the sociological model 54, namely, model 521.

[0076] As a result, this embodiment allows for the estimation of visual behavior by taking into account the influence of internal and external stimuli. Furthermore, this embodiment is practical for implementation in real environments because the sociological model 54 relies on audiovisual data. Furthermore, this embodiment relies on multimodal input, making it possible to improve the robustness of the sociological model 54 compared to conventional methods.

[0077] [Processing Procedure] Next, an example of a processing procedure performed by the control device 4 will be described. First, an example of a procedure for training a model (modeling unit 52, sociological model 54) will be described. Fig. 6 is a flowchart of a processing procedure for training a model according to this embodiment.

[0078] (Step S101) The detection unit 51 acquires the detection data detected by the sensor 2.

[0079] (Step S102) The detection unit 51 acquires relationship data (for example, information about relationships between people).

[0080] (Step S103) The detection unit 51 detects the feature information (first feature information to fourth feature information) from the detection data and the relational data.

[0081] (Step S104) The modeling unit 52 converts the information about the environment included in the feature information (information based on the relational data) into text.

[0082] (Step S105) The modeling unit 52 receives environmental information from the sound pickup device 21, the image capture device 22, the detection unit 51, and the modeling unit 52. These inputs can also be encoded as features to better represent multimodal temporal interaction dynamics. The output is linguistic information and embedded representations, or linguistic information or embedded representations.

[0083] (Step S106) The modeling unit 52 trains the generated linguistic information and embedded expression, linguistic information, or embedded model.

[0084] In this embodiment, the above steps S104 to S106 are referred to as a modeling data generation process.

[0085] Next, an example of a procedure for making an estimation using the trained model (modeling unit 52, sociological model 54) will be described. Fig. 7 is a flowchart of the processing procedure when using the model according to this embodiment.

[0086] (Step S201) The detection unit 51 acquires the detection data detected by the sensor 2.

[0087] (Step S202) The detection unit 51 detects feature information (first feature information to fourth feature information) from the detection data.

[0088] (Step S203) The modeling unit 52 assigns identification information to a plurality of people.

[0089] (Step S204) The modeling unit 52 converts environmental information (information based on relational data) included in the feature information into text. The modeling unit 52 receives environmental information from the sound pickup device 21, the image pickup device 22, the detection unit 51, and the modeling unit 52. The modeling unit 52 then outputs the linguistic information and the embedded expression, or the linguistic information or the embedded expression, to the sociological model 54.

[0090] (Step S205) The sociological model 54 receives information output from the first modeling device 521 to the fourth modeling device 524 for one time or multiple consecutive times, and converts the information into linguistic information and embedded expressions, or linguistic information or embedded expressions.

[0091] (Step S206) The first generation unit 6 generates a non-verbal cue (NOC) using the output of the sociological model 54. Furthermore, the first generation unit 6 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the non-verbal cue. Furthermore, the second generation unit 7 generates speech information (text, audio signal) based on the output of the sociological model 54 using a method similar to that of the first generation unit 6. Furthermore, the second generation unit 7 inputs the output from the Gaussian model to a decoder, and creates a behavioral command for the robot 3 based on the speech information.

[0092] (Step S207) The first generating unit 6 and the second generating unit 7 each output the generated action instructions to the robot 3 to control the action of the robot 3.

[0093] 6 and 7 are merely examples, and are not limiting. The control device 4 may, for example, perform several processes in parallel.

[0094] In the above embodiment, the modeling units of the modeling unit 52 are TGN models, but the present invention is not limited to this. The modeling units (first modeling unit 521 to fourth modeling unit 524) may be other models, and the models of the respective modeling units may be different. Furthermore, although the sociological model 54 has been described as an example of an LCBM, the sociological model 54 may be another model.

[0095] [Evaluation results] Next, an example of the evaluation results will be described. FIG. 8 is a diagram showing an example of an evaluation result. In FIG. 8, reference symbol g601 indicates a correct joint attention area (the focus point of multiple people), and reference symbol g602 indicates a predicted joint attention area. Furthermore, the text indicated by reference symbol g604 is an example in which text indicating that a person (or object) other than the person being watched is displayed around the person (or object) other than the person being watched, along with a person who is not speaking. Furthermore, the text indicated by reference symbol g603 is an example in which text indicating that a person is clapping their hands (or making a gesture like that) while talking is displayed around the person (or object) being watched.

[0096] As shown in Fig. 8, the prediction is correct when the distance between the correct joint gaze area and the predicted joint gaze area is within a threshold defined according to the size of the face bounding box. In this embodiment, the center point of the head bounding box is defined as the correct position and the estimated position.

[0097] 9 is a diagram showing an example of evaluation results for thresholds. The thresholds are 30, 60, 90, and 100. The number of correctly estimated instances (Count) can be expressed as in the following equation (2).

[0098]

number

[0099] The detection rate (DR) was calculated using the following formula (3).

[0100]

number

[0101] The model was trained using supervised learning to optimize the mean squared error (MSE) loss function, and the model was fine-tuned over training epochs using a learning rate schedule with an initial learning rate of 0.001.

[0102] As shown in FIG. 9, compared to the cases where LSTM is not used (first comparative example and second comparative example), better evaluation results were obtained for all thresholds in this embodiment where LSTM is used. The first line in Figure 9, "Configuration of the embodiment," is a case where an LSTM is provided and four inputs (speech (t), gaze (q), action (a), and location (l)) for each person are received. The second line does not use an LSTM, and J JAt Use only (in this case, J JAt-1 and J JAt-2 The third row shows the case without LSTM and excluding gaze angle. The fourth row shows the case with LSTM and three inputs for each person: speech (t), action (a), and location (l).

[0103] As shown by these evaluation results, this embodiment makes it possible to estimate social gaze saliency by combining temporal attributes such as the subject's voice activity, behavior, and location during multiparty facilitation.

[0104] In addition, a program for implementing some or all of the functions of the control device 4 in the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the control device 4. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. Alternatively, some or all of these components may be realized by LSI (Large Scale Integration) hardware (including circuitry) such as an ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), or SOC (System On Chip), or may be realized by a combination of software and hardware.

[0105] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0106] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0107] 1...communication system, 2...sensor, 3...robot, 4...control device, 5...learning device, 6...first generation unit, 7...second generation unit, 50...visual behavior estimation device, 21...sound pickup device, 22...image capture device, 23...communication unit, 31...control unit, 32...communication unit, 51...detection unit, 52...modeling unit, 53...state model, 54...sociological model, 511...first feature information detection unit, 512...second feature information detection unit, 513...third feature information detection unit, 514...fourth feature information detection unit, 521...first modeling device, 522...second modeling device, 523...third modeling device, 524...fourth modeling device, g21...first module, g22...second module, g23...third module, g24...fourth module, g25...fifth module, g26...sixth module

Claims

1. a detection unit that acquires a plurality of sensing data; an estimation unit that connects a plurality of pieces of sensing data that are multimodal data, combines the connected plurality of pieces of sensing data with a learnable token, integrates the connected data at different times, and estimates a location where a plurality of people are paying attention using the integrated data; A visual behavior estimation device comprising:

2. The estimation unit The connected plurality of sensing data are input to a fully connected layer; inputting the concatenated plurality of sensing data and the trainable token into a transformer encoder and combining them; The combined data from different times is integrated using a Long Short-Term Memory (LSTM) network to take into account temporal dynamics; The data integrated by the long short-term memory network is passed through a fully connected layer, and then an activation function is used to estimate the locations where the plurality of people are paying attention. The visual behavior inference device according to claim 1 .

3. The plurality of sensing data At least two of the following information are included: information indicating whether a person is speaking; body language information of the person; and position information of either the person or an object. The visual behavior estimation device according to claim 1 or 2.

4. The detection unit acquires a plurality of sensing data, An estimation unit connects the plurality of sensing data, which are multimodal data, combines the connected plurality of sensing data with a learnable token, integrates the connected data at different times, and estimates the location where a plurality of people are paying attention using the integrated data. Visual action estimation methods.

5. The computer of the visual behavior estimation device Acquire multiple sensing data, A plurality of pieces of sensing data, which are multimodal data, are linked together, the linked pieces of sensing data are combined with a learnable token, the linked pieces of data at different times are integrated, and the integrated data is used to estimate the location where a plurality of people are paying attention. program.