Robot Social Group Detection Using Deep Learning Skeletons

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting F-formations in social groups are largely rule-based, leading to accuracy challenges in dynamic environments and fail when parts of the group are occluded, and they rely on head pose and orientation without considering temporal information.

Innovation Solution

A processor-implemented method using a deep learning-based model to identify human body skeletons, predict key-points, associate confidence scores, and utilize a conditional random field (CRF) and multi-class Support Vector Machine (SVM) with a Gaussian Radial Basis Function (RBF) kernel to detect social groups and predict approach angles for a robot to join the group effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If rule-based methods are used for detecting F-formations, then the system is simpler to implement, but the accuracy deteriorates in dynamic environments and under occlusion

Engineering Contradiction:
Improvesystem complexityVSAvoiddetection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces rule-based mechanical detection systems with a deep learning-based computational model. The deep learning model processes visual data to detect F-formations, substituting traditional mechanical/rules-based approaches with intelligent computational methods that achieve higher accuracy in dynamic environments while maintaining manageable system complexity through automated learning processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the detection parameters from fixed rules to learned parameters through deep learning. The model adapts to different environmental conditions and occlusion scenarios by learning patterns from data, transforming the detection approach from static rules to dynamic, data-driven parameters that can handle variability in real-world scenarios.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If head pose and orientation are used for detection, then the method is simpler, but temporal information is ignored leading to reduced accuracy

Engineering Contradiction:
Improvedetection method complexityVSAvoiddetection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements continuous temporal analysis through deep learning models that process sequences of data over time. The system continuously monitors and analyzes F-formation patterns, capturing temporal dynamics and changes in group configurations, thereby maintaining detection accuracy across varying time conditions and dynamic social interactions.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent adds the temporal dimension to the detection problem by transitioning from static head pose analysis to dynamic sequence-based detection. The deep learning model processes temporal sequences of spatial data, introducing time as an additional dimension that enables the system to capture evolution of social interactions and improve detection accuracy in dynamic environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If deep learning-based models are used to identify human skeletons and predict key-points, then detection accuracy improves, but computational complexity and processing time increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex detection task into multiple specialized components: deep learning models for skeleton identification, separate modules for key-point prediction, and distinct processing stages for different detection objectives. This segmentation allows each component to be optimized independently, managing computational complexity while maintaining high overall accuracy through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3929803B1System and method for enabling robot to perceive and detect socially interacting groups
Publication Date: 2023.06.28 TATA CONSULTANCY SERVICES LTD
  • EP3929803B1 patent drawingFigure 1A~2
  • EP3929803B1 patent drawingFigure 3
  • EP3929803B1 patent drawingFigure 4

AI summary

This disclosure relates to system and method for enabling a robot to perceive and detect socially interacting groups. Various known systems have limited accuracy due to prevalent rule-driven methods. In case of few data-driven learning methods, they lack datasets with varied conditions of light, occlusion, and backgrounds. The disclosed method and system detect the formation of a social group of people, or, f-formation in real-time in a given scene. The system also detects outliers in the process, i.e., people who are visible but not part of the interacting group. This plays a key role in correct f-formation detection in a real-life crowded environment. Additionally, when a collocated robot plans to join the group it has to detect a pose for itself along with detecting the formation. Thus, the system provides the approach angle for the robot, which can help it to determine the final pose in a socially acceptable manner.