Skeleton Graph Multi-View Behavior Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human behavior recognition systems face challenges in accurately identifying interactive behaviors due to noise interference and the complexity and diversity of skeleton data, which affects the performance of graph convolution models.

Innovation Solution

A method and system that utilize multi-view comparison by adaptively deleting nodes and edges in skeleton spatio-temporal graphs, applying an information bottleneck principle to maximize relevant information for behavior recognition, and integrating different views into a compact representation for improved robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If graph convolution models are used to recognize interactive behaviors, then behavior recognition capability is provided, but noise interference from sensor errors and occlusions degrades recognition accuracy

Engineering Contradiction:
Improveinteractive behavior recognition capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the skeleton graph into multiple views by selectively deleting different nodes and edges to create node-dropping views and edge-dropping views. This segmentation allows the model to process different aspects of the interaction separately, reducing the impact of noise in any single view while maintaining comprehensive behavior recognition capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating different views with different levels of noise resistance. Node-dropping views and edge-dropping views have different local characteristics that make them resistant to different types of noise, allowing the system to maintain reliable recognition under various noise conditions.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If graph convolution models process diverse skeleton data from different people, then comprehensive behavior coverage is achieved, but data distribution diversity causes model bias and hinders learning

Engineering Contradiction:
Improvebehavior coverageVSAvoidlearning accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms the diverse skeleton data into multiple views by adding the dimension of view variation through node and edge deletion. This dimensional transformation allows the model to learn from different perspectives of the same behavior, reducing bias from data distribution diversity while maintaining comprehensive behavior coverage.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the parameters of the skeleton graph by selectively deleting nodes and edges to create different views. This parameter variation allows the model to learn robust behavior representations that are invariant to specific parameter configurations, improving learning accuracy across diverse data distributions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240395067A1Method and system for identifying human interactive behavior based on multi-view comparison
Publication Date: 2024.11.28 SHANDONG NORMAL UNIV
  • US20240395067A1 patent drawing
  • US20240395067A1 patent drawing
  • US20240395067A1 patent drawing

AI summary

A method and a system for identifying human interactive behavior based on multi-view comparison are provided, belonging to the field of computer vision. The method includes acquiring position information of human joints in each frame of video data; constructing a skeleton spatio-temporal graph based on the position information of the human joints in each frame; based on skeleton spatio-temporal graph, adaptively deleting an edge or a node of the skeleton spatio-temporal graph through a graph convolution neural network, and constructing enhanced node-dropping and edge-dropping views; using an information bottleneck principle, increasing a difference between the enhanced view and an original skeleton spatio-temporal graph, simultaneously maximizing information related to a behavior recognition task, and reserving minimum enough information for the behavior recognition task in each view to obtain a multi-view representation; and obtaining a human interactive behavior recognition result based on the multi-view representation.