Multi-view Action Recognition via 3D Pose Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing action recognition systems face challenges in efficiently analyzing human actions captured by camera videos, particularly in complex environments with varying camera positions and angles, leading to unreliable and time-consuming manual monitoring.

Innovation Solution

A system that uses one or more processors to obtain multiple videos of subjects, track target subjects across multiple cameras, reconstruct a 3D model of the target subjects, and recognize their actions based on the reconstructed 3D poses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If multiple cameras are used to capture videos in complex environments, then the coverage and detection capability are improved, but the complexity of tracking and recognizing actions across multiple views increases

Engineering Contradiction:
Improveaction recognition accuracyVSAvoidmulti-camera system complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent transforms 2D video data from multiple cameras into 3D pose information by reconstructing spatial coordinates of body keypoints. This dimensional transformation enables the system to resolve ambiguities in action recognition by adding depth information, thereby improving detection accuracy while managing the complexity of multi-camera coordination through geometric reconstruction algorithms.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces 3D pose reconstruction as an intermediary representation between raw multi-camera video data and final action recognition. This intermediate 3D skeletal model serves as a unified representation that simplifies the integration of data from multiple camera views and provides a standardized format for action classification, reducing the overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual monitoring is used to identify human actions in videos, then the system is simple to implement, but the process becomes time-consuming and unreliable

Engineering Contradiction:
Improveaction identification speedVSAvoidmanual monitoring reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements an automated system that performs action recognition without human intervention. The system automatically processes video data from multiple cameras, reconstructs 3D poses, and classifies actions using machine learning models. This self-service approach eliminates the time-consuming and unreliable manual monitoring process while maintaining simplicity through automated pipelines.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual video review with an automated computational system. Instead of human observers watching and interpreting videos, the system uses computer vision algorithms to automatically extract features, reconstruct poses, and classify actions, thereby dramatically improving both speed and reliability of action identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If deep learning techniques are used for action recognition, then the accuracy is improved, but the requirement for additional training data increases

Engineering Contradiction:
Improveaction recognition accuracyVSAvoidtraining data requirement
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs 3D pose reconstruction as a preliminary step before action recognition. By pre-processing the video data into structured 3D skeletal representations, the system creates a standardized and enriched input format that contains explicit spatial and temporal information. This preliminary transformation reduces the amount of raw training data needed because the critical features are already extracted and organized, allowing deep learning models to learn more efficiently from compact representations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12299929B2Multi-view multi-target action recognition
Publication Date: 2025.05.13 SONY GROUP CORP
  • US12299929B2 patent drawing
  • US12299929B2 patent drawing
  • US12299929B2 patent drawing

AI summary

Implementations generally perform robust multi-view multi-target action recognition using reconstructed 3-dimensional (3D) poses. In some implementations, a method includes obtaining a plurality of videos of a plurality of subjects in an environment, where at least one target subject of the plurality of subjects performs one or more actions in the environment. The method further includes tracking the at least one target subject across at least two cameras. The method further includes reconstructing a 3-dimensional (3D) model of the at least one target subject based on the plurality of videos and the tracking of the at least one target subject. The method further includes recognizing the one or more actions of the at least one target subject based on the reconstructing of the 3D model.