Latent Space Decomposition for Overlapping Action Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods struggle to effectively recognize patterns in untrimmed videos where multiple actions overlap, as they are not well-captured by existing techniques, leading to challenges in action recognition and segmentation.

Innovation Solution

A generation module that encodes input signals into a latent space, decomposes them into pattern-related and unrelated components, combines pattern-related components from multiple signals, and decodes them to generate a composed output signal, enabling better detection and segmentation of complex patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If current methods are used for pattern recognition in untrimmed videos, then the system is simple to operate, but the reliability of pattern detection deteriorates when multiple actions overlap

Engineering Contradiction:
Improvepattern detection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the pattern recognition task into distinct components: encoding input signals into latent representations, decomposing these representations into action-specific and non-action-specific components, and separately processing overlapping actions. This segmentation enables reliable detection of multiple simultaneous actions by treating them as independent detectable entities rather than a confused mixture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces latent space representations as an intermediary between raw input signals and final pattern recognition. This intermediate representation layer transforms complex overlapping action signals into disentangled latent components, making it easier to reliably detect and separate multiple simultaneous actions without direct complexity in the final recognition system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If current methods are used for action recognition, then the device complexity is low, but the measurement precision of overlapping actions deteriorates

Engineering Contradiction:
Improveaction recognition precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the action recognition process into encoding, decomposition, and detection stages. By decomposing latent representations into action-specific components, the system achieves precise measurement of overlapping actions through structured processing rather than direct analysis, improving precision while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from analyzing raw input signals directly to operating in a latent space dimension. This dimensional transformation allows for better separation and precise measurement of overlapping actions by representing them in a transformed feature space where their distinct characteristics are more easily distinguished and measured.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If pattern segmentation is performed on untrimmed videos with overlapping actions, then the productivity is reduced due to processing complexity, but the reliability of multi-label detection can be improved

Engineering Contradiction:
Improvemulti-label detection reliabilityVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary encoding of input signals into latent representations and pre-decomposition into action-specific components before the actual pattern detection and segmentation tasks. This preliminary processing organizes the data in advance, making subsequent multi-label detection more reliable and efficient by presenting pre-processed, disintangled features to the detection algorithms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuous processing through the encoding-decomposition-detection pipeline, ensuring that useful information is preserved and transformed continuously rather than through discrete, interruptive steps. This continuous transformation through latent space representations maintains processing efficiency while improving detection reliability through systematic feature disentanglement.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4468253A1Generation module and method for generating a composed output signal
Publication Date: 2024.11.27 TOYOTA JIDOSHA KK
  • EP4468253A1 patent drawingFigure 1~2(IV)
  • EP4468253A1 patent drawingFigure 3
  • EP4468253A1 patent drawing

AI summary

A generation module (10) for generating a composed output signal (pmm',c, pmm',c') from two or more input signals (pm,c,pm',c') representing respective patterns, the generation module being configured to: - encode (i) the two or more input signals (pm,c,pm',c') onto a latent space; - decompose (ii) the encoded input signals into a first component (rm, rm') along one or more pattern-related directions (dm1,dm2,...,dmJ) of the latent space, and a second component (rc, rc') along one or more pattern-unrelated directions (dc1,dc2,...,dcK) of the latent space; - generate (iv) a latent output signal having, as the first component, a combination of the respective first components (rm, rm') of at least two of the encoded input signals; - decode (iv) the latent output signal to obtain the composed output signal (pmm',c,pmm',c')..