Action Recognition for Precise Video Scene Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video recording technologies make it difficult for users to efficiently search and view specific scenes within recorded moving images, as users must often fast-forward or skip through unnecessary footage to find desired moments, leading to boredom and inefficiency.

Innovation Solution

An information processing device and method that recognizes the actions of a photographer or subject during image capture using sensors, allowing users to associate and select specific actions for targeted reproduction, thereby enabling precise scene selection and playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If users manually fast-forward or skip through recorded moving images to find specific scenes, then they can view the entire recorded content, but the time required to find desired scenes increases significantly and user experience deteriorates

Engineering Contradiction:
Improvetime to find specific sceneVSAvoidease of scene selection
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The system performs preliminary analysis of the recorded moving image content before reproduction, automatically detecting and tagging specific scenes (such as faces, objects, or actions) during the recording phase. This preliminary processing enables users to directly access desired scenes without manual fast-forwarding, resolving the contradiction by preparing the data structure in advance to facilitate rapid scene retrieval.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary indexing system that acts as a mediator between the raw video content and user search queries. This intermediary layer automatically generates scene descriptors and metadata during recording, allowing users to search for and navigate to specific scenes through keywords or tags rather than manually scanning through the entire video, thus reducing time loss while maintaining ease of operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If action recognition technology is implemented to enable precise scene selection, then scene search efficiency improves, but device complexity increases

Engineering Contradiction:
Improvescene search efficiencyVSAvoidcomplexity of action recognition system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent integrates action recognition functionality into existing video recording devices by utilizing their existing sensors (accelerometers, gyroscopes, microphones) and processing units for multiple purposes: both standard video recording and action/scene recognition. This multi-functionality approach enables scene search efficiency improvement without requiring entirely separate dedicated hardware, thus limiting the increase in device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs self-service mechanisms where the recording device automatically performs action recognition and scene tagging without requiring external processing equipment. The device uses its own sensors and processors to analyze recorded content and generate searchable metadata, eliminating the need for complex external analysis systems and keeping the overall system complexity manageable while maintaining high scene search efficiency.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS7917020B2Information processing device and method, photographing device, and program
Publication Date: 2011.03.29 SONY GROUP CORP
  • US7917020B2 patent drawing
  • US7917020B2 patent drawing
  • US7917020B2 patent drawing

AI summary

An information processing device includes a processing unit that recognizes an action based on sensor data output by a sensor and obtained in a same timing as a photographing timing of an image string. The processing unit associates information indicating the action as information for a selection of a reproduction position at a time of a reproduction of the image string with the image string. In addition, the processing unit extracts a plurality of features from the sensor data. The processing unit recognizes the action based on a time series of the plurality of features using a model for a recognition.