Dynamic Image Window for Action Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image recognition techniques struggle to accurately determine the appropriate range of a moving image for estimating the action of a subject, which is crucial for precise action recognition in moving images.
Innovation Solution
An image processing apparatus and method that acquires and processes a sequence of image data to specify the appropriate range of a moving image around the timing of a still image capture, using an estimation model to add information indicating the subject's action to the still image data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a fixed range of moving image data is used for action estimation, then the processing is simple, but the action recognition accuracy is insufficient
Solution Approach 1:
The patent applies dynamics by making the range of moving image data dynamic rather than fixed. The specification unit dynamically adjusts the time range of image data to be processed based on the timing of still image capture, selecting an appropriate window of video frames that includes the capture moment. This allows the system to adapt the data range to each capture event, improving action recognition accuracy without requiring complex manual configuration.
Solution Approach 2:
The patent changes the parameter of the time range for processing image data. Instead of using a predetermined fixed range, the system adjusts the time window parameter dynamically based on the still image capture timing. The specification unit determines the appropriate time range (e.g., several seconds before and after capture) to ensure the action context is fully captured, thereby improving recognition accuracy while keeping the adjustment mechanism relatively simple.
2Measurement precision
If a larger range of moving image data is processed, then the action estimation accuracy improves, but the processing time increases
Solution Approach 1:
The patent applies preliminary action by pre-defining the time range window for processing. The specification unit establishes a predetermined time window (e.g., 2 seconds before and after capture) that is sufficient to capture the action context. This preliminary setup allows the system to process only the necessary data range, avoiding unnecessary processing of excessive frames while still capturing enough context for accurate action estimation.
Solution Approach 2:
The patent uses partial action by processing only the essential portion of the moving image data needed for action recognition. Instead of processing the entire video sequence or an excessively large time range, the system selectively processes a specific time window around the capture moment that contains the relevant action information. This partial processing approach achieves sufficient accuracy while minimizing processing time.
3Loss of information
If the time range for processing image data is extended, then the action context is better captured, but the data processing load increases
Solution Approach 1:
The patent applies extraction by selecting and extracting only the relevant time window of image data from the complete moving image sequence. The specification unit extracts the specific时间段 (time range) that contains the action context needed for recognition, excluding unnecessary data before and after this window. This extraction approach ensures complete action context is captured while keeping the processed data volume manageable.
Data Source
AI summary
An image processing apparatus acquires a plurality of pieces of image data sequentially outputted from an imager and, in accordance with reception of an image capturing instruction to capture a still image, specify as image data to be processed a plurality of pieces of image data in a period that includes a timing at which the still image is captured. The image processing apparatus, based on an action of a subject estimated using the image data to be processed, add information that indicates the action of the subject to data of the still image.


