Machine-Learning Image Processing for Main-Subject Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In capturing scenes where multiple subjects are moving in the same direction, such as athletics events or races, it is challenging to consistently determine the positional relationship and identify the main subject due to varying angles of view and photographer-subject distances, making it difficult to track a specific subject like the front.
Innovation Solution
An image processing apparatus and method using a trained machine-learning model to detect the positional relationship and decide a main subject region based on the reliability of the acquired information, without requiring constant angles of view or photographer-subject distances, by employing a digital single-lens reflex camera with components like a computing device, image sensor, and machine-learning models like CNN for subject detection and tracking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If movement direction detection is used to identify the front subject, then subject tracking capability is improved, but detection reliability deteriorates when angle of view or photographer-subject distance varies
Solution Approach 1:
The patent introduces a trained machine-learning model as an intermediary between the input image and the subject detection process. This model serves as a mediator that processes the image data and outputs reliable positional relationship information even when viewing conditions vary, thereby resolving the contradiction between adaptability and reliability.
Solution Approach 2:
The patent replaces traditional mechanical or rule-based movement direction detection methods with a machine-learning-based system. This substitution enables the system to maintain high detection reliability across varying angles and distances by learning from training data rather than relying on fixed geometric relationships.
2Device complexity
If traditional movement direction detection is used, then system complexity is reduced, but measurement precision of positional relationship deteriorates
Solution Approach 1:
The patent changes the fundamental parameters of the detection system by transitioning from simple geometric calculations to a machine-learning model that processes complex image features. This parameter change enables high measurement precision of positional relationships while accepting increased system complexity in the form of trained models and computational resources.
3Measurement precision
If machine-learning model is introduced to detect positional relationship, then detection precision is improved, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by training the machine-learning model in advance with labeled data before deployment. This preliminary training phase allows the model to learn complex patterns and relationships, enabling high detection precision during actual operation without requiring complex real-time processing logic.
Solution Approach 2:
The patent uses copying by training the machine-learning model on a large dataset of example images with known positional relationships. The model learns to replicate the positional relationships from training examples, enabling accurate detection in new situations without requiring explicit programming for each scenario.
Data Source
AI summary
An image processing apparatus acquires, from an input image based on image data capturing a scene in which a plurality of subjects are moving in a substantially same direction, information regarding a positional relationship of the plurality of subjects in the direction, with use of a trained machine-learning model. In a case where it is determined, based on a reliability of the acquired information, to decide the main subject region, the apparatus decides a main subject region from among subject regions respectively corresponding to the plurality of subjects in the image data, based on the acquired information.


