Region Extraction Model Learning Apparatus Using Composited Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning technologies face challenges in achieving high extraction precision for person regions in images, particularly in unusual poses, due to the high cost and effort required in preparing specialized learning data with consistent backgrounds, which is prohibitive for applications like sports videos.
Innovation Solution
A region extraction model learning device generates composited learning data by combining existing learning data with background images using parameters for enlargement, translation, and rotation, creating new images and masks to reduce preparation costs and improve extraction precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized learning data with consistent backgrounds is prepared for deep learning, then extraction precision is improved, but preparation cost and effort increase prohibitively
Solution Approach 1:
The patent uses image synthesis technology to generate synthetic learning data by copying and compositing person regions from existing images onto standardized background templates. This creates artificial but realistic training images without requiring actual physical setup or manual image collection, dramatically reducing data preparation costs while maintaining extraction precision.
Solution Approach 2:
The patent applies geometric transformation parameters (scaling, rotation, translation) to person regions and adjusts composition parameters to vary background conditions. By changing these parameters systematically, diverse training samples are generated from limited source images, improving model generalization without proportional increases in data collection effort.
2Measurement precision
If images with same backgrounds are prepared for learning data to improve extraction precision, then extraction precision is improved, but preparation effort becomes extremely costly
Solution Approach 1:
Instead of manually collecting images with consistent backgrounds, the patent copies person regions from various source images and composites them onto standardized background templates. This automated copying process eliminates the time-consuming manual image collection and background matching process while ensuring consistent background conditions across all training samples.
Solution Approach 2:
The patent prepares standardized background templates and person region libraries in advance. These pre-processed elements can be quickly combined through parameter adjustment to generate diverse training samples, eliminating the need for time-consuming real-time image processing and background matching during data preparation.
3Adaptability or versatility
If deep learning is used for semantic segmentation, then region extraction capability is improved, but requirement for large amounts of learning data increases preparation complexity
Solution Approach 1:
The patent uses automated image synthesis to generate synthetic training data by copying person regions and compositing them onto standardized backgrounds. This approach provides the large volume of diverse training data required by deep learning models without the manual effort of collecting and annotating real images, simplifying the data preparation process while maintaining extraction versatility.
Solution Approach 2:
The patent creates a universal framework for generating training data that can accommodate various person poses, backgrounds, and conditions through parameter adjustment. This single synthetic data generation system replaces the need for multiple specialized data collection processes, reducing preparation complexity while maintaining broad extraction capability across different scenarios.
Data Source
AI summary
Provided is technology for extracting a person region from an image, that can suppress preparation costs of learning data. Included are a composited learning data generating unit that generates, from already-existing learning data that is a set of an image including a person region and a mask indicating the person region, and a background image to serve as a background of a composited image, composited learning data that is a set of a composited image and a compositing mask indicating a person region in the composited image, and a learning unit that learns model parameters using the composited learning data. The composited learning data generating unit includes a compositing parameter generating unit that generates compositing parameters that are a set of an enlargement factor, a degree of translation, and a degree of rotation, using the mask of the learning data, and a composited image and compositing mask generating unit that extracts a compositing person region from an image in the learning data using a mask of the learning data, generates the composited image from the background image and the compositing person region, using the compositing parameters, generates the compositing mask from a mask generating image and the compositing person region that are the same size as the composited image, using the compositing parameters, and generates the composited learning data.


