Region Extraction Model Learning Apparatus Using Composited Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning technologies face challenges in achieving high extraction precision for person regions in images, particularly in unusual poses, due to the high cost and effort required in preparing specialized learning data with consistent backgrounds, which is prohibitive for applications like sports videos.

Innovation Solution

A region extraction model learning device generates composited learning data by combining existing learning data with background images using parameters for enlargement, translation, and rotation, creating new images and masks to reduce preparation costs and improve extraction precision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specialized learning data with consistent backgrounds is prepared for deep learning, then extraction precision is improved, but preparation cost and effort increase prohibitively

Engineering Contradiction:
Improveextraction precisionVSAvoiddata preparation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses image synthesis technology to generate synthetic learning data by copying and compositing person regions from existing images onto standardized background templates. This creates artificial but realistic training images without requiring actual physical setup or manual image collection, dramatically reducing data preparation costs while maintaining extraction precision.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies geometric transformation parameters (scaling, rotation, translation) to person regions and adjusts composition parameters to vary background conditions. By changing these parameters systematically, diverse training samples are generated from limited source images, improving model generalization without proportional increases in data collection effort.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If images with same backgrounds are prepared for learning data to improve extraction precision, then extraction precision is improved, but preparation effort becomes extremely costly

Engineering Contradiction:
Improveextraction precisionVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of manually collecting images with consistent backgrounds, the patent copies person regions from various source images and composites them onto standardized background templates. This automated copying process eliminates the time-consuming manual image collection and background matching process while ensuring consistent background conditions across all training samples.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent prepares standardized background templates and person region libraries in advance. These pre-processed elements can be quickly combined through parameter adjustment to generate diverse training samples, eliminating the need for time-consuming real-time image processing and background matching during data preparation.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If deep learning is used for semantic segmentation, then region extraction capability is improved, but requirement for large amounts of learning data increases preparation complexity

Engineering Contradiction:
Improveregion extraction capabilityVSAvoidlearning data preparation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses automated image synthesis to generate synthetic training data by copying person regions and compositing them onto standardized backgrounds. This approach provides the large volume of diverse training data required by deep learning models without the manual effort of collecting and annotating real images, simplifying the data preparation process while maintaining extraction versatility.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates a universal framework for generating training data that can accommodate various person poses, backgrounds, and conditions through parameter adjustment. This single synthetic data generation system replaces the need for multiple specialized data collection processes, reducing preparation complexity while maintaining broad extraction capability across different scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11816839B2Region extraction model learning apparatus, region extraction model learning method, and program
Publication Date: 2023.11.14 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11816839B2 patent drawing
  • US11816839B2 patent drawing
  • US11816839B2 patent drawing

AI summary

Provided is technology for extracting a person region from an image, that can suppress preparation costs of learning data. Included are a composited learning data generating unit that generates, from already-existing learning data that is a set of an image including a person region and a mask indicating the person region, and a background image to serve as a background of a composited image, composited learning data that is a set of a composited image and a compositing mask indicating a person region in the composited image, and a learning unit that learns model parameters using the composited learning data. The composited learning data generating unit includes a compositing parameter generating unit that generates compositing parameters that are a set of an enlargement factor, a degree of translation, and a degree of rotation, using the mask of the learning data, and a composited image and compositing mask generating unit that extracts a compositing person region from an image in the learning data using a mask of the learning data, generates the composited image from the background image and the compositing person region, using the compositing parameters, generates the compositing mask from a mask generating image and the compositing person region that are the same size as the composited image, using the compositing parameters, and generates the composited learning data.