Discriminator Data Labeling via Boundary Region Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating labeled data for discriminator learning are inefficient due to the lack of consideration for regions effective for learning, leading to high work costs and ineffective labeling.

Innovation Solution

An apparatus comprising an obtaining unit, a setting unit, a determining unit, and a display control unit that identifies and displays partial regions of target data for labeling, based on the distribution of local data, to efficiently select and label data near the discrimination boundary, thereby reducing labeling costs and improving learning efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data is selected for labeling based on discrimination boundary proximity alone, then learning effectiveness is improved, but labeling efficiency deteriorates due to lack of region consideration

Engineering Contradiction:
Improvelearning effectivenessVSAvoidlabeling efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the data space into multiple partial regions based on the distribution of local data points near the discrimination boundary. Instead of treating all boundary-proximity data uniformly, the system divides the data into distinct regions (e.g., Region A, Region B, Region C) and selects partial regions for labeling based on their specific characteristics and densities, thereby improving labeling efficiency while maintaining learning effectiveness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by considering the specific characteristics of different data regions. The determining unit evaluates the distribution density and spatial characteristics of local data points in different areas and selects partial regions accordingly. Regions with higher density or specific spatial patterns are prioritized for labeling, allowing the system to adapt its labeling strategy to the local quality of each region rather than applying a uniform approach.

Inventive Principle:
Principle #3Local quality

2Reliability

If all data is labeled to ensure comprehensive learning, then learning completeness is improved, but work cost increases significantly

Engineering Contradiction:
Improvelearning completenessVSAvoidwork cost
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the most valuable portions of the data for labeling. By identifying local data points near the discrimination boundary and determining which partial regions contain the most informative data, the system extracts and labels only those critical regions rather than processing all data. This extraction approach maintains learning completeness by focusing on high-value data while significantly reducing the overall work cost.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by selecting only the necessary partial regions for labeling rather than processing the entire dataset. The determining unit identifies and labels only the partial regions that contain local data points with the highest learning value, accepting that not all data needs to be labeled to achieve effective learning. This partial approach reduces work cost while maintaining sufficient learning completeness.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11741683B2Apparatus for processing labeled data to be used in learning of discriminator, method of controlling the apparatus, and non-transitory computer-readable recording medium
Publication Date: 2023.08.29 CANON KK
  • US11741683B2 patent drawing
  • US11741683B2 patent drawing
  • US11741683B2 patent drawing

AI summary

An apparatus comprising: an obtaining unit configured to obtain target data as a result of discrimination of each portion of input data performed by a discriminator having learned in advance by using existing labeled data; a setting unit configured to set each portion of the target data, which is effective for additional learning of the discriminator, as local data; a determining unit configured to determine not less than one partial region of the target data, which accepts labeling by a user, based on a distribution of the set local data in the target data; and a display control unit configured to cause a display unit to display the determined not less than one partial region.