Learning Device With Symmetry-Based Feature Point Label Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting feature points from images require large amounts of training data, which is time-consuming, and necessitate preparing multiple extractors for each label, leading to significant labor costs.

Innovation Solution

A learning device and method that converts unique feature point labels into shared labels based on congruence or mirror symmetry relations, using a first inference engine to identify feature points and a second inference engine to refine positions, reducing the number of labels needed and increasing training efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large amount of training data is used to increase accuracy and robustness of the extraction model, then the model performance is improved, but the time required for data collection and processing enormously increases

Engineering Contradiction:
Improveaccuracy and robustness of extraction modelVSAvoidtime required for data collection
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges multiple unique labels (first labels) that correspond to feature points with congruence or mirror symmetry relations into a single shared label (second label). This consolidation allows training data to be shared across multiple feature point types, effectively increasing the amount of training data available for each label without requiring additional data collection. The inference device uses this label conversion mechanism to improve model accuracy and robustness while avoiding the time-consuming process of collecting separate large datasets for each unique feature point label.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If the same number of extractors are prepared as the number of labels for specifying feature point positions, then the extraction accuracy is maintained, but the labor for preparing extractors enormously increases

Engineering Contradiction:
Improvefeature point position extraction accuracyVSAvoidnumber of extractors to be prepared
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal inference engine that can handle multiple feature point types through label conversion. Instead of preparing separate extractors for each unique label, the system uses a single inference engine that converts first labels to second labels based on congruence or mirror symmetry relations. This allows one extractor to perform the function of multiple extractors would otherwise be needed, significantly reducing preparation labor while maintaining extraction accuracy through the learned relationships between different feature point types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12412298B2Learning device, control method, and storage medium
Publication Date: 2025.09.09 NEC CORP
  • US12412298B2 patent drawing
  • US12412298B2 patent drawing
  • US12412298B2 patent drawing

AI summary

The learning device 1A includes an acquiring means 23A, a conversion means 24A, and a learning means 25A. The acquiring means 23A is configured to acquire a combination of a first label that is a unique label for each feature point of an object and a feature point image in which a feature point corresponding to the first label is imaged. The conversion means 24A is configured to convert the first label to a second label that is set to a same label for feature points of the object with at least one of a congruence relation in appearance or a mirror symmetry relation to one another. The learning means 25A is configured to learn an inference engine based on the second label, the feature point image, and correct answer data regarding a position of the feature point, the inference engine being configured to perform an inference on the position of the feature point included in an image that is inputted to the inference engine.