Energy Distance Feature Vector Similarity for Deep Learning Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data labeling in deep learning applications is cumbersome and inefficient, particularly when dealing with large datasets such as images, audio recordings, and text files, as existing labeling models are difficult to train and generate.

Innovation Solution

A method and system for measuring similarities between sets of feature vectors using an Energy Distance measure, which generates feature representations from existing and target datasets, and calculates metric distances to facilitate efficient similarity measurement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional data labeling models are used, then data labeling can be automated, but the models are difficult to train and generate

Engineering Contradiction:
Improvedata labeling automationVSAvoidmodel training difficulty
Core Design Contradiction:
Extent of automationVSEase of manufacture

Solution Approach 1:

The patent introduces an intermediary energy distance metric that mediates between the source dataset feature representations and target dataset feature representations. This energy distance serves as a bridge to assess domain compatibility and guide the labeling process, making the automated labeling more feasible by providing a measurable criterion for model training and generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter space by transforming data into feature representations using deep learning models and measuring similarities through energy distance calculations. This parameter transformation enables automated labeling by converting complex data into a standardized feature space where automated decision-making becomes feasible.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual data labeling is performed, then labeling accuracy can be maintained, but the process is cumbersome and time-consuming

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual labeling process with an automated system that uses energy distance calculations to determine label assignments. This substitution eliminates the time-consuming manual process while maintaining accuracy through rigorous feature representation and distance-based decision-making.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service labeling by automatically assigning labels to data points based on their feature representations and energy distance measurements to labeled examples. This self-service capability eliminates the need for manual intervention while maintaining high labeling accuracy through intelligent automated decision-making.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If existing labeling models are used for large datasets, then labeling can be performed, but the models are difficult to train and generate

Engineering Contradiction:
Improvedataset sizeVSAvoidmodel training difficulty
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent segments the large dataset into source datasets with known feature representations and target datasets requiring labeling. By dividing the data into these segments and using energy distance to bridge them, the system can handle large datasets more effectively, as the energy distance metric provides a scalable way to assess similarities across different data segments without requiring retraining of complex models.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250157188A1Measure similarities between sets of feature vectors in deep learning applications
Publication Date: 2025.05.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250157188A1 patent drawing
  • US20250157188A1 patent drawing
  • US20250157188A1 patent drawing

AI summary

A method, system and apparatus for measuring similarities between sets of feature vectors, including generating a feature representation from existing sources of datasets, generating representation of metric distances of the existing sources to each other using energy distance measure from the feature representation, generating feature representation of a target dataset, generating representation of the metric distances of each of the target dataset using the energy distance measure, and providing a choice of existing sources which further extremizes a geometric content of a hypervolume described by a pseudolabel sequence.