Adaptive Label Histogram Sampling for Cost-Aware Diversity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for creating label histograms through crowdsourcing are costly and inefficient due to the concentration of votes on specific labels, leading to wasted resources and reduced diversity in the final set of label histograms.

Innovation Solution

A label histogram creating device that performs a first sampling process on all data points with a limited number of votes, followed by a second sampling process on selectively chosen data points with increased votes, based on information entropy, to enhance diversity while minimizing overall sampling costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of times of sampling is increased to increase diversity of label histograms, then the accuracy of performance evaluation is improved, but the monetary cost increases

Engineering Contradiction:
Improveaccuracy of performance evaluationVSAvoidmonetary cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent applies local quality by differentiating the sampling strategy across different data points based on their individual characteristics. Specifically, data points are categorized into two groups: those with concentrated votes (high confidence) and those with dispersed votes (low confidence). The sampling process is then optimized locally for each group - concentrated data points undergo minimal or no additional sampling, while dispersed data points receive intensified sampling. This localized approach ensures that resources are allocated efficiently to where they are most needed, improving evaluation accuracy without proportionally increasing overall cost.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the sampling parameter (number of times of sampling) based on the confidence level of initial sampling results. For data points showing concentrated votes, the sampling parameter is kept low or set to zero for subsequent rounds. For data points with dispersed votes, the sampling parameter is increased in subsequent sampling rounds. This dynamic parameter adjustment allows the system to adapt the sampling intensity to the actual needs of each data point, resolving the contradiction between accuracy improvement and cost control.

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If votes are concentrated on a specific label for certain data points, then the sampling cost is reduced, but the diversity of label histograms decreases

Engineering Contradiction:
Improvesampling costVSAvoiddiversity of label histograms
Core Design Contradiction:
Loss of energyVSAdaptability or versatility

Solution Approach 1:

The patent performs a preliminary sampling round to assess the vote distribution characteristics of each data point before committing to a final sampling strategy. This preliminary action allows the system to identify which data points have concentrated votes and which have dispersed votes. Based on this preliminary assessment, the system then decides whether to invest additional sampling resources in each data point, thereby preliminarily determining the cost-diversity trade-off for the entire dataset.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by performing additional sampling only on the subset of data points that exhibit dispersed votes, rather than uniformly increasing sampling for all data points. This partial approach ensures that diversity is enhanced where needed while avoiding unnecessary costs for data points that already show concentrated votes. The sampling effort is partially applied based on actual requirement, resolving the contradiction between cost reduction and diversity maintenance.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If a label histogram is discarded when votes are concentrated, then the diversity of the final set is increased, but the sampling cost is wasted

Engineering Contradiction:
Improvediversity of label histogramsVSAvoidwasted sampling cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent introduces dynamics into the sampling process by making the sampling strategy adaptive rather than static. Instead of predetermined uniform sampling or simple discarding rules, the system dynamically adjusts the sampling intensity for each data point based on real-time feedback from previous sampling rounds. Data points with concentrated votes dynamically receive reduced or zero additional sampling, while those with dispersed votes dynamically receive increased sampling. This dynamic adaptation prevents waste while maintaining diversity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback mechanisms where the results of each sampling round are used to inform subsequent sampling decisions. The vote distribution pattern from initial sampling serves as feedback that triggers different follow-up actions: concentrated votes feedback leads to reduced sampling or data point selection for the final set, while dispersed votes feedback leads to intensified sampling. This feedback loop ensures that sampling resources are not wasted on data points that don't need them, while still achieving the desired diversity in the final label histogram set.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250299392A1Label histogram creating device, label histogram creating method and label histogram creating program
Publication Date: 2025.09.25 NT T INC
  • US20250299392A1 patent drawing
  • US20250299392A1 patent drawing
  • US20250299392A1 patent drawing

AI summary

A label histogram creating part (14) of a label histogram creating device (1) sets the number of times of sampling (β) for each piece of data (x) for a data set (X) including N pieces of data (x) and performs a first sampling process on the data set (X) by using a crowdsourcing (2) to create a set (L) of label histograms. A pick out part (16) performs a pick out process of picking out pieces of data (x) that are targets of a second sampling process from the data set (X) on the basis of uncertainty of information included in the label histograms. The label histogram creating part (14) performs the second sampling process on the pieces of data (x) picked out by the pick out part (16) with the number of times of sampling (β) increased compared to the number of times of sampling (β) in the first sampling process.