Adaptive Label Histogram Sampling for Cost-Aware Diversity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for creating label histograms through crowdsourcing are costly and inefficient due to the concentration of votes on specific labels, leading to wasted resources and reduced diversity in the final set of label histograms.
Innovation Solution
A label histogram creating device that performs a first sampling process on all data points with a limited number of votes, followed by a second sampling process on selectively chosen data points with increased votes, based on information entropy, to enhance diversity while minimizing overall sampling costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of times of sampling is increased to increase diversity of label histograms, then the accuracy of performance evaluation is improved, but the monetary cost increases
Solution Approach 1:
The patent applies local quality by differentiating the sampling strategy across different data points based on their individual characteristics. Specifically, data points are categorized into two groups: those with concentrated votes (high confidence) and those with dispersed votes (low confidence). The sampling process is then optimized locally for each group - concentrated data points undergo minimal or no additional sampling, while dispersed data points receive intensified sampling. This localized approach ensures that resources are allocated efficiently to where they are most needed, improving evaluation accuracy without proportionally increasing overall cost.
Solution Approach 2:
The patent changes the sampling parameter (number of times of sampling) based on the confidence level of initial sampling results. For data points showing concentrated votes, the sampling parameter is kept low or set to zero for subsequent rounds. For data points with dispersed votes, the sampling parameter is increased in subsequent sampling rounds. This dynamic parameter adjustment allows the system to adapt the sampling intensity to the actual needs of each data point, resolving the contradiction between accuracy improvement and cost control.
2Loss of energy
If votes are concentrated on a specific label for certain data points, then the sampling cost is reduced, but the diversity of label histograms decreases
Solution Approach 1:
The patent performs a preliminary sampling round to assess the vote distribution characteristics of each data point before committing to a final sampling strategy. This preliminary action allows the system to identify which data points have concentrated votes and which have dispersed votes. Based on this preliminary assessment, the system then decides whether to invest additional sampling resources in each data point, thereby preliminarily determining the cost-diversity trade-off for the entire dataset.
Solution Approach 2:
The patent applies partial action by performing additional sampling only on the subset of data points that exhibit dispersed votes, rather than uniformly increasing sampling for all data points. This partial approach ensures that diversity is enhanced where needed while avoiding unnecessary costs for data points that already show concentrated votes. The sampling effort is partially applied based on actual requirement, resolving the contradiction between cost reduction and diversity maintenance.
3Adaptability or versatility
If a label histogram is discarded when votes are concentrated, then the diversity of the final set is increased, but the sampling cost is wasted
Solution Approach 1:
The patent introduces dynamics into the sampling process by making the sampling strategy adaptive rather than static. Instead of predetermined uniform sampling or simple discarding rules, the system dynamically adjusts the sampling intensity for each data point based on real-time feedback from previous sampling rounds. Data points with concentrated votes dynamically receive reduced or zero additional sampling, while those with dispersed votes dynamically receive increased sampling. This dynamic adaptation prevents waste while maintaining diversity.
Solution Approach 2:
The patent implements feedback mechanisms where the results of each sampling round are used to inform subsequent sampling decisions. The vote distribution pattern from initial sampling serves as feedback that triggers different follow-up actions: concentrated votes feedback leads to reduced sampling or data point selection for the final set, while dispersed votes feedback leads to intensified sampling. This feedback loop ensures that sampling resources are not wasted on data points that don't need them, while still achieving the desired diversity in the final label histogram set.
Data Source
AI summary
A label histogram creating part (14) of a label histogram creating device (1) sets the number of times of sampling (β) for each piece of data (x) for a data set (X) including N pieces of data (x) and performs a first sampling process on the data set (X) by using a crowdsourcing (2) to create a set (L) of label histograms. A pick out part (16) performs a pick out process of picking out pieces of data (x) that are targets of a second sampling process from the data set (X) on the basis of uncertainty of information included in the label histograms. The label histogram creating part (14) performs the second sampling process on the pieces of data (x) picked out by the pick out part (16) with the number of times of sampling (β) increased compared to the number of times of sampling (β) in the first sampling process.


