Dual-Model Annotation Data Selection for Precise Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The evaluation of online classification models faces challenges with high labeling costs and inadequate precision due to the use of random evaluation sets that do not accurately reflect real-time data distribution, leading to fluctuations in recall estimation.
Innovation Solution
A method and device for determining labeled data using two text recognition models to identify candidate data that meets a labeling condition, filtering out data recognized by at least one model as belonging to a target category, and adjusting a score threshold to optimize the number of data to be labeled, thereby reducing labeling costs and improving precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a full labeled fixed test set is used for model evaluation, then the evaluation can be performed, but the labeling cost is high and the precision is inadequate due to random sampling not reflecting real-time data distribution
Solution Approach 1:
The patent changes the parameter of data selection from random sampling to structured sampling based on score thresholds. By adjusting the threshold parameter, the system selects data with specific characteristics (uncertain predictions) that better reflect real-time data distribution, improving evaluation precision while reducing the number of data points that need labeling
Solution Approach 2:
The patent performs preliminary filtering of candidate data using multiple text recognition models before the actual labeling process. By pre-identifying and scoring candidate data based on model uncertainty, the system prepares a refined subset of data that requires labeling, reducing the overall labeling cost while maintaining evaluation quality
2Ease of operation
If random evaluation sets are used for model evaluation, then the evaluation process is simple, but the recall estimation fluctuates due to inadequate representation of real-time data distribution
Solution Approach 1:
The patent replaces random sampling with threshold-based sampling where data is selected based on prediction score thresholds from multiple models. This parameter change ensures that selected data represents uncertain cases that better reflect real-time data distribution, stabilizing recall estimation while maintaining operational simplicity through automated threshold application
Solution Approach 2:
The patent implements a feedback mechanism where multiple text recognition models evaluate candidate data and generate scores based on prediction confidence. This feedback loop identifies data points with uncertain predictions, which are then selected for labeling, creating a self-correcting system that improves recall estimation stability
3Measurement precision
If multiple text recognition models are used to identify candidate data, then the precision of data selection is improved, but the device complexity increases
Solution Approach 1:
The patent segments the data selection process into distinct stages: candidate data generation by multiple models, scoring based on prediction confidence, threshold-based filtering, and final selection. This segmentation allows each model to perform a specific function and makes the overall complex process manageable and systematic
Solution Approach 2:
The patent makes the text recognition models serve multiple functions: they perform their primary text recognition task and simultaneously generate confidence scores for candidate data selection. This multi-functionality reduces the need for separate evaluation mechanisms, managing complexity while maintaining data selection precision
Data Source
AI summary
The present disclosure relates to an annotation data determination method and apparatus, and a readable medium and an electronic device. By means of the present disclosure, high-quality data to be annotated is obtained for model performance evaluation. The method includes: acquiring candidate data from a candidate data set; respectively inputting the candidate data into a first text recognition model and a second text recognition model, so as to obtain a first recognition result output by the first text recognition model and a second recognition result output by the second text recognition model, wherein both the first text recognition model and the second text recognition model can recognize whether text data is of a target category; according to the first recognition result and the second recognition result, determining whether the candidate data meets an annotation condition, wherein the annotation condition is the category of the candidate data being recognized by the first text recognition model or the second text recognition model as at least one target category among target categories; and if it is determined that the candidate data meets the annotation condition, determining the candidate data as text data to be annotated.


