Label data processing method, device and equipment, and storage medium
By evaluating and dividing the data to be submitted for labeling, and selecting the target dataset according to the iteration requirements of the algorithm model, the problems of high labeling cost and sample set imbalance were solved, thereby improving the iteration effect of the algorithm model and the data quality.
Patent Information
- Authority / Receiving Office
- CN Β· China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
- Filing Date
- 2022-07-14
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, data mining methods result in high labeling costs and an imbalance in the number of categories in the sample set. The training sample set has a long tail distribution, and the collection of bad case data samples is limited and time-consuming, which affects the iterative effect of the algorithm model.
By utilizing multiple historical versions of the algorithm model to evaluate the data to be submitted for bidding, the dataset is divided according to the evaluation results, and the target dataset is determined according to the iteration requirements, so as to selectively mine more valuable data.
It improved the quality of the submitted data, reduced the labeling cost, enhanced the iteration effect of the algorithm model, and shortened the iteration time.
Smart Images

Figure CN115329979B_ABST