Label data processing method, device and equipment, and storage medium

By evaluating and dividing the data to be submitted for labeling, and selecting the target dataset according to the iteration requirements of the algorithm model, the problems of high labeling cost and sample set imbalance were solved, thereby improving the iteration effect of the algorithm model and the data quality.

CN115329979BActive Publication Date: 2026-07-24APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN Β· China
Patent Type
Patents(China)
Current Assignee / Owner
APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
Filing Date
2022-07-14
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, data mining methods result in high labeling costs and an imbalance in the number of categories in the sample set. The training sample set has a long tail distribution, and the collection of bad case data samples is limited and time-consuming, which affects the iterative effect of the algorithm model.

Method used

By utilizing multiple historical versions of the algorithm model to evaluate the data to be submitted for bidding, the dataset is divided according to the evaluation results, and the target dataset is determined according to the iteration requirements, so as to selectively mine more valuable data.

Benefits of technology

It improved the quality of the submitted data, reduced the labeling cost, enhanced the iteration effect of the algorithm model, and shortened the iteration time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329979B_ABST
    Figure CN115329979B_ABST
Patent Text Reader

Abstract

The disclosure provides a kind of marking data processing method, device and equipment and storage medium. It is related to artificial intelligence field, especially it is related to intelligent transportation and other fields. The specific implementation scheme is: obtaining multiple to be sent data;Using the N historical version algorithm of algorithm model to the multiple to be sent data is evaluated, obtains the N evaluation results of multiple to be sent data;According to the N evaluation results of multiple to be sent data, multiple to be sent data is divided into M data sets;According to the iteration requirement of algorithm model, the target data set is determined from M data sets.According to the scheme of the disclosure, it can be targeted to the different iteration requirements of algorithm model at different time, more valuable data is mined, the quality of to be sent data is improved, and then the annotation cost can be reduced, and the iteration effect of algorithm model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art