Data labeling method, device and system based on incremental learning and confidence learning

CN122347716APending Publication Date: 2026-07-07CHINA AUTOMOTIVE INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AUTOMOTIVE INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD
Filing Date
2026-06-09
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

In existing technologies, manual annotation and model annotation methods suffer from high costs and inefficiencies in the field of data annotation, especially in large-scale datasets and frequently updated scenarios, making it difficult to guarantee the accuracy and high quality of labels.

Method used

We employ a combination of incremental learning and confidence learning. We train a supervised learning model using an initial dataset with manual annotations to generate predicted labels. Then, we use confidence learning to filter high-confidence labels and incremental learning to optimize the model, thereby reducing the workload of manual annotation and improving label accuracy.

Benefits of technology

This approach reduces manual intervention while ensuring the accuracy and high quality of labels, thereby improving the efficiency and precision of model annotation and lowering annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347716A_ABST
    Figure CN122347716A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electric digital processing, in particular to a data labeling method, device and system based on incremental learning and confidence learning. The method comprises the following steps: training a supervised learning model in a first stage through a first data set labeled by manual labeling; inputting an unlabeled second data set into the supervised learning model trained in the first stage; mixing the first data set and the second data set to divide them into multiple data subsets; training a target model by using a part of the subsets to obtain labels and prediction probabilities of another part of the subsets; selecting data belonging to the first data set in the another part of the subsets and calculating a benchmark value of correct labels; selecting data belonging to the second data set in the another part of the subsets and retaining data with prediction probabilities greater than or equal to the benchmark value as incremental training samples. The application can reduce manual operation, ensure the accuracy and high quality of labels and improve the accuracy of model labeling.
Need to check novelty before this filing date? Find Prior Art