一种数据处理方法及装置

By calculating the semantic similarity in the content set and correcting the labels, the problem of low accuracy in label recognition models is solved, and a more efficient label recognition effect is achieved.

CN115577270BActive Publication Date: 2026-07-17BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2022-09-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The accuracy of tag recognition models trained in existing technologies is not high, mainly because the accuracy of content tags is difficult to guarantee.

Method used

By acquiring a set of content and a set of tags to be corrected, the semantic similarity between the content is calculated. When the similarity meets the preset conditions and the tags are inconsistent, a prompt message is output to indicate that the tags may be incorrect, and the tags are corrected. Finally, the second tag recognition model is trained.

Benefits of technology

This improves the accuracy of the label recognition model, avoids training the model based on inaccurate labels, and enhances the accuracy of label recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115577270B_ABST
    Figure CN115577270B_ABST
Patent Text Reader

Abstract

本申请公开了一种数据处理方法,该方法包括:获取第一内容集合和第一待校正标签集合,第一待校正标签集合包括第一内容集合中的各个内容分别对应的第一待校正标签;确定第一内容集合中的第一内容和第二内容的含义相似度;若第一内容和第二内容的含义相似度符合预设条件、并且第一内容的第一待校正标签不同于第二内容的第一待校正标签,则输出提示信息,提示信息用于指示第一内容的第一待校正标签和第二内容的第一待校正标签中至少存在一个错误的标签。由此可见,利用本方案,可以筛选出第一待校正标签可能不准确的内容,从而避免基于第一待校正标签可能不准确的内容来训练标签识别模型,相应的,可以避免标签识别模型的识别准确性不高。
Need to check novelty before this filing date? Find Prior Art