语料数据的序号识别处理方法、装置、设备和存储介质

By performing multi-level sequence recognition processing on parallel corpus data, high-quality target corpus data is selected for training the translation model, which solves the problem of low quality of traditional training corpus data and improves the accuracy and precision of the translation model.

CN118821794BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-04-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In the training process of traditional machine translation models, the training corpus data used contains a lot of invalid data or translation errors, resulting in low accuracy of translation results.

Method used

By performing multi-level sequence recognition processing on parallel corpus data, including first-level, second-level, and third-level sequence recognition, target corpus data that meets the conditions is selected for training the translation model, ensuring the accuracy of sequence recognition and the quality of the data.

Benefits of technology

It improves the model accuracy and translation result accuracy of the translation model, and reduces errors and erroneous data during the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118821794B_ABST
    Figure CN118821794B_ABST
Patent Text Reader

Abstract

本申请涉及一种语料数据的序号识别处理方法、装置、设备和存储介质。所述方法涉及人工智能,包括:对待识别平行语料数据进行一级序号识别处理,获得待识别平行语料数据的序号分布属性,对序号分布属性满足二级序号识别处理条件的待识别平行语料数据,进行二级序号识别处理,获得二级序号识别结果。根据序号分布属性和二级序号识别结果,确定满足三级序号识别处理条件的第一目标语料数据端,对第一目标语料数据端进行三级序号识别处理,获得三级序号识别结果。基于二级序号识别结果或三级序号识别结果,对语料数据进行匹配处理和数据筛选,获得目标语料数据。采用本方法能够全面准确获得语料数据在不同端的序号,剔除序号不对等的错误数据。
Need to check novelty before this filing date? Find Prior Art