语料数据的序号识别处理方法、装置、设备和存储介质
By performing multi-level sequence recognition processing on parallel corpus data, high-quality target corpus data is selected for training the translation model, which solves the problem of low quality of traditional training corpus data and improves the accuracy and precision of the translation model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2023-04-19
- Publication Date
- 2026-07-17
AI Technical Summary
In the training process of traditional machine translation models, the training corpus data used contains a lot of invalid data or translation errors, resulting in low accuracy of translation results.
By performing multi-level sequence recognition processing on parallel corpus data, including first-level, second-level, and third-level sequence recognition, target corpus data that meets the conditions is selected for training the translation model, ensuring the accuracy of sequence recognition and the quality of the data.
It improves the model accuracy and translation result accuracy of the translation model, and reduces errors and erroneous data during the training process.
Smart Images

Figure CN118821794B_ABST