基于预训练模型的中文跨语言知识增强方法

By constructing a bilingual dictionary to acquire initial consonant and translation knowledge, and combining a tree-structured attention mechanism and BERT to generate embeddings, the shortcomings of Chinese pre-trained models in the fusion of polyphonic characters and cross-linguistic knowledge are addressed, thereby improving the performance of Chinese sentiment analysis and other tasks.

CN117648935BActive Publication Date: 2026-07-17GUILIN UNIV OF ELECTRONIC TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUILIN UNIV OF ELECTRONIC TECH
Filing Date
2023-10-26
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing Chinese pre-trained models have shortcomings in understanding text semantics and Chinese character features, especially in the problem of polyphonic characters. At the same time, they ignore the characteristics of Chinese itself when integrating cross-language knowledge, which limits the effectiveness of sentiment analysis and other natural language processing tasks.

Method used

By constructing a bilingual dictionary to acquire initial consonant and translation knowledge, we calculate attention scores using tree-based attention and variable attention mechanisms, and combine BERT to generate embeddings. We then use singular value decomposition and focus loss function for fine-tuning to enhance the semantic information of the Chinese distributed representation.

Benefits of technology

It significantly improves the model's performance in sentiment analysis, named entity recognition, natural language inference, and domain question answering tasks, reaching state-of-the-art levels. It also solves the problem of polyphonic characters and enhances the semantic representation of Chinese word embeddings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117648935B_ABST
    Figure CN117648935B_ABST
Patent Text Reader

Abstract

本发明涉及自然语言处理技术领域,涉及一种基于预训练模型的中文跨语言知识增强方法,其包括以下步骤:步骤1、句子进行分词;步骤2、分词后的句子通过多语言知识增强模块中的声母知识层和翻译知识层获得翻译知识和声母知识;步骤3、将翻译知识和声母知识与原句子拼接成一棵树,使用可变注意力机制来计算树上的注意力分数;步骤4、使用BERT生成嵌入,包括标记嵌入、位置嵌入和段嵌入,同时在每个标记之后插入相应的初始辅音和翻译知识;步骤5、基于生成的注意力分数和嵌入,对编码器的输出进行了微调,以实现各种下游任务。本发明能较佳地实现中文跨语言知识增强。
Need to check novelty before this filing date? Find Prior Art