基于迭代匹配和未登录词识别的航空领域文本分词方法

By constructing a relevance matrix and a word formation weight dictionary, and combining semantic similarity and contextual similarity, iterative matching and out-of-vocabulary word recognition are performed, solving the problem of insufficient new word recognition capability and efficiency in Chinese word segmentation in the aviation field, and achieving efficient and accurate word segmentation results.

CN118070799BActive Publication Date: 2026-07-17CHINA AERO POLYTECH ESTAB

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AERO POLYTECH ESTAB
Filing Date
2023-11-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing Chinese word segmentation algorithms are insufficient in terms of ability and efficiency in handling new word recognition, especially in the aviation field. Traditional methods rely on dictionaries, have high computational requirements, are prone to word association, and have weak new word recognition capabilities.

Method used

By constructing a relevance matrix and a word formation weight dictionary, and combining semantic similarity and contextual similarity, iterative matching and out-of-vocabulary word identification are performed to establish segmentation rules, identify out-of-vocabulary words, and improve word segmentation efficiency.

Benefits of technology

It effectively alleviates the dependence of traditional word segmentation algorithms on dictionaries, reduces the amount of computation, avoids word concatenation, improves the ability to identify new words and the accuracy of word segmentation, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118070799B_ABST
    Figure CN118070799B_ABST
Patent Text Reader

Abstract

本发明涉及一种基于迭代匹配和未登录词识别的航空领域文本分词方法,其包括:S1、利用航空领域文本的训练集语料构建相关性矩阵,获得构词权值词典;S2、使用语义相似度制定切分规则,切分航空领域文本中的长词得到基础词;S3、分析基础词的组成,找出未登录单字进行未登录词识别,完成航空领域文本的分词操作。本发明提出的航空领域文本分词方法,通过使用语义相似度均值制定长词的切分规则,对组合词进行迭代切分,解决传统最大匹配算法词语粘连的缺陷。建立了由未登录单字和后续词组成的相关性矩阵,实现了未登录词的自动识别,提高算法效率;重新确定语境相似度,得到未登录词判定式,提高未登录词的识别能力,丰富已有词典,降低词典依赖性。
Need to check novelty before this filing date? Find Prior Art