基于深度学习和表意文字描述序列的多种类汉字识别方法

By using deep learning and ideographic character description sequences, a large number of data samples were generated and a recognition network was trained, which solved the problem of recognizing rare characters and characters in the official script. This enabled accurate recognition of rare characters and differentiation of characters in the official script, thus improving the accuracy of ancient character recognition.

CN117333883BActive Publication Date: 2026-07-17HUAZHONG UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2023-10-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing character recognition methods struggle to accurately identify rare characters and distinguish between characters with established script, especially in the study of ancient characters, where current technology cannot effectively address the problem of recognizing and differentiating between rare characters and characters with established script.

Method used

This study employs a deep learning-based approach using ideographic character description sequences. By generating a large number of data samples, including images of non-existent Chinese characters, the recognition network is trained. The model is then optimized using residual networks and cross-entropy loss functions, thereby improving the accuracy of recognizing rare characters and enhancing the ability to distinguish between characters from the official script.

Benefits of technology

It significantly improves the variety and accuracy of Chinese character recognition, especially the recognition rate of rare characters, and can accurately distinguish clerical script characters, thus improving the accuracy of ancient character recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117333883B_ABST
    Figure CN117333883B_ABST
Patent Text Reader

Abstract

本发明提出了一种基于深度学习和表意文字描述序列的多种类汉字识别方法,包括以下步骤:首先利用汉字表意文字描述序列,生成已有近九万种汉字以及随机生成不存在的汉字的图像数据,然后将图像数据经过大量数据增强后通过残差网络,并采用改进后的交叉损失函数进行训练,最后对于输入图片进行多种类汉字的识别。本发明通过输入种类繁多的汉字图像以及不断随机生成不存在的新汉字图像,利用深度的残差网络和改进后的交叉熵损失函数进行训练,这样的训练方式不仅增强了对生僻字的识别能力,还实现了对隶定字的有效区分。
Need to check novelty before this filing date? Find Prior Art