基于掩码下一尺度预测的自监督场景文字识别方法
By constructing a self-supervised method with multi-scale views and loss function optimization, the problems of lack of multi-scale modeling and attention diffusion in self-supervised scene text recognition are solved, achieving efficient recognition and improved robustness of scene text.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANKAI UNIV
- Filing Date
- 2026-05-06
- Publication Date
- 2026-07-17
AI Technical Summary
Existing self-supervised scene text recognition methods lack the ability to model multi-scale hierarchical structures, leading to problems such as attention diffusion, limited field of view, and cross-scale semantic inconsistency.
A self-supervised method based on masked next-scale prediction is adopted. By constructing small-scale enhanced views, large-scale enhanced views and local magnified views, and combining a shared coding network, a next-scale prediction decoder and a masked reconstruction decoder, multi-scale feature extraction and reconstruction are performed, and multi-scale language alignment loss is used for joint optimization.
It significantly improves the model's ability to recognize extreme scale variations and blurred text, solves the problem of attention diffusion, ensures the semantic consistency of cross-scale features, and improves the accuracy and robustness of text recognition in complex scenarios.
Smart Images

Figure CN122157225B_ABST