一种基于重点区域破坏与重构学习的唐卡佛像识别方法
By employing a learning method based on key region destruction and reconstruction, the key regions of Thangka images are cropped and scrambled using the YOLOv10 algorithm. An attention mechanism is introduced into the classification network, which solves the problems of coarse recognition granularity and low accuracy in Thangka Buddha image recognition, and achieves efficient fine-grained recognition results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIBET UNIVERSITY FOR NATIONALITIES
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for Thangka Buddha image recognition suffer from problems such as coarse recognition granularity, low accuracy, and low computational efficiency. In particular, due to the complex background and high dimensionality of Thangka images, traditional convolutional neural networks face sparsity issues when processing high-dimensional features.
We employ a learning approach based on key region destruction and reconstruction. We use the YOLOv10 algorithm to identify and crop key regions of Thangka images, scramble the spatial layout of sub-blocks through affine transformation, and introduce the Convolutional Block Attention (CBAM) and Spatial-Frequency and Channel Transpose Attention (SFCT) mechanisms into the classification network. We jointly optimize the classification, adversarial, and region alignment loss functions for end-to-end training.
It significantly improves the fine-grained recognition accuracy to 85.92%, which is superior to many existing fine-grained and Thangka-specific recognition methods. It also improves computational efficiency and achieves efficient recognition on self-built datasets.
Smart Images

Figure CN122416004A_ABST