一种基于重点区域破坏与重构学习的唐卡佛像识别方法

By employing a learning method based on key region destruction and reconstruction, the key regions of Thangka images are cropped and scrambled using the YOLOv10 algorithm. An attention mechanism is introduced into the classification network, which solves the problems of coarse recognition granularity and low accuracy in Thangka Buddha image recognition, and achieves efficient fine-grained recognition results.

CN122416004APending Publication Date: 2026-07-17TIBET UNIVERSITY FOR NATIONALITIES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIBET UNIVERSITY FOR NATIONALITIES
Filing Date
2026-06-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for Thangka Buddha image recognition suffer from problems such as coarse recognition granularity, low accuracy, and low computational efficiency. In particular, due to the complex background and high dimensionality of Thangka images, traditional convolutional neural networks face sparsity issues when processing high-dimensional features.

Method used

We employ a learning approach based on key region destruction and reconstruction. We use the YOLOv10 algorithm to identify and crop key regions of Thangka images, scramble the spatial layout of sub-blocks through affine transformation, and introduce the Convolutional Block Attention (CBAM) and Spatial-Frequency and Channel Transpose Attention (SFCT) mechanisms into the classification network. We jointly optimize the classification, adversarial, and region alignment loss functions for end-to-end training.

Benefits of technology

It significantly improves the fine-grained recognition accuracy to 85.92%, which is superior to many existing fine-grained and Thangka-specific recognition methods. It also improves computational efficiency and achieves efficient recognition on self-built datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416004A_ABST
    Figure CN122416004A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于重点区域破坏与重构学习的唐卡佛像识别方法,涉及图像识别技术领域,包括:获取待识别唐卡图像;利用YOLOv10算法裁剪出重点区域子图,均匀划分为k×k个子块后经仿射变换打乱空间布局;将原始重点区域子图与破坏后图像共同输入分类网络,该网络融合卷积块注意力机制CBAM和空间‑频率与通道转置注意力机制SFCT;计算并最小化包含更新后分类损失、对抗学习损失和区域对齐学习损失的总损失函数进行端到端训练;利用训练好的网络输出佛像类别。本发明能有效去除背景噪声干扰,聚焦佛像细粒度判别区域,显著提升唐卡佛像识别的准确率和计算效率。
Need to check novelty before this filing date? Find Prior Art