SYSTEMS, METHODS, AND TECHNIQUES FOR LEARNING AND USING INSTANCE-DEPENDENT HOLLOW ATTENTION FOR EFFECTIVE VISION TRANSFORMERS

FR3153920B3Active Publication Date: 2025-11-07LOREAL SA
0 Cites -1 Cited by

Patent Information

Application Number
FR2023010590
Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2023-10-04
Publication Date
2025-11-07
Estimated Expiration
2033-10-04
Patent Text Reader

Abstract

Systems, Methods, and Techniques for Learning and Using Instance-Dependent Shallow Attention for Efficient Vision Transformers. Vision transformers (ViTs) have demonstrated competitive performance advantages over convolutional neural networks (CNNs), although they are often associated with high computational costs. The methods, systems, and techniques described here learn instance-dependent attention patterns, using a lightweight connectivity predictor module to estimate a connectivity score for each pair of tokens. Intuitively, two tokens have high connectivity scores if their features are considered spatially or semantically relevant. Since each token supports only a small number of other tokens, the resulting binary connectivity masks are often very sparse, providing an opportunity to speed up the network through sparse computations.Equipped with the learned unstructured attention pattern, the hollow attention ViT produces a superior Pareto-optimal trade-off between FLOPs and top-1 accuracy on ImageNet compared to token hollowness (48%–69% reduction in MHSA FLOPs; accuracy drop to 0.4%). A combination of attention hollowness and tokens reduces FLOP ViTs by more than 60%. Figure for the abstract: none.
Need to check novelty before this filing date? Find Prior Art