SYSTEMS, METHODS, AND TECHNIQUES FOR LEARNING AND USING INSTANCE-DEPENDENT HOLLOW ATTENTION FOR EFFECTIVE VISION TRANSFORMERS
FR3153920B3Active Publication Date: 2025-11-07LOREAL SA
0 Cites -1 Cited by
Patent Information
- Application Number
- FR2023010590
- Authority / Receiving Office
- FR · FR
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2023-10-04
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2033-10-04
Abstract
Systems, Methods, and Techniques for Learning and Using Instance-Dependent Shallow Attention for Efficient Vision Transformers. Vision transformers (ViTs) have demonstrated competitive performance advantages over convolutional neural networks (CNNs), although they are often associated with high computational costs. The methods, systems, and techniques described here learn instance-dependent attention patterns, using a lightweight connectivity predictor module to estimate a connectivity score for each pair of tokens. Intuitively, two tokens have high connectivity scores if their features are considered spatially or semantically relevant. Since each token supports only a small number of other tokens, the resulting binary connectivity masks are often very sparse, providing an opportunity to speed up the network through sparse computations.Equipped with the learned unstructured attention pattern, the hollow attention ViT produces a superior Pareto-optimal trade-off between FLOPs and top-1 accuracy on ImageNet compared to token hollowness (48%–69% reduction in MHSA FLOPs; accuracy drop to 0.4%). A combination of attention hollowness and tokens reduces FLOP ViTs by more than 60%. Figure for the abstract: none.
Need to check novelty before this filing date? Find Prior Art