Rearranging Feed Forward Networks in Transformer Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transformer-based models, despite their performance improvements, face challenges in efficient deployment on resource-constrained devices due to high latency, increased energy consumption, and large memory footprint, which is exacerbated by the need for more layers that further increase computation.
Innovation Solution
Rearranging feed forward networks (FFNs) of a transformer-based model by incorporating a multilayer perceptron operation on the channel dimension of a feature map and using point-wise convolutions, along with a token interaction block for spatial mixing of input features, implemented using depth-wise convolutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If transformer-based models use more layers to improve performance, then model accuracy is improved, but computation time and memory footprint increase
Solution Approach 1:
The patent changes the parameter configuration of FFN blocks by applying point-wise convolutions and redistributing parameters across layers. Specifically, it modifies the hidden dimension sizes and parameter distributions to achieve better performance with fewer computational resources, directly addressing the trade-off between accuracy and inference time.
Solution Approach 2:
The patent introduces spatial mixing operations that operate on the channel dimension of feature maps. By adding this new dimensional operation (spatial mixing via point-wise convolutions), the model achieves enhanced feature representation without increasing the traditional sequence length or number of attention heads, thus improving accuracy without proportionally increasing computation time.
2Measurement precision
If transformer-based models increase computational resources to improve accuracy, then model performance is improved, but energy consumption increases
Solution Approach 1:
The patent optimizes parameter distribution across layers by applying point-wise convolutions with specific kernel sizes and channel configurations. This parameter reconfiguration enables the model to achieve high accuracy with reduced computational intensity, thereby lowering energy consumption while maintaining performance.
Solution Approach 2:
The patent segments the FFN blocks into multiple layers with different parameter configurations, where each layer performs specialized transformations. This segmentation allows the model to distribute computational workload more efficiently across layers, reducing peak energy consumption while maintaining overall accuracy through cumulative feature transformation.
3Productivity
If transformer-based models reduce the number of parameters to improve efficiency on edge devices, then deployment efficiency is improved, but model accuracy may decrease
Solution Approach 1:
The patent applies point-wise convolutions that change the parameter density and distribution pattern. By concentrating parameters in specific channels and positions rather than uniformly distributing them, the model achieves high accuracy with fewer total parameters, making it suitable for edge devices while maintaining performance.
Solution Approach 2:
The patent creates a composite architecture combining traditional FFN blocks with point-wise convolution operations and spatial mixing mechanisms. This composite structure leverages the strengths of each component to achieve high accuracy with reduced parameter count, as the point-wise convolutions provide efficient feature transformation with lower computational overhead than standard FFN layers.
Data Source
AI summary
A processor-implemented method for image or text processing includes receiving, by an artificial neural network (ANN) model, a set of tokens corresponding to an input. A token interaction block of the ANN model processes the set of tokens according to each channel of the input to generate a spatial mixture of a set of features for each channel of the input. A feed forward network block of the ANN model generates a mixture of channel features based on the spatial mixture of the set of features for each channel of the input. An attention block of the ANN model determines a set of attended features of the mixture of channel features according to a set of attention weights. In turn, the ANN model generates an inference based on the set of attend features of the mixture of channel features.


