Propagating attention information in efficient machine learning model
By introducing an attention propagation mechanism into the converter architecture, the problem of high computing costs of conventional vision converters is solved, and efficient computing on low-power devices is achieved while maintaining the accuracy of the model.
Patent Information
- Application Number
- CN202380076525.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2023-08-23
- Publication Date
- 2025-06-20
AI Technical Summary
Conventional vision converters are cost-effective in computing, especially in low-power devices, because their self-attention calculations are squared with the size of the input data.
By introducing an attention propagation mechanism into the converter architecture, the self-attention output of one converter block is propagated to other converter blocks, reducing repeated calculations and thus reducing calculation overhead.
This approach significantly reduces the computational overhead of the model while maintaining or improving the accuracy of the model, and is suitable for resource-constrained devices.
Smart Images

Figure CN120188166A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Non - Provisional Patent Application No. 18 / 335,685, filed on June 15, 2023, which claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 424,789, filed on November 11, 2022, the entire contents of both of which are incorporated herein by reference.
[0003] Introduction
[0004] Aspects of the present disclosure relate to machine learning.
[0005] A variety of machine - learning architectures have been used to provide solutions to a wide variety of computational problems. There are a wide variety of machine - learning model architectures, such as artificial neural networks (which may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), generative adversarial networks (GANs), etc.), random forest models, etc. Transformer architectures have been applied in natural language processing (NLP) and computer vision. Vision transformers have been increasingly widely used in a variety of image and video processing tasks.
[0006] However, some conventional vision transformers tend to be computationally expensive. For example, because vision transformers typically compute self - attention at each block, the computational and memory requirements grow quadratically with respect to the size of the input data. Thus, despite their accuracy and effectiveness, vision transformers have historically had limited or no applicability to low - power (or otherwise constrained) devices. Some conventional solutions have failed to provide an efficient and accurate transformer architecture. Summary of the Invention
[0007] Certain aspects provide a method that includes: using a first transformer block of a plurality of transformer blocks to generate a first attention propagation output, the generating including processing input data for the first transformer block using a first self - attention sub - block of the first transformer block; propagating the first attention propagation output to a second transformer block of the plurality of transformer blocks; and generating an output for the second transformer block, the generating the output for the second transformer block including generating output features for the second transformer block based on the first attention propagation output.
[0008] In other aspects, provided are: a processing system configured to perform the methods mentioned above and those described herein; a non-transitory computer-readable medium including instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the methods mentioned above and those described herein; a computer program product embodied on a computer-readable storage medium, the computer program product including code for performing the methods mentioned above and those further described herein; and a processing system including components for performing the methods mentioned above and those further described herein.
[0009] The following description and the related drawings set forth in detail certain illustrative features of one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings depict example features of certain aspects of the present disclosure and are thus not to be considered as limiting the scope of the present disclosure.
[0011] Figure 1 Depicts an example architecture for the propagation of attention outputs generated by a transducer.
[0012] Figure 2A and Figure 2B Depicts an example transducer architecture for propagating attention data.
[0013] Figure 3A and Figure 3B Depicts an example transducer architecture for propagating attention feature data.
[0014] Figure 4 Depicts an example encoder / decoder architecture using a transducer and attention propagation.
[0015] Figure 5 Is a flowchart depicting an example method for generating and providing an attention propagation output.
[0016] Figure 6 Is a flowchart depicting an example method for generating an output using propagated attention information.
[0017] Figure 7 Is a flowchart depicting an example method for generating an output using propagated feature information.
[0018] Figure 8 Is a flowchart depicting an example method for propagating attention outputs in a transducer architecture.
[0019] Figure 9 Depicts an example processing system configured to perform various aspects of the present disclosure.
[0020] For ease of understanding, the same reference numerals have been used, where possible, to designate the same elements common to the various figures. Elements and features of one aspect can be beneficially incorporated into other aspects without further recitation. Detailed Description
[0021] Aspects of the present disclosure provide apparatus, methods, processing systems, and non-transitory computer-readable media for improved transformer-based models that use propagated attention information.
[0022] Transformers (and in particular vision transformers) have become increasingly prevalent in a wide variety of machine learning tasks. Transformer-based architectures are generally configured to generate an output based on a sequence of data (e.g., a sequence of frames in a video, a sequence of patches from a frame or image, a sequence of words, a sequence of audio data, etc.). Generally speaking, a machine learning model can use any number of transformer blocks (each transformer block providing self-attention) and any other components (e.g., one or more neural network layers).
[0023] In some aspects, a transformer includes a self-attention component for generating self-attention. In some aspects, a transformer may also include a feed-forward component (e.g., a small neural network such as a multi-layer perceptron (MLP)). In the self-attention component, the input data can be linearly projected (e.g., multiplied using learned parameters) into three matrices: a query matrix Q (also referred to in some aspects as a query representation or simply "query"), a key matrix K (also referred to in some aspects as a key representation or simply "key"), and a value matrix V (also referred to in some aspects as a value representation or simply "value"). For example, during training, one or more query weights, key weights, and value weights are learned based on training data, and the query Q, key K, and value V can be generated by multiplying the input data by the learned weights.
[0024] In some aspects, an attention matrix A (also referred to in some aspects as an attention map or simply "attention") is subsequently generated based on the query and the key. For example, a self-attention block can compute the dot product of the query matrix and the transposed key matrix (e.g., Q·K T ). In some aspects, a self-attention block can apply one or more operations (e.g., a row-wise softmax operation) to the dot product to produce the attention matrix. That is, the attention matrix can be defined as A = σ(Q·K T ), where σ is the softmax function.
[0025] In some aspects, the features f generated by the self-attention block can subsequently be computed as the dot product of the attention matrix A and the value matrix V. These features can subsequently be provided as input to a feed-forward block (e.g., a neural network or sub-network) to generate the output from the transformer block.
[0026] Such transformer blocks are computationally expensive. For example, the computation and memory of self-attention components scale quadratically with the size of the input. In a conventional architecture that typically has multiple transformers arranged sequentially, each transformer can independently compute self-attention at each block, incurring significant computational costs.
[0027] In aspects of the present disclosure, information related to the self-attention output of one or more transformers in an architecture can be propagated to one or more other transformers in the architecture and reused by these receiving transformers. Such a dependency on the self-attention map across transformer blocks can significantly reduce the computational overhead of the model while maintaining high accuracy. That is, some or all of the computations performed in computing self-attention at one transformer can be propagated and reused by one or more other transformers, thereby improving the computational efficiency of the model. Additionally, in some aspects, this propagation can act as a form of regularization, thereby improving model generalization and robustness. Additionally, in some aspects, the independently computed self-attention at two transformers can be highly similar (e.g., having a high cosine similarity). In such cases, propagating the attention can have a negligible impact on model accuracy (and in some deployments can even improve accuracy).
[0028] As used herein, the attention propagation output (also referred to as attention information) can refer to any part of the self-attention information, including the attention matrix A, the feature f, etc. For example, a first (e.g., upstream) transformer can compute the self-attention feature as f = A·V, where f is the feature at the first transformer, A is the attention matrix at the first transformer, and V is the value matrix at the first transformer. A second (e.g., downstream) transformer can reuse some or all of the attention propagation output. For example, the second transformer can compute the self-attention feature as f′ = A·V′ or as f′ = Φ(A)·V′, where f′ represents the feature at the second transformer, A is the attention matrix received from the first transformer, Φ is a transformation function (e.g., upsampling, convolution, etc.), and V′ is the value matrix at the second transformer. By avoiding at least some of the computations of the attention matrix at the second transformer, the computational overhead is significantly reduced. As another example, the second transformer can compute the self-attention feature as f′ = f or as f′ = Φ(f), where Φ is a transformation function (e.g., upsampling, convolution, etc.). That is, the second transformer can reuse all of the attention features generated by the first transformer module rather than just using the attention map.
[0029] In some aspects, rather than directly propagating the attention information (e.g., using an identity mapping), the architecture can use various propagation operations (such as one or more convolution operations) to dynamically modify or update the information before providing the attention information to the downstream transformer.
[0030] By using such attention propagation, aspects of the present disclosure enable a significant reduction in computational overhead and an improvement in transformer efficiency. This can enable transformer-based models to be effectively used on resource-constrained devices such as wearable devices and other mobile devices.
[0031] Example architecture for the propagation of attention outputs generated by a transducer
[0032] Figure 1 An example architecture 100 for the propagation of attention outputs generated by a transformer is depicted.
[0033] In the illustrated example, a machine learning model architecture including a series of transformers 110A through 110N (collectively referred to as transformers 110, which in some aspects are also referred to as transformer blocks and / or transformer components) is depicted. Generally, any number of transformers 110 may be present in architecture 100. Additionally, although not included in the illustrated example, in various aspects, various other components or blocks (e.g., one or more neural networks, multi-layer perceptrons, etc.) may be present in architecture 100.
[0034] As illustrated, input data 105 is provided as input to architecture 100 to generate output data 135. The specific form and content of input data 105 and output data 135 may vary depending on the particular implementation. For example, in some aspects, input data 105 and output data 135 include image data. Output data 135 may also be used to generate processing results such as image classification, machine translation, object detection, speech recognition, etc.
[0035] In the depicted architecture 100, each of transformers 110A through 110N includes a corresponding self-attention sub-block 115 (which in some aspects is also referred to as a self-attention block and / or self-attention component) and a feed-forward sub-block 120 (which in some aspects is also referred to as an MLP block, MLP component, MLP sub-block, feed-forward block, and / or feed-forward component). Specifically, transformer 110A includes self-attention sub-block 115A and feed-forward sub-block 120A, transformer 110B includes self-attention sub-block 115B and feed-forward sub-block 120B, and transformer 110N includes self-attention sub-block 115N and feed-forward sub-block 120N. Although not depicted in the illustrated example, in some aspects, each transformer 110 may include one or more residual connections, such as to perform an element-wise sum of the input to the self-attention sub-block 115 and the output from the same self-attention sub-block 115, to perform an element-wise sum of the input to the feed-forward sub-block 120 and the output from the same feed-forward sub-block 120, and so on.
[0036] In various aspects of the present disclosure, operations performed by a "self-attention" component or block generally may include any variant of attention, including factored self-attention, local self-attention, hybrid self-attention, window self-attention, etc. In some aspects, the self-attention sub-block 115 may implement multi-head self-attention (MSA).
[0037] Although not depicted in the illustrated example, in some aspects, the input data 105 (e.g., an image) is first preprocessed to generate a tokenized image (also referred to as a set of features) before being used as an input to the first transformer 110A. For example, assume the input data 105 is an image where h×w is the spatial resolution (e.g., height and width in pixels) and c is the number of channels. In some aspects, the input data 105 may first be tokenized into n = hw / p 2 non-overlapping patches, where p×p is the patch size. Each patch may then be projected into an embedding such as by using a linear layer to obtain the tokenized image where ";" represents stacking row by row, and d represents the dimension of the vector space The tokenized image may then be used as an input to the first transformer 110A. In some aspects, positional embeddings may additionally be added to the tokenized image to preserve positional information. In some aspects (e.g., in a supervised setting), learned tokens may also be prepended to the tokenized image.
[0038] In a conventional architecture, as discussed above, each self-attention sub-block 115A to 115N (collectively referred to as self-attention sub-block 115) may compute self-attention independently of each other. However, in the illustrated example, the transformer 110A outputs an attention propagation output 125, which is propagated to one or more downstream transformers (e.g., propagated to transformer 110B and / or transformer 110N) using a propagation operation 130. Although the illustrated example depicts the attention propagation output 125 being propagated to both transformer 110B and transformer 110N, in various aspects, one or more of the transformers 110 may receive propagated attention information from one or more other transformers, or may also compute their own self-attention instead of receiving propagated attention information.
[0039] In some aspects, the transformer 110 that propagates the attention output and the transformer 110 that receives the propagated attention information may vary depending on the specific implementation. In at least one aspect, the architecture 100 may propagate attention information from one or more transformer blocks in the sequence of transformer blocks of the model to one or more subsequent intermediate transformer blocks. That is, since the attention within such intermediate transformers (e.g., after one or more initial transformers but before one or more final or output transformers) can generally be relatively similar, the architecture 100 may include propagating attention information from a first intermediate block (e.g., the second or third transformer in the sequence) to one or more subsequent (intermediate) blocks.
[0040] In some aspects, the specific content of the attention propagation output 125 may vary depending on the specific implementation. For example, in at least one aspect, the attention propagation output 125 corresponds to the attention values (e.g., attention matrix A) generated by the self-attention sub-block 115A of the transformer 110A and used to generate the output features. In some aspects, the attention propagation output 125 corresponds to the output features themselves (e.g., feature f) generated by the self-attention sub-block 115A of the transformer 110A and provided as input to the feed-forward sub-block 120A.
[0041] In some aspects, the specific operation or transformation applied by the propagation operation 130 may vary depending on the specific implementation. For example, in some aspects, the propagation operation 130 is an identity operation that simply maps or provides the attention propagation output 125 to the downstream transformer 110 without modification. In some aspects, the propagation operation 130 may include one or more transformations. In some aspects, the transformation applied by the propagation operation 130 can ensure that directly reusing the attention and / or features from a previous transformer 110 does not affect the translational invariance and equivariance in subsequent blocks, and can further act as a strong regularizer to improve model generalization. For example, the transformation can be implemented by a parametric function, where one or more learned parameters are refined during training to improve generalization and maintain invariance and equivariance (compared to simply copying features).
[0042] In some aspects, the propagation operation 130 may be implemented using parameters or parameterized functions. That is, in at least one aspect, the propagation operation 130 may be performed based, at least in part, on one or more learned parameters (such as weights). In some aspects, the propagation operation 130 may include one or more convolutional operations. In some aspects, the propagation operation 130 may be implemented as an inverted bottleneck operation (e.g., a fully connected linear layer of a neural network, followed by a depthwise convolution, and then followed by another fully connected layer). For example, the propagation operation 130 may upsample the input data in the channel dimension (e.g., from d to 4d), apply a depthwise convolution (e.g., using a 5×5 kernel), and then downsample back to the d dimension. In some aspects, different propagation operations 130 may be performed for each downstream transformer, and / or the same operation may be used with different learned parameters for each downstream transformer.
[0043] In some aspects, the propagation operation 130 may include other upsampling and / or downsampling operations. For example, in a window self-attention implementation (or other architecture) where the output features of the self-attention sub-block 115A are different from the output features received from the self-attention sub-blocks 115B to 115N and / or the feed-forward sub-blocks 120B to 120N, the propagation operation 130 may appropriately upsample the attention propagation output 125 (e.g., to match the size and dimensions of the output of the received self-attention sub-block), or the attention propagation output 125 may be concatenated with the input to the downstream transformer.
[0044] In the illustrated architecture 100, the propagated attention information (e.g., the attention propagation output 125 and / or the attention propagation output 125 modified via the propagation operation 130 (also referred to as the modified attention propagation output)) is provided to each of the downstream transformers 110B to 110N. Generally, the specific techniques for providing and / or using the propagated attention information may vary depending on the specific content of the information (e.g., whether the content includes an attention matrix or the output features of the self-attention sub-block 115A), the architecture 100, and the specific implementation.
[0045] For example, in some aspects, if the attention propagation output 125 corresponds to the attention matrix of the self-attention sub-block 115A, the propagated attention information may be used as the attention matrix for the self-attention sub-blocks 115B to 115N. That is, the architecture 100 may only compute the value matrices for the self-attention sub-blocks 115B to 115N and compute the dot product with the propagated attention matrix from the self-attention sub-block 115A. In this way, the self-attention sub-blocks 115B to 115N do not need to generate their own attention matrices, thus significantly reducing the computational overhead of the architecture 100.
[0046] As another example, in some aspects, if the attention propagation output 125 corresponds to the output features of the self-attention sub-block 115A, then the propagated attention information can be used as the output features of the self-attention sub-blocks 115B to 115N, and / or can be combined with the output features of the self-attention sub-blocks 115B to 115N (or with the inputs to the self-attention sub-blocks 115B to 115N). That is, the transformers 110B to 110N can simply use the propagated output feature information as the features originally generated by the self-attention sub-blocks 115B to 115N, and these self-attention sub-blocks can thus be inhibited from performing any operations (or can be completely omitted).
[0047] In some aspects, the output features propagated from the self-attention sub-block 115A can be aggregated or combined with the input features to the transformers 110B to 110N. For example, the features output by the feed-forward sub-block 120A (which are conventionally provided to the self-attention sub-block 115B) can instead be combined with the propagated output features from the self-attention sub-block 115A (e.g., concatenated together). Subsequently, this combined data can be provided as an input to the feed-forward sub-block 120B. Thus, in some aspects, the self-attention sub-block 115B may not be needed. Similarly, the features output by the transformer immediately preceding the transformer 110N (which are conventionally provided to the self-attention sub-block 115N) can instead be combined with the propagated output features from the self-attention sub-block 115A (e.g., concatenated together). Subsequently, this combined data can be provided as an input to the feed-forward sub-block 120N, and the self-attention sub-block 115N may not be needed.
[0048] Generally speaking, as discussed above, the attention propagation implemented in the architecture 100 can thus achieve a significantly reduced computational burden (e.g., reduced memory footprint, reduced number of operations, etc.), while maintaining or enhancing model accuracy.
[0049] Example transducer architecture for propagating attention data
[0050] Figure 2A and Figure 2B depicts an example transformer architecture for propagating attention data. In some aspects, Figure 2A the architecture 200A provides additional details of the transformer that propagates or provides attention information to other transformers. For example, the architecture 200A can provide Figure 1 additional details of the transformer 110A. Similarly, Figure 2B the architecture 200B provides additional details of the transformer that receives or accesses the propagated attention information from other transformers. For example, the architecture 200B can provide Figure 1 additional details of the transformer 110B.
[0051] AsFigure 2A As illustrated, the input data 205A is accessed by the transformer 110A. As used herein, "accessing" data generally can include receiving, retrieving, requesting, or otherwise obtaining access to the data. As discussed above, the input data 205A can correspond to the input to the first transformer of the model (e.g., raw or preprocessed input data), the output of a previous transformer or other model component or block, and so on. For example, as discussed above, the input data 205A can correspond to a tokenized input image (which can optionally include positional embeddings and / or learned tokens).
[0052] As discussed above, the transformer 110A includes a self-attention sub-block 115A and a feed-forward sub-block 120A. Within the self-attention sub-block 115A, the input data 205A is used to generate a query matrix 210A, a key matrix 215A, and a value matrix 220A. For example, the self-attention sub-block 115A can use learned weights or other parameters to generate the matrices based on the input data (e.g., using linear projections).
[0053] As indicated by operation 225A (which can correspond to a dot product and / or transpose operation), the attention matrix 230A for the self-attention sub-block 115A is generated based on the query matrix 210A and the key matrix 215A. For example, as discussed above, the attention matrix 230A can be generated by computing the dot product between the transposed key matrix 215A and the query matrix 210A. Although not depicted in the illustrated example, in some aspects, the result of the dot product can be further processed or transformed, such as using a softmax operation (or other operation), to form the attention matrix 230A.
[0054] As indicated by operation 235A (which can correspond to a dot product), the output features 240A (also referred to as attention features) for the self-attention sub-block 115A can be generated based on the value matrix 220A and the attention matrix 230A.
[0055] These output features 240A are then provided as input to the feed-forward sub-block 120A, which generates an output 250A from the transformer 110A. This output 250A can then be used as input to a subsequent block or component, such as input to Figure 1 the transformer 110B. In some aspects, as discussed above, the feed-forward sub-block 120A includes an MLP that includes two linear layers separated by a Gaussian error linear unit (GeLU) activation function.
[0056] Although not depicted in the illustrated example, in some aspects, the transformer 110A may include one or more skip or residual connections (with or without layer normalization). For example, the output feature 240A may be generated by summing the output of operation 235A with the input data 205A. As another example, the output 250A of the transformer 110A may be generated by summing the output of the last layer of the feed-forward sub-block 120A with the output feature 240A.
[0057] In the illustrated example, the attention matrix 230A is further provided to one or more downstream transformers via link 245. That is, the attention matrix 230A may correspond to Figure 1 the attention propagation output 125. In some aspects, as discussed above, the attention matrix 230A is processed using one or more propagation operations before being provided to the downstream transformer. For example, the propagation operation may include an identity operation, one or more convolution operations, bottleneck operations, upsampling or downsampling operations, etc.
[0058] In this way, the architecture 200A enables the attention matrix 230A to be used by one or more subsequent transformers, thereby enabling the attention matrix to bypass this calculation and significantly improving the efficiency of the model.
[0059] Now turning to Figure 2B , the input data 205B is accessed by the transformer 110B. Generally speaking, the transformer 110B may be downstream or after the transformer 110A. That is, if the model includes a sequence of transformer blocks, the transformation performed by the operation of the transformer 110A may be executed relatively early in the sequence compared to the operation of the transformer 110B. However, it should be understood that the transformer 110B does not need to be the immediately adjacent or the immediately following transformer relative to the transformer 110A. In some aspects, the input data 205B (e.g., the output of the feed-forward sub-block of the previous transformer) may be received from a previous model component such as a previous transformer.
[0060] As discussed above, the transformer 110B includes a self-attention sub-block 115B and a feed-forward sub-block 120B. As illustrated, the self-attention sub-block 115B may receive the propagated attention matrix 230A from the transformer 110A via link 245, rather than independently computing the attention (e.g., generating an attention matrix based on the input data 205B). Although for the sake of conceptual clarity, the illustrated example depicts receiving the attention matrix 230A itself (e.g., via an identity operation), as discussed above, the propagation operation may additionally or alternatively include one or more transformations or modifications of the attention matrix 230A before providing the attention matrix to the self-attention sub-block 115B.
[0061] In this way, within the self-attention sub-block 115B, the input data 205B is used to generate the value matrix 220B. Thus, in some aspects, there is no need to generate a query matrix and a key matrix for the self-attention sub-block 115B. For example, the self-attention sub-block 115B can use learned weights or other parameters to generate the value matrix 220B based on the input data (e.g., using linear projection).
[0062] Furthermore, the self-attention sub-block 115B does not need to perform dot product, transpose, or softmax operations on any query or key matrix, because the (potentially transformed) attention matrix 230A from the transformer 110A is reused.
[0063] As indicated by operation 235B (which can correspond to a dot product), the output features 240B for the self-attention sub-block 115B can thus be generated based on the value matrix 220B and the propagated attention matrix 230A. These output features 240B can then be provided as input to the feed-forward sub-block 120B, which generates the output 250B from the transformer 110B. This output 250B can then be used as input to subsequent blocks or components, or as the output from the model.
[0064] Although not depicted in the illustrated example, in some aspects, the transformer 110B can include one or more skip or residual connections (with or without layer normalization), as discussed above. For example, the output features 240B can be generated by summing the output of operation 235B with the input data 205B. As another example, the output 250B of the transformer 110B can be generated by summing the output of the last layer of the feed-forward sub-block 120B with the output features 240B.
[0065] Thus, in the illustrated example, compared with a conventional architecture that independently calculates attention at each transformer, the computational resources consumed when processing the input data 205B in the transformer 110B are significantly reduced. In this way, the architecture 200B enables a significant improvement in the efficiency of the model.
[0066] Example transducer architecture for propagating attention feature data
[0067] Figure 3A and Figure 3B depicts an example transformer architecture for propagating attention feature data. In some aspects, Figure 3A the architecture 300A provides additional details of the transformer that propagates or provides attention information to other transformers. For example, the architecture 300A can provide Figure 1 additional details of the transformer 110A. Similarly, Figure 3B the architecture 300B provides additional details of the transformer that receives or accesses the propagated attention information from other transformers. For example, the architecture 300B can provideFigure 1 Additional details of the transformer 110B.
[0068] As Figure 3A illustrated, the input data 305A is accessed by a transformer 110A having a self-attention sub-block 115A and a feed-forward sub-block 120A. For example, as discussed above, the input data 305A may correspond to a tokenized input image (which may optionally include positional embeddings and / or learned tokens). In some aspects, the operations of the self-attention sub-block 115A and the feed-forward sub-block 120A may generally correspond to the operations discussed above with reference to Figure 2A For example, the input data 305A can be used to generate a query matrix 310A, a key matrix 315A, and a value matrix 320A, which can then be used (via operations 325A and 335A, respectively) to generate an attention matrix 330A and output features 340A. The output features 340A are then provided to the feed-forward sub-block 120A, which processes the features 340A to generate an output 350A for the transformer 110A. Although not depicted in the illustrated example, in some aspects, the transformer 110A may also include one or more residual or skip connections (e.g., for the self-attention sub-block 115A and / or for the feed-forward sub-block 120A), as discussed above.
[0069] In the illustrated example, the output features 340A are further provided to one or more downstream transformers via a link 345. That is, the output features 340A may correspond to Figure 1 the attention propagation output 125. In some aspects, as discussed above, the output features 340A are processed using one or more propagation operations before being provided to the downstream transformers. For example, the propagation operations may include an identity operation, one or more convolutional operations, bottleneck operations, upsampling or downsampling operations, etc.
[0070] In some aspects, the link 345 includes parameters or a parameterized function that maps the output of the self-attention sub-block 115A (e.g., the output features 340A) for use by one or more subsequent or downstream transformers. In some aspects, as discussed above, the link 345 is an identity function. In some aspects, the link 345 may encode local relationships between tokens. For example, the link 345 may include a series of convolutions, such as two linear layers with a depthwise convolution between the linear layers. For example, if the link 345 includes an identity function, this may be useful because the representation learning may be affected due to the lack of self-attention at the downstream transformers (e.g., due to the fact that the downstream transformers reuse the output features 340A without modification and without generating their own attention), as the relationships across tokens may no longer be encoded in the attention matrix.
[0071] In some aspects, with respect to supervised learning, the learned tokens (discussed above, which may be pre-placed in the input data) can be divided into class embeddings and patch embeddings, where the patch embeddings are input into the first linear layer of link 345. This first layer can expand the channel dimension, and subsequently each depthwise convolutional layer can convolve the data using an r×r kernel to capture cross-token relationships. Note that in some aspects, prior to the depthwise convolutional operation, the input matrix can be spatially reshaped into a feature tensor. The output of the depthwise operation can then be flattened back into a vector and provided to the last linear / full-connected layer of link 345, which reduces the channel dimension back to the initial depth.
[0072] In this way, architecture 300A enables output features 340A to be used by one or more subsequent transformers, thereby enabling these output features to bypass such computations and significantly improving the efficiency of the model.
[0073] Now turning to Figure 3B , the input data 305B is accessed by a transformer 110B having a self-attention sub-block 115B and a feed-forward sub-block 120B. Generally speaking, transformer 110B can be downstream or after transformer 110A. That is, if the model includes a sequence of transformer blocks, then transformer 110A can be executed relatively early in the sequence compared to transformer 110B. However, it should be understood that transformer 110B does not need to be the immediately adjacent or immediately following transformer relative to transformer 110A.
[0074] As discussed above, transformer 110B includes a self-attention sub-block 115B and a feed-forward sub-block 120B. As illustrated, rather than independently computing the attention or output features, the self-attention sub-block 115B can be not used (and in some embodiments may be omitted) and instead the propagated output features 340A from transformer 110A (accessed via link 345) can be used. Although for clarity of concept the illustrated example depicts receiving the output features 340A themselves (e.g., via an identity operation), as discussed above, the propagation operation can additionally or alternatively include one or more transformations or modifications to the output features 340A prior to providing the (modified) output features to transformer 110B.
[0075] As illustrated, instead of using the self-attention sub-block 115B to compute self-attention, the transformer 110B can combine the input data 305B and the propagated output features 340A via operation 355. For example, operation 355 can correspond to concatenation, element-wise addition or summation, averaging, etc. As illustrated, the resulting output is then provided as input to the feed-forward sub-block 120B, which generates the output data 350B from the transformer 110B. In some aspects, instead of combining the input data 305B and the propagated output features 340A, the propagated output features 340A themselves can be used as input to the feed-forward sub-block 120B, and the input data 305B can be not used or discarded. For example, the output features 340A can be upsampled and provided to the feed-forward sub-block 120B to generate the output data 350B, and both the input data 305B and the self-attention sub-block 115B can be not used, discarded, or omitted. As illustrated, the output data 350B can then be used as input to subsequent blocks or components, or as the output from the model.
[0076] Although not depicted in the illustrated example, as discussed above, in some aspects, the transformer 110B can include one or more skip or residual connections (with or without layer normalization). For example, the summation operation 355 can effectively implement a residual connection to generate the input to the feed-forward sub-block 120B. Additionally, the output data 350B of the transformer 110B can be generated by summing the output of the last layer of the feed-forward sub-block 120B with the output of operation 355. In this way, due to the presence of such residual connections, the transformer 110 that re-uses previous attention data can still independently learn representations and still provide useful computations, even without performing independent self-attention.
[0077] Thus, in the illustrated example, compared to conventional architectures that independently compute attention at each transformer, the computational resources consumed when processing the input data 305B in the transformer 110B are significantly reduced. In this way, the architecture 300B enables a significant improvement in the efficiency of the model.
[0078] Example machine learning architecture using a transducer and attention propagation
[0079] Figure 4 An example machine learning architecture 400 that uses transformers and attention propagation is depicted. In one aspect, the architecture 400 corresponds to an encoder-decoder architecture, where the input data 405 is successively processed using one or more encoder transformers 410A to 410N (collectively referred to as encoder transformers 410), and then using one or more decoder transformers 450A to 450N (collectively referred to as decoder transformers 450) to generate the output data 435. The output data 435 can be used to obtain a wide variety of processing results, such as image classification, machine translation, object detection, speech recognition, etc.
[0080] As illustrated, the input data 405 is accessed by the first encoder transformer 410A and processed using the self-attention sub-block 415A to generate output features (also referred to as attention features in some aspects). These features are then processed by the feed-forward sub-block 420A to generate the output from the encoder transformer 410A. In the illustrated example, the output from the encoder transformer 410A is provided as input to the encoder transformer 410B (including the self-attention sub-block 415B and the feed-forward sub-block 420B). Although not depicted in the illustrated example, in some aspects, the transformers 410 and 450 may include one or more residual connections for the self-attention sub-blocks 415 and 455 and / or for the feed-forward sub-blocks 420 and 460, as discussed above.
[0081] In the illustrated example, as indicated by the link 430A, the output from the encoder transformer 410A is also concatenated with the input to the corresponding decoder transformer 450A. That is, the output of the encoder transformer 410A is concatenated with the output of the decoder transformer 450B, and the concatenated data is provided as input to the decoder transformer 450A (which includes the self-attention sub-block 455A and the feed-forward sub-block 460A).
[0082] In addition, as illustrated by the link 425A, the attention propagation output is further provided from the self-attention sub-block 415A to the self-attention sub-block 455A. For example, the attention map or matrix generated by the self-attention sub-block 415A can be propagated to the self-attention sub-block 455A, allowing the self-attention sub-block 455A to reuse the attention information and suppress independently computing self-attention. In some aspects, as discussed above, this propagation operation may include an identity mapping or may include various transformations such as upsampling, convolution, etc. Although the illustrated example depicts the propagation of attention information from the encoder to the decoder, in some aspects, attention may additionally or alternatively be propagated to other transformers (e.g., from the first encoder to subsequent encoders).
[0083] In the illustrated example, each encoder transformer 410 propagates its output to a corresponding decoder transformer 450 (as indicated by links 430A through 430N), and further propagates attention information to the corresponding decoder transformer 450 (as indicated by links 425A through 425N). Specifically, in addition to encoder transformer 410A and decoder transformer 450A discussed above, encoder transformer 410B (including self-attention sub-block 415B and feed-forward sub-block 420B) also propagates attention information via link 425B and outputs to decoder transformer 450B (including self-attention sub-block 455B and feed-forward sub-block 460B) via link 430B, and encoder transformer 410N (including self-attention sub-block 415N and feed-forward sub-block 420N) propagates attention information via link 425N and outputs to decoder transformer 450N (including self-attention sub-block 455N and feed-forward sub-block 460N) via link 430N.
[0084] Although three encoder transformers 410 and three decoder transformers 450 are depicted for clarity of concept, any number of encoders and decoders may be present in architecture 400. Additionally, in various aspects, there may be one or more components or blocks between the final encoder transformer and the first decoder transformer, or the final encoder transformer may provide its output directly to the first decoder transformer 450.
[0085] Generally speaking, in the illustrated example, the corresponding decoder transformer 450 for each encoder transformer 410 is defined based on the sequence of transformers used. For example, the first encoder transformer 410A corresponds to the final decoder transformer 450A, the second encoder transformer 410B corresponds to the penultimate decoder transformer 450B, the final encoder transformer 410N corresponds to the first decoder transformer 450N, and so on.
[0086] In this way, architecture 400 enables the attention information from one or more encoder transformers 410 to be propagated and reused by one or more subsequent decoder transformers 450 (or by other encoder transformers 410), and / or enables the attention information from one or more decoder transformers 450 to be propagated and reused by one or more subsequent decoder transformers 450. As discussed above, this significantly reduces the computational overhead of architecture 400 while further maintaining or improving model accuracy.
[0087] Example method for generating and providing attention propagation output
[0088] Figure 5is a flowchart depicting an example method 500 for generating and providing an attention propagation output. In some aspects, method 500 is performed by an upstream transformer that generates self-attention that will be propagated to one or more downstream transformers. For example, method 500 may be performed by the transformer 110A of Figure 1 , Figure 2A and / or Figure 3A , by the encoder transformers 410A to 410N of Figure 4 , and / or by the decoder transformers 450A to 450N of Figure 4 .
[0089] At block 505, the transformer accesses input data. As discussed above, this input may include the input to the model itself, the output from previous blocks or components of the model (e.g., previous transformers), etc. For example, the accessed input may correspond to the input data 105 of Figure 1 , the input data 205A of Figure 2A , the input data 305A of Figure 3A , the input data 405 of Figure 4 , the output of the encoder transformer 410 of Figure 4 , etc. In some aspects, the accessed input data corresponds to tokenized image data, as discussed above. The accessed data is typically used as the input to the transformer and is processed to generate the output from the transformer (e.g., using self-attention and / or a feed-forward network).
[0090] At block 510, the transformer generates key, query, and value matrices for the transformer based on the input. For example, as discussed above with reference to Figure 2A and Figure 3A , the transformer may use learned parameters and / or linear projections to generate the matrices (e.g., by multiplying the accessed input by learned values).
[0091] At block 515, the transformer generates an attention (e.g., an attention map or matrix) based on the generated keys and queries. For example, as discussed above, the transformer may compute the dot product of the key matrix and the transposed query matrix and apply a softmax operation (or other operation) to the result to generate the attention matrix.
[0092] At block 520, the transformer generates output features (also referred to as attention features or self-attention features, as discussed above) based on the generated attention and values. For example, as discussed above, the transformer may compute the dot product of the attention matrix and the value matrix. In some aspects, as discussed above, the output features may further be generated using a residual connection (such as by aggregating the input data (accessed at block 505) and the features (generated based on the attention and values) using element-wise summation).
[0093] At block 525, the transducer may generate an output of the transducer based on the generated features. For example, as discussed above, the transducer may use a feed-forward sub-block (e.g., MLP) of the transducer to process the features to generate an output. In some aspects, as discussed above, the output of the transducer may be further generated using a residual connection (such as by aggregating the output features (generated at block 520) and the output of the feed-forward sub-block using element-wise summation).
[0094] At block 530, the transducer provides an attention propagation output (e.g., an attention matrix and / or output features) as an output. For example, via a propagation operation (e.g., an identity mapping or a skip connection, one or more convolutional operations, or one or more other transformations), the attention propagation output may be provided to one or more downstream transducers and reused by one or more downstream transducers.
[0095] At block 535, the transducer similarly provides the generated output (generated at block 525) as an output from the transducer and inputs it to a subsequent or neighboring downstream component (e.g., inputs it to the next transducer in a sequence of transducers).
[0096] In this way, method 500 enables the attention information generated by the transducer to be propagated and shared and / or reused by one or more downstream components, thereby significantly reducing computational overhead and maintaining or improving model accuracy and robustness.
[0097] Example method for generating an output using propagated attention information
[0098] Figure 6 is a flowchart depicting an example method 600 for generating an output using propagated attention information. In some aspects, method 600 is performed by a downstream transducer that receives and uses the propagated attention information from one or more upstream transducers. For example, method 600 may be performed by Figure 1 , Figure 2B and / or Figure 3B of the transducer 110B and / or by Figure 4 of the encoder transducers 410A to 410N and / or by Figure 4 of the decoder transducers 450A to 450N.
[0099] At block 605, the transducer accesses input data. As discussed above, the input may include the output from a previous block or component of the model (e.g., a previous transducer). For example, the accessed input may correspond to Figure 1 of the transducer 110A / 110B, Figure 2A of the output 250A, Figure 2B of the input data 205B, Figure 3A of the output 350A, Figure 3Binput data 305B, Figure 4 the output of decoder-transformer 450, and so on. The accessed data is typically used as input to the transformer and processed to generate an output from the transformer (e.g., using self-attention and / or a feed-forward network).
[0100] At block 610, the transformer accesses attention information (e.g., attention propagation output) from one or more previous transformers. In some aspects, the attention propagation output corresponds to an attention map or matrix generated by a previous transformer, as discussed above. For example, the accessed input may correspond to Figure 1 the attention propagation output 125, Figure 2A and Figure 2B the attention matrix 230A, and so on. In some aspects, as discussed above, rather than receiving the attention matrix itself (e.g., via a skip connection), the received attention propagation output corresponds to a transformed or modified attention matrix (e.g., upsampled, processed using one or more convolutional operations, etc.).
[0101] At block 615, the transformer generates a value matrix based on the input accessed / received at block 605. For example, as discussed above with reference to Figure 2B the transformer may use learned parameters and / or a linear projection to generate the value matrix (e.g., by multiplying the accessed input by learned values).
[0102] At block 620, the transformer generates output features (e.g., self-attention features) based on the propagated attention matrix and the generated value matrix. For example, as discussed above, the transformer may compute the dot product of the propagated attention matrix and the generated value matrix. In some aspects, as discussed above, the output features may further be generated using a residual connection (such as by aggregating (using element-wise summation) the input data (accessed at block 605) and the features (generated based on attention and values)).
[0103] At block 625, the transformer may then generate an output of the transformer based on the generated features. For example, as discussed above, the transformer may use a feed-forward sub-block of the transformer (e.g., an MLP) to process the features to generate an output. In some aspects, as discussed above, the output of the transformer may further be generated using a residual connection (such as by aggregating (using element-wise summation) the output features (generated at block 620) and the output of the feed-forward sub-block).
[0104] At block 630, the transformer then provides the generated output (generated at block 625) as an output from the transformer and inputs it to a subsequent or neighboring downstream component (e.g., input to the next transformer in a sequence of transformers), or as an output from the model.
[0105] In this way, method 600 enables the attention information generated by the upstream transformer to be propagated and shared and / or reused by the downstream transformer, thereby significantly reducing the computational overhead and maintaining or improving the model accuracy and robustness.
[0106] Example method for generating an output using propagated feature information
[0107] Figure 7 is a flowchart depicting an example method 700 for generating an output using the propagated feature information. In some aspects, method 700 is performed by a downstream transformer that receives and uses the propagated attention information from one or more upstream transformers. For example, method 700 may be performed by Figure 1 , Figure 2B and / or Figure 3B transformer 110B and / or by Figure 4 encoder transformers 410A to 410N and / or by Figure 4 decoder transformers 450A to 450N.
[0108] At block 705, the transformer accesses the input data. As discussed above, this input may include the output from a previous block or component of the model (e.g., a previous transformer). For example, the accessed input may correspond to Figure 1 the output of transformer 110A / 110B, Figure 2A the output 250A, Figure 2B the input data 205B, Figure 3A the output 350A, Figure 3B the input data 305B, Figure 4 the output of decoder transformer 450, etc. The accessed data is typically used as the input to the transformer and is processed to generate an output from the transformer (e.g., using self-attention and / or a feed-forward network).
[0109] At block 710, the transformer accesses the attention information (e.g., attention propagation output) from one or more previous transformers. In some aspects, the attention propagation output corresponds to the output attention features generated by the previous transformer, as discussed above. For example, the accessed input may correspond to Figure 1 the attention propagation output 125, Figure 3A and Figure 3B the output features 340A, etc. of
[0110] At block 720, the transformer optionally generates output features (e.g., self-attention features) based on the propagated features. For example, as discussed above, the transformer may concatenate or sum the input data (received at block 705) with the propagated output features. In some aspects, as discussed above, the transformer may instead simply use the propagated features as its own attention output.
[0111] At block 725, the transformer may generate the output of the transformer based on the generated features (e.g., based on the concatenated or summed input and propagated features, or only on the propagated features). For example, as discussed above, the transformer may use a feed-forward sub-block (e.g., MLP) of the transformer to process the features to generate an output. In some aspects, as discussed above, the output of the transformer may be further generated using a residual connection (such as by aggregating the output features (generated at block 720) and the output of the feed-forward sub-block using element-wise summation).
[0112] At block 730, the transformer then provides the generated output (generated at block 725) as the output from the transformer and inputs it to a subsequent or neighboring downstream component (e.g., to the next transformer in a sequence of transformers), or as the output from the model.
[0113] In this way, method 700 enables the attention information generated by an upstream transformer to be propagated and shared and / or reused by downstream transformers, thereby significantly reducing computational overhead and maintaining or improving model accuracy and robustness.
[0114] Example method for propagating attention outputs in a transducer architecture
[0115] Figure 8 is a flowchart depicting an example method 800 for propagating attention outputs in a transformer architecture. In some aspects, method 800 is performed by one or more upstream transformers (such as Figure 1 , Figure 2A and / or Figure 3A 's transformer 110A, and / or by Figure 4 's encoder transformers 410A to 410N) that generate self-attention to be propagated to one or more downstream transformers, and / or is performed by one or more downstream transformers that receive and use the propagated attention information from one or more upstream transformers (e.g., Figure 1 , Figure 2B and / or Figure 3B 's transformer 110B, and / or by Figure 4 's decoder transformers 450A to 450N).
[0116] At block 805, a first attention propagation output is generated using a first transformer block among a plurality of transformer blocks, and the generation includes processing input data for the first transformer block using a first self-attention sub-block of the first transformer block.
[0117] At block 810, the first attention propagation output is propagated to a second transformer block among the plurality of transformer blocks.
[0118] At block 815, an output for the second transformer block is generated, and generating the output for the second transformer block includes generating output features for the second transformer block based on the first attention propagation output.
[0119] In some aspects, generating the first attention propagation output further includes generating an attention matrix using the first transformer block, and generating the attention matrix includes processing query representations and key representations of the input data for the first transformer block using the first self-attention sub-block.
[0120] In some aspects, the first attention propagation output includes the attention matrix.
[0121] In some aspects, generating the output features for the second transformer block further includes accessing the output for a third transformer block among the plurality of transformer blocks, where the third transformer block is immediately before the second transformer block, and generating the output features for the second transformer block using a second self-attention sub-block of the second transformer based on the first attention propagation output and value representations of the output for the third transformer block.
[0122] In some aspects, generating the first attention propagation output includes generating output features for the first transformer block using the first transformer block by processing the attention matrix and value representations of the input data for the first transformer block using the first self-attention sub-block.
[0123] In some aspects, the first attention propagation output includes the output features for the first transformer block.
[0124] In some aspects, method 800 further includes generating an output for the first transformer block using the first transformer block, and generating the output for the first transformer block includes processing output features of the first self-attention sub-block using a first feed-forward sub-block of the first transformer block.
[0125] In some aspects, generating the output for the second transformer block includes processing output features of the second self-attention sub-block using a second feed-forward sub-block of the second transformer block.
[0126] In some aspects, the first transformer block includes an encoder block and the second transformer block includes a decoder block.
[0127] In some aspects, the plurality of transformer blocks includes a sequence of transformer blocks, the sequence of transformer blocks includes one or more initial blocks, a plurality of intermediate blocks, and one or more final blocks, and the plurality of intermediate blocks includes a first transformer block and a second transformer block.
[0128] In some aspects, generating the first attention propagation output includes processing input data for the first transformer block using a plurality of window self-attention operations to generate output features for the first transformer block.
[0129] In some aspects, the first attention propagation output includes the output features for the first transformer block.
[0130] In some aspects, propagating the first attention propagation output to the second transformer block includes propagating the first attention propagation output using a propagation operation that includes transforming the first attention propagation output by concatenating the output features for a third transformer block in the plurality of transformer blocks to the first attention propagation output, and the third transformer block is immediately before the second transformer block.
[0131] In some aspects, propagating the first attention propagation output to the second transformer block includes propagating the first attention propagation output using a propagation operation, and the propagation operation includes transforming the first attention propagation output using an upsampling operation.
[0132] In some aspects, propagating the first attention propagation output to the second transformer block includes propagating the first attention propagation output using a propagation operation.
[0133] In some aspects, the propagation operation includes transforming the first attention propagation output by performing one or more convolution operations on the first attention propagation output.
[0134] In some aspects, the propagation operation includes an identity operation.
[0135] In some aspects, when generating the output features for the second transformer block, the second self-attention sub-block does not compute the attention matrix.
[0136] Example processing system for an efficient transducer architecture with attention propagation output
[0137] In some aspects, reference Figures 1 to 8 The workflows, techniques, and methods described herein may be implemented on one or more devices or systems. Figure 9 Depicted is an example processing system 900 configured to perform various aspects of the present disclosure, including, for example, the techniques and methods described with reference to Figures 1 to 8 In one aspect, the processing system 900 may use a transformer-based architecture to train, implement, or provide machine learning models, such as the architectures Figure 1 of 100, Figure 2Aarchitecture 200A, Figure 2B architecture 200B, Figure 3A architecture 300A, Figure 3B architecture 300B, and / or Figure 4 architecture 400. Although depicted as a single system for conceptual clarity, in at least some aspects, the operations described below with respect to processing system 900 can be distributed across any number of devices, as discussed above.
[0138] Processing system 900 includes a central processing unit (CPU) 902, which in some examples can be a multi-core CPU. Instructions executed at CPU 902 can be loaded, for example, from a program memory associated with CPU 902 or from a partition of memory 924.
[0139] Processing system 900 also includes additional processing components customized for specific functions, such as a graphics processing unit (GPU) 904, a digital signal processor (DSP) 906, a neural processing unit (NPU) 908, a multimedia processing unit 910, and a wireless connection component 912.
[0140] An NPU (such as NPU 908) is generally a dedicated circuit configured to implement the control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligent processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.
[0141] An NPU (such as NPU 908) is configured to accelerate the execution of common machine learning tasks, such as image classification, machine translation, object detection, and various other prediction models. In some examples, multiple NPUs can be instantiated on a single chip (such as a system-on-chip (SoC)), while in other examples, an NPU can be part of a dedicated neural network accelerator.
[0142] An NPU can be optimized for training or inference, or in some cases is configured to balance the performance between the two. For an NPU capable of performing both training and inference, these two tasks can generally still be performed independently.
[0143] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation involving inputting an existing dataset (usually labeled or tagged), iterating over the dataset, and then adjusting model parameters such as weights and biases to improve model performance. Generally speaking, optimizing based on incorrect predictions involves passing back through the layers of the model and determining gradients to reduce prediction errors.
[0144] NPUs designed to accelerate inference are generally configured to operate on a complete model. Such NPUs can thus be configured to input a new data slice and quickly process the new data through the already trained model to generate a model output (e.g., an inference).
[0145] In a particular implementation, the NPU 908 is part of one or more of the CPU 902, GPU 904, and / or DSP 906.
[0146] In some examples, the wireless connection component 912 can include sub-components for, for example, third-generation (3G) connections, fourth-generation (4G) connections (e.g., 4G LTE), fifth-generation connections (e.g., 5G or NR), Wi-Fi connections, Bluetooth connections, and other wireless data transmission standards. The wireless connection component 912 is further connected to one or more antennas 914.
[0147] The processing system 900 can also include one or more sensor processing units 916 associated with sensors in any way, one or more image signal processors (ISPs) 918 associated with image sensors in any way, and / or a navigation component 920, which can include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.
[0148] The processing system 900 can also include one or more input and / or output devices 922, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, etc.
[0149] In some examples, one or more of the processors in the processing system 900 can be based on the ARM, RISC-V, MIPS, or X86 instruction sets.
[0150] The processing system 900 also includes a memory 924, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 924 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 900.
[0151] Specifically, in this example, the memory 924 includes a self-attention component 924A, a feed-forward component 924B, and a propagation component 924C. Although the illustrated components (and other components not depicted) are depicted as discrete components for clarity of concept, the illustrated components (and other components not depicted) in various aspects may be implemented jointly or individually. Figure 9 The components illustrated (and other components not depicted) in
[0152] In the illustrated example, the memory 924 also includes model parameters 924D. The model parameters 924D may generally correspond to learnable or trainable parameters of one or more machine learning models, such as those for generating self-attention, controlling attention propagation operations, processing attention features to generate transformer outputs, and so on.
[0153] Although depicted as residing in the memory 924 for clarity of concept, in some aspects, some or all of the model parameters 924D may reside in any other suitable location.
[0154] The processing system 900 also includes a self-attention circuit 926, a feed-forward circuit 927, and a propagation circuit 928. The depicted circuits and other circuits not depicted may be configured to perform various aspects of the techniques described herein.
[0155] In one aspect, the self-attention component 924A and the self-attention circuit 926 may be used to generate or perform self-attention in one or more transformer blocks, as discussed above. For example, the self-attention component 924A and the self-attention circuit 926 may implement Figure 1 , Figure 2A , Figure 2B , Figure 3A and / or Figure 3B one or more self-attention sub-blocks 115 of Figure 4 and / or the operations of the self-attention sub-blocks 415 and / or 455 of
[0156] The feed-forward component 924B and the feed-forward circuit 927 may be used to generate output data based on attention features in one or more transformer blocks, as discussed above. For example, the feed-forward component 924B and the feed-forward circuit 927 may implement Figure 1 , Figure 2A , Figure 2B , Figure 3A and / or Figure 3B one or more feed-forward sub-blocks 120 of Figure 4 and / or the operations of the feed-forward sub-blocks 420 and / or 460 of
[0157] The propagation component 924C and the propagation circuit 928 can be used to propagate attention information (e.g., an attention matrix and / or output attention features) from one or more upstream transformer blocks to one or more downstream transformer blocks, as discussed above. For example, the propagation component 924C and the propagation circuit 928 can implement Figure 1 the propagation operation 130 of Figure 2A and Figure 2B the link 245 of Figure 3A and Figure 3B the link 345 of Figure 4 and / or the operation of the link 425 of
[0158] Although depicted as separate components and circuits in Figure 9 for clarity, the self-attention circuit 926, the feed-forward circuit 927, and the propagation circuit 928 can be implemented jointly or separately in other processing devices of the processing system 900, such as within the CPU 902, GPU 904, DSP 906, NPU 908, etc.
[0159] Generally speaking, the processing system 900 and / or its components can be configured to execute the methods described herein.
[0160] It is worth noting that in other aspects, aspects of the processing system 900 can be omitted, such as in the case where the processing system 900 is a server computer, etc. For example, in other aspects, the multimedia processing unit 910, the wireless connection component 912, the sensor processing unit 916, the ISP 918, and / or the navigation component 920 can be omitted. Additionally, aspects of the processing system 900 can be distributed among multiple devices.
[0161] Example clause
[0162] Specific implementation examples are described in the following numbered clauses:
[0163] Clause 1: Using a first transformer block among a plurality of transformer blocks to generate a first attention propagation output, the generating including using a first self-attention sub-block of the first transformer block to process input data for the first transformer block; propagating the first attention propagation output to a second transformer block among the plurality of transformer blocks; and generating an output for the second transformer block, the generating the output for the second transformer block including generating output features for the second transformer block based on the first attention propagation output. In some aspects, using a second self-attention sub-block of the second transformer block to generate the output features for the second transformer block. One advantage of Clause 1 is that attention information can be reused, thereby reducing computational overhead.
[0164] Clause 2: The method according to Clause 1, wherein: generating the first attention propagation output further includes using the first transformer block to generate an attention matrix; and generating the attention matrix includes using the first self-attention sub-block to process the query representation and key representation of the input data for the first transformer block. An advantage of Clause 2 is that the attention matrix can be generated by the first transformer and reused by downstream transformers.
[0165] Clause 3: The method according to Clause 1 or 2, wherein the first attention propagation output includes the attention matrix. An advantage of Clause 3 is that the second transformer does not need to generate its own attention matrix, thereby reducing overhead.
[0166] Clause 4: The method according to any one of Clauses 1 to 3, wherein generating the output features for the second transformer block further includes: accessing the output of a third transformer block among the plurality of transformer blocks, wherein the third transformer block is immediately before the second transformer block; and using the second self-attention sub-block of the second transformer block to generate the output features for the second transformer block based on the first attention propagation output and the value representation of the output of the third transformer block. An advantage of Clause 4 is that the second transformer can generate its own value representation, thereby improving model accuracy.
[0167] Clause 5: The method according to any one of Clauses 1 to 4, wherein generating the first attention propagation output includes using the first transformer block to generate output features for the first transformer block by using the first self-attention sub-block to process the attention matrix and the value representation of the input data for the first transformer block. An advantage of Clause 5 is that the transformer can compute self-attention to improve model performance.
[0168] Clause 6: The method according to any one of Clauses 1 to 5, wherein the first attention propagation output includes the output features for the first transformer block. An advantage of Clause 6 is that the output features of the first transformer block can be reused.
[0169] Clause 7: The method according to any one of Clauses 1 to 6, the method further includes using the first transformer block to generate an output for the first transformer block, wherein generating the output for the first transformer block includes using the first feed-forward sub-block of the first transformer block to process the output features of the first self-attention sub-block. An advantage of Clause 7 is that the output of the first transformer block can be used to improve or provide one or more predictions.
[0170] Clause 8: The method according to any one of Clauses 1 to 7, wherein generating the output for the second transformer block includes processing the output features of the second self-attention sub-block using a second feed-forward sub-block of the second transformer block. An advantage of Clause 8 is that the output of the second transformer block can be used to improve or provide one or more predictions.
[0171] Clause 9: The method according to any one of Clauses 1 to 8, wherein: the first transformer block includes an encoder block, and the second transformer block includes a decoder block. An advantage of Clause 9 is that the decoder can reuse the attention from the encoder.
[0172] Clause 10: The method according to any one of Clauses 1 to 9, wherein: the plurality of transformer blocks includes a sequence of transformer blocks, the sequence of transformer blocks includes one or more initial blocks, a plurality of intermediate blocks, and one or more final blocks, and the plurality of intermediate blocks includes the first transformer block and the second transformer block. An advantage of Clause 10 is that the attention can be shared among the intermediate transformers in the sequence.
[0173] Clause 11: The method according to any one of Clauses 1 to 10, wherein generating the first attention propagation output includes processing the input data for the first transformer block using a plurality of window self-attention operations to generate the output features for the first transformer block. An advantage of Clause 11 is that multiple self-attentions such as window self-attention can be used in combination with attention propagation.
[0174] Clause 12: The method according to any one of Clauses 1 to 11, wherein the first attention propagation output includes the output features for the first transformer block. An advantage of Clause 12 is that the window self-attention features can be shared.
[0175] Clause 13: The method according to any one of Clauses 1 to 12, wherein: propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output, the propagation operation including transforming the first attention propagation output by concatenating the output features for a third transformer block in the plurality of transformer blocks to the first attention propagation output, and the third transformer block is immediately before the second transformer block. An advantage of Clause 13 is that the propagation operation can include transforming features to improve model performance.
[0176] Clause 14: The method according to any one of Clauses 1 to 13, wherein: Propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output, and the propagation operation includes using an upsampling operation to transform the first attention propagation output. An advantage of Clause 14 is that the attention propagation can be upsampled to improve compatibility and model performance.
[0177] Clause 15: The method according to any one of Clauses 1 to 14, wherein propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output. An advantage of Clause 15 is that the attention can be propagated using various operations.
[0178] Clause 16: The method according to any one of Clauses 1 to 15, wherein the propagation operation includes transforming the first attention propagation output by performing one or more convolution operations on the first attention propagation output. An advantage of Clause 16 is that the convolution operation can improve model accuracy.
[0179] Clause 17: The method according to any one of Clauses 1 to 16, wherein the propagation operation includes an identity operation. An advantage of Clause 17 is that the attention can be accurately propagated and has reduced computational overhead.
[0180] Clause 18: The method according to any one of Clauses 1 to 17, wherein when generating the output features for the second transformer block, the second self-attention sub-block does not calculate the attention matrix. An advantage of Clause 18 is that the computational overhead and latency introduced by the second transformer block are reduced.
[0181] Clause 19: A processing system, the processing system comprising: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform the method according to any one of Clauses 1 to 18.
[0182] Clause 20: A processing system, the processing system comprising: components for performing the method according to any one of Clauses 1 to 18.
[0183] Clause 21: A non-transitory computer-readable medium, the non-transitory computer-readable medium including computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method according to any one of Clauses 1 to 18.
[0184] Clause 22: A computer program product embodied on a computer-readable storage medium, the computer program product comprising: code for performing the method according to any one of Clauses 1 to 18.
[0185] Additional Notes
[0186] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein do not limit the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, the functions and arrangements of the elements discussed may be changed without departing from the scope of the disclosure. Each example may omit, substitute, or add various procedures or components as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described for some examples may be combined in some other examples. For example, any number of the aspects set forth herein may be used to implement a device or practice a method. In addition, the scope of the disclosure is intended to cover such devices or methods practiced using other structures, functionality, or a combination of structure and functionality in addition to or different from the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure herein may be embodied by one or more elements of the claims.
[0187] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or having an advantage over other aspects.
[0188] As used herein, the phrase referring to a list of items "at least one of" refers to any combination of those items (including a single member). For example, "at least one of a, b, or c" is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination having multiple identical elements (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c, or any other ordering of a, b, and c).
[0189] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" may include computing, calculating, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, and the like. Additionally, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determine" may include parsing, selecting, picking, establishing, etc.
[0190] The methods disclosed herein include one or more steps or acts for implementing the methods. The method steps and / or acts may be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the order and / or use of the specific steps and / or acts may be modified without departing from the scope of the claims. In addition, the various operations of the methods described above may be performed by any suitable component capable of performing the corresponding functions. The component may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally speaking, in the presence of the operations illustrated in the drawings, those operations may have corresponding components plus functional components with similar numbers.
[0191] The following claims are not intended to be limited to the aspects shown herein, but should be accorded the full scope consistent with the claim language. In the claims, unless specifically stated otherwise, the use of the singular form of an element is not intended to mean "one and only one" but "one or more". Unless specifically stated otherwise, the term "some" means one or more. No claim element shall be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the phrase "step for". All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later will be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is explicitly recited in the claims.
Claims
1. A computer-implemented method, the computer-implemented method comprising: Generate a first attention propagation output using a first transformer block among a plurality of transformer blocks, the generating including processing input data for the first transformer block using a first self-attention sub-block of the first transformer block; Propagate the first attention propagation output to a second transformer block among the plurality of transformer blocks; And Generate an output for the second transformer block, the generating of the output for the second transformer block including generating output features for the second transformer block based on the first attention propagation output.
2. The computer-implemented method according to claim 1, wherein: Generating the first attention propagation output further includes generating an attention matrix using the first transformer block; and Generating the attention matrix includes processing query representations and key representations of the input data for the first transformer block using the first self-attention sub-block.
3. The computer-implemented method according to claim 2, wherein the first attention propagation output comprises the attention matrix.
4. The computer-implemented method according to claim 3, wherein generating the output features for the second transformer block further comprises: Access an output for a third transformer block among the plurality of transformer blocks, where the third transformer block is immediately before the second transformer block; And Generate the output features for the second transformer block using a second self-attention sub-block of the second transformer block, based on the first attention propagation output and value representations of the output for the third transformer block.
5. The computer-implemented method according to claim 2, wherein generating the first attention propagation output comprises using the first transformer block to generate output features for the first transformer block by processing the attention matrix and the value representation of the input data for the first transformer block using the first self-attention sub-block.
6. The computer-implemented method according to claim 5, wherein the first attention propagation output comprises the output features for the first transformer block.
7. The computer-implemented method according to claim 1, the computer-implemented method further comprising using the first transformer block to generate an output for the first transformer block, wherein generating the output for the first transformer block comprises using a first feed-forward sub-block of the first transformer block to process the output features of the first self-attention sub-block.
8. The computer-implemented method according to claim 1, wherein generating the output for the second transformer block comprises using a second feed-forward sub-block of the second transformer block to process the output features of the second self-attention sub-block.
9. The computer-implemented method according to claim 1, wherein: The first transformer block includes an encoder block, and The second transformer block includes a decoder block.
10. The computer-implemented method according to claim 1, wherein: The plurality of transformer blocks includes a sequence of transformer blocks, The sequence of transformer blocks includes one or more initial blocks, a plurality of intermediate blocks, and one or more final blocks, and The plurality of intermediate blocks includes the first transformer block and the second transformer block.
11. The computer-implemented method according to claim 1, wherein generating the first attention propagation output includes using a plurality of window self-attention operations to process the input data for the first transformer block to generate the output features for the first transformer block.
12. The computer-implemented method according to claim 11, wherein the first attention propagation output includes the output features for the first transformer block.
13. The computer-implemented method according to claim 12, wherein: Propagating the first attention propagation output to the second transformer block includes propagating the first attention propagation output using a propagation operation, The propagation operation includes transforming the first attention propagation output by concatenating output features of a third transformer block among the plurality of transformer blocks to the first attention propagation output, and The third transformer block is immediately before the second transformer block.
14. The computer-implemented method according to claim 12, wherein: Propagating the first attention propagation output to the second transformer block includes propagating the first attention propagation output using a propagation operation, and The propagation operation includes transforming the first attention propagation output using an upsampling operation.
15. The computer-implemented method according to claim 1, wherein propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output.
16. The computer-implemented method according to claim 15, wherein the propagation operation includes transforming the first attention propagation output by performing one or more convolution operations on the first attention propagation output.
17. The computer-implemented method according to claim 1, wherein when generating the output features for the second transformer block, the second self-attention sub-block does not calculate the attention matrix.
18. A processing system, the processing system comprising: A memory, the memory including computer-executable instructions; And One or more processors configured to execute the computer-executable instructions and cause the processing system to perform operations including the following: Generate a first attention propagation output using a first transformer block among a plurality of transformer blocks, the generating including processing input data for the first transformer block using a first self-attention sub-block of the first transformer block; Propagate the first attention propagation output to a second transformer block among the plurality of transformer blocks; And Generate an output for the second transformer block, the generating of the output for the second transformer block including generating output features for the second transformer block based on the first attention propagation output.
19. The processing system according to claim 18, wherein: Generating the first attention propagation output further includes using the first transformer block to generate an attention matrix; and Generating the attention matrix includes using the first self-attention sub-block to process query representations and key representations of the input data for the first transformer block.
20. The processing system according to claim 19, wherein the first attention propagation output includes the attention matrix.
21. The processing system according to claim 20, wherein generating the output features for the second transformer block further comprises: Accessing the output of a third transformer block among the plurality of transformer blocks, where the third transformer block is immediately before the second transformer block; and Using the second self-attention sub-block of the second transformer block to generate the output features for the second transformer block based on the first attention propagation output and the value representation of the output for the third transformer block.
22. The processing system according to claim 19, wherein generating the first attention propagation output comprises using the first transformer block to generate output features for the first transformer block by processing the attention matrix and the value representation of the input data for the first transformer block using the first self-attention sub-block.
23. The processing system according to claim 22, wherein the first attention propagation output comprises the output features for the first transformer block.
24. The processing system according to claim 18, the operation further comprising using the first transformer block to generate an output for the first transformer block, wherein generating the output for the first transformer block comprises using a first feed-forward sub-block of the first transformer block to process the output features of the first self-attention sub-block.
25. The processing system according to claim 18, wherein generating the output for the second transformer block comprises using a second feed-forward sub-block of the second transformer block to process the output features of the second self-attention sub-block.
26. The processing system according to claim 18, wherein: The first transformer block includes an encoder block, and The second transformer block includes a decoder block.
27. The processing system according to claim 18, wherein: The plurality of transformer blocks includes a sequence of transformer blocks, The sequence of transformer blocks includes one or more initial blocks, a plurality of intermediate blocks, and one or more final blocks, and The plurality of intermediate blocks includes the first transformer block and the second transformer block.
28. The processing system according to claim 18, wherein generating the first attention propagation output comprises using a plurality of window self-attention operations to process the input data for the first transformer block to generate the output features for the first transformer block.
29. The processing system according to claim 28, wherein the first attention propagation output comprises the output features for the first transformer block.
30. The processing system according to claim 29, wherein: Propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output, The propagation operation includes transforming the first attention propagation output by concatenating the output features of a third transformer block among the plurality of transformer blocks to the first attention propagation output, and The third transformer block is immediately before the second transformer block.
31. The processing system according to claim 29, wherein: Propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output, and The propagation operation includes using an upsampling operation to transform the first attention propagation output.
32. The processing system according to claim 18, wherein propagating the first attention propagation output to the second transformer block includes using a propagation operation to propagate the first attention propagation output.
33. The processing system according to claim 32, wherein the propagation operation includes transforming the first attention propagation output by performing one or more convolution operations on the first attention propagation output.
34. The processing system according to claim 18, wherein when generating the output features for the second transformer block, the second self-attention sub-block does not calculate an attention matrix.
35. A processing system, the processing system comprising: Components for using a first transformer block among a plurality of transformer blocks to generate a first attention propagation output, the components for generating are configured to use the first self-attention sub-block of the first transformer block to process the input data for the first transformer block; Components for propagating the first attention propagation output to a second transformer block among the plurality of transformer blocks; and Components for generating an output for the second transformer block, the components for generating the output for the second transformer block are configured to output features for the second transformer block based on the first attention propagation output.