Efficient inference using adapted transformer with token abstractor
By integrating learnable token abstractors into transformer layers, the adapted transformer model reduces computational complexity and enhances efficiency, addressing the scalability issues of traditional transformer models while maintaining accuracy.
Patent Information
- Application Number
- PCT/CN2023/136760
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-12
AI Technical Summary
Transformer models face significant computational complexity due to global attention, which scales exponentially with input token length, leading to high processing overhead and limited deployment efficiency, especially on resource-constrained devices.
The implementation of an adapted transformer model with learnable token abstractors inserted in certain transformer layers to reduce the number of tokens, preserve informative tokens, and create contextual summarization, thereby reducing computational complexity while maintaining accuracy.
This approach significantly reduces latency and increases throughput of the transformer network with minimal loss in accuracy, and requires only lightweight adaptation by training the added token abstractors without modifying the original pre-trained model.
Smart Images

Figure CN2023136760_12062025_PF_FP_ABST
Abstract
Description
EFFICIENT INFERENCE USING ADAPTED TRANSFORMER WITH TOKEN ABSTRACTORBACKGROUND
[0001] While transformer models have achieved remarkable performance in deep-learning tasks across many data modalities, the global attention computational complexity of these models scales exponentially with the input token length. For example, natural language processing (NLP) transformers –e.g. BERT or generative pre-trained transformers (GPT) –scale quadratically with the number of tokens, and vision transformers (ViT) scale in quadruple due to the quadratic token length growth in each spatial dimension (height and width axis) . Various techniques have been used with trained transformers to adaptively remove tokens during inference to reduce computational footprint for deployment efficiency and resource-constrained devices. However, existing methods incur high processing overhead –thus limiting adoption and the ease of deployment, and / or are tailored to a given task or domain –thus lacking breadth in applicability. As one example, the Transkimmer injects token classifiers at the beginning of each transformer layer to learn what tokens to be pruned during inference, but this approach is expensive in optimization, as it requires re-training both the classifiers and the original transformer model to retain accuracy. As another example, Token Merging gradually merges similar image tokens as one (instead of dropping them) within a transformer layer; while this approach works well with some visual recognition tasks (e.g., using ViT) , it generalizes poorly to language tasks.BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The various advantages of the embodiments will become apparent to one skilled in the art by reading the following specification and appended claims, and by referencing the following drawings, in which:
[0003] FIG. 1 provides a block diagram illustrating a conventional transformer network;
[0004] FIG. 2 provides a block diagram illustrating an example of an adapted transformer network according to one or more embodiments;
[0005] FIG. 3 provides a block diagram illustrating an example of an adapted transformer layer for a performance-enhanced transformer network according to one or more embodiments;
[0006] FIG. 4 provides a flow diagram illustrating an example process of optimizing a transformer network with adapted transformer layers according to one or more embodiments;
[0007] FIG. 5A provides a flow diagram illustrating an example method of producing a performance-enhanced transformer network according to one or more embodiments;
[0008] FIG. 5B provides a flow diagram illustrating an example method of filtering tokens according to one or more embodiments;
[0009] FIG. 5C provides a flow diagram illustrating an example method of generating a contextual token according to one or more embodiments;
[0010] FIG. 6 provides a block diagram illustrating an example performance-enhanced computing system according to one or more embodiments;
[0011] FIG. 7 provides a block diagram illustrating an example semiconductor apparatus for accelerating compute tasks according to one or more embodiments;
[0012] FIG. 8 is a block diagram illustrating an example processor core according to one or more embodiments; and
[0013] FIG. 9 is a block diagram illustrating an example of a multi-processor based computing system according to one or more embodiments.DESCRIPTION OF EMBODIMENTS
[0014] Embodiments relate generally to artificial intelligence (AI) systems and deep learning (DL) technology. More particularly, embodiments relate to a performance-enhanced transformer model in which certain transformer layers are adapted to reduce the number of tokens and, hence, the computational complexity of the model. The technology provided herein recognizes that not all tokens are informative, and less important tokens do carry residual information. As described herein, the improved transformer model uses a learnable token abstractor inserted in certain transformer layers of a pre-trained transformer network to preserve the most informative tokens, remove the remaining (e.g., less-informative) tokens, and create new auxiliary tokens as contextual summarization –thereby reducing the total number of tokens required for efficient inference. As a result, latency and throughput of the transformer network is increased significantly with minimal loss in accuracy. Further, by utilizing the original pre-trained transformer model (i.e., the original pre-trained model is “frozen” such that the weights and parameters of the original model are unchanged and not modified) , only the added token abstractors need to be trained, which presents a lightweight adaptation.
[0015] FIG. 1 provides a block diagram illustrating a conventional transformer network 100. The conventional transformer network 100 can be of a variety of transformer types, such as, e.g., BERT, ViT, Wav2Vec, etc. The conventional transformer network 100 receives a set of input tokens 115 representative of an input 110 (e.g., an image or text) . The input tokens 115 can be provided, e.g., via a text or image encoder, and typically are in the form of fixed-length vectors. The conventional transformer network 100 includes a base transformer 120 having a series of transformer layers. For purposes of illustration, 5 layers are shown --including a transformer layer 1 (label 121) , a transformer layer 2, and other transformer layers 3-5. The transformer layers are trained to transform input tokens to intermediate token representations, such as tokens 122 produced by the transformer layer 1 (label 121) . The final layer (transformer layer 5 in the illustrated example) produces output tokens 125 of length L (L=10 in the illustrated example) , and a downstream classifier 130 processes the output tokens 125 to produce a classification result. The downstream classifier 130 can include, e.g., one or more neural network layers, or a separate neural network, etc.
[0016] For the conventional transformer network 100, the token length L (i.e., the number of tokens) of intermediate representation remains the same at all stages. That is, at each layer the number of tokens produced is the same as the number of tokens presented to that layer. In the example as illustrated in FIG. 1, each layer produces 10 tokens, the same number as in the input tokens 115. Given there are 5 layers in the illustrated example, this means that the conventional transformer network 100 processes a total of 50 tokens for a forward pass.
[0017] FIG. 2 provides a block diagram illustrating an example of an adapted transformer network 200 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The adapted transformer network 200 can generally be implemented in a computing system, such as the performance-enhanced computing system 10 (FIG. 6, discussed herein) , and / or in an AI / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc.
[0018] The adapted transformer network 200 receives a set of input tokens 215 representative of an input 210 (e.g., an image or text) . The input tokens 215 can be provided, e.g., via a text or image encoder. As shown in FIG. 2, the adapted transformer network 200 includes a base transformer 220 having a series of transformer layers including a transformer layer 1 (label 221) , a transformer layer 2 (label 223) , and other transformer layers 3-5. The base transformer 220 includes layers derived from a pre-trained transformer (such as, e.g., the conventional transformer layers of the conventional transformer network 100) , where the transformer layers are trained to transform input tokens to intermediate token representations –such as tokens 222 produced by the transformer layer 1 (label 221) .
[0019] According to embodiments, one or more of the transformer layers (e.g., conventional transformer layers) in the base transformer 220 are adapted, as described further herein with reference to FIG. 3, to insert a token abstractor into the transformer layer to form an adapted transformer layer. The token abstractor is a small neural network structure trained to preserve the most informative tokens, remove the remaining (e.g., less-informative) tokens, and create one or more new auxiliary tokens as contextual summarization. For example, if the input is an image of a cat, the background pixels for classifying the cat image carry less feature about the cat, and thus some tokens relating to the background are less informative and can be removed. As another example, if the input is a text phrase “I really like this film, ” for sentiment analysis the word “this” has less to do with feeling other than correspondence and, thus, some tokens relating to “this” can be removed. While these tokens can be removed, they do carry residual information which impacts transformer prediction quality. Therefore, retaining residual information is important. Overall, each token abstractor reduces the number of tokens processed by succeeding transformer layers in the base transformer 220 and abstracts the residual information by creating one or a small number of auxiliary contextual tokens.
[0020] In the example of FIG. 2, two transformer layers, transformer layer 2 (label 223) and transformer layer 4, are each adapted to include a token abstractor inserted into the layer. As shown in the figure, transformer layer 2 operates to reduce the number of tokens from 10 tokens to 5 tokens. That is, the token abstractor has a reduction factor r of 0.5. The 5 tokens output by the adapted transformer layer 2 include 4 tokens 224 along with an added contextual token 224a. The transformer layer 3 then operates on the 5 tokens (output by the adapted transformer layer 2) to produce 5 tokens. The adapted transformer layer 4 operates to reduce the number of tokens from 5 to 2, where the 2 tokens include a contextual token. The transformer layer 5 then operates on the 2 tokens (output by the adapted transformer layer 4) to produce 2 tokens –a token 225 and a contextual token 225a. The two output tokens (225 and 225a) are then presented to a downstream classifier 230, which processes the output tokens to produce a classification result. The downstream classifier 230 can include, e.g., one or more neural network layers, or a separate neural network, etc.
[0021] In the example as illustrated in FIG. 2, the transformer layer 1 processes 10 tokens (producing 10 tokens) ; the transformer layer 2 processes 10 tokens (producing 5 tokens) ; the transformer layer 3 processes 5 tokens (producing 5 tokens) ; the transformer layer 4 processes 5 tokens (producing 2 tokens) ; and the transformer layer 5 processes 2 tokens (producing 2 tokens) . This means that the adapted transformer network 200 processes a total of 32 tokens for a forward pass, representing a reduction of 36%in tokens processed in comparison to the conventional transformer network 100 (which processes 50 tokens) discussed above with reference to FIG. 1. The reduction percentage will, in embodiments, be based on the number of transformer layers overall, the number of adapted transformer layers and / or the reduction factor r.
[0022] While FIG. 2 shows, for illustrative purposes, the adapted transformer network 200 with five transformer layers, two of which are adapted transformer layers, and reductions from 10 to 5 tokens and 5 to 2 tokens, it will be understood that the adapted transformer network 200 can include a different number of transformer layers, a different number of adapted transformer layers, and / or different token reductions. Additionally, while the token abstractor in the example has a reduction factor r of 0.5, other reduction values can be employed. Also, while the adapted transformer network 200 is illustrated for use in conjunction with a downstream classifier 230 for handling classification tasks, in some embodiments the adapted transformer network 200 can be used in neural network applications that employ a transformer model for tasks other than classification.
[0023] Some or all components and / or features in the adapted transformer network 200 can be implemented using one or more of a central processing unit (CPU) , a graphics processing unit (GPU) , an artificial intelligence (AI) accelerator, a field programmable gate array (FPGA) accelerator, an application specific integrated circuit (ASIC) , and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC. More particularly, components of the adapted transformer network 200 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as random access memory (RAM) , read only memory (ROM) , programmable ROM (PROM) , firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured programmable logic arrays (PLAs) , FPGAs, complex programmable logic devices (CPLDs) , and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with complementary metal oxide semiconductor (CMOS) logic circuits, transistor-transistor logic (TTL) logic circuits, or other circuits.
[0024] For example, computer program code to carry out operations by the adapted transformer network 200 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0025] FIG. 3 provides a block diagram illustrating an example of an adapted transformer layer 300 for a performance-enhanced transformer network according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. As shown in FIG. 3, the adapted transformer layer 300 includes a multi-head self attention network (MHSA) 310, a feed-forward network (FFN) 320, and a token abstractor 330 arranged between the MHSA 310 and the FFN 320. The MHSA 310 and the FFN 320 are components of a trained transformer layer (e.g., part of a pre-trained conventional transformer network) and, thus, the adapted transformer layer 300 is essentially a trained transformer layer with the token abstractor 330 inserted (i.e., arranged) between the MHSA 310 and the FFN 320. It will be understood that a multi-head attention network (MHA) and MHSA are very similar (MHSA is actually a special case of MHA) , and in some embodiments the adapted transformer layer 300 uses a MHA instead of MHSA for element 310.
[0026] As shown in FIG. 3, the adapted transformer layer 300 receives as input a set of tokens xIN and produces as output a set of tokens xOUT. The MHSA 310 processes the input tokens xIN (length L tokens) and produces tokens (length L) . The tokens produced by the MHSA 310 are summed with the input tokens xIN to produce tokens xMID (length L) , which are input to the token abstractor 330. The token abstractor 330 produces a set of tokens x’MID (length L’, which is less than L) which are fed to the FFN 320 and also summed with the output of the FFN 320 to produce the output tokens xOUT (length L’) . The abstractor tokens x’MID thus include a reduced set of tokens (based on removing tokens) and one or more contextual token (s) . In equation form, the various token sets are represented as follows: xMID=xIN +MHSA (xIN) (1) x′MID=Abstractor (xMID) (2) xIUT=x′MID+FFN (x′MID) . (3)
[0027] As shown in FIG. 3, the token abstractor 330 includes two components, a token filter 332 and a contextual token generator 333. The token filter 332 prunes (i.e., removes) uninformative tokens and retains the top-T tokens. The contextual token generator 333 summarizes the entire sequence representation with newly created token (s) , which can be considered as an approximation of information loss due to the removed tokens.
[0028] Token Filter:
[0029] The core of conventional transformer models lies in the attention mechanism of the MHSA, which provides for learning the dependencies among tokens, i.e., multi-head attention scores, to transform their representations. The multi-head attention scores can be considered as providing information regarding the importance of each token, and a trained transformer can thus be considered as providing guidance, via the multi-head attention scores, as to which tokens are more important. While a single token has many scores to other tokens, compounded by the number of attention heads in the MHSA, the token abstractor 330 uses a reduction summation to arrive at global importance of a given token.
[0030] A typical MHSA is composed of a total of H heads, where each head uses a query Q (h) , a key K (h) , and a value V (h) in shape to calculate the head attention score A (h) :
[0031] where D is the dimension of a token vector.
[0032] Each element in A (h) is the attention magnitude of token j to token i, which is used by the token filter 332 as a metric for importance. The token filter 332 computes a column sum of attention scores across all heads of the MHSA 310 to yield the global importance, IGLIBAL, for all L tokens:
[0033] The token filter 332 determines a sorted (e.g., ranked) set of indices based on the global importance (IGLIBAL) values: indices = argsort_descend (IGLIBAL) [: T] (6)
[0034] The token filter 332 then selects the top T attentive tokens, xTOP, from the tokens 331 in xMID based on the sorted indices: xTIP=index_select (xMID, indices) ∈RT×D (7)
[0035] where the value of T depends on the hyperparameter r (reduction factor) , described further herein with reference to the optimization process of FIG. 4. While the hyperparameter r is set as part of the optimization process, the token filter 332 is otherwise untrained.
[0036] Contextual Token Generator:
[0037] While the token filter 332 retains the top T attentive tokens, the discarding of the remaining tokens (L-T tokens) does have an impact, both in terms of loss of information (although the discarded tokens are less important, they carry some residual information) and effect on training the token abstractor 330 (due to the loss in information) . For example, to maximize the potential increase in performance / acceleration, the number of tokens retained, T, should be as small as possible relative to L (the number of tokens input to the token abstractor 330) . Yet making T as small as possible risks being overly-aggressive in token reduction, which can result in excessive loss of information or increased training / tuning for the token abstractor 330.
[0038] By capturing residual information of the discarded tokens via the contextual token, these potentially negative impacts can be reduced. The contextual token generator 333 includes a multi-head attention (MHA) 334 (e.g., a MHA sub-layer) to construct a new contextual token, xCONTEXT, that abstracts the information of the entire sequence representation –including the residual information in the discarded tokens. The resulting new contextual token is produced as shown in Equation (8a) : xCONTEXT= MHA (query=qSEQ, key=xMID, value=xMID) ∈R1×D (8a)
[0039] where
[0040] The MHA 334 is trained, as described herein with reference to the optimization process of FIG. 4. In embodiments, the top T attentive tokens are not masked out, to allow the module flexibility to decide the composition. In such embodiments, the query, denoted by qSEQ, is the average of all L tokens in xMID, serving as the D-dimension representation for the sequence. The key and value come from xMID, and are part of the standard output (per set of tokens) of the MHSA 310 of the original trained transformer. It will be understood that a MHA and MHSA are very similar (MHSA is actually a special case of MHA) , and each provide the same input query / key / value tensor. In some embodiments, an MHSA is used (instead of MHA) in the contextual token generator 333. In some embodiments, the token generator 333 is implemented using other structures, such as, e.g., average pooling with a multi-layer perceptron (MLP) , or a one-dimensional convolution layer.
[0041] In some embodiments, the top T attentive tokens are masked out in the contextual token generator 333, such that the resulting contextual token is focused on the residual information in the discarded tokens. As one example, the masking is accomplished by using the top T token result as a mask –e.g., by assigning a mask value of “0” corresponding to the top T tokens that are retained by the token filter 332, and assigning a mask value of “1” corresponding to the tokens that are to be discarded. By applying a mask value of “0” to the top tokens, these values do not contribute to the computations used in generating the contextual token.
[0042] In embodiments, the contextual token generator 333 produces a single contextual token. In some embodiments, the contextual token generator 333 produces a plurality of contextual tokens (which comes at higher computational cost) . As one example, a plurality of contextual token generators 333 are used, each individually trained and each providing a contextual token. As another example, the column mean process in the contextual token generator 333 is split into parts (e.g., two parts) , each part generating its own qSEQ (effectively formulating qSEQ differently for each part by splitting the column mean) . Each qSEQ is then fed separately into the MHA 334 to produce separate contextual tokens.
[0043] Finally, the token abstractor 330 produces its output, x’MID, by concatenating the output of the token filter 332 (T top tokens, xTOP) and the output of the contextual token generator 333 (xCONTEXT) : x′MID=concat (xTIP, xCINTEXT) (9)
[0044] where x’MID is a sequence of tokens of length L’ = T+1 (for a single contextual token) . The token reduction factor r is related to (e.g., can be determined from) the input token length, L, and the output token length, L’: r = (L –L’) / L (10)
[0045] The reduction factor r can be set as part of the optimization process for the token abstractor, as described herein with reference to FIG. 4. In the example illustrated in FIG. 3, the MHSA 310 is shown as producing 5 tokens (thus xMID is comprised of 5 tokens 331) , while token abstractor produces 3 tokens 335. Of the 3 tokens 335 produced by the token abstractor, 2 tokens are provided by the token filter 332, and 1 added token is provided by the contextual token generator 333 and concatenated with the 2 tokens provided by the token filter 332 to provide the 3 resulting tokens 335.
[0046] The adapted transformer layer 300 can generally be implemented in a layer of the adapted transformer network 200 (e.g., by inserting a token abstractor 330 into the transformer layer 1, label 221 in FIG. 2, already discussed) . Some or all components and / or features in the adapted transformer layer 300 can be implemented using one or more of a CPU, a GPU, an AI accelerator, a FPGA accelerator, an ASIC, and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC, and / or in an AI / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc. More particularly, components of the adapted transformer layer 300 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0047] For example, computer program code to carry out operations by the adapted transformer layer 300 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, ISA instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0048] Transformer Network Optimization:
[0049] FIG. 4 provides a flow diagram illustrating an example process (e.g., workflow) for optimizing a transformer network with adapted transformer layers according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The process 400 can generally be implemented in a computing system, such as the performance-enhanced computing system 10 (FIG. 6, discussed herein) , and / or in an AI / deep learning framework, such as, e.g., PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc. More particularly, the process 400 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0050] For example, computer program code to carry out operations shown in the process 400 and / or functions associated therewith can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, ISA instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0051] The process 400 begins at block 402 with a pre-trained transformer network, either by taking a standard transformer network and pre-training the network for a specific type of task, or by obtaining a pre-trained transformer network that has been trained for a specific type of task. Importantly, once the transformer network is trained, the trained parameters of the network are generally frozen (i.e., will remain unchanged) ; only the parameters of the token abstractors that are inserted into transformer layers of the transformer network are trained in addition.
[0052] At block 404, the number N of adapted transformer layers is initialized or adjusted. The number N represents the number of token abstractors 330. For a first pass through, the value of N is initialized to an initial value. In some embodiments, the initial value of N is set to approximately one-fourth of the total number of transformer layers (e.g., 1 adapted transformer layer for every 4 transformer layers) , and other values of N can be used as an initial value. For subsequent passes, the value of N is adjusted. At block 406, the N token abstractors 330 are inserted into N adapted transformer layers (i.e., one token abstractor 330 is to be inserted in each adapted transformer layer) , at essentially uniform intervals (e.g., uniform intervals or within one or two layers of uniform intervals) across the transformer layers. As one example, for a transformer network having 12 transformer layers, if N is initially set to 2, then there are to be 2 adapted transformer layers, and token abstractors 330 are inserted into 2 layers, such as, e.g., layer 4 and layer 8. As another example, for a transformer network having 12 transformer layers, if N is initially set to 4, then there are to be 4 adapted transformer layers, and token abstractors 330 are inserted into 4 layers, such as, e.g., layer 3, layer 5, layer 8 and layer 10.
[0053] At block 408, the reduction factor r is initialized or adjusted to a value within the range of (0, 1) , where 0 < r < 1. For a first pass through, the value of r is initialized to an initial value. In some embodiments, the initial value of r is set to be within the range of 0.05-0.40. For subsequent passes with the same value of N, the value of r is adjusted. In some embodiments, after N is adjusted, the first pass through r is re-initialized; in other embodiments after N is adjusted, the first pass through r remains the same as the previous pass, or is adjusted.
[0054] The reduction factor r determines the relative length of input and output token sequences at each token abstractor 330. The output sequence length L' at each adapted transformer layer is: L' = floor (L –r*L) = floor (L* (1 -r) ) (11)
[0055] A higher value of r means more aggressive token reduction for acceleration. Assuming the contextual token generator 333 is to produce a single contextual token, the number of top attentive tokens produced by the token filter 332 is T = L'-1. In some embodiments, the number of new contextual tokens to be produced by the contextual token generator is parameterized as part of the optimization workflow.
[0056] At block 410, the token abstractors 330 in the adapted transformer layers are trained. This involves training the parameters of the MHA of each contextual token generator 333. In embodiments the training process is the same process, and uses the same training data, as was used in training the original transformer network (i.e., as used to produce the pre-trained transformer network) . In some embodiments any conventional neural network training process can be used for training the token abstractors 330. Only the parameters of the MHA of the contextual token generator 333 are trained as part of the optimization process 400; as previously stated, the parameters of the original transformer network are generally frozen / unchanged to keep training cost low. In some embodiments, however, this condition is relaxed such that the parameters of the MHA are trained together with training (or re-training) the original transformer model weights --either in the same manner as pre-training, or integrating with other efficient training methods, e.g., Low-Rank Adaptation (LoRA) , etc.
[0057] At block 412, the transformer network as presently configured with the adapted transformer layers is operated on test / sample data to determine results. The performance is monitored to determine processing time required for operation of the transformer network as presently configured.
[0058] At block 414, the results are evaluated to determine if the adapted transformer network meets optimization criteria, such as, e.g., accuracy and / or performance goals. For example, the results are are evaluated to determine if the accuracy meets desired goals, and to determine if the performance meets processing time (e.g., speed) goals. The respective accuracy and speed goals involve a tradeoff between higher accuracy (typically requiring longer processing time) and higher speed (typically resulting in lower accuracy) . For example, where the transformer is used for a classification task, the results are evaluated to determine if the prediction results (e.g., percentage of correct classification) is within a predetermined threshold (e.g., a predetermined percentage) . As another example, the performance, e.g., the number of operations required, is evaluated to determine if the performance is within a predetermined threshold (e.g., predetermined number of operations) . Alternatively, the accuracy and performance are evaluated against baseline accuracy and performance values (e.g., accuracy and performance of the pre-trained transformer network before adapted layers are added) to determine if within acceptable ranges. For example, in some embodiments an accuracy goal is to achieve a 1%or less relative drop in accuracy compared to baseline (e.g., accuracy of the pre-trained transformer) . As one example, if the baseline accuracy is 95%, the accuracy goal can be to achieve 95%*0.99=94.05%after adaptation.
[0059] If YES at block 414 (e.g., the transformer network as presently configured meets both accuracy and performance goals) , the process continues to block 416, where the transformer network configuration is stored. In embodiments, this includes the number of adapted transformer layers (N) , the location of abstractors (e.g., which layers of the transformer are adapted) , the value of r for each abstractor, and the configuration / structure and weights of the MHA for each token abstractor. In some embodiments, if original weights of the transformer model are trainer (or re-trained) , the new weights are also stored. At block 418, the transformer network according to the present configuration is deployed (e.g., via an AI accelerator) . The process then ends at block 430.
[0060] If NO at block 414 (e.g., the transformer network as presently configured does not meet both accuracy and performance goals) , the process goes back to block 408 or block 404, depending on whether one of the goals is met, or none of the goals are met. If the transformer network as presently configured meets one but not both of accuracy and performance goals, the process reverts back to block 408, where the reduction factor r is adjusted, and the process then continues to block 410. For example, the reduction factor r is increased to improve performance (e.g., increased throughput / speed) . As another example, the reduction factor r is decreased to improve accuracy. In embodiments, when the process reverts back to block 408, the weights of the abstractors are re-initialized before restarting training. In some embodiments, the abstractor weights are not re-initialized but are continued to be updated during ongoing training.
[0061] If the transformer network as presently configured does not meet either accuracy or performance goals, the process reverts back to block 404, where the number of adapted transformer layers, N, is adjusted, and the process then continues to block 406. For example, if the number of adapted transformer layers, N, is decreased, typically accuracy will be improved for the same value of r. If the number of adapted transformer layers, N, is increased, the impact on performance or accuracy can be different based on the current value of r and the base transformer model. Generally, increasing N to add abstractors can increase inference cost (e.g., negatively impact performance) , depending on the value of r. The net acceleration effect (e.g., performance improvement) comes from the cost savings (due to r) being higher than the overhead cost of N abstractors.
[0062] The optimization workflow 400 as described herein provides for a heuristic search over N and r. In some embodiments, a single reduction factor r is used for all token abstractors 330 to limit the search space. In some embodiments, the reduction factor r is set separately for each adapted transformer layer / token abstractor 330. In some embodiments, advanced meta-algorithms (such as, e.g., binary-search, gradient-based, evolutionary search, Bayesian optimization, reinforcement learning, etc. ) can be used to arrive at optimal hyperparameters (including N and r) .
[0063] FIG. 5A provides a flow diagram illustrating an example method 500 of producing a performance-enhanced transformer network according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The method 500 can generally be implemented in a computing system, such as the performance-enhanced computing system 10 (FIG. 6, discussed herein) , and / or in an AI / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc. More particularly, the method 500 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0064] For example, computer program code to carry out operations shown in the method 500 and / or functions associated therewith can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, ISA instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0065] Illustrated processing block 510 provides for generating a neural network (e.g., a transformer network) comprising a plurality of transformer layers. Illustrated processing block 520 provides for adapting one or more of the plurality of transformer layers to form one or more adapted transformer layers, where at block 520a each adapted transformer layer is to include a token abstractor coupled between a first attention network and a feed forward network, the first attention network comprising one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network, and at block 520b the token abstractor includes a token filter to filter out one or more of a plurality of tokens and a contextual token generator to generate a contextual token representing residual information from the plurality of tokens.
[0066] FIG. 5B provides a flow diagram illustrating an example method 540 of filtering tokens according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The method 540 can generally be implemented in a computing system, such as the performance-enhanced computing system 10 (FIG. 6, discussed herein) , and / or in an AI / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc. For example, all or portions of the method 540 can be implemented in the token filter 332 (FIG. 3, already discussed) . More particularly, the method 540 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0067] For example, computer program code to carry out operations shown in the method 540 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, ISA instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0068] Illustrated processing block 542 provides for generating a column sum based on a set of attention scores from the first attention network. Illustrated processing block 544 provides for sorting, based on the column sum, indices of the plurality of tokens to form a set of sorted indices. Illustrated processing block 546 provides for selecting, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.
[0069] FIG. 5C provides a flow diagram illustrating an example method 560 of generating a contextual token according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The method 560 can generally be implemented in a computing system, such as the performance-enhanced computing system 10 (FIG. 6, discussed herein) , and / or in an AI / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow or any similar framework or derivative framework such as HuggingFace’s Transformer, OpenVINO’s training extension, etc. For example, all or portions of the method 560 can be implemented in the contextual token generator 333 (FIG. 3, already discussed) . More particularly, the method 560 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0070] For example, computer program code to carry out operations shown in the method 560 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, ISA instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .
[0071] Illustrated processing block 562 provides for generating a column mean based on the set of attention scores from the first attention network (i.e., the first attention network of the respective one or more adapted transformer layers) . Illustrated processing block 564 provides for generating, via a second attention network of the contextual token generator, a contextual token based on the column mean and a key and a value from the first attention network, where the second attention network comprises one of a MHSA network or a MHA network. In some embodiments, the contextual token generator uses an output of the token filter to mask out the subset of tokens from an input to the second attention network.
[0072] In some embodiments, the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers. In some embodiments, each of the plurality of adapted transformer layers uses the same reduction factor. In some embodiments, at least two of the plurality of adapted transformer layers use different reduction factors. In embodiments, the plurality of transformer layers is pre-trained.
[0073] Embodiments of each of the above systems, devices, components and / or methods, including : the adapted transformer network 200, the adapted transformer layer 300, the token abstractor 330, the token filter 332, the contextual token generator 333, the process 400, the method 500, the method 540, and / or the method 560, and / or any other system components, can be implemented in hardware, software, or any suitable combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits. For example, embodiments of each of the above systems, devices, components and / or methods can be implemented via the system 10 (FIG. 6, discussed further below) , the semiconductor apparatus 30 (FIG. 7, discussed further below) , the processor 40 (FIG. 8, discussed further below) , and / or the computing system 60 (FIG. 9, discussed further below) .
[0074] Alternatively, or additionally, all or portions of the foregoing systems and / or devices and / or components and / or methods can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., to be executed by a processor or computing device. For example, computer program code to carry out the operations of the components can be written in any combination of one or more operating system (OS) applicable / appropriate programming languages, including an object-oriented programming language such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages.
[0075] FIG. 6 shows a block diagram illustrating an example performance-enhanced computing system 10 for image and / or text analysis according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The system 10 can generally be part of an electronic device / platform having computing and / or communications functionality (e.g., a server, cloud infrastructure controller, database controller, notebook computer, desktop computer, personal digital assistant / PDA, tablet computer, convertible tablet, smart phone, etc. ) , imaging functionality (e.g., camera, camcorder) , media playing functionality (e.g., smart television / TV) , wearable functionality (e.g., watch, eyewear, headwear, footwear, jewelry, or other wearable devices) , vehicular functionality (e.g., car, truck, motorcycle) , robotic functionality (e.g., robot or autonomous robot) , Internet of Things (IoT) functionality, etc., or any combination thereof. In the illustrated example, the system 10 can include a host processor 12 (e.g., central processing unit / CPU) having an integrated memory controller (IMC) 14 that can be coupled to system memory 20. The host processor 12 can include any type of processing device, such as, e.g., microcontroller, microprocessor, RISC processor, ASIC, etc., along with associated processing modules or circuitry. The system memory 20 can include any non-transitory machine-or computer-readable storage medium such as RAM, ROM, PROM, EEPROM, firmware, flash memory, etc., configurable logic such as, for example, PLAs, FPGAs, CPLDs, fixed-functionality hardware logic using circuit technology such as, for example, ASIC, CMOS or TTL technology, or any combination thereof suitable for storing instructions 28.
[0076] The system 10 can also include an input / output (I / O) module 16. The I / O module 16 can communicate with for example, one or more input / output (I / O) devices 17, a network controller 24 (e.g., wired and / or wireless NIC) , and storage 22. The storage 22 can be comprised of any appropriate non-transitory machine-or computer-readable memory type (e.g., flash memory, DRAM, SRAM (static random access memory) , solid state drive (SSD) , hard disk drive (HDD) , optical disk, etc. ) . The storage 22 can include mass storage. In some embodiments, the host processor 12 and / or the I / O module 16 can communicate with the storage 22 (all or portions thereof) via a network controller 24. In some embodiments, the system 10 can also include a graphics processor 26 (e.g., a graphics processing unit / GPU) and / or an AI accelerator 27. In some embodiments, the system 10 can also include a perception subsystem (e.g., including one or more sensors and / or cameras, not shown) and / or an actuation subsystem (not shown) . In an embodiment, the system 10 can also include a vision processing unit (VPU, not shown) .
[0077] The host processor 12 and the I / O module 16 can be implemented together on a semiconductor die as a system on chip (SoC) 11, shown encased in a solid line. The SoC 11 can therefore operate as a computing apparatus for image and / or text analysis. In some embodiments, the SoC 11 can also include one or more of the system memory 20, the network controller 24, and / or the graphics processor 26 (shown encased in dotted lines) . In some embodiments, the SoC 11 can also include other components of the system 10.
[0078] The host processor 12 and / or the I / O module 16 can execute program instructions 28 retrieved from the system memory 20 and / or the storage 22 to perform one or more aspects of the process 400, the method 500, the method 540, and / or the method 560, as described herein with reference to FIGs. 4 and 5A-5C. The system 10 can implement one or more aspects of the adapted transformer network 200, the adapted transformer layer 300, the token abstractor 330, the token filter 332 and / or the contextual token generator 333, as described herein with reference to FIGs. 2 and 3. The system 10 is therefore considered to be performance-enhanced at least to the extent that the technology provides an improved transformer network that improves throughput / speed while maintaining acceptable accuracy.
[0079] Computer program code to carry out the processes described above can be written in any combination of one or more programming languages, including an object-oriented programming language such as JAVA, JAVASCRIPT, PYTHON, SMALLTALK, C++ or the like and / or conventional procedural programming languages, such as the “C” programming language or similar programming languages, and implemented as program instructions 28. Additionally, program instructions 28 can include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, microprocessor, etc. ) .
[0080] I / O devices 17 can include one or more of input devices, such as a touchscreen, keyboard, mouse, cursor-control device, microphone, digital camera, video recorder, camcorder, biometric scanners and / or sensors; input devices can be used to enter information and interact with system 10 and / or with other devices. The I / O devices 17 can also include one or more of output devices, such as a display (e.g., touchscreen, liquid crystal display / LCD, light emitting diode / LED display, plasma panels, etc. ) , speakers and / or other visual or audio output devices. The input and / or output devices can be used, e.g., to provide a user interface.
[0081] FIG. 7 shows a block diagram illustrating an example semiconductor apparatus 30 for image and / or text analysis according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The semiconductor apparatus 30 can be implemented, e.g., as a chip, die, or other semiconductor package. The semiconductor apparatus 30 can include one or more substrates 32 comprised of, e.g., silicon, sapphire, gallium arsenide, etc. The semiconductor apparatus 30 can also include logic 34 comprised of, e.g., transistor array (s) and other integrated circuit (IC) components) coupled to the substrate (s) 32. The logic 34 can be implemented at least partly in configurable logic or fixed-functionality logic hardware. The logic 34 can implement the system on chip (SoC) 11 described above with reference to FIG. 6. The logic 34 can implement one or more aspects of the processes described above, including the process 400, the method 500, the method 540, and / or the method 560. The logic 34 can implement one or more aspects of the adapted transformer network 200, the adapted transformer layer 300, the token abstractor 330, the token filter 332 and / or the contextual token generator 333, as described herein with reference to FIGs. 2 and 3. The apparatus 30 is therefore considered to be performance-enhanced at least to the extent that the technology provides an improved transformer network that improves throughput / speed while maintaining acceptable accuracy.
[0082] The semiconductor apparatus 30 can be constructed using any appropriate semiconductor manufacturing processes or techniques. For example, the logic 34 can include transistor channel regions that are positioned (e.g., embedded) within the substrate (s) 32. Thus, the interface between the logic 34 and the substrate (s) 32 may not be an abrupt junction. The logic 34 can also be considered to include an epitaxial layer that is grown on an initial wafer of the substrate (s) 32.
[0083] FIG. 8 is a block diagram illustrating an example processor core 40 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The processor core 40 can be the core for any type of processor, such as a micro-processor, an embedded processor, a digital signal processor (DSP) , a network processor, a graphics processing unit (GPU) , or other device to execute code. Although only one processor core 40 is illustrated in FIG. 8, a processing element can alternatively include more than one of the processor core 40 illustrated in FIG. 8. The processor core 40 can be a single-threaded core or, for at least one embodiment, the processor core 40 can be multithreaded in that it can include more than one hardware thread context (or “logical processor” ) per core.
[0084] FIG. 8 also illustrates a memory 41 coupled to the processor core 40. The memory 41 can be any of a wide variety of memories (including various layers of memory hierarchy) as are known or otherwise available to those of skill in the art. The memory 41 can include one or more code 42 instruction (s) to be executed by the processor core 40. The code 42 can implement one or more aspects of the processes described above (including the process 400, the method 500, the method 540, and / or the method 560) . The processor core 40 can implement one or more aspects of the adapted transformer network 200, the adapted transformer layer 300, the token abstractor 330, the token filter 332 and / or the contextual token generator 333as described herein with reference to FIGs. 2 and 3. The processor core 40 can follow a program sequence of instructions indicated by the code 42. Each instruction can enter a front end portion 43 and be processed by one or more decoders 44. The decoder 44 can generate as its output a micro operation such as a fixed width micro operation in a predefined format, or can generate other instructions, microinstructions, or control signals which reflect the original code instruction. The illustrated front end portion 43 also includes register renaming logic 46 and scheduling logic 48, which generally allocate resources and queue the operation corresponding to the convert instruction for execution.
[0085] The processor core 40 is shown including execution logic 50 having a set of execution units 55-1 through 55-N. Some embodiments can include a number of execution units dedicated to specific functions or sets of functions. Other embodiments can include only one execution unit or one execution unit that can perform a particular function. The illustrated execution logic 50 performs the operations specified by code instructions.
[0086] After completion of execution of the operations specified by the code instructions, back end logic 58 retires the instructions of code 42. In one embodiment, the processor core 40 allows out of order execution but requires in order retirement of instructions. Retirement logic 59 can take a variety of forms as known to those of skill in the art (e.g., re-order buffers or the like) . In this manner, the processor core 40 is transformed during execution of the code 42, at least in terms of the output generated by the decoder, the hardware registers and tables utilized by the register renaming logic 46, and any registers (not shown) modified by the execution logic 50.
[0087] Although not illustrated in FIG. 8, a processing element can include other elements on chip with the processor core 40. For example, a processing element can include memory control logic along with the processor core 40. The processing element can include I / O control logic and / or can include I / O control logic integrated with memory control logic. The processing element can also include one or more caches.
[0088] FIG. 9 is a block diagram illustrating an example of a multi-processor based computing system 60 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The multiprocessor system 60 includes a first processing element 70 and a second processing element 80. While two processing elements 70 and 80 are shown, it is to be understood that an embodiment of the system 60 can also include only one such processing element.
[0089] The system 60 is illustrated as a point-to-point interconnect system, wherein the first processing element 70 and the second processing element 80 are coupled via a point-to-point interconnect 71. It should be understood that any or all of the interconnects illustrated in FIG. 9 can be implemented as a multi-drop bus rather than point-to-point interconnect.
[0090] As shown in FIG. 9, each of the processing elements 70 and 80 can be multicore processors, including first and second processor cores (i.e., processor cores 74a and 74b and processor cores 84a and 84b) . Such cores 74a, 74b, 84a, 84b can be configured to execute instruction code in a manner similar to that discussed above in connection with FIG. 8.
[0091] Each processing element 70, 80 can include at least one shared cache 99a, 99b. The shared cache 99a, 99b can store data (e.g., instructions) that are utilized by one or more components of the processor, such as the cores 74a, 74b and 84a, 84b, respectively. For example, the shared cache 99a, 99b can locally cache data stored in a memory 62, 63 for faster access by components of the processor. In one or more embodiments, the shared cache 99a, 99b can include one or more mid-level caches, such as level 2 (L2) , level 3 (L3) , level 4 (L4) , or other levels of cache, a last level cache (LLC) , and / or combinations thereof.
[0092] While shown with only two processing elements 70, 80, it is to be understood that the scope of the embodiments is not so limited. In other embodiments, one or more additional processing elements can be present in a given processor. Alternatively, one or more of the processing elements 70, 80 can be an element other than a processor, such as an accelerator or a field programmable gate array. For example, additional processing element (s) can include additional processors (s) that are the same as a first processor 70, additional processor (s) that are heterogeneous or asymmetric to processor a first processor 70, accelerators (such as, e.g., graphics accelerators or digital signal processing (DSP) units) , field programmable gate arrays, or any other processing element. There can be a variety of differences between the processing elements 70, 80 in terms of a spectrum of metrics of merit including architectural, micro architectural, thermal, power consumption characteristics, and the like. These differences can effectively manifest themselves as asymmetry and heterogeneity amongst the processing elements 70, 80. For at least one embodiment, the various processing elements 70, 80 can reside in the same die package.
[0093] The first processing element 70 can further include memory controller logic (MC) 72 and point-to-point (P-P) interfaces 76 and 78. Similarly, the second processing element 80 can include a MC 82 and P-P interfaces 86 and 88. As shown in FIG. 9, MC’s 72 and 82 couple the processors to respective memories, namely a memory 62 and a memory 63, which can be portions of main memory locally attached to the respective processors. While the MC 72 and 82 is illustrated as integrated into the processing elements 70, 80, for alternative embodiments the MC logic can be discrete logic outside the processing elements 70, 80 rather than integrated therein.
[0094] The first processing element 70 and the second processing element 80 can be coupled to an I / O subsystem 90 via P-P interconnects 76 and 86, respectively. As shown in FIG. 9, the I / O subsystem 90 includes P-P interfaces 94 and 98. Furthermore, the I / O subsystem 90 includes an interface 92 to couple I / O subsystem 90 with a high performance graphics engine 64. In one embodiment, a bus 73 can be used to couple the graphics engine 64 to the I / O subsystem 90. Alternately, a point-to-point interconnect can couple these components.
[0095] In turn, the I / O subsystem 90 can be coupled to a first bus 65 via an interface 96. In one embodiment, the first bus 65 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the embodiments are not so limited.
[0096] As shown in FIG. 9, various I / O devices 65a (e.g., biometric scanners, speakers, cameras, and / or sensors) can be coupled to the first bus 65, along with a bus bridge 66 which can couple the first bus 65 to a second bus 67. In one embodiment, the second bus 67 can be a low pin count (LPC) bus. Various devices can be coupled to the second bus 67 including, for example, a keyboard / mouse 67a, communication device (s) 67b, and a data storage unit 68 such as a disk drive or other mass storage device which can include code 69, in one embodiment. The illustrated code 69 can implement one or more aspects of the processes described above, including the process 400, the method 500, the method 540, and / or the method 560. The illustrated code 69 can be similar to the code 42 (FIG. 8) , already discussed. Further, an audio I / O 67c can be coupled to second bus 67 and a battery 61 can supply power to the computing system 60. The system 60 can implement one or more aspects of the adapted transformer network 200, the adapted transformer layer 300, the token abstractor 330, the token filter 332 and / or the contextual token generator 333 as described herein with reference to FIGs. 2 and 3.
[0097] Note that other embodiments are contemplated. For example, instead of the point-to-point architecture of FIG. 9, a system can implement a multi-drop bus or another such communication topology. Also, the elements of FIG. 9 can alternatively be partitioned using more or fewer integrated chips than shown in FIG. 9.
[0098] Additional Notes and Examples:
[0099] Example S1 includes a performance-enhanced computing system, comprising a processor, and a memory coupled to the processor, the memory storing a neural network, the neural network comprising a plurality of transformer layers, and one or more adapted transformer layers arranged among the plurality of transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including a token filter to filter out one or more of a plurality of tokens, and a contextual token generator to generate a contextual token representing residual information from the plurality of tokens, wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.
[0100] Example S2 includes the computing system of Example S1, wherein the token filter is to generate a column sum based on a set of attention scores from the first attention network, sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices, and select, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.
[0101] Example S3 includes the computing system of Example S1 or S2, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to generate a column mean based on the set of attention scores from the first attention network, and generate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.
[0102] Example S4 includes the computing system of any of Examples S1-S3, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.
[0103] Example S5 includes the computing system of any of Examples S1-S4, wherein each of the plurality of adapted transformer layers uses the same reduction factor.
[0104] Example S6 includes the computing system of any of Examples S1-S4, wherein at least two of the plurality of adapted transformer layers use different reduction factors.
[0105] Example S7 includes the computing system of any of Examples S1-S6, wherein the contextual token generator is to use an output of the token filter to mask out the subset of tokens from an input to the second attention network.
[0106] Example S8 includes the computing system of any of Examples S1-S7, wherein the plurality of transformer layers is pre-trained.
[0107] Example A1 includes a semiconductor apparatus comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable logic or fixed-functionality hardware logic, the logic coupled to the one or more substrates comprising a neural network, the neural network comprising a plurality of transformer layers, and one or more adapted transformer layers arranged among the plurality of transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including a token filter to filter out one or more of a plurality of tokens, and a contextual token generator to generate a contextual token representing residual information from the plurality of tokens, wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.
[0108] Example A2 includes the semiconductor apparatus of Example A1, wherein the token filter is to generate a column sum based on a set of attention scores from the first attention network, sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices, and select, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.
[0109] Example A3 includes the semiconductor apparatus of Example A1 or A2, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to generate a column mean based on the set of attention scores from the first attention network, and generate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.
[0110] Example A4 includes the semiconductor apparatus of any of Examples A1-A3, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.
[0111] Example A5 includes the semiconductor apparatus of any of Examples A1-A4, wherein each of the plurality of adapted transformer layers uses the same reduction factor.
[0112] Example A6 includes the semiconductor apparatus of any of Examples A1-A4, wherein at least two of the plurality of adapted transformer layers use different reduction factors.
[0113] Example A7 includes the semiconductor apparatus of any of Examples A1-A6, wherein the contextual token generator is to use an output of the token filter to mask out the subset of tokens from an input to the second attention network.
[0114] Example A8 includes the semiconductor apparatus of any of Examples A1-A7, wherein the plurality of transformer layers is pre-trained.
[0115] Example C1 includes at least one computer readable storage medium comprising a set of instructions which, when executed by a computing system, cause the computing system to generate a neural network comprising a plurality of transformer layers, and adapt one or more of the plurality of transformer layers to form one or more adapted transformer layers, wherein each adapted transformer layer is to include a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including a token filter to filter out one or more of a plurality of tokens, and a contextual token generator to generate a contextual token representing residual information from the plurality of tokens, wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.
[0116] Example C2 includes the at least one computer readable storage medium of Example C1, wherein the token filter is to generate a column sum based on a set of attention scores from the first attention network, sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices, and select, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.
[0117] Example C3 includes the at least one computer readable storage medium of Example C1 or C2, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to generate a column mean based on the set of attention scores from the first attention network, and generate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.
[0118] Example C4 includes the at least one computer readable storage medium of any of Examples C1-C3, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.
[0119] Example C5 includes the at least one computer readable storage medium of any of Examples C1-C4, wherein each of the plurality of adapted transformer layers uses the same reduction factor.
[0120] Example C6 includes the at least one computer readable storage medium of any of Examples C1-C4, wherein at least two of the plurality of adapted transformer layers use different reduction factors.
[0121] Example C7 includes the at least one computer readable storage medium of any of Examples C1-C6, wherein the contextual token generator is to use an output of the token filter to mask out the subset of tokens from an input to the second attention network.
[0122] Example C8 includes the at least one computer readable storage medium of any of Examples C1-C7, wherein the plurality of transformer layers is pre-trained.
[0123] Example M1 includes a method comprising generating a neural network comprising a plurality of transformer layers, and adapting one or more of the plurality of transformer layers to form one or more adapted transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including a token filter that filters out one or more of a plurality of tokens, and a contextual token generator that generates a contextual token representing residual information from the plurality of tokens, wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.
[0124] Example M2 includes the method of Example M1, wherein the token filter performs operations comprising generating a column sum based on a set of attention scores from the first attention network, sorting, based on the column sum, indices of the plurality of tokens to form a set of sorted indices, and selecting, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.
[0125] Example M3 includes the method of Example M1 or M2, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator performs operations comprising generating a column mean based on the set of attention scores from the first attention network, and generating, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.
[0126] Example M4 includes the method of any of Examples M1-M3, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.
[0127] Example M5 includes the method of any of Examples M1-M4, wherein each of the plurality of adapted transformer layers uses the same reduction factor.
[0128] Example M6 includes the method of any of Examples M1-M4, wherein at least two of the plurality of adapted transformer layers use different reduction factors.
[0129] Example M7 includes the method of any of Examples M1-M6, wherein the contextual token generator uses an output of the token filter to mask out the subset of tokens from an input to the second attention network.
[0130] Example M8 includes the computing system of any of Examples M1-M7, wherein the plurality of transformer layers is pre-trained.
[0131] Example M9 includes the method of any of Examples M1-M8, wherein a number N of transformer layers are adapted, and wherein the method further comprises determining one or more of the number N or the reduction factor based on meeting optimization criteria.
[0132] Example M10 includes the method of any of Examples M1-M9, wherein determining one or more of the number N or the reduction factor includes initializing the number N, inserting N token abstractors into the plurality of transformer layers to form N adapted transformer layers, initializing the reduction factor, training the token abstractors, and evaluating accuracy and performance of the neural network with the trained token abstractors.
[0133] Example M11 includes the method of any of Examples M1-M10, wherein determining one or more of the number N or the reduction factor further includes adjusting one or more of the number N or the reduction factor if the accuracy and performance of the neural network do not meet accuracy and performance tradeoff criteria.
[0134] Example M12 includes the method of any of Examples M1-M11, wherein the plurality of transformer layers is trained or re-trained together with training of the token abstractors.
[0135] Example AM1 includes an apparatus comprising means for performing the method of any of Examples M1 to M12.
[0136] Embodiments are applicable for use with all types of semiconductor integrated circuit ( “IC” ) chips. Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs) , memory chips, network chips, systems on chip (SoCs) , solid state drive (SSD) / NAND drive controller ASICs, and the like. In addition, in some of the drawings, signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and / or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner. Rather, such added detail may be used in connection with one or more exemplary embodiments to facilitate easier understanding of a circuit. Any represented signal lines, whether or not having additional information, may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and / or single-ended lines.
[0137] Example sizes / models / values / ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured. In addition, well known power / ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the platform within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art. Where specific details (e.g., circuits) are set forth in order to describe example embodiments, it should be apparent to one skilled in the art that embodiments can be practiced without, or with variation of, these specific details. The description is thus to be regarded as illustrative instead of limiting.
[0138] The term “coupled” may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections, including logical connections via intermediate components (e.g., device A may be coupled to device C via device B) . In addition, the terms “first” , “second” , etc. may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.
[0139] As used in this application and in the claims, a list of items joined by the term “one or more of” may mean any combination of the listed terms. For example, the phrases “one or more of A, B or C” may mean A, B, C; A and B; A and C; B and C; or A, B and C.
[0140] Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Therefore, while the embodiments have been described in connection with particular examples thereof, the true scope of the embodiments should not be so limited since other modifications will become apparent to the skilled practitioner upon a study of the drawings, specification, and following claims.
Claims
1.A computing system, comprising:a processor; anda memory coupled to the processor, the memory storing a neural network, the neural network comprising:a plurality of transformer layers; andone or more adapted transformer layers arranged among the plurality of transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including:a token filter to filter out one or more of a plurality of tokens; anda contextual token generator to generate a contextual token representing residual information from the plurality of tokens,wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.2.The computing system of claim 1, wherein the token filter is to:generate a column sum based on a set of attention scores from the first attention network;sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices; andselect, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.3.The computing system of claim 2, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to:generate a column mean based on the set of attention scores from the first attention network; andgenerate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.4.The computing system of any one of claims 1-3, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.5.The computing system of claim 4, wherein each of the plurality of adapted transformer layers uses the same reduction factor.6.The computing system of claim 4, wherein at least two of the plurality of adapted transformer layers use different reduction factors.7.A semiconductor apparatus comprising:one or more substrates; andlogic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable logic or fixed-functionality hardware logic, the logic comprising a neural network, the neural network comprising:a plurality of transformer layers; andone or more adapted transformer layers arranged among the plurality of transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including:a token filter to filter out one or more of a plurality of tokens; anda contextual token generator to generate a contextual token representing residual information from the plurality of tokens,wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.8.The semiconductor apparatus of claim 7, wherein the token filter is to:generate a column sum based on a set of attention scores from the first attention network;sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices; andselect, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.9.The semiconductor apparatus of claim 8, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to:generate a column mean based on the set of attention scores from the first attention network; andgenerate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.10.The semiconductor apparatus of any one of claims 7-9, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.11.The semiconductor apparatus of claim 10, wherein each of the plurality of adapted transformer layers uses the same reduction factor.12.The semiconductor apparatus of claim 10, wherein at least two of the plurality of adapted transformer layers use different reduction factors.13.At least one computer readable storage medium comprising a set of instructions which, when executed by a computing system, cause the computing system to:generate a neural network comprising a plurality of transformer layers; andadapt one or more of the plurality of transformer layers to form one or more adapted transformer layers, wherein each adapted transformer layer is to include a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including:a token filter to filter out one or more of a plurality of tokens; anda contextual token generator to generate a contextual token representing residual information from the plurality of tokens,wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.14.The at least one computer readable storage medium of claim 13, wherein the token filter is to:generate a column sum based on a set of attention scores from the first attention network;sort, based on the column sum, indices of the plurality of tokens to form a set of sorted indices; andselect, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.15.The at least one computer readable storage medium of claim 14, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator is to:generate a column mean based on the set of attention scores from the first attention network; andgenerate, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.16.The at least one computer readable storage medium of any one of claims 13-15, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.17.The at least one computer readable storage medium of claim 16, wherein each of the plurality of adapted transformer layers uses the same reduction factor.18.The at least one computer readable storage medium of claim 16, wherein at least two of the plurality of adapted transformer layers use different reduction factors.19.A method comprising:generating a neural network comprising a plurality of transformer layers; andadapting one or more of the plurality of transformer layers to form one or more adapted transformer layers, wherein each adapted transformer layer includes a token abstractor coupled between a first attention network and a feed forward network, the token abstractor including:a token filter that filters out one or more of a plurality of tokens; anda contextual token generator that generates a contextual token representing residual information from the plurality of tokens,wherein the first attention network comprises one of a multi-head self-attention (MHSA) network or a multi-head attention (MHA) network.20.The method of claim 19, wherein the token filter performs operations comprising:generating a column sum based on a set of attention scores from the first attention network;sorting, based on the column sum, indices of the plurality of tokens to form a set of sorted indices; andselecting, based on a reduction factor and the set of sorted indices, a subset of tokens from the plurality of tokens.21.The method of claim 20, wherein the contextual token generator includes a second attention network, wherein the second attention network comprises one of a MHSA network or a MHA network, and wherein the contextual token generator performs operations comprising:generating a column mean based on the set of attention scores from the first attention network; andgenerating, via the second attention network, a contextual token based on the column mean and a key and a value from the first attention network.22.The method of any one of claims 19-21, wherein the one or more adapted transformer layers includes a plurality of adapted transformer layers arranged at essentially uniform intervals among the plurality of transformer layers.23.The method of claim 22, wherein each of the plurality of adapted transformer layers uses the same reduction factor.24.The method of claim 22, wherein at least two of the plurality of adapted transformer layers use different reduction factors.25.The method of claim 22, wherein a number N of transformer layers are adapted, and wherein the method further comprises determining one or more of the number N or the reduction factor based on meeting optimization criteria.
Citation Information
Patent Citations
Method and apparatus for speaking time estimation
CN113674733A
Use of extracellular vesicles for targeted delivery of immune checkpoint inhibitor
KR1020240154736A
Method for translate sign language gloss using transformer, and computer program recorded on record-medium for executing method thereof
KR102571902B1
Systems and methods for translating with limited attention
US11410015B1