Inference operation method and inference operation device

CN122819461APending Publication Date: 2026-09-25HANGZHOU SUPERACME MICROELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610958217.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0092]此外,本公开的技术方案在克服了Flash Attention等方案中缺点的基础上,还可以产生以下额外的积极效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819461A_ABST
    Figure CN122819461A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for performing inference operation. In one aspect, the method for performing inference operation includes operations performed by each of a plurality of first computing components to obtain input hidden state of a Transformer layer and one or more sets of head weight data associated with one or more heads associated with the first computing component, to compute one or more sets of head branch data corresponding to the one or more heads based on the input hidden state of the Transformer layer and the one or more sets of head weight data, and to compute one or more attention head outputs corresponding to the one or more heads based on the one or more sets of head branch data. The method further includes operations performed by a second computing component to reduce the plurality of attention head outputs from the plurality of first computing components to generate a multi-head attention output of the multi-head attention sub-layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence (AI) technology, and more specifically, to a reasoning operation method, reasoning operation device, computer-readable storage medium, and computer program product for an attention-based model. Background Technology

[0002] In the field of AI technology, with the rapid growth (e.g., exponential growth) of the parameter scale and computing power requirements of large models, the performance requirements for hardware storage, communication, and computing are becoming increasingly demanding. Therefore, it is necessary to improve the inference and computation methods of these models. Summary of the Invention

[0003] One of the purposes of this disclosure is to provide an inference operation method, inference operation device, computer-readable storage medium, and computer program product for attention-based models.

[0004] According to a first aspect of this disclosure, an inference operation method for an attention-based model is proposed, wherein the attention-based model includes a Transformer layer, the Transformer layer includes a multi-head attention sublayer, and the inference operation method includes: Each of the plurality of first computing units performs the following operations: Obtain the input hidden state of the Transformer layer and one or more sets of head weight data corresponding one-to-one with one or more heads associated with the first computation unit. Each set of head weight data corresponding to each head includes query sub-weight data corresponding to that head in the query weight data, key sub-weight data corresponding to that head in the key weight data, and value sub-weight data corresponding to that head in the value weight data. Based on the input hidden state of the Transformer layer and the one or more sets of head weight data, calculate one or more sets of head branch data corresponding one-to-one with the one or more heads, wherein each set of head branch data corresponding to each head includes query sub-data, key data, and value sub-data corresponding to that head, and Based on the set or more sets of head branch data, calculate one or more attention head outputs corresponding one-to-one with the one or more heads; and The second computing unit reduces the multiple attention head outputs from the plurality of first computing units to generate the multi-head attention output of the multi-head attention sublayer.

[0005] In some embodiments, the input hidden state of the Transformer layer is broadcast to each of the plurality of first computing units.

[0006] In some embodiments, the calculation of each group of head branch data corresponding to each head is performed independently of each other; and / or

[0007] The calculation of the output of each attention head corresponding to each head is performed independently.

[0008] In some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, includes: Based on the input hidden state and the query sub-weight data corresponding to the head, calculate the query sub-data corresponding to the head; or Based on the sub-hidden state corresponding to the newly added lexical unit in the input hidden state and the key weight data corresponding to the head, calculate the new key data corresponding to the newly added lexical unit; and based on the historical key data and the newly added key data, calculate the key data corresponding to the head, wherein the historical key data is generated based on the sub-hidden state corresponding to the historical lexical unit in the input hidden state; or Based on the sub-hidden state corresponding to the newly added word in the input hidden state and the value sub-weight data corresponding to the head, the newly added value sub-data corresponding to the newly added word is calculated, and based on the historical value sub-data and the newly added value sub-data, the value data corresponding to the head is calculated, wherein the historical value sub-data is generated based on the sub-hidden state corresponding to the historical word in the input hidden state.

[0009] In some embodiments, calculating the query sub-data corresponding to the head based on the input hidden state and the query sub-weight data corresponding to the head includes: splitting the input hidden state into first sub-state data and second sub-state data according to the feature dimension, and splitting the query sub-weight data into first query sub-weight data and second query sub-weight data; calculating the first query sub-data based on the first sub-state data and the first query sub-weight data; calculating the second query sub-data based on the second sub-state data and the second query sub-weight data; and calculating the query sub-data based on the first query sub-data and the second query sub-weight data; or

[0010] Calculating the new key data corresponding to the new word element based on the sub-hidden state corresponding to the new word element in the input hidden state and the key weight data corresponding to the head includes: splitting the sub-hidden state corresponding to the new word element in the input hidden state into third sub-state data and fourth sub-state data according to the feature dimension, and splitting the key weight data into first key weight data and second key weight data; calculating the first new key data based on the third sub-state data and the first key weight data; calculating the second new key data based on the fourth sub-state data and the second key weight data; and calculating the new key data based on the first new key data and the second new key data; or

[0011] Calculating the new value sub-data corresponding to the new word in the input hidden state based on the sub-hidden state corresponding to the new word and the value sub-weight data corresponding to the head includes: splitting the sub-hidden state corresponding to the new word in the input hidden state into fifth sub-state data and sixth sub-state data according to the feature dimension, and splitting the value sub-weight data into first value sub-weight data and second value sub-weight data; calculating the first new value sub-data based on the fifth sub-state data and the first value sub-weight data; calculating the second new value sub-data based on the sixth sub-state data and the second value sub-weight data; and calculating the new value sub-data based on the first new value sub-data and the second new value sub-data.

[0012] In some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, further includes: The calculated query sub-data is cached in a first cache unit corresponding to a first calculation unit that performs the calculation; or The newly calculated key data is cached in a second cache unit corresponding to a first calculation unit that performs the calculation, wherein the second cache unit is further configured to cache at least a portion of the historical key data; or The newly calculated value-added sub-data is cached in a third cache unit corresponding to a first calculation unit that performs the calculation, wherein the third cache unit is also configured to cache at least a portion of the historical value sub-data.

[0013] In some embodiments, at least one of the first cache component, the second cache component, and the third cache component is integrated into the same chip as a corresponding first computing component; or

[0014] At least one of the first cache component, the second cache component, and the third cache component is further configured to perform a transpose operation in the computation.

[0015] In some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, further includes: In response to the remaining cache space of the second cache component being less than or equal to a first preset threshold, the data cached in the second cache component is written page by page to the storage component corresponding to a first computing component performing the computation; or In response to the remaining cache space of the third cache component being less than or equal to a second preset threshold, the data cached in the third cache component is written page by page into the storage component corresponding to a first computing component that performs the calculation.

[0016] In some embodiments, the storage component and the corresponding first computing component are disposed separately from each other; or

[0017] All query sub-weight data, all key sub-weight data, all value sub-weight data, historical key data in all key data, and historical value sub-data in all value data corresponding to one or more heads associated with each first calculation unit are stored in the same storage unit corresponding to that first calculation unit.

[0018] In some embodiments, reducing the multiple attention head outputs from the plurality of first computing units by a second computing unit to generate the multi-head attention output of the multi-head attention sublayer includes: The second computing unit performs at least one of the following calculations on all attention head outputs from the plurality of first computing units: summation, concatenation, weighted averaging, and maximum aggregation, to generate the multi-head attention output of the multi-head attention sublayer.

[0019] In some embodiments, the plurality of first computing units and the second computing unit are respectively configured; or

[0020] The second computing unit is implemented by one of the plurality of first computing units; or

[0021] The plurality of first computing units are configured to operate in parallel with each other.

[0022] According to a second aspect of this disclosure, an inference computing apparatus for an attention-based model is provided, wherein the attention-based model includes a Transformer layer, the Transformer layer includes a multi-head attention sublayer, and the inference computing apparatus includes a plurality of first computing units and second computing units, wherein each first computing unit is configured to perform the following operations: Obtain the input hidden state of the Transformer layer and one or more sets of head weight data corresponding one-to-one with one or more heads associated with the first computation unit. Each set of head weight data corresponding to each head includes query sub-weight data corresponding to that head in the query weight data, key sub-weight data corresponding to that head in the key weight data, and value sub-weight data corresponding to that head in the value weight data. Based on the input hidden state of the Transformer layer and the one or more sets of head weight data, calculate one or more sets of head branch data corresponding one-to-one with the one or more heads, wherein each set of head branch data corresponding to each head includes query sub-data, key data, and value sub-data corresponding to that head, and Based on the one or more sets of head branch data, calculate one or more attention head outputs that correspond one-to-one with the one or more heads; The second computing unit is configured to reduce the multiple attention head outputs from the plurality of first computing units to generate the multi-head attention output of the multi-head attention sublayer; The plurality of first computing components and the second computing component are respectively configured, or the second computing component is implemented by one of the plurality of first computing components.

[0023] In some embodiments, the inference computing device further includes: One or more cache units corresponding to each first computing unit, wherein each cache unit is configured to cache at least a portion of the data related to the computation performed by the first computing unit; One or more storage units corresponding to each first computing unit, wherein each storage unit is configured to store at least a portion of the data related to the computation performed by the first computing unit.

[0024] In some embodiments, a first computing unit and one or more cache units corresponding to the first computing unit are integrated into the same chip; or

[0025] A first computing unit and one or more storage units corresponding to the first computing unit are disposed separately from each other; or

[0026] Cache components include static random access memory (SRAM); or

[0027] Storage components include dynamic random access memory (DRAM); or

[0028] The granularity at which data is split or transposed during computation is configured to dynamically adjust based on at least one of the real-time remaining capacity of the cache component and the current access bandwidth of the storage component.

[0029] According to a third aspect of this disclosure, a computer-readable storage medium is provided, on which instructions are stored, which, when executed by a processor, implement the reasoning operation method described above.

[0030] According to a fourth aspect of this disclosure, a computer program product is proposed, including instructions that, when executed by a processor, implement the reasoning operation method as described above.

[0031] Other features and advantages of this disclosure will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0032] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.

[0033] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein: Figure 1 A schematic diagram of the structure of a model based on an attention mechanism is shown; Figure 2 A schematic diagram of the structure of an inference computing apparatus for an attention-based model, according to an exemplary embodiment of the present disclosure, is shown. Figure 3 A schematic diagram of the structure of an inference computing apparatus for an attention-based model is shown according to another exemplary embodiment of the present disclosure; Figure 4 A flowchart illustrating an inference operation method for an attention-based model according to an exemplary embodiment of the present disclosure is shown. Figure 5 A schematic diagram of the computation process of an inference operation method for an attention-based model according to a specific embodiment of the present disclosure is shown. Figure 6 A schematic diagram is shown illustrating the splitting of query weight data in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure; Figure 7 A schematic diagram of splitting key weight data in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure is shown. Figure 8 A schematic diagram of split value weight data in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure is shown. Figure 9 A schematic diagram is shown illustrating the splitting of input hidden states and query sub-weight data in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure; Figure 10 A schematic diagram is shown illustrating the splitting of sub-hidden states and key weight data of the input hidden state in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure. Figure 11 A schematic diagram is shown illustrating the splitting of sub-hidden states and value sub-weight data of the input hidden state in an inference operation method for an attention-based model according to a specific embodiment of the present disclosure.

[0034] Note that in the embodiments described below, the same reference numerals are sometimes used across different figures to denote the same parts or parts having the same function, and repeated descriptions are omitted. In this specification, similar reference numerals and letters are used to denote similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0035] For ease of understanding, the positions, dimensions, and extents of the structures shown in the accompanying drawings and other materials may not represent actual positions, dimensions, and extents. Therefore, the disclosed invention is not limited to the positions, dimensions, and extents disclosed in the accompanying drawings and other materials. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components. Detailed Implementation

[0036] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0037] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this disclosure or its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in this disclosure. However, those skilled in the art will understand that they merely illustrate exemplary ways that can be used to implement this disclosure, and not exhaustive ways. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components.

[0038] In addition, techniques, methods and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods and equipment should be considered part of the specification.

[0039] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0040] Attention mechanisms are a core computational mechanism in AI fields such as large language models and computer vision. They achieve weighted fusion of features through corresponding calculations on query (Q) data, key (K) data, and value (V) data. During attention computation, the design and segmentation of the data flow directly determine the computational efficiency, hardware resource consumption, and communication overhead of the attention layer.

[0041] One attention computation method employs Flash Attention as a dataflow optimization scheme for the attention mechanism. Its core is a tile-based memory access approach. By dividing Q, K, and V data into multiple smaller blocks, on-chip Static Random-Access Memory (SRAM) is used for block computation and online softmax merging. This effectively reduces the access volume to off-chip Dynamic Random-Access Memory (DRAM) or High Bandwidth Memory (HBM), improving memory access efficiency. However, this Flash Attention scheme has at least the following drawbacks, making it difficult to implement in scenarios with limited SRAM resources: First, it requires a large on-chip SRAM cache capacity. Specifically, Tile-based partitioning needs to cache multiple Q blocks, K blocks, and V blocks, as well as intermediate calculation results, simultaneously. If the on-chip SRAM capacity in the inference processing unit is insufficient, it will be difficult to support the block caching requirements, resulting in a significant reduction in the optimization effect of the solution or even its unusability.

[0042] Second, there is a communication bottleneck across computing components. Specifically, if multiple computing components are used for parallel computation, the tile-based segmentation in Flash Attention may require cross-component data transfer operations for, for example, K data and V data. When the model's inference performance requirements are high (e.g., a high number of tokens processed per unit time) or in long context scenarios, the transfer of K data and V data will consume a large amount of hardware communication bandwidth, becoming a bottleneck for overall computational performance.

[0043] Third, softmax has high computational complexity. Specifically, in Tile-style partitioning, the softmax intermediate variables of each block need to be split, passed, and merged, which increases the complexity of the computational logic and introduces coupling dependencies between intermediate data.

[0044] Fourth, data read / write efficiency is low. Specifically, in tile-based partitioning, the length of the lowest dimension is limited by the partitioning rules. In practical applications, a stride-based memory access method must be used to attempt to improve performance. However, this method still cannot improve the memory access efficiency loss caused by the excessive proportion of DRAM activation and precharge operations. Especially in the application scenario of three-dimensional (3D) DRAM, this defect means that the bandwidth improvement advantage brought by the large number of through-silicon vias (TSVs) added in 3D DRAM cannot be fully utilized.

[0045] Fifth, there is a waste of parallel resources. Specifically, there is some data flow between the weight data, K data, and V data of each attention head, and the intermediate hidden state needs to be synchronized multiple times, which increases the communication overhead between computing components and makes it difficult to achieve true isolation.

[0046] To address at least one of the aforementioned problems, this disclosure proposes a reasoning method and apparatus for attention-based models. For example... Figure 1 As shown, the attention-based model 100 may include a Transformer layer 110. The number of Transformer layers 110 included in the model 100 may be one or more. Further, as... Figure 1 As shown, when model 100 includes multiple Transformer layers 110, the multiple Transformer layers 110 can be cascaded, and the output of the previous Transformer layer 110 can be provided as the input of the next Transformer layer 110 to participate in further computation. In some embodiments, at least two Transformer layers 110 can have the same structure, only the values ​​of the specific parameters involved can be different. Alternatively, in other embodiments, at least two Transformer layers 110 can have different structures, which is not limited here. Further, as Figure 1 As shown, at least one Transformer layer 110 may include a multi-head attention sub-layer 111, which can generate a corresponding multi-head attention output.

[0047] The inference operation method disclosed herein can be run on an inference operation device, or in other words, the inference operation device disclosed herein can be configured to execute the inference operation method. The inference operation method and the inference operation device will be described in detail below.

[0048] In one exemplary embodiment of this disclosure, such as Figure 2 and Figure 3As shown, the inference processing device 200 may include a plurality of first computing units 211 and second computing units 212. The first computing units 211 and second computing units 212 may be configured to perform corresponding calculations, as will be described in detail below. Furthermore, in some embodiments, such as... Figure 2 As shown, multiple first computing units 211 and second computing units 212 can be separately configured, meaning that the second computing unit 212 can be a different computing unit than any of the first computing units 211. Alternatively, in other embodiments, such as... Figure 3 As shown, the second computing unit 212 can be implemented by one of the multiple first computing units 211, that is, a certain computing unit can serve as both a first computing unit 211 and a second computing unit 212, so as to save hardware resources.

[0049] In some embodiments, such as Figure 2 and Figure 3 As shown, the inference processing device 200 may further include one or more cache units 220 corresponding to each of the first computing units 211. Each cache unit 220 may be configured to cache at least a portion of the data related to the computation performed by the first computing unit 211, as will be described in detail later. In some embodiments, such as Figure 2 and Figure 3 As shown, the cache unit 220 corresponding to the same first computing unit 211 may include a first cache unit 221, a second cache unit 222, and a third cache unit 223, to respectively store the corresponding data involved in the calculation process performed by the first computing unit 211, as will be described in detail later. Further, in some embodiments, the first computing unit 211 and one or more cache units 220 corresponding to the first computing unit 211 may be integrated in the same chip. In addition, the cache unit 220 may include static random access memory (SRAM). Therefore, in some embodiments, the cache unit 220 may be in the form of on-chip SRAM relative to the first computing unit 211.

[0050] In some embodiments, such as Figure 2 and Figure 3As shown, the inference processing device 200 may further include one or more storage units 230 corresponding to each of the first computing units 211. Each storage unit 230 may be configured to store at least a portion of the data related to the computation performed by the first computing unit 211, as will be described in detail later. Further, in some embodiments, the first computing unit 211 and the one or more storage units 230 corresponding to it may be disposed separately from each other. In addition, the storage unit 230 may include dynamic random access memory (DRAM), such as 3D DRAM. Thus, in some embodiments, the storage unit 230 may be manifested as off-chip DRAM relative to the first computing unit 211. It should also be noted that the various cache or storage units, cache or storage sub-units, and cache or storage cells described herein are intended to include, but are not limited to, these and any other suitable types of memory.

[0051] During inference operations, the amount of data involved may be large, necessitating data splitting. The granularity of data splitting during computation can be dynamically adjusted based on at least one of the real-time remaining capacity of cache component 220 and the current access bandwidth of storage component 230, as will be explained in detail later. Furthermore, inference operations may also involve the transposition of data (e.g., tensors, matrices, or vectors). At least partial transposition can be performed in cache component 220. Accordingly, the granularity of data transposition during computation can be dynamically adjusted based on at least one of the real-time remaining capacity of cache component 220 and the current access bandwidth of storage component 230, as will be explained in detail later.

[0052] In one exemplary embodiment of this disclosure, such as Figure 4 As shown, the inference operation method 400 may include each of the plurality of first computing units 211 performing the following operations: step S410, obtaining the input hidden state of the Transformer layer and a set or more sets of head weight data corresponding one-to-one with one or more heads associated with the first computing unit 211; step S420, calculating a set or more sets of head branch data corresponding one-to-one with one or more heads based on the input hidden state of the Transformer layer and the set or more sets of head weight data; step S430, calculating one or more attention head outputs corresponding one-to-one with one or more heads based on the set or more sets of head branch data; and step S440, reducing the multiple attention head outputs from the plurality of first computing units 211 by the second computing unit 212 to generate multi-head attention outputs of the multi-head attention sub-layer.

[0053] In other words, in the exemplary embodiments of this disclosure, the relevant data and computations can be split according to the head dimension, so that each first computation unit 211 can be configured to acquire only one or more sets of head weight data corresponding one-to-one with one or more heads associated with it, without acquiring head weight data corresponding to other heads processed by other first computation units 211. Combined with the acquired input hidden state of the Transformer layer, the first computation unit 211 can be configured to perform only the computation corresponding to one or more heads associated with it, that is, to calculate one or more sets of head branch data corresponding one-to-one with the one or more heads, and then calculate one or more attention head outputs corresponding one-to-one with the one or more heads, without processing computations related to other heads, thereby effectively reducing the amount of data transmission and computation associated with each first computation unit 211, and reducing the communication and computational burden. In addition, it should be noted that, depending on the specific operation rules of the multi-head attention mechanism, the various data involved in this disclosure, such as input hidden state, weight data, head weight data, head branch data, attention head output, and multi-head attention output, may be represented in the form of tensors, matrices, or vectors.

[0054] Furthermore, in some embodiments, although the same first computing unit 211 may be configured to perform calculations related to multiple heads, the calculations related to each head can be performed independently, thereby minimizing the interaction of related data and coupling of calculations between different heads. Specifically, in some embodiments, the calculations of each group of head branch data corresponding to each head can be performed independently. Additionally, in some embodiments, the calculations of each attention head output corresponding to each head can be performed independently. In the following sections, the technical solutions of this disclosure will be described in detail using the calculation related to one head in the first computing unit 211 as an example. Those skilled in the art will understand that specific calculations related to other heads can be performed with reference to this method.

[0055] In some embodiments, such as Figure 5 As shown, the input hidden state of the Transformer layer can be broadcast to each of the multiple first computation units 211, so that each first computation unit 211 can obtain the input hidden state of the Transformer layer for subsequent computation. In some embodiments, the input hidden state can be in the form of a tensor with dimensions (B, S, H), where B represents the batch size, i.e., the number of samples processed simultaneously in one forward / backward propagation, S represents the sequence length, i.e., the number of tokens contained in a single sample, and H represents the hidden dimension, which is related to the structure of the model.

[0056] Furthermore, as described above, the first computing unit 211 is also configured to acquire one or more sets of head weight data corresponding one-to-one with one or more heads associated with it. Each set of head weight data corresponding to each head may include query weight data W. Q The query subweight data W corresponding to this header Qh Key weight data W K The key weight data W corresponding to this head Kh Sum-weighted data W V The sub-weight data W corresponding to this head Vh In some embodiments, the weight data W is queried. Q Key weight data W K Sum-weighted data W V Either of these can be in the form of a tensor or a matrix, weighted according to the head dimension of the query data W. Q Key weight data W K Sum-weighted data W V By splitting the data, we can obtain query sub-weight data W that corresponds one-to-one with each head. Qh Key weight data W Kh Sum of sub-weights data W Vh The query subweight data W here Qh Key weight data W Kh Sum of sub-weights data W Vh Either of these can be in matrix or vector form. In a specific example, such as Figure 6 As shown, query weight matrix W Q It can be split into multiple query sub-weight vectors W according to the head dimension. Qh (For example, Figure 6 W shown Qh1 ~W Qh8 ), where each query sub-weight vector W Qh To query the weight matrix W Q The corresponding column in. Similarly, such as Figure 7 As shown, the key weight matrix W K It can be split into multiple key weight vectors W according to the head dimension. Kh (For example, Figure 7 The W shown Kh1 ~W Kh8 ), where each key weight vector W Kh The key weight matrix W K The corresponding column in the table. For example... Figure 8 As shown, the value weight matrix W V It can be split into multiple value sub-weight vectors W according to the head dimension. Vh (For example, Figure 8The W shown Vh1 ~W Vh8 ), where each value sub-weight vector W Vh Value weight matrix W V The corresponding column in the table.

[0057] In some embodiments, each first computation unit 211 can be configured to perform attention computation related to only one head. In this case, the number of query sub-weight data (matrices or vectors) obtained from the split can be equal to the number of heads; similarly, the number of key sub-weight data (matrices or vectors) obtained from the split can be equal to the number of heads, and the number of value sub-weight data (matrices or vectors) obtained from the split can be equal to the number of heads. Alternatively, in other embodiments, each first computation unit 211 can be configured to perform attention computation related to N1 heads, where N1 is an integer greater than 1 and less than the total number of heads N2, and N2 is divisible by N1. In this case, the number of query sub-weight data (matrices or vectors) obtained from the split can be equal to N2 / N1; similarly, the number of key sub-weight data (matrices or vectors) obtained from the split can be equal to N2 / N1, and the number of value sub-weight data (matrices or vectors) obtained from the split can be equal to N2 / N1. In still other embodiments, at least two of the plurality of first computation units 211 can be configured to perform attention computation related to different numbers of heads, respectively. In this case, the query weight data, key weight data, and value weight data can be split according to the header weight data required by each first computing unit 211, without limitation. In some embodiments, such as under the Group Query Attention (GQA) mechanism, the query (Q) header can be divided into several groups, and all Q headers in each group share the same set of key (K) headers and value (V) headers. That is, the key weight data and value weight data can be split according to the number of K or V headers, and Q is combined with the corresponding K and V. In this way, for the same first computing unit 211, the number of query sub-weight data (matrices or vectors) obtained by splitting may not be equal to the number of key sub-weight data (matrices or vectors) or value sub-weight data (matrices or vectors) obtained by splitting. It is understood that other attention computing mechanisms can also be used to split the query weight data, key weight data, and value weight data according to the headers associated with each first computing unit 211, without limitation.

[0058] Each first computation unit 211 can calculate one or more sets of head branch data corresponding to the one or more heads based on the input hidden state of the Transformer layer and one or more sets of head weight data associated with the first computation unit 211. Each set of head branch data corresponding to each head may include the query sub-data Q corresponding to that head.h Key data K h Sum of subdata V h .

[0059] Specifically, in some embodiments, calculating a set of head branch data corresponding to a head based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head may include: calculating the query sub-weight data W corresponding to the head based on the input hidden state and the set of head weight data corresponding to the head. Qh Calculate the query sub-data Q corresponding to this head. h For example, this can be achieved by calculating the input hidden state and the query subweight data W corresponding to that head. Qh The result of multiplication yields the query sub-data Q corresponding to that head. h .

[0060] Furthermore, in some embodiments, if the cache space of the cache unit 220 corresponding to the first calculation unit 211 (e.g., the first cache unit 221 related to the calculation of query sub-data) is very limited, then the input hidden state and the query sub-weight data W corresponding to the head can be further processed according to the feature dimension. Qh This is done by splitting the computation to reduce the cache space required during the calculation process. Specifically, such as... Figure 9 As shown, based on the input hidden state IN and the query sub-weight data W corresponding to that head... Qh Calculate the query sub-data Q corresponding to this head. h This can include: splitting the input hidden state IN into first sub-state data IN-1 and second sub-state data IN-2 according to the feature dimension, and separating the query sub-weight data W Qh Split into the first query sub-weight data W Qh-1 Second query sub-weight data W Qh-2 Based on the first sub-state data IN-1 and the first query sub-weight data W Qh-1 Calculate the first query subdata Q h-1 And based on the second sub-state data IN-2 and the second query sub-weight data W Qh-2 Calculate the second query subdata Q h-2 And based on the first query sub-data Q h-1 Second query sub-data Q h-2 Calculate the query subdata Q hIn other words, by further splitting the input hidden state and query sub-weight data according to feature dimensions, the dimensions of the data involved in the calculation can be reduced, thereby ensuring that the performance of the first calculation unit 211 and the cache unit 220 (e.g., the first cache unit 221) meets the requirements. Finally, the query sub-data can be obtained from the first query sub-data and the second query sub-data through reduction operations such as summation and concatenation. It is understood that in some embodiments, the input hidden state or query sub-weight data can be split into more parts to further reduce the required cache space. In addition, as mentioned above, the granularity of splitting the input hidden state or query sub-weight data in the calculation can be dynamically adjusted according to at least one of the real-time remaining capacity of the cache unit 220 and the current access bandwidth of the storage unit 230 to balance computational efficiency and resource consumption.

[0061] In some embodiments, calculating a set of head branch data corresponding to a head based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head may further include: processing the calculated query sub-data Q h The cache is stored in a first cache unit 221 corresponding to a first computation unit 211 that performs the computation, so that it can be called in subsequent computations.

[0062] In some embodiments, similar calculation methods can be used to calculate key data and value data, referring to the calculation methods described above regarding query sub-data.

[0063] However, considering that a portion of the key data and a portion of the value data may have already been calculated, in order to reduce redundant calculations, in some embodiments, the key data and value data can be calculated in the manner described below. Specifically, calculating a set of head branch data corresponding to a head based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head may include: based on the sub-hidden state corresponding to the newly added lexical unit in the input hidden state and the key weight data W corresponding to the head. Kh Calculate the new key data corresponding to the newly added word element, and calculate the key data K corresponding to the head based on the historical key data and the newly added key data. h For example, this can be achieved by calculating the sub-hidden state corresponding to the newly added word in the input hidden state and the key weight data W corresponding to that head. Kh The result of multiplication yields the new key data corresponding to the newly added word element. Then, by concatenating the historical key data and the new key data, the key data K corresponding to that header can be obtained. h The historical key data can be generated based on the sub-hidden states corresponding to the historical tokens in the input hidden state.

[0064] Furthermore, in some embodiments, if the cache space of the cache unit 220 corresponding to the first computing unit 211 (e.g., the second cache unit 222 related to the calculation of key data) is very limited, then the input hidden state and the key weight data W corresponding to the head can be further processed according to the feature dimension. Kh To reduce the cache space required during computation, the tensor is split. Splitting by feature dimension means keeping the feature dimension unchanged while splitting along other dimensions. In other words, the tensor's feature-dimension portion remains intact, while its other-dimension portions are split into multiple parts, reducing storage space and transmission bandwidth. As mentioned above, the input hidden state can be in the form of a tensor with dimensions (B, S, H). Splitting the input hidden state by feature dimension means splitting at least one of B and S, while keeping H unchanged, to create multiple tensors. Specifically, as... Figure 10 As shown, based on the sub-hidden state INp corresponding to the newly added term in the input hidden state and the key weight data W corresponding to the head, Kh Calculate the new key data K corresponding to the newly added word element. hp This can include: splitting the sub-hidden state INp corresponding to the newly added word in the input hidden state into the third sub-state data INp-3 and the fourth sub-state data INp-4 according to the feature dimension, and separating the key weight data W Kh Split into first key weight data W Kh-1 Second bond weight data W Kh-2 Based on the third sub-state data INp-3 and the first bond weight data W Kh-1 Calculate the first newly added key data K hp-1 And based on the fourth sub-state data INp-4 and the second bond weight data W Kh-2 Calculate the second newly added bond data K hp-2 And based on the first newly added key data K hp-1 Second newly added key data K hp-2 Calculate the newly added key data K hpIn other words, by further splitting the sub-hidden states corresponding to the newly added lexical units and the key weight data corresponding to the head in the input hidden state according to the feature dimension, the dimensionality of each data involved in the calculation can be reduced, thereby ensuring that the performance of the first calculation unit 211 and the cache unit 220 (e.g., the second cache unit 222) meets the requirements. Finally, the newly added key data can be obtained from the first and second newly added key data through reduction operations such as summation and concatenation. It is understood that in some embodiments, the sub-hidden states corresponding to the newly added lexical units or the key weight data corresponding to the head in the input hidden state can be split into more parts to further reduce the required cache space. In addition, as mentioned above, the granularity of splitting the sub-hidden states corresponding to the newly added lexical units or the key weight data corresponding to the head in the input hidden state during the calculation can be dynamically adjusted according to at least one of the real-time remaining capacity of the cache unit 220 and the current access bandwidth of the storage unit 230 to balance computational efficiency and resource consumption. For example, in response to the current access bandwidth being greater than or equal to a first preset bandwidth threshold, the splitting granularity can be reduced to alleviate data transmission pressure; while in response to the current access bandwidth being less than a second preset bandwidth threshold, the splitting granularity can be increased to reduce the number of split parts, wherein the first preset bandwidth threshold can be greater than or equal to the second preset bandwidth threshold.

[0065] In some embodiments, calculating a set of head branch data corresponding to a head based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head may further include: adding the calculated new key data K. hpThe cached data is stored in a second cache unit 222 corresponding to a first computation unit 211 performing computation. This second cache unit 222 can also be configured to cache at least a portion of historical key data. Further, in some embodiments, calculating a set of head branch data corresponding to a head based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head may include: in response to the remaining cache space of the second cache unit 222 being less than or equal to a first preset threshold, writing the cached data in the second cache unit 222 page by page into the storage unit 230 corresponding to the first computation unit 211 performing computation. That is, when the second cache unit 222 is about to be full or is full, key data or a portion of the key data can be written page by page into the storage unit 230 to free up space in the second cache unit 222 for computation. The first preset threshold is used to evaluate the empty / full state of the second cache unit 222, and it can be set as needed, without limitation here. Furthermore, in subsequent calculations, if at least some of the key data in storage component 230 is required, it can be read from storage component 230 with a larger granularity, thereby helping to improve the efficiency of data reading and transmission.

[0066] Similar to the calculation of key data, in some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, may include: based on the sub-hidden state corresponding to the newly added lexical unit in the input hidden state and the value sub-weight data W corresponding to the head. Vh Calculate the new value sub-data corresponding to the newly added word element, and calculate the value sub-data V corresponding to the head based on the historical value sub-data and the new value sub-data. h For example, this can be achieved by calculating the sub-hidden state corresponding to the newly added word in the input hidden state and the sub-weight data W corresponding to the value of the head. Vh The result of multiplication yields the new value sub-data corresponding to the newly added word element. Then, by concatenating the historical value sub-data and the new value sub-data, the value data V corresponding to that head can be obtained. h The historical value sub-data can be generated based on the sub-hidden states corresponding to the historical tokens in the input hidden state.

[0067] Furthermore, in some embodiments, if the cache space of the cache unit 220 corresponding to the first calculation unit 211 (e.g., the third cache unit 223 related to the calculation of the value data) is very limited, then the input hidden state and the value weight data W corresponding to the head can be further processed according to the feature dimension. Vh This is done by splitting the computation to reduce the cache space required during the calculation process. Specifically, such as... Figure 11As shown, based on the sub-hidden state INp corresponding to the newly added word in the input hidden state and the sub-weight data W corresponding to the head value, Vh Calculate the new value sub-data V corresponding to the newly added word element. hp This can include: splitting the sub-hidden state INp corresponding to the newly added word in the input hidden state into the fifth sub-state data INp-5 and the sixth sub-state data INp-6 according to the feature dimension, and dividing the value sub-weight data W Vh Split into first-value sub-weight data W Vh-1 Second-valued sub-weight data W Vh-2 Based on the fifth sub-state data INp-5 and the first sub-weight data W Vh-1 Calculate the first newly added sub-data V hp-1 And based on the sixth sub-state data INp-6 and the second sub-weight data W Vh-2 Calculate the second newly added sub-data V hp-2 And based on the first newly added sub-data V hp-1 Second newly added sub-data V hp-2 Calculate the new value sub-data V hp In other words, by further splitting the sub-hidden states corresponding to the newly added tokens and the value sub-weights corresponding to the head in the input hidden state according to the feature dimension, the dimensionality of each data involved in the calculation can be reduced, thereby ensuring that the performance of the first calculation unit 211 and the cache unit 220 (e.g., the third cache unit 223) meets the requirements. Finally, the newly added value sub-data can be obtained from the first and second newly added value sub-data through reduction operations such as summation and concatenation. It is understood that in some embodiments, the sub-hidden states corresponding to the newly added tokens or the value sub-weights corresponding to the head in the input hidden state can be split into more parts to further reduce the required cache space. In addition, as mentioned above, the granularity of splitting the sub-hidden states corresponding to the newly added tokens or the value sub-weights corresponding to the head in the input hidden state during the calculation can be dynamically adjusted according to at least one of the real-time remaining capacity of the cache unit 220 and the current access bandwidth of the storage unit 230 to balance computational efficiency and resource consumption.

[0068] In some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, may further include: applying the calculated new value-added sub-data V... hpThe data is cached in a third cache unit 223 corresponding to a first computation unit 211 performing computation. This third cache unit 223 can also be configured to cache at least a portion of historical value sub-data. Further, in some embodiments, calculating a set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, may include: in response to the remaining cache space of the third cache unit 223 being less than or equal to a second preset threshold, writing the cached data in the third cache unit 223 page by page into a storage unit 230 corresponding to the first computation unit 221 performing computation. That is, when the third cache unit 223 is about to be full or is full, the value sub-data or a portion of the value sub-data can be written page by page into the storage unit 230 to free up space in the third cache unit 223 for computation. The second preset threshold is used to evaluate the empty / full state of the third cache unit 223, and it may be equal to or different from the first preset threshold described above. Furthermore, in subsequent calculations, if at least some of the sub-data in storage component 230 is required, it can be read from storage component 230 with a larger granularity, thereby helping to improve the efficiency of data reading and transmission.

[0069] In some embodiments, as mentioned in the above description of the inference computing device, at least one of the first cache unit 221, the second cache unit 222, and the third cache unit 223 may be integrated in the same chip as the corresponding first computing unit 211. For example, the first cache unit 221, the second cache unit 222, or the third cache unit 223 may be on-chip SRAM relative to the first computing unit 211.

[0070] Furthermore, in some embodiments, depending on the specific operational rules of the attention mechanism, the computation performed by the first computation unit 211 may involve transposing data (e.g., tensors, matrices, or vectors). At least one of the first cache unit 221, the second cache unit 222, and the third cache unit 223 can also be configured to perform the transpose operation in the computation to simplify the computation. Additionally, as mentioned above, the granularity of data transposition during computation can be configured to be dynamically adjusted based on at least one of the real-time remaining capacity of the cache unit 220 and the current access bandwidth of the storage unit 230, to balance computational efficiency and resource consumption.

[0071] In some embodiments, as mentioned in the above description of the inference apparatus, the storage unit 230 may be disposed separately from the corresponding first computing unit 211. For example, the storage unit 230 may be an off-chip DRAM relative to the first computing unit 211. Furthermore, in some embodiments, all query sub-weight data, all key sub-weight data, all value sub-weight data, historical key data in all key data, and historical value sub-data in all value data corresponding to one or more heads associated with each first computing unit 211 may be stored in the same storage unit 230 corresponding to that first computing unit 211. This allows for reading directly from the storage unit 230 or via the cache unit 220 during the computation process at a larger granularity (e.g., by page), thereby improving the efficiency of data reading and transmission and effectively avoiding coupling between data of different first computing units 211 or data of different heads.

[0072] In some embodiments, calculating an attention head output corresponding to a head based on a set of head branch data may include calculating the attention head output based on the head branch data using a scaled dot product. Specifically, the attention head output A h The following formula can be satisfied: , Among them, Q h For the query sub-data corresponding to this header, K h For the key data corresponding to this header, V is the transpose of the key data corresponding to this header. h For the value sub-data corresponding to this head, d k For the head dimension, it can satisfy d k =H / N2, where H is the hidden dimension of the input hidden state and N2 is the total number of heads.

[0073] It is understood that in some other embodiments, other methods may be used to calculate the attention head output corresponding to the head, and are not limited to the scaled dot product method described above.

[0074] In some embodiments, in order to improve computational efficiency, a plurality of first computing units 211 may be configured to run in parallel with each other, that is, the related computations of at least some of the plurality of heads may be performed in parallel without being coupled or interfering with each other.

[0075] In the reasoning and calculation method disclosed herein, such as Figure 5 As shown, after each of the first computing units 211 calculates one or more corresponding attention head outputs, the second computing unit 212 can output A to the multiple attention heads from the multiple first computing units 211. hA reduction process is performed to generate the multi-head attention output A of the multi-head attention sublayer. Various reduction methods can be used to derive the multi-head attention output of the multi-head attention sublayer; for example, the second computation unit 212 can perform a reduction on the attention heads A from multiple first computation units 211. h Perform at least one of the following calculations: summation, concatenation, weighted averaging, and maximum aggregation, to produce the multi-head attention output A of the multi-head attention sublayer.

[0076] In a specific example, an inference computing device based on a Field Programmable Gate Array (FPGA) is used as the application platform. This inference computing device is equipped with four first computing units 211, a cache unit 220 (a total of 512KB of on-chip distributed SRAM, of which each first computing unit 211 corresponds to an SRAM capacity of 128KB), and four sets of storage units 230 (off-chip DDR4 DRAM). The technical solution of this disclosure is specifically implemented using this inference computing device for a large model inference scenario with an 8-head attention mechanism (where the sequence length S=1024 and the hidden dimension H=512 of the input hidden state). The detailed operation is as follows: First, the hardware resources are initialized. Specifically, the four first computing units 211 of the FPGA can be bound to the eight attention heads in a 2-head / first computing unit configuration. Each first computing unit 211 is allocated 128KB of on-chip distributed SRAM as cache space for the two attention heads within that first computing unit 211. Each set of DDR4 DRAM is configured to store the KV cache and attention weights corresponding to the two attention heads.

[0077] Then, the input hidden state is broadcast and split into head dimensions. Specifically, the input hidden state output by the previous layer of the large model can be globally broadcast once at the head of the attention layer and passed to the four first computing units 211; the Q, K, and V matrices (dimensions of 512×1024) are split into eight independent head branches (each head has a dimension of 64×1024) according to the head dimension, and the weights and KV caches of the eight head branches are allocated to the four first computing units 211 respectively, so that the hardware resources of each head branch are completely isolated and there is no cross-head data flow.

[0078] Next, linear partitioning and SRAM cache adaptation are performed. Specifically, since the SRAM cache space corresponding to a single first computing unit 211 is only 128KB, it is not enough to cache the complete head branch matrix with a dimension of 64×1024. Therefore, the Q, K, and V matrices in each head branch can be linearly partitioned according to the feature dimension (partitioned into two consecutive data blocks of 64×512). After partitioning, the size of a single data block is 32KB, which is much smaller than the 128KB SRAM cache capacity, thus meeting the caching requirements. Moreover, there is no cross-block sequence S dimension dependency after partitioning.

[0079] In some cases, fine-grained transposition can be supported by SRAM. Specifically, when reading each 64×512 linearly divided block from DDR4 DRAM, the data can be first written to the SRAM cache of the corresponding first computing unit 211, and the fine-grained transposition operation from 64×512 to 512×64 can be completed in SRAM. After the transposition is completed, the data can be directly read from SRAM for calculation. The transposition granularity can be adjusted to 64×64 according to the remaining capacity of SRAM, so that a single transposition only occupies 4KB of SRAM space. This ensures transposition efficiency while controlling SRAM resource consumption.

[0080] Next, independent head computation and naive softmax computation are performed. Specifically, the eight head branches within the four first computation units 211 can perform attention computation in parallel. Since the sequence S dimension (1024) is not split in any way, the Q within each head branch is... h K h T After the similarity calculation is completed, the naive softmax calculation can be performed directly without splitting and merging statistics, which simplifies the calculation logic.

[0081] Finally, cross-head reduction and result output are performed. Specifically, after the softmax and V weighted calculations are completed in all eight head branches, eight attention head outputs (or partial hidden states) with dimensions of 64×1024 are obtained. At the very end of the attention layer, the eight attention head outputs can be passed to the second computation unit 212 to perform reduction operations such as concatenation, thereby obtaining the final multi-head attention output or output hidden state with a dimension of 512×1024. This multi-head attention output or output hidden state can be written back to DDR4 DRAM, thus completing the calculation of the entire attention layer.

[0082] The inference computation method disclosed herein is used for attention-based models. Specifically, such models can be applied to various fields such as image, video, speech, and audio processing. By converting data such as images, videos, speech, or audio into tensor form and deploying the relevant operations of the Transformer layer based on the head-parallel strategy described above, it helps to improve the running speed. For example, attention-based models may include Large Language Models (LLM) for intelligent customer service and office assistance, Vision Language Models (VLM) for scene recognition and environment awareness, Diffusion Transformers (DIT) for image and video generation, and World Action Models (WAM) and Vision-Language-Action (VLA) models, etc.

[0083] Furthermore, this disclosure also proposes a computer-readable storage medium, such as a non-transitory computer-readable storage medium, on which instructions can be stored, which, when executed by a processor, can implement the operation of the inference operation method described above.

[0084] The computer-readable storage medium in the embodiments of this disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. It should be noted that the computer-readable storage medium described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0085] In addition, this disclosure also proposes a computer program product that may include instructions that, when executed by a processor, can implement the inference operation method described above.

[0086] Instructions can be any set of instructions that will be executed directly by one or more processors, such as machine code, or any set of instructions that will be executed indirectly, such as a script. The terms “instruction,” “application,” “process,” “step,” and “program” used herein are interchangeable. Instructions can be stored in object code format for direct processing by one or more processors, or stored in any other computer language, including scripts or sets of independent source code modules that are interpreted on demand or compiled ahead of time. Instructions can include instructions that cause one or more processors to act as the various neural networks described herein. The function, methods, and routines of instructions are explained in more detail in other parts of this document.

[0087] In the technical solution disclosed herein, by decomposing the various data and computation processes involved in multi-head attention computation according to the head dimension, the shortcomings of solutions such as Flash Attention can be overcome, thereby achieving the following beneficial effects: First, it solves the problem of high cache resource requirements. Specifically, through linear partitioning along the head dimension and potentially linear partitioning along the feature dimension, it splits data such as large matrices into small linear data blocks (rather than tile data blocks). Each data block occupies a small cache, thus adapting to low-capacity on-chip SRAM scenarios and significantly reducing SRAM resource costs. By using a linear partitioning mode instead of a block partitioning mode, it reduces the dependence on cache space. More importantly, it reduces the dependence on the stride access method and reduces the overhead of row and page switching in storage components (such as DRAM), significantly improving DRAM-side access efficiency.

[0088] Second, it eliminates communication bottlenecks across computing units. Specifically, by binding attention heads to computing units, the Q, K, and V data of each head are read, written, and computed entirely within a single computing unit, without the need for operations such as KV aggregation, which greatly reduces the bandwidth requirements for cross-unit communication.

[0089] Third, it simplifies the softmax calculation logic. Specifically, since the sequence S dimension is not split, the naive softmax calculation is used directly, thus eliminating the steps of splitting, passing, and merging in block softmax, which can significantly reduce the complexity of the calculation logic.

[0090] Fourth, it offers high data read / write efficiency for storage components. Specifically, due to the linear partitioning along the header dimension, entire pages of data can be easily aggregated for access to storage components (such as DRAM), thereby reducing the proportion of activation and precharging in the overall access process and improving DRAM access efficiency. Simultaneously, fine-grained transposition can be performed using cache components (such as SRAM), avoiding the inefficiency of direct transposition in DRAM, and the resource consumption of SRAM is controllable due to the linear partitioning.

[0091] Fifth, it eliminates the communication overhead of head-parallel processing. Specifically, it achieves complete partitioning and resource isolation of the head dimension, performing the broadcast of the input hidden state only once at the head of the attention layer and the reduction of the partial hidden state only once at the tail, while there is no cross-head data flow or synchronization in between. Compared with traditional head-parallel processing schemes, the cross-head communication overhead is significantly reduced.

[0092] In addition, the technical solution disclosed herein overcomes the shortcomings of solutions such as Flash Attention, and can also produce the following additional positive effects: First, it has strong hardware adaptability. Specifically, the technical solution disclosed herein does not depend on a specific hardware architecture and can be directly adapted to various AI hardware systems such as FPGAs, application-specific integrated circuits (ASICs), AI accelerators, and graphics processing units (GPUs). It is especially suitable for low-computing-power hardware scenarios with limited SRAM resources, such as industrial and consumer-grade applications.

[0093] Second, the model has high compatibility. Specifically, the technical solution disclosed herein can support any number of attention head splits, and the reduction operation is configurable. It can be adapted to the attention layer design of mainstream large language models (such as LLaMA, GPT, GLM, etc.) and computer vision models (such as ViT, etc.) without modifying the model structure.

[0094] Third, improved inference throughput. Specifically, by eliminating communication bottlenecks across computing components and improving computing power utilization, the token / s inference throughput of large models can be significantly improved compared to solutions such as Flash Attention, under the same hardware configuration.

[0095] Fourth, hardware power consumption is reduced. Specifically, the technical solution disclosed herein features high data flow cohesion and no communication power consumption across computing components, which significantly reduces the computational power consumption of the entire attention layer compared to existing solutions.

[0096] Fifth, the engineering implementation is simple. Specifically, the computational logic of linear partitioning and naive softmax is simple, requiring no complex block scheduling and intermediate data management, which reduces the development difficulty of hardware firmware and software drivers and shortens the engineering implementation cycle.

[0097] The terms “left,” “right,” “front,” “back,” “top,” “bottom,” “upper,” “lower,” “high,” “lower,” etc., used in the specification and claims, if present, are for descriptive purposes and not necessarily for describing unchanging relative positions. It should be understood that such terms are interchangeable where appropriate, enabling embodiments of this disclosure described herein to operate, for example, in orientations different from those shown or otherwise described herein. For example, when the device in the drawings is reversed, a feature previously described as “above” other features may now be described as “below” other features. The device may also be oriented in other ways (rotated 90 degrees or in other orientations), in which case the relative spatial relationships will be interpreted accordingly.

[0098] In the specification and claims, when an element is described as being "on top of," "attached to," "connected to," "coupled to," or "in contact with" another element, the element may be directly located on top of, directly attached to, directly connected to, directly coupled to, or directly in contact with the other element, or one or more intermediate elements may be present. Conversely, when an element is described as being "directly" located on top of, directly attached to, directly connected to, directly coupled to, or directly in contact with another element, no intermediate elements are present. In the specification and claims, when a feature is arranged "adjacent" to another feature, it may mean that a feature has a portion overlapping with the adjacent feature or a portion located above or below the adjacent feature.

[0099] As used herein, the term “exemplary” means “serving as an example, instance, or illustration” and not as a “model” to be precisely copied. Any implementation described herein by example is not necessarily to be construed as preferred or advantageous over other implementations. Moreover, this disclosure is not limited to any theory expressed or implied as given in the field of art, background art, summary of invention, or detailed description.

[0100] As used herein, the term "substantially" means any minor variation resulting from design or manufacturing defects, device or component tolerances, environmental influences, and / or other factors. The term "substantially" also allows for differences from the perfect or ideal situation due to parasitic effects, noise, and other practical considerations that may exist in the actual implementation.

[0101] Furthermore, terms such as “first,” “second,” etc., may be used in this document for reference purposes only and are not intended to be limiting. For example, unless the context clearly indicates otherwise, the words “first,” “second,” and other such numerical terms relating to structures or elements do not imply order or sequence.

[0102] It should also be understood that when the term “includes / contains” is used herein, it indicates the presence of the indicated feature, whole, step, operation, sub-part and / or component, but does not preclude the presence or addition of one or more other features, wholes, steps, operations, sub-parts and / or components and / or combinations thereof.

[0103] Additionally, when used in this disclosure, the terms “here,” “above,” “below,” “below,” “in the preceding text,” and similar terms should refer to the entirety of this disclosure and not any particular part thereof. Furthermore, unless expressly stated otherwise or otherwise understood in the context in which they are used, conditional language used herein, such as “may,” “possibly,” “for example,” “like,” etc., is generally intended to express that certain embodiments include, while other embodiments do not, certain features, elements, and / or states. Therefore, such conditional language is not generally intended to imply that one or more embodiments require features, elements, and / or states in any way, or whether such features, elements, and / or states are included or performed in any particular embodiment.

[0104] In this disclosure, the term "provide" is used broadly to cover all ways of obtaining an object, and therefore "providing an object" includes, but is not limited to, "purchasing," "preparing / manufacturing," "arranging / setting up," "installing / assembling," and / or "ordering" an object. Furthermore, in this disclosure, the terms "circuit," "sub-component," and "module" are used interchangeably.

[0105] As used herein, the term “and / or” includes any and all combinations of one or more of the listed items in association. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise.

[0106] Those skilled in the art will recognize that the boundaries between the above operations are merely illustrative. Multiple operations may be combined into a single operation, a single operation may be distributed among additional operations, and operations may be performed with at least partial overlap in time. Moreover, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be changed in various other embodiments. However, other modifications, variations, and substitutions are equally possible. Aspects and elements of all the embodiments disclosed above may be combined in any way and / or in combination with aspects or elements of other embodiments to provide multiple additional embodiments. Therefore, this specification and the accompanying drawings should be considered illustrative rather than restrictive. In fact, the novel devices, methods, and systems described herein may be embodied in various other forms. Furthermore, various omissions, substitutions, and changes may be made to the form of the methods and systems described herein without departing from the spirit of this disclosure. For example, although blocks are presented in a given arrangement, alternative embodiments may perform similar functions with different components and / or circuit topologies, and some blocks may be deleted, moved, added, subdivided, combined, and / or modified. Each of these blocks may be implemented in various different ways.

[0107] The various embodiments of this disclosure can be described in a progressive manner, with references made to similar or identical parts between embodiments. Each embodiment focuses on describing the differences from other embodiments. In this disclosure, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this disclosure. In this disclosure, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

[0108] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. The various embodiments disclosed herein can be combined in any way without departing from the spirit and scope of this disclosure. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A reasoning computation method for attention-based models, wherein, The attention-based model includes a Transformer layer, which in turn includes a multi-head attention sub-layer. The inference operation method includes: Each of the plurality of first computing units performs the following operations: Obtain the input hidden state of the Transformer layer and one or more sets of head weight data corresponding one-to-one with one or more heads associated with the first computation unit. Each set of head weight data corresponding to each head includes query sub-weight data corresponding to that head in the query weight data, key sub-weight data corresponding to that head in the key weight data, and value sub-weight data corresponding to that head in the value weight data. Based on the input hidden state of the Transformer layer and the one or more sets of head weight data, calculate one or more sets of head branch data corresponding one-to-one with the one or more heads, wherein each set of head branch data corresponding to each head includes query sub-data, key data, and value sub-data corresponding to that head, and Based on the set or more sets of head branch data, calculate one or more attention head outputs corresponding one-to-one with the one or more heads; and The second computing unit reduces the multiple attention head outputs from the plurality of first computing units to generate the multi-head attention output of the multi-head attention sublayer.

2. The reasoning operation method according to claim 1, wherein, The input hidden state of the Transformer layer is broadcast to each of the plurality of first computing units.

3. The reasoning operation method according to claim 1, wherein, The calculation of each group of head branch data corresponding to each head is performed independently; and / or The calculation of the output of each attention head corresponding to each head is performed independently.

4. The reasoning operation method according to claim 3, wherein, Based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, the calculation of a set of head branch data corresponding to that head includes: Based on the input hidden state and the query sub-weight data corresponding to the head, calculate the query sub-data corresponding to the head; or Based on the sub-hidden state corresponding to the newly added lexical unit in the input hidden state and the key weight data corresponding to the head, calculate the new key data corresponding to the newly added lexical unit; and based on the historical key data and the newly added key data, calculate the key data corresponding to the head, wherein the historical key data is generated based on the sub-hidden state corresponding to the historical lexical unit in the input hidden state; or Based on the sub-hidden state corresponding to the newly added word in the input hidden state and the value sub-weight data corresponding to the head, the newly added value sub-data corresponding to the newly added word is calculated, and based on the historical value sub-data and the newly added value sub-data, the value data corresponding to the head is calculated, wherein the historical value sub-data is generated based on the sub-hidden state corresponding to the historical word in the input hidden state.

5. The reasoning operation method according to claim 4, wherein, Calculating the query sub-data corresponding to the head based on the input hidden state and the query sub-weight data corresponding to the head includes: splitting the input hidden state into first sub-state data and second sub-state data according to the feature dimension, and splitting the query sub-weight data into first query sub-weight data and second query sub-weight data; calculating the first query sub-data based on the first sub-state data and the first query sub-weight data; calculating the second query sub-data based on the second sub-state data and the second query sub-weight data; and calculating the query sub-data based on the first query sub-data and the second query sub-weight data; or Calculating the new key data corresponding to the new word element based on the sub-hidden state corresponding to the new word element in the input hidden state and the key weight data corresponding to the head includes: splitting the sub-hidden state corresponding to the new word element in the input hidden state into third sub-state data and fourth sub-state data according to the feature dimension, and splitting the key weight data into first key weight data and second key weight data; calculating the first new key data based on the third sub-state data and the first key weight data; calculating the second new key data based on the fourth sub-state data and the second key weight data; and calculating the new key data based on the first new key data and the second new key data; or Calculating the new value sub-data corresponding to the new word in the input hidden state based on the sub-hidden state corresponding to the new word and the value sub-weight data corresponding to the head includes: splitting the sub-hidden state corresponding to the new word in the input hidden state into fifth sub-state data and sixth sub-state data according to the feature dimension, and splitting the value sub-weight data into first value sub-weight data and second value sub-weight data; calculating the first new value sub-data based on the fifth sub-state data and the first value sub-weight data; calculating the second new value sub-data based on the sixth sub-state data and the second value sub-weight data; and calculating the new value sub-data based on the first new value sub-data and the second new value sub-data.

6. The reasoning operation method according to claim 4, wherein, Calculating the set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, further includes: The calculated query sub-data is cached in a first cache unit corresponding to a first calculation unit that performs the calculation; or The newly calculated key data is cached in a second cache unit corresponding to a first calculation unit that performs the calculation, wherein the second cache unit is further configured to cache at least a portion of the historical key data; or The newly calculated value-added sub-data is cached in a third cache unit corresponding to a first calculation unit that performs the calculation, wherein the third cache unit is also configured to cache at least a portion of the historical value sub-data.

7. The reasoning operation method according to claim 6, wherein, At least one of the first cache component, the second cache component, and the third cache component is integrated into the same chip as the corresponding first computing component; or At least one of the first cache component, the second cache component, and the third cache component is further configured to perform a transpose operation in the computation.

8. The reasoning operation method according to claim 6, wherein, Calculating the set of head branch data corresponding to a head, based on the input hidden state of the Transformer layer and a set of head weight data corresponding to a head, further includes: In response to the remaining cache space of the second cache component being less than or equal to a first preset threshold, the data cached in the second cache component is written page by page to the storage component corresponding to a first computing component performing the computation; or In response to the remaining cache space of the third cache component being less than or equal to a second preset threshold, the data cached in the third cache component is written page by page into the storage component corresponding to a first computing component that performs the calculation.

9. The reasoning operation method according to claim 8, wherein, The storage component and the corresponding first computing component are configured separately from each other; or All query sub-weight data, all key sub-weight data, all value sub-weight data, historical key data in all key data, and historical value sub-data in all value data corresponding to one or more heads associated with each first calculation unit are stored in the same storage unit corresponding to that first calculation unit.

10. The reasoning operation method according to claim 1, wherein, The process of reducing the multiple attention head outputs from the plurality of first computing units by the second computing unit to generate the multi-head attention output of the multi-head attention sub-layer includes: The second computing unit performs at least one of the following calculations on all attention head outputs from the plurality of first computing units: summation, concatenation, weighted averaging, and maximum aggregation, to generate the multi-head attention output of the multi-head attention sublayer.

11. The reasoning operation method according to claim 10, wherein, The plurality of first computing units and the second computing units are respectively configured; or The second computing unit is implemented by one of the plurality of first computing units; or The plurality of first computing units are configured to operate in parallel with each other.

12. An inference computing device for attention-based models, wherein, The attention-based model includes a Transformer layer, which in turn includes a multi-head attention sublayer. The inference processing unit includes multiple first computational units and second computational units, wherein each first computational unit is configured to perform the following operations: Obtain the input hidden state of the Transformer layer and one or more sets of head weight data corresponding one-to-one with one or more heads associated with the first computation unit. Each set of head weight data corresponding to each head includes query sub-weight data corresponding to that head in the query weight data, key sub-weight data corresponding to that head in the key weight data, and value sub-weight data corresponding to that head in the value weight data. Based on the input hidden state of the Transformer layer and the one or more sets of head weight data, calculate one or more sets of head branch data corresponding one-to-one with the one or more heads, wherein each set of head branch data corresponding to each head includes query sub-data, key data, and value sub-data corresponding to that head, and Based on the one or more sets of head branch data, calculate one or more attention head outputs that correspond one-to-one with the one or more heads; The second computing unit is configured to reduce the multiple attention head outputs from the plurality of first computing units to generate the multi-head attention output of the multi-head attention sublayer; The plurality of first computing components and the second computing component are respectively configured, or the second computing component is implemented by one of the plurality of first computing components.

13. The reasoning apparatus according to claim 12, further comprising: One or more cache units corresponding to each first computing unit, wherein each cache unit is configured to cache at least a portion of the data related to the computation performed by the first computing unit; One or more storage units corresponding to each first computing unit, wherein each storage unit is configured to store at least a portion of the data related to the computation performed by the first computing unit.

14. The reasoning apparatus according to claim 13, wherein, A first computing unit and one or more cache units corresponding to the first computing unit are integrated into the same chip; or A first computing unit and one or more storage units corresponding to the first computing unit are disposed separately from each other; or The cache components include static random access memory (SRAM); or Storage components include dynamic random access memory (DRAM); or The granularity at which data is split or transposed during computation is configured to dynamically adjust based on at least one of the real-time remaining capacity of the cache component and the current access bandwidth of the storage component.

15. A computer-readable storage medium storing instructions that, when executed by a processor, implement the reasoning operation method according to any one of claims 1 to 11.

16. A computer program product comprising instructions that, when executed by a processor, implement the reasoning operation method according to any one of claims 1 to 11.