Multi-head attention calculation adaptive allocation method and device of Transform model

By using preset power allocation method and preset rules to adaptively allocate attention heads in the multi-head attention calculation of Transformer model, the problem of insufficient use of computing resources on the multi-core processor is solved, and the effect of efficient utilization of computing resources and improving computing speed is achieved.

CN120045308APending Publication Date: 2025-05-27太初(无锡)电子科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411861966.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When performing multi-head attention calculation on multi-core processors, conventional solutions lead to insufficient use of computing resources and waste of chip computing power because the number of attention heads and the number of computing cores do not match.

Method used

By obtaining the number of attention heads, batch processing numbers, and the initial number of core groups and card numbers of the target processor, the preset power distribution method is used to split the number of attention heads of a single core group of the target processor to obtain multiple attention head batches of a single core group, and the preset rules are used to allocate these batches to the calculation core in turn to perform calculations.

Benefits of technology

It realizes efficient utilization of computing resources, ensures that the Transformer models with different configurations can exert their computing capabilities as much as possible on the target processor, improves the speed of multi-headed attention calculations, and maximizes the use of computing core resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045308A_ABST
    Figure CN120045308A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of processors, and discloses a multi-head attention calculation adaptive allocation method and device for a Transform model, and adaptive allocation is carried out according to obtained key parameters of the model and structural parameters of a target processor. Furthermore, the number of attention heads of a single core group is split by adopting a preset two-power allocation method, and the parallel computing core characteristics of a processor and the specific condition of a model are fully considered, so that the problem of idle computing resources caused by mismatching of the number of attention heads and the number of computing cores in a conventional allocation mode is avoided; the efficient utilization of computing resources is realized, so that the models with different configurations can exert the computing power as much as possible on the target processor. Furthermore, through self-adaptive distribution, calculation of multiple attention heads can be reasonably distributed to each calculation core for parallel execution, the parallel calculation advantage of the processor is fully played, and then the speed of model multi-head attention calculation is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of processors, and in particular to an adaptive allocation method and device for multi-head attention calculation of a Transformer model. Background Art

[0002] The Transformer model structure is currently widely used in the field of natural language processing, and the model effect is significantly better than other model methods proposed before. Multi-Head Attention is a core mechanism introduced in the Transformer model. By dividing the hidden space (Hidden Size) of the model into multiple attention heads, parallel computing of the attention mechanism can be achieved, thereby accelerating the model calculation on parallel hardware such as multi-core processors, many-core processors, and GPUs.

[0003] Currently, when calculating multi-head attention on a many-core processor, the conventional solution is to sequentially allocate the attention heads to be calculated to the computing cores, with each attention head corresponding to one computing core. For this method, since the number of computing cores on the processor hardware is fixed, while the number of attention heads included in different models is not fixed, and the total number of attention heads to be processed is proportional to the batch number of problems, the attention heads often cannot be evenly allocated to the computing cores, resulting in insufficient use of computing resources and wasting the computing power of the chip. Summary of the Invention

[0004] In view of this, the present invention provides an adaptive allocation method and device for multi-head attention calculation of a Transformer model to solve the problem in the prior art that when calculating multi-head attention on a many-core processor, the computing resources are not fully utilized and the computing power of the chip is wasted due to evenly allocating the attention to the computing cores.

[0005] In a first aspect, the present invention provides an adaptive allocation method for multi-head attention calculation of a Transformer model, which is used for a target processor including multiple parallel computing cores; the method includes:

[0006] Obtain the number of attention heads, batch processing number of the Transformer model, and the initial number of core groups and the number of cards of the target processor; based on the number of attention heads, batch processing number, initial number of core groups, and the number of cards, use a preset power-of-two allocation method to split the number of attention heads of a single core group of the target processor to obtain multiple attention head batches of a single core group; use a preset rule to sequentially allocate the multiple attention head batches to multiple computing cores of a single core group and perform attention calculation to obtain the multi-head attention calculation result of the Transformer model.

[0007] The multi-head attention calculation adaptive allocation method for the Transformer model provided by the present invention can perform adaptive allocation according to the actual situation by obtaining the key parameters of the Transformer model (the number of attention heads, the number of batches) and the structural parameters of the target processor itself (the initial number of core groups and the number of cards). Further, the preset power-of-two allocation method is used to split the number of attention heads in a single core group, fully considering the characteristics of the parallel computing cores of the processor and the specific situation of the model, avoiding the problem of idle computing resources caused by the mismatch between the number of attention heads and the number of computing cores in the conventional allocation method, realizing the efficient utilization of computing resources, and enabling Transformer models with different configurations to exert their computing capabilities as much as possible on the target processor. Further, through adaptive allocation, the calculations of multiple attention heads can be reasonably allocated to each computing core for parallel execution, giving full play to the parallel computing advantage of the processor, and thus effectively improving the speed of multi-head attention calculation of the Transformer model. Therefore, by implementing the present invention, when performing the multi-head attention calculation of the Transformer model, in the case where the number of model attention heads, the number of batches, and the number of execution cards are variable, it can adaptively support the multi-head attention calculation under various configurations, has excellent generalization, can maximize the use of computing core resources, and approach the peak performance of the processor. And it can be applied without relying on additional technologies such as paged attention, and has a wider scope of application.

[0008] In an optional implementation manner, based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, the preset power-of-two allocation method is used to split the number of attention heads in a single core group of the target processor to obtain multiple attention head batches for a single core group, including:

[0009] Based on the initial number of core groups and the number of cards, calculate the target number of core groups of the target processor; based on the number of attention heads and the number of batches, calculate the global number of attention heads; based on the target number of core groups and the global number of attention heads, use the preset power-of-two allocation method to split the number of attention heads in a single core group of the target processor to obtain multiple attention head batches for a single core group.

[0010] The multi-head attention calculation adaptive allocation method for the Transformer model provided by the present invention first accurately calculates the target number of core groups of the target processor based on the initial number of core groups and the number of cards, clearly determining the number of basic structural units of the processor that can be used for allocating attention heads. Then, the global number of attention heads is calculated based on the number of attention heads and the batch size, overall grasping the scale of attention heads involved in the current calculation task. Further, the number of attention heads for a single core group is split using a preset power-of-two allocation method, ensuring that the multiple attention head batches obtained by the split can better adapt to the computing cores of each core group, fully considering the characteristics of the parallel computing cores of the processor and the specific situation of the model, avoiding the problem of idle computing resources caused by the mismatch between the number of attention heads and the number of computing cores in the conventional allocation method, and achieving efficient utilization of computing resources.

[0011] In an alternative embodiment, based on the target number of core groups and the global number of attention heads, the number of attention heads for a single core group of the target processor is split using a preset power-of-two allocation method to obtain multiple attention head batches for a single core group, including:

[0012] The attention heads for a single core group of the target processor are evenly allocated based on the target number of core groups and the global number of attention heads to obtain the number of attention heads for a single core group; the number of attention heads for a single core group is split using a preset power-of-two allocation method to obtain multiple attention head batches for a single core group.

[0013] The multi-head attention calculation adaptive allocation method for the Transformer model provided by the present invention first evenly allocates the attention heads for a single core group of the target processor based on the target number of core groups and the global number of attention heads, ensuring that each core group is in a relatively fair state during the initial allocation of attention heads, avoiding the situation where some core groups undertake too many or too few attention head calculation tasks, and enabling the overall computational load to be relatively evenly distributed at the level of each core group. Further, the number of attention heads for a single core group after uniform allocation is split using a preset power-of-two allocation method, which can flexibly divide the attention heads into batches of different scales according to the number of computing cores, fully utilize the computing power of each computing core, minimize the idle computing cores to the greatest extent regardless of the change in the number of attention heads, optimize the utilization efficiency of computing resources within a single core group, and improve the contribution of each core group to multi-head attention calculation.

[0014] In an alternative embodiment, multiple attention head batches are sequentially allocated to multiple computing cores of a single core group using a preset rule and attention calculation is performed to obtain the multi-head attention calculation result of the Transformer model, including:

[0015] Allocate multiple attention head batches to multiple computing cores of a single core group in sequence according to preset rules and perform attention calculations to obtain the attention calculation results of each computing core; determine the multi-head attention calculation results of the Transformer model based on the attention calculation results of each computing core.

[0016] The method for adaptively allocating multi-head attention calculations of the Transformer model provided by the present invention allocates multiple attention head batches to multiple computing cores of a single core group in sequence according to preset rules and performs attention calculations, enabling each computing core to participate in the calculation tasks of the corresponding attention heads, truly realizing the implementation of parallel computing and leveraging the parallel advantages of multi-core processors. Further, by reasonably processing the calculation results of each computing core, the multi-head attention calculation results of the Transformer model can be accurately obtained finally, ensuring that the entire calculation process can not only make full use of the efficiency of parallel computing but also guarantee the correctness of the final output results.

[0017] In an optional implementation manner, allocating multiple attention head batches to multiple computing cores of a single core group in sequence according to preset rules and performing attention calculations to obtain the attention calculation results of each computing core includes:

[0018] Allocate multiple attention head batches to multiple computing cores of a single core group in sequence according to preset rules to obtain the target attention head of each computing core; perform attention calculations based on the target attention head of each computing core to obtain the attention calculation results of each computing core.

[0019] The method for adaptively allocating multi-head attention calculations of the Transformer model provided by the present invention allocates multiple attention head batches to multiple computing cores of a single core group in sequence according to preset rules, enabling each computing core to clearly identify the target attention head it needs to process, avoiding the confusion and intersection of calculation tasks, clearly defining the scope of calculation responsibilities of each computing core, allowing the computing core to focus on the corresponding attention head calculation, and improving the accuracy and efficiency of each computing core in performing calculation tasks. Further, due to the clear task allocation, the computing core can perform targeted operations based on the allocated attention heads, and thus can efficiently obtain the attention calculation results of each computing core.

[0020] In an optional implementation manner, determining the multi-head attention calculation results of the Transformer model based on the attention calculation results of each computing core includes:

[0021] Determine the attention head calculation results of each attention head batch based on the attention calculation results of each computing core; determine the multi-head attention calculation results of the Transformer model based on the attention head calculation results of each attention head batch.

[0022] The multi - head attention calculation adaptive allocation method for the Transformer model provided by the present invention determines the calculation results of each attention head batch based on the attention calculation results of each computing core, realizes the integration from the calculation results of a single computing core to the results of each attention head batch, summarizes the results related to the same batch of attention heads scattered in each computing core, and ensures that the calculation results of each attention head batch can accurately reflect the calculation situation of the set of attention heads represented by this batch. Further, based on the calculation results of each attention head batch, the multi - head attention calculation results of the Transformer model are determined, ensuring that the finally obtained multi - head attention calculation results of the Transformer model are accurate, complete, comprehensively reflecting the calculation results of the model under the multi - head attention mechanism, making the entire calculation process logically rigorous and the results reliable.

[0023] In a second aspect, the present invention provides a multi - head attention calculation adaptive allocation device for a Transformer model, which is used for a target processor including multiple parallel computing cores; the device includes:

[0024] An acquisition module, configured to acquire the number of attention heads, the number of batch processes of the Transformer model, the initial number of core groups of the target processor, and the number of cards; a splitting module, configured to split the number of attention heads of a single core group of the target processor by using a preset power - of - two allocation method based on the number of attention heads, the number of batch processes, the initial number of core groups, and the number of cards, to obtain multiple attention head batches of a single core group; an allocation and calculation module, configured to sequentially allocate multiple attention head batches to multiple computing cores of a single core group by using a preset rule and perform attention calculation to obtain the multi - head attention calculation results of the Transformer model.

[0025] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the multi - head attention calculation adaptive allocation method for the Transformer model in the first aspect or any corresponding embodiment thereof.

[0026] In a fourth aspect, the present invention provides a computer - readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the multi - head attention calculation adaptive allocation method for the Transformer model in the first aspect or any corresponding embodiment thereof.

[0027] Fifth aspect, the present invention provides a computer program product, including computer instructions for causing a computer to execute the multi-head attention calculation adaptive allocation method of the Transformer model according to the first aspect or any corresponding embodiment thereof above. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 is a schematic flowchart of the multi-head attention calculation adaptive allocation method of the Transformer model according to an embodiment of the present invention;

[0030] Figure 2 is a schematic flowchart of the multi-head attention calculation adaptive allocation method of another Transformer model according to an embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of splitting batches of attention heads by powers of 2 according to an embodiment of the present invention;

[0032] Figure 4 is a schematic flowchart of the multi-head attention calculation adaptive allocation method of yet another Transformer model according to an embodiment of the present invention;

[0033] Figure 5 is a schematic input-output diagram of multi-head attention calculation according to an embodiment of the present invention;

[0034] Figure 6 is a schematic diagram of 16 computing cores performing attention calculations for a batch composed of 8 attention heads according to an embodiment of the present invention;

[0035] Figure 7 is a schematic flowchart of the multi-head attention calculation adaptive allocation method of the Transformer based on a many-core processor according to an embodiment of the present invention;

[0036] Figure 8 is a structural block diagram of the multi-head attention calculation adaptive allocation device of the Transformer model according to an embodiment of the present invention;

[0037] Figure 9 is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] The embodiment of the present invention provides an adaptive allocation method for multi-head attention calculation of a Transformer model. By splitting the number of attention heads in a single core group using a preset power-of-two allocation method, it can adaptively support multi-head attention calculation under various configurations where the number of attention heads in the model, the number of batch processes, and the number of execution cards are variable. It has excellent generalization and can maximize the utilization of computing core resources, approaching the peak performance of the processor.

[0040] According to the embodiment of the present invention, an embodiment of an adaptive allocation method for multi-head attention calculation of a Transformer model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0041] In this embodiment, an adaptive allocation method for multi-head attention calculation of a Transformer model is provided, which can be used for a target processor including multiple parallel computing cores. Among them, the target processor represents an execution medium that can abstract multiple parallel computing cores, and can be a many-core processor, an AI processor, an NPU (neural network processor), etc.

[0042] Figure 1 is a flowchart of the adaptive allocation method for multi-head attention calculation of a Transformer model according to the embodiment of the present invention. As Figure 1 shown, the process includes the following steps:

[0043] Step S101, obtain the number of attention heads, the number of batch processes of the Transformer model, and the initial number of core groups and the number of cards of the target processor.

[0044] Among them, the Transformer model represents a deep neural network architecture model based on the self-attention mechanism (Self-Attention), which can include multiple attention heads, and the calculation of each attention head is independent of each other and can be executed in parallel.

[0045] Further, the number of attention heads is a parameter of the Transformer model, representing the number of subspaces into which the hidden space (Hidden Size) of the Transformer model is divided, that is, the number of independent "perspectives" for simultaneous attention calculation.

[0046] Further, the batch size represents the number of groups of input samples processed simultaneously during one model operation, that is, how many groups of input data the Transformer model processes at one time.

[0047] Further, the core group of the target processor is a way of grouping the internal computing cores of the processor, and the initial number of core groups is the number of core groups originally divided.

[0048] Further, the number of cards represents the number of target processors in the form of cards used during actual inference execution.

[0049] Specifically, different Transformer models have different numbers of attention heads, the kvcache length of a single attention head is variable, the number of cards used during actual inference execution of a model is variable, and the batch size input during inference execution is also variable. For a selected set of inference configurations, that is, the number of attention heads of the specified model, the number of multi-core processor cards used during inference execution, the batch size of inference execution, the kvcache length of the attention heads will grow dynamically as the inference progresses.

[0050] Among them, kvcache includes kcache and vcache, representing the generated data of the character sequences processed by the Transformer model, stored in high-bandwidth memory (HBM) or other high-speed memories.

[0051] Further, the kvcache length is equal to the total length of the currently processed character sequence, and each character in the character sequence corresponds to an entry in the kvcache.

[0052] Step S102, based on the number of attention heads, batch size, initial number of core groups, and number of cards, use a preset power-of-two allocation method to split the number of attention heads of a single core group of the target processor to obtain multiple attention head batches of a single core group.

[0053] Among them, the preset power-of-two allocation method represents a strategy for splitting the number of attention heads of a single core group according to the power of 2 in a specific computing resource allocation scenario (such as the allocation of multi-head attention calculation of the Transformer model in the target processor).

[0054] In an optional embodiment. Assume that the number of attention heads allocated to a single core group is M. First, try to split out the largest multiple of 32 from M. That is, calculate where represents rounding down, so that the number a that can be split into 32 attention heads in one batch can be obtained. At this time, the remainder is M - 32a.

[0055] Secondly, for the remainder part, try to split out an integer multiple of 16. Calculate to obtain the number b that can be split into 16 attention heads in one batch, and the new remainder is M - 32a - 16b.

[0056] Then, for the new remainder, split out an integer multiple of 8. Calculate to obtain the number c that can be split into 8 attention heads in one batch, and then obtain a new remainder, and so on.

[0057] Finally. Continue to split according to the powers of 4 and 2 until the final remainder is 1 or 0. For example, when the remainder is 1, a batch consisting of one attention head is obtained.

[0058] Specifically, based on the obtained number of attention heads, batch processing quantity, initial number of core groups, and number of cards, the number of attention heads of a single core group of the target processor is split using a preset power-of-two allocation method, which fully considers the characteristics of the parallel computing cores of the processor and the specific situation of the model, avoids the problem of idle computing resources caused by the mismatch between the number of attention heads and the number of computing cores in the conventional allocation method, realizes the efficient utilization of computing resources, and enables Transformer models with different configurations to exert their computing capabilities as much as possible on the target processor.

[0059] Step S103, use a preset rule to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation to obtain the multi-head attention calculation result of the Transformer model.

[0060] Specifically, a preset rule can be used to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation. For example, for 32 attention heads, each computing core is allocated 32 / N attention heads for calculation; for 16 attention heads, each computing core is allocated 16 / N attention heads for calculation, and so on, until the remainder part is 1, and all N cores perform the calculation of this 1 attention head.

[0061] The multi - head attention calculation adaptive allocation method of the Transformer model provided in this embodiment can perform adaptive allocation based on the actual situation by obtaining the key parameters of the Transformer model (the number of attention heads, the number of batches) and the structural parameters of the target processor itself (the initial number of core groups and the number of cards). Further, the preset power - of - two allocation method is used to split the number of attention heads in a single core group, fully considering the characteristics of the parallel computing cores of the processor and the specific situation of the model, avoiding the problem of idle computing resources caused by the mismatch between the number of attention heads and the number of computing cores in the conventional allocation method, realizing the efficient utilization of computing resources, and enabling Transformer models with different configurations to exert their computing capabilities as much as possible on the target processor. Further, through adaptive allocation, the calculations of multiple attention heads can be reasonably distributed to each computing core for parallel execution, giving full play to the parallel computing advantage of the processor, and thus effectively improving the speed of multi - head attention calculation of the Transformer model. Therefore, by implementing the present invention, when performing the multi - head attention calculation of the Transformer model, in the case where the number of attention heads of the model, the number of batches, and the number of execution cards are variable, it can adaptively support the multi - head attention calculation under various configurations, has excellent generalization, can maximize the use of computing core resources, and approach the peak performance of the processor. And it can be applied without relying on additional technologies such as paged attention, and has a wider scope of application.

[0062] In this embodiment, a multi - head attention calculation adaptive allocation method of the Transformer model is provided, which can be used for a target processor including multiple parallel computing cores. Among them, the target processor represents an execution medium that can abstract multiple parallel computing cores, and can be a many - core processor, an AI processor, an NPU (neural network processor), etc.

[0063] Figure 2 It is a flowchart of the multi - head attention calculation adaptive allocation method of the Transformer model according to an embodiment of the present invention. As Figure 2 shown, the process includes the following steps:

[0064] Step S201, obtain the number of attention heads, the number of batches of the Transformer model, and the initial number of core groups and the number of cards of the target processor. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0065] Step S202, based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, use the preset power - of - two allocation method to split the number of attention heads in a single core group of the target processor to obtain multiple attention - head batches of a single core group.

[0066] Specifically, the above step S202 includes:

[0067] Step S2021, calculating the target number of cores of the target processor based on the initial number of core groups and the number of cards.

[0068] Specifically, the target number of cores of the target processor can be calculated using the following relational expression (1):

[0069] N 6 = N 4 × N 5 (1)

[0070] In the formula: N 6 represents the target number of cores; N 4 represents the initial number of core groups; N 5 represents the number of cards.

[0071] Furthermore,

[0072] Step S2022, calculating the global number of attention heads based on the number of attention heads and the number of batches.

[0073] Among them, the global number of attention heads represents the total number of attention heads that need to be processed in a complete calculation process under a specific model configuration (number of attention heads) and data batch processing method (number of batches), and can be used to measure the scale of the multi-head attention calculation task.

[0074] Specifically, the global number of attention heads can be calculated using the following relational expression (2):

[0075] N 3 = N 1 × N 2 (2)

[0076] In the formula: N 3 represents the global number of attention heads; N 1 represents the number of attention heads; N 2 represents the number of batches.

[0077] Furthermore, since different Transformer models may have different settings for the number of attention heads, and in practical applications, the number of batches will also be flexibly adjusted according to factors such as hardware performance and dataset scale, calculating the global number of attention heads can dynamically adapt to these changes. Regardless of how the model structure and data processing scale change, based on this global quantitative indicator, the scale of attention heads that need to be processed can be accurately grasped, and then the calculation resource allocation strategy can be optimized to ensure that in various scenarios, the multi-head attention calculation can make better use of the parallel computing power of the processor, thereby improving the calculation efficiency of the entire model.

[0078] Step S2023: Based on the target number of core groups and the global number of attention heads, use a preset power-of-two allocation method to split the number of attention heads for a single core group of the target processor, obtaining multiple attention head batches for the single core group.

[0079] Specifically, based on the determined target number of core groups and the global number of attention heads, use the preset power-of-two allocation method to more finely split the number of attention heads originally allocated to a single core group, so as to obtain multiple attention head batches suitable for the computing cores within the single core group to process. Through this splitting method, the computing core resources within the core group can be fully utilized, enabling each computing core to efficiently participate in the multi-head attention calculation, and thus improving the efficiency of the entire Transformer model in performing multi-head attention calculation on the target processor.

[0080] In some alternative embodiments, the above step S2023 includes:

[0081] Step a1: Based on the target number of core groups and the global number of attention heads, evenly allocate the attention heads of a single core group of the target processor to obtain the number of attention heads for the single core group.

[0082] Step a2: Use the preset power-of-two allocation method to split the number of attention heads for the single core group, obtaining multiple attention head batches for the single core group.

[0083] Specifically, based on the known target number of core groups and the global number of attention heads, with the core group as the allocation object, evenly allocate the attention heads and initially divide the number of attention heads to be processed by a single core group, expressed as the following relational formula (3):

[0084]

[0085] In the formula: N 7 represents the number of attention heads for a single core group.

[0086] Furthermore, after the initial even allocation of the attention heads for a single core group, in order to better adapt these attention heads to the computing core resources within the single core group, fully utilize each computing core, and avoid resource waste such as computing core idleness, further use the preset power-of-two allocation method to split the number of attention heads for the single core group, refining it into multiple attention head batches of different scales, which is convenient for subsequent targeted allocation and calculation according to the number of computing cores.

[0087] Among them, the specific splitting process of the preset power-of-two allocation method can refer to the description in the above step S102 and will not be elaborated here.

[0088] In an alternative embodiment, such as Figure 3As shown, first split out the integer multiples of 32, then split out the integer multiples of 16 from the remainder part, and then split out the integer multiples of 8 from the remainder part, and so on until the remainder part is 1 or 0.

[0089] Step S203, use the preset rule to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation to obtain the multi-head attention calculation result of the Transformer model. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.

[0090] The multi-head attention calculation adaptive allocation method of the Transformer model provided in this embodiment first accurately calculates the target core group number of the target processor based on the initial core group number and the number of cards, clearly determining the number of basic structural units of the processor available for allocating attention heads. Then, calculate the global number of attention heads through the number of attention heads and the batch processing number, overall grasping the scale of attention heads involved in the current computing task. Further, based on the target core group number and the global number of attention heads, evenly allocate the attention heads of a single core group of the target processor, ensuring that each core group is in a relatively fair state when initially allocating attention heads, avoiding the situation where some core groups undertake too many or too few attention head calculation tasks, and enabling the overall computing load to be relatively evenly distributed at the level of each core group. Further, use the preset power-of-two allocation method to split the number of attention heads of a single core group after uniform allocation, which can flexibly divide the attention heads into batches of different scales according to the number of computing cores, make full use of the computing power of each computing core, and minimize the idleness of computing cores to the greatest extent regardless of the change in the number of attention heads, optimize the utilization efficiency of computing resources within a single core group, and improve the contribution of each core group to multi-head attention calculation.

[0091] In this embodiment, a multi-head attention calculation adaptive allocation method of a Transformer model is provided, which can be used for a target processor including multiple parallel computing cores. Among them, the target processor represents an execution medium that can abstract multiple parallel computing cores, and can be a many-core processor, an AI processor, an NPU (neural network processor), etc.

[0092] Figure 4 It is a flowchart of the multi-head attention calculation adaptive allocation method of the Transformer model according to the embodiment of the present invention. As Figure 4 shown, the process includes the following steps:

[0093] Step S401, obtain the number of attention heads, batch processing number of the Transformer model, and the initial core group number and the number of cards of the target processor. For details, please refer to Figure 1Step S101 of the illustrated embodiment will not be elaborated herein.

[0094] Step S402: Based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, use a preset power-of-two distribution method to split the number of attention heads of a single core group of the target processor, obtaining multiple attention head batches for a single core group. For details, please refer to Figure 2 Step S202 of the illustrated embodiment will not be elaborated herein.

[0095] Step S403: Use a preset rule to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation to obtain the multi-head attention calculation result of the Transformer model.

[0096] Specifically, the above Step S403 includes:

[0097] Step S4031: Use a preset rule to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation to obtain the attention calculation result of each computing core.

[0098] Among them, as Figure 5 shown, when performing the attention calculation of a single head, a q vector and kvcache are required as inputs, and the calculation result is called the attention output. After the calculation of each head is completed, the attention outputs of each head are concatenated as the overall output of the multi-head attention.

[0099] In the calculation of multi-head attention, each head has its own independent q vector and kvcache, and the q vector and kvcache are data that can be directly used when performing multi-head attention calculation.

[0100] Specifically, a preset rule can be used to sequentially allocate multiple attention head batches to multiple computing cores of a single core group and perform attention calculation. As Figure 6 shown, taking the case of 16 computing cores and 8 attention heads as an example, 1 attention head is allocated to every 2 computing cores for calculation. The total length of the kvcache to be calculated by 2 computing cores is S, and each computing core is evenly allocated S / 2 length of the kvcache to perform attention calculation.

[0101] Furthermore, the attention calculation result of each computing core can be obtained through calculation.

[0102] In some alternative embodiments, the above Step S4031 includes:

[0103] Step b1: Use a preset rule to sequentially allocate multiple attention head batches to multiple computing cores of a single core group to obtain the target attention head of each computing core.

[0104] Step b2: Perform attention calculation based on the target attention heads of each computing core to obtain the attention calculation results of each computing core.

[0105] Specifically, multiple batches of attention heads can be sequentially assigned to multiple computing cores of a single core group according to a preset rule to obtain the target attention heads of each computing core. For example, as Figure 6 shown, the target attention head of each computing core is 1 / 2.

[0106] In an optional embodiment, assume there are N computing cores in a single core group. The allocation of attention head batches of different scales is as follows:

[0107] (1) Batch allocation of 32 attention heads: According to the rule, for a batch of 32 attention heads, each computing core is allocated attention heads to perform calculations. Further, these allocated attention heads are the target attention head parts of the corresponding computing cores in this batch.

[0108] (2) Batch allocation of 16 attention heads: For a batch of 16 attention heads, each computing core is allocated attention heads, and further, they are included in the scope of the target attention heads of this computing core.

[0109] (3) Batch allocation of 8 attention heads: Continuing according to the rule, for a batch of 8 attention heads, each computing core is allocated attention heads. For example, when N = 16, every 2 computing cores are allocated 1 attention head to perform calculations (that is, each computing core is allocated attention heads, that is, every 2 computing cores are jointly responsible for 1 attention head). Further, this 1 attention head becomes the target attention head (jointly responsible part) of these 2 computing cores.

[0110] (4) Batch allocation of the remaining 1 attention head: When there is a batch formed by the remaining 1 attention head, all N computing cores jointly perform the calculation task of this 1 attention head, that is, each computing core participates in the calculation of this unique attention head. Further, this attention head is also one of the target attention heads of each computing core.

[0111] Furthermore, each computing core can perform the calculation of a single attention head according to the following steps:

[0112] (1) Calculate the attention score

[0113] First, calculate the similarity between the allocated q vector and the K vector in the kvcache. Usually, a dot product operation or other methods can be used to obtain the attention score. For example, if the dimension of the q vector is d model, the dimension of the K vector is also d model , through the dot product operation q·K T (K T denoting the transpose of the K vector) can obtain a numerical value (in actual applications, it is often a batch operation and will obtain a matrix - form result). This numerical value reflects the degree of correlation between the query represented by the current q vector and the elements represented by the corresponding K vector.

[0114] Then, to prevent problems such as overly large numerical values, the obtained dot - product result can be scaled. Commonly, it is divided by to obtain the scaled attention scores.

[0115] Finally, through normalization operations such as the softmax function, these attention scores are transformed into a probability distribution form, making the sum of probabilities corresponding to each row equal to 1. Thus, the importance weights of each element for the current query can be determined.

[0116] (2) Calculate attention scores

[0117] Specifically, using the calculated weights, a weighted sum of the v vectors in the kvcache is performed. For example, if the weight matrix is A (the result after normalization such as softmax), and the matrix composed of v vectors is V, the result of the weighted sum can be obtained through matrix multiplication AV, that is, a part of the calculation result of a single attention head on this computing core is obtained.

[0118] Furthermore, the above - mentioned calculation results of each attention head can be integrated to finally form the complete attention calculation result of this computing core for all its target attention heads, that is, the attention calculation result of each computing core is obtained.

[0119] Step S4032, determine the multi - head attention calculation result of the Transformer model based on the attention calculation result of each computing core.

[0120] In some alternative embodiments, the above - mentioned step S4032 includes:

[0121] Step c1, determine the attention - head calculation result of each attention - head batch based on the attention calculation result of each computing core.

[0122] Step c2, determine the multi - head attention calculation result of the Transformer model based on the attention - head calculation result of each attention - head batch.

[0123] Specifically, after each computing core completes its own attention calculation, since there are cases where multiple computing cores are jointly responsible for one attention head when allocating attention heads (for example, every 2 computing cores process 1 attention head), the results of multiple computing cores participating in the calculation of the same attention head can be merged and aggregated through a reduction operation, so as to obtain the attention head calculation results of each attention head batch.

[0124] In an alternative embodiment, assume that for a certain attention head in a certain attention head batch, it is jointly calculated by 8 / N neighboring cores (for example, when N = 16, every 2 cores are responsible for 1 attention head).

[0125] Furthermore, each of these 8 / N computing cores has completed the calculation for this attention head and obtained the corresponding partial results.

[0126] Furthermore, the partial results of these 8 / N computing cores are merged according to certain rules. For example, commonly, these results are added together or other merging methods that conform to the model calculation logic are used, so as to obtain the complete calculation result of this attention head, that is, the result integration of a single attention head is completed. After performing such a reduction operation on each attention head in this attention head batch, the attention head calculation results of the entire attention head batch can be obtained.

[0127] Furthermore, the attention head calculation results obtained after reduction for each attention head batch can be concatenated in a certain order or merged using other appropriate integration methods to form the multi-head attention calculation results of the Transformer model. For example, first, the results of all 32 attention head batches are integrated together, then the results of 16 attention head batches are integrated, and so on, until finally a complete result vector or result matrix and other forms are formed, which are the multi-head attention calculation results of the Transformer model.

[0128] The multi-head attention calculation adaptive allocation method of the Transformer model provided in this embodiment uses a preset rule to sequentially allocate multiple attention head batches to multiple computing cores in a single core group, so that each computing core can clearly identify the target attention head it needs to process, avoiding the chaos and intersection of computing tasks, clearly defining the computing responsibility scope of each computing core, enabling the computing core to focus on the corresponding attention head calculation, and improving the accuracy and efficiency of each computing core in executing computing tasks. Further, due to the clear task allocation, the computing core can perform targeted operations based on the allocated attention head, and thus can efficiently obtain the attention calculation results of each computing core. Further, based on the attention calculation results of each computing core, the attention head calculation results of each attention head batch are determined, realizing the integration from the calculation results of a single computing core to the results of each attention head batch, summarizing the results related to the attention heads of the same batch scattered in each computing core, and ensuring that the calculation results of each attention head batch can accurately reflect the calculation situation of the attention head set represented by this batch. Further, based on the attention head calculation results of each attention head batch, the multi-head attention calculation results of the Transformer model are determined, ensuring that the finally obtained multi-head attention calculation results of the Transformer model are accurate, complete, comprehensively reflecting the calculation results of the model under the multi-head attention mechanism, and making the entire calculation process logically rigorous and the results reliable.

[0129] In one example, an adaptive allocation method for Transformer multi-head attention calculation based on a many-core processor is provided.

[0130] Among them, a large language model based on Transformer contains multiple attention heads, and the calculations of each attention head are independent of each other and can be executed in parallel. As Figure 5 shown, when performing the attention calculation of a single head, a q vector and kvcache are required as inputs, and the calculation result is called the attention output. After the calculations of each head are completed, the attention outputs of each head are concatenated as the overall output of the multi-head attention.

[0131] In the calculation of multi-head attention, each head has its own independent q vector and kvcache. The q vector and kvcache are calculated by the previous calculation steps and are directly available data when performing the multi-head attention calculation. The kvcache is stored in high-bandwidth memory (HBM) or other high-speed memories, and its length is equal to the total length of the currently processed character sequence. Each character in the character sequence corresponds to an entry in the kvcache.

[0132] Different large language models have different numbers of attention heads. The length of the kvcache for a single attention head is variable. The number of cards used when a model actually performs inference is variable, and the batch size of the input during inference is also variable.

[0133] For a selected set of inference configurations, that is, specifying the number of attention heads of the model, the number of multi-core processor cards used during inference, and the batch size of inference, the length of the kvcache of the attention heads will grow dynamically as the inference progresses.

[0134] As Figure 7 shown, the calculation of multi-head attention is adaptively allocated according to the following steps:

[0135] 1) First, calculate the global number of attention heads and the number of core groups included in the multi-card multi-core processor based on the number of attention heads and the batch size. Then, with the core group as the allocation object, evenly distribute the attention heads;

[0136] Among them, the number of attention heads is a model parameter, and different models have different numbers, denoted as N1. The batch size is how many groups of input data are processed by the model at one time, denoted as N2. Then the global number of attention heads N3 = N1 * N2. Let the number of core groups included in the multi-core processor be N4, and the number of multi-core processors, that is, the number of cards, be N5. Then the total number of core groups N6 = N4 * N5. With the core group as the allocation object, evenly distribute the attention heads, and the number of attention heads assigned to a single core group N7 = N3 / N6.

[0137] 2) Inside a core group, there are N computing cores. The number of attention heads often cannot be evenly divided by the number of computing cores. To make full use of the computing core resources, split the number of attention heads assigned to the core group according to powers of 2, and multiple attention head batches are obtained after splitting. As Figure 3 shown, for example, first split out multiples of 32, then split out multiples of 16 from the remainder part, and then split out multiples of 8 from the remainder part, and so on until the remainder part is 1 or 0;

[0138] 3) For 32 attention heads, each computing core is assigned 32 / N attention heads to perform calculations; for 16 attention heads, each computing core is assigned 16 / N attention heads to perform calculations, and so on until the remainder part is 1, and all N cores perform the calculation of this 1 attention head;

[0139] Among them, the number of computing cores N within the core group is fixed and unchangeable. To fully accelerate the execution of the model, it is necessary to utilize all computing core resources (no core is idle). The number of attention heads that is a power of 2 can ensure that no core is idle. For example, when N = 16 and the number of attention heads is 32, each core processes 2 heads. When the number of attention heads is 8, every 2 cores cooperate to process one head, thus ensuring that no core is idle. If the number of heads processed in the current round does not meet a power of 2, such as 7 heads, then 2 cores process one head, occupying a total of 14 cores, resulting in a waste of 16 - 14 = 2 computing core resources.

[0140] 4) As Figure 6 shown, taking the case of 16 computing cores and 8 attention heads as an example, every 2 computing cores are allocated 1 attention head to perform calculations. The total length of the kvcache to be calculated by 2 cores is S, and each core is evenly allocated S / 2 length of the kvcache to perform attention calculations;

[0141] Among them, the value of S is determined by the input data of the model (that is, how long the data the user wants the model to process).

[0142] 5) After the attention calculation of a single core is completed, reduction is performed among the 8 / N adjacent cores participating in the calculation of the same attention head to obtain the calculation result of a single attention head, and then the calculation result is written back to the corresponding position in the device main memory;

[0143] Among them, reduction is a specialized term in the field of high-performance computing, referring to combining multiple partial results to obtain the final result. Taking N = 16 cores as an example, there are a total of 8 attention heads, and 2 cores cooperate to process one attention head. After each core finishes processing its own data, it can only obtain a part of the final result, and the results of 2 cores need to be combined to obtain the final result.

[0144] 6) After all batches of attention heads that are powers of 2 are executed, the calculation of all multi-head attentions is completed.

[0145] The method for adaptively allocating Transformer multi-head attention calculations based on a many-core processor provided in this example can adaptively support multi-head attention calculations under various configurations when performing multi-head attention calculations of the Transformer model, with variable numbers of model attention heads, variable batch processing numbers, and variable numbers of execution cards. It has excellent generalization and can utilize all computing core resources to the maximum extent, approaching the peak performance of the processor. And it can be applied without relying on additional technologies such as paged attention, with a wider scope of application.

[0146] In this embodiment, a multi-head attention calculation adaptive allocation device for a Transformer model is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" may be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0147] This embodiment provides a multi-head attention calculation adaptive allocation device for a Transformer model, which can be used for a target processor including multiple parallel computing cores. As Figure 8 shown, the device includes:

[0148] An acquisition module 801, configured to acquire the number of attention heads of the Transformer model, the number of batches, and the initial number of core groups and the number of cards of the target processor.

[0149] A splitting module 802, configured to split the number of attention heads of a single core group of the target processor based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, using a preset power-of-two allocation method, to obtain multiple attention head batches of a single core group.

[0150] An allocation and calculation module 803, configured to sequentially allocate multiple attention head batches to multiple computing cores of a single core group using a preset rule and perform attention calculation, to obtain the multi-head attention calculation result of the Transformer model.

[0151] In some alternative implementation manners, the splitting module 802 includes:

[0152] A first calculation sub-module, configured to calculate the target number of core groups of the target processor based on the initial number of core groups and the number of cards.

[0153] A second calculation sub-module, configured to calculate the global number of attention heads based on the number of attention heads and the number of batches.

[0154] A splitting sub-module, configured to split the number of attention heads of a single core group of the target processor based on the target number of core groups and the global number of attention heads, using a preset power-of-two allocation method, to obtain multiple attention head batches of a single core group.

[0155] In some alternative implementation manners, the splitting sub-module includes:

[0156] A first allocation unit, configured to evenly allocate the attention heads of a single core group of the target processor based on the target number of core groups and the global number of attention heads, to obtain the number of attention heads of a single core group.

[0157] The splitting unit is used to split the number of attention heads of a single core group by using a preset power-of-two allocation method, so as to obtain multiple attention head batches of a single core group.

[0158] In some alternative embodiments, the allocation calculation module 803 includes:

[0159] The allocation calculation sub-module is used to sequentially allocate multiple attention head batches to multiple computing cores of a single core group by using a preset rule and perform attention calculation, so as to obtain the attention calculation result of each computing core.

[0160] The determination sub-module is used to determine the multi-head attention calculation result of the Transformer model based on the attention calculation result of each computing core.

[0161] In some alternative embodiments, the allocation calculation sub-module includes:

[0162] The second allocation unit is used to sequentially allocate multiple attention head batches to multiple computing cores of a single core group by using a preset rule, so as to obtain the target attention head of each computing core.

[0163] The calculation unit is used to perform attention calculation based on the target attention head of each computing core, so as to obtain the attention calculation result of each computing core.

[0164] In some alternative embodiments, the determination sub-module includes:

[0165] The first determination unit is used to determine the attention head calculation result of each attention head batch based on the attention calculation result of each computing core.

[0166] The second determination unit is used to determine the multi-head attention calculation result of the Transformer model based on the attention head calculation result of each attention head batch.

[0167] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0168] The multi-head attention calculation adaptive allocation device of the Transformer model in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0169] The embodiment of the present invention also provides a computer device having the above Figure 8 shown multi-head attention calculation adaptive allocation device of the Transformer model.

[0170] See also Figure 9 , Figure 9 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 9 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 9 A processor 10 is taken as an example.

[0171] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0172] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0173] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0174] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0175] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0176] An embodiment of the present invention further provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.

[0177] A part of the present invention can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present invention can be invoked or provided. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for a computer to execute computer program instructions include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0178] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A multi-head attention calculation adaptive allocation method for a Transformer model, characterized in that: For a target processor including multiple parallel computing cores; the method comprises: Obtain the number of attention heads and batches of the Transformer model, as well as the initial number of core groups and cards of the target processor; Based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, the number of attention heads of a single core group of the target processor is split using a preset power-of-two allocation method to obtain multiple batches of attention heads of the single core group; The multiple attention head batches are sequentially allocated to the multiple computing cores of the single core group using preset rules and attention calculations are performed to obtain the multi-head attention calculation results of the Transformer model.

2. The method according to claim 1, characterized in that Based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, the number of attention heads of a single core group of the target processor is split using a preset power-of-two allocation method to obtain multiple batches of attention heads of the single core group, including: Calculating a target number of core groups of the target processor based on the initial number of core groups and the number of cards; Calculate the number of global attention heads based on the number of attention heads and the number of batches; Based on the target number of core groups and the global number of attention heads, the number of attention heads of a single core group of the target processor is split using a preset power-of-two allocation method to obtain multiple batches of attention heads of the single core group.

3. The method according to claim 2, characterized in that Based on the target number of core groups and the number of global attention heads, the number of attention heads of a single core group of the target processor is split using a preset power-of-two allocation method to obtain multiple batches of attention heads of the single core group, including: Uniformly distributing the attention heads of a single core group of the target processor based on the target core group number and the global attention head number to obtain the number of attention heads of the single core group; The number of attention heads of the single core group is split using a preset binary allocation method to obtain multiple attention head batches of the single core group.

4. The method according to claim 1, characterized in that The plurality of attention head batches are sequentially allocated to the plurality of computing cores of the single core group using a preset rule and attention calculation is performed to obtain the multi-head attention calculation result of the Transformer model, including: Allocating the multiple attention head batches to the multiple computing cores of the single core group in sequence according to a preset rule and performing attention calculation to obtain an attention calculation result of each computing core; Determine the multi-head attention calculation result of the Transformer model based on the attention calculation result of each computing core.

5. The method according to claim 4, characterized in that The plurality of attention head batches are sequentially allocated to the plurality of computing cores of the single core group using a preset rule and attention calculation is performed to obtain an attention calculation result of each computing core, including: Allocating the multiple attention head batches to the multiple computing cores of the single core group in sequence according to a preset rule to obtain a target attention head for each computing core; Attention calculation is performed based on the target attention head of each computing core to obtain the attention calculation result of each computing core.

6. The method according to claim 4, characterized in that Determining the multi-head attention calculation result of the Transformer model based on the attention calculation result of each computing core includes: Determine the attention head calculation result of each attention head batch based on the attention calculation result of each computing core; Based on the attention head calculation results of each attention head batch, determine the multi-head attention calculation results of the Transformer model.

7. A multi-head attention calculation adaptive allocation device for a Transformer model, characterized in that: For a target processor including multiple parallel computing cores; the device comprises: An acquisition module, used to obtain the number of attention heads and batch processing of the Transformer model, as well as the initial number of core groups and the number of cards of the target processor; A splitting module, configured to split the number of attention heads of a single core group of the target processor using a preset power-of-two allocation method based on the number of attention heads, the number of batches, the initial number of core groups, and the number of cards, to obtain multiple batches of attention heads of the single core group; The allocation calculation module is used to allocate the multiple attention head batches to the multiple computing cores of the single core group in sequence according to preset rules and perform attention calculation to obtain the multi-head attention calculation result of the Transformer model.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the multi-head attention calculation adaptive allocation method of the Transformer model according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the multi-head attention calculation adaptive allocation method of the Transformer model described in any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the multi-head attention calculation adaptive allocation method of the Transformer model described in any one of claims 1 to 6.

Citation Information

Cited By

  • Distributed task processing method, controller, controller cluster and electronic equipment

    CN120492130A