Large language model acceleration method and device based on hierarchical grouping attention, equipment and medium

By adopting the hierarchical grouping attention mechanism in the large language model, grouping the input sequences and performing hierarchical attention fusion, the problem of high computational complexity when processing ultra-long texts is solved, and a significant improvement in computing efficiency is achieved.

CN119940433AInactive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH +1

Patent Information

Application Number
CN202411964485.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing ultra-long text sequences, the calculation complexity of large language models increases quadratically, resulting in a sharp increase in video memory usage and inference time, affecting processing efficiency.

Method used

Using an acceleration method based on hierarchical grouping attention, hierarchical attention fusion is performed by grouping input sequences and using intra- and inter-group attention mechanisms to reduce the complexity of attention calculations.

Benefits of technology

It significantly reduces the computational complexity of large language models when processing ultra-long text, reduces video memory usage and inference time, and improves inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940433A_ABST
    Figure CN119940433A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model acceleration method and device based on hierarchical grouping attention, equipment and a medium, and the method comprises the following steps: carrying out the grouping processing of an input sequence in the reasoning process of a large language model; an intra-group attention mechanism is used for the grouped sequences, and intra-group attention is obtained; an inter-group attention mechanism is used for the grouped sequences, and inter-group attention is obtained; and performing hierarchical attention fusion on the intra-group attention and the inter-group attention to obtain a final result of the current attention module. According to the method, the attention calculation complexity of the basic module of the large language model can be greatly reduced, and the video memory and reasoning time consumed by the large language model for processing the super-long sequence text are greatly reduced, so that the reasoning efficiency is greatly improved. The method can be widely applied to the technical field of natural languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language technology, and in particular to a large language model acceleration method, device, equipment and medium based on hierarchical grouping attention. Background Art

[0002] In recent years, the large language model based on the Transformer module has achieved remarkable results in multiple tasks in the field of natural language processing. Among them, the large language model's ability to read ultra-long texts is particularly outstanding. It can read ultra-long texts containing tens of thousands of word sequences in parallel in a short period of time, and perform a series of language understanding tasks, such as information extraction, text translation, sentiment analysis, cloze test, dialogue generation, etc. At present, the large language model has been gradually applied to the web pages and programs of major artificial intelligence manufacturers at home and abroad, bringing great convenience to human society.

[0003] One of the key components of the large language model based on Transformer is the self-attention module. The original operation process of the self-attention module is as follows. First, the input long sequence X consisting of multiple word vectors (Tokens) is represented as X = [x1, x2, ..., x n ], where each word vector The self-attention process first converts the input long sequence X table into a key-value matrix K = XW through linear projection. K 、Query matrix Q = XW Q and the numerical matrix V = XW V . Then, the inner product of the query matrix Q and the key matrix K is calculated as the attention weights. These attention weights quantify the influence of each word vector on other word vectors, which can be expressed mathematically as:

[0004]

[0005] However, in formula (1), as the length n of the long sequence X increases, the computational complexity increases quadratically, that is, O(n 2). This poses a major challenge to large language models in processing very long text sequences, greatly affecting the processing time and memory usage of long texts. For example, on NVIDIA's TITAN XP graphics card, when the sequence length n is 2048, the video memory usage of a single self-attention module during inference is 36.13MB, and the inference time per word vector is 5.06 milliseconds. When the sequence length n is 8192, the video memory usage of a single self-attention module during inference is as high as 300MB, and the inference time per word vector is 88 milliseconds. It can be seen that as the length of the text increases, the memory and time required for inference increase dramatically. Large language models often contain dozens or hundreds of self-attention modules that are executed sequentially, requiring huge computing resources to support their processing of very long texts. Summary of the invention

[0006] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a large language model acceleration method, device, equipment and medium based on hierarchical grouping attention.

[0007] The first technical solution adopted by the present invention is:

[0008] A large language model acceleration method based on hierarchical grouping attention includes the following steps:

[0009] During the inference process of the large language model, the input sequence is processed in groups;

[0010] Use the intra-group attention mechanism on the grouped sequences to obtain intra-group attention;

[0011] Use the inter-group attention mechanism on the grouped sequences to obtain inter-group attention;

[0012] Perform hierarchical attention fusion on the intra-group attention and inter-group attention to obtain the final result of the current attention module.

[0013] Furthermore, in the inference process of the large language model, the input sequence is grouped and processed, including:

[0014] The input sequence X = [x1, x2, ..., x n ] are grouped, each group contains k word vectors (Token), so there are a total of groups; where n is the number of word vectors contained in the input sequence X, and the i-th group is represented by X i =[x (i-1)*k+1 , x (i-1)*k+2 , ..., x min(i*k,n) ];

[0015] For the grouped sequence X i , perform adaptive filling to solve the last group Xm The problem of having fewer than k word vectors.

[0016] Furthermore, for the grouped sequence X i , perform adaptive filling, including:

[0017] Copy part of the word vectors connected to the last group from the second-to-last group to the last group to fill the last group to k word vectors.

[0018] Furthermore, the grouped sequences are subjected to an intra-group attention mechanism to obtain intra-group attention, including:

[0019] For the sequence X of the i-th group i , intra-group attention G i The calculation method is:

[0020] G i =Attention(Q i , K i , V i )

[0021] In the formula, Q i , K i and V i are the sequences X from the i-th group respectively. i The query matrix, key matrix and value matrix obtained by linear transformation of matrix multiplication;

[0022] Concatenate all the outputs of self-attention of each group to get the final output A Intra =[G1, G2, ... G m ], as the in-group attention.

[0023] Furthermore, the query matrix Q i , key matrix K i and the numerical matrix V i The calculation formula is as follows:

[0024] Q i =X i W Q

[0025] K i =X i W k

[0026] V i =X i W V

[0027] Where W Q , W k and W VThese are the parameters learned by the model;

[0028] In-group attention Intra The dimension of is the same as that of the original input sequence X, that is, it is consistent with the input and output dimensions of the original self-attention.

[0029] Furthermore, the grouped sequences are subjected to an inter-group attention mechanism to obtain inter-group attention, including:

[0030] Concatenate each grouped sequence X i The word vectors (Token) in are used to create the group vector u i ;

[0031] Each group vector is concatenated to form U = [u1u2, ..., u m ], the key matrix for calculating the attention between groups and the numerical matrix Calculate the query matrix of inter-group attention based on the input sequence X Where W K , W V and W Q These are the parameters learned by the model;

[0032] based on The inter-group attention is calculated as follows:

[0033]

[0034] Furthermore, the final result of the current attention module is expressed as:

[0035] A=(A Intra +A Inter ) / 2

[0036] In the formula, A Intra is the in-group attention, A Inter For inter-group attention.

[0037] The second technical solution adopted by the present invention is:

[0038] A large language model acceleration device based on hierarchical grouping attention, comprising:

[0039] The sequence grouping module is used to group the input sequence during the reasoning process of the large language model;

[0040] The intra-group calculation module is used to apply the intra-group attention mechanism to the grouped sequences to obtain the intra-group attention;

[0041] The inter-group calculation module is used to apply the inter-group attention mechanism to the grouped sequences to obtain the inter-group attention;

[0042] The fusion calculation module is used to perform hierarchical attention fusion on the intra-group attention and the inter-group attention to obtain the final result of the current attention module.

[0043] The third technical solution adopted by the present invention is:

[0044] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a large language model acceleration method based on hierarchical grouping attention as described above.

[0045] The fourth technical solution adopted by the present invention is:

[0046] A computer-readable storage medium, wherein at least one instruction, at least one program, code set or instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement a large language model acceleration method based on hierarchical grouping attention as described above.

[0047] The fifth technical solution adopted by the present invention is:

[0048] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned large language model acceleration method based on hierarchical grouping attention.

[0049] The beneficial effect of the present invention is that the solution proposed by the present invention can calculate the attention of the basic module of the large language model with significantly lower complexity when processing ultra-long texts, greatly reducing the video memory and reasoning time required for the large language model to process ultra-long sequence texts, thereby greatly improving the reasoning efficiency. The present invention is applicable to all natural language models based on Transformer, and can also be applied to computer vision models based on Transformer after appropriate modification. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 It is a flowchart of the steps of a large language model acceleration method based on hierarchical grouping attention in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0053] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0054] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0055] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0056] Example 1

[0057] like Figure 1 As shown, this embodiment provides a large language model acceleration method based on hierarchical grouping attention, which can greatly reduce the attention calculation complexity of the basic module of the large language model, and greatly reduce the video memory and reasoning time required for the large language model to process ultra-long sequence texts, thereby greatly improving the reasoning efficiency. The method specifically includes the following steps:

[0058] S 1. In the reasoning process of the large language model, the input sequence is grouped and processed.

[0059] In order to reduce the computational complexity of the original self-attention mechanism O(n 2 ), the embodiment of the present invention proposes a hierarchical grouping attention method. The attention calculation method can realize attention calculation with lower complexity. When the large language model accepts a long text sequence, the input sequence is first divided into several groups evenly.

[0060] Exemplarily, step S1 specifically includes the following steps:

[0061] S11. Given an input sequence X = [x1, x2, ..., x n ], this embodiment divides the input sequence X into several groups, each group contains k word vectors (Token), so there are a total of groups.

[0062] For the sake of clarity and convenience of description, this embodiment uses represents the i-th group. Where n is the number of word vectors contained in the input sequence, and d is the feature dimension of each word vector. At this time, the i-th group can be represented as X i =[x (i-1)*k+1 , x (i-1)*k+2 , …, x min(i*k,n) ].

[0063] S12, for the grouped X obtained in step S11 i , and perform adaptive filling.

[0064] Because large language models that make predictions in an autoregressive fashion usually predict each subsequent word vector based on the features of the previous word vector. Therefore, if the length n of the input sequence is not divisible by the group size k, this will result in the last group X being m It may consist of less than k word vectors, or even only one or two word vectors in extreme cases. If the number of word vectors in the last group is too small, it may limit the ability of the group to provide important contextual information. This important contextual information helps the model obtain stable and effective intermediate layer feature representations and affects the overall coherence and fluency of the model's prediction results.

[0065] In order to solve this problem, the filling scheme adopted in this embodiment is: copy a part of the word vectors connected to the last group from the second to last group to the last group, and fill the last group to k word vectors. It should be noted that the present invention is not limited to the above filling scheme, and other schemes that can fill the last group to k word vectors are also applicable to the present invention application and should fall within the protection scope of the present invention.

[0066] S2. Use the intra-group attention mechanism on the grouped sequences to obtain intra-group attention.

[0067] This embodiment uses the intra-group attention mechanism based on the grouped input sequence obtained in step S1 to calculate the attention within each group, which is called intra-group attention.

[0068] As an optional implementation, step S2 specifically includes the following steps:

[0069] S21, for the sequence X of the i-th group i , intra-group attention G i The calculation method is:

[0070] G i =Attention(Q i , K i , V i ) (2)

[0071] Among them, Q i =X i W Q , K i =X i W k and V i =X i W V are the sequences X from the i-th group respectively. i The query matrix, key matrix and value matrix are obtained by matrix multiplication and linear transformation. Here W Q , W k and W V are the parameters learned by the model.

[0072] S22. Concatenate all the outputs of self-attention of each group to get the final output AI nt ra=[G1,G2,...G m ].

[0073] Specifically, The original input sequence is maintained The same dimension, that is, the input and output dimensions are consistent with the original self-attention. In actual calculations, different grouped self-attentions can be executed in parallel on the computing device, thereby improving the computing efficiency.

[0074] S3. Use the inter-group attention mechanism on the grouped sequences to obtain inter-group attention.

[0075] In the above formula (2), the intra-group attention can only capture the relationship between the tokens in each group. When the length of the input sequence exceeds the size k of each group, it will hinder the model's ability to effectively use long-range context information to complete complex language tasks (such as question-answering tasks). To solve this problem, the embodiment of the present invention proposes an inter-group attention mechanism that can enhance the model's ability to use information from different groups and ensure a comprehensive understanding of the entire long sequence.

[0076] In some embodiments, step S3 specifically includes the following steps:

[0077] S31, first splice each group X i The word vectors (Token) in are used to create the group vector u i Then, concatenate each group vector to form U = [u1u2, ..., u m ]. These group vectors U encapsulate the key information of the corresponding group. Therefore, the embodiment of the present invention uses U to replace X in the original solution to calculate the key matrix of inter-group attention and the numerical matrix And the query matrix for calculating the attention between groups Then continue to use the X calculation in the original solution.

[0078] S32. At this time, the dimension of U is significantly less than the input sequence X, thus greatly reducing the computational cost. The inter-group attention is calculated as follows:

[0079]

[0080] In order to avoid increasing the number of model parameters, the embodiment of the present invention reuses the linear projection parameter W from the intra-group attention. Q , W K and W V Therefore, the inter-group attention mechanism proposed in this invention does not introduce new parameters. In addition, the inter-group attention A calculated Inter The dimension of is the same as that of the input sequence X, that is, n × d. Among them, n is the number of word vectors contained in the input sequence, and d is the feature dimension of each word vector.

[0081] S4. Perform hierarchical attention fusion on the intra-group attention and inter-group attention to obtain the final result of the current attention module.

[0082] In order to fuse the information from the intra-group and inter-group attention, this embodiment calculates their average to obtain the final attention A = (A Intra +A Inter) / 2. At this time, the hierarchical grouping attention calculation method proposed in the embodiment of the present invention can be used as an efficient plug-and-play substitute for the traditional self-attention calculation method, and maintains the consistency of the input and output dimensions of the self-attention calculation module without introducing additional parameters.

[0083] It should be noted that in this embodiment, the final attention is obtained by averaging the intra-group attention and the inter-group attention, but it is not limited to this. The final attention can also be calculated by assigning a weight value between the intra-group attention and the inter-group attention.

[0084] In general, the differences between the present invention and the prior art are as follows:

[0085] Computational complexity of existing technologies: One of the key components of the existing large language model based on Transformer is the self-attention module. The original operation process of the self-attention module is shown in formula (1). Its computational complexity increases quadratically with the increase of the length n of the long sequence X, that is, O(n 2 ), which brings great challenges to large language models in processing very long text sequences, and greatly affects the processing time and memory usage of long texts.

[0086] Summary of existing technologies: As the length of text increases, the memory and time required for inference in existing technologies increase quadratically. Large language models often contain dozens or hundreds of self-attention modules that are executed sequentially, so huge computing resources are required to support their processing of ultra-long texts.

[0087] Based on this, the present invention proposes a large language model acceleration method based on hierarchical grouping attention, which is O(n 2 ) computational complexity, the method of the present invention can significantly lower the complexity O(n 2 / k+n*k) calculates the attention of the large language model basic module when processing ultra-long text (n>>1024), where k is a constant, generally 256. This is significantly different from the prior art.

[0088] The complexity of the intra-group attention and inter-group attention mechanisms proposed in this invention will be analyzed below. For intra-group attention, each group contains k word vectors, and the computational complexity of each group is O(k 2 ). groups, the total intra-group attention complexity is O(nk). Since k is a hyperparameter that is usually considered a constant (the recommended value is 256), the overall complexity can be simplified to O(n). As for inter-group attention, the query vector contains n word vectors, and the key vector is composed of word vectors, so the computational complexity of the inter-group attention is O(n2 / k). In summary, the overall complexity of the hierarchical grouping attention mechanism proposed in this invention is O(n 2 / k+nk). In fact, by setting k to a large constant (e.g., 256), the computational cost of processing a long input sequence (e.g., n=128,000) can be reduced to approximately 1 / 256 of the computational cost required by the traditional self-attention mechanism. This demonstrates the efficiency of the present invention in processing very long texts.

[0089] Example 2

[0090] This embodiment provides a large language model acceleration device based on hierarchical grouping attention, including:

[0091] The sequence grouping module is used to group the input sequence during the reasoning process of the large language model;

[0092] The intra-group calculation module is used to apply the intra-group attention mechanism to the grouped sequences to obtain the intra-group attention;

[0093] The inter-group calculation module is used to apply the inter-group attention mechanism to the grouped sequences to obtain the inter-group attention;

[0094] The fusion calculation module is used to perform hierarchical attention fusion on the intra-group attention and the inter-group attention to obtain the final result of the current attention module.

[0095] Since the device is a large language model acceleration device based on hierarchical grouping attention according to an embodiment of the present invention, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0096] Example 3

[0097] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A large language model acceleration method based on hierarchical group attention is shown.

[0098] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0099] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.

[0100] Since the electronic device is an electronic device corresponding to a large language model acceleration method based on hierarchical grouping attention in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0101] Example 4

[0102] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A large language model acceleration method based on hierarchical group attention is shown.

[0103] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, the storage medium including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0104] Since the storage medium is a storage medium corresponding to a large language model acceleration method based on hierarchical grouping attention in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0105] Example 5

[0106] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a large language model acceleration method based on hierarchical grouping attention according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing various embodiments can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0107] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0108] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0109] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.

Claims

1. A large language model acceleration method based on hierarchical grouping attention, characterized in that: The following steps are involved: During the inference process of the large language model, the input sequence is processed in groups; Use the intra-group attention mechanism on the grouped sequences to obtain intra-group attention; Use the inter-group attention mechanism on the grouped sequences to obtain inter-group attention; Perform hierarchical attention fusion on the intra-group attention and inter-group attention to obtain the final result of the current attention module.

2. According to claim 1, a large language model acceleration method based on hierarchical grouping attention is characterized in that: In the inference process of the large language model, the input sequence is grouped and processed, including: The input sequence X=[x1,x2,…,x n ] are grouped, each group contains k word vectors, so there are a total of groups; where n is the number of word vectors contained in the input sequence X, and the i-th group is represented by X i =[x (i-1)*k+1 ,x (i-1)*k+2 ,...,x min(i*k,n) ]; For the grouped sequence X i , perform adaptive filling to solve the last group X m The problem of having fewer than k word vectors.

3. A large language model acceleration method based on hierarchical grouping attention according to claim 2, characterized in that: For the grouped sequence X i , perform adaptive filling, including: Copy part of the word vectors connected to the last group from the second-to-last group to the last group to fill the last group to k word vectors.

4. The large language model acceleration method based on hierarchical grouping attention according to claim 1 is characterized in that: The grouped sequences are subjected to an intra-group attention mechanism to obtain intra-group attention, including: For the sequence X of the i-th group i , intra-group attention G i The calculation method is: G i =Attention(Q i ,K i ,V i ) In the formula, Q i , K i and V i are the sequences X from the i-th group respectively. i The query matrix, key matrix and value matrix obtained by linear transformation of matrix multiplication; Concatenate all the outputs of self-attention of each group to get the final output A Intra =[G1,G2,…G m ], as the in-group attention.

5. The large language model acceleration method based on hierarchical grouping attention according to claim 4 is characterized in that: Query Matrix Q i , key matrix K i and the numerical matrix V i The calculation formula is as follows: Q i =X i W Q K i =X i W k V i =X i W V Where W Q , W k and W V These are the parameters learned by the model; In-group attention Intra The dimension of is the same as that of the original input sequence X, that is, it is consistent with the input and output dimensions of the original self-attention.

6. The large language model acceleration method based on hierarchical grouping attention according to claim 1, characterized in that: The inter-group attention mechanism is used on the grouped sequences to obtain the inter-group attention, including: Concatenate each grouped sequence X i The word vectors in the group vector u are created i ; Each group vector is concatenated to form U = [u1u2, ···, u m ], the key matrix for calculating the attention between groups and the numerical matrix Calculate the query matrix of inter-group attention based on the input sequence X Where W K , W V and W Q These are the parameters learned by the model; based on The inter-group attention is calculated as follows:

7. The large language model acceleration method based on hierarchical grouping attention according to claim 1, characterized in that: The final result of the current attention module is expressed as: A=(A Intra +A Inter ) / 2 In the formula, A Intra is the in-group attention, A Inter For inter-group attention.

8. A large language model acceleration device based on hierarchical grouping attention, characterized in that: include: The sequence grouping module is used to group the input sequence during the reasoning process of the large language model; The intra-group calculation module is used to apply the intra-group attention mechanism to the grouped sequences to obtain the intra-group attention; The inter-group calculation module is used to apply the inter-group attention mechanism to the grouped sequences to obtain the inter-group attention; The fusion calculation module is used to perform hierarchical attention fusion on the intra-group attention and the inter-group attention to obtain the final result of the current attention module.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large language model reasoning acceleration method and device based on sparse sliding window

    CN118132682A

  • Medical long text question and answer method and device, electronic equipment and storage medium

    CN118520882A

Cited By

  • Petrochemical static equipment multi-view depth representation learning method and system

    CN121259513A