Sequence processing method based on MLA, related device, system and storage medium

By performing feature processing of low-rank compression and rotation position embedding in multi-head potential attention (MLA), the traffic volume of the computing node is reduced, the problem of excessive traffic volume in MLA sequence processing is solved, and the processing efficiency is improved.

CN120296366AActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510781985.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing multi-potential attention (MLA) has too high traffic to compute nodes during sequence processing, affecting processing efficiency.

Method used

Rotating position embedding through low-rank compression query and key-value features, and communication between computing nodes is performed, parallel aggregation operations and multi-head attention mechanism are separated, and communication volume of computing nodes is reduced.

Benefits of technology

It effectively reduces the traffic of the computing node and improves the processing efficiency of the computing node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296366A_ABST
    Figure CN120296366A_ABST
Patent Text Reader

Abstract

The invention discloses an MLA-based sequence processing method, a related device, a system and a storage medium, and the method comprises the steps: carrying out the decompression of a first query feature based on low-rank compression, and then carrying out the rotation position embedding, and obtaining a target query feature; segmenting the first key value features based on low-rank compression to obtain to-be-processed first sub-key value features of different calculation units; performing rotation position embedding on the distributed and processed first sub-key value features by different calculation units to obtain second sub-key value features; performing communication based on each computing node to obtain a second key value feature embedded by the first key value feature through the rotation position; obtaining a target key feature and a target value feature based on the second key value feature; a multi-head attention mechanism is performed based on the target query feature, the target key feature, and the target value feature. According to the scheme, the communication traffic can be reduced during sequence processing based on the MLA, so that the processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a sequence processing method based on MLA, and related devices, systems, and storage media. Background Art

[0002] In recent years, with the remarkable development of large language models, many excellent network architectures have emerged. For example, compared with ordinary multi-head attention (MHA), multi-head latent attention (MLA) benefits from performing low-rank compression on Q, K, and V before executing MHA, which greatly reduces the video memory requirements of the KV cache during inference.

[0003] This application's research found that in existing MLA, its parallel aggregation operation is carried out in MHA, resulting in a sharp increase in the communication volume of computing nodes during sequence processing based on MLA, thereby affecting the final processing efficiency of computing nodes. In view of this, how to reduce the communication volume of computing nodes during sequence processing based on MLA to improve the final processing efficiency of computing nodes has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a sequence processing method based on MLA, and related devices, systems, and storage media, which can reduce the communication volume of computing nodes during sequence processing based on MLA to improve the final processing efficiency of computing nodes.

[0005] To solve the above technical problem, the first aspect of this application provides a sequence processing method based on MLA, including: after decompressing the first query feature based on low-rank compression, performing rotary position embedding to obtain the target query feature of the multi-head attention mechanism; wherein, the first query feature is obtained from the text sequence to be responded; splitting the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node; respectively performing rotary position embedding on the first sub-key-value features assigned to be processed by different computing units in each computing node to obtain the second sub-key-value features; based on communication among each computing node, obtaining the second key-value feature after rotary position embedding of the first key-value feature; based on the second key-value feature, obtaining the target key feature and target value feature of the multi-head attention mechanism; and performing the multi-head attention mechanism based on the target query feature, target key feature, and target value feature to complete the MLA processing.

[0006] To solve the above technical problems, a second aspect of the present application provides a sequence processing device based on MLA, including: a position embedding module, a grouping and splitting module, a distribution processing module, a sequence aggregation module, a feature processing module, and a multi-head processing module. The position embedding module is configured to perform rotational position embedding on the first query feature after decompression based on low-rank compression to obtain the target query feature of the multi-head attention mechanism; wherein, the first query feature is obtained from the text sequence to be responded. The grouping and splitting module is configured to split the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node. The distribution processing module is configured to perform rotational position embedding on the allocated first sub-key-value features by different computing units in each computing node respectively to obtain the second sub-key-value features. The sequence aggregation module is configured to communicate based on each computing node to obtain the second key-value feature after rotational position embedding of the first key-value feature. The feature processing module is configured to obtain the target key feature and the target value feature of the multi-head attention mechanism based on the second key-value feature. The multi-head processing module is configured to execute the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing.

[0007] To solve the above technical problems, a third aspect of the present application provides a sequence processing system, including a plurality of computing nodes, each computing node respectively containing a plurality of computing units, and the sequence processing system is configured to execute program instructions to implement the MLA-based sequence processing method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium, storing program instructions that can be run by a processor, and the program instructions are used to implement the MLA-based sequence processing method in the first aspect above.

[0009] In the above solution, the first query feature based on low-rank compression is rotated and position-embedded after compression to obtain the target query feature of the multi-head attention mechanism. The first query feature is obtained from the text sequence to be responded to, and the first key-value feature based on low-rank compression is segmented to obtain the first sub-key-value features to be processed by different computing units in each computing node. Then, different computing units in each computing node perform rotation and position-embedding on the assigned first sub-key-value features to obtain the second sub-key-value features. Thus, communication is performed based on each computing node to obtain the second key-value feature with rotation and position-embedding of the first key-value feature. Based on the second key-value feature, the target key feature and target value feature of the multi-head attention mechanism are obtained. Furthermore, the multi-head attention mechanism is executed based on the target query feature, target key feature, and target value feature to complete the MLA processing. Since the parallel aggregation operation is separated from the multi-head attention (MHA), and the parallel aggregation operation is pre-positioned before the multi-head attention (MHA) after obtaining the first key-value feature based on low-rank compression, at this time, only the first key-value feature based on low-rank compression needs to be communicated. Compared with performing the parallel aggregation operation in the multi-head attention (MHA), the communication volume of the computing node can be greatly reduced. Therefore, the communication volume of the computing node can be reduced when performing sequence processing based on MLA, so as to improve the final processing efficiency of the computing node. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of an embodiment of the sequence processing method based on MLA of the present application; Figure 2a is a schematic diagram of the process of an embodiment of the sequence processing method based on MLA of the present application; Figure 2b is a schematic diagram of an embodiment of dimension segmentation of the first key-value feature of the present application; Figure 2c is a schematic diagram of an embodiment of CP group communication of the present application; Figure 2d is a schematic diagram of an embodiment of SP group communication of the present application; Figure 2e is a schematic diagram of the effect of an embodiment of the sequence processing method based on MLA of the present application; Figure 3 is a schematic framework diagram of an embodiment of the sequence processing device based on MLA of the present application; Figure 4 is a schematic framework diagram of an embodiment of the sequence processing system of the present application; Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments

[0011] The following will describe the solution of the embodiment of the present application in detail with reference to the accompanying drawings of the specification.

[0012] In the following description, specific details such as specific system architectures, interfaces, and technologies are proposed for illustration rather than limitation in order to thoroughly understand the present application.

[0013] The terms "system" and "network" in this article are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the front and back associated objects. In addition, "multiple" in this article means two or more than two.

[0014] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the sequence processing method based on MLA of the present application. Specifically, it may include the following steps: Step S11: After the first query feature based on low-rank compression is decompressed, rotation position embedding is performed to obtain the target query feature of the multi-head attention mechanism.

[0015] In the embodiments of the present disclosure, the first query feature can be obtained from the text sequence to be responded to. It should be noted that according to different actual application scenarios, the specific content of the text to be responded to may also be different. For example, in a teaching scenario, the text to be responded to may include, but is not limited to: "Please explain a certain knowledge point", "Please analyze the following test questions", "Please generate a lecture on a certain knowledge point", etc.; or, in a medical scenario, the text to be responded to may include, but is not limited to: "What are the indications of a certain drug", "Whether two certain drugs can be taken simultaneously", "Whether fasting is required before a certain medical examination", etc. Of course, the above examples are only several possible examples of the text to be responded to in the actual application process, and the specific content of the text to be responded to is not limited herein, nor will it be exemplified one by one.

[0016] In an implementation scenario, after obtaining the text to be responded to, it can be first embedded to obtain the embedded representation of the text to be responded to, and then the embedded representation of the text to be responded to can be sent into an intelligent dialogue model including MLA (i.e., multi-head latent attention) so that each network layer in the intelligent dialogue model processes it in turn until the previous network layer of MLA finishes processing, then the hidden layer features of the previous network layer of MLA can be obtained, and then a query low-rank projection matrix can be used to perform low-rank compression on the hidden layer features to obtain the first query feature. Please refer to Figure 2a , Figure 2a is a schematic diagram of the process of an embodiment of the sequence processing method based on MLA of the present application. As Figure 2aAs shown in the figure, for the convenience of description, the query-related low-rank compression can be denoted as q_down, and the aforementioned hidden layer features can be denoted as h t , then the first query feature can be expressed as:

[0017] In the above formula, w DQ represents the query low-rank projection matrix, represents the first query feature of the low-rank compression. It should be noted that the above is only a brief description of the low-rank projection. For the specific process, refer to the technical details of MLA, which will not be elaborated here.

[0018] In an implementation scenario, after obtaining the first query feature of the low-rank compression, the query restoration projection matrix can be used to perform a projection operation on the first query feature and then perform dimension slicing to obtain the sliced query features of each attention head in the multi-head attention. After that, for each sliced query feature, on the one hand, the sliced query feature can perform rotary position embedding, and on the other hand, the sliced query feature can be fused with its own sliced query feature after rotary position embedding (such as concatenation, etc.) to obtain the target query features of each attention head. For the convenience of understanding, the aforementioned projection operation can be denoted as q_up, and the sliced query feature after the projection operation and dimension slicing can be expressed as:

[0019] In the above formula, split represents dimension slicing, both Q_nope and Q_rope represent the sliced query features after the projection operation and dimension slicing, that is, they are essentially the same. The main difference is that the latter also needs to go through rotary position embedding, that is, it can be understood that the latter is an identical backup of the former. On this basis, for the convenience of description, the rotary position encoding can be denoted as Rotary_Embed, and the sliced query feature after rotary position embedding can be expressed as: Q_rope = Rotary_Embed(Q_rope) In the above formula, Q_rope on the right side of the equal sign represents the sliced query feature that needs to perform rotary position embedding after the projection operation and dimension slicing, and Q_rope on the left side of the equal sign represents the sliced query feature after rotary position embedding. It should be noted that for the specific process of rotary position embedding, refer to the technical details of the rotary position encoding (Rotary Position Embedding, RoPE), which will not be elaborated here. On this basis, the sliced query feature after the projection operation and dimension slicing and the sliced query feature after rotary position embedding can be fused to obtain the target query feature. For the convenience of description, taking the fusion achieved by feature concatenation as an example, the target query feature Q can be expressed as: Q = concate(Q_nope, Q_rope) In the above formula, concate represents concatenation, Q_nope represents the sharded query feature after projection operation and dimension slicing, and Q_rope represents the sharded query feature after rotational position embedding. It should be noted that the above process of obtaining the target query matrix can be completed in any computing node in the sequence processing system. For example, the sequence processing system may include multiple computing nodes (such as two computing nodes, three computing nodes, four computing nodes, etc.), and each computing node may contain several computing units (such as GPUs, etc.). In addition, the sequence processing system may also include computing resources other than computing nodes, and the above process of obtaining the target query matrix can be completed in the computing resources other than computing nodes in the sequence processing system, which is not limited here. For the specific structure of the sequence processing system, reference can be made to the embodiments of the following sequence processing system, which will not be elaborated here. In addition, for the convenience of subsequent multi-head attention mechanism processing, the feature dimension of the target query feature in the hidden layer size can be represented in the form of n * d1, where n represents the number of multi-heads, d1 represents the target dimension, and the product of the number of multi-heads n and the target dimension d1 is equal to the feature dimension of the target query feature in the hidden layer size.

[0020] Step S12: Split the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node.

[0021] In an implementation scenario, the first key-value feature of low-rank compression can be obtained by referring to the aforementioned first query feature of low-rank compression. For example, in the multi-round conversation process of an intelligent dialogue model including MLA (i.e., multi-head latent attention), the hidden layer features of the dialogue text sequence in the previous network layer of MLA can be obtained in each round of conversation, so that the hidden layer features can be low-rank compressed by a key-value low-rank projection matrix to obtain the first key-value feature and cache it, so as to greatly reduce the video memory space required for KV caching. Then, the cached first key-value feature can be obtained and split in the current round of conversation. For the convenience of description, the low-rank compression related to key-value can be denoted as kv_down, and the aforementioned hidden layer feature can be denoted as h t , then the first key-value feature can be expressed as:

[0022] In the above formula, represents the key-value low-rank projection matrix, represents the first key-value feature. It should be noted that the above is only a brief description of low-rank projection, and the specific process can refer to the technical details of MLA, which will not be elaborated here.

[0023] In an implementation scenario, as a possible implementation example, the first key-value feature of low-rank compression can be segmented according to the total number of different computing units in each computing node. Exemplarily, the first key-value feature can be evenly segmented into the above-mentioned total number of parts and distributed to different computing units in each computing node respectively. For example, in the case of having M computing nodes and each computing node having N computing units, the first key-value feature can be segmented into M*N first sub-key-value features. Of course, the above example is only a possible example of segmenting the first key-value feature in the actual application process, and other possible situations are not limited here, nor will they be exemplified one by one.

[0024] In another implementation scenario, as another possible implementation example, different from the foregoing segmentation method, the first key-value feature of low-rank compression can also be segmented successively based on the number of CP groups and the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node. Exemplarily, the first key-value feature is characterized as a feature tensor whose feature size is the product of the batch size, sequence length, and hidden layer size, and for the convenience of description, it can be denoted as B*S*H, or simply denoted as (B, S, H), where B represents the Batch size (i.e., the batch size), S represents the Sequence length (i.e., the sequence length), and H represents the Hidden size (i.e., the hidden layer size). Then, after obtaining the first key-value feature of low-rank compression, the first key-value feature can be first segmented in the feature dimension of the sequence length based on the number of CP groups for the purpose of being distributed to different computing nodes respectively after the final segmentation is completed, and then the first key-value feature after the first segmentation can be segmented again in the feature dimension of the sequence length for each sub-feature based on the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node. That is to say, the first key-value feature with a feature size of (B, S, H) is segmented for the first time according to the number of CP groups CP_size, and CP_size sub-features with a feature size of (B, S / CP_size, H) can be obtained. These sub-features will be distributed to different computing nodes respectively after the final segmentation is completed. Then, for each sub-feature, it can be further segmented according to the number of SP groups SP_size for the second time, and SP_size sub-features with a feature size of (B, S / CP_size / SP_size, H) can be obtained, which are used as the first sub-key-value features distributed to different computing units. Please refer to Figure 2b , Figure 2b is a schematic diagram of an embodiment of dimension segmentation of the first key-value feature in the present application. As Figure 2bAs shown in the figure, taking the number of CP groups CP_size and the number of SP groups SP_size as 2 respectively, the first key value feature with a feature size of (B, S, H) can first be split into 2 sub-features with a feature size of (B, S / 2, H) according to the number of CP groups, and will be allocated to different computing nodes respectively after the final splitting is completed, such as computing node 0 (i.e., Figure 2b Node0 in Figure 2b ), and computing node 1 (such as Figure 2b Node1 in Figure 2b ). Then, the first sub-feature is further split into 2 sub-features with a feature size of (B, S / 2 / 2, H) as the first sub-key value feature according to the number of SP groups, and is respectively allocated to computing unit 0 (i.e., Figure 2b GPU0 in Figure 2b Node0 in Figure 2b ), and computing unit 1 (i.e., Figure 2b GPU1 in

[0025] In the above formula, both represent the first sub-key value feature, that is, they are essentially the same, and the main difference is that the latter also needs to continue to perform rotational position embedding, that is, a backup of the former is made and renamed as the latter for further performing rotational position embedding.

[0026] Step S13: The first sub-key value features allocated and processed by different computing units in each computing node are respectively subjected to rotational position embedding to obtain second sub-key value features.

[0027] Specifically, for the specific process of performing rotational position embedding on the first sub-key value feature, reference can be made to the specific process of performing rotational position embedding on the first query feature after decompression. For example, please continue to refer to Figure 2b , computing unit 0 (i.e., Figure 2b GPU0 in Figure 2b Node0 in Figure 2b ), computing unit 1 (i.e., Figure 2b GPU1 in Figure 2bNode1) in computing unit 0 (i.e. Figure 2b GPU0), computing node 1 (i.e. Figure 2b Node1 in computing unit 1 (i.e. Figure 2b The GPU 1 in the GPU 1 can perform data processing independently, that is, the first sub-key value features assigned to each can be rotated and embedded to obtain the corresponding second sub-key value features. The specific process of each computing unit independently performing the rotation position embedding is briefly described below. For the convenience of description, for any first sub-key value feature, the first sub-key value feature after the rotation position embedding can be expressed as: K_rope=Rotary_Embed(K_rope) In the above formula, Rotary_Embed represents rotational position embedding, K_rope on the right side of the equal sign is the first sub-key value feature, and K_rope on the left side of the equal sign is the first sub-key value feature after rotational position embedding. It should be noted that for the specific process of rotational position embedding, please refer to the technical details of Rotary Position Embedding (RoPE), which will not be repeated here. On this basis, the first sub-key value feature before rotational position embedding can be fused with the first sub-key value feature after rotational position embedding to obtain the second sub-key value feature. Taking the fusion achieved through feature splicing as an example, the second sub-key value feature can be expressed as:

[0028] In the above formula, Indicates the second sub-key value feature, concate indicates concatenation, K_rope represents the first sub-key value feature before the rotation position embedding is performed, and K_rope represents the first sub-key value feature after the rotation position embedding is performed.

[0029] Step S14: communicate based on each computing node to obtain a second key-value feature in which the first key-value feature is embedded in a rotated position.

[0030] In one implementation scenario, as described above, the low-rank compressed first key-value feature can be divided according to the total number of different computing units in each computing node. In this case, different computing nodes in each computing node can exchange their respective calculated second sub-key-value features with each other, so that different computing units in each computing node can obtain the second sub-key-value features calculated by other computing units, and then each computing unit can combine the second sub-key-value features to obtain the second key-value feature embedded in the first key-value feature through rotational position encoding.

[0031] In another implementation scenario, as described above, different from the foregoing situation, the first key-value feature after low-rank compression can also be segmented successively based on the number of CP groups and the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node. In this case, CP group communication can be performed between different computing nodes and then SP group communication can be performed within the same computing node to obtain the second key-value feature after the first key-value feature is rotationally position-embedded. In the above manner, since CP group communication is performed first and then SP group communication during the execution of the parallel aggregation operation, that is, communication is first performed between computing nodes and then within computing nodes, the communication volume between computing nodes can be reduced compared with performing SP group communication first and then CP group communication. The following will be described separately from two aspects: CP group communication and SP communication: When performing CP group communication, specifically, based on the configuration information about the CP group between different computing nodes, the allgather function can be used to perform CP group communication on the second sub-key-value features of different computing nodes. It should be noted that allgather is a communication primitive, and its core function is to aggregate the relevant data originally distributed on each computing unit to all computing units, and finally each computing unit will have a complete data copy. Of course, the above introduction is only a brief description of allgather, and the technical details of allgather can be specifically referred to, which will not be elaborated here.

[0032] In a specific implementation scenario, CP group communication occurs between computing nodes, and the configuration information of the CP group can include the pairs of computing units that communicate between different computing nodes. Still taking Figure 2b the situation shown as an example, the configuration information of the CP group can include: (Node0GPU0, Node1GPU0), (Node0GPU1, Node1GPU1), that is, during CP group communication, communication occurs between computing unit 0 in computing node 0 and computing unit 0 in computing node 1, and communication occurs between computing unit 1 in computing node 0 and computing unit 1 in computing node 1. Of course, the above example is only a possible example in the actual application process, and other possible situations will not be exemplified one by one here.

[0033] In a specific implementation scenario, please refer to Figure 2c , Figure 2c which is a schematic diagram of an embodiment of CP group communication in this application. Computing node 0 (i.e., Figure 2c Node0 in Figure 2c ) contains 8 computing units (such as, 8 GPUs), which are respectively abbreviated as 0 to 7 in Figure 2c , and computing node 1 (i.e., Figure 2cThey are briefly recorded as 0 to 7 respectively. In addition, the first computing unit in computing node 0 contains the second sub-key value feature e00, the second computing unit in computing node 0 contains the second sub-key value feature e01, the third computing unit in computing node 0 contains the second sub-key value feature e02, the fourth computing unit in computing node 0 contains the second sub-key value feature e03, the fifth computing unit in computing node 0 contains the second sub-key value feature e04, the sixth computing unit in computing node 0 contains the second sub-key value feature e05, the seventh computing unit in computing node 0 contains the second sub-key value feature e06, and the eighth computing unit in computing node 0 contains the second sub-key value feature e07; similarly, the first computing unit in computing node 1 contains the second sub-key value feature e10, the second computing unit in computing node 1 contains the second sub-key value feature e11, the third computing unit in computing node 1 contains the second sub-key value feature e12, the fourth computing unit in computing node 1 contains the second sub-key value feature e13, the fifth computing unit in computing node 1 contains the second sub-key value feature e14, the sixth computing unit in computing node 1 contains the second sub-key value feature e15, the seventh computing unit in computing node 1 contains the second sub-key value feature e16, and the eighth computing unit in computing node 1 contains the second sub-key value feature e17. On this basis, if the configuration information of the CP group includes: (Node0GPU0, Node1GPU0), (Node0GPU1, Node1GPU1), (Node0GPU2, Node1GPU2), (Node0GPU3, Node1GPU3), (Node0GPU4, Node1GPU4), (Node0GPU5, Node1GPU5), (Node0GPU6, Node1GPU6), (Node0GPU7, Node1GPU7), then after executing the allgather function based on the above configuration information of the CP group, the data owned by computing unit 0 in computing node 0 (i.e., Figure 2c Node0 in Figure 2c is the second sub-key value feature e00 and the second sub-key value feature e10, and the data owned by computing unit 1 in computing node 0 (i.e., Figure 2c Node0 in Figure 2c is the second sub-key value feature e01 and the second sub-key value feature e11, and the data owned by computing unit 2 in computing node 0 (i.e., Figure 2c Node0 in Figure 2cIn computing node 0 (i.e., Node0), the data owned by computing unit 5 are the second sub-key value features e05 and e15, and the data owned by computing unit 6 are the second sub-key value features e06 and e16, and the data owned by computing unit 7 are the second sub-key value features e07 and e17. Similarly, in computing node 1 (i.e., Node1), the data owned by computing unit 0 are the second sub-key value features e10 and e00, the data owned by computing unit 1 are the second sub-key value features e11 and e01, the data owned by computing unit 2 are the second sub-key value features e12 and e02, the data owned by computing unit 3 are the second sub-key value features e13 and e03, the data owned by computing unit 4 are the second sub-key value features e14 and e04, the data owned by computing unit 5 are the second sub-key value features e15 and e05, the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 0 (i.e., Node0), the data owned by computing unit 6 are the second sub-key value features e06 and e16, and the data owned by computing unit 7 are the second sub-key value features e07 and e17. Similarly, in computing node 1 (i.e., Node1), the data owned by computing unit 0 are the second sub-key value features e10 and e00, the data owned by computing unit 1 are the second sub-key value features e11 and e01, the data owned by computing unit 2 are the second sub-key value features e12 and e02, the data owned by computing unit 3 are the second sub-key value features e13 and e03, the data owned by computing unit 4 are the second sub-key value features e14 and e04, the data owned by computing unit 5 are the second sub-key value features e15 and e05, the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 0 (i.e., Node0), the data owned by computing unit 7 are the second sub-key value features e07 and e17; similarly, in computing node 1 (i.e., Node1), the data owned by computing unit 0 are the second sub-key value features e10 and e00, the data owned by computing unit 1 are the second sub-key value features e11 and e01, the data owned by computing unit 2 are the second sub-key value features e12 and e02, the data owned by computing unit 3 are the second sub-key value features e13 and e03, the data owned by computing unit 4 are the second sub-key value features e14 and e04, the data owned by computing unit 5 are the second sub-key value features e15 and e05, the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 0 are the second sub-key value features e10 and e00, and the data owned by computing unit 1 are the second sub-key value features e11 and e01, and the data owned by computing unit 2 are the second sub-key value features e12 and e02, and the data owned by computing unit 3 are the second sub-key value features e13 and e03, and the data owned by computing unit 4 are the second sub-key value features e14 and e04, and the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 1 are the second sub-key value features e11 and e01, and the data owned by computing unit 2 are the second sub-key value features e12 and e02, and the data owned by computing unit 3 are the second sub-key value features e13 and e03, and the data owned by computing unit 4 are the second sub-key value features e14 and e04, and the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 2 are the second sub-key value features e12 and e02, and the data owned by computing unit 3 are the second sub-key value features e13 and e03, and the data owned by computing unit 4 are the second sub-key value features e14 and e04, and the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 3 are the second sub-key value features e13 and e03, and the data owned by computing unit 4 are the second sub-key value features e14 and e04, and the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 4 are the second sub-key value features e14 and e04, and the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 5 are the second sub-key value features e15 and e05, and the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 6 are the second sub-key value features e16 and e06, and the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c In computing node 1 (i.e., Node1), the data owned by computing unit 7 are the second sub-key value features e17 and e07. That is to say, the communication volume of the CP group communication is 16 * G (where G is the data volume of each second sub-key value feature). Of course, Figure 2c The above is only a possible example of the CP group communication in the actual application process. Other possible situations are not limited here and will not be listed one by one.

[0034] When performing SP group communication, specifically, based on the configuration information of the SP group within the same computing node, the allgather function can be used to perform SP group communication on the second sub-key value features of the same computing node. It should be noted that different from the allgather operation during the aforementioned CP group communication, which occurs between different computing nodes, the allgather operation during SP group communication occurs within the same computing node.

[0035] In a specific implementation scenario, SP group communication occurs within a computing node. The configuration information of the SP group may include pairs of computing units that communicate between the same computing nodes. Still taking Figure 2b the situation shown as an example, the configuration information of the CP group may include: (Node0GPU0, Node0GPU1), (Node1GPU0, Node1GPU1). That is, during CP group communication, communication occurs between computing unit 0 and computing unit 1 in computing node 0, and between computing unit 0 and computing unit 1 in computing node 1. Of course, the above example is only a possible example in the actual application process, and other possible situations will not be exemplified one by one here.

[0036] In a specific implementation scenario, please refer to Figure 2d , Figure 2d which is a schematic diagram of an embodiment of SP group communication in this application. Computing node 0 (i.e., Figure 2d Node0 in Figure 2d ) contains 8 computing units (such as 8 GPUs), which are abbreviated as 0 to 7 respectively in Figure 2d . Computing node 1 (i.e., Figure 2d Node1 in Figure 2d ) contains 8 computing units (such as 8 GPUs), which are abbreviated as 0 to 7 respectively in Figure 2dIn Node1), the configuration information about the SP group includes: (Node1GPU0, Node1GPU1), (Node1GPU1, Node1GPU2), (Node1GPU2, Node1GPU3), (Node1GPU3, Node1GPU4), (Node1GPU4, Node1GPU5), (Node1GPU5, Node1GPU6), (Node1GPU6, Node1GPU7), (Node1GPU7, Node1GPU0). After executing the allgather function based on the above configuration information of the SP group, computing node 0 (i.e., Figure 2d In Node0) and computing node 1 (i.e., Figure 2d In Node1), the data owned by each computing unit is e00~e17 (i.e., the second key value feature after aggregating all the second sub-key value features). It should be noted that in this example, if the SP group communication is executed first and then the CP group communication is executed, then after the SP communication is executed first, the data owned by each computing unit in computing node 0 and computing node 1 is 8*G (where G is the data volume of each second sub-key value feature). Therefore, when the CP group communication is executed again, the communication volume of the CP group will reach 16*8*G (where G is the data volume of each second sub-key value feature), which is much larger than the communication volume of the aforementioned CP group 8*G (where G is the data volume of each second sub-key value feature). Therefore, in the embodiments of the present disclosure, executing the CP group communication first and then the SP group communication can significantly reduce the communication volume between computing nodes.

[0037] In one implementation scenario, in order to facilitate the processing of the second sub-key value features when executing the CP group communication first and then the SP group communication, the unsqueeze operation can be performed on the second sub-key value features first, then the CP group communication is performed on the second sub-key value features after the unsqueeze operation is executed, and finally the SP group communication is performed on the second sub-key value features after the CP group communication is executed, so that each computing unit in different computing nodes can obtain the complete second key value feature. For the sake of description, the above process can be expressed as:

[0038]

[0039]

[0040] In the above formula, represents the second sub-key value feature, represents the second sub-key value feature after the unsqueeze operation is executed, represents the allgather function, and CP_GROUP represents the configuration information of the CP group. represents the second sub-key value feature after performing CP communication, and SP_GROUP represents the configuration information of the SP group. represents the key value feature after performing SP group communication. As a possible example, after successively performing CP group communication and SP group communication, each computing unit in different computing nodes can obtain a third sub-key value feature (as described above ). It can be characterized as a feature tensor whose feature size is the product of the batch size, the number of CP groups, the sequence length divided by the number of SP groups, and the hidden layer size, that is, it can be expressed as (B, CP_size, S / SP_size, H). For the specific meanings of each parameter, please refer to the relevant descriptions above and will not be elaborated here. On this basis, the third sub-key value feature can be reshaped in the target dimension to obtain the second key value feature. For example, the CP dimension of the third sub-key value feature can be flattened on the SP dimension through a reshape operation. It should be noted that the target dimension includes the feature dimension of the number of CP groups and the feature dimension of the sequence length divided by the number of SP groups, and the second key value feature is characterized as a feature tensor whose feature size is the product of the batch size, the sequence length, and the hidden layer size (i.e., B*S*H). For ease of description, the second key value feature can be expressed as:

[0041] In the above formula, represents the third sub-key value feature, and RESHAPE represents the reshaping operation. represents the second key value feature.

[0042] Step S15: Based on the second key value feature, obtain the target key feature and target value feature of the multi-head attention mechanism.

[0043] Specifically, after obtaining the second key value feature, based on the feature dimension of the rotary position embedding, the third key value feature representing low-rank compression and the first key feature representing the rotary position embedding can be sliced from the second key value feature. Then, based on the third key value feature, decompression is performed to obtain the second key feature and the target value feature of the multi-head attention mechanism. Finally, based on the first key feature and the second key feature, fusion is performed to obtain the target key feature of the multi-head attention mechanism. It should be noted that the feature dimension of the rotary position embedding represents the feature dimension after the key feature vector performs rotary dimension embedding in the MLA. Exemplarily, taking the second key value feature expressed as (B, S, 576) as an example (i.e., the hidden layer size H is 576), if the feature dimension of the rotary position embedding is 64, it indicates that 64 dimensions in the feature dimension of the hidden layer size in the second key value feature belong to the first key feature. Of course, the above example is only a possible situation in the actual application process, and other possible situations are not limited here and will not be exemplified one by one.

[0044] In a specific implementation scenario, based on the feature dimension of the rotary position embedding, the feature dimension of the second key-value feature in the hidden layer size can be sliced to obtain a third key-value feature and a first key feature. It should be noted that the feature dimension of the rotary position embedding is the same as the feature dimension of the first key feature in the hidden layer size. Still taking the second key-value feature expressed as (B, S, 576) as an example, if the feature dimension of the rotary position embedding is 64, then a 64-dimensional first key feature (B, S, 64) can be sliced from the feature dimension of the second key-value feature in the hidden layer size. For example, the last 64 dimensions of the feature dimension of the second key-value feature in the hidden layer size can be considered as belonging to the first key feature. On this basis, the remaining part can be regarded as the third key-value feature (B, S, 512) after low-rank compression. Of course, the above example is only one possible situation in the actual application process, and other possible situations are not limited here, nor will they be listed one by one. For the sake of convenience of description, the third key-value feature and the first key feature can be expressed as:

[0045] In the above formula, represents the second key-value feature, split represents slicing, K_rope represents the first key feature, represents the third key-value feature.

[0046] In a specific implementation scenario, after obtaining the third key-value feature after low-rank compression, it can be decompressed to separate K and V. Specifically, the third key-value feature can be processed based on the key-value decompression parameter to obtain a decompressed key-value feature, and then the key-value is split based on the decompressed key-value feature to obtain a second key feature and a target value feature. For the sake of convenience of description, the decompression operation on the key-value feature can be denoted as kv_up, such as it can be implemented through the decompression parameter w UKV For specific implementation details, please refer to the technical details of MLA, which will not be elaborated here. On this basis, the specific process of decompressing the third key-value feature to obtain the second key feature and the target value feature can be expressed as:

[0047] In the above formula, represents the third key-value feature, split represents key-value splitting, K_nope represents the second key feature, and V represents the target value feature. In addition, for the convenience of subsequent multi-head attention mechanism processing, the feature dimension of the value feature obtained after key-value splitting in the hidden layer size can be expressed in the form of n*d2, where n represents the number of multi-heads, d2 represents the target dimension, and the product of the number of multi-heads n and the target dimension d2 is the same as the feature dimension of the value feature obtained after key-value splitting in the hidden layer size.

[0048] In a specific implementation scenario, after obtaining the first key feature and the second key feature, the two can be fused to incorporate position encoding information into the key feature information. Specifically, a replication operation can be performed on the feature dimension of the first key feature in the hidden layer size based on the number of heads of the multi-head attention mechanism to obtain a third key feature, and the second key feature can be reshaped into the number of heads multiplied by the target dimension in the feature dimension of the hidden layer size to obtain a fourth key feature. It should be noted that the number of heads multiplied by the target dimension is equal to the number of feature dimensions of the second key feature in the feature dimension of the hidden layer size. Taking the number of heads denoted as n as an example, if the number of feature dimensions of the first key feature in the hidden layer size is 64, then the number of feature dimensions of the third key feature in the hidden layer size is n * 64. In addition, if the number of feature dimensions of the second key feature in the hidden layer size is H, it can be reshaped into n * d3, where d3 is the target dimension, and the number of heads n multiplied by the target dimension d3 is equal to the number of feature dimensions H of the second key feature in the hidden layer size. On this basis, the target key feature of the multi-head attention mechanism can be obtained by concatenating the third key feature and the fourth key feature in the feature dimension of the hidden layer size. Still taking the above example, the number of feature dimensions of the target key feature obtained by concatenating in the feature dimension of the hidden layer size can be expressed as n * (d3 + 64). It should be noted that the number of feature dimensions of other dimensions (i.e., batch size, sequence length) can still be expressed as B and S respectively. Of course, the above examples are only several possible examples in the actual application process, and other possible situations are not limited here, nor will they be listed one by one. For ease of description, the third key feature can be expressed as: HK_rope = K_rope.repeat(num.heads) In the above formula, num.heads represents the number of heads, repeat represents the replication operation, K_rope represents the first key feature, and HK_rope represents the third key feature. On this basis, the target key feature of the multi-head attention mechanism can be expressed as: K = concate(K_nope, HK_rope) In the above formula, HK_rope represents the third key feature, K_nope represents the fourth key feature, and K represents the target key feature.

[0049] Step S16: Execute the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing.

[0050] In an implementation scenario, after obtaining the target query feature, the target key feature, and the target value feature of the multi-head attention mechanism, the multi-head attention mechanism can be executed accordingly to complete the MLA processing. For ease of description, the output result of the multi-head attention mechanism can be expressed as: Atten_out = MHA(Q, K, V) In the above formula, Atten_out represents the output result of the multi-head attention mechanism, MHA represents the multi-head attention mechanism, and Q, K, and V represent the target query feature, the target key feature, and the target value feature respectively. For the specific processing process, refer to the technical details of the multi-head attention mechanism MHA, which will not be elaborated here.

[0051] In an implementation scenario, after completing the MLA, decoding can be performed based on the output features after MLA processing (such as the output result Atten_out of the aforementioned multi-head attention mechanism) to obtain the response result of the text sequence to be responded. It should be noted that the response result can include but is not limited to at least one data type among text, image, video, audio, and table, which is not limited here. For example, when the text sequence to be responded is the text sequence "Please explain a certain knowledge point" in an educational scenario, its response result can be the text sequence "The meaning of a certain knowledge point is..., for example,..."; or, when the text sequence to be responded is the text sequence "What are the indications of a certain drug" in a medical scenario, its response result can be the text sequence "The indications of a certain drug can include..., but it should be noted especially that it does not include...". Of course, the above examples are only several possible examples in the actual application process, and other possible situations are not limited here, nor will they be listed one by one. It should be noted that for the specific process of decoding based on the output features, refer to the technical details of decoding techniques such as autoregressive decoding, and the decoding process will not be elaborated here. Please refer to Figure 2e , Figure 2e is a schematic diagram of the effect of an embodiment of the sequence processing method based on MLA in this application. As Figure 2e shown, for test samples with a sequence length of 32k under the 73B MoE (i.e., Mixture of Experts) model, if the parallel configuration of CP2 and SP8 and the hardware configuration of 64 NPU910B cards are adopted, the wps (words per second) index ( Figure 2e the blue line in Figure 2e ) of the processing method of the disclosed embodiment of this application is significantly improved compared to the wps index (

[0052] In the above solution, the first query feature based on low-rank compression is rotated and position-embedded after compression to obtain the target query feature of the multi-head attention mechanism. The first query feature is obtained from the text sequence to be responded to, and the first key-value feature based on low-rank compression is segmented to obtain the first sub-key-value features to be processed by different computing units in each computing node. Then, different computing units in each computing node perform rotation and position embedding on the allocated first sub-key-value features to obtain the second sub-key-value features. Thus, communication is performed based on each computing node to obtain the second key-value feature after the first key-value feature is rotated and position-embedded. Based on the second key-value feature, the target key feature and the target value feature of the multi-head attention mechanism are obtained. Furthermore, the multi-head attention mechanism is executed based on the target query feature, the target key feature, and the target value feature to complete the MLA processing. Since the parallel aggregation operation is separated from the multi-head attention MHA and the parallel aggregation operation is pre-placed before the multi-head attention MHA after the first key-value feature based on low-rank compression is obtained, only the first key-value feature based on low-rank compression needs to be communicated at this time. Compared with performing the parallel aggregation operation in the multi-head attention MHA, the communication volume of the computing nodes can be greatly reduced. Therefore, the communication volume of the computing nodes can be reduced when performing sequence processing based on MLA, so as to improve the final processing efficiency of the computing nodes.

[0053] Please refer to Figure 3 , Figure 3 FIG. is a schematic framework diagram of an embodiment of the sequence processing device based on MLA of the present application. The sequence processing device 30 based on MLA includes: a position embedding module 31, a grouping and segmentation module 32, a distributed processing module 33, a sequence aggregation module 34, a feature processing module 35, and a multi-head processing module 36. The position embedding module 31 is configured to perform rotation and position embedding on the first query feature based on low-rank compression after decompression to obtain the target query feature of the multi-head attention mechanism; wherein, the first query feature is obtained from the text sequence to be responded to; the grouping and segmentation module 32 is configured to segment the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node; the distributed processing module 33 is configured to perform rotation and position embedding on the allocated first sub-key-value features by different computing units in each computing node to obtain the second sub-key-value features; the sequence aggregation module 34 is configured to perform communication based on each computing node to obtain the second key-value feature after the first key-value feature is rotated and position-embedded; the feature processing module 35 is configured to obtain the target key feature and the target value feature of the multi-head attention mechanism based on the second key-value feature; the multi-head processing module 36 is configured to execute the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing.

[0054] In the above solution, the sequence processing device 30 based on MLA performs rotational position embedding on the first query feature based on low-rank compression after compression to obtain the target query feature of the multi-head attention mechanism. The first query feature is obtained from the text sequence to be responded to, and the first key-value feature based on low-rank compression is segmented to obtain the first sub-key-value features to be processed by different computing units in each computing node. Then, different computing units in each computing node perform rotational position embedding on the allocated first sub-key-value features to obtain the second sub-key-value features. Thus, based on each computing node for communication, the second key-value feature after rotational position embedding of the first key-value feature is obtained, and based on the second key-value feature, the target key feature and the target value feature of the multi-head attention mechanism are obtained. Furthermore, based on the target query feature, the target key feature, and the target value feature, the multi-head attention mechanism is executed to complete the MLA processing. Since the parallel aggregation operation is separated from the multi-head attention MHA and the parallel aggregation operation is pre-positioned before the multi-head attention MHA after obtaining the first key-value feature based on low-rank compression, only the first key-value feature based on low-rank compression needs to be communicated at this time. Compared with performing the parallel aggregation operation in the multi-head attention MHA, the communication volume of the computing nodes can be greatly reduced. Therefore, the communication volume of the computing nodes can be reduced when performing sequence processing based on MLA, so as to improve the final processing efficiency of the computing nodes.

[0055] In some disclosed embodiments, the grouping and segmentation module 32 is specifically configured to sequentially segment the first value feature based on low-rank compression according to the number of CP groups and the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node. The sequence aggregation module 34 is specifically configured to perform CP group communication between different computing nodes and then perform SP group communication within the same computing node to obtain the second key-value feature after rotational position embedding of the first key-value feature.

[0056] In some disclosed embodiments, the first key-value feature is characterized as a feature tensor whose feature size is the product of the batch size, the sequence length, and the hidden layer size. The grouping and segmentation module 32 includes a first segmentation sub-module, which is configured to perform a first segmentation on the first key-value feature in the feature dimension of the sequence length according to the number of CP groups for subsequent distribution to different computing nodes after the final segmentation is completed. The grouping and segmentation module 32 includes a second segmentation sub-module, which is configured to perform a second segmentation on each sub-feature of the first key-value feature after the first segmentation in the feature dimension of the sequence length to obtain the first sub-key-value features to be processed by different computing units in each computing node.

[0057] In some disclosed embodiments, when performing CP group communication between different computing nodes, the sequence aggregation module 34 is specifically configured to perform CP group communication on the second sub-key-value features of different computing nodes by using the allgather function based on the configuration information about the CP groups between different computing nodes.

[0058] In some disclosed embodiments, when the sequence aggregation module 34 performs SP group communication within the same computing node, it is specifically configured to perform SP group communication on the second sub-key value features of the same computing node by using the allgather function based on the configuration information about the SP group within the same computing node.

[0059] In some disclosed embodiments, after successively performing CP group communication and SP group communication, each computing unit in different computing nodes obtains third sub-key value features. The third sub-key value features are characterized as a feature tensor whose feature size is the product of the batch size, the number of CP groups, the sequence length divided by the number of SP groups, and the hidden layer size. The sequence processing device 30 based on MLA further includes a feature reshaping module, which is configured to reshape based on the third sub-key value features in the target dimension to obtain second key value features; wherein, the target dimension includes the feature dimension of the number of CP groups and the feature dimension of the sequence length divided by the number of SP groups, and the second key value features are characterized as a feature tensor whose feature size is the product of the batch size, the sequence length, and the hidden layer size.

[0060] In some disclosed embodiments, the feature processing module 35 includes a slicing sub-module, which is configured to slice from the second key value features a third key value feature representing low-rank compression and a first key feature representing rotated position embedding based on the feature dimension of the rotated position embedding; the feature processing module 35 includes a decompression sub-module, which is configured to decompress based on the third key value features to obtain second key features and target value features of the multi-head attention mechanism; the feature processing module 35 includes a fusion sub-module, which is configured to fuse the first key features and the second key features to obtain target key features.

[0061] In some disclosed embodiments, the slicing sub-module is specifically configured to slice the second key value features in the feature dimension of the hidden layer size based on the feature dimension of the rotated position embedding to obtain third key value features and first key features; wherein, the feature dimension of the rotated position embedding is the same as the feature dimension of the first key features in the feature dimension of the hidden layer size.

[0062] In some disclosed embodiments, the decompression sub-module includes a processing unit, which is configured to process the third key value features based on the key value decompression parameters to obtain decompressed key value features; the decompression sub-module includes a splitting unit, which is configured to perform key value splitting based on the decompressed key value features to obtain second key features and target value features.

[0063] In some disclosed embodiments, the fusion sub-module includes a replication unit configured to perform a replication operation on the feature dimension of the first key feature in the hidden layer size based on the number of heads of the multi-head attention mechanism to obtain a third key feature; the fusion sub-module includes a reshaping unit configured to reshape the feature dimension of the second key feature in the hidden layer size into the number of heads multiplied by the target dimension to obtain a fourth key feature; wherein, the number of heads multiplied by the target dimension is equal to the number of feature dimensions of the second key feature in the feature dimension of the hidden layer size; the fusion sub-module includes a splicing unit configured to splice the third key feature and the fourth key feature in the feature dimension of the hidden layer size to obtain a target key feature.

[0064] In some disclosed embodiments, the sequence processing device 30 based on MLA further includes a feature decoding module configured to, after performing the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing, decode based on the output feature after the MLA processing to obtain a response result of the text sequence to be responded; wherein, the response result includes at least one data type of text, image, video, audio, and table.

[0065] Please refer to Figure 4 , Figure 4 FIG. is a schematic framework diagram of an embodiment of the sequence processing system of the present application. The sequence processing system 40 includes a plurality of computing nodes 41, and each computing node 41 contains a number of computing units 411. The sequence processing system 40 is configured to execute program instructions to implement the steps in any of the foregoing embodiments of the sequence processing method based on MLA. Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the computing unit 411 may include, but is not limited to, GPU, NPU, CPU, etc., and the specific type of the computing unit 411 is not limited herein. In addition, each computing node 41 may adopt a distributed architecture; or, each computing node 41 may also adopt a centralized architecture. The deployment manner of each computing node 41 is not limited herein. It should be noted that although Figure 4 only three computing nodes 41 are schematically drawn, in actual application, the specific number of the plurality of computing nodes 41 in the sequence processing system 40 may be set according to actual application needs. For example, in the case of relatively light computing load, the specific number of the plurality of computing nodes 41 in the sequence processing system 40 may be relatively small, such as 2, 3, etc., or in the case of relatively heavy computing load, the specific number of the plurality of computing nodes 41 in the sequence processing system 40 may be relatively large, such as 7, 8, etc., and the specific number of the plurality of computing nodes 40 in the sequence processing system 40 is not limited herein. It can be understood that the specific number of the plurality of computing units in each computing node 41 may also be set with reference to the foregoing computing load and other factors such as parallel efficiency, which will not be elaborated herein.

[0066] In the above solution, the sequence processing system 40 performs rotation position embedding on the first query feature based on low-rank compression after compression to obtain the target query feature of the multi-head attention mechanism. The first query feature is obtained from the text sequence to be responded, and is segmented based on the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node. Then, different computing units in each computing node perform rotation position embedding on the allocated first sub-key-value features to obtain the second sub-key-value features. Thus, communication is performed based on each computing node to obtain the second key-value feature with rotation position embedding of the first key-value feature, and based on the second key-value feature, the target key feature and the target value feature of the multi-head attention mechanism are obtained. Furthermore, the multi-head attention mechanism is executed based on the target query feature, the target key feature, and the target value feature to complete the MLA processing. Since the parallel aggregation operation is separated from the multi-head attention MHA and the parallel aggregation operation is preposed before the multi-head attention MHA after obtaining the first key-value feature based on low-rank compression, only the first key-value feature based on low-rank compression needs to be communicated at this time. Compared with performing the parallel aggregation operation in the multi-head attention MHA, the communication volume of the computing nodes can be greatly reduced. Therefore, the communication volume of the computing nodes can be reduced when performing sequence processing based on MLA to improve the final processing efficiency of the computing nodes.

[0067] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above-described method embodiments for sequence processing based on MLA.

[0068] In the above solution, the computer-readable storage medium 50 performs rotational position embedding on the first query feature based on low-rank compression after compression to obtain the target query feature of the multi-head attention mechanism. The first query feature is obtained from the text sequence to be responded to, and is segmented based on the first key-value feature of low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node. Then, different computing units in each computing node perform rotational position embedding on the assigned first sub-key-value features to obtain the second sub-key-value features. Thus, communication is performed based on each computing node to obtain the second key-value feature with rotational position embedding of the first key-value feature, and based on the second key-value feature, the target key feature and target value feature of the multi-head attention mechanism are obtained. Furthermore, the multi-head attention mechanism is executed based on the target query feature, target key feature, and target value feature to complete the MLA processing. Since the parallel aggregation operation is separated from the multi-head attention MHA and the parallel aggregation operation is pre-positioned before the multi-head attention MHA after obtaining the first key-value feature of low-rank compression, only the first key-value feature of low-rank compression needs to be communicated at this time. Compared with performing the parallel aggregation operation in the multi-head attention MHA, the communication volume of the computing nodes can be greatly reduced. Therefore, the communication volume of the computing nodes can be reduced when performing sequence processing based on MLA to improve the final processing efficiency of the computing nodes.

[0069] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.

[0070] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0071] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0072] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0073] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0074] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0075] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A sequence processing method based on MLA, characterized in that, Including: After decompressing the first query feature based on low-rank compression, perform rotational position embedding to obtain the target query feature of the multi-head attention mechanism; wherein, the first query feature is obtained from the text sequence to be responded to. Slice the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node. Perform rotational position embedding on the first sub-key-value features assigned to be processed by different computing units in each computing node respectively to obtain the second sub-key-value features. Based on communication among the computing nodes, obtain the second key-value feature after the rotational position embedding of the first key-value feature. Based on the second key-value feature, obtain the target key feature and the target value feature of the multi-head attention mechanism. Execute the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing.

2. The method according to claim 1, characterized in that, The slicing of the first key-value feature based on low-rank compression to obtain the first sub-key-value features to be processed by different computing units in each computing node includes: Slice the first key-value feature based on low-rank compression successively according to the number of CP groups and the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node. The communication based on the computing nodes to obtain the second key-value feature after the rotational position embedding of the first key-value feature includes: Based on CP group communication among different computing nodes and then SP group communication within the same computing node, obtain the second key-value feature after the rotational position embedding of the first key-value feature.

3. The method according to claim 2, wherein The first key-value feature is characterized as a feature tensor with a feature size being the product of the batch size, the sequence length, and the hidden layer size. The slicing of the first key-value feature based on low-rank compression successively according to the number of CP groups and the number of SP groups to obtain the first sub-key-value features to be processed by different computing units in each computing node includes: Based on the number of CP groups, perform a first slice on the first key-value feature in the feature dimension of the sequence length for subsequent distribution to different computing nodes after the final slicing is completed. Based on the number of SP groups, perform a second slice on each sub-feature of the first key-value feature after the first slice in the feature dimension of the sequence length to obtain the first sub-key-value features to be processed by different computing units in each computing node.

4. The method according to claim 2, characterized in that, The CP group communication based on different computing nodes includes: Based on the configuration information about CP groups among different computing nodes, use the allgather function to perform CP group communication on the second sub-key-value features of different computing nodes. And / or, the SP group communication within the same computing node includes: Based on the configuration information about SP groups within the same computing node, use the allgather function to perform SP group communication on the second sub-key-value features of the same computing node.

5. The method according to claim 2, wherein After successively performing the CP group communication and the SP group communication, each of the computing units in different computing nodes obtains a third sub-key value feature, where the third sub-key value feature is characterized as a feature tensor with a feature size being the product of the batch size, the number of CP groups, the sequence length divided by the number of SP groups, and the hidden layer size. The method further includes: Reshaping based on the third sub-key value feature in a target dimension to obtain the second key value feature; where the target dimension includes the feature dimension of the number of CP groups and the feature dimension of the sequence length divided by the number of SP groups, and the second key value feature is characterized as a feature tensor with a feature size being the product of the batch size, the sequence length, and the hidden layer size.

6. The method according to claim 1, characterized in that, The obtaining of the target key feature and the target value feature of the multi-head attention mechanism based on the second key value feature includes: Based on the feature dimension of the rotational position embedding, splitting from the second key value feature a third key value feature representing low-rank compression and a first key feature representing the rotational position embedding; Decompressing based on the third key value feature to obtain a second key feature and the target value feature of the multi-head attention mechanism; Fusing based on the first key feature and the second key feature to obtain the target key feature.

7. The method according to claim 6, wherein The splitting from the second key value feature a third key value feature representing low-rank compression and a first key feature representing the rotational position embedding based on the feature dimension of the rotational position embedding includes: Based on the feature dimension of the rotational position embedding, splitting the second key value feature in the feature dimension of the hidden layer size to obtain the third key value feature and the first key feature; where the feature dimension of the rotational position embedding is the same as the feature dimension of the first key feature in the feature dimension of the hidden layer size.

8. The method according to claim 6, characterized in that The decompressing based on the third key value feature to obtain a second key feature and the target value feature of the multi-head attention mechanism includes: Processing the third key value feature based on key value decompression parameters to obtain a decompressed key value feature; Performing key value splitting based on the decompressed key value feature to obtain the second key feature and the target value feature.

9. The method according to claim 6, characterized in that, The fusing based on the first key feature and the second key feature to obtain the target key feature includes: Performing a replication operation on the first key feature in the feature dimension of the hidden layer size based on the number of heads of the multi-head attention mechanism to obtain a third key feature, and reshaping the second key feature in the feature dimension of the hidden layer size to the number of heads multiplied by the target dimension to obtain a fourth key feature; where the number of heads multiplied by the target dimension is equal to the feature dimension of the second key feature in the feature dimension of the hidden layer size; Performing splicing on the third key feature and the fourth key feature in the feature dimension of the hidden layer size to obtain the target key feature.

10. The method according to claim 1, characterized in that, After performing the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing, the method further includes: Decode based on the output features after the MLA processing to obtain the response result of the text sequence to be responded; wherein, the response result includes at least one data type of text, image, video, audio, and table.

11. A sequence processing device based on MLA, characterized in that, Comprising: A position embedding module, configured to perform rotational position embedding on the decompressed first query feature based on low-rank compression to obtain the target query feature of the multi-head attention mechanism; wherein, the first query feature is obtained from the text sequence to be responded. A grouping and splitting module, configured to split the first key-value feature based on low-rank compression to obtain the first sub key-value features to be processed by different computing units in each computing node. A distribution processing module, configured to respectively perform rotational position embedding on the allocated and processed first sub key-value features by different computing units in each computing node to obtain the second sub key-value features. A sequence aggregation module, configured to communicate based on each computing node to obtain the second key-value feature of the first key-value feature after the rotational position embedding. A feature processing module, configured to obtain the target key feature and the target value feature of the multi-head attention mechanism based on the second key-value feature. A multi-head processing module, configured to execute the multi-head attention mechanism based on the target query feature, the target key feature, and the target value feature to complete the MLA processing.

12. A sequence processing system, characterized in that, The sequence processing system includes a plurality of computing nodes, each of the computing nodes respectively includes a plurality of computing units, and the sequence processing system is configured to execute program instructions to implement the MLA-based sequence processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Store program instructions that can be run by a processor, and the program instructions are configured to implement the MLA-based sequence processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Parallel processing method and device based on sequence model

    CN116128021A

  • Calculation method and device of neural network model, electronic equipment and storage medium

    CN117273084A

  • Sequence parallelization method and device for linear attention, equipment and medium

    CN118132155A

  • Data processing method and device, equipment and storage medium

    CN118227671A

  • Neural network model compression method and device, storage medium and program product

    CN118504643A