Method for cooperatively accelerating Transform algorithm based on CPU + NPU

By co-accelerating between CPU and NPU, distributing various module tasks in the Transformer algorithm, and processing CPU and NPU calculations in parallel, the Transformer algorithm's performance degradation in large-scale data processing and real-time processing is solved, and efficient inference operation performance and big data processing support are achieved.

CN119990189APending Publication Date: 2025-05-13西安翔腾微电子科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411750018.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When processing large-scale data sets or real-time processing requirements, the Transformer algorithm has high computational complexity, resulting in a degradation inference performance.

Method used

By synergistically accelerating between the CPU and NPU, each module task in the Transformer algorithm is finely distributed, and the CPU calculation and NPU calculation are processed in parallel by taking advantage of the respective advantages of the CPU and NPU.

Benefits of technology

It effectively improves the inference operation performance of the Transformer algorithm, reduces dependence on a single hardware resource, and supports big data processing and complex model inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990189A_ABST
    Figure CN119990189A_ABST
Patent Text Reader

Abstract

The invention relates to a method for cooperatively accelerating a Transform algorithm based on CPU + NPU. The method comprises the following steps: 1) determining the core number M (N is gt; 0) and the number N of cores of the NPU (N is gt; 0 is any positive integer); 2) determining the calculation attribution of each module task in the Transform algorithm; 3) calculating each module according to the distribution of the calculation tasks in the step 2), and combining results through a CPU calculation residual layer; and 4) after the calculation tasks in the step 2) are distributed, CPU calculation and NPU calculation are processed in parallel in a parallel pipeline processing mode. According to the method, through fine task distribution and cooperative calculation, respective advantages of the CPU and the NPU are fully utilized, and efficient acceleration of the Transform algorithm is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to a method for accelerating a Transformer algorithm based on a CPU (central processing unit) + a NPU (neural network processor) collaboration. Background Art

[0002] Since its advent, the Transformer model has completely revolutionized the research and application of deep learning with its outstanding ability to process sequence data, especially in the field of natural language processing, and has become the core of modern deep learning applications.

[0003] In practical applications, especially when facing large-scale data sets or real-time processing needs, the high computational complexity of the Transformer algorithm has become a bottleneck that cannot be ignored. As the amount of processed data increases, the required computing resources and memory will increase dramatically, resulting in a decrease in the inference performance of the Transformer algorithm. Summary of the invention

[0004] In order to solve the technical problems existing in the background technology, the present invention provides a method for accelerating the Transformer algorithm based on CPU+NPU collaboration. Through sophisticated task distribution and collaborative computing, the present invention fully utilizes the respective advantages of CPU and NPU to achieve efficient acceleration of the Transformer algorithm.

[0005] The technical solution of the present invention is: the present invention is a method for accelerating the Transformer algorithm based on CPU+NPU collaboration, characterized in that the method comprises the following steps:

[0006] 1) Determine the number of CPU cores M (N is any positive integer > 0) and the number of NPU cores N (N is any positive integer > 0);

[0007] 2) Determine the computational ownership of each module task in the Transformer algorithm;

[0008] 3) Calculate each module according to the distribution of computing tasks in step 2), and merge the results through the CPU calculation residual layer;

[0009] 4) After the computing tasks are distributed according to step 2), the CPU computing and the NPU computing are processed in parallel through parallel pipeline processing.

[0010] Furthermore, the specific steps of step 2) are to determine the execution mode of each module task in the Transformer algorithm and obtain the calculation attribution of each module task.

[0011] Furthermore, the Transformer algorithm in step 2) includes four module tasks: layer normalization module, self-attention module, matrix multiplication module and activation function module layer, among which the normalization module and self-attention module are calculated by the CPU, and the matrix multiplication module and activation function module are calculated by the NPU.

[0012] Furthermore, the specific steps of step 3) are: according to the calculation allocation of the Transformer algorithm module, each module is divided into multiple tasks for calculation, and the calculation results are merged and output through the CPU calculation residual layer; the data transmission and result merging rules are: the CPU calculates the results completed by the normalization module and the self-attention module as input to the NPU for calculation, and after the NPU calculation is completed, it is transmitted to the CPU to complete the calculation of the residual layer.

[0013] Furthermore, in step 3), each module is divided into multiple tasks for calculation, specifically: Task S1 consists of a layer normalization module and a self-attention module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S2 consists of a matrix multiplication module and an activation function module, and the calculation is completed in the NPU, and the result is transmitted to the CPU to complete the merging of the residual layer; Task S3 consists of a layer normalization module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S4 consists of two groups of matrix multiplication modules and activation function modules, and the calculation is completed in the NPU, and the result is output to the CPU to complete the merging of the residual layer.

[0014] Furthermore, the specific steps of step 4) are:

[0015] 4.1) By decomposing the computational tasks of the Transformer algorithm, the algorithm's inputs A, B, C, and D are processed in parallel;

[0016] 4.2) When task S1 completes processing of input A, task S2 continues to complete processing of input A, while task S1 processes input B. This sequence of steps can complete parallel pipeline processing of multiple inputs.

[0017] 4.3) After the final task S4 completes the processing of input D, it merges the results and outputs them.

[0018] Furthermore, in step 4), if CPU core M>1, NPU core N>1, that is, the CPU and NPU have two or more cores, the four tasks can be processed in parallel, that is, when task S4 processes input A, task S1 processes input D, and when task S4 has not completed processing task A, task S1 should wait; among them, task S2, task S3, and task S4 all have three buffers for storing the processing results of the previous task.

[0019] Furthermore, in step 4), if CPU core M>1, NPU core=1, then tasks S2, S3, and S4 are combined into a large task and processed in parallel with task S1, that is, when task S1 processes input B, the large task processes input A. When input A is not processed, the large task should wait after processing the next task of input B, and two buffers are configured for the large task to store the processing results of the previous task.

[0020] The present invention provides a method for accelerating the Transformer algorithm based on CPU+NPU collaboration, which takes the computing capabilities of the underlying hardware architecture CPU and NPU as the core, distributes tasks to each module in the Transformer algorithm by studying the composition structure of the Transformer algorithm and the operations during calculation in the computing platform, uses parallel pipeline processing during calculation and finally summarizes and outputs the results, thereby improving the reasoning and operation performance of the Transformer algorithm. Therefore, the advantages of the present invention are:

[0021] 1) The present invention splits the structure of the Transformer algorithm, reasonably allocates computing tasks, and provides a parallel pipeline computing method, thereby effectively improving the performance of the reasoning Transformer algorithm.

[0022] 2) The collaborative acceleration technology of the present invention not only improves the execution efficiency of the algorithm, but also reduces the dependence on single hardware resources, providing strong support for big data processing and complex model reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of computing task allocation of the Transformer algorithm of the present invention;

[0024] Figure 2 This is a flow chart of a first embodiment of the parallel computing of the present invention;

[0025] Figure 3 This is a flow chart of the second parallel computing embodiment of the present invention. DETAILED DESCRIPTION

[0026] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] The method of the specific embodiment of the present invention is as follows:

[0028] 1) Determine the number of CPU cores M (N is any positive integer > 0) and the number of NPU cores N (N is any positive integer > 0); so as to subsequently determine the operation mode of parallel pipeline processing.

[0029] 2) Determine the computational ownership of each module task in the Transformer algorithm; specifically: determine the execution method of each module task in the Transformer algorithm, obtain the computational ownership of each module task, and different algorithm sub-modules have different processing accuracy and speed on different processors.

[0030] Determine the algorithm submodule allocation rules, and allocate computing tasks according to the processing accuracy and speed of different operator modules on different processors. The algorithm submodules include layer normalization module, self-attention module, matrix multiplication module, and activation function module. The layer normalization module and self-attention module are calculated by the CPU, and the matrix multiplication module and activation function module are calculated by the NPU module.

[0031] 3) Calculate each module according to the distribution of computing tasks in step 2), and merge the results through the CPU computing residual layer; specifically, according to the calculation distribution of the Transformer algorithm module, divide each module into multiple tasks for calculation, and merge the calculation results through the CPU computing residual layer for output;

[0032] Each module is divided into multiple tasks for calculation, specifically: Task S1 consists of a layer normalization module and a self-attention module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S2 consists of a matrix multiplication module and an activation function module, and the calculation is completed in the NPU, and the result is transmitted to the CPU to complete the merging of the residual layer; Task S3 consists of a layer normalization module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S4 consists of two groups of matrix multiplication modules and activation function modules, and the calculation is completed in the NPU, and the result is output to the CPU to complete the merging of the residual layer.

[0033] Determine the data transmission and result merging rules: The results of the normalization module and the self-attention module of the CPU-side calculation layer are used as input to the NPU side for calculation. After the NPU side completes the calculation, it is transmitted to the CPU side to complete the calculation of the residual layer.

[0034] 4) After the computing tasks are distributed according to step 2), the CPU computing and the NPU computing are processed in parallel through parallel pipeline processing.

[0035] 4.1) By decomposing the computational tasks of the Transformer algorithm, the algorithm's inputs A, B, C, and D are processed in parallel;

[0036] 4.2) When task S1 completes processing of input A, task S2 continues to complete processing of input A, while task S1 processes input B. This sequence of steps can complete parallel pipeline processing of multiple inputs.

[0037] 4.3) After finally processing task S4, the results are merged and output.

[0038] See also Figure 1 The Transformer algorithm contains four module tasks, namely: layer normalization module; self-attention module; matrix multiplication module and activation function module. The layer normalization module and self-attention module have smaller computational and data volumes, and are more suitable for calculation on the CPU. The matrix multiplication module and activation function module are more in line with the computing architecture of the NPU, which can improve its computing speed.

[0039] Task S1 consists of a layer normalization module and a self-attention module, and the calculation is completed in the CPU, and the result is transmitted to the NPU. Task S2 consists of a matrix multiplication module and an activation function module, and the calculation is completed in the NPU, and the result is transmitted to the CPU to complete the merging of the residual layer. Task S3 consists of a layer normalization module, and the calculation is completed in the CPU, and the result is transmitted to the NPU. Task S4 consists of two groups of matrix multiplication and activation function modules, and the calculation is completed in the NPU, and the result is output to the CPU to complete the merging of the residual layer.

[0040] See also Figure 2 In the first specific embodiment of the present invention, the four assigned computing tasks can be processed in parallel to speed up the algorithm operation performance. If the CPU core M>1, the NPU core N>1, that is, the CPU and the NPU have two or more cores. The four tasks can be processed in parallel, that is, when task S4 processes input A, task S1 processes input D. When task S4 has not finished processing task A, task S1 should wait. Task S2, task S3, and task S4 all have three buffers for storing the processing results of the previous task.

[0041] See also Figure 3 In the second specific embodiment of the present invention, if CPU core M>1, NPU core=1, then tasks S2, S3, and S4 are combined into a large task and processed in parallel with task S1, that is, when task S1 processes input B, the large task processes input A. When input A is not processed, the large task should wait after processing the next task of input B, and two buffers are configured for the large task to store the processing results of the previous task.

[0042] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

[0043] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.

[0044] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.

Claims

1. A method for accelerating the Transformer algorithm based on CPU+NPU collaboration, characterized in that: The method comprises the following steps: 1) Determine the number of CPU cores M (N is any positive integer > 0) and the number of NPU cores N (N is any positive integer > 0); 2) Determine the computational ownership of each module task in the Transformer algorithm; 3) Calculate each module according to the distribution of computing tasks in step 2), and merge the results through the CPU calculation residual layer; 4) After the computing tasks are distributed according to step 2), the CPU computing and the NPU computing are processed in parallel through parallel pipeline processing.

2. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 1, characterized in that: The specific steps of step 2) are to determine the execution mode of each module task in the Transformer algorithm and obtain the calculation attribution of each module task.

3. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 2, characterized in that: The Transformer algorithm in step 2) includes four module tasks: layer normalization module, self-attention module, matrix multiplication module and activation function module layer, wherein the normalization module and the self-attention module are calculated by the CPU, and the matrix multiplication module and the activation function module are calculated by the NPU.

4. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 3, characterized in that: The specific steps of step 3) are: according to the calculation allocation of the Transformer algorithm module, each module is divided into multiple tasks for calculation, and the calculation results are merged and output through the CPU calculation residual layer; the data transmission and result merging rules are: the CPU calculates the results completed by the normalization module and the self-attention module as input to the NPU for calculation, and after the NPU calculation is completed, it is transmitted to the CPU to complete the calculation of the residual layer.

5. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 4, characterized in that: In the step 3), each module is divided into multiple tasks for calculation, specifically: Task S1 consists of a layer normalization module and a self-attention module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S2 consists of a matrix multiplication module and an activation function module, and the calculation is completed in the NPU, and the result is transmitted to the CPU to complete the merging of the residual layer; Task S3 consists of a layer normalization module, and the calculation is completed in the CPU, and the result is transmitted to the NPU; Task S4 consists of two groups of matrix multiplication modules and activation function modules, and the calculation is completed in the NPU, and the result is output to the CPU to complete the merging of the residual layer.

6. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 5, characterized in that: The specific steps of step 4) are: 4.1) By decomposing the computational tasks of the Transformer algorithm, the algorithm's inputs A, B, C, and D are processed in parallel; 4.2) When task S1 completes processing of input A, task S2 continues to complete processing of input A, while task S1 processes input B. This sequence of steps can complete parallel pipeline processing of multiple inputs. 4.3) After finally processing task S4, the results are merged and output.

7. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 5, characterized in that: In the step 4), if the CPU core M>1, the NPU core N>1, that is, the CPU and the NPU have two or more cores, the four tasks can be processed in parallel, that is, when task S4 processes input A, task S1 processes input D, and when task S4 has not completed processing task A, task S1 should wait; among them, task S2, task S3, and task S4 all have three buffers for storing the processing results of the previous task.

8. The method for accelerating the Transformer algorithm based on CPU+NPU collaboration according to claim 5, characterized in that: In the step 4), if CPU core M>1, NPU core=1, then tasks S2, S3, and S4 are combined into a large task and processed in parallel with task S1, that is, when task S1 processes input B, the large task processes input A. When input A is not processed, the large task should wait after processing the next task of input B, and two buffers are configured for the large task to store the processing results of the previous task.