A compilation-based neural network heterogeneous many-core multi-level resource mapping method
Through the compilation-based multi-level resource mapping method of heterogeneous multi-cores, circular splitting and resource binding technology, the problem that the deep learning compiler TVM cannot effectively map resources is solved, and the high-performance operation of deep learning load on domestic heterogeneous multi-core processors is achieved.
Patent Information
- Application Number
- CN202110381428.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-04-09
AI Technical Summary
The current deep learning compiler TVM cannot effectively map resource, resulting in the inability to fully utilize the performance of domestic heterogeneous multi-core processors.
The compilation-based multi-level resource mapping method of heterogeneous multi-core cores is adopted, and flexible resource mapping of multi-core groups, kernel threads and vector computing components is realized through circular splitting and resource binding technology.
Fully tap the parallel potential of neural network operators, give full play to the advantages of on-chip multi-level parallelism, and improve the performance of deep learning loads on heterogeneous multi-core platforms.
Smart Images

Figure CN114253545B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a neural network heterogeneous multi-core multi-level resource mapping method based on compilation, and belongs to the technical field of compilation optimization. Background Art
[0002] The role of the deep learning compiler is to deploy deep learning workloads on specific hardware platforms to efficiently complete training and reasoning tasks. It can fully tap the algorithm characteristics and pattern features in the field of artificial intelligence, and convert the models of various typical deep learning frameworks into a unified computational graph. Then, through a series of domain algorithm-guided compilation optimization technologies and architecture-related underlying optimization technologies, it generates efficient code for different hardware platforms to accelerate the reasoning process in deep learning.
[0003] TVM (Tensor Virtual Machine) is the current mainstream deep learning compiler. It implements a unified software stack for different deep learning frameworks and hardware platforms, and deploys deep learning models under different frameworks to hardware platforms in the most efficient way possible. The supported hardware platforms mainly include CPU architectures such as X86 and ARM, and GPU architectures such as NVIDIA and AMD. For multi-core CPU architectures, the deep learning compiler uses multi-OpenMP multi-threaded parallelism to map computing tasks to multiple cores for execution; for GPU architectures, it uses programming models such as CUDA or OpenCL to map computing tasks to resources.
[0004] Domestic heterogeneous multi-core processors use a new on-chip heterogeneous fusion architecture that is completely different from existing CPUs and GPUs. The on-chip computing resources are divided into multiple core groups. The core groups use a master-slave heterogeneous multi-core architecture with shared memory. In addition, a vector extension instruction system is added on the basis of the basic instruction system, which is very suitable for accelerating deep learning tasks. However, the hardware platforms supported by current deep learning compilers mainly include CPU architectures such as X86 and ARM, and GPU architectures such as NVIDIA and AMD. For this domestic heterogeneous multi-core architecture with multi-level hardware resources, the current deep learning compiler TVM cannot perform effective resource mapping, making it impossible for deep learning loads to fully utilize the performance of heterogeneous multi-core platforms. Summary of the invention
[0005] The purpose of the present invention is to provide a neural network heterogeneous many-core multi-level resource mapping method based on compilation, so as to solve the problem that the current deep learning compiler TVM cannot perform effective resource mapping for domestic heterogeneous many-core processors, resulting in the deep learning load being unable to fully utilize the performance of the heterogeneous many-core platform.
[0006] To achieve the above object, the technical solution adopted by the present invention is: to provide a neural network heterogeneous multi-core multi-level resource mapping method based on compilation, comprising the following steps:
[0007] S1. Perform multi-core group resource mapping, as follows:
[0008] S11, splitting the outermost loop x of the neural network operator to obtain an outer loop xo and an inner loop xi, and setting the number of loops of the split outer loop xo to be equal to the number N of many-core core groups;
[0009] S12, for the outer loop xo obtained in S11, bind its calculation process to the many-core core group resources;
[0010] S2. Perform slave core thread resource mapping, as follows:
[0011] S21, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 2, execute S23, otherwise, execute S22;
[0012] S22, loop splitting the inner loop xi obtained in step S11 to obtain an outer loop yo and an inner loop yi, and setting the number of loops of the outer loop yo after the splitting to be equal to the number T of slave core threads in the core group;
[0013] S23, split the inner loop y of the outermost loop x to obtain an outer loop yo and an inner loop yi, and set the number of loops of the outer loop yo after the split to be equal to the number T of slave core threads in the core group;
[0014] S24, for the outer loop yo obtained in S22 or S23, bind its calculation process to the slave core thread resources;
[0015] S3. Perform vector component resource mapping, as follows:
[0016] S31, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 3, execute step S33, otherwise, execute step S32;
[0017] S32, performing loop splitting on the inner loop yi obtained in step S2 to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the splitting to be equal to the operation core vector width K;
[0018] S33, performing loop splitting on the inner loop z of the inner loop y to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the splitting to be equal to the operation core vector width K;
[0019] S34. Bind the calculation process of the inner loop zi obtained in S32 or S33 to the operation core vector calculation component.
[0020] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0021] The present invention provides a neural network heterogeneous many-core multi-level resource mapping method based on compilation. On the basis of a deep learning compiler, the method aims at the characteristics of the domestic heterogeneous many-core multi-level hardware resource architecture, and realizes flexible resource mapping of computing tasks on multi-core groups, multi-slave cores and vector computing components through a DSL interface, fully taps the parallel potential of neural network operators, and gives play to the advantages of on-chip multi-level parallelism, thereby improving the performance of deep learning loads on heterogeneous many-core platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Attached Figure 1 This is a flow chart of the neural network heterogeneous multi-core multi-level resource mapping method based on compilation of the present invention. DETAILED DESCRIPTION
[0023] Embodiment: The present invention provides a neural network heterogeneous multi-core multi-level resource mapping method based on compilation, which specifically includes the following steps:
[0024] S1. Perform multi-core group resource mapping, as follows:
[0025] S11, splitting the outermost loop x of the neural network operator to obtain an outer loop xo and an inner loop xi, and setting the number of loops of the split outer loop xo to be equal to the number N of many-core core groups;
[0026] S12, for the outer loop xo obtained in S11, bind its calculation process to the many-core core group resources;
[0027] S2. Perform slave core thread resource mapping, as follows:
[0028] S21, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 2, execute S23, otherwise, execute S22;
[0029] S22, loop splitting the inner loop xi obtained in step S11 to obtain an outer loop yo and an inner loop yi, and setting the number of loops of the outer loop yo after the splitting to be equal to the number T of slave core threads in the core group;
[0030] S23, split the inner loop y of the outermost loop x to obtain an outer loop yo and an inner loop yi, and set the number of loops of the outer loop yo after the split to be equal to the number T of slave core threads in the core group;
[0031] S24, for the outer loop yo obtained in S22 or S23, bind its calculation process to the slave core thread resources;
[0032] S3. Perform vector component resource mapping, as follows:
[0033] S31, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 3, execute step S33, otherwise, execute step S32;
[0034] S32, performing loop splitting on the inner loop yi obtained in step S2 to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the splitting to be equal to the operation core vector width K;
[0035] S33, splitting the inner loop z of the inner loop y to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the split to be equal to the operation core vector width K;
[0036] S34. Bind the calculation process of the inner loop zi obtained in S32 or S33 to the operation core vector calculation component.
[0037] The above embodiment is further explained as follows:
[0038] The present invention proposes a neural network heterogeneous multi-core multi-level resource mapping method based on compilation, and implements the above resource mapping method in the TVM deep learning compiler based on domestic processors. The specific process is as follows Figure 1 As shown, it mainly includes three steps: multi-core group resource mapping, slave core thread resource mapping, and vector component resource mapping:
[0039] S1. First, perform multi-core group resource mapping, as follows:
[0040] S11. Split the outermost loop x of the neural network operator, and set the number of loops of the split outer loop xo to be equal to the number of many-core core groups N. The corresponding DSL is as follows: xo, xi = s[A].split(x, nparts=N);
[0041] S12. For the outer loop xo obtained in S11, bind its computation process to the many-core core group resources. The corresponding DSL is as follows: s[A].bind(xo, tvm.thread_axis(“_CGN”));
[0042] S2. Secondly, perform slave core thread resource mapping, as follows:
[0043] S21, if the number of loop layers of the neural network operator before splitting in step S1 is greater than or equal to 2, execute step S23, otherwise, execute step S22;
[0044] S22, split the loop xi, and set the number of loops of the outer loop yo after the split to be equal to the number of slave core threads T in the core group. The corresponding DSL is as follows: yo, yi = s[A].split(xi, nparts=T);
[0045] S23, split the inner loop y of loop x, and set the number of loops of the outer loop yo after the split to be equal to the number of slave core threads T in the core group. The corresponding DSL is as follows: yo, yi = s[A].split(y, nparts=T);
[0046] S24. For the outer loop yo obtained in S22 or S23, bind its computation process to the slave core thread resource. The corresponding DSL is as follows: s[A].bind(yo, tvm.thread_axis(“_PEN”));
[0047] S3. Finally, vector component resource mapping is performed, as follows:
[0048] S31, if the number of loop layers of the neural network operator before splitting in step S1 is greater than or equal to 3, execute step S33, otherwise, execute step S32;
[0049] S32, loop splitting is performed on loop yi, and the number of loops of the inner loop zi after the splitting is set to be equal to the operation core vector width K. The corresponding DSL is as follows: zo, zi = s[A].split(yi, factor=K);
[0050] S33, split the inner loop z of loop y, and set the number of loops of the inner loop zi after the split to be equal to the operation core vector width K. The corresponding DSL is as follows: zo, zi = s[A].split(z, factor=K);
[0051] S34. For the inner loop zi obtained in S32 or S33, its calculation process is bound to the operation core vector calculation component. The corresponding DSL is as follows: s[A].vectorize(zi).
[0052] When the above-mentioned compilation-based neural network heterogeneous multi-core multi-level resource mapping method is adopted, its.
[0053] In order to facilitate a better understanding of the present invention, the terms used in this article are briefly explained below:
[0054] Compilation: The process of translating a source program (high-level language) into a target program (low-level language or machine language).
[0055] Heterogeneous many-core: adopts a new on-chip heterogeneous fusion architecture.
[0056] Neural Networks: Neural networks with many hidden layers, also called deep feedforward networks or multilayer perceptrons.
[0057] TVM: Tensor Virtual Machine, a deep learning compiler launched by Amazon, can deploy deep learning workloads on specific hardware platforms to efficiently complete reasoning tasks.
[0058] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A compilation-based neural network heterogeneous multi-core multi-level resource mapping method, It is characterized in that The following steps are involved: S1. Perform multi-core group resource mapping, as follows: S11, splitting the outermost loop x of the neural network operator to obtain an outer loop xo and an inner loop xi, and setting the number of loops of the split outer loop xo to be equal to the number N of many-core core groups; S12, for the outer loop xo obtained in S11, bind its calculation process to the many-core core group resources; S2. Perform slave core thread resource mapping, as follows: S21, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 2, execute S23, otherwise, execute S22; S22, loop splitting the inner loop xi obtained in step S11 to obtain an outer loop yo and an inner loop yi, and setting the number of loops of the outer loop yo after the splitting to be equal to the number T of slave core threads in the core group; S23, split the inner loop y of the outermost loop x to obtain an outer loop yo and an inner loop yi, and set the number of loops of the outer loop yo after the split to be equal to the number T of slave core threads in the core group; S24, for the outer loop yo obtained in S22 or S23, bind its calculation process to the slave core thread resources; S3. Perform vector component resource mapping, as follows: S31, if the number of loop layers of the neural network operator before splitting in step S11 is greater than or equal to 3, execute step S33, otherwise, execute step S32; S32, performing loop splitting on the inner loop yi obtained in step S2 to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the splitting to be equal to the operation core vector width K; S33, performing loop splitting on the inner loop z of the inner loop y to obtain an outer loop zo and an inner loop zi, and setting the number of loops of the inner loop zi after the splitting to be equal to the operation core vector width K; S34. Bind the calculation process of the inner loop zi obtained in S32 or S33 to the operation core vector calculation component.
Citation Information
Patent Citations
Accelerated programming and compiling method for supporting heterogeneous many-core full-chip view angle
CN112558978A
Convolution calculation data reuse method based on heterogeneous many-core processor
CN112559197A