Two-level parallel acceleration method for AI frameworks on heterogeneous many-core processors
By designing a two-level parallel acceleration method for AI frameworks for heterogeneous multi-core processors, using model optimization and core group management modules, the problem that traditional AI frameworks cannot efficiently utilize heterogeneous multi-core processor resources is solved, and high-performance and efficient AI framework operation is achieved.
Patent Information
- Application Number
- CN202210136541.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-15
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2042-02-15
AI Technical Summary
Traditional AI frameworks cannot efficiently utilize their multi-level storage hierarchy and powerful computing power on heterogeneous multi-core processors, resulting in poor performance and utility of AI frameworks in this architecture.
Design a two-level parallel acceleration method for AI frameworks for heterogeneous multi-core processors. The deep learning model is converted into a tree computing graph through the model optimization module. The thread management module and the kernel group management module work together to realize multi-threaded automatic scheduling between core groups and parallel computing between cores, making full use of the multi-level storage resources and computing capabilities of heterogeneous multi-cores.
The two-level parallel acceleration of the AI framework on heterogeneous multi-core processors has been achieved, which has significantly improved the performance and effectiveness of the AI framework on heterogeneous multi-cores, and made full use of multi-level storage and computing resources.
Smart Images

Figure CN114661460B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a two-level parallel acceleration method for an AI framework for heterogeneous multi-core processors, which is used for an AI framework for heterogeneous multi-core processors and belongs to the field of computing technology. Background Art
[0002] Currently, the AI framework is an important basic software that supports applications in the field of deep learning. After years of development, the AI framework has achieved remarkable optimization effects on architectures such as traditional CPU multi-core processors, GPUs, and AI special-purpose chips. On the multi-core CPU architecture, the AI framework generally achieves parallel computing through multi-threading, organizes the core computations (i.e., operators) in the operation into a thread pool for automatic scheduling, and also performs parallel computing through multi-threading inside the operator; on the GPU architecture, the AI framework sends the operators to the acceleration card for acceleration, the framework runs on the host, only for scheduling and control, and the main computations are all performed on the acceleration card.
[0003] Different from traditional processor architectures, the heterogeneous multi-core processor adopts a heterogeneous fusion architecture. The main core and the slave core array form a core group, multiple core groups are integrated on one processor, and it has multiple levels of memory hierarchy. The main core is mainly responsible for control, and the slave core array is responsible for collaborative acceleration computing. Under this special architecture, the design methods of traditional AI frameworks cannot adapt well and cannot efficiently utilize this two-level parallel architecture. Therefore, on heterogeneous multi-core processors, how the AI framework can efficiently utilize its multi-level memory hierarchy and powerful computing power is a challenge. Summary of the Invention
[0004] The purpose of the present invention is to provide a two-level parallel acceleration method for an AI framework for heterogeneous multi-core processors, which can make full use of the multi-level memory resources and computing power of heterogeneous multi-cores, realize automatic two-level parallel acceleration of the AI framework, and significantly improve the usability and high performance of the AI framework on heterogeneous multi-cores.
[0005] To achieve the above object, the present invention provides a two-level parallel acceleration method for an AI framework for heterogeneous multi-core processors, based on the following functional modules:
[0006] Model Optimization Module: used to appropriately split the model and convert some serial computation graphs into tree-like computation graphs;
[0007] Thread Management Module: used to manage the thread pool of the AI framework and schedule each thread to an idle core group;
[0008] Core Group Management Module: used to parallelly allocate computation tasks to each slave core for execution, and after detecting that all slave cores have completed the computation, report to the thread management module that the computation task is completed and set itself to the idle state;
[0009] The acceleration method includes the following steps:
[0010] Step 1: The AI framework calls the model optimization module to optimize the deep learning model or the pre-trained model, and organizes it into a more parallelizable tree-shaped computation graph;
[0011] Step 2: The AI framework converts the optimized tree-shaped computation graph in Step 1 into a thread pool composed of computing tasks;
[0012] Step 3: The thread management module organizes the thread pool obtained in Step 2 into different thread queues according to the relevance. The computing tasks between different queues are not relevant and can be executed in parallel, while the computing tasks within the same queue must be executed according to the tasks;
[0013] Step 4: The thread management module monitors the status of each core group. If it finds that a certain core group is in an idle state, it schedules the computing tasks of a certain queue to be executed on this core group, and ensures that the computing tasks running on each core group belong to different computing queues;
[0014] When a certain core group completes the computing task and is in an idle state, the thread management module continues to allocate computing tasks to this core group until all the computing tasks of all queues are completed;
[0015] When a certain core group receives a computing task, the core group management module distributes the computing task in parallel to each slave core for execution, and monitors the running status of each slave core. After all slave cores complete the computing task and are in an idle state, the core group management module reports the completion of the task to the thread management module and waits for the allocation of the next computing task;
[0016] Step 5: When the entire running task is completed, both the thread management module and the core group management module stop running. After the next running task starts, both the thread management module and the core group management module enter the monitoring state, waiting for the allocation of thread tasks and core group computing tasks.
[0017] The further improved solutions in the above technical solutions are as follows:
[0018] 1. In the above solution, the thread management module is configured to:
[0019] Automatically detect all core group resources. When it finds a core group in an idle state, it selects a computing task in the thread pool and allocates it to this thread, and sets the core group resources to the busy state;
[0020] Continue to schedule parallelizable computing tasks to other core groups until all core groups are in a busy state;
[0021] Enter the monitoring state and wait for the next task allocation.
[0022] 2. In the above solution, when a certain core group receives the computing task allocated by the thread management module, the main core is responsible for running a core group management module.
[0023] 3. In step 3 of the above solution, the thread management module constructs corresponding thread queues according to the number of core groups of the heterogeneous many-core processor.
[0024] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0025] The present invention provides a two-level parallel acceleration method for an AI framework for a heterogeneous many-core processor, which realizes data-level parallelism through multi-thread automatic scheduling among core groups at a larger scale, and realizes collaborative acceleration of operator-level operations through heterogeneous slave core computing resources within the core group, fully utilizes the multi-level storage resources and computing power of the heterogeneous many-core, realizes automatic two-level parallel acceleration of the AI framework, and significantly improves the usability and high performance of the AI framework on the heterogeneous many-core. Brief Description of the Drawings
[0026] Appendix Figure 1 is a flowchart of the two-level parallel acceleration method of the AI framework of the present invention. Detailed Embodiments
[0027] Embodiment: The present invention provides a two-level parallel acceleration method for an AI framework for a heterogeneous many-core processor, based on the following functional modules:
[0028] Model optimization module: used to appropriately split the model and convert some serial computational graphs into tree-shaped computational graphs;
[0029] Thread management module: used to manage the thread pool of the AI framework and schedule each thread to an idle core group;
[0030] Core group management module: used to parallelly allocate computing tasks to each slave core for execution, and after detecting that all slave cores have completed the calculation, report to the thread management module that the computing task is completed and place itself in an idle state;
[0031] The acceleration method includes the following steps:
[0032] Step 1, the AI framework calls the model optimization module to optimize the deep learning model or the pre-trained model, and organizes it into a more parallelizable tree-shaped computational graph;
[0033] Step 2, the AI framework converts the optimized tree-shaped computational graph in step 1 into a thread pool composed of computing tasks;
[0034] Step 3. The thread management module organizes the thread pools obtained in Step 2 into different thread queues according to the relevance. The computing tasks between different queues are not relevant and can be executed in parallel, while the computing tasks within the same queue must be executed according to the tasks.
[0035] Step 4. The thread management module monitors the status of each core group. If it finds that a certain core group is in an idle state, it schedules the computing tasks of a certain queue to be executed on this core group, and ensures that the computing tasks running on each core group belong to different computing queues.
[0036] When a certain core group completes the computing task and is in an idle state, the thread management module continues to allocate computing tasks to this core group until all the computing tasks of all queues are completely completed.
[0037] When a certain core group receives a computing task, the core group management module distributes the computing task to each slave core in parallel and monitors the running status of each slave core. After all slave cores complete the computing task and are in an idle state, the core group management module reports the completion of the task to the thread management module and waits for the allocation of the next computing task.
[0038] Step 5. When the entire running task is completed, both the thread management module and the core group management module stop running. After the next running task starts, both the thread management module and the core group management module enter the monitoring state, waiting for the allocation of thread tasks and core group computing tasks.
[0039] The above thread management module is configured as follows:
[0040] Automatically detect all core group resources. When a core group in an idle state is found, select a computing task from the thread pool and allocate it to this thread, and set the core group resources to the busy state.
[0041] Continue to schedule parallelizable computing tasks to other core groups until all core groups are in the busy state.
[0042] Enter the monitoring state and wait for the next task allocation.
[0043] When a certain core group receives the computing task allocated by the thread management module, the main core is responsible for running a core group management module.
[0044] In the above Step 3, the thread management module constructs corresponding thread queues according to the number of core groups of the heterogeneous many-core processor.
[0045] The further explanation of the above embodiment is as follows:
[0046] The parallel modes of existing AI frameworks can be mainly divided into three types:
[0047] Parallelism between processors. This parallel mode is generally carried out through communication message libraries such as MPI. Each processor runs as an independent process as part of the model or part of the training data, and then reduction is performed through global communication to form a unified result. This parallel mode is not the parallelism inside the processor, but a higher-level and coarser-grained data parallelism or model parallelism, which can be directly supported on a large-scale system composed of heterogeneous many-core processors.
[0048] Parallelism between cores inside a multi-core processor. This parallel mode is generally implemented through the pthread library. A complete running process is composed of a thread pool that can be parallelized, and a unified scheduler schedules each thread in the thread pool to each core for running. Fine-grained parallelism can also be achieved through multiple threads inside a single operator. However, the slave cores of heterogeneous many-core processors cannot support the thread mode. The multi-thread mode can only achieve thread parallelism between the main cores and cannot utilize the computing resources of the slave cores.
[0049] For the separate heterogeneous many-core acceleration mode of the GPU acceleration card, this mode will load the operator calculation process onto the GPU for execution, and the host part only performs operations such as IO and scheduling, and it also cannot adapt to the heterogeneous integrated many-core chip.
[0050] The present invention combines the above three parallel modes, supports distributed operation through the MPI library between chips, implements a two-level parallel mode for heterogeneous many-cores inside the chip, uses multiple threads to support the calculation scheduling between core groups, and each thread no longer takes a single core but the entire core group as the scheduling unit, and then realizes the parallel acceleration calculation of the slave core array through the acceleration library inside the thread. This two-level parallel mode is designed and tailored for the heterogeneous many-core architecture, can make full use of its multi-level storage hierarchy and computing resources, and give play to the efficiency of the AI framework.
[0051] The overall architecture is as Figure 1 shown, and the main functional modules include:
[0052] Model optimization module: In order to improve the parallelizability of the model, the model optimization module is responsible for appropriately splitting the model. The number of splits is adapted to the number of core groups of the heterogeneous chip, and some serial computation graphs are converted into tree-shaped computation graphs, which is beneficial to subsequent thread allocation and scheduling.
[0053] Thread management module: Responsible for managing the thread pool of the AI framework and scheduling each thread to an idle core group.
[0054] This module will automatically detect all core group resources. When a core group in an idle state is found, it will select a computing task from the thread pool, allocate it to this thread, and set the core group resource to the busy state. Then, this module will continue to schedule parallel computing tasks to other core groups. After all core groups are in the busy state, the thread management module enters the monitoring state and waits for the next task allocation;
[0055] Core group management module: When a certain core group receives the computing task allocated by the thread management module, the main core is responsible for running a core group management module. This module will parallelly allocate the computing task to each slave core for execution. Each slave core executes a part of the computing task and uses its own local storage space to accelerate the memory access speed;
[0056] After the core group management module detects that all slave cores have completed the calculation, it reports the completion of the computing task to the thread management module and then sets itself to the idle state.
[0057] The design of the two-layer complementary parallel management module. The upper-layer thread management module relies on pthread multi-threading to implement the management, allocation, and scheduling of computing tasks, thereby supporting parallel operations between core groups; the lower-layer core group management module is responsible for managing the slave core array within the core group, can efficiently allocate computing tasks to each slave core, and cooperate with the upper-layer thread management module; the two-level parallel method can efficiently utilize the multi-level storage hierarchy and powerful computing resources of heterogeneous processors, providing the heterogeneous many-core backend optimization support lacking in current open-source AI frameworks;
[0058] The overall operation process is as follows:
[0059] The AI framework first calls the model optimization module to optimize the deep learning model written by the user or the pre-trained model, and organizes it into a more parallelizable tree-shaped computational graph. This optimization step is beneficial for the thread management module to perform better thread allocation and scheduling;
[0060] The optimized computational graph is converted by the AI framework into a thread pool composed of computing tasks. The thread pool will enter the thread management module. The thread management module will organize the thread pool into different thread queues according to the relevance. The computing tasks between different queues are not relevant and can be executed in parallel. The computing tasks within the same queue must be executed according to the tasks;
[0061] The thread management module constructs corresponding thread queues according to the number of core groups of the heterogeneous many-core processor, and at the same time monitors the status of each core group. If a certain core group is found to be in the idle state, it will schedule the computing tasks of a certain queue to this core group for execution, and ensure that the computing tasks running on each core group belong to different computing queues;
[0062] When a certain core group has completed its computing tasks and is in an idle state, the thread management module will continue to allocate computing tasks to this core group until all the computing tasks in all queues are completely finished;
[0063] When a certain core group receives a computing task, it will enter the core group management module, which is responsible for parallelly allocating the computing task to each slave core for execution, and then monitoring the running status of each slave core. After all slave cores have completed their computing tasks and are in an idle state, the core group management module reports the completion of the task to the thread management module and waits for the allocation of the next computing task;
[0064] When the entire running task is completed, both the thread management module and the core group management module will abort their operations. After the next running task starts, both of these modules will enter the monitoring state, waiting for the allocation of thread tasks and core group computing tasks.
[0065] Design of a model optimization method adapted to a heterogeneous multi-core architecture to improve the model's parallelizability and enhance the parallel efficiency of multiple core groups; The overall technical solution is independent of the specific implementation of the AI framework, so it can support multiple mainstream AI frameworks upwards, such as TensorFlow, Pytorch, etc., and has good portability; It is a dedicated architecture design adapted to heterogeneous multi-core processors, and the overall implementation method is transparent to users.
[0066] To facilitate a better understanding of the present invention, the terms used in this article will be briefly explained below:
[0067] Heterogeneous multi-core processor: A high-performance heterogeneous central processing unit that integrates a small number of general-purpose master core cores responsible for management, communication, and computing functions and a large number of reduced slave core cores responsible for computing functions on a single complete chip. The general-purpose master core core runs a general operating system, mainly undertakes the management and control functions of the entire chip, and also undertakes certain computing functions and the communication functions between the chip and the outside; The slave core core plays the role of accelerating computing. Each processor generally contains multiple core groups, and each core group consists of a master core core and multiple reduced slave core cores.
[0068] AI framework: The AI framework is a very important basic software in the field of artificial intelligence, providing a series of upper-layer interfaces with rich functions to support the programming implementation of various machine learning algorithms; Open-source AI frameworks only support CPUs, GPUs, and some AI-specific chips. Common AI frameworks include TensorFlow, Pytorch, MXNet, Caffe, etc.
[0069] Two-level parallelism: A parallel mode adapted to a heterogeneous multi-core architecture, including parallelism between single-chip core groups and parallelism between slave cores within a single core group.
[0070] When the above-mentioned AI framework two-level parallel acceleration method for heterogeneous many-core processors is adopted, data-level parallelism is achieved through multi-thread automatic scheduling among larger-scale core groups, operator-level operations are synergistically accelerated through heterogeneous slave-core computing resources within the core group, message interaction is realized through full-chip shared memory resources, and efficient memory access optimization is achieved through slave-core local memory, so as to make full use of the multi-level storage resources and computing capabilities of heterogeneous many-cores, realize the two-level parallel acceleration of the AI framework automatically, and significantly improve the usability and high performance of the AI framework on heterogeneous many-cores.
[0071] The above embodiments are only used to illustrate the technical concept and characteristics of the present invention, and their purpose is to enable those familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A two-level parallel acceleration method for AI framework for heterogeneous multi-core processors, characterized in that: Based on the following functional modules: Model optimization module: used to split the model appropriately and convert some serial computation graphs into tree-like computation graphs; Thread management module: used to manage the thread pool of the AI framework and schedule each thread to the idle core group; Core group management module: used to distribute computing tasks to each slave core in parallel for execution, and after detecting that all slave cores have completed computing, report the completion of computing tasks to the thread management module and put itself into an idle state; The acceleration method comprises the following steps: Step 1: The AI framework calls the model optimization module to optimize the deep learning model or pre-trained model and organize it into a more parallelizable tree-shaped computational graph; Step 2: The AI framework converts the tree-shaped computation graph optimized in step 1 into a thread pool consisting of computational tasks. Step 3: The thread management module organizes the thread pool obtained in step 2 into different thread queues according to the correlation. The computing tasks between different queues are not correlated and can be executed in parallel. The computing tasks within the same queue must be executed task by task. Step 4: The thread management module monitors the status of each core group. If a core group is found to be idle, the computing tasks of a queue are scheduled to be executed on the core group, and the computing tasks running on each core group at the same time are guaranteed to belong to different computing queues. When a core group completes its computing tasks and is in an idle state, the thread management module continues to assign computing tasks to the core group until all computing tasks in all queues are completed; When a core group receives a computing task, the core group management module distributes the computing task to each slave core in parallel for execution and monitors the running status of each slave core. After all slave cores have completed the computing task and are in an idle state, the core group management module reports the task completion to the thread management module and waits for the next computing task to be assigned. Step 5: When the entire running task is completed, the thread management module and the core group management module both stop running. When the next running task starts, the thread management module and the core group management module both enter the monitoring state and wait for the allocation of thread tasks and core group computing tasks.
2. The AI framework two-level parallel acceleration method for heterogeneous multi-core processors according to claim 1 is characterized in that: The thread management module is configured as follows: Automatically detect all core group resources. When an idle core group is found, select a computing task from the thread pool and assign it to the thread, and set the core group resource to a busy state. Continue to schedule parallel computing tasks to other core groups until all core groups are busy; Enter the monitoring state and wait for the next task assignment.
3. The two-level parallel acceleration method for AI framework for heterogeneous multi-core processors according to claim 1 is characterized in that: When a core group receives a computing task assigned by a thread management module, the main core is responsible for running a core group management module.
4. The two-level parallel acceleration method for AI framework for heterogeneous multi-core processors according to claim 1 is characterized in that: In step 3, the thread management module constructs a corresponding thread queue according to the number of core groups of the heterogeneous many-core processor.
Citation Information
Patent Citations
Method and system for two-stage partition and two-time polycondensation parallel computing of any banded linear equation set of heterogeneous many-core processor
CN110347967A
Heterogeneous many-core processor-oriented multi-task parallel scheduling method
CN112416539A