Hybrid deployment method and system for multiple acceleration cards of single machine
By using the PCIe interface to connect multiple acceleration devices in a stand-alone machine and abstracting them into vGPU software devices, and implementing interfaces based on MLIR, the problem of inability to effectively deploy multiple acceleration devices in the existing technology is solved, and efficient acceleration device utilization and deep learning framework support is achieved.
Patent Information
- Application Number
- CN202510171080.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to effectively deploy multiple acceleration devices, such as GPUs, FPGAs, and ASICs, in a stand-alone machine, resulting in the inability to reasonably utilize existing devices.
Connect all acceleration devices through the PCIe interface, install and configure drivers, and abstract the acceleration card into a vGPU software device. The vGPU programming interface is encapsulated through the acceleration device driver, and the vGPU back-end dialect interface is realized based on MLIR, supporting mainstream deep learning frameworks.
It realizes high bandwidth and low latency communication between different acceleration devices, supports mainstream deep learning frameworks, and enables acceleration devices to be rationally utilized through software abstraction, and establishes a unified hybrid deployment solution.
Smart Images

Figure CN120068964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heterogeneous acceleration technology, and more specifically, to a method and system for hybrid deployment of multiple acceleration cards on a single machine. Background Art
[0002] Mainstream deep learning frameworks such as PyTorch and TensorFlow have widely supported multi - card deployment on a single machine. Through built - in parallel computing modules, efficient multi - card parallelism is achieved. Multi - card deployment on a single machine is widely used in the training and inference of large language models (LLMs), such as the deployment practices of models like Llama and Qwen.
[0003] However, in existing mainstream single - machine deployments, a single manufacturer or a single device is mostly used, and scenarios of hybrid deployment for multiple acceleration devices (such as GPUs, FPGAs, ASICs, etc.) are rarely supported, making it impossible to rationally utilize existing devices. Summary of the Invention
[0004] The technical task of the present invention is to address the above - mentioned deficiencies by providing a method and system for hybrid deployment of multiple acceleration cards on a single machine, establishing a unified hybrid deployment solution to achieve rational utilization of existing devices.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] A method for hybrid deployment of multiple acceleration cards on a single machine, the implementation of this method includes a hybrid deployment hardware platform and a hybrid deployment software stack.
[0007] The hybrid deployment hardware platform:
[0008] All acceleration devices are connected to the host through PCIe interfaces;
[0009] Driver programs are installed and configured for each acceleration device;
[0010] The hybrid deployment software stack:
[0011] Abstract the acceleration card as a vGPU software device;
[0012] Encapsulate the vGPU programming interface through the acceleration device driver;
[0013] Implement the vGPU backend dialect interface based on MLIR.
[0014] Furthermore, the acceleration devices include GPUs, FPGAs, and ASICs, and different acceleration devices achieve high - bandwidth and low - latency communication with the host side through PCIe interfaces.
[0015] Furthermore, the implementation of the vGPU backend dialect interface based on MLIR
[0016] Software abstraction of acceleration devices is performed based on the MLIR compiler infrastructure. As the backend implementation of MLIR, it enables acceleration devices to support mainstream deep learning frameworks; open source projects are used to provide acceleration services to users.
[0017] Further, the mainstream deep learning frameworks include PyTorch and TensorFlow learning frameworks;
[0018] torch-mlir and TensorFlow MLIR open source projects are used to provide acceleration services to users.
[0019] Further, the acceleration card is abstracted as a vGPU software device, and the application layer completes data and computing offloading tasks through the vGPU software interface. The vGPU software interface communicates with specific devices through device drivers.
[0020] Further, by implementing the vGPU device dialect module, the operator selection and scheduling functions in computing tasks are completed;
[0021] By implementing a dialect module for each acceleration device, the differences in different implementations of operators for different devices are isolated.
[0022] Further, the vGPU programming interface is encapsulated, and the implementation methods include a third-party compilation toolchain, an online reconfigurable configuration file, and an ASIC custom operator interface library.
[0023] The present invention also claims to protect a single-machine hybrid deployment system for multiple acceleration cards, including a hybrid deployment hardware platform part and a hybrid deployment software stack part.
[0024] In the hybrid deployment hardware platform part, all acceleration devices are connected to the host through the PCIe interface; driver programs are installed and configured for each acceleration device.
[0025] In the hybrid deployment software stack part, the acceleration card is abstracted as a vGPU software device; the vGPU programming interface is encapsulated through the acceleration device driver; and the vGPU backend dialect (Dialect) interface is implemented based on MLIR.
[0026] This system realizes the hybrid deployment of multiple acceleration cards on a single machine through the above method.
[0027] The present invention also claims to protect a device for realizing the hybrid deployment of multiple acceleration cards on a single machine, including: at least one memory and at least one processor;
[0028] The at least one memory is used to store machine-readable programs;
[0029] The at least one processor is used to call the machine-readable programs to implement the above method.
[0030] The present invention also claims protection for a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above-mentioned method can be implemented.
[0031] Compared with the prior art, the method and system for hybrid deployment of multiple acceleration cards on a single machine according to the present invention have the following beneficial effects:
[0032] Different acceleration devices achieve high-bandwidth and low-latency communication with the host side through the PCIe interface; and based on the MLIR compiler infrastructure, software abstraction of the acceleration devices is performed, enabling the acceleration devices to support mainstream deep learning frameworks, and providing acceleration services to users using open-source projects such as torch-mlir and TensorFlow MLIR; thereby realizing the rational utilization of existing acceleration devices and establishing a unified hybrid deployment solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic diagram of the hybrid deployment hardware platform architecture provided by an embodiment of the present invention;
[0034] Figure 2 is a schematic diagram of the hybrid deployment software stack architecture provided by an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the implementation principle of the software module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present invention will be further described below in conjunction with specific embodiments.
[0037] The embodiment of the present invention provides a method for hybrid deployment of multiple acceleration cards on a single machine, and the implementation of this method includes a hybrid deployment hardware platform and a hybrid deployment software stack.
[0038] The hybrid deployment hardware platform:
[0039] All acceleration devices are connected to the host through the PCIe interface;
[0040] Install and configure driver programs for each acceleration device;
[0041] The hybrid deployment software stack:
[0042] Abstract the acceleration card as a vGPU software device;
[0043] Encapsulate the vGPU programming interface through the acceleration device driver;
[0044] Implement the vGPU backend dialect interface based on MLIR.
[0045] Among them, as Figure 1 shown, the acceleration devices include GPUs, FPGAs, and ASICs. Different acceleration devices achieve high-bandwidth and low-latency communication with the host side through PCIe interfaces.
[0046] The vGPU backend dialect interface is implemented based on MLIR.
[0047] As Figure 2 shown, software abstraction of the acceleration devices is performed based on the MLIR compiler infrastructure. As the backend implementation of MLIR, the acceleration devices support mainstream deep learning frameworks such as the PyTorch and TensorFlow learning frameworks. The torch-mlir and TensorFlow MLIR open-source projects are used to provide acceleration services to users.
[0048] Figure 3 Shows the implementation modules of the vGPU in the abstract device layer of this method and the relationships between the modules:
[0049] The application layer completes data and computing offloading tasks through the vGPU software interface. The vGPU software interface communicates with specific devices through device drivers.
[0050] In addition, by implementing the vGPU device dialect module, the operator selection and scheduling functions for computing tasks are completed;
[0051] In this method, a dialect module is implemented for each type of acceleration device to isolate differences in different implementations of device operators.
[0052] Among them, the vGPU programming interface is encapsulated, and the implementation methods include a third-party compilation toolchain, an online reconfigurable configuration file, and an ASIC custom operator interface library.
[0053] The embodiment of the present invention also provides a system for hybrid deployment of multiple acceleration cards on a single machine, including a hybrid deployment hardware platform part and a hybrid deployment software stack part.
[0054] In the hybrid deployment hardware platform part, all acceleration devices are connected to the host through PCIe interfaces; driver programs are installed and configured for each acceleration device;
[0055] In the hybrid deployment software stack part, the acceleration cards are abstracted as a vGPU software device; the vGPU programming interface is encapsulated through the acceleration device driver; and the vGPU backend dialect interface is implemented based on MLIR.
[0056] This system realizes the hybrid deployment of multiple acceleration cards on a single machine through the method for hybrid deployment of multiple acceleration cards on a single machine described in the above embodiment.
[0057] The acceleration devices include GPUs, FPGAs, and ASICs. Different acceleration devices achieve high-bandwidth and low-latency communication with the host side through the PCIe interface.
[0058] The vGPU backend dialect interface is implemented based on MLIR.
[0059] The acceleration devices are software-abstracted based on the MLIR compiler infrastructure. As the backend implementation of MLIR, the acceleration devices support mainstream deep learning frameworks such as the PyTorch and TensorFlow learning frameworks. The torch-mlir and TensorFlow MLIR open-source projects are used to provide acceleration services to users.
[0060] For the hybrid deployment software stack part, the implementation modules of the abstract device layer vGPU and the relationships between the modules are as follows:
[0061] The application layer completes data and computing offloading tasks through the vGPU software interface. The vGPU software interface communicates with specific devices through device drivers.
[0062] In addition, by implementing the vGPU device dialect module, the operator selection and scheduling functions for computing tasks are completed;
[0063] By implementing a dialect module for each acceleration device, the differences in the different implementations of operators for different devices are isolated.
[0064] The vGPU programming interface is encapsulated, and the implementation methods include a third-party compilation toolchain, an online reconfigurable configuration file, and an ASIC custom operator interface library.
[0065] An embodiment of the present invention also provides a device for realizing hybrid deployment of multiple acceleration cards in a single machine, including: at least one memory and at least one processor;
[0066] The at least one memory is used to store machine-readable programs;
[0067] The at least one processor is used to call the machine-readable program to implement the method for hybrid deployment of multiple acceleration cards in a single machine described in the above embodiment.
[0068] An embodiment of the present invention also provides a computer-readable medium. Computer instructions are stored on the computer-readable medium. When the computer instructions are executed by a processor, the processor executes the method for hybrid deployment of multiple acceleration cards in a single machine described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided. Software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored on the storage medium.
[0069] In this case, the program code read from the storage medium itself can implement the functions of any one of the above-described embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0070] Examples of the storage medium for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as a CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code may be downloaded from a server computer via a communication network.
[0071] In addition, it should be clear that not only can the functions of any one of the above-described embodiments be implemented by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0072] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, a CPU or the like installed on the expansion board or the expansion unit is caused to execute part or all of the actual operations, thereby implementing the functions of any one of the above-described embodiments.
[0073] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-described multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for hybrid deployment of multiple accelerator cards on a single machine, characterized in that: The implementation of this method includes hybrid deployment of hardware platforms and hybrid deployment of software stacks. The hybrid deployment hardware platform: All acceleration devices are connected to the host through the PCIe interface; Install and configure drivers for each acceleration device; The hybrid deployment software stack: Abstract the accelerator card into a vGPU software device; Encapsulate the vGPU programming interface through the acceleration device driver; Implement the vGPU backend dialect interface based on MLIR.
2. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 1, characterized in that: The acceleration devices include GPU, FPGA, and ASIC. Different acceleration devices achieve high-bandwidth and low-latency communication with the host side through the PCIe interface.
3. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 1, characterized in that: The vGPU backend dialect interface is implemented based on MLIR. Based on the MLIR compiler infrastructure, software abstraction is performed on the acceleration device as the backend implementation of MLIR, so that the acceleration device supports mainstream deep learning frameworks; open source projects are used to provide acceleration services to users.
4. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 3, characterized in that: The mainstream deep learning frameworks include PyTorch and TensorFlow learning frameworks; Use torch-mlir and TensorFlow MLIR open source projects to provide acceleration services to users.
5. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 1 or 3, characterized in that: The accelerator card is abstracted as a vGPU software device, and the application layer completes data and computing offload tasks through the vGPU software interface, and the vGPU software interface communicates with the specific device through the device driver.
6. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 5, characterized in that: By implementing the vGPU device dialect module, the operator selection and scheduling functions of computing tasks are completed; By implementing a dialect module for each acceleration device, different implementations of operators on different devices are isolated.
7. The method for hybrid deployment of multiple acceleration cards on a single machine according to claim 6, characterized in that: The encapsulated vGPU programming interface is implemented by a third-party compilation tool chain, an online reconfigurable configuration file, and an ASIC customized operator interface library.
8. A single machine mixed deployment system with multiple accelerator cards, characterized in that: Including hybrid deployment hardware platform part and hybrid deployment software stack part, In the hybrid deployment hardware platform, all acceleration devices are connected to the host through the PCIe interface; a driver is installed and configured for each acceleration device; The hybrid deployment software stack abstracts the accelerator card into a vGPU software device; encapsulates the vGPU programming interface through the accelerator device driver; and implements the vGPU backend dialect interface based on MLIR; The system realizes mixed deployment of multiple acceleration cards on a single machine through the method described in any one of claims 1 to 7.
9. A device for implementing mixed deployment of multiple accelerator cards on a single machine, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method described in any one of claims 1 to 7.