A Hybrid Acceleration Card Management Method, Device, Electronic Device, and Storage Medium
By binding the runtime library interfaces of multiple accelerator cards with a unified hardware abstract interface and using remote direct memory access technology for data transmission, the unified processing problem of data storage and transmission between GPU accelerator cards of different manufacturers is solved, and efficient data processing and transmission efficiency is achieved.
Patent Information
- Application Number
- CN202510059248.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The existing technology is difficult to compatible with GPU accelerator cards from different manufacturers, which leads to challenges in the development and optimization of cross-platform applications, and cannot completely solve the problem of unified processing of data storage and transmission between GPU accelerator cards from different manufacturers.
By binding the runtime library interfaces of multiple accelerator cards with a preset unified hardware abstraction interface, remote direct memory access technology is used to register the memory address of the hybrid accelerator card to the hardware abstraction layer, data transmission operations between hybrid accelerator cards are realized.
It realizes unified management of GPU accelerator cards compatible with different manufacturers, improves data processing and transmission efficiency, reduces the adaptation workload of developers, and improves the computing performance and resource utilization of heterogeneous computing systems.
Smart Images

Figure CN119473994B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method, apparatus, electronic device, and storage medium for managing hybrid acceleration cards. Background Art
[0002] In the fields of high performance computing (HPC), artificial intelligence (AI), scientific computing, and large-scale data processing, the demand for the performance of computing devices is increasing. To meet these demands, due to its excellent parallel computing capabilities and high throughput, the GPU (Graphics Processing Unit) has gradually become a key acceleration device in high performance computing systems. Compared with traditional CPUs (Central Processing Units), GPUs can process large amounts of data streams at the same time, and are particularly suitable for tasks such as deep learning training, image processing, and complex scientific computing. However, existing GPU acceleration cards are usually provided by multiple manufacturers, and there are significant differences in hardware architectures, software development interfaces (Application Programming Interfaces, APIs), and data transmission methods among different manufacturers, resulting in many challenges in the development and optimization of cross-platform applications.
[0003] Some existing solutions attempt to simplify the compatibility issues of different manufacturers' hardware through virtualization technology or containerization technology. However, these methods can only achieve hardware abstraction and resource management to a certain extent, and cannot completely solve the problem of unified processing of data storage and transmission between different manufacturers' GPU acceleration cards. In addition, virtualization technology usually incurs certain performance overheads, and existing solutions will further limit the efficiency of high performance computing applications. Therefore, there is an urgent need for a unified processing method that can be compatible with different manufacturers' GPUs and efficiently process data storage and transmission.
[0004] Regarding the problem in the related art of how to improve data processing and transmission efficiency while being compatible with different manufacturers' GPU acceleration cards, no effective solution has been proposed yet. Summary of the Invention
[0005] In this embodiment, a method, apparatus, electronic device, and storage medium for managing hybrid acceleration cards are provided to solve the problem in the related art of how to improve data processing and transmission efficiency while being compatible with different manufacturers' GPU acceleration cards.
[0006] In a first aspect, in this embodiment, a method for managing hybrid acceleration cards is provided, which is applied to a server system. The server system includes a server. The method includes:
[0007] In response to the received server data transfer request, determine the runtime library interface of the hybrid acceleration card;
[0008] Based on the Remote Direct Memory Access (RDMA) technology, register the memory address of the hybrid acceleration card to a preset hardware abstraction layer; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface;
[0009] Through the preset hardware abstraction layer, call the runtime library interface to execute the data transfer operation between the hybrid acceleration cards.
[0010] In some embodiments, the hybrid acceleration card includes multiple acceleration cards; before determining the runtime library interface of the hybrid acceleration card in response to the received server data transfer request, it includes:
[0011] Determine multiple acceleration cards corresponding to the runtime library linked to the server;
[0012] Establish an adapter management component to manage the hardware adapters of the acceleration cards through the adapter management component; wherein, the hardware adapter is used to link the acceleration card and the server.
[0013] In some embodiments, determining the runtime library interface of the hybrid acceleration card in response to the received server data transfer request includes:
[0014] In response to the received server data transfer request, determine the runtime library type of the hybrid acceleration card according to the environment information of the server system environment;
[0015] Based on the runtime library type of the hybrid acceleration card, determine the runtime library interface of the hybrid acceleration card.
[0016] In some embodiments, after determining the runtime library type of the hybrid acceleration card according to the environment information of the server system environment in response to the received server data transfer request, it further includes:
[0017] Judge whether the server system environment of the hybrid acceleration card includes a single server or multiple servers;
[0018] When it is judged that the server system environment of the hybrid acceleration card includes a single server, number the runtime library type of the hybrid acceleration card to obtain a library type number;
[0019] Match the library type number and the current thread identifier corresponding to the server data transfer request in the form of a key-value pair.
[0020] In some of these embodiments, the method further includes:
[0021] Based on a preset calling method, load the dynamic link library of the runtime library of the acceleration card in the server system;
[0022] Obtain the handle of the runtime library of the acceleration card;
[0023] Determine the runtime library interface function of the acceleration card from the dynamic link library according to the handle of the runtime library;
[0024] Bind the runtime library interface function to the preset unified hardware abstraction interface.
[0025] In some of these embodiments, the registering the memory address of the hybrid acceleration card to a preset hardware abstraction layer based on remote direct memory access technology includes:
[0026] Obtain the data memory addresses of multiple acceleration cards; the multiple acceleration cards include the memory addresses for receiving data and / or sending data;
[0027] Determine whether there is a memory direct access module corresponding to the acceleration card in the server system;
[0028] When there is a memory direct access module corresponding to the acceleration card in the server system, perform memory registration in the hardware abstraction layer according to the data memory address of the acceleration card.
[0029] In some of these embodiments, the determining whether there is a memory direct access module corresponding to the acceleration card in the server system further includes:
[0030] When there is no memory direct access module corresponding to the acceleration card in the server system, convert the data memory address of the acceleration card into a virtual address according to a preset mapping method;
[0031] Perform memory registration in the hardware abstraction layer according to the virtual address.
[0032] In a second aspect, a hybrid acceleration card management device is provided in this embodiment. The device includes: a response module, a registration module, and a transmission module;
[0033] The response module is configured to determine the runtime library interface of the hybrid acceleration card in response to a server data transmission request;
[0034] The registration module is configured to perform remote direct memory registration of the memory address of the hybrid acceleration card to a preset hardware abstraction layer based on remote direct memory access technology; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface;
[0035] The transmission module is used to call the runtime library interface to perform data transmission operations between the hybrid acceleration cards according to the preset hardware abstraction layer.
[0036] In a third aspect, an electronic device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the hybrid acceleration card management method described in the first aspect above is implemented.
[0037] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the hybrid acceleration card management method described in the first aspect above is implemented.
[0038] Compared with the related art, a hybrid acceleration card management method, device, electronic device, and storage medium provided in this embodiment bind the runtime library interfaces of multiple acceleration cards to a preset unified hardware abstraction interface, manage different hardware acceleration cards through the unified interface to be compatible with GPU acceleration cards from different manufacturers; at the same time, based on the method of remote direct memory access registration, the data transmission of the hybrid acceleration card device memory is optimized to improve data processing and transmission efficiency.
[0039] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0041] Figure 1 is a hardware structure block diagram of a terminal of the hybrid acceleration card management method provided by an embodiment of the present application;
[0042] Figure 2 is a flowchart of the hybrid acceleration card management method provided by an embodiment of the present application;
[0043] Figure 3 is a schematic diagram of a unified abstract device operation access interface for hybrid acceleration cards provided by this specific embodiment;
[0044] Figure 4 is a schematic diagram of a method for unified processing of hybrid acceleration card hardware abstraction and data transmission in this specific embodiment;
[0045] Figure 5 is a flowchart of a unified method for the hardware abstraction layer interface provided by this specific embodiment;
[0046] Figure 6 It is a schematic diagram of the data transmission operation of the memory data of the acceleration card device provided in this specific embodiment;
[0047] Figure 7 It is a schematic diagram of the data transmission of the cross-server hybrid acceleration card device provided in this specific embodiment;
[0048] Figure 8 It is a structural block diagram of the hybrid acceleration card management device provided in an embodiment of the present application. Detailed implementation manners
[0049] To more clearly understand the purpose, technical solution and advantages of the present application, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.
[0050] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The term "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0051] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 It is a hardware structural block diagram of the terminal of the hybrid acceleration card management method provided in an embodiment of the present application. As Figure 1 shown, the terminal may include one or more ( Figure 1a processor 102 (only one is shown in the figure) and a memory 104 for storing data. The processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those Figure 1 shown in the figure, or have a different configuration from that Figure 1 shown in the figure.
[0052] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the hybrid acceleration card management method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0053] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0054] Currently, in the fields of computer hardware acceleration, data storage and transmission, GPU acceleration cards with excellent parallel computing capabilities and high throughput are usually adopted. However, the hardware architectures and programming interfaces of GPUs from different manufacturers have their own characteristics, and the underlying implementation mechanisms of the platforms vary greatly. Developers need to write specific codes to adapt to the APIs of each manufacturer when using different GPUs in application programs, which increases the complexity of development. At the same time, the hardware abstraction layers of different manufacturers lack generality, making it more difficult to achieve unified hardware abstraction and data transmission processing in a multi-vendor hybrid acceleration card environment, resulting in poor portability of applications and high maintenance costs.
[0055] To address the compatibility issues of multi-vendor hybrid GPUs, the industry has proposed various cross-platform hardware abstraction schemes and programming interfaces. To some extent, these technologies have achieved the abstraction and encapsulation of different GPU acceleration cards. However, they often require a large amount of adaptation work in actual use and still have performance bottlenecks, making it difficult to meet the requirements of efficient parallel computing. In addition, some cross-platform hardware abstraction schemes and programming interfaces exhibit unstable performance when executing complex deep learning tasks, limiting their application in some high-performance computing scenarios.
[0056] In addition, in terms of data transmission, the data transmission efficiency in cross-platform GPU systems has a significant impact on the overall computing performance. Especially in the case of mixed use of multiple acceleration cards, data transmission between different GPUs is often restricted by factors such as interface differences, inconsistent data formats, and different memory management mechanisms, resulting in low data transmission efficiency and low resource utilization in the system. Taking the training process of a deep learning model as an example, data needs to be frequently transmitted between different acceleration cards. If the data transmission efficiency is not high, it will directly affect the training speed and overall performance of the model, and may even cause data transmission to become the bottleneck of the system. Therefore, implementing a unified processing mechanism for data transmission and resource management between different acceleration cards is crucial for improving the resource utilization and overall efficiency of multi-acceleration card systems.
[0057] Some existing solutions attempt to simplify the compatibility issues of hardware from different vendors through virtualization technology or containerization technology. However, these methods can only achieve hardware abstraction and resource management to a certain extent and cannot completely solve the problem of unified processing of data storage and transmission between GPU acceleration cards from different vendors. In addition, virtualization technology usually has a certain performance overhead, which may further limit the efficiency of high-performance computing applications. Therefore, there is an urgent need for a unified processing method and device that can be compatible with GPUs from different vendors and efficiently process data storage and transmission to reduce the adaptation workload of developers and improve the computing performance and resource utilization of heterogeneous computing systems.
[0058] Therefore, in the existing technology, when using multiple hybrid acceleration cards, it is necessary to call the drivers and runtime libraries of their respective acceleration cards, which is rather inconvenient. For example, the memory of each acceleration card cannot achieve unified memory management. When performing operations such as memory application and release, it is necessary to separately link and call the runtime library interfaces of each acceleration card, which results in very low efficiency in the use of hybrid acceleration cards and a large bottleneck in the data transmission rate between hybrid cards.
[0059] To solve the above problems, this paper proposes a method and device for unified processing of hybrid accelerator card hardware abstraction and data transmission. By constructing a unified runtime API - runtime API, this method not only supports GPU accelerator cards from multiple vendors, but also provides an efficient and unified processing mechanism for data storage and transmission. In this way, developers can simplify the development process, reduce the complexity of data transmission and resource management between different accelerator cards, and improve the overall performance and computing efficiency of multi - accelerator - card heterogeneous systems.
[0060] Based on the above problems, in this embodiment, a method for managing hybrid accelerator cards is provided, which is applied to a server system including a server. Figure 2 It is a flowchart of the hybrid accelerator card management method provided by an embodiment of this application, as Figure 2 shown, and this process includes the following steps:
[0061] Step S210, in response to a received server data transmission request, determine the runtime library interface of the hybrid accelerator card.
[0062] In this step, when the server system receives a server data transmission request, in response to this server data transmission request, determine the hybrid accelerator card corresponding to the server for data transmission according to the data transmission request. The hybrid accelerator card includes accelerator cards on multiple different servers. For different types of accelerator cards on each server, install their corresponding SDK (Software Development Kit) development kits on the server, such as accelerator card drivers and runtime libraries, etc. At the same time, determine the runtime library interface of the hybrid accelerator card according to the data type and transmission type in the data transmission request, so as to facilitate subsequent data transmission between servers by calling the runtime library interface.
[0063] Among them, before determining the runtime library interface of the hybrid accelerator card, it includes: determining multiple accelerator cards corresponding to the runtime library linked to the server; establishing an adapter management component, and managing the hardware adapters of the accelerator cards through the adapter management component; wherein, the hardware adapter is used to link the accelerator card and the server.
[0064] Specifically, first, sort out the API interfaces of the runtime libraries of different types of accelerator cards on the server in the server system, classify and summarize them according to the types of different accelerator cards, so as to facilitate subsequent binding to the unified hardware abstraction interface. Then, check whether the dynamic link libraries of the runtime libraries corresponding to each accelerator card can be correctly linked by the server system, and establish an adapter management component. Manage the hardware adapters corresponding to the accelerator cards correctly linked to the server system through the adapter management component name. Manage the accelerator card runtime library interface based on this adapter management component to realize the initialization of the unified hardware abstraction interface. Subsequently, if more accelerator card runtime libraries need to be supported in the future, only need to add the hardware adapters of the new accelerator card runtime library in the adapter management component, without modifying the existing interface functions, to achieve seamless switching and compatible expansion of interface calls for different types of accelerator cards.
[0065] Step S220: Based on the remote direct memory access technology, register the memory address of the hybrid accelerator card to the preset hardware abstraction layer; there is a unified hardware abstraction interface preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface.
[0066] In this step, after determining the corresponding hybrid accelerator card based on the server data transfer request, based on the remote direct memory access technology, register the memory address of the hybrid accelerator card to the preset hardware abstraction layer; establish a remote direct memory access connection between the hybrid accelerator cards in the server through the hardware abstraction layer, and execute the write or read operation of the remote direct memory access data, that is, the data to be transferred in the server data transfer request, according to the data source and target address after the remote direct memory access memory registration. By encapsulating the write and read related operations of the remote direct memory access data into the hardware abstraction layer, and specifically encapsulating them in the unified hardware abstraction interface preset in the hardware abstraction layer, so as to call the remote direct memory access related interfaces corresponding to the hybrid accelerator card through the hardware abstraction layer, thereby optimizing the data transfer of the hybrid accelerator card device memory, which is beneficial to improving the efficiency of subsequent data transfer between hybrid accelerator cards.
[0067] Among them, Remote Direct Memory Access (RDMA for short) is a technology that allows one computer to directly access the memory of another computer without involving the host processor. This technology is mainly used for high-speed network communication, especially in high-performance computing (HPC) and data center environments to achieve efficient data transmission. The runtime library is a set of program libraries that provide various services and functions required by a program during runtime. These services usually include memory management, input / output operations, system call interfaces, exception handling, string processing, mathematical calculations, etc. The runtime library is an indispensable part of a program's runtime. They are usually provided by the compiler and are closely related to a specific programming language or compiler.
[0068] Step S230: Through a preset hardware abstraction layer, call the runtime library interface to perform data transfer operations between hybrid acceleration cards.
[0069] In this step, in the hardware abstraction layer, call the runtime library interface of the hybrid acceleration card through the hardware abstraction layer interface to establish a connection between the remote direct memory access network cards corresponding to the hybrid acceleration cards, and perform data transfer operations between the hybrid acceleration cards to achieve unified processing of data transfer between different hybrid acceleration cards.
[0070] Through the above steps, by binding the runtime library interfaces of multiple acceleration cards to a preset unified hardware abstraction interface, manage different hardware acceleration cards through a unified interface to be compatible with GPU acceleration cards from different manufacturers; at the same time, based on the method of remote direct access memory registration, optimize the data transfer of the hybrid acceleration card device memory to improve data processing and transfer efficiency.
[0071] In some of these embodiments, determining the runtime library interface of the hybrid acceleration card in step S210 in response to the received server data transfer request includes: in response to the received server data transfer request, determine the runtime library type of the hybrid acceleration card according to the environmental information of the server system environment; based on the runtime library type of the hybrid acceleration card, determine the runtime library interface of the hybrid acceleration card.
[0072] In this embodiment, it is determined whether the server system environment of the hybrid acceleration card includes a single server or multiple servers; when it is determined that the server system environment of the hybrid acceleration card includes multiple servers, that is, when the server system includes multiple servers with multiple hybrid types of acceleration cards and there is only one type of acceleration card in a single server among the multiple servers, the type of the hybrid acceleration card runtime library is automatically selected according to the environment information of the server system environment. Exemplarily, the environment information of the server system environment may be PCIe (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) device information or the driver information of the system kernel.
[0073] Further, in some of these embodiments, after determining the runtime library type of the hybrid acceleration card according to the environment information of the server system environment in response to the received server data transmission request, it further includes: determining whether the server system environment of the hybrid acceleration card includes a single server or multiple servers; when it is determined that the server system environment of the hybrid acceleration card includes a single server, numbering the runtime library type of the hybrid acceleration card to obtain a library type number; and matching the library type number and the current thread identifier corresponding to the server data transmission request in the form of a key-value pair.
[0074] Among them, when it is determined that the server system environment of the hybrid acceleration card includes a single server, that is, when the usage scenario is that there are multiple hybrid types of acceleration cards in a single server, the type of the corresponding acceleration card runtime library needs to be manually selected. Specifically, after selecting the acceleration card runtime library corresponding to the hybrid acceleration card, a hash table is established, and the current runtime library type of the hybrid acceleration card and the current thread identifier number corresponding to the data transmission request are matched in the form of a key-value pair; when calling the relevant interface of the runtime library, the unified hardware abstraction interface of the corresponding hardware abstraction layer will look up the currently selected runtime library type of the hybrid acceleration card according to the hash table.
[0075] In some of these embodiments, a hybrid acceleration card management method further includes: loading the dynamic link library of the runtime library of the acceleration card in the server system based on a preset calling method; obtaining the handle of the runtime library of the acceleration card; determining the runtime library interface function of the acceleration card from the dynamic link library according to the runtime library handle; and binding the runtime library interface function to a preset unified hardware abstraction interface.
[0076] In this embodiment, the method for binding the runtime interface of the unified hardware abstraction interface of the preset hardware abstraction layer to the selected accelerator card runtime library interface includes loading the dynamic link library of the accelerator card runtime library in the server system path by means of explicit dynamic library call to obtain the runtime library handle of the accelerator card; searching for and obtaining the runtime library interface function of the corresponding accelerator card from the runtime library according to the runtime library handle; and then binding the obtained runtime interface of the corresponding accelerator card to the predefined unified hardware abstraction runtime interface, so as to facilitate subsequent calling of the corresponding accelerator card runtime library interface through the unified hardware abstraction layer runtime interface, and achieve the effect of high-speed data transmission between different types of accelerator cards.
[0077] In some of these embodiments, in step S220, based on the Remote Direct Memory Access (RDMA) technology, registering the memory address of the hybrid accelerator card to the preset hardware abstraction layer includes: obtaining the data memory addresses of multiple accelerator cards; the multiple accelerator cards include the memory addresses for receiving data and / or sending data; determining whether there is a memory direct access module corresponding to the accelerator card in the server system; and when there is a memory direct access module corresponding to the accelerator card in the server system, performing memory registration in the hardware abstraction layer according to the data memory address of the accelerator card.
[0078] Among them, determining whether there is a memory direct access module corresponding to the accelerator card in the server system further includes: when there is no memory direct access module corresponding to the accelerator card in the server system, converting the data memory address of the accelerator card into a virtual address according to the preset mapping method; and performing memory registration in the hardware abstraction layer according to the virtual address.
[0079] In this embodiment, when performing data transmission between the device memories of the hybrid accelerator card, if there is a kernel driver module for direct access to the accelerator card device memory in the server system for the selected accelerator card, the virtual address corresponding to the memory allocated by the accelerator card is registered for RDMA (Remote Direct Memory Access) memory through the hardware abstraction layer interface; if there is no kernel driver module for direct access to the accelerator card device memory in the system for the selected accelerator card, the memory allocated by the accelerator card is mapped to the CPU of the server through the hardware abstraction layer interface to obtain its virtual address, and then RDMA memory registration is performed according to the virtual address. Then, an RDMA connection is established through the hardware abstraction layer interface, and an RDMA data write operation or read operation is performed according to the data source address and target address after RDMA memory registration.
[0080] In one possible embodiment, when there is a single server in the server system and there is only one type of acceleration card in the single server, the type of the acceleration card runtime library is automatically selected by the system environment information of the server system. The system environment can be PCIe device information or the driver information of the system kernel.
[0081] The following describes and illustrates this embodiment through specific examples.
[0082] This specific embodiment proposes a method for unified processing of hybrid acceleration hardware abstraction and data transmission, which can enable developers to improve the usage efficiency of hybrid acceleration cards and the data transmission speed to a certain extent, and reduce the burden on developers to handle hardware differences. First, for different types of acceleration cards on each server in the server system, install their corresponding SDK (Software Development Kit) development kits on the server, such as acceleration card drivers and runtime libraries. Secondly, sort out and classify the runtime library function interfaces of each acceleration card, which can be roughly divided into device interfaces, error reporting interfaces, memory interfaces, stream and event interfaces, graph interfaces, driver interfaces, context driver interfaces, memory and device driver interfaces, etc. After sorting out various types of hybrid acceleration card interface categories, define a unified hardware abstraction interface. Figure 3 It is a schematic diagram of the unified abstract device runtime access interface of the hybrid acceleration card provided in this specific embodiment. Refer to Figure 3 , the unified hardware abstraction device runtime interface includes a memory management interface, a device query interface, an image processing interface, an exception handling interface, and a stream and event interface; at the same time, when calling the hardware abstraction initialization interface, an adapter management component will be established to provide a unified interface to manage different hardware acceleration card adapters. Subsequently, if more acceleration card runtime libraries need to be supported, only new adapters need to be added to the adapter management component without modifying the existing interface functions.
[0083] Figure 4 It is a schematic diagram of the method for unified processing of hybrid acceleration card hardware abstraction and data transmission in this specific embodiment. Refer to Figure 4 , this method includes steps S11 to S15.
[0084] Step S11: Sort out the interfaces of each acceleration card runtime library to facilitate subsequent binding with the unified hardware abstraction runtime interface.
[0085] Step S12: Call the hardware abstraction initialization interface.
[0086] Specifically, the invocation of the hardware abstraction initialization interface specifically includes: checking whether the dynamic link libraries of the runtime libraries corresponding to each accelerator card can be correctly linked to the system; establishing an adapter management component to provide a unified interface for managing different hardware accelerator card adapters. Refer to Figure 3 , and manage the A-type computing accelerator card, B-type computing accelerator card, and X-type computing accelerator card, the A-type accelerator card software development kit (SDK), B-type accelerator card software development kit (SDK), and X-type computing accelerator card software development kit (SDK) respectively through the adapter management component.
[0087] Step S13: Select the type of the required accelerator card runtime library according to the data processing request.
[0088] Specifically, according to the server data transmission request, determine the type of the hybrid accelerator card runtime library to be used. Further, for the hybrid accelerator cards in the server, that is, the situation where multiple accelerator cards coexist and are compatible in a single server, after invoking the hardware abstraction initialization interface, it is necessary to select the corresponding runtime library type according to the type of accelerator card required by the current thread. And after the accelerator card initialization interface is invoked, it takes effect for the entire process, and each platform only needs to execute it once. The accelerator card types that are not needed can be not initialized. The accelerator card runtime library type selection function interface is for the current thread. If necessary after each thread is created, this function can be called, and the runtime library type can be changed at any time. The implementation principle is that when a runtime library type is selected, a hash table is established, and then the currently selected library type number and the thread ID number are matched as key-value pairs. When invoking the interfaces related to the runtime library, such as the memory allocation interface, the corresponding abstract layer memory allocation interface will look up the currently selected runtime library type according to the hash table, so as to achieve seamless switching of the hybrid accelerator card runtime library.
[0089] In addition, for the hybrid accelerator cards with multiple servers, but for the situation where there is only one type of accelerator card in a single server, the hardware abstraction layer and its corresponding runtime library can be directly bound by querying the system environment information on the server. This environment information can be PCIe device information or can be judged through system driver information.
[0090] Step S14: Bind the defined hardware abstraction layer runtime interface to the selected accelerator card runtime library interface.
[0091] Specifically, when the matching accelerator card runs the runtime library, the interfaces of the hardware abstraction layer are bound to the interfaces of the matching accelerator card runtime library. By means of explicit dynamic library calls, the dynamic link library of the corresponding accelerator card runtime library is loaded in the server system path, the handle of the runtime library is obtained, and the interface function of the corresponding accelerator card runtime library is searched for and obtained from the runtime library according to the handle, and then this function is bound to the predefined unified hardware abstraction runtime interface.
[0092] Step S15: Call the corresponding accelerator card runtime library interface through the unified hardware abstraction layer runtime interface to perform high-speed data transmission for data between different types of accelerator cards.
[0093] Among them, in the prior art, there are usually certain speed bottlenecks in the data transmission between the memory of hybrid accelerator card devices. In scenarios with relatively high requirements for computing real-time performance and efficiency, ordinary Ethernet data transmission often cannot meet the requirements. For example, for the data transmission of the memory of accelerator card devices between cross-server machines, the traditional approach, whether it is the sending end or the receiving end, generally copies the data in the accelerator card device memory to the CPU memory first and then performs Ethernet data transmission. In the case of a large amount of data, due to the two additional processes of copying data from the accelerator card device memory to the CPU memory, the transmission rate is relatively low.
[0094] This application optimizes the data transmission of hybrid accelerator card devices based on RDMA (Remote Direct Memory Access). Among them, RDMA-related operations, such as the establishment of RDMA connections, the write operation and read operation of the accelerator card device memory of RDMA, are directly implemented at the hardware abstraction layer. The RDMA-related interfaces can be directly called through the hardware abstraction layer, and the implementation of these interfaces also depends on the RDMA-related core libraries and tool sets.
[0095] In one specific embodiment, the specific implementation steps for memory data transmission between hybrid acceleration card devices in a server system are as follows: For the data sender, first obtain the data address and size on the acceleration card that need to be transmitted across servers, and then call the hardware abstraction layer runtime interface for RDMA memory registration. If the currently selected acceleration card has a kernel driver module for direct access to the memory of the acceleration card device installed on its server, when using the hardware abstraction layer runtime interface, the dynamically allocated memory address will be directly registered for RDMA memory; if there is no kernel driver module for direct access to the memory of the acceleration card device, when calling the hardware abstraction layer runtime interface, the memory address of the data to be transmitted on the acceleration card device will be converted into a virtual address in the user space through memory mapping, and then RDMA memory registration will be performed; similarly, for the data receiver, it is necessary to call the hardware abstraction layer runtime interface to dynamically allocate a block of acceleration card device memory for reception, and at the same time, this memory address also needs to perform corresponding RDMA memory registration operations. After obtaining the data source address and target address after RDMA memory registration, call the hardware abstraction layer interface to execute the establishment of an RDMA connection and the RDMA write operation, and the high-speed transmission of data between hybrid acceleration cards across servers can be achieved.
[0096] Further, on the premise of the above steps, after each server node completes data copying using the hardware abstraction layer, it can execute the corresponding kernel function to complete the parallel computing task. At the same time, when the hybrid acceleration cards on each server node are executing the parallel computing task, they can also exchange data with each other through RDMA to complete the synchronization of calculation results.
[0097] Through the method for unified processing of hybrid acceleration card hardware abstraction and data transmission provided by this specific embodiment, in the current development environment where there are multiple hybrid acceleration cards inside a single server machine, when initializing through the runtime library of the hardware abstraction layer, only the type of acceleration card currently needed to be used needs to be selected. After selection, the memory copy function interface of the hardware abstraction layer runtime library can be directly called to copy the data that needs to be parallel computed into the memory of the acceleration card device. Optionally, the type of acceleration card can also be switched according to its own needs, so as to simultaneously complete the memory copy operation for the hybrid acceleration card. Through the data processing method of this application, it can be compatible with a variety of hardware accelerators, improve the data transmission speed between hybrid acceleration cards, reduce the burden on developers to handle hardware differences, and thus improve the overall system performance and efficiency.
[0098] Figure 5 is the flowchart of the unified method of the hardware abstraction layer interface provided by this specific embodiment, as Figure 5As shown, the unified method for the hardware abstraction layer interface includes: First, initialize the hardware abstraction interface and determine whether there is a hybrid accelerator card in the server system, that is, whether there is a hybrid accelerator card in a single server; if so, select the type of runtime library currently needed, and at the same time match the current runtime library type and the current thread number, that is, the thread ID number, where the current thread ID number is the current thread identifier in the foregoing embodiment. Thereafter, bind the hardware abstraction layer runtime interface to the selected accelerator card runtime library interface. Among them, each hardware abstraction layer runtime interface will match the corresponding accelerator card runtime library according to the thread ID, and at the same time will check the initialization status of the runtime library. If not, there is no need to select the accelerator card runtime library type, and directly select the accelerator card runtime library type automatically according to the system environment information of the server system, and then bind the hardware abstraction layer runtime interface to the selected accelerator card runtime library interface. Finally, regardless of whether there is a hybrid accelerator card in the machine, it is necessary to call the runtime interface of the corresponding accelerator card through the hardware abstraction layer interface to achieve the data transmission of the hybrid accelerator card; when the data transmission is completed, that is, after multiple related interfaces are used, the relevant resources will be released to avoid resource occupation.
[0099] Figure 6 is a schematic diagram of the data transmission operation of the accelerator card device memory in this specific embodiment. Refer to Figure 6 , when performing a data transmission operation through a hybrid accelerator card device, first obtain the memory addresses for sending data of the type A accelerator card and the type B accelerator card respectively, and determine whether there is a kernel driver module for direct memory access of the type A accelerator card and / or the type B accelerator card memory in the server system; among them, the type A accelerator card is used to send data, and the type B accelerator card is used to receive data. If there is a kernel driver module for direct memory access of the accelerator card memory, then directly use the obtained memory address for remote direct memory access (RDMA) memory registration, and call the hardware abstraction layer to perform the data transmission operation of remote direct memory access (RDMA). If there is no kernel driver module for direct memory access of the accelerator card memory, then it is necessary to use the memory mapping method to convert both the sending data memory address and the receiving data memory address into virtual addresses, and then perform remote direct memory access (RDMA) memory registration, and call the hardware abstraction layer to perform the data transmission operation of remote direct memory access (RDMA).
[0100] Figure 7 is a schematic diagram of cross-server hybrid accelerator card device data transmission in this specific embodiment. Refer to Figure 7, wherein, different types of Class A acceleration cards and Class B acceleration cards are respectively connected to the RDMA network card through the expansion card pcie switch; the RDMA network card obtains the system memory of the server system through the processor CPU; an executable code is stored in the processor, and when the executable code is executed, it is used to implement the hybrid acceleration card hardware abstraction and data transmission unified processing method in the above embodiments.
[0101] In this embodiment, a hybrid acceleration card management device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0102] Figure 8 is the structural block diagram of the hybrid acceleration card management device provided by the embodiment of the present application, as Figure 8 shown, the device includes: a response module 10, a registration module 20, and a transmission module 30.
[0103] The response module 10 is used to determine the runtime library interface of the hybrid acceleration card in response to a server data transmission request.
[0104] The registration module 20 is used to remotely directly access the memory address of the hybrid acceleration card to register it in the preset hardware abstraction layer based on the remote direct memory access technology; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface.
[0105] The transmission module 30 is used to call the runtime library interface to execute the data transmission operation between the hybrid acceleration cards according to the preset hardware abstraction layer.
[0106] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0107] In this embodiment, an electronic device is also provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0108] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, wherein, the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0109] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0110] S1. In response to a received server data transmission request, determine the runtime library interface of the hybrid acceleration card.
[0111] S2. Based on the remote direct memory access technology, register the memory address of the hybrid acceleration card to a preset hardware abstraction layer; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface.
[0112] S3. Through the preset hardware abstraction layer, call the runtime library interface to execute data transmission operations between hybrid acceleration cards.
[0113] It should be noted that for specific examples in this embodiment, reference may be made to the examples described in the above-mentioned embodiment and optional implementation manners, and details will not be elaborated herein.
[0114] In addition, in combination with the hybrid acceleration card management method provided in the above-mentioned embodiment, a storage medium may also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by the processor, any one of the hybrid acceleration card management methods in the above-mentioned embodiment is implemented.
[0115] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0116] Obviously, the drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations according to these drawings without creative work. In addition, it can be understood that although the work done during the development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.
[0117] The term "embodiment" in the present application means that the specific features, structures, or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0118] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A hybrid accelerator card management method, characterized in that: Applied in a server system, the server system includes a server; the method includes: Determine multiple accelerator cards corresponding to the runtime library linked to the server; establish an adapter management component, and manage the hardware adapter of the accelerator card through the adapter management component; wherein the hardware adapter is used to link the accelerator card and the server; the establishment of the adapter management component includes: sorting out the API interfaces of the runtime libraries of different types of accelerator cards in the server system, and classifying and summarizing them according to the types of different accelerator cards; checking whether the dynamic link library of the runtime library corresponding to each accelerator card can be correctly linked by the server system, and establishing the adapter management component; the management of the hardware adapter of the accelerator card through the adapter management component includes: managing the hardware adapter corresponding to the accelerator card correctly linked to the server system through the adapter management component name, and managing the accelerator card runtime library interface based on the adapter management component; In response to a received server data transmission request, determining a runtime library type of a hybrid acceleration card according to environmental information of the server system environment; determining whether the server system environment of the hybrid acceleration card includes a single server or multiple servers; when determining that the server system environment of the hybrid acceleration card includes multiple servers, automatically selecting the type of the runtime library of the hybrid acceleration card according to the environmental information of the server system environment; when determining that the server system environment of the hybrid acceleration card includes a single server, numbering the runtime library type of the hybrid acceleration card to obtain a library type number; matching the library type number and the current thread identifier corresponding to the server data transmission request in a key-value pair manner; determining a runtime library interface of the hybrid acceleration card based on the runtime library type of the hybrid acceleration card; Based on remote direct memory access technology, the memory address of the hybrid acceleration card is registered to a preset hardware abstraction layer; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface; The preset hardware abstraction layer is used to call the runtime library interface to perform data transmission operations between hybrid acceleration cards.
2. The hybrid accelerator card management method according to claim 1, characterized in that: The method further comprises: Based on a preset calling method, loading the dynamic link library of the runtime library of the accelerator card in the server system; Obtaining the runtime library handle of the accelerator card; Determining a runtime library interface function of the accelerator card from the dynamic link library according to the runtime library handle; The runtime library interface function is bound to the preset unified hardware abstract interface.
3. The hybrid accelerator card management method according to claim 1, characterized in that: The method of registering the memory address of the hybrid accelerator card to a preset hardware abstraction layer based on remote direct memory access technology includes: Acquire data memory addresses of multiple acceleration cards; the multiple acceleration cards include memory addresses for receiving data and / or sending data; Determine whether there is a memory direct access module corresponding to the acceleration card in the server system; When a memory direct access module corresponding to the acceleration card exists in the server system, memory registration is performed in the hardware abstraction layer according to the data memory address of the acceleration card.
4. The hybrid accelerator card management method according to claim 3, characterized in that: The determining whether there is a memory direct access module corresponding to the acceleration card in the server system also includes: When the memory direct access module corresponding to the acceleration card does not exist in the server system, the data memory address of the acceleration card is converted into a virtual address according to a preset mapping method; Memory registration is performed at the hardware abstraction layer according to the virtual address.
5. A hybrid accelerator card management device, characterized in that: The device comprises: a response module, a registration module and a transmission module; The response module is used to determine multiple accelerator cards corresponding to the runtime library linked to the server; establish an adapter management component, and manage the hardware adapter of the accelerator card through the adapter management component; wherein the hardware adapter is used to link the accelerator card and the server; the establishment of the adapter management component includes: sorting out the API interfaces of the runtime libraries of different types of accelerator cards in the server system, and classifying and summarizing them according to the types of different accelerator cards; checking whether the dynamic link library of the runtime library corresponding to each accelerator card can be correctly linked by the server system, and establishing an adapter management component; the management of the hardware adapter of the accelerator card through the adapter management component includes: managing the hardware adapter corresponding to the accelerator card correctly linked to the server system through the adapter management component name, and based on the adapter management component, The hybrid acceleration card is configured to manage the runtime library interface of the hybrid acceleration card; it is also configured to respond to a server data transmission request and determine the runtime library type of the hybrid acceleration card according to the environmental information of the server system environment; it is determined whether the server system environment of the hybrid acceleration card includes a single server or multiple servers; when it is determined that the server system environment of the hybrid acceleration card includes multiple servers, the type of the runtime library of the hybrid acceleration card is automatically selected according to the environmental information of the server system environment; when it is determined that the server system environment of the hybrid acceleration card includes a single server, the runtime library type of the hybrid acceleration card is numbered to obtain a library type number; the library type number and the current thread identifier corresponding to the server data transmission request are matched in a key-value pair manner; based on the runtime library type of the hybrid acceleration card, the runtime library interface of the hybrid acceleration card is determined; The registration module is used to register the memory address of the hybrid acceleration card to a preset hardware abstraction layer for remote direct memory access based on remote direct memory access technology; a unified hardware abstraction interface is preset in the hardware abstraction layer; the runtime library interface is bound to the unified hardware abstraction interface; The transmission module is used to call the runtime library interface to perform data transmission operations between hybrid acceleration cards according to the preset hardware abstraction layer.
6. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the hybrid accelerator card management method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the hybrid accelerator card management method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Hardware calculation module, device and method, electronic device and storage medium
CN116627888A
Data access method, electronic equipment and computer readable storage medium
CN117785498A