Data transmission method and device

By using a network card for data transmission between the CPU and GPU, the latency and bandwidth issues caused by the long distance of the PCIe bus are resolved, and data transmission efficiency is improved.

CN115549858BActive Publication Date: 2025-09-30ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211066913.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-09-30
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

After the CPU and GPU are decoupled, the PCIe bus is extended, resulting in longer data transmission delays and lower bandwidth, affecting data transmission efficiency.

Method used

By using a network card for data transmission between the CPU and GPU, calling the semantically rewritten transfer function, avoiding direct transmission through the PCIe bus, using the network card for network transmission, combined with remote direct memory access technology, dynamically adjusting the transmission method to optimize the network status.

Benefits of technology

It improves data transmission efficiency, reduces the impact of PCIe bus latency on bandwidth, and improves data transmission performance from CPU to GPU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115549858B_ABST
    Figure CN115549858B_ABST
Patent Text Reader

Abstract

Embodiments of this specification provide a data transmission method and apparatus, wherein the data transmission method includes: determining, in response to a data processing request, the amount of to-be-processed data carried in the data processing request; and, if, based on the amount of data, the to-be-processed data is to be transmitted via a network interface card (NIC), invoking a corresponding first transfer function to send the to-be-processed data to a GPU via the NIC. By invoking the corresponding first transfer function when, based on the amount of data, the to-be-processed data is to be transmitted via the NIC, the method enables the to-be-processed data to be transmitted from the CPU to the GPU via the NIC, thereby avoiding data transmission via a bus and thereby improving data transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a data transmission method. Background Art

[0002] The development of artificial intelligence technology in applications such as video surveillance has led to an increase in the use of GPUs (graphics processing units). However, CPUs (central processing units) and GPUs are typically physically coupled in a fixed ratio, which cannot meet the growing demand for GPU computing power. Therefore, decoupling the CPU and GPU to achieve GPU pooling is commonly adopted, allowing a single CPU to access multiple GPUs within a GPU resource pool.

[0003] However, when the GPU is decomposed to form a GPU resource pool, the bus connecting the CPU and GPU will also be stretched further, which will lead to longer data transmission delays, reduced bandwidth, and increased data transmission time, thereby affecting data transmission efficiency. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a data transmission method. One or more embodiments of this specification also relate to a data transmission device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a data transmission method is provided for use by a CPU, the method comprising:

[0006] In response to a data processing request, determining the amount of data to be processed carried in the data processing request;

[0007] In the case where it is determined according to the data amount that the data to be processed is transmitted through the network card, a corresponding first transmission function is called to send the data to be processed to the GPU through the network card.

[0008] According to a second aspect of the embodiments of this specification, there is provided a data transmission device for a CPU, the device comprising:

[0009] a determination module configured to determine, in response to a data processing request, the amount of to-be-processed data carried in the data processing request;

[0010] The sending module is configured to call a corresponding first transmission function when it is determined according to the data volume that the data to be processed is to be transmitted through the network card, and send the data to be processed to the GPU through the network card.

[0011] According to a third aspect of an embodiment of this specification, a computing device is provided, including:

[0012] memory and processor;

[0013] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data transmission method are implemented.

[0014] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data transmission method are implemented.

[0015] According to a fifth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data transmission method.

[0016] An embodiment of the present specification provides a data transmission method, which determines, in response to a data processing request, the amount of data to be processed carried in the data processing request; when it is determined based on the data amount that the data to be processed is to be transmitted through a network card, calls a corresponding first transmission function to send the data to be processed to a GPU through the network card.

[0017] The above method determines that the data to be processed is transmitted through the network card, and calls the first transmission function corresponding to the transmission mode of the network card. Since the second transmission function that defines the transmission mode as transmission through the bus is updated, the CPU can call the updated first transmission function, so that the transmission mode of the data to be processed can be changed without the CPU's perception, and the data to be processed can be transmitted from the CPU to the GPU through the network card, changing the original bus transmission mode for data transmission from the CPU to the GPU, avoiding data transmission through the bus, and thus avoiding the problems of increased data transmission delay, reduced bandwidth, and increased data transmission time when data is transmitted through the bus, thereby improving data transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic diagram of GPU pooling provided by one embodiment of this specification;

[0019] Figure 2 This is a schematic diagram of the structure of a parallel computing program provided by one embodiment of this specification;

[0020] Figure 3 is a schematic diagram of GPU remote direct memory access communication provided by one embodiment of this specification;

[0021] Figure 4 This is a schematic diagram of GPU resource virtual pooling provided by one embodiment of this specification;

[0022] Figure 5 is a schematic diagram of a wired network technology provided by one embodiment of this specification;

[0023] Figure 6 is a schematic diagram of a data copy technology provided by an embodiment of this specification;

[0024] Figure 7 This is a schematic diagram of a specific application scenario of a data transmission method provided by an embodiment of this specification;

[0025] Figure 8 This is a flow chart of a data transmission method provided by one embodiment of this specification;

[0026] Figure 9 This is a flowchart of a data transmission method provided by an embodiment of this specification;

[0027] Figure 10 This is a structural diagram of a data transmission method provided by an embodiment of this specification;

[0028] Figure 11 is a schematic diagram of a transmission method provided by an embodiment of this specification;

[0029] Figure 12 This is a flowchart of a data transmission method according to an embodiment of the present invention.

[0030] Figure 13 This is a structural diagram of a data transmission device provided by an embodiment of this specification;

[0031] Figure 14 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0032] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0033] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0034] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0035] First, the terms involved in one or more embodiments of this specification are explained.

[0036] GPU pooling: Generally speaking, GPUs and CPUs (servers) are physically coupled in fixed ratios (such as 2, 4, 6, and 8). To meet the diverse needs of different users, cloud computing providers must prepare additional servers in various ratios. For example, when a user's computing power requirements are low, cloud computing providers can prepare GPUs and CPUs in a ratio of 2 to 1. However, when a user's computing power requirements are high, cloud computing providers may need to prepare GPUs and CPUs in a ratio of 6 to 1 or 8 to 1. This allows for full utilization of computing resources while meeting the user's computing power requirements and avoiding waste. However, a fixed combination of CPUs and GPUs is extremely inefficient in terms of resource utilization, upgrades, and maintenance. Furthermore, when high performance is required, the computing power of a single server is far from sufficient.

[0037] Figure 1 This is a schematic diagram of GPU pooling provided by an embodiment of this specification, see Figure 1 , Figure 1 In A, A represents a GPU server cluster, which consists of multiple GPU server nodes B. Each GPU server node consists of a CPU, a network card, a memory, and a graphics card, which are connected by a bus (PCIe). In A, the ratio of the number of CPUs and GPUs on each GPU server node B is fixed, such as Figure 1As shown in the figure, a GPU server node has one CPU and two GPUs. Each GPU server node B is connected via a network card. C represents a GPU resource pool. The process from A to C illustrates the GPU pooling process, which involves removing the GPUs from each GPU server node, decoupling them from the CPU, and placing them into the GPU resource pool. Each node after GPU decomposition is connected to the GPU resource pool via a bus (PCIe). This allows for dynamic allocation of GPU nodes based on actual user needs, enabling dynamic scaling and flexible release. Pooling, from top to bottom, includes pooling at the framework, application, driver, and bus layers. Bus pooling has the least reliance on upper-layer software, offers the highest adaptability, and requires no additional software overhead.

[0038] PCIe: PCI-Express, the full name of which is peripheral component interconnect express, is a high-speed serial computer expansion bus standard.

[0039] TLP (Transaction Layer Packet) is a data packet used for data transmission between the CPU and the PCIe bus, or between the PCIe bus and the target device. Data is transmitted from the transaction layer at the sender to the transaction layer at the receiver in the form of packets. Requests from the upper layer (software layer or application layer) are packaged into transaction layer packets based on their type, destination address, and other relevant attributes. These packets then pass through the data link layer and physical layer before reaching the target device.

[0040] Parallel computing program: By utilizing the processing power of the GPU, computing performance can be greatly improved. The CPU can use this parallel computing program to control the GPU.

[0041] Specifically, the parallel computing program is executed collaboratively by the CPU and GPU, with the CPU being called the host and the GPU being called the device. The CPU executes the host-side code, responsible for serially performing logical operations and transaction processing tasks, while the GPU executes the device-side code, responsible for highly parallel, large-scale computing tasks. The parallel computing program can include a library of various functions. When the user controls the GPU through the CPU, they can call the functions contained in the parallel computing program to implement the corresponding functionality.

[0042] Figure 2 This is a schematic diagram of the structure of a parallel computing program provided by an embodiment of this specification, see Figure 2, the parallel computing program includes a kernel-mode driver control library, a user-mode driver library, a package library for the user-mode driver library, and a general package library. The kernel-mode driver control library directly controls the GPU hardware and cannot be called directly by users; the user-mode driver library contains various GPU operation functions for the upper layer to call directly, and the kernel-mode driver control library is indirectly called by calling the user-mode driver library to control the GPU; the package of the user-mode driver library can be understood as a layer of package of the user-mode driver library to simplify the user development difficulty. It contains most of the GPU operation functions and omits some functions that directly operate the GPU context. All functions contained in the package library for the user-mode driver library can be found from the user-mode driver library. The general package library is a general term for all libraries that use the package library for the user-mode driver library to encapsulate functions; the application can be understood as user-written code that can be used to control the GPU, such as user-written code to control the GPU to implement deep learning tasks or image processing tasks.

[0043] Parallel computing programs provide two dynamic libraries within the user-mode driver library: one responsible for runtime functions (RuntimeAPI), and the other interpreting driver-level functions (DriverAPI). Applications can directly call either the wrapper library or the user-mode driver library, and the functions contained in both are largely interchangeable. Driver-level functions provide finer-grained control over context and module loading, requiring explicit function calls to specify execution configuration and kernel parameters, resulting in greater operability.

[0044] DMA: Direct Memory Access allows hardware devices of different speeds to communicate without relying on a large interrupt load from the CPU. It is the ability of a device to directly access the host memory without CPU intervention.

[0045] RDMA (Remote Direct Memory Access) is the ability to access (read and write) memory on a remote machine through network technology without interrupting CPU processing on that system. It can resolve server-side data processing delays during network transmission.

[0046] GDR: The full name of the English name is GPUDirectRDMA, which means GPU remote direct memory access. Figure 3 This is a schematic diagram of GPU remote direct memory access communication provided by an embodiment of this specification, see Figure 3Data between the CPU and GPU is transferred via a PCIe switch. GDR leverages the standard features of PCI Express to enable direct data exchange (P2P) between the GPU and third-party peer devices. Third-party peer devices include network interfaces (Network Interfaces), video acquisition devices (Video Acquisition Devices), and storage adapters (Storage Adapters). GDR enables peripheral PCIe devices (RDMA network cards) to directly access GPU memory, avoiding data copying within system memory, eliminating CPU bandwidth and latency bottlenecks, and providing direct communication capabilities for remote GPUs.

[0047] API: The full name of the English name is Application Program Interface, which can be understood as the function interface called when the CPU controls the GPU.

[0048] HtoD: Host (host side, refers to CPU) to Device (device side, refers to GPU), refers to the CPU sending data to the GPU. The larger the HtoD bandwidth, the faster the data transmission and the stronger the communication capability.

[0049] DtoH: Sends data from GPU to CPU.

[0050] Pageable memory: This refers to memory that can be paged and swapped by the operating system. Pageable memory is memory space allocated by the operating system on the host.

[0051] Page-locked memory: Pinned memory refers to memory that the operating system cannot page or swap, ensuring it always resides in physical memory. Page-locked memory can be allocated on the host by built-in functions of parallel computing programs.

[0052] In actual applications, the GPU and CPU in the server are physically coupled in a fixed ratio. Simple hardware stacking cannot meet the demand for increased computing power. In large-scale distributed training, the acceleration ratio will drop sharply as the number of machines increases. In addition, the fixed CPU / GPU ratio is prone to GPU resource fragmentation. The GPU and CPU cannot be decoupled or repaired separately during generation changes.

[0053] In current deployments, GPUs rely on the PCIe bus to connect to the CPU and use communication protocols to connect to other GPUs in the server for internal pooling. The current PCIe bus and its communication protocols only operate over short distances, limiting GPU pooling to within the server or at the rack level. Current approaches to GPU pooling, from top to bottom, primarily include: framework-level pooling, parallel computing program-level pooling, driver-level pooling, and PCIe bus-level pooling. The closer a pooling solution is to the underlying hardware, the less dependent it is on the upper-level software stack. PCIe bus-level pooling offers advantages in terms of maintenance workload, software overhead, and device compatibility.

[0054] However, during data transfer between the CPU and GPU pooled at the PCIe bus layer, direct memory access is initiated by the GPU. Data is transmitted from the CPU to the GPU via a bus read (PCIeRead). Data is transmitted via non-posted data packets (PCIeNon-postedTLPs), and the sender must receive a completion notification message from the receiver. However, as the bus is extended, the transmission latency of a single data packet (TLP) increases. Limited by the maximum number of data packets in transit, bandwidth is significantly reduced. Furthermore, each GPU in the GPU resource pool shares a single PCIe bus, significantly increasing the throughput from the CPU to the GPU and requiring higher bandwidth requirements.

[0055] In the prior art, different methods for GPU pooling are provided, such as: Figure 4 This is a schematic diagram of GPU resource virtual pooling provided by an embodiment of this specification, see Figure 4 , providing a network-accessible shared resource pool through the Bitfusion architecture to support artificial intelligence. The Bitfusion architecture is an architecture for pooling GPU resources to form a GPU resource pool, which is then shared for everyone's use. It includes a server and a client. The server can virtualize physical GPU resources and share them with multiple users. The client can be understood as a virtual machine, which can transmit the virtual machine's GPU service request to the server over the network. The server then sends the service request to the user-mode driver library in the parallel computing program for processing. Specifically, during the process of sending the service request to the user-mode driver library in the parallel computing program, the service request is intercepted by an inserted agent in the client and transmitted to the server over the network, realizing virtual pooling of GPU resources. However, this software-level pooling has many limitations, such as difficulty in adapting to different scenarios and high version maintenance costs.

[0056] or, Figure 5 This is a schematic diagram of a wired network technology provided by an embodiment of this specification, see Figure 5 When virtualizing the PCIe bus over Ethernet using wired network technology, data packets transmitted on the PCIe bus can be transmitted over Ethernet to achieve the effect of physically distanced GPUs and thus implement GPU pooling. However, this approach increases CPU-to-GPU latency and reduces the bandwidth for data transmission from the CPU to the GPU, thus affecting overall performance.

[0057] as well as, Figure 6 This is a schematic diagram of a data copy technology provided by an embodiment of this specification, see Figure 6 , provides a data copy technology for fast data transmission, Figure 6 The left image in the figure is a schematic diagram of data transmission using a parallel computing program, and the right image is a schematic diagram of data transmission using data copy technology. In comparison, the low-latency GPU memory copy library based on GPU remote direct memory access technology allows the CPU to directly map and access GPU memory. Data copy allows the CPU to directly access GPU memory through mapping, allowing low-latency copying between GPU and CPU memory. This data copy technology converts data transfer from CPU to GPU (HtoD) from reading data to writing data to achieve low latency. However, this method consumes CPU resources and only has a beneficial effect when the data being transferred is within 64KB. When implementing this technology, an additional kernel-mode driver must be installed and loaded on the target machine, which increases complexity and is opaque to the calling upper software. The transfer function must be changed to a specific function.

[0058] Therefore, an effective technical solution is urgently needed to solve the above problems.

[0059] In this specification, a data transmission method is provided. This specification also relates to a data transmission device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0060] See also Figure 7 , Figure 7 This is a schematic diagram of a specific application scenario of a data transmission method provided according to an embodiment of this specification.

[0061] Figure 7It includes CPU (host side, host) and GPU resource pool (device side, device) as well as user-level driver library. The GPU resource pool can be understood as physically decoupling the GPU from the CPU and merging it into the GPU resource pool. The GPU resource pool can include at least one GPU. The CPU is connected to the GPU resource pool so that the CPU can call any one or more GPUs in the GPU resource pool to dynamically adjust the resources that the CPU can call, thereby meeting the growing demand for GPU computing power. The CPU and GPU are each equipped with at least one network card. The CPU and GPU can transmit data through the PCIe bus, or perform remote direct memory access to the GPU through the network card. The user-level driver library can be understood as a function library, which contains all the functions that the user needs to call when controlling the GPU through the CPU.

[0062] During specific implementation, the user sends a data processing request to the CPU, and the data processing request carries the data to be processed. Alternatively, the CPU can obtain the data to be processed corresponding to the data processing request from the database based on the data processing request. According to the current network status of the CPU and GPU, when it is determined that the transmission method is to be transmitted through the network card, the corresponding transmission function is called from the user-layer driver library. The transmission function is a function that has been semantically rewritten and is used to indicate network transmission through the network card. The processed data is transmitted to the GPU through the network card, thereby avoiding data transmission through the PCIe bus in a direct memory access reading manner.

[0063] exist Figure 7 In the embodiment, the user sends an image processing request to the CPU, and the image processing request is used to control the GPU to perform image processing. The image processing request may carry image data corresponding to the image to be processed, or the CPU may obtain the image data corresponding to the image processing request from the database according to the image processing request. When it is determined that the transmission mode is network card transmission, the corresponding transmission function is called from the user layer driver library. The transmission function is a function that has been semantically rewritten, and is used to indicate remote direct data access through the network card, and transmit the image data to be processed to the GPU through the network card, thereby avoiding data transmission through the PCIe bus in a direct memory access reading mode, achieving an increase in bandwidth, and thus achieving an improvement in data transmission efficiency.

[0064] See also Figure 8 , Figure 8 A flowchart of a data transmission method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0065] Step 802: In response to a data processing request, determine the amount of to-be-processed data carried in the data processing request.

[0066] Among them, the data processing request can be understood as a request sent by the user to the CPU to control the GPU to process image data, such as a request to enhance image data or sharpen image data; the data to be processed can be understood as the image data corresponding to the image requested by the user to be processed; the data volume of the data to be processed can be understood as the size of the data to be processed, for example, the data volume of the data to be processed can be 4kb (Kilobyte, kilobyte) in size, etc.

[0067] Based on this, the CPU may determine the data size of the to-be-processed data carried in the request sent by the user to the CPU for processing image data.

[0068] For example, the CPU determines the data size m of the image data AA corresponding to the image A in response to a request sent by the user to perform grayscale processing on the image A.

[0069] Step 804: When it is determined based on the data volume that the data to be processed is to be transmitted through the network card, a corresponding first transmission function is called to send the data to be processed to the GPU through the network card.

[0070] Among them, the network card can be understood as computer hardware used to allow computers to communicate on a computer network. At least one network card can be set on the CPU side and the GPU side respectively, such as a network card with RDMA function; transmission through the network card can be understood as network transmission through the network card; the first transmission function can be understood as the transmission function corresponding to the network card transmission mode, and the second transmission function can be understood as the transmission function corresponding to the bus transmission mode. The first transmission function can be a transmission function obtained after semantic rewriting of the second transmission function, and semantic rewriting can be understood as changing the transmission mode corresponding to the transmission function; calling the first transmission function can enable the data to be processed to be transmitted through the network card transmission mode corresponding to the first transmission function.

[0071] Based on this, when it is determined according to the data volume that the data to be processed is transmitted through the network card, the first transmission function corresponding to the network card transmission mode is called so that the data to be processed can be transmitted to the GPU through the network card.

[0072] Furthermore, when determining to transmit the data to be processed through the network card based on the data volume, in order to further improve transmission efficiency and avoid transmission loss, a transmission threshold of the data to be processed can be determined based on the current network status. The specific implementation method is as follows:

[0073] Detecting the current network status between the CPU and GPU;

[0074] Determining a transmission threshold of data to be processed according to the current network state;

[0075] When the data volume is greater than the transmission threshold, it is determined to transmit the data to be processed through the network card.

[0076] Among them, the current network status can be understood as the network connection status between the CPU and the GPU; the transmission threshold can be understood as the threshold of the size of the pending data that can currently be transmitted from the CPU to the GPU; transmission through the network card can be understood as network transmission through the network card set on the CPU and GPU. For example, the network card can be a network card with a remote direct data access function, then remote direct data access communication can be performed through the network card with a remote direct data access function.

[0077] Based on this, after the CPU receives the data processing request sent by the user, it detects the network connection status between the CPU and the GPU, and determines the threshold of the size of the data to be processed that can be transmitted based on the current network connection status. If the data volume is greater than the transmission threshold, it determines to transmit the data to be processed through the network card.

[0078] Continuing with the above example, after the CPU receives a request from the user to grayscale image A, it checks the network connection status between the CPU and GPU. When the network connection is good, it determines that the threshold for the size of the currently transmittable data to be processed is 4 KB. If the image data AA corresponding to image A is larger than 4 KB, it determines that the image data is transmitted through the network card.

[0079] In summary, by determining the transmission threshold of the data to be processed according to the network status, the optimal data transmission method can be selected in different network environments to ensure the maximum bandwidth, thereby improving data transmission efficiency.

[0080] Specifically, when the data volume is less than the transmission threshold, it can be determined that the data to be processed is transmitted via bus transmission. The specific implementation method is as follows:

[0081] When the data amount is less than the transmission threshold, it is determined to transmit the data to be processed through a bus.

[0082] The bus can be understood as a PCIe bus, which is used to connect the CPU and GPU.

[0083] Based on this, when the size of the data to be processed is smaller than the transmission threshold, the transmission mode is determined to be transmission through the bus.

[0084] Continuing with the above example, when the size of the image data AA is less than 4 kb, it is determined that the image data is transmitted via the PCIe bus.

[0085] In summary, by determining the transmission mode according to whether the size of the data to be processed exceeds the transmission threshold, a suitable transmission mode can be determined, thereby further improving the transmission efficiency.

[0086] In actual applications, when it is determined that the data to be processed is to be transmitted via a bus, the corresponding second transmission function is called to send the data to be processed to the GPU via the bus.

[0087] Among them, the second transfer function can be understood as a transfer function corresponding to the bus transmission mode. The semantics of the second transfer function is to transfer data to the GPU through the bus. When the second transfer function is called, the data to be processed will be transmitted to the GPU according to the bus transmission mode corresponding to the second transfer function.

[0088] Based on this, when it is determined that the transmission mode of the data to be processed is bus transmission, the second transmission function corresponding to the bus transmission mode can be called to enable the data to be processed to be transmitted to the GPU through the bus.

[0089] Specifically, when it is determined based on the data volume that the data to be processed is transmitted through a network card, a corresponding first transmission function may be called in a pre-created function library.

[0090] Here, the function library is understood to be a library that contains all the call functions used to control the GPU through the CPU, such as the user-level driver library, which includes transmission functions corresponding to the transmission mode, such as the first transmission function corresponding to the network card transmission mode, and the second transmission function corresponding to the bus transmission mode.

[0091] In summary, by pre-creating a function library, users are provided with transmission functions with rewritten semantics. After determining the transmission method, they can find and call the transmission function corresponding to the transmission method from the function library, realize network transmission through the network card, and thus improve transmission efficiency.

[0092] In actual application, the specific implementation method of creating this function library is as follows:

[0093] Determine a second transfer function corresponding to the bus transmission mode;

[0094] semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode;

[0095] A function library is created according to the first transfer function and the second transfer function.

[0096] Based on this, the second transmission function corresponding to the bus transmission mode can be determined, the second transmission function can be semantically rewritten, and the corresponding transmission mode can be changed, the transmission mode can be changed from the bus transmission mode to the network card transmission mode, and the first transmission function corresponding to the network card transmission mode can be obtained. According to the first transmission function and the second transmission function, a function library can be constructed.

[0097] In summary, the pre-created function library includes transmission functions corresponding to two transmission modes, one is the second transmission function corresponding to the bus transmission mode, and the other is the first transmission function corresponding to the network card transmission mode. When determining one of the two transmission modes, the corresponding transmission function can be found in the function library to realize the determined transmission mode, thereby improving transmission efficiency.

[0098] In addition, when creating a function library, you can also determine all the call functions that the CPU can use to control the GPU, and semantically rewrite these call functions, such as initialization functions, context functions, etc., to obtain two different sets of call functions for user selection.

[0099] In practical applications, Figure 9 This is a flow chart of a data transmission method provided by an embodiment of this specification, see Figure 9 After physically decoupling the CPU and GPU coupled at a fixed ratio, the GPU is merged into the GPU resource pool. Figure 9 The GPU resource pool in the GPU resource pool includes multiple GPUs, and the CPU can call any one or more GPUs in the GPU resource pool through the connection between the CPU and the GPU resource pool. Specifically, when using the data transmission method provided by the embodiments of this specification, when it is determined that the transmission mode is PCIe bus transmission, the user can call the second transmission function through the application on the CPU side to realize the transmission of data to the GPU in the GPU resource pool in the PCIe bus transmission mode; when it is determined that the transmission mode is GPU remote direct memory access transmission, the user can call the first transmission function through the application on the CPU side to realize the transmission mode of GPU remote direct memory access transmission to the GPU in the GPU resource pool. The transmission mode corresponding to the second transmission function is PCIe bus transmission, for example, it can be the initial transmission function contained in the user-mode driver layer; the first transmission function can be understood as the transmission function after the semantic rewriting of the second transmission function, and its corresponding transmission mode is GPU remote direct memory access transmission.

[0100] Specifically, Figure 10 This is a schematic diagram of a data transmission method according to an embodiment of the present invention. Figure 10In the process of semantically rewriting the second transfer function to obtain the first transfer function, a hijacking library can be inserted between the user-state driver library and the encapsulation library of the user-state driver library. The hijacking library can be used to hijack the user-state driver library. That is to say, the hijacking library can semantically rewrite the functions related to data transmission in the user-state driver library, so that its transmission mode is changed from the original PCIe bus transmission to GPU remote direct memory access network transmission through the network card, providing functions that are exactly the same as those of the user-state driver library. Users can directly call the hijacking library through the upper-level application, or they can make an implicit call through the encapsulation library of the user-state driver library through the upper-level application.

[0101] Since the original user-mode driver library is used through the dynamic library in the parallel computing program, the hijacking library still presents the dynamic library to the outside world. At this time, the user will not perceive the existence of the hijacking library when calling functions through the upper-level application. In other words, it is transparent to the upper-level software application. The user can directly call this library to achieve the improvement of the bandwidth of remote data transmission without perception.

[0102] In summary, by inserting the hijacking library, the original PCIe bus transmission method of transmitting data from the CPU to the GPU is transformed into GPU remote direct memory access network transmission through the network card. The remote direct data access function can reduce CPU participation and significantly increase the HtoD bandwidth without reducing the DtoH bandwidth. At the same time, the hijacking library can ensure that the upper-layer applications are unaware.

[0103] Furthermore, after receiving the data processing request, the determined second transfer function may be hijacked to obtain the first transfer function. The specific implementation is as follows:

[0104] In the case where it is determined according to the data amount that the data to be processed is transmitted through the network card, determining a second transmission function;

[0105] semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode;

[0106] Call the first transmission function corresponding to the network card transmission mode.

[0107] Based on this, when it is determined that the transmission method is through the network card, the second transmission function related to the data transmission can be determined, the second transmission function can be hijacked and semantically rewritten, so that the bus transmission method corresponding to the second transmission function is changed to the network card transmission method, and the first transmission function corresponding to the network card transmission method is obtained. The first transmission function is called to transmit the data to be processed to the GPU through the network card.

[0108] In summary, by dynamically hijacking the transmission function related to data transmission corresponding to the data processing request, the transmission mode is changed, thereby improving the transmission efficiency.

[0109] Furthermore, according to the working principle of remote direct data access, before data transmission, both the CPU-side memory and the GPU-side video memory need to be registered with the network card in advance. The most direct implementation of GPU remote direct memory access transmission is to pin the memory before each transmission and unpin it immediately after the transmission is completed. After the semantic rewriting of the transfer function related to data transmission, each time the CPU transmits data to the GPU, the memory and video memory need to be temporarily registered with the corresponding network card respectively. This operation is very time-consuming. For example, for sending 64MByte data, this time accounts for more than 50% of the entire communication time. Therefore, before receiving the data request, the CPU's memory address and the GPU's video memory address can be registered with the network card. The network card can include a CPU-side network card and a GPU-side network card. The specific implementation method is as follows:

[0110] The CPU is provided with at least one CPU side network card, and the GPU is provided with at least one GPU side network card;

[0111] Correspondingly, the memory address of the CPU is registered to the CPU side network card, and the video memory address of the GPU is registered to the GPU side network card.

[0112] Among them, the CPU-side network card refers to the network card set on the CPU, and the GPU-side network card refers to the network card set on the GPU; registering the CPU's memory address to the CPU-side network card can be understood as providing the CPU's memory address to the network card; registering the GPU's video memory address to the GPU-side network card can be understood as providing the GPU's video memory address to the network card.

[0113] Based on this, the memory address of the CPU and the video memory address of the GPU can be pre-registered with the network card.

[0114] In summary, by pre-registering the CPU memory address and GPU video memory address with the network card, the time for registering the addresses can be saved, thereby improving transmission efficiency.

[0115] In addition, when the CPU transmits data to be processed to the GPU, there are two transmission methods: pageable memory and locked page memory. When the CPU is performing data transmission, when data is transmitted through pageable memory, the GPU cannot directly access this pageable memory, that is, the data in the pageable memory cannot be transmitted to the GPU, resulting in failure of data transmission. Therefore, Figure 11 This is a schematic diagram of a transmission method provided by an embodiment of this specification, such as Figure 11As shown, before receiving a data processing request, a cache area can be registered on the GPU side so that the data transmitted to the GPU is first transmitted to the cache area, and then transmitted from the cache area to the GPU. Specifically, in the process of transmitting the data to be processed, when it is determined that the memory requested by the CPU is pageable memory, a temporary page-locked memory is applied on the CPU side, and the data to be processed is first transferred from the pageable memory to the page-locked memory, and then transferred to the cache area on the GPU side, and then transmitted to the GPU through the cache area. When it is determined that the memory requested by the CPU is page-locked memory, the data to be processed is directly transferred from the page-locked memory to the cache area on the GPU side, and then transmitted to the GPU through the cache area. This realizes data transmission between the CPU and the GPU.

[0116] Furthermore, when GPUs are pooled to obtain a GPU resource pool, the GPU resource pool includes at least two GPUs. When data to be transmitted is sent to the GPU via the network card, at least one target GPU corresponding to the data processing request can be determined. The specific implementation method is as follows:

[0117] Determining at least one target GPU corresponding to the data processing request;

[0118] The data to be processed is sent to the at least one target GPU through the network card.

[0119] The target GPU can be understood as the GPU that needs to be called for the data processing request.

[0120] It is understandable that based on the increase in computing power, one GPU may not be able to meet the computing power demand. At this time, the data processing request can control at least one target GPU in the GPU resource pool to meet the increase in computing power demand.

[0121] In addition, to prevent multiple GPUs from interfering with each other during GPU remote direct memory access, before data transmission, the PCIe topology structure can be perceived and a depth-first search algorithm can be used to provide each GPU in the GPU resource pool with the network card binding with the closest communication distance and remote direct data access function, so as to maximize the efficiency of GPU remote direct memory access communication.

[0122] In summary, an embodiment of the present specification provides a data transmission method, wherein the CPU determines a transmission method for the data to be processed in response to a data processing request, wherein the data processing request carries the data to be processed; when it is determined that the data to be processed is to be transmitted through a network card based on the data volume, the corresponding first transmission function is called to send the data to be processed to the GPU through the network card.

[0123] The above method determines that the transmission method of the data to be processed is through the network card, and calls the corresponding first transmission function, so that the data to be processed can be transmitted from the CPU to the GPU through the network card, avoiding data transmission through the bus, thereby improving data transmission efficiency and reducing the impact of the increased PCIe bus latency on the bandwidth part.

[0124] Furthermore, subsequent tests may be performed on the parallel computing program applied to the above method, and bandwidth performance tests and AI performance tests may be performed on the parallel computing program.

[0125] Specifically, during the bandwidth performance test, according to the test results, the bandwidth performance of data transmission from the CPU to the GPU has been improved, while the bandwidth of data transmission from the GPU to the CPU has not been affected; during the AI ​​performance test, it can be seen that the gain of increasing bandwidth on AI performance is affected by the following factors: the bandwidth performance of data transmission from the CPU to the GPU of the parallel computing program applied to the above method is improved, the performance loss is reduced, and the bandwidth of data transmission from the GPU to the CPU remains basically unchanged.

[0126] The benefit of increasing bandwidth on AI training performance is primarily affected by two factors: the ratio of memory copy operations to the runtime of code running on the GPU. Specifically, the larger the amount of data transferred in a single task, the higher the proportion of memory copy time, and the more significant the performance improvement from increasing bandwidth. Furthermore, the number of concurrent tasks. Specifically, because multiple data transfer tasks on a single GPU share a single bandwidth for transmitting data from the CPU to the GPU, the higher the number of concurrent tasks, the greater the data throughput and the greater the benefit from increasing bandwidth.

[0127] In practical applications, a neural network model or a deep learning recommendation model can be used as a test model to test the parallel computing program applied to the above-mentioned data transmission method.

[0128] The test results show that the number of parameters passed to the program for training in a single task increased from 16 to 256, with significant performance improvements. In the process of multi-tasking concurrency, the performance increased by 4%-25.4%, with significant performance improvements in all aspects.

[0129] In summary, the results of subsequent tests show that the data transmission method provided in this specification can significantly improve the bandwidth performance of data transmission from the CPU to the GPU, reduce the impact of increased latency on bandwidth caused by the distance between the PCIe bus when the CPU and GPU are decoupled, and change the semantics of some functions through the hijacking layer to replace the transmission mode without the user's perception. Through topology awareness and dynamic configuration of transmission thresholds, it is possible to select the network card closest to the target GPU for communication, and adopt different transmission methods for different sizes of data to be processed to ensure the optimality of transmission, thereby improving data transmission efficiency.

[0130] The following combined Figure 12 , taking the application of the data transmission method provided in this specification in image sharpening as an example, the data transmission method is further explained. Figure 12 A flowchart of a data transmission method according to an embodiment of the present disclosure is shown, which specifically includes the following steps.

[0131] Step 1202: Register the memory address of the CPU to the CPU-side network card, and register the video memory address of the GPU to the GPU-side network card.

[0132] Step 1204: Receive an image sharpening request from the user, where the image sharpening request carries image data.

[0133] Step 1206: In response to the image sharpening request, detect the network status between the CPU and the GPU.

[0134] Step 1208: Determine the transmission threshold according to the network status.

[0135] Step 1210: When the size of the image data is greater than the transmission threshold, determine the GPU remote direct memory access mode.

[0136] Step 1212: Determine a transfer function related to data transmission, and rewrite its semantics to change its transmission mode to GPU remote direct memory access mode.

[0137] Step 1214: calling the semantically rewritten transfer function to transfer the image data from the CPU memory to the GPU video memory via the GPU remote direct memory access method.

[0138] The above method determines that the transmission mode of the data to be processed is through the network card, and calls the corresponding first transmission function, so that the data to be processed can be transmitted from the CPU to the GPU through the network card, avoiding data transmission through the bus, thereby improving data transmission efficiency.

[0139] Corresponding to the above method embodiment, this specification also provides a data transmission device embodiment, Figure 13FIG1 shows a schematic diagram of the structure of a data transmission device provided by an embodiment of this specification. Figure 13 As shown, the device includes:

[0140] The determining module 1302 is configured as a CPU to determine a transmission mode of the data to be processed in response to a data processing request, wherein the data processing request carries the data to be processed;

[0141] The sending module 1304 is configured to, when it is determined according to the data volume that the data to be processed is to be transmitted through the network card, call a corresponding first transmission function and send the data to be processed to the GPU through the network card.

[0142] In an optional embodiment, the sending module 1304 is further configured to:

[0143] When it is determined according to the data volume that the data to be processed is transmitted through a network card, a corresponding first transmission function is called in a pre-created function library.

[0144] In an optional embodiment, the apparatus further includes a creation module, and the creation module is further configured to:

[0145] Determine a second transfer function corresponding to the bus transmission mode;

[0146] semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode;

[0147] A function library is created according to the first transfer function and the second transfer function.

[0148] In an optional embodiment, the sending module 1304 is further configured to:

[0149] In the case where it is determined according to the data amount that the data to be processed is transmitted through the network card, determining a second transmission function;

[0150] semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode;

[0151] Call the first transmission function corresponding to the network card transmission mode.

[0152] In an optional embodiment, the determining module 1302 is further configured to:

[0153] Detecting the current network status between the CPU and GPU;

[0154] Determining a transmission threshold of data to be processed according to the current network state;

[0155] When the data volume is greater than the transmission threshold, it is determined to transmit the data to be processed through the network card.

[0156] In an optional embodiment, the determining module 1302 is further configured to:

[0157] When the data amount is less than the transmission threshold, it is determined to transmit the data to be processed through a bus.

[0158] In an optional embodiment, the sending module 1304 is further configured to:

[0159] In the case where it is determined that the data to be processed is to be transmitted via a bus, a corresponding second transmission function is called to send the data to be processed to the GPU via the bus.

[0160] In an optional embodiment, the apparatus further includes a registration module, and the registration module is further configured to:

[0161] The CPU is provided with at least one CPU side network card, and the GPU is provided with at least one GPU side network card;

[0162] Correspondingly, the memory address of the CPU is registered to the CPU side network card, and the video memory address of the GPU is registered to the GPU side network card.

[0163] In an optional embodiment, the sending module 1304 is further configured to:

[0164] Determining at least one target GPU corresponding to the data processing request;

[0165] The data to be processed is sent to the at least one target GPU through the network card.

[0166] One embodiment of the present specification provides a data transmission device, wherein a CPU determines a transmission method for data to be processed in response to a data processing request, wherein the data processing request carries the data to be processed; when it is determined that the data to be processed is to be transmitted through a network card based on the data volume, a corresponding first transmission function is called to send the data to be processed to a GPU through the network card.

[0167] The above-mentioned device calls the corresponding first transmission function when determining that the transmission method of the data to be processed is through the network card, so that the data to be processed can be transmitted from the CPU to the GPU through the network card, avoiding data transmission through the bus, thereby improving data transmission efficiency.

[0168] The above is a schematic scheme of a data transmission device of this embodiment. It should be noted that the technical scheme of the data transmission device and the technical scheme of the above-mentioned data transmission method are of the same concept. For details not described in detail in the technical scheme of the data transmission device, please refer to the description of the technical scheme of the above-mentioned data transmission method.

[0169] Figure 14 14 shows a block diagram of a computing device 1400 according to one embodiment of the present disclosure. Components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.

[0170] Computing device 1400 also includes an access device 1440 that enables computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1440 may include one or more of any type of network interface (e.g., a network interface card (NIC)), whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0171] In one embodiment of the present specification, the above components of the computing device 1400 and Figure 14 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 14 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0172] Computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. Computing device 1400 can also be a mobile or stationary server.

[0173] The processor 1420 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data transmission method when executed by the processor.

[0174] The above is a schematic solution of a computing device of this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above-mentioned data transmission method are of the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned data transmission method.

[0175] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data transmission method when executed by a processor.

[0176] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data transmission method are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned data transmission method.

[0177] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data transmission method.

[0178] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above-mentioned data transmission method are of the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the above-mentioned data transmission method.

[0179] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0180] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0181] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0182] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0183] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data transmission method for a CPU, the method comprising: In response to a data processing request, determining the amount of data to be processed carried in the data processing request; When it is determined according to the data volume that the data to be processed is transmitted through a network card, a first transmission function corresponding to a transmission mode of the network card is called to send the data to be processed to the GPU through the network card.

2. The method according to claim 1, wherein, when it is determined based on the data volume that the data to be processed is to be transmitted through a network card, calling a first transmission function corresponding to a network card transmission mode comprises: When it is determined according to the data volume that the data to be processed is transmitted through a network card, a corresponding first transmission function is called in a pre-created function library.

3. The method according to claim 2, wherein the function library is created by the following steps: Determine a second transfer function corresponding to the bus transmission mode; semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode; A function library is created according to the first transfer function and the second transfer function.

4. The method according to claim 1, wherein, when it is determined based on the data volume that the data to be processed is to be transmitted via a network card, calling a first transmission function corresponding to a network card transmission mode comprises: In the case where it is determined according to the data amount that the data to be processed is transmitted through the network card, determining a second transmission function; semantically rewriting the second transmission function to obtain a first transmission function corresponding to the network card transmission mode; Call the first transmission function corresponding to the network card transmission mode.

5. The method according to claim 1, wherein determining, based on the data volume, whether to transmit the data to be processed through a network card comprises: Detecting the current network status between the CPU and GPU; Determining a transmission threshold of data to be processed according to the current network state; When the data volume is greater than the transmission threshold, it is determined to transmit the data to be processed through the network card.

6. The method according to claim 5, further comprising: When the data amount is less than the transmission threshold, it is determined to transmit the data to be processed through a bus.

7. The method according to claim 6, further comprising: In the case where it is determined that the data to be processed is to be transmitted via a bus, a corresponding second transmission function is called to send the data to be processed to the GPU via the bus.

8. The method according to claim 1, further comprising: The CPU is provided with at least one CPU side network card, and the GPU is provided with at least one GPU side network card; Correspondingly, the memory address of the CPU is registered to the CPU side network card, and the video memory address of the GPU is registered to the GPU side network card.

9. The method according to claim 1, wherein the number of the GPUs is at least two; Accordingly, sending the data to be processed to the GPU through the network card includes: Determining at least one target GPU corresponding to the data processing request; The data to be processed is sent to the at least one target GPU through the network card.

10. A data transmission device for a CPU, comprising: a determination module configured to determine, in response to a data processing request, the amount of to-be-processed data carried in the data processing request; The sending module is configured to call a first transmission function corresponding to the network card transmission mode when it is determined according to the data volume that the data to be processed is transmitted through the network card, and send the data to be processed to the GPU through the network card.

11. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • GPU server and data transmission method

    CN111782565A

  • Heterogeneous server cluster and data forwarding method, device and equipment

    CN114827151A