GPU box, and data processing method and system

Through GPU BOX, multiple GPUs are decoupled from computing devices to form an independent GPU area network, solving the problem of underutilization of GPU resources, and achieving efficient GPU resource management and long-distance access across devices, improving resource utilization and reducing costs.

WO2025179936A1PCT designated stage Publication Date: 2025-09-04XFUSION DIGITAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129187
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2024-10-31
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

In the prior art, the GPU resources of the GPU server are not fully utilized by the computing node, resulting in waste of resources, and the GPU resources can only be used in the local server and cannot be accessed across chassis, computer rooms or data centers.

Method used

Through GPU BOX, multiple GPUs are decoupled from computing devices to form an independent GPU area network, and network devices communicate with computing devices to realize GPU resource allocation and management across devices, including processors, expansion chips and GPU bridge chips, supporting multiple computing devices to share GPU resources.

Benefits of technology

Improves the utilization rate of GPU resources, avoids resource waste, reduces usage costs, and allows long-distance access to GPU resources across chassis, computer rooms and even data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129187_04092025_PF_FP_ABST
    Figure CN2024129187_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computing. Disclosed are a GPU BOX, and a data processing method and system, which can provide GPU resources to a plurality of computing devices, thereby increasing the utilization rate of the GPU resources. The GPU BOX provided in the present application comprises a processor, a first expansion chip, a plurality of GPUs and a GPU bridge chip, wherein the processor is connected to the plurality of GPUs by means of the first expansion chip, and the plurality of GPUs are connected by means of the GPU bridge chip. The GPU BOX is used for receiving data processing tasks sent by one or more computing devices, processing, on the basis of the data processing tasks, data to be processed, and sending data processing results to the one or more computing devices, wherein the data processing tasks comprise the data to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

A GPU BOX, data processing method and system

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 28, 2024, with application number 202410224289.9 and application name “A GPU BOX, Data Processing Method and System”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of computing technology, and in particular to a GPU BOX, a data processing method, and a system. Background Art

[0003] A graphics processing unit (GPU) is a microprocessor specialized for image and graphics-related computing. With the advancement of modern technology, applications like machine vision and artificial intelligence (AI) are placing increasing demands on GPUs. In related technologies, a compute node containing a central processing unit (CPU) is typically integrated with a GPU server within a single chassis. The GPU resources of the GPU server can only be used by that compute node. If the GPU server's GPU resources are not fully utilized by the compute node, this results in wasted GPU resources.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a GPU BOX, a data processing method, and a system, which can provide GPU resources for multiple computing devices and improve GPU resource utilization.

[0006] To achieve the above technical objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a GPU BOX, including a processor, a first extension chip, multiple image processors GPUs, and a GPU bridge chip; the processor is connected to the multiple GPUs through the first extension chip; the multiple GPUs are connected through the GPU bridge chip; the GPU BOX is used to receive data processing tasks sent by one or more computing devices, process the data to be processed based on the data processing tasks, and send the data processing results to one or more computing devices; the data processing tasks include the data to be processed.

[0008] In one possible implementation, a GPU BOX is communicatively connected to one or more computing devices via a network device; a processor is used for data processing tasks of a target computing device, which is any computing device among multiple computing devices; the data processing tasks are sent to one or more target GPUs, and the data processing tasks instruct the one or more target GPUs to process the data to be processed; data processing results returned by the one or more target GPUs are received, and the data processing results are sent to the target computing device via the network device.

[0009] In another possible embodiment, the processor includes one or more data processors DPUs, and the one or more DPUs are connected to multiple GPUs through a high-speed serial computer expansion bus switch PCIE Switch. The one or more DPUs send data processing tasks to one or more target GPUs through the PCIE Switch.

[0010] In another possible implementation, before sending the data processing task to one or more target GPUs, the processor is further used to obtain the GPU resource requirements of the data processing task; based on the GPU resource requirements of the data processing task, determine one or more target GPUs, and the available resources of the one or more target GPUs meet the GPU resource requirements of the data processing task.

[0011] In another possible implementation, the GPU BOX further includes a memory storing a GPU resource record table, wherein the GPU resource record table includes available resource information of multiple GPUs. The processor is specifically configured to determine one or more target GPUs based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirements of the data processing task; wherein the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirements of the data processing task.

[0012] In another possible implementation, the processor is further configured to update available resource information of one or more target GPUs in the GPU resource record table based on GPU resources occupied by the one or more target GPUs for executing the data processing task.

[0013] In another possible implementation, the processor is specifically configured to split the data processing task into multiple subtasks to be executed; and send task information to one or more target GPUs, where the task information includes the subtasks to be executed and GPU resources required to execute the subtasks to be executed.

[0014] In another possible embodiment, the GPU BOX also includes a memory, which stores a GPU resource record table. The GPU resource record table includes available resource information of multiple GPUs. After receiving the data processing results returned by one or more target GPUs, the processor is also used to reclaim the GPU resources allocated to the data processing tasks on the one or more target GPUs, and update the available resource information of the one or more target GPUs in the GPU resource record table.

[0015] In another possible implementation, one or more GPUs are configured to process the data to be processed under the instruction of the processor.

[0016] In a second aspect, an embodiment of the present application provides a data processing method, which is applied to a processor of a GPU BOX, where the GPU BOX includes multiple GPUs, and the GPU BOX is communicatively connected to multiple computing devices via a network device. The method includes: receiving a data processing task from a target computing device forwarded by the network device, where the target computing device is any computing device among multiple computing devices, and the data processing task includes data to be processed; sending the data processing task to one or more target GPUs, where the data processing task instructs processing of the data to be processed; receiving data processing results returned by one or more target GPUs, and sending the data processing results to the target computing device via the network device.

[0017] In this method, multiple computing devices can use the GPU resources in an external GPU BOX. Compared to related art GPU servers, which only provide GPU resources to the server's internal computing nodes, this method improves GPU resource utilization, avoids waste, and reduces GPU resource costs. Furthermore, because the GPU BOX connects to multiple computing devices via a network device, this method is not restricted by location, allowing computing devices to access GPU resources remotely via the network.

[0018] In one possible implementation, the processor includes one or more DPUs, and the one or more DPUs are connected to multiple GPUs through a PCIE switch, and send data processing tasks to one or more target GPUs, including: one or more DPUs send data processing tasks to one or more target GPUs through a PCIE switch.

[0019] The DPU is a chip specifically designed for data processing, with a high degree of data parallelism and optimization capabilities, enabling more efficient processing and acceleration of various data-intensive tasks. A processor with multiple DPUs can increase processor bandwidth and improve data processing efficiency. Furthermore, a PCIE switch contains multiple upstream and downstream ports, allowing data to be transferred between one or more DPUs and one or more GPUs via the PCIE switch, enabling multiple DPUs to transfer data to multiple GPUs.

[0020] In another possible implementation, before sending the data processing task to one or more target GPUs, the method further includes: obtaining the GPU resource requirements of the data processing task; and determining one or more target GPUs based on the GPU resource requirements of the data processing task, wherein the available resources of the one or more target GPUs meet the GPU resource requirements of the data processing task.

[0021] Among them, allocating corresponding GPU resources to the data processing task according to the GPU resource requirements of the data processing task can achieve on-demand allocation, avoid GPU resource waste, and improve GPU resource utilization.

[0022] In another possible implementation, the GPU BOX includes a GPU resource record table, which includes available resource information of multiple GPUs. Determining one or more target GPUs based on the GPU resource requirements of the data processing task includes: determining one or more target GPUs based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirements of the data processing task; wherein the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirements of the data processing task.

[0023] Among them, the above-mentioned processor manages the resources of multiple GPUs through the GPU resource record table, can clearly grasp the resource allocation status of multiple GPUs and the amount of remaining available resources, and can allocate tasks to GPUs with sufficient available resources based on the GPU resource record table, which is conducive to improving the processor's management efficiency of GPU resources on the GPU.

[0024] In another possible implementation, the method further includes: updating available resource information of the one or more target GPUs in the GPU resource record table based on GPU resources occupied by the one or more target GPUs for executing the data processing task.

[0025] Among them, if the processor determines the target GPU, it indicates that the processor instructs to allocate the resources of the target GPU to the data processing task, and updates the available resource information of the target GPU in the GPU resource record table. This is conducive to the timely allocation of tasks based on the latest GPU available resource information when there are other tasks in the future, avoiding task execution failure due to insufficient GPU available resources.

[0026] In another possible implementation, sending a data processing task to one or more target GPUs includes: splitting the data processing task into multiple subtasks to be executed; sending task information to the one or more target GPUs, where the task information includes the subtasks to be executed and the GPU resources required to execute the subtasks to be executed.

[0027] Among them, if multiple target GPUs are determined, the data processing task is divided into multiple subtasks based on the available resources of each target GPU to avoid subtask execution failure due to insufficient resources of the target GPU. In addition, splitting into multiple subtasks can also realize parallel execution of multiple subtasks, thereby improving the processing efficiency of the data processing task.

[0028] In another possible implementation, the GPU BOX includes a GPU resource record table, which includes available resource information of multiple GPUs. After receiving data processing results returned by one or more target GPUs, the method further includes: recovering GPU resources allocated to data processing tasks on one or more target GPUs, and updating the available resource information of one or more target GPUs in the GPU resource record table.

[0029] The processor promptly updates the resource information of the target GPU recorded in the GPU resource record table, which is beneficial for obtaining the latest resource information of multiple GPUs when there are new tasks in the future, thereby improving the resource utilization of the GPU.

[0030] In a third aspect, an embodiment of the present application provides a data processing method, which is applied to a computing device, wherein the computing device is communicatively connected to a GPU BOX via a network device. The method includes: sending a data processing task to a processor of the GPU BOX via the network device, so that one or more GPUs of the GPU BOX processes the data processing task; and receiving a data processing result sent by the processor of the GPU BOX via the network device.

[0031] In the above method, the computing device sends data processing tasks to the external GPU BOX through the network device, and the GPU BOX performs the data processing tasks. The computing device can use remote GPU resources across chassis, computer rooms, and even data centers, thereby improving the utilization rate of GPU resources. In addition, the computing device does not need to deploy GPU resources inside, reducing hardware costs.

[0032] In one possible implementation, the processor of the GPU BOX includes one or more data processors DPUs, and sending a data processing task to the processor of the GPU BOX through a network device includes: determining one or more target DPUs based on DPU resource requirements of the data processing task; and sending the data processing task to the one or more target DPUs of the GPU BOX through the network device.

[0033] Among them, the computing device allocates data processing tasks to one or more target DPUs based on the data processing tasks' demand for DPU resources, realizing on-demand allocation. While ensuring that the data processing tasks have sufficient DPU resources, it can also avoid wasting DPU resources.

[0034] In a fourth aspect, an embodiment of the present application provides a data processing system, comprising multiple computing devices, a network device, and a GPU BOX; the multiple computing devices are communicatively connected to the GPU BOX through the network device; and the GPU BOX is used to process data processing tasks sent by the multiple computing devices through the network device.

[0035] In a fifth aspect, an embodiment of the present application provides a GPU BOX, comprising a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the GPU BOX executes the data processing method as described in the second aspect and any possible implementation thereof.

[0036] In a sixth aspect, an embodiment of the present application provides a processor, wherein the processor is applied to each module of the data processing method of the second aspect or any possible implementation method of the second aspect.

[0037] In a seventh aspect, an embodiment of the present application provides a computing device, wherein the computing device is applied to various modules of the data processing method of the third aspect or any possible implementation method of the third aspect.

[0038] In an eighth aspect, embodiments of the present application provide a computer-readable storage medium comprising computer instructions. When the computer instructions are executed on a GPU BOX, the GPU BOX executes the data processing method according to the second aspect and any possible implementation thereof.

[0039] In a ninth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions. When the computer instructions are executed on a GPU BOX, the GPU BOX executes the data processing method according to the second aspect and any possible implementation thereof.

[0040] For the specific description of the fourth to ninth aspects and their various implementations in the embodiments of the present application, reference can be made to the detailed description in the second or third aspect and their various implementations; and for the beneficial effects of the second to third aspects and their various implementations, reference can be made to the beneficial effect analysis in the second or third aspect and their various implementations, which will not be repeated here.

[0041] These and other aspects of the embodiments of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is an architecture diagram of a GPU server provided in an embodiment of the present application;

[0043] FIG2 is a schematic diagram of the internal logical structure of a GPU server provided in an embodiment of the present application;

[0044] FIG3 is a diagram illustrating an architecture of a data processing system provided in an embodiment of the present application;

[0045] FIG4 is an architecture diagram of another data processing system provided in an embodiment of the present application;

[0046] FIG5 is a schematic diagram of the internal logic structure of a GPU BOX provided in an embodiment of the present application;

[0047] FIG6 is a schematic diagram of the internal structure of a DPU provided in an embodiment of the present application;

[0048] FIG7 is a flow chart of a data processing method provided in an embodiment of the present application;

[0049] FIG8 is a schematic diagram of the structure of a processor provided in an embodiment of the present application;

[0050] FIG9 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] For ease of understanding, the following briefly introduces the relevant terms involved in the embodiments of this application:

[0052] (1) NVLink: A bus and its communication protocol. NVLink uses a point-to-point structure and serial transmission for connecting CPUs and GPUs, and can also be used to connect multiple graphics processors.

[0053] (2) Graphics processing unit (GPU): also known as display core, visual processor or display chip, is a microprocessor specifically used for image and graphics-related computing tasks.

[0054] (3) NV Switch: A GPU bridging device (chip) that provides the required NVLink cross-network. Multiple GPUs can be interconnected at high speed through the NVLink interface, thereby improving the communication efficiency and bandwidth between multiple GPUs.

[0055] (4) PCIE Switch: includes multiple upstream ports and multiple downstream ports, providing PCIE data transmission links for devices or chips connected to the upstream ports and downstream ports. For example, it is used to provide many-to-many data transmission links between multiple data processing units (DPUs) and multiple GPUs, hereinafter referred to as processors.

[0056] In the following, the terms "first," "second," and "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, a feature designated as "first," "second," or "third," etc., may explicitly or implicitly include one or more of the features.

[0057] In the description of the embodiments of this application, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" in this document is simply a description of an association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more.

[0058] In the related art, as shown in Figure 1, taking the computing device as a server as an example, a computing device may include: a chassis, a GPU frame, a computing node frame and a power module; the GPU frame, the computing node frame and the power module are connected to the chassis of the GPU server through the internal hardware connection backplane of the chassis, the GPU frame contains the GPU chip and its peripheral circuits, and the computing node frame contains the CPU chip, memory stick, baseboard management controller (baseboard management controller, BMC) and other components. The GPU resources of the GPU server can only be used by the computing nodes in the chassis. If the GPU resources of the GPU server are not fully utilized by the computing nodes, it will cause a waste of GPU resources.

[0059] In one embodiment, for the GPU server shown in Figure 1, the computing nodes include CPU0 and CPU1, and the GPU chassis contains 8 GPU chips, namely GPU0-GPU7, as shown in Figure 2, which shows a schematic diagram of the logical structure between the two CPU chips and the 8 GPU chips. As can be seen from Figure 2, the two CPU chips can be connected via a bus, and the 8 GPU chips are connected to the NV Switch to form a GPU module. The two CPU chips can share access to the GPU resources of the 8 GPU chips through a high-speed serial computer expansion bus switch (peripheral component interconnect express switch, PCIE Switch).

[0060] Based on this, an embodiment of the present application proposes a data processing method, which provides image data processing services for multiple computing devices through an external GPU BOX, wherein one computing device may include all components except the GPU frame (including the GPU frame and internal GPU chips and other components) as shown in Figure 1. Specifically, the GPU BOX proposed in this method is set outside the computing device and communicates with the computing device through a network device. Each computing device (target computing device) can send data processing tasks to the GPU BOX through the network device. The processor of the GPU BOX then sends the data processing tasks to one or more target GPUs. After the one or more target GPUs have processed the data processing tasks, the processor of the GPU BOX returns the data processing results to the computing device (target computing device).

[0061] In the above method, multiple computing devices can use the GPU resources in the external GPU BOX. Compared with the related art GPU server in which the GPU resources can only be provided to the computing nodes inside its own server chassis, this method can improve the utilization rate of GPU resources, avoid waste, and reduce the cost of using GPU resources.

[0062] In addition, since the GPU BOX is connected to multiple computing devices through a network device, this method is not restricted by the site, and the computing devices can access remote GPU resources through the network.

[0063] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0064] Please refer to Figure 3, which shows a data processing system architecture diagram involved in the data processing method provided in an embodiment of the present application. As shown in Figure 3, the data processing system may include: a computing device 110, a GPU BOX 120 and a network device 130.

[0065] In one scenario, the data processing system may include one or more computing devices 110, where the one or more computing devices 110, GPU BOX 120, and network device 130 are all in the same data center, and the one or more computing devices 110 are connected to the network device 130, and a network connection is established with the GPU BOX 120 through the network device 130.

[0066] Of course, in some other scenarios, the data processing system may include one or more computing devices 110, one or more computing devices 110 are set in a data center, the GPU BOX 120 may be set outside the data center, and the one or more computing devices 110 are connected to the network device 130, and a network connection is established with the GPU BOX 120 through the network device 130.

[0067] As shown in FIG4 , FIG4 shows another data processing system architecture diagram.

[0068] Computing device 110 is a computing device with data processing, logical operation, and storage capabilities. For example, computing device 110 may include a server, tablet computer, desktop computer, laptop computer, notebook computer, computing node, or netbook computer. A server may be a rack server, blade server, tower server, cabinet server, or other different types of server. A server may include one or more computing nodes, each of which includes at least one CPU. When a server includes multiple computing nodes, the CPUs in the multiple computing nodes share a common operating system.

[0069] GPU BOX120 can be understood as decoupling one or more GPUs from the computing device to form an independent GPU area network (GPU Area Network), which has image data processing capabilities, used to process image data sent by one or more computing devices 110, and feed back the processed image data to one or more computing devices 110.

[0070] As shown in FIG5 , the internal structure of the GPU BOX 120 includes a DPU 121 and a GPU 122 , and optionally, a PCIE Switch 123 and an NV Switch 124 .

[0071] The number of DPUs 121 and GPUs 122 in GPU BOX 120 can be one or more respectively. When there are multiple DPUs 121 and multiple GPUs 122, multiple DPUs 121 can be connected to multiple GPUs 122 through PCIE Switch 123, and multiple GPUs 122 are also connected to NV Switch 124 respectively to form a GPU module.

[0072] DPU121, a data processor with a high degree of data parallelism and optimization capabilities, can more efficiently process and accelerate various data-intensive tasks.

[0073] In an embodiment of the present application, DPU121 is used to receive an image data processing task sent by the computing device 110 through the network device 130. The image data processing task is carried by a network signal, and after converting the network signal representing the image data processing task into a PCIE signal, it is sent to one or more GPUs122 through the PCIE Switch123; when the GPU122 completes processing the image data, it converts the image data processing result into a PCIE signal representation, and sends it to the DPU121 through the PCIE Switch123. The DPU121 converts the PCIE signal representing the image data processing result into a network signal, and sends it to the computing device 110 through the network device 130.

[0074] The above-mentioned network signal refers to data information transmitted through the network, and the PCIE signal refers to data information transmitted through the PCIE bus and complies with the PCIE transmission protocol, for example, complies with the PCIE4.0 or PCIE5.0 protocol.

[0075] In one embodiment, as shown in FIG6 , the internal structure of the DPU 121 includes a field programmable gate array (FPGA) 1211 , a CPU 1212 , a first type of memory 1213 , a second type of memory 1214 , and a network interface 1215 .

[0076] FPGA1211 is a semi-custom circuit in a dedicated integrated circuit and is a programmable logic array. In the embodiment of the present application, FPGA1211 is used to convert network signals into PCIE signals.

[0077] CPU 1212, the computing and control core, is the final execution unit for information processing and program execution. CPU 1212 can be a single-core CPU (single-CPU) or a multi-core CPU (multi-CPU). In the embodiment of the present application, CPU 1212 is used to receive PCIE signals converted by FPGA 1211, determine GPU 122 based on the PCIE signals, and send data processing tasks to GPU 122.

[0078] The first type of memory 1213 can provide fast read and write storage space for temporarily caching data processed by the FPGA 1211 or the CPU 1212. Due to its fast read and write capabilities, it can also accelerate the processing speed of the FPGA 1211 or the CPU 1212. The first type of memory 1213 includes, but is not limited to, a dual inline memory module (DIMM). A DIMM can include multiple memory chips, wherein the memory chips on the DIMM can be dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc. For example, it can be DDR4 (fourth generation double data rate synchronous dynamic random access memory) or DDR5 (fifth generation double data rate synchronous dynamic random access memory).

[0079] The second type of memory 1214 is a persistent storage space used to store execution instructions of the FPGA 1211 or the CPU 1212. In the embodiment of the present application, it can also be used to store GPU resource records. The second type of memory 1214 includes, but is not limited to, a solid state disk (SSD), a hard disk drive (HDD), or persistent memory (PMEM).

[0080] In some implementations, FPGA 1211 or CPU 1212 reads instructions stored in the second type of memory 1214 into the first type of memory 1213 to implement the data processing method provided in the embodiments of the present application for DPU 121 .

[0081] Network interface 1215, a device including a transmitter and a receiver, is used to communicate with other devices or communication networks. It can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface. Alternatively, network interface 1215 can be a wireless interface. It should be understood that network interface 1215 includes multiple physical ports and is used for communication, etc.

[0082] In some embodiments, the DPU 121 further includes a bus 1216. For example, the FPGA 1211 and the CPU 1212 are connected via a PCIE bus, the CPU 1212 and the second type of memory 1214 are connected via a serial advanced technology attachment (SATA), and the first type of memory 1213 and the network interface 1215 are typically connected to the FPGA 1211 via a system bus.

[0083] Based on the DPU121 structure shown in Figure 6, the network data sent by the computing device 110 through the network device 130 is transmitted to the FPGA1211 chip inside the DPU121 through the network interface 1215. The FPGA1211 chip converts the network data into PCIE signals and hands them over to the CPU1212 for processing. After the CPU1212 determines the GPU122 for the data processing task to be processed, it distributes it to each GPU122 for processing through the PCIE Switch123. After the GPU122 completes the processing, it sends the processing result to the DPU121 through the PCIE Switch123. The DPU121 then returns the data processing result to the computing device 110 through the original path.

[0084] The network device 130 is used to provide a network signal transmission link between the computing device 110 and the GPU BOX 120. For example, the network device 130 can be a network switch, such as a fabric chanel (FC) switch, a PCIE switch, or an Ethernet switch.

[0085] It should be noted that if the transmission distance between the computing device 110 and the GPU BOX 120 is too long, a PCIE retimer repeater can be added between the PCIE Switch 123 and the GPU 122 inside the GPU BOX 120 according to the signal link loss requirements to increase the data transmission signal and avoid data signal attenuation due to the long transmission distance, which may cause data damage.

[0086] In an application scenario, the computing device 110, GPU BPX120 and network device 130 in Figure 3 or Figure 4 are set in a data center. A computing device 110 in the data center sends the image data to be processed to GPU BOX120 through the network device 130. After receiving the image data, one or more DPUs121 in GPU BOX120 convert the image data into PCIE signals and send them to one or more GPUs122 for processing through PCIE Switch123; after one or more GPUs122 process the image data, they convert the image data processing results into PCIE signals and send them to one or more DPUs121 through PCIE Switch123. The one or more DPUs121 convert the image data processing results into network signals and send them to the computing device 110 through the network device 130.

[0087] The following describes the data processing method provided in the embodiment of the present application:

[0088] Please refer to Figure 7, which is a flowchart of a data processing method provided in an embodiment of the present application. This method is applied to a processor of a GPU Box, such as DPU 121 in Figure 5. The GPU Box includes multiple GPUs, which are communicatively connected to multiple computing devices via a network device. As shown in Figure 7, the method may include S101-S103.

[0089] S101: GPU BOX receives a data processing task from a target computing device.

[0090] Specifically, a processor of the GPU BOX, such as DPU 121 , may receive a data processing task from a target computing device that is forwarded by a network device.

[0091] The target computing device is any computing device among the plurality of computing devices; the data processing task instructs processing of data to be processed.

[0092] The data processing task request includes the data to be processed, which is the image or graphic data that the target computing device needs to process when processing AI or machine vision services.

[0093] In one embodiment, the target computing device sends a data processing task to a processor of the GPU BOX through a network device, so that one or more GPUs of the GPU BOX process the data processing task.

[0094] Specifically, the processor of the GPU BOX includes one or more DPUs. The target computing device first determines one or more target DPUs based on the DPU resource requirements of the data processing task; then sends the data processing task to the one or more target DPUs of the GPU BOX through the network device.

[0095] The DPU resource requirement of a data processing task includes the amount of data processing resources required for data conversion of the data to be processed (for example, converting network signals into PCIE signals). For example, if the amount of data to be processed is 1G and the required DPU resources are 2 DPUs, then the DPU resource requirement of the data processing task is 2 DPUs.

[0096] For example, if the GPU BOX contains 8 DPUs, identified as DPU0-DPU7 respectively, the target computing device determines DPU0-DPU1 in the GPU BOX as the target DPU based on the DPU resource requirements of the data processing task. When the target computing device receives the data processing task through the network device, it indicates that the DPU identified as receiving the data processing task is DPU0-DPU1, so as to send the data processing task to DPU0-DPU1 of the GPU BOX. Subsequently, the two target DPUs jointly perform data conversion and sending and other processing on the data processing task.

[0097] S102: The GPU BOX processes the data processing task through one or more target GPUs.

[0098] Specifically, the processor of the GPU BOX sends the data processing task to one or more target GPUs.

[0099] In one embodiment, the processor includes one or more DPUs, which are connected to multiple GPUs via a PCIE switch. One or more target DPUs in the processor that receive data processing tasks send the data processing tasks to one or more target GPUs via the PCIE switch.

[0100] In this implementation process, if there are multiple target DPUs, each target DPU is responsible for issuing part of the data processing tasks, and each GPU in the multiple target GPUs is also responsible for processing part of the received data processing tasks.

[0101] A DPU is a chip specifically designed for data processing. It features a high degree of data parallelism and optimization capabilities, enabling it to more efficiently process and accelerate various data-intensive tasks. If a processor includes multiple DPUs, they can operate in parallel, increasing data processing bandwidth and improving data conversion efficiency.

[0102] The PCIE Switch contains multiple upstream ports and multiple downstream ports, so one or more DPUs can be connected to one or more GPUs through the PCIE Switch, and any DPU can transmit data to any GPU.

[0103] Optionally, before sending the data processing task to one or more target GPUs, the processor first determines one or more target GPUs. The specific determination process includes S102a-S102b:

[0104] S102a: The GPU BOX processor obtains the GPU resource requirements of the data processing task.

[0105] GPU resource requirements include the number of GPU cores required.

[0106] In one embodiment, the data processing task carries the number of GPU cores required for the data to be processed, for example, 200 GPU cores, and the processor obtains the GPU resources required for the data processing task from the data processing task.

[0107] S102b: The GPU BOX processor determines one or more target GPUs based on the GPU resource requirements of the data processing task.

[0108] The available resources of the one or more target GPUs meet the GPU resource requirement of the data processing task, that is, the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirement of the data processing task.

[0109] If the processor determines one target GPU, the available resources of the target GPU meet the GPU resource requirements of the data processing task. If the processor determines multiple target GPUs, the sum of the available resources of the multiple target GPUs meets the GPU resource requirements of the data processing task.

[0110] In one embodiment, the GPU BOX includes a memory, such as the second type of memory 1214 in Figure 6, which stores a GPU resource record table. The GPU resource record table includes available resource information of multiple GPUs. The specific method for the processor to determine one or more target GPUs includes: the processor determines one or more target GPUs based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirements of the data processing task.

[0111] The available resource information of the GPU includes the number of available GPU cores. Optionally, the GPU resource record table also includes the number of allocated GPU cores.

[0112] If there is a GPU whose available resources are greater than or equal to the GPU resource requirements of the data processing task, then the GPU is the target GPU.

[0113] If the available resources of each GPU in multiple GPUs are less than the GPU resource requirement of the data processing task, but the sum of the available resources of the multiple GPUs is greater than or equal to the GPU resource requirement of the data processing task, and the difference between the sum of the available resources of the multiple GPUs and the GPU resource requirement of the data processing task is less than a first threshold, then the multiple GPUs are all target GPUs.

[0114] For example, as shown in Table 1, Table 1 shows a GPU resource record table, which records the available resource information of multiple GPUs. Table 1 includes GPU identification, the number of available GPU cores, and the number of allocated GPU cores.

[0115] Table 1

[0116] The GPU identifier in Table 1 is used to indicate the GPU identity. The identifier includes but is not limited to a number or a symbol. Table 1 uses the number as an example for explanation.

[0117] If the GPU resource requirement of the data processing task is 200 GPU cores, the processor determines from the GPU resource record table that the number of available GPU cores of GPU4 meets the GPU resource requirement of the data processing task, and GPU4 can be used as the target GPU; if the GPU resource requirement of the data processing task is 300 GPU cores, the number of available cores of each GPU in the GPU resource record table is less than the GPU resource requirement of the data processing task, and the sum of the number of available GPU cores of the processor GPU1 and GPU4 is 300, or the sum of the number of available GPU cores of GPU2 and GPU4 is 350, both of which can meet the GPU resource requirement of the data processing task, so the processor can determine GPU1 and GPU4 as the target GPU, or determine GPU2 and GPU4 as the target GPU.

[0118] The above example is only one possible implementation method. In other implementation methods, if the number of available cores of multiple GPUs in the GPU resource record table can meet the GPU resource requirements of the data processing task, the processor can select any one as the target GPU, or the processor can determine multiple target GPUs to jointly process the data processing task based on the load balancing principle to improve the processing efficiency of the data processing task. The embodiment of the present application does not limit how to select the target GPU.

[0119] If the processor determines a target GPU, the processor sends the data processing task to the target GPU.

[0120] Correspondingly, the target GPU receives the data processing task sent by the processor and executes the data processing task.

[0121] If the processor determines multiple target GPUs, the processor executes S102A-S102B to send the data processing task to the multiple target GPUs.

[0122] It should be noted that before the target computing device sends the data processing task to the GPU BOX through the network device, it first converts the data processing task into a network signal. Therefore, the data processing task received by the GPU BOX processor is a network signal. Before the processor sends the data processing task to one or more target GPUs, it needs to convert the network signal into a PCIE signal.

[0123] S102A: The GPU BOX processor splits the data processing task into multiple subtasks to be executed.

[0124] Specifically, when splitting a data processing task, the GPU BOX processor splits the data processing task into multiple subtasks to be executed based on the available resources of each target GPU, wherein the available resources of the target GPU executing each subtask are greater than or equal to the GPU resources required by the subtask to be executed.

[0125] For example, if the processor determines that two GPUs jointly process the data processing task, the processor can split the data processing task into two subtasks, wherein the processor determines that subtask 1 is performed by target GPU1 and subtask 2 is performed by target GPU2, the available resources of target GPU1 are greater than or equal to the GPU resources required for subtask 1, and the available resources of target GPU2 are greater than or equal to the GPU resources required for subtask 2.

[0126] S102B: The processor of the GPU BOX sends task information to one or more target GPUs. The task information includes subtasks to be executed and the amount of GPU resources required to execute the subtasks to be executed.

[0127] Correspondingly, the target GPU receives the task information sent by the processor and allocates required GPU resources to the subtask indicated by the task information to execute the subtask.

[0128] For example, if the GPU resources required for a data processing task include 300 GPU cores, the processor splits the data processing task into two subtasks, Subtask 1 and Subtask 2. Subtask 1 requires 100 GPU cores, and Subtask 2 requires 200 GPU cores. The processor sends task information to target GPU 1, including "Subtask 1" and the GPU resources required to execute "Subtask 1": 100 GPU cores. The processor also sends task information to target GPU 2, including "Subtask 2" and the GPU resources required to execute "Subtask 2": 200 GPU cores.

[0129] In S102A-S102B, if the processor determines multiple target GPUs, the data processing task is divided into multiple subtasks based on the available resources of each target GPU to avoid subtask execution failure due to insufficient resources of some target GPUs. In addition, splitting into multiple subtasks can also achieve parallel execution of multiple subtasks, thereby improving the processing efficiency of the data processing task.

[0130] Optionally, after determining the one or more target GPUs, the processor updates the available resource information of the one or more target GPUs in the GPU resource record table based on the GPU resources occupied by executing the data processing task in the one or more target GPUs.

[0131] In one embodiment, the processor updates the GPU resources required by the data processing task from the number of available GPU cores of the target GPU to the number of allocated GPU cores in the GPU resource record table.

[0132] For example, if the data processing task requires 200 GPU cores, as shown in the GPU resource record table in Table 1, where the number of available GPU cores of GPU4 is 200, the processor determines that GPU4 is the target GPU and updates the number of available GPU cores of GPU4 from 200 to 0 and the number of allocated GPU cores from 100 to 300. Table 2 shows the updated GPU resource record table.

[0133] Table 2

[0134] It can be understood that if the processor determines the target GPU, it indicates that the processor instructs to allocate the resources of the target GPU to the data processing task, and updates the available resource information of the target GPU in the GPU resource record table. This is beneficial for subsequent other tasks to be allocated based on the latest GPU available resource information in a timely manner, avoiding task execution failure due to insufficient GPU available resources.

[0135] In the above optional implementation method, the processor can update the GPU resource record table after determining the target GPU, or can update the GPU resource record table after issuing the data processing task. The embodiment of the present application does not limit the execution order.

[0136] The above-mentioned processor manages the resources of multiple GPUs through the GPU resource record table, can clearly grasp the resource allocation status of multiple GPUs and the amount of remaining available resources, and can allocate tasks to GPUs with sufficient available resources based on the GPU resource record table, which is conducive to improving the processor's management efficiency of GPU resources on the GPU.

[0137] S103: The GPU BOX obtains the data processing result and sends the data processing result to the target computing device through the network device.

[0138] Specifically, the processor of the GPU BOX receives data processing results returned by one or more target GPUs, and sends the data processing results to the target computing device through the network device.

[0139] The data processing result is the processing result of the data to be processed in the data processing task.

[0140] In one embodiment, one or more target GPUs convert the data processing results into PCIE signals before returning the data processing results. After the processor DPU receives the PCIE signals, it converts the PCIE signals into network signals. The data processing results are sent to the target computing device through the network device in the format of network signals.

[0141] After receiving the data processing results from one or more target GPUs, the processor reclaims the GPU resources allocated for the data processing task on the one or more target GPUs and updates the available resource information for the one or more target GPUs in the GPU resource record table. Correspondingly, after the one or more target GPUs complete the data processing task, they release the GPU resources allocated for the data processing task.

[0142] In one embodiment, the processor updates the GPU resources allocated to the data processing task from the number of GPU cores allocated to the target GPU to the number of GPU cores available in the GPU resource record table.

[0143] For example, if the target GPU allocates 200 GPU cores for the data processing task, as shown in the GPU resource record table in Table 2, where GPU4 is the target GPU, after receiving the data processing results returned by one or more target GPUs, the processor updates the number of allocated GPU cores from 300 to 100 and the number of available GPU cores of GPU4 from 0 to 200. Table 3 shows the updated GPU resource record table.

[0144] Table 3

[0145] In this method, the target GPU releases GPU resources promptly after completing the data processing task, preventing GPU resource occupancy and reducing resource waste. The processor promptly updates the target GPU's resource information recorded in the GPU resource record table, enabling subsequent new tasks to obtain the latest resource information from multiple GPUs, thereby improving GPU resource utilization.

[0146] In a data processing method proposed in an embodiment of the present application, a GPU Box is located outside a computing device and communicates with the computing device via a network device. Each computing device (target computing device) can flexibly use the GPU resources within the GPU Box through the network device. After use, the GPU Box promptly releases the GPU resources used by the computing device. Compared with the related art GPU server, which only provides GPU resources to internal computing nodes, this method can improve GPU resource utilization, avoid waste, and reduce the cost of GPU resource use.

[0147] In addition, since the GPU BOX is connected to multiple computing devices through network devices, this method is not restricted by the site. Computing devices can access remote GPU resources through the network across chassis, computer rooms, and even data centers.

[0148] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed herein, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0149] The embodiment of the present application further provides a processor 200, such as the DPU 121 in Figure 5. As shown in Figure 8, it is a schematic diagram of the structure of a processor 200 provided in the embodiment of the present application.

[0150] Among them, the processor 200 includes: a receiving unit 201, which is used to receive a data processing task request from a target computing device forwarded by a network device, and the target computing device is any computing device among multiple computing devices; the data processing task indicates processing of the data to be processed; a sending unit 202, which is used to send the data processing task to one or more target GPUs so that the data processing task is executed by one or more target GPUs; the receiving unit 201 is also used to receive data processing results returned by one or more target GPUs, and send the data processing results to the target computing device through the network device.

[0151] In some embodiments, the processor includes one or more DPUs, and the one or more DPUs are connected to multiple GPUs through a PCIE Switch. The sending unit 202 is specifically used to send the PCIE signal between the DPU and the PCIE Switch to one or more target GPUs, wherein the PCIE signal includes a data processing task received from the target computing device.

[0152] In some embodiments, the processor 200 also includes an acquisition unit 203, which is used to obtain the GPU resource requirements of the data processing task before sending the data processing task to one or more target GPUs; the processor 200 also includes a determination unit 204, which is used to determine one or more target GPUs based on the GPU resource requirements of the data processing task, and the available resources of the one or more target GPUs meet the GPU resource requirements of the data processing task.

[0153] In some embodiments, the GPU BOX includes a GPU resource record table, which includes available resource information of multiple GPUs. The determination unit 204 is specifically used to determine one or more target GPUs based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirements of the data processing task; wherein the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirements of the data processing task.

[0154] In some implementations, the processor 200 further includes an updating unit 205 configured to update available resource information of one or more target GPUs in the GPU resource record table based on GPU resources occupied by the one or more target GPUs for executing data processing tasks.

[0155] In some implementations, the sending unit 202 is specifically configured to split the data processing task into multiple subtasks to be executed; and send task information to one or more target GPUs, where the task information includes the subtasks to be executed and GPU resources required to execute the subtasks to be executed.

[0156] In some embodiments, the GPU BOX includes a GPU resource record table, which includes available resource information of multiple GPUs. After receiving the data processing results returned by one or more target GPUs, the update unit 205 is also used to recycle the GPU resources allocated to the data processing tasks on the one or more target GPUs, and update the available resource information of the one or more target GPUs in the GPU resource record table.

[0157] Of course, the processor 200 provided in the embodiment of the present application includes but is not limited to the above modules.

[0158] FIG9 is a schematic diagram of the structure of a computing device 300 provided in an embodiment of the present application, and the computing device 300 may be the computing device 110 in FIG3 or FIG4 . As shown in FIG9 , the computing device 300 includes a processor 301 , a memory 302 , and a network interface 303 .

[0159] The processor 301 includes one or more CPUs, which may be single-core CPUs or multi-core CPUs. An operating system may be run in the processor.

[0160] The memory 302 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical storage.

[0161] In some implementations, the processor 301 implements the data processing method provided in the embodiments of the present application by reading instructions stored in the memory 302, or the processor 301 implements the data processing method provided in the embodiments of the present application by internally stored instructions. In the case where the processor 301 implements the method in the above embodiment by reading instructions stored in the memory 302, the memory 302 stores instructions for implementing the data processing method provided in the embodiments of the present application.

[0162] Network interface 303, a device comprising a transmitter and a receiver, is used to communicate with other devices or a communication network. It can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface. Alternatively, network interface 303 can be a wireless interface. It should be understood that network interface 303 includes multiple physical ports and is used for communication, etc.

[0163] In some implementations, the computing device 300 further includes a bus 304 , and the processor 301 , memory 302 , and network interface 303 are typically interconnected via the bus 304 , or are interconnected in other ways.

[0164] Another embodiment of the present application provides a GPU box comprising a memory and a processor. The memory and processor are coupled; the memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the GPU box performs the steps of the data processing method described in the above method embodiment.

[0165] Another embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on the GPU BOX, the GPU BOX executes each step executed by the GPU BOX in the data processing method flow shown in the above method embodiment.

[0166] Another embodiment of the present application provides a chip system, which is applied to a GPU Box. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuits. The interface circuits are used to receive signals from the GPU Box's memory and send signals to the processors. The signals include computer instructions stored in the memory. When the GPU Box processor executes the computer instructions, the GPU Box performs each step performed by the GPU Box in the data processing method flow shown in the above method embodiment.

[0167] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions are executed on the GPU BOX, the GPU BOX executes each step executed by the GPU BOX in the data processing method flow shown in the above method embodiment.

[0168] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer-executable instructions are loaded and executed on a computer, the process or function according to the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a server, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD).

[0169] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, and all such changes or substitutions shall fall within the scope of protection of this application.

Claims

1. An image processing resource pool device GPU BOX, characterized in that: The system comprises a processor, a first expansion chip, a plurality of image processors (GPUs), and a GPU bridge chip; the processor is connected to the plurality of GPUs via the first expansion chip; the plurality of GPUs are connected via the GPU bridge chip; The GPU BOX is used to receive data processing tasks sent by one or more computing devices, process the data to be processed based on the data processing tasks, and send the data processing results to the one or more computing devices; the data processing tasks include the data to be processed.

2. The GPU BOX according to claim 1, wherein: The GPU BOX is communicatively connected to one or more computing devices via a network device; the processor is used to receive a data processing task from a target computing device, where the target computing device is any computing device among the multiple computing devices; the data processing task is sent to one or more target GPUs, where the data processing task instructs the one or more target GPUs to process the data to be processed; the processor is also used to receive data processing results returned by the one or more target GPUs, and send the data processing results to the target computing device via the network device.

3. The GPU BOX according to claim 2, wherein: The processor includes one or more data processors DPUs, which are connected to the multiple GPUs via a high-speed serial computer expansion bus switch PCIE Switch. The one or more DPUs send the data processing tasks to one or more target GPUs via the PCIE Switch.

4. The GPU BOX according to claim 2 or 3, characterized in that: Before sending the data processing task to one or more target GPUs, the processor is further used to obtain the GPU resource requirements of the data processing task; based on the GPU resource requirements of the data processing task, determine one or more target GPUs, and the available resources of the one or more target GPUs meet the GPU resource requirements of the data processing task.

5. The GPU BOX according to claim 4, wherein: The GPU BOX also includes a memory, in which a GPU resource record table is stored. The GPU resource record table includes available resource information of the multiple GPUs. The processor is specifically configured to determine one or more target GPUs based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirement of the data processing task; wherein the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirement of the data processing task.

6. The GPU BOX according to claim 5, characterized in that: The processor is further configured to update available resource information of the one or more target GPUs in the GPU resource record table based on GPU resources occupied by the one or more target GPUs for executing the data processing task.

7. The GPU BOX according to any one of claims 4 to 6, wherein: The processor is specifically configured to split the data processing task into a plurality of subtasks to be executed; and send task information to the one or more target GPUs, where the task information includes the subtasks to be executed and GPU resources required to execute the subtasks to be executed.

8. The GPU BOX according to any one of claims 2 to 7, wherein: The GPU BOX also includes a memory, in which a GPU resource record table is stored. The GPU resource record table includes available resource information of the multiple GPUs. After receiving the data processing results returned by the one or more target GPUs, the processor is further used to reclaim the GPU resources allocated to the data processing task on the one or more target GPUs, and update the available resource information of the one or more target GPUs in the GPU resource record table.

9. The GPU BOX according to any one of claims 1 to 8, wherein: The one or more GPUs are used to process the data to be processed under the instruction of the processor.

10. A data processing method, characterized in that: A processor applied to a GPU BOX device in an image processing resource pool, wherein the GPU BOX includes multiple GPUs and is communicatively connected to multiple computing devices via a network device, wherein the method includes: receiving a data processing task from the target computing device, where the target computing device is any computing device among the plurality of computing devices, and the data processing task includes data to be processed; Sending the data processing task to one or more target GPUs, wherein the data processing task instructs the one or more target GPUs to process the data to be processed; Receive data processing results returned by the one or more target GPUs, and send the data processing results to the target computing device through the network device.

11. The method according to claim 10, characterized in that The processor includes one or more data processors (DPUs), and the one or more DPUs are connected to the multiple GPUs via a high-speed serial computer expansion bus (PCIE) switch. The sending of the data processing task to the one or more target GPUs includes: The one or more DPUs send the data processing task to one or more target GPUs through the PCIE Switch.

12. The method according to claim 10 or 11, characterized in that Before sending the data processing task to one or more target GPUs, the method further includes: Obtaining GPU resource requirements for the data processing task; Based on the GPU resource requirement of the data processing task, one or more target GPUs are determined, where available resources of the one or more target GPUs meet the GPU resource requirement of the data processing task.

13. The method according to claim 12, characterized in that The GPU BOX further includes a memory, wherein a GPU resource record table is stored in the memory, wherein the GPU resource record table includes available resource information of the multiple GPUs. Determining one or more target GPUs based on the GPU resource requirements of the data processing task includes: Based on the available resource information of each GPU in the GPU resource record table and the GPU resource requirement of the data processing task, one or more target GPUs are determined; wherein the sum of the available resources of the one or more target GPUs is greater than or equal to the GPU resource requirement of the data processing task.

14. The method according to claim 13, characterized in that The method further comprises: Based on the GPU resources occupied by the one or more target GPUs for executing the data processing task, available resource information of the one or more target GPUs in the GPU resource record table is updated.

15. The method according to any one of claims 12 to 14, characterized in that The sending of the data processing task to one or more target GPUs includes: Splitting the data processing task into multiple subtasks to be executed; Sending task information to the one or more target GPUs, where the task information includes the subtasks to be executed and GPU resources required to execute the subtasks to be executed.

16. The method according to any one of claims 10 to 15, characterized in that The GPU BOX further includes a memory, wherein a GPU resource record table is stored in the memory, wherein the GPU resource record table includes available resource information of the multiple GPUs. After receiving the data processing results returned by the one or more target GPUs, the method further includes: Reclaim the GPU resources allocated to the data processing task on the one or more target GPUs, and update the available resource information of the one or more target GPUs in the GPU resource record table.

17. A data processing method, characterized in that: Applied to a computing device, the computing device is communicatively connected to an image processing resource pool device GPU BOX via a network device, and the method includes: Sending a data processing task to a processor of the GPU BOX through the network device, so that one or more graphics processors GPU of the GPU BOX processes the data processing task; Receive a data processing result sent by the processor of the GPU BOX through the network device.

18. The method according to claim 17, characterized in that The processor of the GPU BOX includes one or more data processors DPU, and sending the data processing task to the processor of the GPU BOX through the network device includes: Determining one or more target DPUs based on the DPU resource requirements of the data processing task; The data processing task is sent to the one or more target DPUs of the GPU BOX through the network device.

19. A data processing system, characterized in that: It includes multiple computing devices, network devices and image processing resource pool device GPU BOX; the multiple computing devices are connected to the GPU BOX through the network device; the GPU BOX is used to process data processing tasks sent by the multiple computing devices through the network device.

20. An image processing resource pool device GPU BOX, characterized in that: The GPU BOX comprises a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the GPU BOX executes the method according to any one of claims 10 to 16.

Citation Information

Patent Citations

  • Using data processing unit (DPU) as preprocessor for graphics processing unit (GPU)-based machine learning

    CN115150278A

  • Data resource scheduling method and system, electronic equipment and storage medium

    CN115543617A

  • Processor allocation method, system and device, storage medium and electronic equipment

    CN116166434A

  • GPU BOX and data processing method and system

    CN118295955A

  • Video Processing Across Multiple Graphics Processing Units

    US20100053176A1