Virtual GPU system and application method therefor, and device and storage medium
Through the virtual GPU system, multiple CPU service nodes and schedulers are used to simulate high-performance GPUs, solving the problems of insufficient GPU service nodes and the CPU service nodes being unable to meet the large-scale data computing performance, achieving the improvement of the number of GPU cores and the improvement of computing performance, while maintaining a low cost.
Patent Information
- Application Number
- PCT/IB2024/062951
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, there are few GPU service nodes, resulting in insufficient processing performance. At the same time, the CPU service node cannot meet the performance requirements of large-scale data computing.
A virtual GPU system is proposed, through a server cluster composed of multiple CPU service nodes, load balancing is used by a scheduler, vector driver components and algorithm plug-ins are deployed on each CPU service node to realize remote direct data access network communication between CPU service nodes, and simulate a single high-performance GPU.
A significant improvement in the number of GPU cores has been achieved. Each simulated CPU core performance is equivalent to 10% of GPU performance. When the number of CPU service nodes exceeds 10, the system performance exceeds that of ordinary GPUs, and the computing performance can be further improved by adding CPU service nodes while maintaining low costs.
Smart Images

Figure IB2024062951_26062025_PF_FP_ABST
Abstract
Description
[0001] Virtual GPU System, Application Method, Device, and Storage Medium Thereof This disclosure claims priority from Chinese patent application No. 202311777298.2, filed with the China Patent Office on December 2, 2023, entitled "Virtual GPU System, Application Method, Device, and Storage Medium Thereof," the entire contents of which are incorporated herein by reference. TECHNICAL FIELD This disclosure belongs to the field of computer technology and, more specifically, relates to a virtual GPU system, application method, device, and storage medium thereof. BACKGROUND In the era of artificial intelligence, the use of graphics processing units (GPUs) is becoming increasingly widespread. A GPU is a microprocessor used in personal computers, workstations, game consoles, and mobile devices such as tablets and smartphones to perform image processing. Although GPU service nodes deploying GPUs offer relatively good processing performance in the image processing field, due to their high cost, they are typically not deployed in large numbers. Furthermore, using CPU (Central Processing Unit) service nodes for computing cannot meet the performance requirements of large-scale data computations. SUMMARY OF THE INVENTION This disclosure proposes a virtual GPU system and its application method, device, and storage medium, alleviating the technical issues in related technologies related to the limited number of GPU service nodes and the inability of CPU service nodes to meet large-scale data computing performance requirements. A first embodiment of this disclosure proposes a virtual GPU system comprising: multiple CPU service nodes and a scheduler communicating with each of the CPU service nodes; each CPU service node is deployed with a vector driver component and an algorithm plug-in for implementing a vector algorithm; the CPU service nodes communicate with each other via a remote direct data access network; the scheduler is configured to schedule each CPU service node to execute pending tasks in parallel; each CPU service node is configured to execute tasks scheduled by the scheduler using the locally deployed vector driver component to drive the locally deployed algorithm plug-in, and during task execution, share data with CPU service nodes requiring collaborative computing via the remote direct data access network.In some embodiments, the scheduler is configured to split a pending task into multiple subtasks; split the data set to be processed by the pending task into multiple subdata sets based on the task type of each subtask; determine a target CPU service node from the multiple CPU service nodes for executing each subtask; distribute the subtask and the corresponding subdata set to the determined target CPU service node; receive processing results returned by each target CPU service node, and aggregate the processing results to obtain a task processing result for the pending task; the target CPU service node is configured to drive the locally deployed algorithm plug-in via the locally deployed vector driver component to execute the received subtask on the received subdata set to obtain a processing result; and transmit the processing result to the scheduler. In some embodiments, the memory on each CPU service node includes a cache and residual storage space; the cache is configured to store the subdata sets corresponding to the subtasks to be executed by the CPU service node and intermediate results during task execution; and the residual storage space is configured to store the processing results of the target CPU service node. A second aspect of the present disclosure provides an application method for a virtual GPU system, which is applied to the virtual GPU system described in the first aspect. The method includes: splitting a task to be processed into multiple subtasks; splitting a data set to be processed by the task to be processed into multiple subdata sets based on the task type of each subtask; determining a target CPU service node for executing each subtask from the multiple CPU service nodes; distributing the subtask and the corresponding subdata set to the determined target CPU service node; receiving processing results returned by each target CPU service node; and aggregating the processing results to obtain a task processing result for the task to be processed. In some embodiments, splitting a task to be processed into multiple subtasks includes: obtaining multiple computing operations included in the task to be processed; obtaining an estimated execution time corresponding to each computing operation, where the estimated execution time is used to represent the time required to execute each computing operation; selecting a minimum estimated execution time from the estimated execution times; splitting a first computing operation among the multiple computing operations into multiple sub-computing operations, where the estimated execution time of each sub-computing operation is less than or equal to the minimum estimated execution time, and the estimated execution time of the first computing operation is greater than the minimum estimated execution time; generating the multiple subtasks based on each sub-computing operation and the second computing operation, where the execution time of the second computing operation is less than or equal to the minimum estimated execution time, and each sub-computing operation or the second computing operation corresponds to a subtask.In some embodiments, based on the task type of each subtask, splitting the data set to be processed by the pending task into multiple subdata sets includes: determining a task type corresponding to the task type of each subtask; splitting the data set into multiple intermediate data sets according to the task type; each task type uniquely corresponds to one intermediate data set; and evenly dividing the data in the intermediate data set corresponding to each task type according to the number of task types included in each task type to obtain multiple subdata sets. In some embodiments, obtaining the estimated execution time corresponding to each computing operation includes: obtaining the estimated execution time of the vector operation corresponding to each computing operation; and using the estimated execution time of the vector operation as the estimated execution time corresponding to each computing operation. In some embodiments, the method further includes: when executing each subtask on the target CPU service node, obtaining a cache usage rate and a remaining storage space usage rate of the target CPU service node; and adjusting the cache size based on the cache usage rate and the remaining storage space usage rate. A third aspect of the present disclosure provides an application device for a virtual GPU system, applicable to the virtual GPU system described in the first aspect. The device comprises: a first splitting module for splitting a task to be processed into multiple subtasks; a second splitting module for splitting a data set to be processed by the task to be processed into multiple subdata sets based on the task type of each subtask; a determination module for determining, from among the multiple CPU service nodes, a target CPU service node for executing each subtask; a distribution module for distributing the subtask and the corresponding subdata set to the determined target CPU service node; a receiving module for receiving processing results returned by each target CPU service node; and a aggregation module for aggregating the processing results to obtain a task processing result for the task to be processed. A fourth aspect of the present disclosure provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the second aspect. A fifth aspect of the present disclosure provides a computer-readable storage medium, wherein the computer program is stored, and the program is executed by the processor to implement the method described in the second aspect. An embodiment of the sixth aspect of the present disclosure provides a computer program product, including a computer program, which implements the method of the second aspect when executed by a processor.The technical solutions provided in the embodiments of the present disclosure have at least the following technical effects or advantages: In the embodiments of the present disclosure, a server cluster virtualizes a GPU, consisting of multiple CPU service nodes. A scheduler performs load balancing among these CPU service nodes. Each CPU service node is deployed with a vector driver component and an algorithm plug-in. The vector driver component can call the algorithm plug-in to execute tasks assigned by the scheduler, thereby enabling the server cluster to simulate a GPU. Simulating multiple CPU service nodes as a single GPU significantly increases the number of GPU cores. According to test data, the performance of each simulated CPU core is equivalent to 10% of the GPU performance. When the number of CPU service nodes in a cluster exceeds 10, the performance of the entire system begins to surpass that of a standard GPU. Furthermore, by adding more CPU service nodes to the cluster, the system's computing performance can be further improved while maintaining a relatively low cost. The high-speed interconnection of the RDMA network effectively resolves data transmission bottlenecks and further improves overall system performance. Finally, the deployment of a vector driver component on each CPU service node helps fully utilize the CPU's vector instructions to call the algorithm plug-in, efficiently completing scheduler tasks through vector computing. Additional aspects and advantages of the present disclosure will be described in part in the following description, and in part will become apparent from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The drawings are provided for illustration purposes only and are not intended to limit the present disclosure. Like reference symbols denote like components throughout the drawings. In the drawings: FIG1 illustrates a schematic structural diagram of a virtual GPU system provided in one embodiment of the present disclosure; FIG2 illustrates a flow diagram of a method for applying a virtual GPU system provided in one embodiment of the present disclosure; FIG3 illustrates a schematic structural diagram of an application device for a virtual GPU system provided in one embodiment of the present disclosure; FIG4 illustrates a schematic structural diagram of an electronic device provided in one embodiment of the present disclosure; and FIG5 illustrates a schematic diagram of a storage medium provided in one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE DRAWINGS Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.It should be noted that, unless otherwise specified, technical or scientific terms used in this disclosure should have the common meanings understood by persons skilled in the art to which this disclosure relates. In the era of artificial intelligence, the application of graphics processing units (GPUs) is becoming increasingly widespread. A GPU is a microprocessor used in personal computers, workstations, game consoles, and some mobile devices such as tablets and smartphones to perform image processing. Although GPU service nodes deployed with GPUs offer relatively good processing performance in the image processing field, due to the high cost of GPU service nodes, they are typically not deployed in large numbers. However, computing using CPU (Central Processing Unit) service nodes cannot meet the performance requirements of large-scale data computations. To alleviate the problems existing in the related art, embodiments of the present disclosure provide a virtual GPU system. This system comprises a server cluster virtual GPU consisting of multiple CPU service nodes. A scheduler performs load balancing among these CPU service nodes. Each CPU service node is deployed with a vector driver component and an algorithm plug-in. The vector driver component can call the algorithm plug-in to execute tasks called by the scheduler, thereby enabling the server cluster to simulate a GPU. Multiple CPU service nodes simulate a single GPU, significantly increasing the number of GPU cores. According to test data, the performance of each simulated CPU core is equivalent to 10% of the GPU performance. When the number of CPU service nodes in a cluster exceeds 10, the overall system performance begins to surpass that of a standard GPU. Furthermore, by adding more CPU service nodes to the cluster, the system's computing performance can be further improved while maintaining a relatively low cost. High-speed interconnection via the RDMA (Remote Direct Memory Access) network effectively resolves data transmission bottlenecks and further improves overall system performance. Finally, the deployment of a vector driver component on each CPU service node helps fully utilize the CPU's vector instruction call algorithm plug-in, enabling efficient completion of scheduler tasks through vector computing. The CPU service nodes in the disclosed embodiments include, but are not limited to, CPU servers deployed in the cloud. A CPU server is a server whose internal processor includes a CPU.As shown in Figure 1, the virtual GPU system includes: multiple CPU service nodes 11 and a scheduler 12 that communicates with each CPU service node 11; each CPU service node 11 is deployed with a vector driver component 13 and an algorithm plug-in 14 for implementing a vector algorithm; the CPU service nodes 11 communicate with each other via a remote direct data access network; the scheduler 12 is configured to schedule each CPU service node 11 to execute pending tasks in parallel; each CPU service node 11 is configured to execute tasks scheduled by the scheduler 12 via a locally deployed vector driver component 13 and a locally deployed algorithm plug-in 14, and to share data with CPU service nodes requiring collaborative computing via a remote direct memory access (RDMA) network during task execution. In this embodiment, the scheduler 12 can be deployed on any one of the multiple CPU service nodes 11, or on a service node other than the multiple CPU service nodes 11, and this embodiment does not specifically limit this. In this embodiment, the vector driver component (Pytorch cvGPU Driver) 13 is configured with a driver algorithm, which enables communication between the scheduler 12 and the CPU service node 11. In this embodiment, the algorithm plugin (Driver Plugin) 14 provides a Pytorch algorithm framework for implementing vector algorithms. In applications, the scheduler 12 sends vectorization instructions to the algorithm plugin 14 via the vector driver component 13. In response to these instructions, the algorithm plugin 14 executes the corresponding vector algorithm. Vectorization instructions include, but are not limited to, instructions such as SSE (Streaming SIMD Extensions) and AVX (Advanced Vector Extensions). The algorithm plugin 14 encapsulates the vector algorithm, facilitating efficient use of the CPU's vector computing capabilities and significantly accelerating its use. In this embodiment, a remote direct data access network enables high-speed interconnection between the CPU service nodes 11. Through the RDMA network, the GPU service nodes can communicate with each other with low latency and high throughput, enabling collaborative computing and data sharing. Remote Direct Data Access Network (RDDN) is a data transmission technology that allows direct memory access between CPU service nodes 11 without CPU intervention. By bypassing the operating system kernel and transmitting data directly between network adapters, RDDN improves data transmission efficiency and bandwidth utilization.Using remote direct data access network technology can optimize memory latency between CPU service nodes 11 to hundreds of times that of non-RDMA, significantly improving inter-cluster performance. In an optional embodiment, the scheduler 12 is configured to split a pending task into multiple subtasks; split the data set to be processed by the pending task into multiple subdata sets based on the task type of each subtask; determine a CPU service node 11 from multiple CPU service nodes 11 to execute each subtask; distribute the subtask and corresponding subdata set to the determined CPU service node 11; receive processing results returned by each CPU service node 11, and aggregate the processing results to obtain a task processing result for the pending task.
[0002] The CPU service node 11 is configured to drive the locally deployed algorithm plug-in 14 via the locally deployed vector driver component 13 to execute the received subtasks on the received subdata sets, obtain processing results, and transmit the processing results to the scheduler 12. In this embodiment, pending tasks include, but are not limited to, computing tasks corresponding to each network layer included in the neural network model. For example, the pending task may be a convolution algorithm task corresponding to a convolution layer, or a pooling algorithm task corresponding to a pooling layer. In this embodiment, if the pending task is a computing task related to the neural network model, the data set to be processed by the pending task may be training data, test data, etc. for the neural network model. In this embodiment, the pending task can be split into multiple subtasks based on user instructions. For example, for a network layer with 784 computing units, the pending task corresponding to the network layer can be split into 784 subtasks based on user instructions, with each computing unit corresponding to one subtask. Splitting rules can also be pre-configured in the application, so that the scheduler 12 can split the pending task into subtasks based on the pre-configured splitting rules. In an optional embodiment, when a pending task is split into multiple subtasks, the scheduler 12 is specifically configured to: obtain multiple computing operations included in the pending task; obtain an estimated execution time corresponding to each computing operation, where the estimated execution time represents the time required to execute each computing operation; select a minimum estimated execution time from the estimated execution times; split a first computing operation among the multiple computing operations into multiple sub-computing operations, where the estimated execution time of each sub-computing operation is less than or equal to the minimum estimated execution time, and the estimated execution time of the first computing operation is greater than the minimum estimated execution time; and generate multiple subtasks based on each sub-computing operation and a second computing operation, where the execution time of the second computing operation is less than or equal to the minimum estimated execution time, and each sub-computing operation or the second computing operation corresponds to a subtask. The computing operations included in the pending task can be basic operations such as addition, subtraction, multiplication, and division, or complex operations such as integration and differentiation, which are not specifically limited in this embodiment. In this embodiment, the estimated execution times of different computing operations can be pre-configured based on empirical data. The estimated execution time here can be the estimated execution time of a common operation corresponding to the operation, or the estimated execution time of a corresponding vector algorithm. In this embodiment, the multiple sub-operations obtained by splitting the first operation can correspond to the same algorithm. For example, the multiple sub-operations obtained by splitting an addition operation can still correspond to addition. The difference lies in that the sub-operations process different data.As an example, a pending task includes two operations, operation A and operation B. The estimated execution time of operation A is 4 seconds, and the estimated execution time of operation B is 2 seconds. Therefore, the minimum estimated execution time is 2 seconds. When splitting the pending task, operation A is the first operation, and operation B is the second operation. Therefore, operation A can be split into two sub-operations, and the estimated execution time of each sub-operation is controlled to be 2 seconds. These two sub-operations and operation B are then executed by different CPU service nodes 11. Since the estimated execution time of the two sub-operations and operation B is both 2 seconds, after this sub-task splitting in this embodiment, the total execution time of operation A and operation B is only 2 seconds, that is, both operations A and B can be completed within 2 seconds. If the pending task is executed according to the relevant technology, the minimum execution time for operation A and operation B is (4s + 2s) / 2 = 3. In some embodiments, the data set to be processed by the pending task can be evenly divided according to the number of subtasks, thereby obtaining multiple subdata sets. Specifically, in some embodiments, when performing the step of splitting the data set to be processed by the pending task into multiple subdata sets based on the task type of each subtask, the scheduler 12 can be configured to: determine a task type corresponding to the task type of each subtask; split the data set into multiple intermediate data sets according to the task type; each task type uniquely corresponds to an intermediate data set; and evenly divide the data in the intermediate data set corresponding to each task type according to the number of task types included in each task type, thereby obtaining multiple subdata sets. In this embodiment, the task type is used to indicate the algorithmic operation of the subtask. For example, for a subtask type of addition, the addition type can be directly used as its task type. Each task type corresponds to a unique task type. Tasks of the same task type belong to the same task type, while different task types belong to different task types. In this embodiment, subtasks with the same task type can be configured to have the same major task type. For example, subtasks whose task types indicate that their algorithmic operation is addition can have the same major task type. Alternatively, subtasks whose task types indicate that their algorithmic operation is addition can have their major task type assigned to them as addition. In this embodiment, the user can pre-configure the data type and data volume of the intermediate data sets corresponding to different major task types. In this way, the scheduler 12 can split the data set according to the user-pre-configured data type and data volume to obtain multiple intermediate data sets.The data included in each major task type is then divided equally to obtain multiple sub-data sets. For example, if a major task type includes 10,000 data items, and this major task type includes 10 task types, then the number of data items in the sub-data set corresponding to each task type is determined to be 10,000 / 10 = 1,000. The 10,000 data items are then divided equally into 10 portions, each containing 1,000 data items. Each portion constitutes a sub-data set. In applications, GPUs have significant video memory for storing data and intermediate results required for computation. When simulating a GPU with a CPU, appropriate memory management techniques need to be developed to emulate the functions of GPU video memory, including data allocation, access speed optimization, and memory replication. Therefore, in this embodiment, the memory of the CPU service node is improved, and a portion of the storage space in the CPU service node 11 is used as cache to simulate the GPU's video memory. In some optional embodiments, the memory on each CPU service node 11 includes a cache and remaining storage space. The cache is used to store sub-data sets corresponding to the subtasks to be executed by the CPU service node 11, as well as intermediate results during task execution. The remaining storage space is used to store the processing results of the target CPU service node 11. In this embodiment, to improve memory utilization on the CPU service node 11, the cache size on the CPU service node 11 can also be adjusted. In specific implementation, in one optional embodiment, the scheduler 12 is further configured to: when executing each subtask on the target CPU service node 11, obtain the cache usage and remaining storage space usage of the target CPU service node 11; and adjust the cache size based on the cache usage and remaining storage space usage. Memory usage = a * cache usage + (1 - a) * remaining storage space usage. a is an adjustment parameter used to control the balance between cache and remaining storage space. Cache usage = used cache size / total cache size. It should be understood that if the cache usage rate shows an increasing trend and the remaining storage space usage rate shows a decreasing trend, the cache size is increased and the remaining storage space size is correspondingly decreased. If the cache usage rate shows a decreasing trend and the remaining storage space usage rate shows an increasing trend, the cache size is decreased and the remaining storage space size is correspondingly increased.In the solution provided in this embodiment, a server cluster virtual GPU is constructed from multiple CPU service nodes. The scheduler performs load balancing among these CPU service nodes. Each CPU service node is deployed with a vector driver component and an algorithm plug-in. The vector driver component can call the algorithm plug-in to execute tasks assigned by the scheduler, thereby enabling the server cluster to simulate a GPU. By simulating multiple CPU service nodes to function as a single GPU, the number of GPU cores can be significantly increased. According to test data, the performance of each simulated CPU core is equivalent to 10% of the GPU performance. When the number of CPU service nodes in the cluster exceeds 10, the overall system performance begins to surpass that of a standard GPU. Furthermore, by adding more CPU service nodes to the cluster, the system's computing performance can be further improved while maintaining a relatively low cost. The high-speed interconnection of the RDMA network effectively resolves data transmission bottlenecks and further improves overall system performance. Finally, the deployment of a vector driver component on each CPU service node helps fully utilize the CPU's vector instructions to call the algorithm plug-in, efficiently completing scheduler tasks through vector computing. The present disclosure also provides an application method for a virtual GPU system. The method can be applied to the scheduler in the aforementioned embodiment. As shown in FIG2 , the method may include the following steps: Step 201: splitting a task to be processed into multiple subtasks; Step 202: splitting a data set to be processed by the task to be processed into multiple subdata sets based on the task type of each subtask; Step 203: determining a target CPU service node for executing each subtask from multiple CPU service nodes; Step 204: distributing the subtask and the corresponding subdata set to the determined target CPU service node; Step 205: receiving processing results returned by each target CPU service node; and Step 206: aggregating the processing results to obtain a task processing result for the task to be processed.In some embodiments, splitting a task to be processed into multiple subtasks includes: obtaining multiple computing operations included in the task to be processed; obtaining an estimated execution time corresponding to each computing operation, where the estimated execution time is used to represent the time required to execute each computing operation; selecting a minimum estimated execution time from the estimated execution times; splitting a first computing operation among the multiple computing operations into multiple sub-computing operations, where the estimated execution time of each sub-computing operation is less than or equal to the minimum estimated execution time, and the estimated execution time of the first computing operation is greater than the minimum estimated execution time; generating the multiple subtasks based on each sub-computing operation and the second computing operation, where the execution time of the second computing operation is less than or equal to the minimum estimated execution time, and each sub-computing operation or the second computing operation corresponds to a subtask. Based on the task type of each subtask, splitting the data set to be processed by the pending task into multiple subdata sets includes: determining a task type corresponding to the task type of each subtask; splitting the data set into multiple intermediate data sets according to the task type; each task type uniquely corresponds to one intermediate data set; and evenly dividing the data in the intermediate data set corresponding to each task type according to the number of task types included in each task type to obtain multiple subdata sets. In some embodiments, obtaining the estimated execution time corresponding to each computing operation includes: obtaining the estimated execution time of a vector operation corresponding to each computing operation; and using the estimated execution time of the vector operation as the estimated execution time corresponding to each computing operation. In some embodiments, the method further includes: when executing each subtask on the target CPU service node, obtaining a cache usage rate and a remaining storage space usage rate of the target CPU service node; and adjusting the cache size based on the cache usage rate and the remaining storage space usage rate. In some embodiments, the method further includes: calculating a cache hit rate for the target CPU service node when executing each subtask on the target CPU service node; and adjusting the data size of the intermediate data set corresponding to each task type based on the cache hit rate. The virtual GPU system application device and the virtual GPU system application method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods employed, executed, or implemented therein.The presently disclosed embodiments further provide an application device for a virtual GPU system, which may be applied to the scheduler in the aforementioned embodiments. As shown in FIG3 , the device may include the following steps: a first splitting module 31 for splitting a task to be processed into multiple subtasks; a second splitting module 32 for splitting a data set to be processed by the task to be processed into multiple subdata sets based on the task type of each subtask; a determination module 33 for determining a target CPU service node for executing each subtask from the multiple CPU service nodes; a distribution module 34 for distributing the subtask and the corresponding subdata set to the determined target CPU service node; a receiving module 35 for receiving processing results returned by each target CPU service node; and a summarizing module 36 for summarizing the processing results to obtain a task processing result for the task to be processed. In some embodiments, the first splitting module 31 is used to: obtain multiple computing operations included in the task to be processed; obtain the estimated execution time corresponding to each computing operation, the estimated execution time is used to represent the time required to execute each computing operation; select the minimum estimated execution time from the estimated execution times; split the first computing operation among the multiple computing operations into multiple sub-computing operations, the estimated execution time of each sub-computing operation is less than or equal to the minimum estimated execution time, and the estimated execution time of the first computing operation is greater than the minimum estimated execution time; generate the multiple subtasks based on each sub-computing operation and the second computing operation, the execution time of the second computing operation is less than or equal to the minimum estimated execution time, and each sub-computing operation or the second computing operation corresponds to a subtask. In some embodiments, the second splitting module 32 is configured to: determine a major task type corresponding to the task type of each subtask; split the data set into multiple intermediate data sets according to the major task type; each major task type uniquely corresponds to one intermediate data set; and evenly divide the data in the intermediate data set corresponding to each major task type according to the number of task types included in each major task type to obtain multiple sub-data sets. In some embodiments, the first splitting module 31 is configured to: obtain an estimated execution time of a vector operation corresponding to each computing operation; and use the estimated execution time of the vector operation as the estimated execution time corresponding to each computing operation. In some embodiments, the device is further configured to: when executing each subtask on the target CPU service node, obtain a cache usage rate and a remaining storage space usage rate of the target CPU service node; and adjust the cache size based on the cache usage rate and the remaining storage space usage rate.In some embodiments, the apparatus is further configured to: calculate a cache hit rate for the target CPU service node when executing each of the subtasks on the target CPU service node; and adjust the data size of the intermediate data set corresponding to each of the major task types based on the cache hit rate. Embodiments of the present disclosure also provide an electronic device for executing the aforementioned virtual GPU system application method. Please refer to FIG4 , which illustrates a schematic diagram of an electronic device provided in some embodiments of the present disclosure. As shown in FIG4 , electronic device 4 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected via the bus 402. Memory 401 stores a computer program executable on processor 400. When processor 400 executes the computer program, the virtual GPU system application method provided in any of the aforementioned embodiments of the present disclosure is executed. Memory 401 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk drive. Communication between the device network element and at least one other network element is achieved through at least one communication interface 403 (which may be wired or wireless). The Internet, wide area network, local area network, metropolitan area network, etc. may be used. Bus 402 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Buses may be classified as address buses, data buses, control buses, etc. Memory 401 is used to store programs. Processor 400 executes the programs upon receiving execution instructions. The application method of the virtual GPU system disclosed in any of the aforementioned embodiments of the present disclosure may be applied to or implemented by processor 400. Processor 400 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits in processor 400 or by software instructions.The processor 400 described above can be a general-purpose processor, including a CPU (Central Processing Unit), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of this disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in storage media well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 401. Processor 400 reads information from memory 401 and, in conjunction with its hardware, completes the steps of the aforementioned method. The electronic device provided in the embodiments of the present disclosure and the virtual GPU system application method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods employed, executed, or implemented therein. The embodiments of the present disclosure also provide a computer-readable storage medium corresponding to the virtual GPU system application method provided in the aforementioned embodiments. Referring to FIG. 5 , the computer-readable storage medium shown is an optical disc 30 storing a computer program (i.e., a program product). When executed by the processor, the computer program executes the virtual GPU system application method provided in any of the aforementioned embodiments.It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical or magnetic storage media, which are not further described here. The computer-readable storage media provided in the above-mentioned embodiments of the present disclosure and the application methods of the virtual GPU system provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods employed, executed, or implemented by the application programs stored therein. The embodiments of the present disclosure also provide a computer program product including a computer program that, when executed by a processor, implements the method described in any of the above-mentioned embodiments. It should be noted that the description provided herein describes numerous specific details. However, it should be understood that the embodiments of the present disclosure can be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail to avoid obscuring the understanding of this description. Similarly, it should be understood that, in order to streamline the present disclosure and facilitate understanding of one or more of the various inventive aspects, in the above description of exemplary embodiments of the present disclosure, various features of the present disclosure are sometimes grouped together in a single embodiment, figure, or description thereof. However, this disclosure should not be interpreted as reflecting a schematic representation that the claimed disclosure requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single embodiment disclosed above. The claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present disclosure. Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the present disclosure and to form distinct embodiments.For example, in the following claims, any of the claimed embodiments may be used in any combination. The above are merely preferred embodiments of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed herein are intended to be encompassed by the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.
Claims
Claims 1. A virtual GPU system, wherein: include: A plurality of CPU service nodes and a scheduler communicating with each of the CPU service nodes; each of the CPU service nodes is deployed with a vector driver component and an algorithm plug-in for implementing a vector algorithm; each of the CPU service nodes communicates with each other through a remote direct data access network; the scheduler is used to schedule each of the CPU service nodes to execute pending tasks in parallel; each of the CPU service nodes is used to drive the algorithm plug-in deployed locally to execute the tasks scheduled by the scheduler through the locally deployed vector driver component, and share data with the CPU service nodes that need collaborative computing through the remote direct data access network during the execution of the tasks.
2. The system according to claim 1, wherein: The scheduler is used to split the task to be processed into multiple subtasks; split the data set to be processed by the task to be processed into multiple subdata sets based on the task type of each subtask; determine the target CPU service node for executing each subtask from the multiple CPU service nodes; distribute the subtask and the corresponding subdata set to the determined target CPU service node; receive the processing results returned by each target CPU service node, and summarize the processing results to obtain the task processing result of the task to be processed; the target CPU service node is used to drive the locally deployed algorithm plug-in through the locally deployed vector drive component, execute the received subtask on the received subdata set, and obtain the processing result; transmit the processing result to the scheduler.
3. The system according to claim 2, wherein: The memory on each of the CPU service nodes includes a cache and a remaining storage space; the cache is used to store a sub-data set corresponding to a sub-task that the CPU service node needs to execute and an intermediate result during the task execution; the remaining storage space is used to store the processing result of the CPU service node.
4. An application method of a virtual GPU system, wherein: Applied to the virtual GPU system described in any one of claims 1-3, the method includes: splitting a task to be processed into multiple subtasks; based on the task type of each subtask, splitting the data set to be processed by the task to be processed into multiple subdata sets; determining a target CPU service node for executing each subtask from the multiple CPU service nodes; distributing the subtask and the corresponding subdata set to the determined target CPU service node; receiving processing results returned by each target CPU service node; and aggregating the processing results to obtain the task processing result of the task to be processed.
5. The method according to claim 4, wherein: Splitting the task to be processed into multiple subtasks includes: obtaining multiple computing operations included in the task to be processed; obtaining an estimated execution time corresponding to each computing operation, wherein the estimated execution time is used to characterize the time required to execute each computing operation; selecting a minimum estimated execution time from the estimated execution times; splitting a first computing operation among the multiple computing operations into multiple sub-computing operations, wherein the estimated execution time of each sub-computing operation is less than or equal to the minimum estimated execution time, and the estimated execution time of the first computing operation is greater than the minimum estimated execution time; generating the multiple subtasks based on each sub-computing operation and a second computing operation, wherein the execution time of the second computing operation is less than or equal to the minimum estimated execution time, and each sub-computing operation or the second computing operation corresponds to a subtask.
6. The method according to claim 4 or 5, wherein: Based on the task type of each of the subtasks, the data set to be processed by the task to be processed is split into multiple sub-data sets, including: determining a task type corresponding to the task type of each of the subtasks; splitting the data set into multiple intermediate data sets according to the task type; each of the task types uniquely corresponds to an intermediate data set; and according to the number of task types included in each of the task types, evenly dividing the data in the intermediate data set corresponding to each of the task types to obtain multiple sub-data sets.
7. The method according to claim 5, wherein: Obtaining the estimated execution time corresponding to each of the computing operations, including: obtaining the estimated execution time of the vector computing operation corresponding to each of the computing operations; and using the estimated execution time of the vector computing operation as the estimated execution time corresponding to each of the computing operations.
8. The method according to any one of claims 4 to 7, wherein: Also includes: When each of the subtasks is executed on the target CPU service node, obtaining a usage rate of a cache and a usage rate of a remaining storage space included in the target CPU service node; The size of the cache is adjusted based on the usage rate of the cache and the usage rate of the remaining storage space.
9. An application device of a virtual GPU system, wherein: The virtual GPU system applied to any one of claims 1-3, the device comprising: a first splitting module for splitting a task to be processed into a plurality of subtasks; a second splitting module for splitting a data set to be processed by the task to be processed into a plurality of sub-data sets based on the task type of each subtask; a determination module for respectively determining a target CPU service node for executing each of the subtasks from the plurality of CPU service nodes; a distribution module, used to distribute the subtasks and the corresponding sub-data sets to the determined target CPU service nodes; a receiving module, used to receive the processing results returned by each of the target CPU service nodes; and a summarizing module, used to summarize the processing results to obtain the task processing result of the task to be processed.
10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor runs the computer program to implement the method according to any one of claims 4 to 8.
11. A computer-readable storage medium having a computer program stored thereon, wherein: The program is executed by a processor to implement the method according to any one of claims 4 to 8.
12. A computer program product, comprising a computer program, wherein: When the computer program is executed by a processor, the method according to claims 4 to 8 is implemented. 16
Citation Information
Patent Citations
Fusion calculation method and readable storage medium
CN111258655A
System for deep learning, method for processing data and electronic equipment
CN116069511A
Resource management method for heterogeneous multi-core system
CN117112169A
Central processing unit, GPU simulation method thereof, and computing system including the same
US20130207983A1
Memory mat as a register file
US20220269645A1
Cited By
GPU (Graphic Processing Unit) task creation system and method, graphics processing unit and electronic equipment
CN120909745A