GPGPU Computing Task Processing Method, Apparatus, Device, and Medium
By using CXL switch and resource evaluation algorithm in the GPGPU scheduling control end, the coordinated processing of multiple GPGPU computing cards is realized, solving the problems of high computing delay and excessive host load, and improving task processing efficiency and flexibility.
Patent Information
- Application Number
- CN202510329294.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the prior art, multiple GPGPU computing cards cannot coordinate the completion of computing tasks, resulting in high computing delays and low efficiency, and excessive load on the central processor of the host, which cannot effectively handle high concurrent computing tasks.
The task allocation terminal and the GPGPU resource pool are connected through the CXL switch, and the optimal task processing mode is filtered out using the resource evaluation algorithm, and the calculation task is split into multiple subtasks, which are allocated to multiple GPGPU computing cards for collaborative execution.
It significantly improves the processing efficiency of GPGPU computing tasks, reduces the load on the host side, supports unlimited expansion of the number of calculation cards, and improves data communication efficiency and task execution flexibility.
Smart Images

Figure CN119829300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a method, device, equipment and medium for processing GPGPU computing tasks. Background Art
[0002] GPGPU (General Purpose GPU, i.e., General Purpose Graphics Processing Unit) is a powerful computing tool. Different from GPU (Graphics Processing Unit), the computing core of GPGPU has a higher parallelism and is more suitable for some non-graphic related program operations, such as repetitive tasks with a large amount of data, like large-scale data encryption, decryption, data calculation, AI (Artificial Intelligence) computing acceleration, etc.
[0003] With the rapid development of the current technological level, the application demand for GPGPU has increased significantly. To meet the computing-intensive application requirements such as data centers, an architecture with multiple GPGPUs installed on a single host is proposed. This design supports the host to simultaneously call multiple concurrent heterogeneous GPGPUs to execute large-scale computations, maximizing the computing efficiency without affecting the computing requirements.
[0004] However, there are still some problems in the practical application of this design. On the one hand, multiple GPGPU computing cards are connected to the host side (Host) in a concurrent form through the PCIe interface (peripheral component interconnect express, a high-speed serial computer expansion bus standard). Different GPGPU computing cards do not support collaborative completion of computing tasks, and a single computing task can only be assigned to one GPGPU computing card. Due to the limited on-chip resources of a single GPGPU computing card, this directly leads to extremely high computing latency and low computing efficiency for some complex computing tasks. On the other hand, the computing task scheduling of the traditional architecture is all executed by the CPU (Central Processing Unit) on the host side. For the task scheduling with high concurrency in a single time period, it requires a large amount of computing power of the host central processor, which will inevitably lead to an excessive load on the host and affect the execution of other important applications.
[0005] In summary, how to improve the processing efficiency of GPGPU computing tasks and reduce the load on the host side is an issue to be solved in this field. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a GPGPU computing task processing method, device, equipment and medium, which can improve the GPGPU computing task processing efficiency and reduce the load on the host side. The specific scheme is as follows:
[0007] In a first aspect, the present application discloses a GPGPU computing task processing method, which is applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host end. The task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; the method includes:
[0008] Receiving the GPGPU computing task sent by the host end;
[0009] Evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table;
[0010] Using a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task;
[0011] If the target task processing mode is a multi-card pipelining processing mode, determining each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and splitting the GPGPU computing task into multiple computing subtasks based on the current resource status information;
[0012] Transmitting each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and returning the processing results of the computing subtasks output by the target GPGPU computing cards to the host end through the CXL switch.
[0013] Optionally, the task allocation terminal further includes DDR5 and a first network interface, and the host end includes a central processing unit, a second network interface and a cache. Among them, the cache is connected to the DDR5 through a PCIe bus, and the first network interface and the second network interface are respectively connected to the CXL switch through a CAT6e network cable.
[0014] Optionally, the GPGPU resource pool includes multiple GPGPU computing cards. Among them, the GPGPU computing cards with adjacent identification information are connected by optical fibers, and each GPGPU computing card is connected to the CXL switch through a CAT6e network cable; the method further includes:
[0015] Determine the positioning information of each of the target GPGPU computing cards;
[0016] Correspondingly, the transmitting of each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch includes:
[0017] Transmit each of the computing subtasks to each of the target GPGPU computing cards corresponding to the positioning information through the CXL switch.
[0018] Optionally, the task description information includes data to be processed, task-related instructions, and task operation data;
[0019] Correspondingly, the evaluating of the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table includes:
[0020] Evaluate the data to be processed to obtain the total amount of data to be processed;
[0021] Determine the storage location of the task-related instructions in the instruction task comparison table according to the hash result of the task-related instructions; wherein, the instruction task comparison table includes each preset task execution strategy and the strategy complexity corresponding to each of the preset task execution strategies;
[0022] Based on the storage location, read the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table.
[0023] Optionally, the processing of the total amount of data to be processed, the target task execution strategy, and the target strategy complexity by using a resource evaluation algorithm to screen out the target task processing mode of the GPGPU computing task includes:
[0024] Read the number of splittable computing subtasks of the GPGPU computing task from the instruction task comparison table based on the task description information;
[0025] Obtain the task execution evaluation factor of the GPGPU computing task; wherein, the task execution evaluation factor includes the strategy complexity reference value corresponding to the target task execution strategy, the total resources of a single GPGPU computing card, and the optical fiber transmission speed between GPGPU computing cards;
[0026] Process the total amount of data to be processed, the target task execution strategy, the target strategy complexity, the number of splittable computing subtasks of the GPGPU computing task, and the task execution evaluation factor by using a resource evaluation algorithm to obtain a first estimated time for processing the GPGPU computing task in a single-card independent processing mode and a second estimated time for processing the GPGPU computing task in a multi-card pipelined processing mode;
[0027] Select a target task processing mode for the GPGPU computing task from the single-card independent processing mode and the multi-card pipelined processing mode according to the first estimated time and the second estimated time.
[0028] Optionally, the selecting a target task processing mode for the GPGPU computing task from the single-card independent processing mode and the multi-card pipelined processing mode according to the first estimated time and the second estimated time includes:
[0029] If the first estimated time is less than the second estimated time, determine that the single-card independent processing mode is the target task processing mode;
[0030] If the first estimated time is not less than the second estimated time, determine that the multi-card pipelined processing mode is the target task processing mode.
[0031] Optionally, the expression of the resource evaluation algorithm is:
[0032] ;
[0033] ;
[0034] Wherein, is the number of splittable computing subtasks of the GPGPU computing task, is the maximum value of the number of splittable computing subtasks, is the target strategy complexity corresponding to the target task execution strategy, is the strategy complexity reference value corresponding to the target task execution strategy, is the total amount of data to be processed, is the total resources of a single GPGPU computing card, is the first estimated time for processing the GPGPU computing task in the single-card independent processing mode, is the optical fiber transmission speed between GPGPU computing cards, is the second estimated time for processing the GPGPU computing task in the multi-card pipelined processing mode.
[0035] In a second aspect, the present application discloses a GPGPU computing task processing device, which is applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host end. The task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; the device includes:
[0036] A task receiving module, configured to receive the GPGPU computing tasks sent by the host end;
[0037] A policy reading module, configured to evaluate the task description information in the GPGPU computing tasks to obtain the total amount of data to be processed, and read the target task execution policy and the target policy complexity corresponding to the task description information from the instruction task comparison table;
[0038] A mode determination module, configured to process the total amount of data to be processed, the target task execution policy, and the target policy complexity by using a resource evaluation algorithm to screen out the target task processing mode of the GPGPU computing tasks;
[0039] A task splitting module, configured to, if the target task processing mode is a multi-card pipelining processing mode, determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing tasks into multiple computing subtasks based on the current resource status information;
[0040] A task processing module, configured to transmit each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and return the processing results of the computing subtasks output by the target GPGPU computing cards to the host end through the CXL switch.
[0041] In a third aspect, the present application discloses an electronic device, including:
[0042] A memory, configured to store a computer program;
[0043] A processor, configured to execute the computer program to implement the steps of the GPGPU computing task processing method disclosed above.
[0044] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the GPGPU computing task processing method disclosed above are implemented.
[0045] The beneficial effects of this application are as follows: This application is applied to the field-programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host end. The task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; The method includes: receiving the GPGPU computing task issued by the host end; evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table; using a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task; if the target task processing mode is the multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information; transmit each computing subtask to each target GPGPU computing card through the CXL switch, and return the processing results of the computing subtasks output by the target GPGPU computing card to the host end through the CXL switch. Thus, it can be seen that this application is applied to the field-programmable gate array of the task allocation terminal in the GPGPU scheduling control end, rather than the host end. That is to say, when there is a GPGPU computing task, the allocation of the GPGPU computing task and the scheduling of the GPGPU computing card are completed by the field-programmable gate array of the task allocation terminal. Based on the concurrent processing ability of the field-programmable gate array, the load of the central processing unit of the host end can be significantly reduced; Secondly, the task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch, which supports infinitely expanding the number of GPGPU computing cards according to requirements through a network cable, has quite strong scalability, and based on this connection relationship, the CXL protocol is used to complete data transmission, improving data communication efficiency; Further, a resource evaluation algorithm is used to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task. Therefore, the task execution strategy in the target task processing mode not only has low complexity but also high efficiency; Next, if the target task processing mode is the multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks. In this way, the target GPGPU computing cards with suitable resource status are used to process the computing subtasks, forming an inter-card pipeline, and multiple cards cooperate to execute the calculation, with high application flexibility and significantly improving the task processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0047] Figure 1 It is a flowchart of a method for processing GPGPU computing tasks disclosed in the present application;
[0048] Figure 2 It is a schematic diagram of a specific GPGPU scheduling control end disclosed in the present application;
[0049] Figure 3 It is a schematic diagram of a specific GPGPU computing task processing disclosed in the present application;
[0050] Figure 4 It is a schematic diagram of a specific field-programmable gate array kernel architecture disclosed in the present application;
[0051] Figure 5 It is a schematic diagram of a specific GPGPU resource pool disclosed in the present application;
[0052] Figure 6 It is a schematic diagram of a specific field-programmable gate array kernel processing flow disclosed in the present application;
[0053] Figure 7 It is a schematic diagram of the structure of a GPGPU computing task processing device disclosed in the present application;
[0054] Figure 8 It is a schematic diagram of the structure of an electronic device disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0056] GPGPU is a powerful computing tool. Different from GPUs, GPGPU has a higher degree of parallelism in its computing cores and is more proficient in some non-graphic related program operations, repetitive tasks with a large amount of data, such as large-scale data encryption, decryption, data calculation, AI computing acceleration, etc.
[0057] With the rapid advancement of current technology levels, the application demand for GPGPUs has increased significantly. To meet the computational intensive application demands such as in data centers, an architecture with multiple GPGPUs on a single host is proposed. This design supports the host to simultaneously call multiple concurrent heterogeneous GPGPUs to execute large-scale computations, maximizing the computational efficiency without affecting the computational requirements.
[0058] However, there are also some problems in the practical application of this design. On the one hand, multiple GPGPU computing cards are connected to the host side in a concurrent form through the PCIe interface. Different GPGPU computing cards do not support collaborative completion of computational tasks. A single computational task can only be assigned to one GPGPU computing card. Due to the limited on-chip resources of a single GPGPU computing card, this directly leads to extremely high computational latency and low computational efficiency for some complex computational tasks. On the other hand, the computational task scheduling work of the traditional architecture is all executed by the central processor on the host side. For the task scheduling of some high-concurrency tasks in a single time period, it requires a large amount of computing power of the host central processor, which is bound to cause an excessive load on the host and affect the execution of other important applications.
[0059] Therefore, the present application correspondingly provides a GPGPU computational task processing solution to improve the GPGPU computational task processing efficiency and reduce the load on the host side.
[0060] See Figure 1 As shown, an embodiment of the present application discloses a GPGPU computational task processing method, which is applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host side. The task allocation terminal and the host side are respectively connected to the GPGPU resource pool through a CXL switch; the method includes:
[0061] Step S11: Receive the GPGPU computational task sent by the host side.
[0062] In this embodiment, the task allocation terminal further includes DDR5 and a first network interface, and the host side includes a central processor, a second network interface, and a cache. Among them, the cache is connected to the DDR5 through a PCIe bus, and the first network interface and the second network interface are respectively connected to the CXL switch through CAT6e network cables.
[0063] For example Figure 2A schematic diagram of a specific GPGPU scheduling control terminal is shown. The GPGPU scheduling control terminal includes a task allocation terminal and a host terminal. Among them, the task allocation terminal includes DDR5 (Double Data Rate 5 Synchronous Dynamic Random Access Memory, that is, the fifth-generation double data rate synchronous dynamic random access memory), a first network interface, and a field programmable gate array (Field Programmable Gate Array, that is, a field programmable gate array). The host terminal includes a central processing unit, a second network interface, and a cache. The cache of the host terminal is connected to the DDR5 of the task allocation terminal through a PCIe bus. The first network interface and the second network interface are respectively connected to a CXL switch (Switch) through a CAT6e (Category 6 Enhanced) network cable.
[0064] For example Figure 3 A schematic diagram of a specific GPGPU computing task processing is shown. The central processing unit of the host terminal receives the GPGPU computing task issued by the user. The cache of the host terminal saves the GPGPU computing task and at the same time informs the task allocation terminal through the CXL switch that there is a task allocation request. Further, the DDR5 of the task allocation terminal is responsible for completing data exchange with the cache of the host terminal, that is, the DDR5 of the task allocation terminal obtains the GPGPU computing task issued by the host terminal. That is to say, the task allocation terminal maps the GPGPU computing task to the DDR5 memory in the form of DMA (Direct Memory Access, that is, direct memory access) through a PCIe interface. In this way, the field programmable gate array of the task allocation terminal obtains the GPGPU computing task from the DDR5.
[0065] Step S12: Evaluate the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and read the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table.
[0066] In this embodiment, the task description information includes data to be processed, task-related instructions, and task operation data. The field programmable gate array reads the task description information of the GPGPU computing task from the DDR5, including data to be processed, task-related instructions, and task operation data. The data to be processed is the original processing data of the GPGPU computing task.
[0067] In this embodiment, evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table includes: evaluating the data to be processed to obtain the total amount of data to be processed; determining the storage location of the task-related instructions in the instruction task comparison table according to the hash result of the task-related instructions; where the instruction task comparison table includes each preset task execution strategy and the strategy complexity corresponding to each preset task execution strategy; and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table based on the storage location.
[0068] For example Figure 4 A schematic diagram of a specific field-programmable gate array kernel architecture is shown. Among them, the field-programmable gate array kernel architecture includes a management information configuration module, a DDR5 read / write module, a user instruction and data to be processed decomposition module, a user instruction cache module, a data to be processed cache module, an instruction comparison table reading module, a data volume evaluation module, an instruction task comparison table, a GPGPU computing resource evaluation module required for the task, a GPGPU computing card allocation FSM (Finite State Machine), an allocation result cache module, a task sending cache, a GPGPU computing card status comparison table read / write module, a CXL protocol conversion module, and a GPGPU pooled resource status table.
[0069] First, parameterized configuration is performed on each module in the field-programmable gate array. Specifically, the management information configuration module completes the corresponding necessary parameterized settings for modules such as the DDR5 read / write module, the user instruction and data to be processed decomposition module, the user instruction cache module, the data to be processed cache module, the instruction comparison table reading module, the data volume evaluation module, the instruction task comparison table, the GPGPU computing resource evaluation module required for the task, and the GPGPU computing card allocation FSM (Finite State Machine). The DDR5 read / write module reads the user's GPGPU computing task from the DDR5 memory, and sequentially sends the task description information of the GPGPU computing task to the user instruction and data to be processed decomposition module in a specific format. The user instruction and data to be processed decomposition module decomposes all the received task description information, and the task-related instructions are cached by the user instruction cache module, the data to be processed is cached by the data to be processed cache module, and the GPGPU computing task is cached by the task sending cache module.
[0070] The data volume evaluation module evaluates the total amount of data to be processed to obtain the total amount of data to be processed; the instruction comparison table reading module determines the storage location of the task-related instructions in the instruction task comparison table according to the hash result of the task-related instructions. The instruction task comparison table includes each preset task execution strategy and the corresponding strategy complexity. Then, the target task execution strategy and the target strategy complexity corresponding to the task description information are read from this storage location.
[0071] Step S13: Use a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task.
[0072] The GPGPU computing resource evaluation module required for the task executes a resource evaluation algorithm according to the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to obtain the optimal target task processing mode of the GPGPU computing task, that is, whether the GPGPU computing task is more suitable for the single-card independent processing mode or the multi-card pipelined processing mode. Among them, the single-card independent processing mode means that a single GPGPU computing card independently completes the entire GPGPU computing task, and the multi-card pipelined processing mode means that multiple GPGPU computing cards cooperate to complete the GPGPU computing task in a pipeline form.
[0073] In this embodiment, the use of a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task includes: reading the number of splittable computing subtasks of the GPGPU computing task from the instruction task comparison table based on the task description information; obtaining the task execution evaluation factor of the GPGPU computing task; where the task execution evaluation factor includes the reference value of the strategy complexity corresponding to the target task execution strategy, the total resources of a single GPGPU computing card, and the optical fiber transmission speed between GPGPU computing cards; using a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, the target strategy complexity, the number of splittable computing subtasks of the GPGPU computing task, and the task execution evaluation factor to obtain the first estimated time for processing the GPGPU computing task in the single-card independent processing mode and the second estimated time for processing the GPGPU computing task in the multi-card pipelined processing mode; and screening out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelined processing mode according to the first estimated time and the second estimated time.
[0074] Specifically, the number n of splittable computing subtasks of the GPGPU computing task can be read from the instruction task comparison table based on the task-related instructions in the task description information, that is, this number represents how many computing subtasks the GPGPU computing task can be split into. For example, it can be split into 10 computing subtasks, 8 computing subtasks, etc.; obtain the task execution evaluation factor of the GPGPU computing task, and the task execution evaluation factor includes the policy complexity reference value corresponding to the target task execution strategy, the total resources of a single GPGPU computing card, and the optical fiber transmission speed between GPGPU computing cards. It can be understood that the task execution evaluation factor includes the influence factor of the computing card characteristics on the computing task; use the resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, the target policy complexity, the number of splittable computing subtasks of the GPGPU computing task, and the task execution evaluation factor to obtain the first estimated time for processing the GPGPU computing task in the single-card independent processing mode and the second estimated time for processing the GPGPU computing task in the multi-card pipelining processing mode; then screen out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelining processing mode according to the first estimated time and the second estimated time. That is to say, in this embodiment, when screening the processing mode, various factors are considered, so as to obtain more accurate first and second estimated times, providing a strong basis for mode selection.
[0075] In this embodiment, the expression of the resource evaluation algorithm is:
[0076] ;
[0077] ;
[0078] Among them, is the number of splittable computing subtasks of the GPGPU computing task, is the maximum value of the number of splittable computing subtasks, is the target policy complexity corresponding to the target task execution strategy, is the policy complexity reference value corresponding to the target task execution strategy, is the total amount of data to be processed, is the total resources of a single GPGPU computing card, is the first estimated time for processing the GPGPU computing task in the single-card independent processing mode, is the optical fiber transmission speed between GPGPU computing cards, is the second estimated time for processing the GPGPU computing task in the multi-card pipelining processing mode.
[0079] It can be understood that in the single-card independent processing mode, a computing card independently processes tasks. Therefore, there is no transmission between computing cards in the single-card independent processing mode, and the total resources of a single computing card are used to process each subtask. So the parameters of the Intelligent resource evaluation algorithm include the number of splittable computing subtasks of the GPGPU computing task, the maximum value of the number of splittable computing subtasks, the target policy complexity corresponding to the target task execution policy, the reference value of the policy complexity corresponding to the target task execution policy, the total amount of data to be processed, and the total resources of a single GPGPU computing card, so as to obtain the first estimated time for processing the GPGPU computing task in the single-card independent processing mode; further, because in the multi-card pipelining processing mode, multiple computing cards cooperate to process tasks, involving transmission between computing cards, and the total resources of a single computing card are used to process a single subtask, so in addition to the parameters in the single-card independent processing mode, the Intelligent resource evaluation algorithm also includes the optical fiber transmission speed between GPGPU computing cards, so as to obtain the second estimated time for processing the GPGPU computing task in the multi-card pipelining processing mode.
[0080] In this embodiment, the method of screening out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelining processing mode according to the first estimated time and the second estimated time includes: if the first estimated time is less than the second estimated time, it is determined that the single-card independent processing mode is the target task processing mode; if the first estimated time is not less than the second estimated time, it is determined that the multi-card pipelining processing mode is the target task processing mode.
[0081] It can be understood that if the first estimated time is less than the second estimated time, that is less than , it means that the single-card independent processing mode takes less time to process tasks and has higher efficiency. Therefore, it is determined that the single-card independent processing mode is the target task processing mode; if the first estimated time is not less than the second estimated time, greater than or equal to , it means that the multi-card pipelining processing mode takes less time to process tasks, or the load on each computing card is smaller. Therefore, it is determined that the multi-card pipelining processing mode is the target task processing mode.
[0082] Step S14: If the target task processing mode is the multi-card pipelining processing mode, determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information.
[0083] The GPGPU pooling resource status table stores the current resource status information of the GPGPU resource pool, i.e., the GPGPU computing cards in the idle state, the GPGPU computing cards in the busy state, the GPGPU computing cards in the fault state, etc. And the read-write module of the GPGPU computing card status comparison table updates the GPGPU pooling resource status table in real time after the GPGPU computing task is completed and distributed and when the resource status information of the GPGPU resource pool is updated, so that the various information recorded in the GPGPU pooling resource status table is real-time and accurate.
[0084] In a specific embodiment, if the target task processing mode is the multi-card pipelining processing mode, the GPGPU computing card allocation FSM controls the read-write module of the GPGPU computing card status comparison table to obtain the current resource status information of the GPGPU resource pool in the GPGPU pooling resource status table, so as to determine each target GPGPU computing card corresponding to the multi-card pipelining processing mode from the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information, so as to subsequently send each computing subtask to each target GPGPU computing card respectively.
[0085] This embodiment also includes another specific embodiment. If the target task processing mode is the single-card independent processing mode, the GPGPU computing card allocation FSM controls the read-write module of the GPGPU computing card status comparison table to obtain the current resource status information of the GPGPU resource pool in the GPGPU pooling resource status table, so as to determine a single target GPGPU computing card corresponding to the single-card independent processing mode from the GPGPU resource pool, and there is no need to split the GPGPU computing task, so as to subsequently send the GPGPU computing task to the target GPGPU computing card.
[0086] In this embodiment, the GPGPU resource pool includes multiple GPGPU computing cards. Among them, the GPGPU computing cards with adjacent identification information are connected by optical fibers, and each GPGPU computing card is connected to the CXL switch through a CAT6e network cable.
[0087] For example Figure 5A schematic diagram of a specific GPGPU resource pool is shown. The GPGPU resource pool includes multiple GPGPU computing cards. Among them, the GPGPU computing cards with adjacent identification information are connected by optical fibers. In this way, data transmission can be carried out between the computing cards, and each GPGPU computing card is connected to the CXL switch through a CAT6e network cable. Therefore, each computing card can interact with the task allocation terminal and the host terminal in the GPGPU scheduling control end through the CXL switch. The identification information is, for example, the ID (Identity document) of the computing card. When the computing cards with adjacent identification information are connected by optical fibers, it is the connection between the optical port transmitting device and the optical port receiving device. Specifically, there is a fiber connection between the optical port transmitting device of the computing card with ID No.1 and the optical port receiving device of the computing card with ID No.2, and there is also a fiber connection between the optical port transmitting device of the computing card with ID No.2 and the optical port receiving device of the computing card with ID No.3, and there is also a fiber connection between the optical port transmitting device of the computing card with ID No.3 and the optical port receiving device of the computing card with ID No.4, etc.
[0088] In this embodiment, it further includes: determining the positioning information of each of the target GPGPU computing cards. When determining each target GPGPU computing card, the positioning information of each target GPGPU computing card is also determined, where the positioning information includes identification information and CXL address.
[0089] The target GPGPU computing card to which the GPGPU computing task is allocated is determined, that is, the allocation result is obtained. Subsequently, the GPGPU computing card allocation FSM sends the task description information and the positioning information of the target GPGPU computing card to the task sending buffer module according to the allocation result. The GPGPU computing card allocation FSM also sends the allocation result to the allocation result buffer module, and caches the allocation result of the GPGPU computing task to DDR5 through the DDR5 read / write module, and then transmits it to the central processing unit of the host end, so that the central processing unit of the host end can accurately read the calculation result from the target computing card from the CXL switch after the calculation is completed.
[0090] Step S15: Transmit each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and return the processing result of the computing subtask output by the target GPGPU computing card to the host end through the CXL switch.
[0091] The task sending buffer splices the task description information and the positioning information of the target GPGPU computing card according to the target task processing mode, and sends the spliced information to the CXL protocol conversion module. The CXL protocol conversion module performs protocol conversion on the spliced information according to the requirements of the CXL3.2 protocol to obtain the converted information.
[0092] Further, if it is a single - card independent processing mode, directly splice the task - related instructions, task operation data, and the positioning information of the target GPGPU computing card. If it is a multi - card pipelining processing mode, based on the allocation requirements of the multi - card pipelining processing mode, configure different task - related instructions and task operation data for the positioning information of each target GPGPU computing card, and then splice the positioning information of each target GPGPU computing card with the corresponding task - related instructions and task operation data; there is a sequential order of task processing among the target GPGPU computing cards. For example, the sequential order of each target GPGPU computing card is computing card No.1, computing card No.2, and computing card No.3. Configure task - related instruction 1, task operation data 1 for computing card No.1, task - related instruction 2, task operation data 2 for computing card No.2, and task - related instruction 3, task operation data 3 for computing card No.3. Splice the positioning information of computing card No.1 with task - related instruction 1 and task operation data 1 to obtain the spliced information 1 of computing card No.1, the spliced information 2 of computing card No.2, and the spliced information 2 of computing card No.2. Then, for subsequent protocol conversion, obtain the converted information 1 of computing card No.1, the converted information 2 of computing card No.2, and the converted information 3 of computing card No.3.
[0093] In this embodiment, the step of transmitting each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch includes: transmitting each of the computing subtasks to each of the target GPGPU computing cards corresponding to the positioning information through the CXL switch. If it is a multi - card pipelining processing mode, transmit each of the computing subtasks and the converted information to each of the target GPGPU computing cards corresponding to the positioning information through the CXL switch according to the positioning information of the target GPGPU computing cards. Specifically, send the data to be processed, the first computing subtask, and the converted information corresponding to the first computing subtask to the first target GPGPU computing card through the CXL switch. For example, send the data to be processed, computing subtask 1, and the converted information 1 corresponding to computing subtask 1 to computing card No.1, while sending the converted information 2 and computing subtask 2 to computing card No.2, and sending the converted information 3 and computing subtask 3 to computing card No.3.
[0094] If it is a single - card independent processing mode, transmit the GPGPU computing task and the converted information to the target GPGPU computing card corresponding to the positioning information through the CXL switch according to the positioning information of the target GPGPU computing card.
[0095] In the multi-card pipelined processing mode, each target GPGPU computing card performs pipelined processing on its corresponding computing subtask, and the last target GPGPU computing card outputs the processing result of the final GPGPU computing task, and returns the processing result of the final GPGPU computing task to the host through the CXL switch. Specifically, for example, the GPGPU computing task is split into computing subtask 1, computing subtask 2, and computing subtask 3. Among them, the target GPGPU computing cards are computing card No. 1, computing card No. 2, and computing card No. 3 respectively. After computing card No. 1 finishes processing computing subtask 1, it transmits the processing result of computing subtask 1 to computing card No. 2 through optical fiber. Computing card No. 2 processes computing subtask 2 based on the processing result of computing subtask 1, thereby obtaining the processing result of computing subtask 2, and transmits the processing result of computing subtask 2 to computing card No. 3 through optical fiber. Computing card No. 3 processes computing subtask 3 based on the processing result of computing subtask 2, thereby obtaining the processing result of computing subtask 3. The processing result of this computing subtask 3 is the processing result of the final GPGPU computing task. Therefore, the processing result of computing subtask 3 is returned to the host through the CXL switch.
[0096] In the single-card independent processing mode, the target GPGPU computing card independently processes the GPGPU computing task to obtain the processing result of the final GPGPU computing task, and returns the processing result of the final GPGPU computing task to the host through the CXL switch.
[0097] It should be noted that it can be that the host actively reads the processing result of the final GPGPU computing task according to the positioning information, or it can be that the target GPGPU computing card actively returns the processing result of the final GPGPU computing task to the host. Among them, whether the host actively obtains the processing result or the computing card actively sends the processing result to the host, the processing result of the GPGPU computing task is returned to the second network port of the host through the CXL switch.
[0098] The beneficial effects of this application are as follows: This application is applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host end. The task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; The method includes: receiving the GPGPU computing task issued by the host end; evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table; using a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task; if the target task processing mode is a multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information; transmit each computing subtask to each target GPGPU computing card through the CXL switch, and return the processing results of the computing subtasks output by the target GPGPU computing card to the host end through the CXL switch. It can be seen that this application is applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end, rather than the host end. That is to say, when there is a GPGPU computing task, the allocation of the GPGPU computing task and the scheduling of the GPGPU computing card are completed by the field programmable gate array of the task allocation terminal. Based on the concurrent processing ability of the field programmable gate array, the load of the central processor of the host end can be significantly reduced; Secondly, the task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch, which supports infinitely expanding the number of GPGPU computing cards through network cables according to requirements, has quite strong scalability, and based on this connection relationship, the CXL protocol is used to complete data transmission, improving data communication efficiency; Further, a resource evaluation algorithm is used to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task. Therefore, the task execution strategy in the target task processing mode not only has low complexity but also high efficiency; Next, if the target task processing mode is a multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks. In this way, the target GPGPU computing cards with suitable resource status are used to process the computing subtasks, forming an inter-card pipeline, and multiple cards cooperate to execute the calculation, with high application flexibility and significantly improving the task processing efficiency.
[0099] For example Figure 6 As shown in the schematic diagram of a specific field-programmable gate array (FPGA) kernel processing flow, the FPGA kernel of the GPGPU task allocation terminal includes a management information configuration module, a DDR5 read / write module, a user instruction and data to be processed decomposition module, a user instruction cache module, a data to be processed cache module, an instruction look-up table reading module, a data volume evaluation module, an instruction task look-up table, a GPGPU computing resource evaluation module for tasks, a GPGPU computing card allocation FSM state machine, an allocation result cache module, a task sending cache, a GPGPU computing card state look-up table read / write module, a CXL protocol conversion module, and a GPGPU pooled resource status table. Each module is used to determine the target task processing mode of the GPGPU computing task, and the determined GPGPU computing task is transmitted to the target GPGPU computing card through a CXL switch. The specific process is as follows:
[0100] 1) The management information configuration module completes the corresponding necessary parameter settings for modules such as the DDR5 read / write module, the user instruction and data to be processed decomposition module, the user instruction cache module, the data to be processed cache module, the instruction look-up table reading module, the data volume evaluation module, the instruction task look-up table, the GPGPU computing resource evaluation module for tasks, and the GPGPU computing card allocation FSM;
[0101] 2) The DDR5 read / write module reads the GPGPU computing task from the DDR5 memory, and sequentially sends the data to be processed, task-related instructions, and task operation data to the user instruction and data to be processed decomposition module in a specific format. This module caches all the received relevant instruction data and data to be processed, which are cached by the user instruction cache module, the data to be processed cache module, and the task sending cache module respectively;
[0102] 3) The data volume evaluation module evaluates the total amount of data to be processed in the GPGPU computing task; the instruction look-up table reading module performs a hash calculation on the instruction encoding according to the task-related instructions and obtains the storage location of the current instruction in the instruction task look-up table, and reads the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task look-up table; subsequently, the total amount of data to be processed, the target task execution strategy, and the target strategy complexity are sent to the GPGPU computing resource evaluation module for tasks;
[0103] 4) The GPGPU computing resource evaluation module for tasks executes a resource evaluation algorithm according to the total amount of data to be processed, the target task execution strategy, and the target strategy complexity, and obtains the target task processing mode of the GPGPU computing task;
[0104] 5) The GPGPU computing card allocation FSM controls the read / write module of the GPGPU computing card status comparison table, and reads information such as the current idle GPGPU computing card ID and CXL address stored in the GPGPU pooling resource status table. Subsequently, according to the requirements of the target task processing mode, the FSM state machine determines the target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool based on the current resource status information of the GPGPU resource pool. It can be understood that if it is a single-card independent processing mode, the number of target GPGPU computing cards is 1; if it is a multi-card pipelined processing mode, the number of target GPGPU computing cards is multiple. Then, it sends the task-related instructions, the target GPGPU computing card ID, and the CXL address and other information to the task sending buffer module.
[0105] 6) After the allocation is completed, the GPGPU computing card allocation FSM instructs the read / write module of the GPGPU computing card status comparison table to update the current resource status information of the GPGPU resource pool in the GPGPU pooling resource status table in real time. The GPGPU computing card allocation FSM also sends the allocation result to the allocation result buffer module, and caches the current task allocation result to the DDR5 through the DDR5 read / write module, and then transmits it to the host central processing unit, so that the host central processing unit can accurately read the calculation result from the target GPGPU computing card after the calculation is completed.
[0106] 7) The task sending buffer splices the task description information and the positioning information of the target GPGPU computing card according to the target task processing mode, and sends the spliced information to the CXL protocol conversion module.
[0107] 8) The CXL protocol conversion module converts the spliced information according to the requirements of the CXL3.2 protocol to obtain the converted information, and then sends the converted information to the target GPGPU computing card through the network cable via the CXL switch.
[0108] See Figure 7 As shown in
[0109] A task receiving module 11, configured to receive the GPGPU computing task issued by the host end.
[0110] A policy reading module 12 is configured to evaluate task description information in the GPGPU computing task to obtain the total amount of data to be processed, and read a target task execution policy and a target policy complexity corresponding to the task description information from an instruction task comparison table;
[0111] A mode determination module 13 is configured to process the total amount of data to be processed, the target task execution policy, and the target policy complexity by using a resource evaluation algorithm to screen out a target task processing mode of the GPGPU computing task;
[0112] A task splitting module 14 is configured to, if the target task processing mode is a multi-card pipelining processing mode, determine respective target GPGPU computing cards corresponding to the target task processing mode from the GPGPU resource pool according to current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information;
[0113] A task processing module 15 is configured to transmit each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and return a processing result of the computing subtasks output by the target GPGPU computing cards to the host through the CXL switch.
[0114] The beneficial effects of this application are as follows: This application is applied to the field-programmable gate array of the task allocation terminal in the GPGPU scheduling control end. The GPGPU scheduling control end further includes a host end. The task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; The method includes: receiving the GPGPU computing task sent by the host end; evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table; using a resource evaluation algorithm to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task; if the target task processing mode is a multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information; transmit each computing subtask to each target GPGPU computing card through the CXL switch, and return the processing results of the computing subtasks output by the target GPGPU computing card to the host end through the CXL switch. Thus, it can be seen that this application is applied to the field-programmable gate array of the task allocation terminal in the GPGPU scheduling control end, rather than the host end. That is to say, when there is a GPGPU computing task, the allocation of the GPGPU computing task and the scheduling of the GPGPU computing card are completed by the field-programmable gate array of the task allocation terminal. Based on the concurrent processing ability of the field-programmable gate array, the load of the central processing unit of the host end can be significantly reduced; Secondly, the task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch, supporting the unlimited expansion of the number of GPGPU computing cards through network cables according to requirements, with quite strong scalability. And based on this connection relationship, the CXL protocol is used to complete data transmission, improving data communication efficiency; Further, a resource evaluation algorithm is used to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity to screen out the target task processing mode of the GPGPU computing task. Therefore, the task execution strategy in the target task processing mode not only has low complexity but also high efficiency; Next, if the target task processing mode is a multi-card pipelining processing mode, then determine each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks. In this way, the target GPGPU computing cards with suitable resource status are used to process the computing subtasks, forming an inter-card pipeline, and multiple cards cooperate to execute the calculation, with high application flexibility and significantly improving the task processing efficiency.
[0115] Furthermore, an embodiment of the present application also provides an electronic device. Figure 8 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the scope of use of the present application.
[0116] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the GPGPU computing task processing method executed by the electronic device disclosed in any of the foregoing embodiments.
[0117] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0118] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computing operations related to machine learning.
[0119] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon include an operating system 221, a computer program 222, data 223, etc. The storage method can be temporary storage or permanent storage.
[0120] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows, Unix, Linux, etc. In addition to the computer program capable of completing the GPGPU computing task processing method executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of completing other specific tasks. The data 223 may include not only the data transmitted by the external device received by the electronic device, but also the data collected by its own input / output interface 25, etc.
[0121] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the GPGPU computing task processing method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0122] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0123] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), registers, hard disks, removable disks, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.
[0124] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0125] The above has introduced in detail a GPGPU computing task processing method, device, equipment and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A method for processing GPGPU computing tasks, characterized in that, A field programmable gate array applied to a task allocation terminal in a GPGPU scheduling control end, the GPGPU scheduling control end further includes a host end, the task allocation terminal and the host end are respectively connected to a GPGPU resource pool through a CXL switch; the method includes: Receiving a GPGPU computing task issued by the host end; Evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from an instruction task comparison table; wherein, the target task execution strategy includes a strategy for splitting the GPGPU computing task into several computing subtasks; Processing the total amount of data to be processed, the target task execution strategy and the target strategy complexity by using a resource evaluation algorithm to screen out the target task processing mode of the GPGPU computing task; If the target task processing mode is a multi-card pipelining processing mode, determining each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool according to the current resource status information of the GPGPU resource pool, and splitting the GPGPU computing task into multiple computing subtasks based on the current resource status information; Transmitting each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and returning the processing results of the computing subtasks output by the target GPGPU computing cards to the host end through the CXL switch; The processing the total amount of data to be processed, the target task execution strategy and the target strategy complexity by using a resource evaluation algorithm to screen out the target task processing mode of the GPGPU computing task includes: Reading the number of splittable computing subtasks of the GPGPU computing task from the instruction task comparison table based on the task description information; obtaining the task execution evaluation factor of the GPGPU computing task; wherein, the task execution evaluation factor includes a strategy complexity reference value corresponding to the target task execution strategy, the total resources of a single GPGPU computing card, and the optical fiber transmission speed between GPGPU computing cards; processing the total amount of data to be processed, the target task execution strategy, the target strategy complexity, the number of splittable computing subtasks of the GPGPU computing task, and the task execution evaluation factor by using a resource evaluation algorithm to obtain a first estimated time for processing the GPGPU computing task in a single-card independent processing mode and a second estimated time for processing the GPGPU computing task in a multi-card pipelining processing mode; screening out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelining processing mode according to the first estimated time and the second estimated time; The screening out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelining processing mode according to the first estimated time and the second estimated time includes: If the first estimated time is less than the second estimated time, determine that the single-card independent processing mode is the target task processing mode; if the first estimated time is not less than the second estimated time, determine that the multi-card pipelining processing mode is the target task processing mode; The expression of the resource evaluation algorithm is: ; ; Among them, is the number of splittable computing subtasks of the GPGPU computing task, is the maximum value of the number of splittable computing subtasks, is the target policy complexity corresponding to the target task execution policy, is the reference value of the policy complexity corresponding to the target task execution policy, is the total amount of data to be processed, is the total resources of a single GPGPU computing card, is the first estimated time for processing the GPGPU computing task in the single-card independent processing mode, is the optical fiber transmission speed between GPGPU computing cards, is the second estimated time for processing the GPGPU computing task in the multi-card pipelining processing mode.
2. The GPGPU computing task processing method according to claim 1, wherein The task allocation terminal further includes a DDR5 and a first network interface, and the host includes a central processing unit, a second network interface, and a cache. Among them, the cache is connected to the DDR5 through a PCIe bus, and the first network interface and the second network interface are respectively connected to the CXL switch through a CAT6e network cable.
3. The GPGPU computing task processing method according to claim 1, wherein The GPGPU resource pool includes multiple GPGPU computing cards. Among them, the GPGPU computing cards with adjacent identification information are connected by optical fibers, and each GPGPU computing card is connected to the CXL switch through a CAT6e network cable; the method further includes: Determine the positioning information of each of the target GPGPU computing cards; Correspondingly, the step of transmitting each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch includes: Transmit each of the computing subtasks to each of the target GPGPU computing cards corresponding to the positioning information through the CXL switch.
4. The GPGPU computing task processing method according to claim 1, wherein The task description information includes the data to be processed, task-related instructions, and task operation data; Correspondingly, the step of evaluating the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and reading the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table includes: Evaluate the data to be processed to obtain the total amount of data to be processed; Determine the storage location of the task-related instructions in the instruction task comparison table according to the hash result of the task-related instructions; where the instruction task comparison table includes each preset task execution strategy and the strategy complexity corresponding to each preset task execution strategy; Read the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table based on the storage location.
5. A GPGPU computing task processing device, characterized in that, Applied to the field programmable gate array of the task allocation terminal in the GPGPU scheduling control end, the GPGPU scheduling control end further includes a host end, and the task allocation terminal and the host end are respectively connected to the GPGPU resource pool through a CXL switch; the device includes: A task receiving module, configured to receive the GPGPU computing task sent by the host end; A strategy reading module, configured to evaluate the task description information in the GPGPU computing task to obtain the total amount of data to be processed, and read the target task execution strategy and the target strategy complexity corresponding to the task description information from the instruction task comparison table; where the target task execution strategy includes a strategy for splitting the GPGPU computing task into several computing subtasks; A mode determination module, configured to process the total amount of data to be processed, the target task execution strategy, and the target strategy complexity by using a resource evaluation algorithm, so as to screen out the target task processing mode of the GPGPU computing task; A task splitting module, configured to, if the target task processing mode is a multi-card pipelining processing mode, determine, according to the current resource status information of the GPGPU resource pool, each target GPGPU computing card corresponding to the target task processing mode from the GPGPU resource pool, and split the GPGPU computing task into multiple computing subtasks based on the current resource status information; A task processing module, configured to transmit each of the computing subtasks to each of the target GPGPU computing cards through the CXL switch, and return the processing result of the computing subtask output by the target GPGPU computing card to the host through the CXL switch; The mode determination module is specifically configured to: Read the number of splittable computing subtasks of the GPGPU computing task from an instruction task comparison table based on the task description information; obtain a task execution evaluation factor of the GPGPU computing task; wherein, the task execution evaluation factor includes a policy complexity reference value corresponding to the target task execution strategy, the total resources of a single GPGPU computing card, and the optical fiber transmission speed between GPGPU computing cards; process the total amount of data to be processed, the target task execution strategy, the target strategy complexity, the number of splittable computing subtasks of the GPGPU computing task, and the task execution evaluation factor by using a resource evaluation algorithm, so as to obtain a first estimated time for processing the GPGPU computing task in a single-card independent processing mode and a second estimated time for processing the GPGPU computing task in a multi-card pipelining processing mode; screen out the target task processing mode of the GPGPU computing task from the single-card independent processing mode and the multi-card pipelining processing mode according to the first estimated time and the second estimated time; The mode determination module is specifically configured to: If the first estimated time is less than the second estimated time, determine that the single-card independent processing mode is the target task processing mode; if the first estimated time is not less than the second estimated time, determine that the multi-card pipelining processing mode is the target task processing mode; The expression of the resource evaluation algorithm is: ; ; Among them, is the number of splittable computing subtasks of the GPGPU computing task, is the maximum value of the number of splittable computing subtasks, is the target policy complexity corresponding to the target task execution policy, is the reference value of the policy complexity corresponding to the target task execution policy, is the total amount of data to be processed, is the total resources of a single GPGPU computing card, is the first estimated time for processing the GPGPU computing task in the single-card independent processing mode, is the optical fiber transmission speed between GPGPU computing cards, is the second estimated time for processing the GPGPU computing task in the multi-card pipelining processing mode.
6. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the steps of the GPGPU computing task processing method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the steps of the GPGPU computing task processing method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Heterogeneous computing resource integration method and device, electronic equipment, storage medium and computer program
CN118519774A
GPU equipment task execution method and device, equipment and storage medium
CN118672789A