A GPGPU scheduling method, device, equipment and medium

By introducing a target scheduler in GPGPU, the calculation unit is allocated based on user access rights and preset scheduling algorithms, the delay problem caused by the calculation unit occupation during concurrent calls by multiple users is solved, scheduling efficiency and user experience are improved, and the open source ecosystem is supported.

CN119759593BActive Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510266148.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-23
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In GPGPU, when multiple users call in a transient concurrent manner, the computing unit is occupied by other users, resulting in excessive delay in important user processes, affecting the user experience. Moreover, the traditional GPGPU scheduling module cannot adapt to multiple users and multiple needs, resulting in poor user experience.

Method used

By introducing a target scheduler in the GPGPU, after obtaining the target instructions to be processed, the user access rights are determined based on the identification information of the user terminal, and an appropriate computing unit is assigned to each target instruction using a preset scheduling algorithm, and the instructions are sent to the computing unit through the preset bus for execution.

Benefits of technology

It improves the scheduling efficiency of GPGPU calling computing unit to handle multi-user computing tasks, avoids process delays caused by the computing unit being occupied by other users, improves user experience, and supports open source ecosystems to facilitate researchers to develop and optimize.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759593B_ABST
    Figure CN119759593B_ABST
Patent Text Reader

Abstract

The present application discloses a GPGPU scheduling method, apparatus, device and medium, which relates to the technical field of resource scheduling and is applied to a target scheduler in a GPGPU, including: obtaining a plurality of target instructions to be processed; the target instructions are obtained after a host splits a GPGPU access request sent by a user terminal based on a RISC‑V instruction set; determining a computing unit group for processing the instructions using a user access right determined based on an identification of the user terminal; allocating a corresponding target computing unit in a GPGPU to each instruction using a preset scheduling algorithm based on the total number of instructions, the data volume of each instruction and the computing unit group; sending each target instruction to a corresponding target computing unit through a bus, so that each target computing unit returns a processing result to a target host after executing a corresponding processing operation, and merging each processing result through the target host and returning it to the user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource scheduling, and in particular to a GPGPU scheduling method, device, equipment and medium. Background Art

[0002] GPGPU (General-purpose computing on graphics processing units) is a powerful computing tool. Unlike GPU, GPGPU has a higher degree of parallelism in its computing core and is better at computing some non-graphics-related programs and repetitive tasks with massive data, such as large-scale data encryption, decryption, data computing, and AI computing acceleration.

[0003] With the rapid advancement of current technology, the demand for GPGPU applications has increased significantly. Currently, the application model that is usually supported is that multiple clients share GPGPU. This design supports several users to remotely call a heterogeneous GPGPU to perform calculations in the form of a network, which can reduce the number of GPGPU applications and maximize the efficiency of GPGPU applications without affecting computing requirements, thereby improving economic efficiency. In GPGPU, the computing unit is its computing core, and the user's computing tasks are usually completed by multiple computing units in collaboration. Dozens of computing units are usually deployed inside a single GPGPU to provide computing services for all users.

[0004] However, due to the limited number of computing units in GPGPU, when faced with multiple users instantaneously calling GPGPU concurrently, the important processes of important users are often delayed too much because the computing units are occupied by other users, which seriously affects the user experience. In addition, the traditional GPGPU scheduling module usually adopts a single mechanical CTA (Compute Thread Array) scheduling algorithm, that is, it checks the idle status of the computing unit through polling and completes the scheduling of the task CTA block to the computing unit. However, this method is not suitable for the precise scheduling of concurrent computing tasks, and cannot adapt to the flexible application of GPGPU by multiple users and multiple needs, which also seriously affects the user experience. In addition, the traditional GPGPU scheduling control is often the private instruction set of certain companies, does not support the open source ecosystem, and is not conducive to the application and development of a wider range of researchers.

[0005] In summary, how to improve the scheduling efficiency of GPGPU calling computing units to process multi-user computing tasks and improve user experience is a problem that needs to be solved. Summary of the invention

[0006] In view of this, the purpose of the present invention is to provide a GPGPU scheduling method, device, equipment and medium, which can improve the scheduling efficiency of GPGPU calling computing units to process multi-user computing tasks and improve user experience. The specific scheme is as follows:

[0007] In a first aspect, the present application discloses a GPGPU scheduling method, which is applied to a target scheduler in a GPGPU. The method comprises:

[0008] Acquire several target instructions to be processed; the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set;

[0009] Determine the user access rights based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access rights;

[0010] Based on the total number of the target instructions, the data volume of each target instruction and the computing unit group, using a preset scheduling algorithm to allocate a corresponding target computing unit in the GPGPU to each target instruction;

[0011] Each of the target instructions is sent to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing result to the target host after executing the processing operation corresponding to the target instruction, and the processing results are merged and returned to the user terminal through the target host.

[0012] Optionally, the target scheduler is connected to the security verification unit via a preset internal channel;

[0013] Accordingly, the obtaining of several target instructions to be processed includes:

[0014] After acquiring the target instruction, the security verification unit verifies the legitimacy of the identity of the user terminal;

[0015] If the identity authentication of the user terminal is passed, a plurality of the target instructions to be processed sent by the security verification unit through the preset internal channel are obtained.

[0016] Optionally, a user authority record table is provided in the target scheduler, and the user authority record table is used to record the computing unit information that different users are allowed to schedule, and the computing unit information includes the computing unit identifier and the number of computing units;

[0017] Accordingly, determining the user access rights based on the identification information of the user terminal includes:

[0018] Performing a hash calculation on the IP address of the user terminal to obtain a hash result, and determining the permission storage address corresponding to the user terminal according to the hash result;

[0019] The computing unit information stored in the permission storage address is read from the user permission record table to obtain the user access permission.

[0020] Optionally, after allocating a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group, the method further includes:

[0021] Reading the computing unit information stored in the permission storage address from the user permission record table to obtain the user access rights, and verifying the target computing unit assigned to each target instruction based on the user access rights;

[0022] If the verification is successful, the step of sending each target instruction to the corresponding target computing unit through the preset bus is allowed to be executed.

[0023] Optionally, the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain a target instruction based on the computing thread array, including:

[0024] The target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain a plurality of task execution kernels, and splits each of the task execution kernels into a plurality of target instructions based on a computing thread array based on task complexity;

[0025] Correspondingly, the merging of the processing results by the target host and returning the results to the user terminal includes:

[0026] The processing results corresponding to each of the task execution cores are merged through the target host, and the merged result is returned to the user terminal.

[0027] Optionally, the GPGPU scheduling method further includes:

[0028] Predicting the processing time of each of the target instructions based on the data volume of each of the target instructions;

[0029] Accordingly, the method of allocating a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group includes:

[0030] Determine first identification information of a computing unit currently in a working state, and determine second identification information of each computing unit in the computing unit group, so as to judge whether there is a call conflict in the computing unit group based on the first identification information and the second identification information;

[0031] If there is a call conflict, a three-party evolutionary game model is established based on the computing unit currently in working state, the computing unit group and the computing unit load, and the equilibrium point of the three-party evolutionary game model is calculated with the goal of minimizing the total processing time. The recursive algorithm is called based on the equilibrium point to allocate the corresponding target computing unit in GPGPU to each target instruction.

[0032] Optionally, after determining whether there is a call conflict in the computing unit group based on the first identification information and the second identification information, the method further includes:

[0033] If there is no call conflict, determining whether the number of computing units in the computing unit group is less than the total number of the target instructions;

[0034] If the number of computing units in the computing unit group is not less than the total number of the target instructions, then polling the running status of each computing unit in the computing unit group in turn based on a polling scheduling algorithm to allocate a corresponding target computing unit in the GPGPU to each of the target instructions;

[0035] If the number of computing units in the computing unit group is less than the total number of the target instructions, the total processing time is minimized, and the shortest time first algorithm is called based on the processing time of each target instruction to allocate the corresponding target computing unit in GPGPU to each target instruction.

[0036] In a second aspect, the present application discloses a GPGPU scheduling device, which is applied to a target scheduler in a GPGPU, and the device includes:

[0037] An instruction acquisition module is used to acquire a number of target instructions to be processed; the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set;

[0038] A computing unit determination module, configured to determine user access rights based on the identification information of the user terminal, so as to determine a computing unit group for executing the computing task of the target instruction based on the user access rights;

[0039] A computing unit allocation module, configured to allocate a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group;

[0040] The result acquisition module is used to send each of the target instructions to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing results to the target host after executing the processing operation corresponding to the target instruction, and merges each of the processing results through the target host and returns them to the user terminal.

[0041] In a third aspect, the present application discloses an electronic device, including:

[0042] Memory, used to store computer programs;

[0043] The processor is used to execute the computer program to implement the steps of the aforementioned disclosed GPGPU scheduling method.

[0044] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed GPGPU scheduling method are implemented.

[0045] It can be seen that the present application obtains several target instructions to be processed through the target scheduler in the GPGPU; wherein the target instruction is an instruction based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set; the user access right is determined based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access right; based on the total number of the target instructions, the data amount of each target instruction and the computing unit group, a preset scheduling algorithm is used to allocate a corresponding target computing unit in the GPGPU to each target instruction; each target instruction is sent to the corresponding target computing unit through a preset bus, so that each target computing unit returns the processing result to the target host after executing the processing operation corresponding to the target instruction, and each processing result is merged and returned to the user terminal through the target host.

[0046] Beneficial effect: In the present application, after the target host obtains the GPGPU access request sent by the user terminal, it will split the GPGPU access request based on the RISC-V instruction set to obtain the target instruction based on the computing thread array, and then send the several instructions obtained after the splitting to the target scheduler in the GPGPU for processing. By adopting the RISC-V instruction set control, it can support the open source ecology and support all researchers to participate in development and function optimization. After the target scheduler obtains several target instructions to be processed, it first needs to determine the user access rights based on the identification information of the user terminal to determine the computing unit group for executing the computing task of the target instruction based on the user access rights. That is, the present application sets different user access rights for different user terminals, and the computing unit groups for executing the computing tasks of the target instructions corresponding to different user access rights are also different, so that the computing tasks of different users can only be scheduled to a specific group of computing units, and have no right to call other computing units. Then, when multiple users call GPGPU concurrently, the user's process delay will not be too high due to the computing unit being occupied by other users, thereby improving the user experience. Furthermore, the present application needs to allocate the corresponding target computing unit in GPGPU to each target instruction based on the total number of target instructions, the data volume of each target instruction and the computing unit group, using a preset scheduling algorithm. That is, the present application does not adopt a single polling scheduling algorithm to allocate computing units for each target instruction, but needs to allocate based on the total number of target instructions, the data volume of each target instruction and the computing unit group, so as to adapt to the flexible application of GPGPU by multiple users and multiple needs. Finally, each target instruction is sent to the corresponding target computing unit through a preset bus, so that each target computing unit performs the processing operation corresponding to the target instruction, and returns the processing result to the target host after the processing is completed, and then merges each processing result through the target host and returns it to the user terminal, thus completing a complete GPGPU access request processing flow. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0048] Figure 1 A system architecture diagram applicable to a GPGPU scheduling method disclosed in this application;

[0049] Figure 2 A detailed schematic diagram of the architecture of a GPGPU disclosed in this application;

[0050] Figure 3 A flow chart of a GPGPU scheduling method disclosed in this application;

[0051] Figure 4 A schematic diagram of the architecture of a computing unit scheduling module disclosed in this application;

[0052] Figure 5 A specific GPGPU scheduling method flow chart disclosed in this application;

[0053] Figure 6 A flowchart of optimal allocation of computing units disclosed in this application;

[0054] Figure 7 A scheduling diagram under the shortest time priority algorithm disclosed in this application;

[0055] Figure 8 A schematic diagram of the structure of a GPGPU scheduling device disclosed in this application;

[0056] Fig. 9 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0058] In GPGPU, the computing unit is its computing core, and the user's computing tasks are usually completed by multiple computing units in collaboration. Dozens of computing units are usually deployed inside a single GPGPU to provide computing services for all users. However, due to the limited number of computing units in GPGPU, when faced with multiple users instantaneously and concurrently calling GPGPU, the important processes of important users are often delayed too much because the computing units are occupied by other users, which seriously affects the user experience. In addition, the traditional GPGPU scheduling module usually adopts a single mechanical CTA scheduling algorithm, that is, the idle state of the computing unit is checked by polling, and the task CTA block is scheduled to the computing unit. However, this method is not suitable for the precise scheduling of concurrent computing tasks, and cannot adapt to the flexible application of GPGPU by multiple users and multiple needs, which also seriously affects the user experience. In addition, the traditional GPGPU scheduling control is often a private instruction set of certain companies, does not support the open source ecosystem, and is not conducive to the application and development of a wider range of researchers.

[0059] To this end, the embodiments of the present application disclose a GPGPU scheduling method, apparatus, device and medium, which can improve the scheduling efficiency of GPGPU calling computing units to process multi-user computing tasks and improve user experience.

[0060] The system architecture used in the GPGPU scheduling solution of this application can be found in Figure 1 As shown in , it mainly includes user terminals, host computers and GPGPUs, where multiple remote user terminals are independently connected to the host computer through the network, and heterogeneous GPGPUs communicate and exchange information with the host computer through the PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) interface in the form of DMA (Direct Memory Access). Among them, the user terminal can be a smart device such as a mobile phone or a computer.

[0061] GPGPU is mainly composed of security and scheduling modules, several computing units, on-chip cache and permission table. After receiving the GPGPU access request from the user terminal, the host sends encrypted computing instructions and related data to GPGPU through DMA. GPGPU completes IOPMP (Integrated On-chip PhysicalMemory Protection) security verification for the current GPGPU access request and schedules the computing task to the target computing unit. After the target computing unit completes the processing, it returns the processing result to the host in the form of DMA through the PCIe interface, and then feeds back to the requesting user terminal.

[0062] Furthermore, Figure 2 The detailed architecture of GPGPU is demonstrated, including GPGPU interrupt control module, GPGPU configuration module, security and scheduling module, DMA, GPGPU computing module, corresponding data storage module and AXI (Advanced eXtensible Interface, a high-performance, high-bandwidth on-chip bus) bus connecting various modules.

[0063] Among them, the GPGPU interrupt control module is responsible for processing the interrupt processing request from the host;

[0064] The GPGPU configuration module is responsible for completing the host's configuration of relevant parameters involved in the security and scheduling module, GPGPU computing module, etc.

[0065] The security and scheduling module arranges the IOPMP SoC (System on Chip) level security verification unit, permission table, computing unit scheduling module (i.e., target scheduler below), and necessary Icache (Instruction Cache) and Dcache (Data Cache); this module is responsible for completing the IOPMP security verification for the current access request and scheduling the computing task to the corresponding computing unit of the GPGPU computing module; among them, the IOPMP security verification unit and computing unit scheduling module adopt an integrated design, and the user can directly execute CTA task allocation and scheduling after passing the security verification, without the need for cumbersome and time-consuming system bus transmission;

[0066] DMA deploys a DMA data transceiver engine, which is responsible for data exchange between GPGPU and host memory, and real-time sending and receiving of computing tasks and processing results;

[0067] The GPGPU computing module deploys all GPGPU computing units, is responsible for executing all computing tasks sent by the user terminal, and coordinates and controls all computing units in a unified manner;

[0068] The data storage section includes instruction TCM (Tightly-Coupled Memory) storage and data TCM storage, which are responsible for storing relevant instruction data and computing demand data; data exchange with the host is completed through the AXI bus;

[0069] The above modules are connected through the AXI control bus and the AXI data bus. The AXI bus transmits the instructions and related data required for each module to work. All functional modules are uniformly controlled and deployed by GPGPU.

[0070] See also Figure 3 As shown, the embodiment of the present application discloses a GPGPU scheduling method, which is applied to a target scheduler in a GPGPU. The method includes:

[0071] Step S11: Obtain several target instructions to be processed; wherein the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set.

[0072] In this embodiment, the distributed GPGPU user terminal will send a GPGPU access request to the target host through a wired / wireless network. After obtaining the GPGPU access request sent by the user terminal, the target host will split the GPGPU access request based on the RISC-V instruction set to obtain target instructions based on the computing thread array, and then send the several instructions obtained after the split to the target scheduler in the GPGPU in the form of DMA through the PCIe interface for processing.

[0073] In a specific implementation, the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain a target instruction based on the computing thread array, including: the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain a number of task execution kernels, and splits each of the task execution kernels into a number of target instructions based on the computing thread array based on the task complexity. That is, the target host will split the GPGPU access request into a number of task execution kernels based on the RISC-V instruction set, and split each task execution kernel into a number of CTA (Compute Thread Array) instructions, and the specific number of splits depends on the complexity of the task.

[0074] Furthermore, the target scheduler is connected to the security verification unit through a preset internal channel; accordingly, the acquisition of several target instructions to be processed includes: after the security verification unit obtains the target instruction, the identity of the user terminal is verified to be legitimate; if the identity authentication of the user terminal is passed, the several target instructions to be processed sent by the security verification unit through the preset internal channel are obtained. First of all, it should be pointed out that the traditional GPGPU usually adopts an organizational structure in which the security verification unit and the target scheduler are separated, that is, the instruction for the remote user to access the GPGPU needs to be verified by the security verification unit first, and then transmitted to the target scheduler through the on-chip bus, and then transmitted to the computing unit through the bus, and the bus transmission delay is bound to reduce the communication efficiency of the user task to the core of the GPGPU computing unit. The security verification unit and the target scheduler in the present application adopt an integrated design. After the security verification unit verifies the legitimacy of the identity of the user terminal, it can directly execute CTA task scheduling, that is, the security verification unit can directly send several target instructions to be processed to the target scheduler through the preset internal channel, without the need for cumbersome and time-consuming system bus transmission, and the communication efficiency is significantly improved. The security verification unit mainly verifies the legitimacy of the IP address of the user terminal that sends the current access request. It should also be noted that if the current access request fails the legitimacy verification, the security verification unit directly discards the current access request without sending it to the target scheduler.

[0075] Step S12: determining the user access rights based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access rights.

[0076] In this embodiment, after the target scheduler obtains several target instructions to be processed, it first needs to determine the user access rights based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access rights. That is, the present application sets different user access rights for different user terminals, and the computing unit groups for executing the computing tasks of the target instructions corresponding to different user access rights are also different, so that the computing tasks of different users can only be scheduled to a specific group of computing units, and have no right to call other computing units. Then, when multiple users call GPGPU concurrently, the user's process will not be delayed too much because the computing unit is occupied by other users, thereby improving the user experience. It should be pointed out that the number and grouping of computing units authorized for each user will also be different according to the level of authority. This authority can be set by the host administrator, and some computing units will be reserved to allow multiple users to call together.

[0077] Specifically, a user authority record table is provided in the target scheduler, and the user authority record table is used to record the computing unit information that different users are allowed to schedule, and the computing unit information includes the computing unit identification and the number of computing units; accordingly, determining the user access rights based on the identification information of the user terminal includes: performing a hash calculation on the IP address of the user terminal to obtain a hash result, and determining the authority storage address corresponding to the user terminal according to the hash result; reading the computing unit information stored in the authority storage address from the user authority record table to obtain the user access rights.

[0078] That is, a user authority record table (i.e., authority Table) is also provided in the computing unit scheduling module. The user authority record table is used to record the computing unit information that different users are allowed to schedule. The computing unit information may specifically include, but is not limited to, the computing unit identification and the number of computing units, i.e., which computing units each user can call, and how many computing units. In addition, the process of determining the user access rights based on the identification information of the user terminal is to first perform a hash calculation on the IP address of the user terminal to obtain a hash result, and the value of the hash result represents the authority storage address corresponding to the user terminal, and then read the computing unit information stored in the authority storage address from the user authority record table, i.e., the number of computing units and the corresponding computing unit identification, to obtain the user access rights.

[0079] Step S13: Based on the total number of the target instructions, the data volume of each target instruction and the computing unit group, a preset scheduling algorithm is used to allocate a corresponding target computing unit in the GPGPU to each target instruction.

[0080] In this embodiment, it is necessary to allocate the corresponding target computing unit in GPGPU to each target instruction based on the total number of target instructions, the data volume of each target instruction and the computing unit group using a preset scheduling algorithm. That is, this application does not adopt a single polling scheduling algorithm to allocate computing units to each target instruction, but needs to allocate based on the total number of target instructions, the data volume of each target instruction and the computing unit group to adapt to the flexible application of GPGPU for multiple users and multiple needs.

[0081] Specifically, after allocating the corresponding target computing unit in GPGPU to each target instruction based on the total number of the target instructions, the data volume of each target instruction and the computing unit group using a preset scheduling algorithm, it also includes: reading the computing unit information stored in the permission storage address from the user permission record table to obtain user access rights, and verifying the target computing unit allocated to each target instruction based on the user access rights; if the verification passes, the step of sending each target instruction to the corresponding target computing unit through a preset bus is allowed to be executed.

[0082] It is understandable that the present application may also be provided with a secondary arbitration module in a scheduling mode. After the computing unit allocation is completed, the allocation result may be sent to the secondary arbitration module in the scheduling mode. The secondary arbitration module in the scheduling mode will read the user authority record table again to obtain the user access rights corresponding to the current user terminal, thereby performing secondary verification on the target computing units allocated to each target instruction based on the user access rights, ensuring that the allocation of the target instruction is not repeated or missed, that efficient calculation can be completed, and that no target instruction is allocated to a computing unit without calling authority. Only after the secondary verification is passed, the step of sending each target instruction to the corresponding target computing unit through the preset bus is allowed to be executed.

[0083] Step S14: Send each of the target instructions to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing results to the target host after executing the processing operation corresponding to the target instruction, and merges each of the processing results through the target host and returns them to the user terminal.

[0084] In this embodiment, each target instruction is sent to the corresponding target computing unit through a preset bus (i.e., the AXI bus), so that each target computing unit executes the processing operation corresponding to the target instruction, and after the processing is completed, the processing result is temporarily stored in the consistency cache of the GPGPU computing module, and then the processing result is written back to the target host in the form of DMA through the PCIe interface. The target host then merges the processing results through the network and returns them to the user terminal, thus completing a complete GPGPU access request processing flow.

[0085] In a specific implementation, the said merging the processing results through the target host and returning them to the user terminal includes: merging the processing results corresponding to each of the task execution cores through the target host, and returning the merged result to the user terminal. That is, the host will merge the processing results corresponding to each task execution core, and return the merged result to the user terminal after rechecking.

[0086] For details, see Figure 4 As shown, Figure 4 This is an architectural diagram of a computing unit scheduling module (i.e., target scheduler) disclosed in the present application. The function of this module is to receive all target instructions to be executed and related data to be processed from the security verification unit, and send each target instruction and related data to the target computing unit in the GPGPU through the AXI bus according to the preset scheduling algorithm based on the actual situation.

[0087] The computing unit scheduling module of GPGPU specifically includes management information configuration module, CTA receiving module, user authority reading module, data volume survey module, CTA queue cache module, Smart computing unit optimal allocation algorithm module, computing unit allocation FSM (Finite State Machine) state machine, scheduling mode secondary arbitration module, computing unit record table module, MUX (Multiplexer) data merging module and CTA transmission cache module. Among them, the management information configuration module is responsible for completing the necessary parameter settings and initialization configurations for other modules, and configuring the preset information of the management terminal; the CTA receiving module is responsible for unloading the target instructions and related CTA data based on the RISC-V instruction set from the security verification unit in sequence; the user authority reading module is used to determine the IP address of the task request user terminal to which the current target instruction to be assigned belongs, and then read the current request user access rights from the authority table; the data volume survey module is used to pre-calculate the total amount of data to be processed for each target instruction, and determine the computing amount and processing time; the CTA queue cache module is used to cache the instruction data corresponding to each received target instruction in sequence; the Smart computing unit optimal allocation algorithm module is used to allocate the corresponding computing rights to each input target instruction according to the actual situation through the algorithm. Unit ID; The computing unit allocation FSM state machine is responsible for controlling the allocation of computing units and corresponding target instructions, integrating relevant data, and controlling the changes and jumps of allocation status; The scheduling mode secondary arbitration module is responsible for confirming the computing unit access rights of the current user, and performing secondary confirmation and rationality verification on the computing units allocated to the target instructions; The computing unit record table module is responsible for recording the GPGPU computing unit IDs that are currently busy, and updating the status of each computing unit in real time; The MUX data merging module is responsible for merging and splicing the destination addresses of the computing units allocated to each target instruction and the instruction data cached by the CTA queue cache module according to a specific format based on the CTA block encoding; The CTA launch cache module is responsible for sequentially outputting the target computing unit allocation results and original instruction data of each target instruction to the target computing unit.

[0088] Therefore, the specific application process of the computing unit scheduling module in GPGPU is:

[0089] 1. The management information configuration module completes the necessary parameter settings for the CTA receiving module, user authority reading module, data volume survey module, CTA queue cache module, Smart computing unit optimal allocation algorithm module, computing unit allocation FSM state machine, scheduling mode secondary arbitration module, computing unit record table, CTA transmission cache and other modules;

[0090] 2. The CTA receiving module receives the target instructions and related data to be processed from the security verification unit in the order of the split cores of the access request, and then sends these data to the CTA queue buffer module, the data volume survey module, and the user authority reading module in the order of input;

[0091] 3. The user permission reading module extracts the IP address of the corresponding user terminal from the target instruction, and performs hash function calculation on the user IP address to determine the permission storage address corresponding to the current user terminal, and then reads the user access rights of the current user from the user permission table, that is, which GPGPU computing units the current user can use to perform calculations, and extracts the number of computing units and the corresponding computing unit ID, and sends them to the computing unit allocation FSM state machine;

[0092] 4. The data volume survey module pre-calculates the data volume of each target instruction in sequence, and then predicts the calculation volume and related processing time, and transmits the relevant information to the Smart computing unit optimal allocation algorithm module for subsequent use;

[0093] 5. The computing unit allocation FSM state machine calls the Smart computing unit optimal allocation algorithm module to allocate the target instructions of the current process according to the total number of target instructions of the current user process received, the computing unit ID accessible to the current user (i.e. the aforementioned computing unit group), the data volume of each target instruction and other information. A specific computing unit ID is allocated to each target instruction according to the CTA block encoding and data volume and other related information. Usually, one instruction is allocated to one computing unit. If the available number is insufficient, it is necessary to perform pre-allocation according to the algorithm, that is, the current instruction is executed after the previous instruction is executed in a computing unit.

[0094] 6. After the FSM state machine for computing unit allocation is completed, the corresponding CTA block code and the target computing unit ID are sent to the scheduling mode secondary arbitration module in sequence. The scheduling mode secondary arbitration module will read the user permission table again, read the user access rights of the current user terminal, and perform secondary confirmation on the allocated computing unit to verify the rationality of the allocation; ensure that the allocation is not repeated or missed, and can complete efficient computing, and no computing unit without calling authority is allocated;

[0095] 7. After the secondary verification of the computing unit allocation is completed, the MUX data merging module combines and splices the destination addresses of the computing units allocated to each instruction and the instruction data information cached by the CTA queue cache module based on the CTA block code and in the format preset by the user, and then sends it to the CTA transmission cache module;

[0096] 8. The CTA launch cache module sequentially outputs the target computing unit allocation results and original instruction data of each related instruction to complete the computing unit scheduling task.

[0097] It can be seen that in this application, after the target host obtains the GPGPU access request sent by the user terminal, it will split the GPGPU access request based on the RISC-V instruction set to obtain the target instruction based on the computing thread array, and then send the several instructions obtained after the splitting to the target scheduler in the GPGPU for processing. By adopting the RISC-V instruction set control, it can support the open source ecology and support all researchers to participate in development and function optimization. After the target scheduler obtains several target instructions to be processed, it first needs to determine the user access rights based on the identification information of the user terminal to determine the computing unit group for executing the computing task of the target instruction based on the user access rights. That is, this application sets different user access rights for different user terminals, and the computing unit groups for executing the computing tasks of the target instructions corresponding to different user access rights are also different, so that the computing tasks of different users can only be scheduled to a specific group of computing units, and have no right to call other computing units. Then, when multiple users call GPGPU concurrently, the user's process delay will not be too high due to the computing unit being occupied by other users, thereby improving the user experience. Furthermore, the present application needs to allocate the corresponding target computing unit in GPGPU to each target instruction based on the total number of target instructions, the data volume of each target instruction and the computing unit group, using a preset scheduling algorithm. That is, the present application does not adopt a single polling scheduling algorithm to allocate computing units for each target instruction, but needs to allocate based on the total number of target instructions, the data volume of each target instruction and the computing unit group, so as to adapt to the flexible application of GPGPU by multiple users and multiple needs. Finally, each target instruction is sent to the corresponding target computing unit through a preset bus, so that each target computing unit performs the processing operation corresponding to the target instruction, and returns the processing result to the target host after the processing is completed, and then merges each processing result through the target host and returns it to the user terminal, thus completing a complete GPGPU access request processing flow.

[0098] See also Figure 5 As shown, the embodiment of the present application discloses a specific GPGPU scheduling method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically, it includes:

[0099] Step S21: Obtain several target instructions to be processed; wherein the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set.

[0100] Step S22: determining the user access rights based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access rights.

[0101] Step S23: Determine the first identification information of the computing unit currently in working state, and determine the second identification information of each computing unit in the computing unit group, so as to judge whether there is a call conflict in the computing unit group based on the first identification information and the second identification information.

[0102] In this embodiment, after obtaining the currently callable computing unit group based on the user access rights, the computing unit allocation FSM state machine calls the Smart computing unit optimal allocation algorithm module to allocate the target instructions of the current process. The allocation process is as follows: Figure 6 As shown in . First, it is necessary to determine whether the computing unit group has a call conflict with the computing units called by other user processes. Specifically, first determine the first identification information of the computing unit that is currently in working state. Specifically, the ID of the computing unit currently occupied by other processes can be read from the computing unit record table module to obtain the first identification information, and the ID of each computing unit in the computing unit group can be determined to obtain the second identification information. Then, it is determined whether there is a conflict between the ID of the computing unit currently occupied by other processes and the ID of the computing unit group accessible to the current user, that is, whether the same identification number exists in the first identification information and the second identification information.

[0103] Step S24: If there is a call conflict, a three-party evolutionary game model is established based on the computing unit currently in working state, the computing unit group and the computing unit load, and the equilibrium point of the three-party evolutionary game model is calculated with the goal of minimizing the total processing time, so as to call the recursive algorithm based on the equilibrium point to allocate the corresponding target computing unit in GPGPU for each target instruction.

[0104] In this embodiment, if there is a call conflict, it means that there is a conflict in the computing unit calls between different user processes, that is, some computing units are called by both the previous user process and the current user process. In this case, some computing units need to execute processing tasks for multiple users. Therefore, the order of allocation and execution of CTAs for different users must be considered in all aspects. Therefore, this application establishes a three-party evolutionary game model based on the computing unit currently in working state (that is, representing the process of the previous user), the computing unit group (current user process) and the computing unit load, and takes the shortest total processing time as the goal to calculate the equilibrium point of the three-party evolutionary game model, so as to call the recursive algorithm based on the equilibrium point to allocate the corresponding target computing unit in GPGPU for each target instruction.

[0105] It should be noted that the above method further includes: predicting the processing time of each target instruction based on the data volume of each target instruction. Specifically, the data volume of each target instruction is pre-calculated by a data volume survey module, and then the processing time of the target instruction is predicted.

[0106] Step S25: If there is no call conflict, determine whether the number of computing units in the computing unit group is less than the total number of the target instructions; if the number of computing units in the computing unit group is not less than the total number of the target instructions, poll the operating status of each computing unit in the computing unit group in turn based on the polling scheduling algorithm to allocate the corresponding target computing unit in GPGPU to each target instruction.

[0107] In this embodiment, if there is no conflict in the computing unit calls between different user processes, it is necessary to further determine whether the number of computing units in the computing unit group is less than the total number of target instructions, the purpose of which is to determine whether there are currently enough callable computing units to process multiple target instructions. If the number of computing units in the computing unit group is not less than the total number of target instructions, then a single computing unit only needs to process one target instruction. Therefore, the operating status of each computing unit in the computing unit group is polled in turn based on the polling scheduling algorithm to assign the corresponding target computing unit in the GPGPU to each target instruction. In this case, a load balancing algorithm can also be added for scheduling, that is, instructions are assigned to the computing unit with the lightest load to achieve load balancing between the computing units, so that the computing resources of each computing unit can be fully utilized to avoid the situation where some computing units are overloaded and some computing units are idle, thereby improving the overall performance of the system.

[0108] Step S26: If the number of computing units in the computing unit group is less than the total number of the target instructions, the total processing time is minimized, and the shortest time first algorithm is called based on the processing time of each target instruction to allocate the corresponding target computing unit in GPGPU to each target instruction.

[0109] In this embodiment, if the number of computing units in the computing unit group is less than the total number of target instructions, it means that some computing units need to execute the processing tasks of multiple target instructions. Therefore, in this case, the shortest total processing time is taken as the goal, and the computing amount and estimated processing time of each target instruction are comprehensively considered. The shortest time first algorithm (SJF) is called to allocate multiple instructions to different computing units in the GPGPU, and the order of executing multiple instructions by a single computing unit is recorded, so as to obtain the optimal scheduling solution. The scheduling example is as follows: Figure 7 shown.

[0110] Finally, after the current user process completes scheduling, it is necessary to traverse all target instructions until all target instructions are assigned corresponding GPGPU computing units. Then the Smart computing unit optimal allocation algorithm module sends the CTA block code of the target instruction and the corresponding target computing unit ID, the target instruction assigned to each computing unit ID and the corresponding order to the computing unit allocation FSM state machine in sequence, ending the single Smart computing unit optimal allocation algorithm execution process.

[0111] It can be seen that this application divides the scheduling allocation involved in the application of multi-user heterogeneous GPGPU computing units into the following three situations:

[0112] Case 1: There is a conflict in the computational unit calls between different user processes;

[0113] Case 2: There is no conflict in the computational unit calls between different user processes, but the number of computational units accessible to the current process user is less than the total number of target instructions to be processed;

[0114] Case 3: There is no conflict in the computational unit calls between different user processes, and the number of computational units accessible to the current process user is not less than the total number of target instructions to be processed;

[0115] The differences between the three situations and the specific scheduling allocation schemes are shown in Table 1:

[0116] Table 1 Different scheduling allocation schemes

[0117]

[0118] Step S27: Send each of the target instructions to the corresponding target computing unit through the preset bus, so that each of the target computing units returns the processing results to the target host after executing the processing operation corresponding to the target instruction, and merges each of the processing results through the target host and returns them to the user terminal.

[0119] For more specific processing procedures of the above steps S21, S22 and S27, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be described in detail here.

[0120] It can be seen that this application divides the scheduling allocation involved in the multi-user heterogeneous GPGPU computing unit application into three situations based on whether there is a conflict in the calling of computing units between different user processes and whether the number of computing units accessible to the current process user is less than the total number of target instructions to be processed. Therefore, in the corresponding situation, the corresponding scheduling algorithm is adopted to realize the allocation of computing units in GPGPU and achieve the optimal computing unit scheduling.

[0121] See also Figure 8As shown, the embodiment of the present application discloses a GPGPU scheduling device, which is applied to a target scheduler in a GPGPU, and the device includes:

[0122] The instruction acquisition module 11 is used to acquire a number of target instructions to be processed; the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set;

[0123] A computing unit determination module 12, configured to determine user access rights based on the identification information of the user terminal, so as to determine a computing unit group for executing the computing task of the target instruction based on the user access rights;

[0124] A computing unit allocation module 13 is used to allocate a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group;

[0125] The result acquisition module 14 is used to send each of the target instructions to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing results to the target host after executing the processing operation corresponding to the target instruction, and merges each of the processing results through the target host and returns them to the user terminal.

[0126] It can be seen that in this application, after the target host obtains the GPGPU access request sent by the user terminal, it will split the GPGPU access request based on the RISC-V instruction set to obtain the target instruction based on the computing thread array, and then send the several instructions obtained after the splitting to the target scheduler in the GPGPU for processing. By adopting the RISC-V instruction set control, it can support the open source ecology and support all researchers to participate in development and function optimization. After the target scheduler obtains several target instructions to be processed, it first needs to determine the user access rights based on the identification information of the user terminal to determine the computing unit group for executing the computing task of the target instruction based on the user access rights. That is, this application sets different user access rights for different user terminals, and the computing unit groups for executing the computing tasks of the target instructions corresponding to different user access rights are also different, so that the computing tasks of different users can only be scheduled to a specific group of computing units, and have no right to call other computing units. Then, when multiple users call GPGPU concurrently, the user's process delay will not be too high due to the computing unit being occupied by other users, thereby improving the user experience. Furthermore, the present application needs to allocate the corresponding target computing unit in GPGPU to each target instruction based on the total number of target instructions, the data volume of each target instruction and the computing unit group, using a preset scheduling algorithm. That is, the present application does not adopt a single polling scheduling algorithm to allocate computing units for each target instruction, but needs to allocate based on the total number of target instructions, the data volume of each target instruction and the computing unit group, so as to adapt to the flexible application of GPGPU by multiple users and multiple needs. Finally, each target instruction is sent to the corresponding target computing unit through a preset bus, so that each target computing unit performs the processing operation corresponding to the target instruction, and returns the processing result to the target host after the processing is completed, and then merges each processing result through the target host and returns it to the user terminal, thus completing a complete GPGPU access request processing flow.

[0127] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the GPGPU scheduling method performed by the electronic device disclosed in any of the aforementioned embodiments.

[0128] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0129] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0130] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.

[0131] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, which can be Windows, Unix, Linux, etc. In addition to including a computer program that can be used to complete the GPGPU scheduling method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the electronic device, the data 223 can also include data collected by its own input and output interface 25, etc.

[0132] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the steps of the GPGPU scheduling method disclosed in any of the aforementioned embodiments are implemented.

[0133] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0134] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0135] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.

[0136] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0137] The above is a detailed introduction to a GPGPU scheduling method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A GPGPU scheduling method, characterized in that: The target scheduler applied to GPGPU includes: Acquire several target instructions to be processed; wherein the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set; Determine the user access rights based on the identification information of the user terminal, so as to determine the computing unit group for executing the computing task of the target instruction based on the user access rights; Based on the total number of the target instructions, the data volume of each target instruction and the computing unit group, using a preset scheduling algorithm to allocate a corresponding target computing unit in the GPGPU to each target instruction; Each of the target instructions is sent to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing result to the target host after executing the processing operation corresponding to the target instruction, and the processing results are merged and returned to the user terminal through the target host.

2. The GPGPU scheduling method according to claim 1, characterized in that: The target scheduler is connected to the security verification unit via a preset internal channel; Accordingly, the obtaining of several target instructions to be processed includes: After acquiring the target instruction, the security verification unit verifies the legitimacy of the identity of the user terminal; If the identity authentication of the user terminal is passed, a plurality of the target instructions to be processed sent by the security verification unit through the preset internal channel are obtained.

3. The GPGPU scheduling method according to claim 1, characterized in that: The target scheduler is provided with a user authority record table, which is used to record the computing unit information that different users are allowed to schedule, and the computing unit information includes the computing unit identification and the number of computing units; Accordingly, determining the user access rights based on the identification information of the user terminal includes: Performing a hash calculation on the IP address of the user terminal to obtain a hash result, and determining the permission storage address corresponding to the user terminal according to the hash result; The computing unit information stored in the permission storage address is read from the user permission record table to obtain the user access permission.

4. The GPGPU scheduling method according to claim 3, characterized in that: After allocating a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group, the method further includes: Reading the computing unit information stored in the permission storage address from the user permission record table to obtain the user access rights, and verifying the target computing unit assigned to each target instruction based on the user access rights; If the verification is successful, the step of sending each target instruction to the corresponding target computing unit through the preset bus is allowed to be executed.

5. The GPGPU scheduling method according to claim 1, characterized in that: The target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain target instructions based on the computing thread array, including: The target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set to obtain a plurality of task execution kernels, and splits each of the task execution kernels into a plurality of target instructions based on a computing thread array based on task complexity; Correspondingly, the merging of the processing results by the target host and returning the results to the user terminal includes: The processing results corresponding to each of the task execution cores are merged through the target host, and the merged result is returned to the user terminal.

6. The GPGPU scheduling method according to any one of claims 1 to 5, characterized in that: Also includes: Predicting the processing time of each of the target instructions based on the data volume of each of the target instructions; Accordingly, the method of allocating a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group includes: Determine first identification information of a computing unit currently in a working state, and determine second identification information of each computing unit in the computing unit group, so as to judge whether there is a call conflict in the computing unit group based on the first identification information and the second identification information; If there is a call conflict, a three-party evolutionary game model is established based on the computing unit currently in working state, the computing unit group and the computing unit load, and the equilibrium point of the three-party evolutionary game model is calculated with the goal of minimizing the total processing time. The recursive algorithm is called based on the equilibrium point to allocate the corresponding target computing unit in GPGPU to each target instruction.

7. The GPGPU scheduling method according to claim 6, characterized in that: After determining whether the computing unit group has a call conflict based on the first identification information and the second identification information, the method further includes: If there is no call conflict, determining whether the number of computing units in the computing unit group is less than the total number of the target instructions; If the number of computing units in the computing unit group is not less than the total number of the target instructions, then polling the running status of each computing unit in the computing unit group in turn based on a polling scheduling algorithm to allocate a corresponding target computing unit in the GPGPU to each of the target instructions; If the number of computing units in the computing unit group is less than the total number of the target instructions, the total processing time is minimized, and the shortest time first algorithm is called based on the processing time of each target instruction to allocate the corresponding target computing unit in GPGPU to each target instruction.

8. A GPGPU scheduling device, characterized in that: A target scheduler applied to GPGPU, the device comprising: An instruction acquisition module is used to acquire a number of target instructions to be processed; the target instructions are instructions based on a computing thread array obtained after the target host splits the GPGPU access request sent by the user terminal based on the RISC-V instruction set; A computing unit determination module, configured to determine user access rights based on the identification information of the user terminal, so as to determine a computing unit group for executing the computing task of the target instruction based on the user access rights; A computing unit allocation module, configured to allocate a corresponding target computing unit in the GPGPU to each target instruction using a preset scheduling algorithm based on the total number of the target instructions, the data volume of each target instruction and the computing unit group; The result acquisition module is used to send each of the target instructions to the corresponding target computing unit through a preset bus, so that each of the target computing units returns the processing results to the target host after executing the processing operation corresponding to the target instruction, and merges each of the processing results through the target host and returns them to the user terminal.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the GPGPU scheduling method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the GPGPU scheduling method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • GPGPU instruction execution efficiency optimization method based on lookup table

    CN118819641A

  • RISC-V vector optimization method, device and equipment based on thread scheduling

    CN119357124A