Scheduling method and device based on CNN matrix block, equipment and storage medium

CN114546618BActive Publication Date: 2026-09-29FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210168685.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2026-09-29
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

[0004]本申请提供一种基于CNN矩阵分块的调度方法、装置、设备及存储介质,以解决相关技术中针对不同大小的CNN网络层无法灵活改变其分块方案,以达到最优的加速效果,且CNN网络层之间的数据依赖性会造成的FPGA资源浪费等问题

Benefits of technology

[0017]实现CNN不同网络层矩阵的自动分块,使得系统能够根据不同网络层的特性灵活改变分块方案,实现最优的加速效果和最高效的资源利用效率;支持多个CNN计算任务并行执行,有效解决CNN网络层之间的数据依赖性会造成的FPGA资源浪费,同时能够结合CNN矩阵自动分块,高效地使用异构平台计算过程中空闲出来的FPGA资源,以最大化资源利用效率。由此,解决了相关技术中针对不同大小的CNN网络层无法灵活改变其分块方案,以达到最优的加速效果,且CNN网络层之间的数据依赖性会造成的FPGA资源浪费等问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114546618B_ABST
    Figure CN114546618B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer networks, in particular to a scheduling method and device based on CNN matrix blocking, equipment and a storage medium, the method comprising the following steps: receiving at least one computing task of a CNN neural network submitted by a user, determining the priority of each computing task while responding to each computing task according to a preset job priority ordering mode; generating a task publishing request according to the idle resources of one or more FPGA board cards and the receiving condition of the published tasks, responding to the task publishing request according to a preset data dependency relationship and a preset job priority ordering; based on the published computing tasks to be received, performing matrix blocking through a preset FPGA resource scheduling algorithm, allocating corresponding FPGA resources, and deploying the FPGA resources to one or more FPGA board cards to perform parallel computing on the published computing tasks to be received. Therefore, the problems that the optimal acceleration effect cannot be achieved in the related art and FPGA resources are easily wasted are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer networks, in particular to a scheduling method and device based on CNN (Convolutional Neural Networks) matrix blocking, equipment and storage medium. BACKGROUND

[0002] In related technologies, the accelerator architecture of CNN on FPGA (Field Programmable Gate Array) has been very mature. In order to increase the parallelism of calculation, there are various methods for network segmentation of CNN. The input matrix can be blocked from the depth, width and height dimensions. Due to the calculation characteristics of the convolutional layer, blocking the matrix in depth will generate a large amount of data transmission cost, causing waste of communication resources. Therefore, the most commonly used method is to block the width and height of the matrix. After the CNN matrix is blocked, it is generally deployed on a heterogeneous computing platform connected with multiple FPGA board cards.

[0003] However, in related technologies, the block scheme of CNN network layers of different sizes cannot be flexibly changed to achieve optimal acceleration effect. There are network layers in CNN that are not suitable for block calculation. When calculating these network layers, due to the data dependency relationship between network layers, other network layers of the same CNN task cannot be calculated at the same time, resulting in idle FPGA board card resources and resource waste. SUMMARY

[0004] The present application provides a scheduling method, device, equipment and storage medium based on CNN matrix blocking to solve the problems of related technologies, such as the inability to flexibly change the block scheme of CNN network layers of different sizes to achieve optimal acceleration effect, and the waste of FPGA resources caused by the data dependency between CNN network layers.

[0005] The first aspect of the present application provides a scheduling method based on CNN matrix blocking, comprising the following steps: receiving at least one computing job of a convolutional neural network (CNN) neural network submitted by a user, and determining the priority of each computing job while responding to each computing job according to a preset job priority ordering manner; generating a task publishing request according to the idle resources of one or more FPGA board cards and the receiving situation of published tasks, and responding to the task publishing request according to a preset data dependency relationship and the preset job priority ordering; based on the published to-be-received computing tasks, performing matrix blocking through a preset FPGA resource scheduling algorithm, allocating corresponding FPGA resources, and deploying to the one or more FPGA board cards to perform parallel calculation on the published to-be-received computing tasks.

[0006] Further, while responding to each computing job in the preset job priority ordering manner, the priority of each computing job is determined, including: detecting the actual type of each computing job; if the actual type is detected as an emergency type, the computing job of the emergency type is added to an emergency job queue, and the priority of the computing job of the emergency type is determined based on the job submission time; if the actual type is detected as a normal type, the computing job of the normal type is added to a normal job queue, and the priority of the computing job of the normal type is determined based on the job waiting time and the estimated execution time, wherein the priority of the emergency job queue is higher than the priority of the normal job queue.

[0007] Further, the priority of the computing job of the normal type is obtained by the ratio of the job waiting time and the estimated execution time, and the job waiting time is obtained by the difference between the current time and the submission time of the computing job of the normal type.

[0008] Further, the matrix blocking based on the published to-be-received computing task is performed by a preset FPGA resource scheduling algorithm, and the corresponding FPGA resource is allocated, including: judging the network layer type of the current task; if the network layer type is a fully connected layer, a FPGA board card for executing the computing task is allocated for the current task; if the network layer type is a convolution layer, the amount of data to be calculated is determined by the number of multiply-add operations, and the optimal matrix blocking and resource allocation strategy are obtained based on the amount of data to be calculated; if the network layer type is a pooling layer, the total data amount is calculated by the input matrix size, and the optimal matrix blocking and resource allocation strategy are obtained based on the total data amount.

[0009] Further, the allocation priority of the FPGA board card that has reconstructed the circuit of the same network layer type is the highest, and the number of computing tasks to be executed at the same time is less than or equal to two each time.

[0010] The second aspect embodiment of the application provides a scheduling device based on CNN matrix blocking, including: a receiving module configured to receive at least one computing job of a convolutional neural network (CNN) neural network submitted by a user, and determine the priority of each computing job while responding to each computing job in a preset job priority ordering manner; a responding module configured to generate a task publishing request according to the idle resources of one or more FPGA board cards and the receiving situation of published tasks, and respond to the task publishing request according to a preset data dependency relationship and the preset job priority ordering; and a scheduling module configured to perform matrix blocking based on the published to-be-received computing task by a preset FPGA resource scheduling algorithm, allocate the corresponding FPGA resource, and deploy the one or more FPGA board cards to perform parallel computing on the published to-be-received computing task.

[0011] Further, the receiving module is further configured to detect an actual type of each computing job; in response to detecting that the actual type is an urgent type, add the computing job of the urgent type to an urgent job queue, and determine a priority of the computing job of the urgent type based on a job submission time; in response to detecting that the actual type is a normal type, add the computing job of the normal type to a normal job queue, and determine a priority of the computing job of the normal type based on a job waiting time and an estimated execution time, wherein the priority of the urgent job queue is higher than the priority of the normal job queue; wherein the priority of the computing job of the normal type is obtained by a ratio of the job waiting time and the estimated execution time, and the job waiting time is obtained by a difference between a current time and a submission time of the computing job of the normal type.

[0012] Further, the scheduling module is further configured to: determine a network layer type of a current task; if the network layer type is a fully connected layer, allocate a piece of FPGA board card for executing the computing task for the current task; if the network layer type is a convolution layer, determine a data amount to be calculated based on a number of multiply-add operations, and obtain an optimal matrix block and a resource allocation strategy based on the data amount to be calculated; if the network layer type is a pooling layer, calculate a total data amount based on an input matrix size, and obtain the optimal matrix block and the resource allocation strategy based on the total data amount.

[0013] Further, the allocation priority of the FPGA board card having reconstructed the same network layer type circuit is the highest, and the number of computing tasks to be executed at the same time each time is less than or equal to two.

[0014] The third aspect of the embodiments of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor executes the program to implement the scheduling method based on CNN matrix block as described in the above embodiments.

[0015] The fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores computer instructions for causing the computer to execute the scheduling method based on CNN matrix block as described in the above embodiments.

[0016] Therefore, the present application has at least the following beneficial effects:

[0017] The automatic blocking of the matrix of different network layers of the CNN is realized, so that the system can flexibly change the blocking scheme according to the characteristics of different network layers, realize the optimal acceleration effect and the highest efficient resource utilization efficiency; support parallel execution of multiple CNN computing tasks, effectively solve the waste of FPGA resources caused by the data dependency between the CNN network layers, and at the same time, combine the automatic blocking of the CNN matrix, efficiently use the idle FPGA resources in the heterogeneous platform computing process, and maximize the resource utilization efficiency. Therefore, the problems in the related art that the blocking scheme cannot be flexibly changed for CNN network layers of different sizes to achieve the optimal acceleration effect, and the waste of FPGA resources caused by the data dependency between the CNN network layers are solved.

[0018] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0020] Figure 1 A flowchart of a scheduling method based on CNN matrix blocking provided according to an embodiment of the present application;

[0021] Figure 2 A flowchart of a job priority sorting algorithm provided according to an embodiment of the present application;

[0022] Figure 3 A flowchart of a task scheduling algorithm provided according to an embodiment of the present application;

[0023] Figure 4 A flowchart of an FPGA resource scheduling algorithm provided according to an embodiment of the present application;

[0024] Figure 5 A block diagram of a scheduling device based on CNN matrix blocking provided according to an embodiment of the present application;

[0025] Figure 6 A structural diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0027] With reference to the accompanying drawings, the CNN matrix block-based scheduling method, device, equipment and storage medium of the embodiments of the present application are described below. In view of the problem in the related art that the block scheme cannot be flexibly changed for CNN network layers of different sizes to achieve optimal acceleration effect, and that the data dependency between CNN network layers causes waste of FPGA resources, the present application provides a CNN matrix block-based scheduling method. In the method, automatic block of CNN different network layer matrices is implemented, so that the system can flexibly change the block scheme according to the characteristics of different network layers to achieve optimal acceleration effect and highest resource utilization efficiency. Multiple CNN computing tasks are supported to be executed in parallel, effectively solving the problem that the data dependency between CNN network layers causes waste of FPGA resources, and at the same time, the idle FPGA resources in the heterogeneous platform computing process can be efficiently used in combination with the automatic block of CNN matrices to maximize the resource utilization efficiency. Thus, the problem in the related art that the block scheme cannot be flexibly changed for CNN network layers of different sizes to achieve optimal acceleration effect, and that the data dependency between CNN network layers causes waste of FPGA resources is solved.

[0028] Specifically, Figure 1 A flowchart of a CNN matrix block-based scheduling method provided by the embodiments of the present application is shown.

[0029] As Figure 1 shown, the CNN matrix block-based scheduling method includes the following steps:

[0030] In step S101, at least one computing job of a convolutional neural network (CNN) neural network submitted by a user is received, and each computing job is responded simultaneously according to a preset job priority ordering manner, and the priority of each computing job is determined.

[0031] It can be understood that the CNN neural network computing job submitted by the user is received, and then the job is responded and prioritized by an emergency job through a job priority ordering algorithm. The preset job priority ordering manner is the manner corresponding to the job priority ordering algorithm.

[0032] In this embodiment, while responding to each computational job according to a preset job priority sorting method, the priority of each computational job is determined, including: detecting the actual type of each computational job; if the actual type is detected as urgent, the urgent type computational job is added to the urgent job queue, and the priority of the urgent type computational job is determined based on the job submission time; if the actual type is detected as normal, the normal type computational job is added to the normal job queue, and the priority of the normal type computational job is determined based on the job waiting time and the estimated execution time. The priority of the urgent job queue is higher than the priority of the normal job queue. The priority of the normal type computational job is obtained by the ratio of the job waiting time and the estimated execution time, and the job waiting time is the difference between the current time and the submission time of the normal type computational job.

[0033] Specifically, the job priority ranking algorithm in this application supports emergency job response and has a certain degree of real-time performance, such as... Figure 2 As shown, the specific execution process is as follows:

[0034] (1) Receive CNN computation jobs submitted by users.

[0035] (2) First, determine whether it is an urgent task. If it is an urgent task, add it directly to the urgent task queue. The priority of urgent tasks is determined by the submission time. The earlier the submission time, the higher the priority. If it is not an urgent task, define it as a normal task.

[0036] (3) For regular jobs, record the time the job was submitted and estimate its execution time. Since the estimated execution time is only used for job priority ranking, the impact of matrix partitioning is not considered for the time being. Based on the operation type of each network layer (convolutional layer, pooling layer, or fully connected layer) of the job, calculate its computation and communication time on the corresponding circuit type FPGA (the computation and communication speeds of various circuit types of FPGAs have unified test parameters), and then add up the times of all network layers to estimate the execution time of the computation job.

[0037] (4) Add regular jobs to the regular job queue, and calculate the priority of each regular job based on its waiting time and estimated execution time. Sort the regular job queue according to the priority. The job scheduling priority sorting is based on two principles: submitted jobs will not be delayed for a long time; and jobs with shorter estimated execution times will be executed first to reduce the overall waiting time of submitted jobs. The priority calculation formula is as follows: Priority = Job waiting time / Estimated execution time = (Current time - Job submission time) / Estimated execution time.

[0038] In step S102, a task release request is generated based on the idle resources of one or more FPGA boards and the status of received published tasks. The task release request is responded to according to preset data dependencies and preset job priorities.

[0039] It is understood that the embodiments of this application can run a task scheduling algorithm, first generating a task release request based on the availability of idle FPGA board resources and the status of received published tasks, and then responding to the task release request and releasing the network layer computing tasks to be received according to data dependencies and job priorities.

[0040] Specifically, each CNN computation job consists of multiple network layers, and there are data dependencies between layers. Considering this characteristic, the task scheduling algorithm of this invention allows multiple different CNN computation jobs to be executed in parallel, maximizing the utilization efficiency of FPGA resources. Figure 3 As shown, the specific scheduling process is as follows:

[0041] (1) The system continuously queries the FPGA board resource usage. When it detects that there are idle FPGA board resources, it queries whether there are any published but unaccepted computing tasks in the system. If there are, it remains in a waiting state; if not, and the number of pending tasks accepted in the FPGA resource scheduler does not exceed one, it initiates a task publication request.

[0042] (2) When the system receives a task release request, it iterates through the tasks in descending order of priority, starting with urgent tasks and then normal tasks. For each task, it checks whether there is a data dependency between the current network layer task of the current task and the task currently being executed or scheduled by the FPGA (generally, the relationship between the upper and lower network layers of the same CNN task). If there is, it iterates to the next task and repeats the above operation; otherwise, it releases the current network layer task of the current task for the system to receive and removes the task from the task list of the task. If the task is the last task in the task list, it means that all network layer tasks of the task have been processed and the task is removed from the task queue.

[0043] (3) If no computation task is published after traversing all jobs, it means that the current network layer computation task of all jobs has a data dependency relationship with the task being executed or to be executed. Send a task publication waiting signal to the FPGA resource scheduler and do not initiate new task publication requests until a task is completed or a new CNN computation job is received, then start publishing computation tasks.

[0044] Furthermore, the task scheduling method described above enables the parallel execution of network layer computation tasks in multiple CNN jobs. It separates network layer computation tasks with data dependencies, allocating FPGA resources for computation only after the computation of the preceding network layer with data dependencies is completed. This effectively avoids FPGA computation task blocking caused by waiting for input data after the computation task is deployed to the FPGA. When there are idle FPGA resources, the system will continuously issue computation tasks without data dependencies, thereby improving the utilization efficiency of FPGA resources.

[0045] After a network layer computation task is published, the FPGA resource scheduler receives it and then divides the network layer matrix into blocks based on the number of idle FPGA boards, the network layer computation parameters of the current task, and the network layer computation parameters of the next task, with the goal of maximizing the utilization efficiency of the currently idle FPGA resources. The scheduler then fills in the data after network layer division as needed, stores the blocks in blocks, allocates corresponding FPGA resources, generates FPGA configuration files, and finally deploys the task on the FPGA to execute the computation task. The scheduling algorithm in this embodiment splits tasks with data dependencies, so FPGA boards do not need to communicate through a host; the FPGA's communication object is only the host. Theoretically, the communication speed of all FPGA boards is only related to the communication speed between the master node and its corresponding node.

[0046] In step S103, based on the published computing tasks to be received, matrix partitioning is performed using a preset FPGA resource scheduling algorithm, and corresponding FPGA resources are allocated to deploy them on one or more FPGA boards for parallel computing of the published computing tasks to be received.

[0047] Understandably, the FPGA resource scheduler receives published computing tasks to be received, performs matrix partitioning using a custom FPGA resource scheduling algorithm, allocates corresponding FPGA resources, and finally deploys them to multiple FPGA boards for computation.

[0048] In this embodiment, based on the published computing tasks to be received, matrix partitioning is performed using a preset FPGA resource scheduling algorithm, and corresponding FPGA resources are allocated. This includes: determining the network layer type of the current task; if the network layer type is a fully connected layer, then allocating an FPGA board for the current task to execute the computing task; if the network layer type is a convolutional layer, then determining the amount of data to be computed based on the number of multiply-accumulate operations, and obtaining the optimal matrix partitioning and resource allocation strategy based on the amount of data to be computed; if the network layer type is a pooling layer, then calculating the total amount of data based on the size of the input matrix, and obtaining the optimal matrix partitioning and resource allocation strategy based on the total amount of data.

[0049] Specifically, such asFigure 4 As shown, the FPGA resource scheduling algorithm based on matrix block partitioning is described in detail below:

[0050] (1) Check if a task release waiting signal has been received. If received, immediately perform matrix blocking and resource allocation on the received tasks to be executed according to the recorded expected board consumption. If not received, continue to the following steps. The above operation is because if a task release waiting signal is received, it means that no next task will be received in a short time. Therefore, if the tasks to be executed continue to wait, it will only cause more waste of FPGA board resources.

[0051] (2) The FPGA resource scheduler receives the published computing tasks to be received, detects the number of currently idle FPGA boards N, and obtains the network layer-related computing parameters of the computing task. These parameters include the network layer type (convolutional layer, pooling layer, or fully connected layer), the input matrix width W0, height H0, and depth C0, the convolution kernel (or filter for pooling layers) width K0 and height K1, and the output matrix width W1, height H1, and depth C1.

[0052] (2) If there is a task that has been received and is to be executed in the FPGA resource scheduler, then obtain the number of boards that the task is expected to consume N0, and calculate the number of FPGA boards that are currently expected to be available for allocation Nv = N - N0; if there is no task to be executed, then the number of FPGA boards that are currently expected to be available for allocation Nv = N.

[0053] (3) If the network layer type of the current task is a fully connected layer, since its input matrix size is not suitable for block calculation, it may increase the calculation cost. Therefore, the FPGA resource scheduler chooses to directly allocate 1 FPGA board to the current task for calculation, records the number of FPGA boards that the current task is expected to consume N1=1, and directly proceeds to step (6).

[0054] (4) If the current task network layer type is a convolutional layer, the number of multiply-accumulate operations can be used to measure the amount of data Z that needs to be computed. However, if the current task network layer type is a pooling layer, the total amount of data to be computed depends only on the size of the input matrix. The specific formula is as follows:

[0055] Z = K0 × K1 × W1 × H1 × C0 × C1, convolutional layer

[0056] Z = W0 × H0 × C0, pooling layer

[0057] The total amount of communication data in both network layers can be approximated as 2Z.

[0058] Let the number of blocks for the width and height of the matrix be nw and nh, respectively. It is stipulated that the size of the submatrices after matrix partitioning cannot be smaller than the size of the convolution kernel matrix, i.e.

[0059]

[0060] nw and nh must be integers, but W0 and H0 may not be divisible by nw and nh. After block partitioning, the data in the edge submatrix may be less than that in other submatrixes. The system will calculate the specific parameters of the edge submatrix data based on the width and height of the input matrix of the network layer and the matrix partitioning parameters, and process them accordingly, without affecting the normal computation of the task. After matrix partitioning, each submatrix needs to be padded with partial data according to the size of the convolution kernel. The total amount of data M that needs to be padded for all submatrixes can be expressed as:

[0061]

[0062] For FPGA circuits with convolutional and pooling layers, both computation and communication speeds are fixed. Communication speeds are similar, denoted by Vt. Pooling layers are faster than convolutional layers in computation; for simplicity, the following formulas will use Vc for both, but in actual algorithm execution, the computation speeds of convolutional and pooling layers will differ. The total average speed V of processing data for all FPGA resources after matrix partitioning can be expressed as:

[0063]

[0064] Considering the worst-case scenario where no new task is issued until the current task is completed, the data processing speed of all unallocated FPGA board resources during the current task execution can only reach V. Therefore, the size of V reflects the FPGA resource utilization of the matrix partitioning scheme in the worst-case scenario, and we need to maximize V while ensuring sufficient resources.

[0065] (5) Next, iterate through the integer values ​​of nw and nh to find the maximum value of V. Under the condition that the product of the two is less than or equal to N, obtain nw and nh, and calculate the product of the two to obtain the expected number of boards consumed for the task, N1.

[0066] (6) If N1 = Nv, then all tasks (1 or 2) that have been received and are to be executed in the FPGA resource scheduler will be matrix-blocked and resource-allocated according to the recorded expected board consumption and the corresponding matrix block parameters (for convolutional and pooling layers);

[0067] If N1 < Nv, there are two cases: (a) if N1 equals 1, the current network layer is a fully connected layer, or a convolution layer or a pooling layer unsuitable for blocking. Obviously, the best processing method for these types of network layers is no blocking, and they cannot be subject to local optimization of matrix blocking together with the network layer of the next task. Considering maximizing resource utilization efficiency, 1 FPGA board resource is directly allocated to the current task for calculation, but the previously received pending task to be executed still remains in a waiting state. (b) if N1 is not 1, matrix blocking and resource allocation are performed for the previous received pending task to be executed according to the recorded estimated board resource consumption, and the current task waits for the reception of the next task;

[0068] If N1 > Nv, local optimization of matrix blocking is performed on the network layers of the two tasks together to maximize resource utilization. The specific allocation process is as follows: the average speeds of processing data of the two tasks recorded previously obtained from the FPGA resource scheduler are V0 and V1 respectively, and the matrix blocking parameters are nw0, nh0 and nw1, nh1 respectively. Traversing to solve the maximum value of V0+V1 within the integer value ranges of the four parameters nw0, nw1, nh0 and nh1, satisfying that the total number of boards consumed by the two tasks does not exceed the current number of idle boards N, then FPGA resource allocation is performed according to the solved matrix blocking parameters.

[0069] (7) When the FPGA resource scheduler performs resource allocation for tasks, it follows the principle: preferentially allocate FPGA boards that have already been reconfigured with circuits of the same network layer type, which can save part of the FPGA reconfiguration time.

[0070] In this embodiment, the allocation priority of FPGA boards that have been reconfigured with circuits of the same network layer type is the highest, and the number of concurrent computing tasks to be executed each time is less than or equal to two.

[0071] It can be understood that, when performing local optimization solution of resource utilization for matrix blocking in the embodiments of the present application, there are at most two pending tasks to be executed at the same time each time, because traversing the optimal solution by software also consumes a certain amount of time, and with the increase of pending tasks to be executed, the cost of solving the optimal solution will increase exponentially. During the process of the algorithm solving the optimal solution, FPGA board resources have not been allocated for calculation, which will反而 cause greater resource waste. Therefore, arranging two pending tasks to solve the optimal solution not only avoids excessively high solution cost, but also obtains a relatively good result of matrix blocking.

[0072] In summary, the FPGA distributed heterogeneous computing platform scheduling algorithm for automatic CNN matrix partitioning in this application addresses the data dependency between adjacent network layers in CNN computing jobs by splitting the network layer tasks of CNN computing jobs, thereby enabling the parallel execution of multiple CNN computing jobs. At the same time, it achieves automatic matrix partitioning of different network layer tasks, maximizing the utilization of FPGA resources and enabling the FPGA distributed heterogeneous system to efficiently complete the computation of multiple CNN jobs.

[0073] The scheduling method based on CNN matrix partitioning proposed in the embodiments of this application realizes automatic partitioning of the matrices of different CNN network layers, enabling the system to flexibly change the partitioning scheme according to the characteristics of different network layers, achieving the optimal acceleration effect and the most efficient resource utilization; it supports the parallel execution of multiple CNN computing tasks, effectively solving the waste of FPGA resources caused by data dependencies between CNN network layers, and can also combine automatic partitioning of CNN matrices to efficiently use the idle FPGA resources during heterogeneous platform computing, so as to maximize resource utilization efficiency.

[0074] Next, referring to the accompanying drawings, a scheduling device based on CNN matrix block division proposed according to an embodiment of this application is described.

[0075] Figure 5 This is a block diagram of a scheduling device based on CNN matrix block division according to an embodiment of this application.

[0076] like Figure 5 As shown, the scheduling device 10 based on CNN matrix block division includes: a receiving module 100, a response module 200 and a scheduling module 300.

[0077] The receiving module 100 is used to receive at least one computation job of a convolutional neural network (CNN) submitted by the user, and respond to each computation job according to a preset job priority sorting method while determining the priority of each computation job; the response module 200 is used to generate a task release request based on the idle resources of one or more FPGA boards and the status of published task reception, and respond to the task release request according to a preset data dependency relationship and a preset job priority sorting method; the scheduling module 300 is used to perform matrix partitioning based on the published computation tasks to be received, and allocate corresponding FPGA resources through a preset FPGA resource scheduling algorithm, so as to deploy them on one or more FPGA boards to perform parallel computation on the published computation tasks to be received.

[0078] Furthermore, the receiving module 100 is further configured to detect the actual type of each computation job; if the actual type is detected as urgent, the urgent type computation job is added to the urgent job queue, and the priority of the urgent type computation job is determined based on the job submission time; if the actual type is detected as normal, the normal type computation job is added to the normal job queue, and the priority of the normal type computation job is determined based on the job waiting time and the estimated execution time, wherein the priority of the urgent job queue is higher than the priority of the normal job queue; wherein the priority of the normal type computation job is obtained by the ratio of the job waiting time and the estimated execution time, and the job waiting time is the difference between the current time and the submission time of the normal type computation job.

[0079] Furthermore, the scheduling module 300 is further used to: determine the network layer type of the current task; if the network layer type is a fully connected layer, allocate an FPGA board for the current task to perform the computation task; if the network layer type is a convolutional layer, determine the amount of data to be computed based on the number of multiply-accumulate operations, and obtain the optimal matrix partitioning and resource allocation strategy based on the amount of data to be computed; if the network layer type is a pooling layer, calculate the total amount of data based on the size of the input matrix, and obtain the optimal matrix partitioning and resource allocation strategy based on the total amount of data.

[0080] Furthermore, FPGA boards that have been reconfigured with the same network layer type have the highest allocation priority, and the number of computational tasks to be executed simultaneously is less than or equal to two.

[0081] It should be noted that the foregoing explanation of the scheduling method embodiment based on CNN matrix block also applies to the scheduling device based on CNN matrix block in this embodiment, and will not be repeated here.

[0082] The scheduling device based on CNN matrix partitioning proposed in the embodiments of this application realizes automatic partitioning of the matrices of different CNN network layers, enabling the system to flexibly change the partitioning scheme according to the characteristics of different network layers, thereby achieving the optimal acceleration effect and the most efficient resource utilization. It supports the parallel execution of multiple CNN computing tasks, effectively solving the waste of FPGA resources caused by data dependencies between CNN network layers. At the same time, it can combine automatic partitioning of CNN matrices to efficiently use the idle FPGA resources during heterogeneous platform computing, thereby maximizing resource utilization efficiency.

[0083] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:

[0084] The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0085] When the processor 602 executes the program, it implements the scheduling method based on CNN matrix block division provided in the above embodiments.

[0086] Furthermore, electronic devices also include:

[0087] Communication interface 603 is used for communication between memory 601 and processor 602.

[0088] The memory 601 is used to store computer programs that can run on the processor 602.

[0089] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0090] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0091] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0092] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0093] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described scheduling method based on CNN matrix block partitioning.

[0094] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0095] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0096] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0097] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0098] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0099] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0101] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A scheduling method based on CNN matrix block partitioning, characterized in that, Includes the following steps: The system receives at least one computation job of a convolutional neural network (CNN) submitted by a user, and responds to each computation job according to a preset job priority sorting method, while determining the priority of each computation job. Based on the available resources of one or more FPGA boards and the status of published task reception, a task publishing request is generated, and the task publishing request is responded to according to the preset data dependency relationship and the preset job priority. as well as Based on the published computing tasks to be received, matrix partitioning is performed using a preset FPGA resource scheduling algorithm, and corresponding FPGA resources are allocated to be deployed on one or more FPGA boards to perform parallel computing on the published computing tasks to be received. The number of computational tasks to be executed simultaneously is limited to less than or equal to two. When there are two computational tasks to be executed and the total number of FPGA boards required to allocate resources to them individually exceeds the current number of idle boards, joint computation can maximize the sum of the average speeds of the two tasks in processing data, and joint resource allocation is performed accordingly.

2. The method according to claim 1, characterized in that, While responding to each computational job according to a preset job priority sorting method, determining the priority of each computational job includes: Detect the actual type of each computational job; If the actual type is detected to be an emergency type, the computation job of the emergency type is added to the emergency job queue, and the priority of the computation job of the emergency type is determined based on the job submission time. If the actual type is detected as normal, the computation job of the normal type is added to the normal job queue, and the priority of the computation job of the normal type is determined based on the job waiting time and the estimated execution time. The priority of the emergency job queue is higher than the priority of the normal job queue.

3. The method according to claim 2, characterized in that, The priority of the normal type of computation job is obtained by the ratio of the job waiting time and the estimated execution time, and the job waiting time is the difference between the current time and the submission time of the normal type of computation job.

4. The method according to claim 1, characterized in that, The process of allocating corresponding FPGA resources based on the published computing tasks to be received, using a preset FPGA resource scheduling algorithm for matrix partitioning, includes: Determine the network layer type of the current task; If the network layer type is a fully connected layer, then an FPGA board is allocated to the current task to perform the computing task; If the network layer type is a convolutional layer, the amount of data to be computed is determined by the number of multiply-accumulate operations, and the optimal matrix partitioning and resource allocation strategy is obtained based on the amount of data to be computed. If the network layer type is a pooling layer, the total data volume is calculated from the input matrix size, and the optimal matrix partitioning and resource allocation strategy is obtained based on the total data volume.

5. The method according to any one of claims 1-4, characterized in that, FPGA boards that have been reconfigured with the same network layer type have the highest allocation priority.

6. A scheduling device based on CNN matrix block partitioning, characterized in that, include: The receiving module is used to receive at least one computation job of a convolutional neural network (CNN) submitted by the user, and respond to each computation job according to a preset job priority sorting method, while determining the priority of each computation job. The response module is used to generate a task release request based on the idle resources of one or more FPGA boards and the status of received published tasks, and respond to the task release request according to the preset data dependency relationship and the preset job priority. as well as The scheduling module is used to perform matrix partitioning based on the published computing tasks to be received, and allocate corresponding FPGA resources to deploy them on one or more FPGA boards to perform parallel computing on the published computing tasks to be received. The allocation module limits the number of computation tasks to be executed simultaneously to less than or equal to two. When there are two computation tasks to be executed and the total number of FPGA boards required to allocate resources to them individually exceeds the current number of idle boards, joint computation can maximize the sum of the average speeds of the two tasks in processing data using a block-based strategy, and perform joint resource allocation accordingly.

7. The apparatus according to claim 6, characterized in that, The receiving module is further configured to detect the actual type of each computing job; if the actual type is detected to be an emergency type, the computing job of the emergency type is added to the emergency job queue, and the priority of the computing job of the emergency type is determined based on the job submission time; If the actual type is detected as normal, the computational job of that normal type is added to the normal job queue, and the priority of the computational job of that normal type is determined based on the job waiting time and estimated execution time. The priority of the emergency job queue is higher than the priority of the normal job queue. The priority of the normal type of computation job is obtained by the ratio of the job waiting time and the estimated execution time, and the job waiting time is the difference between the current time and the submission time of the normal type of computation job.

8. The apparatus according to claim 6, characterized in that, The scheduling module is further used to: determine the network layer type of the current task; if the network layer type is a fully connected layer, allocate an FPGA board for the current task to perform the computation task; if the network layer type is a convolutional layer, determine the amount of data to be computed based on the number of multiply-accumulate operations, and obtain the optimal matrix partitioning and resource allocation strategy based on the amount of data to be computed. If the network layer type is a pooling layer, the total data volume is calculated from the input matrix size, and the optimal matrix partitioning and resource allocation strategy is obtained based on the total data volume.

9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the scheduling method based on CNN matrix block partitioning as described in any one of claims 1-5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the scheduling method based on CNN matrix block partitioning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • FPGA-based neural network calculator generation method and device

    CN111027688A

  • Service use request processing method, related device and computer program product

    CN113986493A