Method, apparatus, electronic device and computer-readable storage medium for generating a computational flow graph scheduling scheme

By grouping and copying the computational flow graphs of the DL model and building the linear programming problem of integers, the problems of low data reuse or low parallelism in the existing technology are solved, and more efficient computing resource utilization and performance improvement are achieved.

CN115437756BActive Publication Date: 2025-06-20STREAM COMPUTING INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110620358.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-03
Publication Date
2025-06-20
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

The prior art has problems with low data reuse or low parallelism in the computing scheduling of DL models, resulting in waste of computing resources and poor performance.

Method used

By grouping the vertices in the original computing flow graph to form a first computing flow graph, and determining the number of computing units required for parallel processing according to the storage resource requirements, the integer linear programming problem is constructed after copying and adding auxiliary vertices to solve the scheduling scheme.

Benefits of technology

It improves the computational parallelism and data reuse rate of the DL model, reduces waste of computing resources, and improves the overall computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115437756B_ABST
    Figure CN115437756B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure discloses a method, an apparatus, an electronic device, and a computer-readable storage medium for generating a computational flow graph scheduling scheme. The method for generating the computational flow graph scheduling scheme includes: grouping the original vertices in the original computational flow graph to obtain a first computational flow graph; determining the number N of computing units required to process a single batch of computational data in parallel; replicating N copies of the first computational flow graph to obtain a second computational flow graph; adding auxiliary vertices to the second computational flow graph to obtain a third computational flow graph; constructing an integer linear programming problem based on the third computational flow graph; and solving the integer linear programming problem to obtain a scheduling scheme for the third computational flow graph. The above method for generating the computational flow graph scheduling scheme solves the technical problems of low data reuse rate or low parallelism in the prior art by converting the original computational flow graph into a third computational flow graph and constructing an integer linear programming problem to solve the scheduling scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computational flow graph scheduling, and particularly to a method, an apparatus, an electronic device, and a computer-readable storage medium for generating a computational flow graph scheduling scheme. Background Art

[0002] A deep learning (DL) model can be represented as a directed acyclic graph (DAG), where the vertices in the graph represent the computational operations in the model, and the directed edges represent the data flow between different computational operations.

[0003] Deploying a DL model on hardware generally can be divided into two scenarios: training the model and inferring the model. In both scenarios, it is necessary to determine the scheduling scheme for executing the DL model. The content of the scheduling scheme includes: the execution order of the vertices in the DAG, the computing devices and the amount of resources used when each vertex is executed, the storage devices and the amount of resources used for the data produced after the vertex is executed, etc.

[0004] In the scenario of training the model, generally, a storage medium with a very high bandwidth such as HBM (High Bandwidth Memory) is used, and the data transmission speed generally does not constitute a performance bottleneck. However, in the scenario of inferring the model, since inference chips generally use storage media with relatively limited bandwidth such as DDR (Double Data Rate SDRAM), the data transmission speed becomes an important factor affecting the inference performance.

[0005] The main directions of the current computational scheduling algorithms for DL models are as follows:

[0006] Vertex fusion: Fuse the vertices with data dependency relationships in the computational graph into one vertex, that is, the output data of the upstream vertex will be directly given to the downstream vertex for consumption after being produced, without caching operations on the storage resources. In this way, the time-consuming data transfer between the original two vertices can be reduced. However, this scheme generally fuses the specified types of computational vertices in the DAG through expert experience, rewrites the DAG according to the fusion result, and arranges the vertex calculation order based on the topological sorting of the rewritten DAG. It is highly dependent on expert experience and is not applicable to all model structures.

[0007] Multi-device allocation: According to the computing and storage characteristics of the vertices, assign them to different computing devices and storage devices for execution to improve the computing utilization rate of each device and reduce the consumption of data transfer between devices. However, this scheme does not change the original calculation order of the vertices in the DAG and cannot improve the computing parallelism during the model execution process.

[0008] Vertex replication: Recalculate vertex results with low computational requirements but high storage requirements to reserve more cache space for other more frequently reused vertex output data, thereby reducing the total time spent on data transfer between low-speed and high-speed caches during the entire DAG execution. This method is equivalent to inserting a replicated vertex in the DAG at another location. However, this solution does not improve the computational parallelism during model execution. Instead, since new vertices are added to the original DAG, the computational consumption of the entire model is increased. Summary of the Invention

[0009] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the Detailed Implementation section later. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0010] To solve the above technical problems in the prior art, the embodiments of the present disclosure propose the following technical solutions:

[0011] In a first aspect, the embodiments of the present disclosure provide a method for generating a computational flow graph scheduling scheme, including:

[0012] A method for generating a computational flow graph scheduling scheme, characterized by including:

[0013] Group the original vertices in the original computational flow graph to obtain a first computational flow graph, where each group serves as a vertex in the first computational flow graph, and the vertex is a set formed by at least one original vertex in the original computational flow graph;

[0014] Determine the number N of computational units required to process a single batch of computational data in parallel according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computational units, where N is an integer greater than or equal to 1;

[0015] Copy N copies of the first computational flow graph to obtain a second computational flow graph;

[0016] Add auxiliary vertices to the second computational flow graph to obtain a third computational flow graph;

[0017] Construct an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph;

[0018] Solve the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph;

[0019] Simplify the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second flow graph.

[0020] Further, the grouping of the original vertices in the original computational flow graph to obtain the first computational flow graph includes:

[0021] Grouping the original vertices in the original computational flow graph according to the input data and output data of the original vertices in the original computational flow graph to obtain the first computational flow graph.

[0022] Further, the calculating the number N of computing units required for parallel processing of a single batch of computing data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units includes:

[0023] Obtaining the maximum storage requirement of the vertices of the first computational flow graph;

[0024] Calculating the number N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units.

[0025] Further, the calculating the number N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units includes:

[0026] Calculating the number N of the computing units according to the following formula:

[0027] where M represents the maximum storage requirement, and m represents the storage space size of a single computing unit.

[0028] Further, the copying the first computational flow graph N times to obtain the second computational flow graph includes:

[0029] Copying the first computational flow graph N times;

[0030] Combining the N first computational flow graphs to generate the second computational flow graph; wherein the second computational flow graph is used for parallel processing of multiple batches of data.

[0031] Further, the auxiliary vertices include: a first auxiliary vertex representing the input data reading operation of the original computational flow graph, a second auxiliary vertex representing the intermediate result calculation operation of the vertices of the original computational flow graph, and a third auxiliary vertex representing the calculation termination operation in the second computational flow graph.

[0032] Further, the constructing an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph includes:

[0033] Finding R t,i , S t,i , L t,i and F t,i such that the value of the following polynomial is minimized:

[0034]

[0035] Among them, i represents the number of vertices in the third computational flow graph, t represents the time step, and R t,i indicates whether to calculate the result of the i-th vertex at the t-th time step; S t,i indicates whether to store the calculation result of the i-th vertex in the cache at the t-th time step; L t,i indicates whether to read the calculation result of the i-th vertex from the cache to the cache of the computing unit at the t-th time step; F t,i indicates whether to release the space occupied by the calculation result of the i-th vertex in the cache of the computing unit at the t-th time step; C i represents the consumption required to transfer the calculation result of the i-th vertex between the cache and the cache of the computing unit; among them, the R t,i = 0 or 1, S t,i = 0 or 1, L t,i = 0 or 1, F t,i = 0 or 1, where 0 indicates not to perform the corresponding operation, and 1 indicates to perform the corresponding operation; T and N are integers greater than 1; among them, the integer linear programming problem further includes the R t,i , S t,i , L t,i and F t,i 's constraint conditions, and the constraint conditions are determined by the hardware performance of the computing unit.

[0036] Furthermore, solving the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph includes:

[0037] Encoding the integer linear programming problem;

[0038] Solving the encoding to obtain the execution order of the vertices in the third computational flow graph.

[0039] Furthermore, simplifying the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second flow graph includes:

[0040] Deleting the auxiliary vertices in the scheduling scheme of the third computational flow graph to obtain the scheduling scheme of the second flow graph.

[0041] Furthermore, the method further includes:

[0042] Determining the amount of data processed by each vertex in the scheduling scheme according to the number of computing units and the number N. In a second aspect, an embodiment of the present disclosure provides a device for generating a computational flow graph scheduling scheme, including:

[0043] The first computational flow graph generation module is used to group the vertices in the original computational flow graph to obtain a first computational flow graph. Each group serves as a vertex in the first computational flow graph, and the vertex is a set formed by at least one original vertex in the original computational flow graph.

[0044] The computing unit quantity determination module is used to determine the number N of computing units required for parallel processing of a single batch of computing data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units. N is an integer greater than or equal to 1.

[0045] The second computational flow graph generation module is used to copy N copies of the first computational flow graph to obtain a second computational flow graph.

[0046] The third computational flow graph generation module is used to add auxiliary vertices to the second computational flow graph to obtain a third computational flow graph.

[0047] The integer linear programming problem construction module is used to construct an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph.

[0048] The integer linear programming problem solving module is used to solve the integer linear programming problem to obtain a scheduling scheme for the third computational flow graph.

[0049] The simplification module is used to simplify the scheduling scheme of the third computational flow graph to form a scheduling scheme for the second flow graph.

[0050] Furthermore, the first computational flow graph generation module is further used to: group the original vertices in the original computational flow graph according to the input data and output data of the original vertices in the original computational flow graph to obtain a first computational flow graph.

[0051] Furthermore, the computing unit quantity determination module is further used to: obtain the maximum storage requirement of the vertices of the first computational flow graph; calculate the number N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units.

[0052] Furthermore, the computing unit quantity determination module is further used to: calculate the number N of computing units according to the following formula: where M represents the maximum storage requirement, and m represents the storage space size of a single computing unit.

[0053] Furthermore, the second computational flow graph generation module is further used to: copy the first computational flow graph N times; combine the N first computational flow graphs to generate a second computational flow graph; wherein the second computational flow graph is used for parallel processing of multiple batches of data.

[0054] Further, the auxiliary vertices include: a first auxiliary vertex representing the input data reading operation of the original computational flow graph, a second auxiliary vertex representing the intermediate result calculation operation of the vertices of the original computational flow graph, and a third auxiliary vertex representing the calculation termination operation in the second computational flow graph.

[0055] Further, the integer linear programming problem construction module is further configured to: obtain the values of R t,i , S t,i , L t,i and F t,i such that the value of the following polynomial is minimized:

[0056]

[0057] where i represents the vertex number in the third computational flow graph, t represents the time step, R t,i represents whether to calculate the result of the i-th vertex at the t-th time step; S t,i represents whether to store the calculation result of the i-th vertex in the cache at the t-th time step; L t,i represents whether to read the calculation result of the i-th vertex from the cache to the cache of the computing unit at the t-th time step; F t,i represents whether to release the space occupied by the calculation result of the i-th vertex in the cache of the computing unit at the t-th time step; C i represents the consumption required to transfer the calculation result of the i-th vertex between the cache and the cache of the computing unit; where the R t,i = 0 or 1, S t,i = 0 or 1, L t,i = 0 or 1, F t,i = 0 or 1, where 0 represents not performing the corresponding operation and 1 represents performing the corresponding operation; T and N are integers greater than 1; where the integer linear programming problem further includes the R t,i , S t,i , L t,i and F t,i constraint conditions, and the constraint conditions are determined by the hardware performance of the computing unit.

[0058] Further, the integer linear programming problem solving module is further configured to: encode the integer linear programming problem; solve the encoding to obtain the execution order of the vertices in the third computational flow graph. Further, the simplification module is further configured to: delete the auxiliary vertices in the scheduling scheme of the third computational flow graph to obtain the scheduling scheme of the second flow graph.

[0059] Further, the generating device of the computing flow graph scheduling scheme is further configured to: determine the amount of data processed by each vertex in the scheduling scheme according to the number of computing units and the number N.

[0060] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions, so that when the processor runs, it implements any one of the methods in the foregoing first aspect.

[0061] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute any one of the methods in the foregoing first aspect.

[0062] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer instructions, when the computer instructions are executed by a computing device, the computing device can execute any one of the methods in the foregoing first aspect.

[0063] An embodiment of the present disclosure discloses a method, device, electronic device, and computer-readable storage medium for generating a computing flow graph scheduling scheme. The method for generating the computing flow graph scheduling scheme includes: grouping vertices in an original computing flow graph to obtain a first computing flow graph, where the vertices in the first computing flow graph are at least one set formed by the vertices in the original computing flow graph; determining the number N of computing units required to parallel-process a single batch of computing data according to the storage resource requirements of the vertices in the first computing flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1; copying N copies of the first computing flow graph to obtain a second computing flow graph; adding auxiliary vertices to the second computing flow graph to obtain a third computing flow graph; constructing an integer linear programming problem corresponding to the third computing flow graph according to the third computing flow graph; solving the integer linear programming problem to obtain a scheduling scheme for the third computing flow graph; and simplifying the scheduling scheme of the third computing flow graph to form a scheduling scheme for the second flow graph. The above method for generating a computing flow graph scheduling scheme solves the technical problems of low data reuse rate or low parallelism in the prior art by converting the original computing flow graph into a third computing flow graph and constructing an integer linear programming problem to solve the scheduling scheme.

[0064] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. In order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given and described in detail in conjunction with the accompanying drawings. Description of the Drawings

[0065] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original components and elements are not necessarily drawn to scale.

[0066] Figure 1 It is a schematic flowchart of a method for generating a computational flow graph scheduling scheme in an embodiment of the present disclosure;

[0067] Figure 2 It is an exemplary schematic diagram of an original computational flow graph in an embodiment of the present disclosure;

[0068] Figure 3 It is a schematic diagram of a first computational flow graph in an embodiment of the present disclosure;

[0069] Figure 4 It is a further schematic flowchart of a method for generating a computational flow graph scheduling scheme in an embodiment of the present disclosure;

[0070] Figure 5 It is an exemplary schematic diagram of a second computational flow graph in an embodiment of the present disclosure;

[0071] Figure 6 It is an exemplary schematic diagram of a third computational flow graph in an embodiment of the present disclosure;

[0072] Figure 7 It is a schematic diagram of the execution order of vertices in a third computational flow graph in an embodiment of the present disclosure;

[0073] Figure 8 It is a schematic diagram of a scheduling scheme for a second flow graph in an embodiment of the present disclosure. Specific Embodiments

[0074] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the scope of protection of the present disclosure.

[0075] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0076] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0077] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0078] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0079] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0080] Figure 1 It is a schematic flowchart of the method for generating a computational flow graph scheduling scheme provided for the embodiments of this disclosure.

[0081] The method for generating the computational flow graph scheduling scheme is used to generate the execution order of vertices in the computational flow graph of the DL model, the computing devices and resource amounts used when each vertex is executed, the storage devices and resource amounts used for the data produced after the vertex is executed, etc.

[0082] As Figure 1 shown, the method includes the following steps:

[0083] Step S101, grouping the original vertices in the original computational flow graph to obtain a first computational flow graph, each group being used as a vertex in the first computational flow graph, and the vertex being a set formed by at least one original vertex in the original computational flow graph.

[0084] As Figure 2 shown, it is an example of the original computational flow graph. In this example, the original computational flow graph includes multiple original vertices, each original vertex respectively representing a kind of computation or operation, such as convolution operation, activation operation, addition operation, pooling operation, etc., and the directed edges between the vertices represent the data flow direction between the vertices.

[0085] In this step, the original vertices in the original computational flow graph are grouped according to certain rules or a preset algorithm. The original vertices assigned to the same group are fused into a fused vertex, and the fused vertex serves as a vertex in the first computational flow graph. Among them, the fused vertex is a set formed by at least one original vertex in the original computational flow graph, that is, the computations / operations represented by the original vertices in the set, as well as the directed edges between the original vertices, serve as the computations / operations and data flows within the vertex in the first computational flow graph.

[0086] Among them, the grouping of the original vertices in the original computational flow graph to obtain the first computational flow graph includes: grouping the original vertices in the original computational flow graph according to the input data and output data of the original vertices in the original computational flow graph to obtain the first computational flow graph. In this embodiment, the original vertices can be grouped according to the dependency relationship of the input data and output data to obtain a vertex of the first computational flow graph.

[0087] The criteria for grouping can also include the computational resource requirements of the original vertices. For example, if the computational resources required by consecutive original vertices are the same, such as the computational resources required by consecutive multiple original vertices are both 2 copies of computational resources (including computing units and storage spaces) or 4 copies of computational resources, then these consecutive multiple original vertices can be grouped into one group to form a vertex of the first computational flow graph.

[0088] The criteria for grouping can also include whether the original vertices can execute computations or operations in parallel. For example, in the original computational flow graph, for the original vertices in two branches after the same original vertex, if the requirements for the computational resources of the original vertices in these two branches do not change much, then the original vertices in the two branches can be grouped into one group to form a vertex of the first computational flow graph.

[0089] Exemplarily, such as Figure 3 The first computational flow graph formed after grouping the original vertices of resnet50 is shown. Among them, the original vertices of the resnet50 network are divided into 4 vertices according to a preset standard, namely group1, group2, group3, and group4. The 4 vertices have data dependencies, that is, the output data of group1 is the input data of group2, the output data of group2 is the input data of group3, the output data of group3 is the input data of group4, and the output data of group4 is the output data of resnet50.

[0090] In the present disclosure, no limitation is imposed on the specific criteria for grouping. Different grouping algorithms corresponding to different grouping criteria can be used to group the original vertices in the original computational flow graph to obtain the vertices in the first computational flow graph.

[0091] Return Figure 1 The method for generating the computational flow graph scheduling scheme further includes: Step S102, determining the number N of computing units required to process a single batch of computational data in parallel according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1.

[0092] Optionally, the storage resource requirements of the vertices in the first computational flow graph include the storage resource requirements of each computational link of the vertices in the first computational flow graph, such as the storage requirements of input data, the storage requirements of intermediate computational results, and the storage requirements of output data.

[0093] Optionally, Step S102 further includes:

[0094] Step S401, obtaining the maximum storage requirement of the vertices of the first computational flow graph;

[0095] Step S402, calculating the number N of computing units required to process a single batch of computational data in parallel according to the maximum storage requirement and the storage resources of the computing units.

[0096] In Step S401, the maximum storage requirement of the vertices of the first computational flow graph is obtained. Optionally, the maximum requirement is the maximum storage requirement among the above-mentioned storage requirements of input data, the storage requirements of intermediate computational results, and the storage requirements of output data.

[0097] Exemplarily, as in the above example of resnet50, the storage requirements of each vertex are shown in the following table:

[0098]

[0099] If the storage resources can meet the maximum storage requirement, then the storage resources can meet the storage requirements of other vertices. Therefore, in this step, the maximum storage requirement of the vertices of the first computational flow graph is first obtained. In the above example, the maximum storage requirement is 3528 KB of the intermediate computational results of vertex Group1.

[0100] Each deep learning model requires hardware resources for graph scheduling of the model. In one example, the applicant's self-developed NPU STCP920 is used for graph scheduling of the resnet50 model. There are 8 computing units on the NPU that can perform efficient matrix multiplication and convolution operations. Each computing unit has an exclusive 1280KB level-1 cache, and the 8 computing units share a sufficiently large level-2 cache. Then, in this example, the storage resource of the computing unit is 1280KB. In step S402, the number N of computing units required to process a single batch of calculation data is calculated according to the maximum storage requirement and the storage resource of the computing unit.

[0101] Optionally, the number N can be determined according to the storage resource required to store the maximum storage requirement. For example, to store 3528KB of data, at least 3 storage resources are required, then the number N of the required computing units can be determined to be 3.

[0102] However, if N = 3, then when the above 8 computing units process data, 2 computing units will be idle. For this reason, optionally, step S102 further includes:

[0103] Calculating the number N of the computing units according to the following formula:

[0104] where M represents the maximum storage requirement and m represents the storage space size of a single computing unit.

[0105] According to the above example, the maximum storage requirement is 3528KB, and the storage resource of a single computing unit is 1280KB. Substituting into the above formula, N = 4 can be calculated. That is, 4 computing units are required to process a batch of data.

[0106] Return Figure 1 , the method for generating the computing flow graph scheduling scheme further includes: step S103, copying N copies of the first computing flow graph to obtain a second computing flow graph.

[0107] In this step, N copies of the first computing flow graph obtained in step S101 are copied. The N copies of the first computing flow graph can process N batches of data in parallel. It can be understood that the first computing flow graph represents the logic for processing data, and when actually executing data processing, the processing performed by each vertex of the first computing flow graph is executed by corresponding hardware such as computing units.

[0108] Optionally, step S103 further includes:

[0109] Copying the first computing flow graph the number N of times;

[0110] Combine the N first computational flow graphs to generate a second computational flow graph; wherein the second computational flow graph is used for parallel processing of multiple batches of data.

[0111] That is, after copying out N first computational flow graphs, combine the N first computational flow graphs to generate a second computational flow graph. Among them, the combination includes using the vertices of the N first computational flow graphs as the vertices of the second computational flow graph, and using the directed edges between the vertices in the N first computational flow graphs as the directed edges of the vertices in the second computational flow graph. For example Figure 5 is an example schematic diagram of the second computational flow graph.

[0112] Return Figure 1 , the method for generating the computational flow graph scheduling scheme further includes: Step S104, adding auxiliary vertices to the second computational flow graph to obtain a third computational flow graph.

[0113] Among them, the auxiliary vertices include: a first auxiliary vertex representing the input data reading operation of the original computational flow graph, a second auxiliary vertex representing the intermediate result calculation operation of the vertices of the original computational flow graph, and a third auxiliary vertex representing the calculation termination operation in the second computational flow graph.

[0114] In the existing scheme, the original computational flow graph output by the deep learning model in the way of taking computational operations as vertices only focuses on the computational process of the input data in the model, while ignoring the influence of the parameter data of the model itself on the computational and storage requirements during the model execution process. In this step, adding auxiliary vertices to the second computational flow graph supplements the life cycle during the model execution process for the original computational flow graph (such as the first auxiliary vertex represents the start of model calculation, the second auxiliary vertex represents the intermediate execution process of the model, and the third auxiliary vertex represents the termination of model calculation), storage space occupancy and other model parameter data information, providing more complete information for the subsequent design of the model scheduling scheme. These information can help the designer more conveniently produce the model scheduling scheme and reduce the workload of analyzing the feasibility of different model scheduling schemes and comparing the performance advantages and disadvantages.

[0115] For example Figure 6Shown is an exemplary schematic diagram of a third computational flow graph. Among them, vertices other than those in the second computational flow graph are auxiliary vertices. For example, batch1 input and group1 weight are the first auxiliary vertices representing the input data reading operations of the original computational flow graph, where batch1 input represents the input data reading of the first sample data, and group1 weight represents the reading of the weight data of the model. Other first auxiliary vertices follow the same pattern and will not be elaborated further. Among them, batch1group1 internal represents the intermediate result calculation operation of the first sample data in the group1 vertex, which is the second auxiliary vertex. Other second auxiliary vertices follow the same pattern and will not be elaborated further. Among them, termination represents the third auxiliary vertex of the calculation termination operation in the second computational flow graph.

[0116] Return Figure 1 , the method for generating the computational flow graph scheduling scheme further includes: Step S105, constructing an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph.

[0117] An integer programming problem is an optimization problem where both the objective function and the constraints in the integer linear programming problem are linear. An integer linear programming problem is a linear programming problem (integer linear programming, ILP) that requires all unknowns to be integers. That is, by constructing the objective function with the consumption in the calculation process as the unknowns and using the performance of the hardware resources as the constraints, an integer linear programming problem can be constructed, and the solution to this integer linear programming problem is the scheduling scheme.

[0118] Exemplarily, the step S105 includes:

[0119] Find the values of R t,i , S t,i , L t,i and F t,i such that the value of the following polynomial is minimized:

[0120]

[0121] where i represents the vertex number in the third computational flow graph, t represents the time step, R t,i represents whether to calculate the result of the i-th vertex at the t-th time step; S t,i represents whether to store the calculation result of the i-th vertex in the low-speed cache at the t-th time step; L t,i represents whether to read the calculation result of the i-th vertex from the low-speed cache to the cache of the calculation unit at the t-th time step; F t,iIndicates whether to release the space occupied by the calculation result of the \(i\)-th vertex in the cache of the computing unit at the \(t\)-th time step; \(C\) i Indicates the consumption required to transfer the calculation result of the \(i\)-th vertex between the cache memory and the cache of the computing unit; where the \(R\) t,i = 0 or 1, \(S\) t,i = 0 or 1, \(L\) t,i = 0 or 1, \(F\) t,i = 0 or 1, where 0 indicates not to perform the corresponding operation, and 1 indicates to perform the corresponding operation; \(T\) and \(N\) are integers greater than 1; where the integer linear programming problem also includes the \(R\) t,i , \(S\) t,i , \(L\) t,i and \(F\) t,i The constraint conditions of, and the constraint conditions are determined by the hardware performance of the computing unit.

[0122] Taking the above example as an example, according to the structure of the original calculation flow graph and the hardware characteristics of the NPU STCP920, the \(R\) t,i , \(S\) t,i , \(L\) t,i and \(F\) t,i The constraint conditions are as follows:

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] The method for constructing the integer linear programming problem described above uses the method for constructing a binary integer programming problem. In practical applications, other methods for constructing integer programming problems can also be used. Such as the method for constructing a multi-variable integer programming problem, such as using \(R\) i , \(S\) i , \(L\) i , \(F\) ito represent the time step for performing the corresponding operation on vertex i, i.e., R i , S i , L i , F i ∈{0, 1, …, T}^4, R i = t means performing a computational operation on vertex i at time step t, R i = 0 means not performing a computational operation on vertex i in the scheduling scheme, and the definitions of other operations are similar. Thus, a non-binary integer linear programming problem can be constructed correspondingly, which will not be elaborated here.

[0134] The model scheduling problem is represented by the above mathematical formula, and a clear optimization goal is set, enabling the designer to design and optimize the scheduling scheme using mathematical methods in optimization theory.

[0135] Return Figure 1 , the method for generating the computational flow graph scheduling scheme further includes: step S106, solving the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph.

[0136] After obtaining the objective function (such as the above formula 1) and the constraints, the solution of the objective function under these constraints can be calculated. Step S106 is the process of solving the integer linear programming problem to obtain the solution that minimizes the objective function, which is the scheduling scheme of the third computational flow graph.

[0137] Optionally, step S106 includes:

[0138] Encoding the integer linear programming problem;

[0139] Solving the encoding to obtain the execution order of the vertices in the third computational flow graph.

[0140] That is, encoding the objective function and constraints constructed in step S105, and then obtaining the execution order of the vertices in the third computational flow graph by running the encoding for solution.

[0141] Optionally, existing toolkits can be used to solve the integer linear programming problem. For example, the python extension package pulp is used to encode and solve the above problem. Pulp is a python extension package developed for linear programming problems, which provides a language specification for describing linear programming problems and encapsulates interfaces that can be called through python for various linear programming problem solvers. It can be understood that the constructed integer linear programming problem can also be encoded in any other programming language. Any software with the ability to solve linear programming problems can be used to solve the encoded integer linear programming problem, which will not be elaborated here.

[0142] As shown Figure 7 in the schematic diagram of the execution order of the vertices in the third computational flow graph obtained by solving the integer linear programming problem. Among them, since the first auxiliary vertex only represents the input of data, the execution order of the first auxiliary vertex is related to its corresponding vertex, that is, the calculation will only start after reading the data, and its execution order has no impact on the execution order of other vertices in the third computational flow graph, and it is only used to construct the above integer linear programming problem. Therefore, the first auxiliary vertex is not shown.

[0143] Return Figure 1 , the method for generating the computational flow graph scheduling scheme further includes: step S107, simplifying the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second flow graph.

[0144] Since the third computational flow graph includes auxiliary vertices, which only provide complete information for the scheduling scheme. Therefore, in actual use, these vertices can be removed. Therefore, the step S107 includes:

[0145] Deleting the auxiliary vertices in the scheduling scheme of the third computational flow graph to obtain the scheduling scheme of the second flow graph. As shown Figure 8 in the scheduling scheme of the second flow graph obtained by simplifying the scheduling scheme of the third computational flow graph. The scheduling scheme of the second flow graph is the final scheduling scheme.

[0146] In order for the hardware device actually used to exert the maximum performance. The method for generating the computational flow graph scheduling scheme further includes:

[0147] Determining the amount of data processed by each vertex in the scheduling scheme according to the number of computing units and the number N.

[0148] In the above example, the scheduling scheme of the second computational flow graph is to use 4 processing units to process 4 batches of data, but the NPU contains 8 computing units. Therefore, in order to exert the maximum computing power, the amount of data processed by each vertex in the scheduling scheme can be doubled. It can be understood that each vertex represents the logic of processing data, and the actual data processing is performed by the computing unit corresponding to the vertex.

[0149] The above embodiments disclose a method for generating a computational flow graph scheduling scheme. The method for generating the computational flow graph scheduling scheme includes: grouping vertices in an original computational flow graph to obtain a first computational flow graph, where the vertices in the first computational flow graph are at least one set formed by the vertices in the original computational flow graph; determining the number N of computing units required to process a single batch of computational data in parallel according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1; replicating N copies of the first computational flow graph to obtain a second computational flow graph; adding auxiliary vertices to the second computational flow graph to obtain a third computational flow graph; constructing an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph; solving the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph; and simplifying the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second flow graph. The method for generating the computational flow graph scheduling scheme solves the technical problems of low data reuse rate or low parallelism in the prior art by converting the original computational flow graph into a third computational flow graph and constructing an integer linear programming problem to solve the scheduling scheme.

[0150] As can be seen from the above embodiments: The design of traditional model scheduling schemes requires designers to have rich experience in DL model optimization and be able to deeply understand the structural characteristics of the DL model to be scheduled currently. However, the use of the automated algorithm proposed in this disclosure can avoid this dependence on expert experience. When manually designing a model scheduling scheme, designers need to spend a lot of time validating different scheduling schemes and comparing their effects, which requires a large amount of time and manpower. However, the use of the automated algorithm proposed in this disclosure can produce a scheduling scheme for the DL model in a short time, greatly saving manpower and time costs. As mentioned above, for DL models with different structures, due to the limitations of designers' own experience and the different levels of understanding of model characteristics, different designers may produce model scheduling schemes with different performances, and it is difficult to prove whether the scheduling schemes they produce are optimal. The use of the automated method proposed in this disclosure can stably produce a globally optimal scheduling scheme for DL models with different structures.

[0151] In addition, the design of traditional computational flow graph scheduling schemes only focuses on the scheduling optimization problem of a single batch of data running once on the computational flow graph, and cannot solve the problem of waste of computing resources caused by too low computational parallelism of a certain vertex in the computational flow graph when the data volume is small. The method proposed in the present invention converts the problem into a scheduling optimization problem of multiple batches of data running multiple times on the computational flow graph through replication and merging of the computational flow graph, expands the boundary of the set of feasible scheduling schemes, and thus can search for a scheduling scheme in which all vertices have a relatively high computational parallelism in a larger scheduling scheme space, reducing waste of computing resources, and thereby improving the overall performance of the execution of the computational flow graph.

[0152] An embodiment of the present disclosure provides a generating device for a computational flow graph scheduling scheme, including:

[0153] A first computational flow graph generating module, configured to group the original vertices in the original computational flow graph to obtain a first computational flow graph, where each group serves as a vertex in the first computational flow graph, and the vertex is a set formed by at least one original vertex in the original computational flow graph;

[0154] A computing unit quantity determining module, configured to determine the quantity N of computing units required for parallel processing of a single batch of computing data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1;

[0155] A second computational flow graph generating module, configured to copy N copies of the first computational flow graph to obtain a second computational flow graph;

[0156] A third computational flow graph generating module, configured to add auxiliary vertices to the second computational flow graph to obtain a third computational flow graph;

[0157] An integer linear programming problem constructing module, configured to construct an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph;

[0158] An integer linear programming problem solving module, configured to solve the integer linear programming problem to obtain a scheduling scheme for the third computational flow graph;

[0159] A simplifying module, configured to simplify the scheduling scheme of the third computational flow graph to form a scheduling scheme for the second flow graph.

[0160] Further, the first computational flow graph generating module is further configured to: group the original vertices in the original computational flow graph according to the input data and output data of the original vertices in the original computational flow graph to obtain a first computational flow graph.

[0161] Further, the computing unit quantity determining module is further configured to: obtain the maximum storage requirement of the vertices of the first computational flow graph; calculate the quantity N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units.

[0162] Further, the computing unit quantity determining module is further configured to: calculate the quantity N of the computing units according to the following formula: where M represents the maximum storage requirement, and m represents the storage space size of a single computing unit.

[0163] Further, the second computational flow graph generation module is further configured to: copy the first computational flow graph N times with respect to the data; combine the N first computational flow graphs to generate a second computational flow graph; wherein the second computational flow graph is used for parallel processing of multiple batches of data.

[0164] Further, the auxiliary vertices include: a first auxiliary vertex representing the input data reading operation of the original computational flow graph, a second auxiliary vertex representing the intermediate result calculation operation of the vertices of the original computational flow graph, and a third auxiliary vertex representing the calculation termination operation in the second computational flow graph.

[0165] Further, the integer linear programming problem construction module is further configured to: obtain the values of R t,i , S t,i , L t,i and F t,i such that the value of the following polynomial is minimized:

[0166]

[0167] wherein, i represents the number of the vertex in the third computational flow graph, t represents the time step, R t,i represents whether to calculate the result of the i-th vertex at the t-th time step; S t,i represents whether to store the calculation result of the i-th vertex in the low-speed cache at the t-th time step; L t,i represents whether to read the calculation result of the i-th vertex from the low-speed cache to the cache of the computing unit at the t-th time step; F t,i represents whether to release the space occupied by the calculation result of the i-th vertex in the cache of the computing unit at the t-th time step; C i represents the consumption required to transfer the calculation result of the i-th vertex between the low-speed cache and the cache of the computing unit; wherein, the R t,i = 0 or 1, S t,i = 0 or 1, L t,i = 0 or 1, F t,i = 0 or 1, where 0 represents not performing the corresponding operation, and 1 represents performing the corresponding operation; T and N are integers greater than 1; wherein, the integer linear programming problem further includes the R t,i , S t,i , L t,i and F t,i constraint conditions, and the constraint conditions are determined by the hardware performance of the computing unit.

[0168] Further, the integer linear programming problem solving module is further configured to: encode the integer linear programming problem; solve the encoding to obtain the execution order of the vertices in the third computational flow graph. Further, the simplification module is further configured to: delete the auxiliary vertices in the scheduling scheme of the third computational flow graph to obtain the scheduling scheme of the second flow graph.

[0169] Further, the computational flow graph scheduling scheme generating device is further configured to: determine the amount of data processed by each vertex in the scheduling scheme according to the number of computing units and the number N.

[0170] An embodiment of the present disclosure further provides an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions such that when the processor runs, it implements the computational flow graph scheduling scheme generating method in any one of the foregoing embodiments.

[0171] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the computational flow graph scheduling scheme generating method in any one of the foregoing embodiments.

[0172] An embodiment of the present disclosure further provides a computer program product, including computer instructions, when the computer instructions are executed by a computing device, the computing device can execute the computational flow graph scheduling scheme generating method in any one of the foregoing embodiments.

[0173] The flowcharts and block diagrams in the accompanying drawings of the present disclosure illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0174] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation to the unit itself.

[0175] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0176] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

Claims

1. A method for generating a computational flow graph scheduling scheme, characterized in that, Including: Dividing the original vertices in the original computational flow graph into multiple groups to obtain a first computational flow graph, where each group corresponds to a vertex in the first computational flow graph, and the vertex is a set formed by at least one original vertex in the original computational flow graph; Determining the number N of computing units required for parallel processing of a single batch of computing data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1; Copying N copies of the first computational flow graph to obtain a second computational flow graph; Adding auxiliary vertices to the second computational flow graph to obtain a third computational flow graph; Constructing an integer linear programming problem corresponding to the third computational flow graph, where the unknowns in the objective function of the integer linear programming problem are the consumptions in the computing process, the constraint conditions corresponding to the integer linear programming problem are the performance of the hardware resources of the computing units, and the unknowns are determined according to the execution order of the vertices in the third computational flow graph, the computing units used by each vertex during execution, and the consumed storage resources; Solving the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph; Simplifying the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second computational flow graph.

2. The method for generating a computational flow graph scheduling scheme according to claim 1, characterized in that, The dividing the original vertices in the original computational flow graph into multiple groups to obtain a first computational flow graph includes: Grouping the original vertices in the original computational flow graph according to the input data and output data of the original vertices in the original computational flow graph to obtain a first computational flow graph.

3. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The calculating the number N of computing units required for parallel processing of a single batch of computing data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units includes: Obtaining the maximum storage requirement of the vertices of the first computational flow graph; Calculating the number N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units.

4. The method for generating a computational flow graph scheduling scheme according to claim 3, characterized in that, The calculating the number N of computing units required for parallel processing of a single batch of computing data according to the maximum storage requirement and the storage resources of the computing units includes: Calculate the number N of the said calculation units according to the following formula: Where M represents the maximum storage requirement, and m represents the storage space size of a single computing unit.

5. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The copying N copies of the first computational flow graph to obtain a second computational flow graph includes: Copying the first computational flow graph N times; Combining the N first computational flow graphs to generate a second computational flow graph; where the second computational flow graph is used for parallel processing of multiple batches of data.

6. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The auxiliary vertices include: A first auxiliary vertex representing the input data reading operation of the original computational flow graph, a second auxiliary vertex representing the intermediate result calculation operation of the vertices of the original computational flow graph, and a third auxiliary vertex representing the computing termination operation in the second computational flow graph.

7. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The constructing an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph includes: Find the value of S t,i and L t,i such that the value of the following polynomial is minimized: where, i represents the vertex number in the third computational flow graph, t represents the time step, and S t,i represents whether to store the calculation result of the i-th vertex in the cache at the t-th time step; L t,i represents whether to read the calculation result of the i-th vertex from the cache to the cache of the computing unit at the t-th time step; C i represents the consumption required to transfer the calculation result of the i-th vertex between the cache and the cache of the computing unit; T and N are integers greater than 1; Among them, the integer linear programming problem further includes the S t,i , the L t,i , R t,i , and F t,i constraints, and the constraints are determined by the hardware performance of the computing unit. R t,i represents whether to calculate the result of the i-th vertex at the t-th time step; F t,i represents whether to release the space occupied by the calculation result of the i-th vertex in the cache of the computing unit at the t-th time step; among them, the R t,i = 0 or 1, S t,i = 0 or 1, L t,i = 0 or 1, F t,i = 0 or 1, where 0 indicates not to perform the corresponding operation, and 1 indicates to perform the corresponding operation.

8. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The solving the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph includes: Encoding the integer linear programming problem; Solve the encoding to obtain the execution order of the vertices in the third computational flow graph.

9. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The method of simplifying the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second computational flow graph includes: Delete the auxiliary vertices in the scheduling scheme of the third computational flow graph to obtain the scheduling scheme of the second computational flow graph.

10. The method for generating a computational flow graph scheduling scheme according to any one of claims 1 or 2, characterized in that, The method further includes: Determine the amount of data processed by each vertex in the scheduling scheme according to the number of computing units and the number N.

11. A device for generating a computational flow graph scheduling scheme, characterized in that, It includes: A first computational flow graph generation module, configured to divide the original vertices in the original computational flow graph into multiple groups to obtain a first computational flow graph, where each group corresponds to a vertex in the first computational flow graph, and the vertex is a set formed by at least one vertex in the original computational flow graph; A computing unit number determination module, configured to determine the number N of computing units required for parallel processing of a single batch of computational data according to the storage resource requirements of the vertices in the first computational flow graph and the storage resources of the computing units, where N is an integer greater than or equal to 1; A second computational flow graph generation module, configured to copy N copies of the first computational flow graph to obtain a second computational flow graph; A third computational flow graph generation module, configured to add auxiliary vertices to the second computational flow graph to obtain a third computational flow graph; An integer linear programming problem construction module, configured to construct an integer linear programming problem corresponding to the third computational flow graph according to the third computational flow graph, where the unknowns in the objective function of the integer linear programming problem are the consumptions in the computing process, and the constraint conditions corresponding to the integer linear programming problem are the performance of the hardware resources of the computing units, and the unknowns are determined according to the execution order of the vertices in the third computational flow graph, the computing units used when each vertex executes, and the storage resources consumed; An integer linear programming problem solving module, configured to solve the integer linear programming problem to obtain the scheduling scheme of the third computational flow graph; A simplification module, configured to simplify the scheduling scheme of the third computational flow graph to form the scheduling scheme of the second computational flow graph.

Citation Information

Patent Citations

  • Computational flow graph construction method and device and storage medium

    CN109960751A

  • GPU processing method and device for data calculation flow graph

    CN110163791A