A method and apparatus for parallel processing of deep learning models
By automatically partitioning deep learning models based on dependency relationships and distributing them across devices, the method improves the efficiency of model parallel training, addressing the inefficiencies of manual partitioning and communication overhead.
Patent Information
- Application Number
- CN201910916367.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-09-26
AI Technical Summary
In the prior art, the parallel training efficiency of deep learning models is low, manual splitting is time-consuming and labor-intensive, and communication overhead is large, resulting in a long training time and cannot meet the parameter adjustment requirements.
By automatically determining the dependencies between computing nodes in the model, dividing relationship groups, and clustering according to predetermined rules to generate parallel execution sets, the relationship groups are allocated to multiple target devices using a simulated annealing algorithm to minimize the total parallel operation time.
It improves the distributed training efficiency of deep learning models when using models in parallel, reduces training time, and improves parameter adjustment efficiency.
Smart Images

Figure CN112561051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and device for parallel processing of deep learning models. Background Art
[0002] Deep learning models have a large number of parameters and a large scale of training data, resulting in high consumption of computing resources. The time taken for a single training often reaches several days or even months, which is simply unbearable for the staff adjusting the parameters. Therefore, it is very necessary to accelerate model training. However, the improvement of the computing power of a single device is very limited, so distributed training is required.
[0003] Currently, there are mainly two ways of distributed training for deep learning models: data parallelism and model parallelism. Data parallelism means that there is a copy of a complete model on each node, which respectively uses different data, completes the forward and backward calculations to obtain gradients, and then updates the parameters. Model parallelism means that the model is split onto different nodes for training according to certain rules.
[0004] In the related art, when performing model parallelism, the model splitting is usually manually completed. Manual splitting is time-consuming and laborious. If the splitting is unreasonable, coupled with the communication overhead between nodes, model parallelism may not even achieve any acceleration effect. Summary of the Invention
[0005] This article provides a method and device for parallel processing of deep learning models, which can automatically split deep learning models and improve the distributed training efficiency of deep learning models when using model parallelism.
[0006] According to the first aspect of the present application, an embodiment of the present invention provides a method for parallel processing of a deep learning model, including:
[0007] Determine the dependency relationship between computing nodes in the model, and divide the relationship groups according to the dependency relationship;
[0008] Cluster the relationship groups according to a predetermined rule to generate a set of parallel-executable sets; wherein, the relationship groups within each set of parallel-executable sets can run in parallel;
[0009] Allocate the relationship groups within all the sets of parallel-executable sets to multiple target devices so that the total parallel operation time of all the sets of parallel-executable sets is the shortest.
[0010] According to the second aspect of the present application, an embodiment of the present invention provides a device for parallel processing of a deep learning model, including:
[0011] A memory, a processor, and a program for parallel processing of a deep learning model stored on the memory and executable on the processor. When the program for parallel processing of the deep learning model is executed by the processor, the steps of the method for parallel processing of the deep learning model are implemented.
[0012] According to the third aspect of the present application, an embodiment of the present invention provides a computer-readable storage medium, on which a program for parallel processing of a deep learning model is stored. When the program for parallel processing of the deep learning model is executed by a processor, the steps of the method for parallel processing of the deep learning model are implemented.
[0013] Compared with the related art, a method and apparatus for parallel processing of a deep learning model provided by an embodiment of the present invention determine the dependency relationship between computing nodes in the model, and divide relationship groups according to the dependency relationship; cluster the relationship groups according to a predetermined rule to generate a set of parallel-executable units; wherein, the relationship groups within each set of parallel-executable units can be run in parallel; allocate the relationship groups within all the sets of parallel-executable units to multiple target devices so that the total parallel computing time of all the sets of parallel-executable units is the shortest. The embodiment of the present invention can automatically split a deep learning model and improve the distributed training efficiency when the deep learning model adopts model parallelism. Description of the Drawings
[0014] Figure 1 It is a flowchart of a method for parallel processing of a deep learning model according to Embodiment 1 of the present invention;
[0015] Figure 2 It is a schematic diagram of an apparatus for parallel processing of a deep learning model according to Embodiment 2 of the present invention;
[0016] Figure 3 It is a schematic diagram of a computational graph of the Inception-V3 model in Example 1 of the present invention;
[0017] Figure 4 It is a schematic diagram of selecting relationship groups with the top time-consuming rankings in Example 1 of the present invention;
[0018] Figure 5 It is a schematic diagram of dividing relationship groups according to the name scope field in Example 2 of the present invention;
[0019] Figure 6 It is a schematic diagram of four types of aggregation nodes in Example 2 of the present invention;
[0020] Figure 7 It is a schematic diagram of a serial running branch in Example 2 of the present invention. Detailed Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments and features in the embodiments of this application can be combined arbitrarily with each other.
[0022] The steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0023] Embodiment 1
[0024] As Figure 1 shown, the embodiment of the present invention provides a method for parallel processing of a deep learning model, including:
[0025] Step S110, determining the dependency relationship between computing nodes in the model, and dividing relationship groups according to the dependency relationship;
[0026] Step S120, clustering the relationship groups according to a predetermined rule to generate a set of parallel-executable tasks; wherein, the relationship groups within each set of parallel-executable tasks can run in parallel;
[0027] Step S130, allocating the relationship groups within all the sets of parallel-executable tasks to multiple target devices so that the total parallel computing time of all the sets of parallel-executable tasks is the shortest;
[0028] In an implementation manner, the determining the dependency relationship between computing nodes in the model and dividing relationship groups according to the dependency relationship includes:
[0029] Determining the upstream node and downstream node of each computing node, and the attribute of this computing node;
[0030] Dividing relationship groups in at least one of the following ways;
[0031] Way 1: Dividing computing nodes with the same attribute into the same relationship group;
[0032] Way 2: Dividing a computing node that has only one downstream node and no upstream node, and the downstream node of this computing node into the same relationship group.
[0033] In an implementation manner, the clustering the relationship groups according to a predetermined rule to generate a set of parallel-executable tasks includes:
[0034] Counting the running time of all relationship groups on a single device, and sorting all relationship groups from most time-consuming to least time-consuming;
[0035] Select multiple relationship groups with the top-ranked selection time, and search for a set that can be executed in parallel among the selected relationship groups;
[0036] Among them, the set that can be executed in parallel satisfies the following conditions: there is no upstream and downstream relationship within n levels between any two relationship groups in the set that can be executed in parallel; and there is a common upstream node or a common downstream node within n levels between any two relationship groups in the set that can be executed in parallel; where n is a preset value.
[0037] In one implementation, the selecting multiple relationship groups with the top-ranked selection time includes:
[0038] Select a relationship groups with the top-ranked selection time;
[0039] Among them, a is the smallest integer such that the ratio of the total selection time of a relationship groups to the total selection time of all relationship groups is greater than or equal to a predetermined ratio value.
[0040] In one implementation, the determining the dependency relationship between calculations in the model and dividing relationship groups according to the dependency relationship includes:
[0041] Divide the calculation nodes into relationship groups according to a predetermined field in the name of the calculation node, and the calculation nodes with the same predetermined field belong to the same relationship group;
[0042] Among them, the predetermined field includes: the name scope field; when the name scope field includes nested levels, the predetermined field is the outermost name scope field.
[0043] In one implementation, the clustering the relationship groups according to a predetermined rule to generate a set that can be executed in parallel includes:
[0044] Traverse all the relationship groups and search for a convergence node with multiple inputs or multiple outputs;
[0045] Starting from the convergence node with multiple inputs, traverse all the input nodes of the convergence node upstream until another convergence node is encountered, and generate a set that can be executed in parallel from all the serial running branches between the two convergence nodes; or
[0046] Starting from the convergence node with multiple outputs, traverse all the output nodes of the convergence node downstream until another convergence node is encountered, and generate a set that can be executed in parallel from all the serial running branches between the two convergence nodes;
[0047] Among them, the serial running branch is a set of relationship groups with an upstream and downstream relationship between two convergence nodes.
[0048] In one embodiment, for the above-mentioned relationship group clustering method of dividing relationship groups according to Method 1 and / or Method 2 and sorting them according to the time consumption on a single device, the step of allocating the relationship groups within all the parallel-executable sets to multiple target devices to minimize the total parallel operation time consumption of all the parallel-executable sets includes:
[0049] Using the simulated annealing algorithm to allocate the relationship groups within all the parallel-executable sets to multiple target devices to minimize the total parallel operation time consumption of all the parallel-executable sets.
[0050] In one embodiment, for the above-mentioned relationship group clustering method of dividing relationship groups according to Method 1 and / or Method 2 and sorting them according to the time consumption on a single device, the step of using the simulated annealing algorithm to allocate the relationship groups within all the parallel-executable sets to multiple target devices to minimize the total parallel operation time consumption of all the parallel-executable sets includes:
[0051] Step 1: Initialization: Set the initial temperature T0, the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ when updating the solution each time; at the initial temperature T0, randomly generate the initial solution X0 and calculate the initial time consumption E0; where the initial solution X0 refers to randomly allocating all the relationship groups to the target devices according to the initial allocation method; the initial time consumption E0 refers to the total time consumption after performing the model operation in the initial allocation method.
[0052] Step 2: Perform K perturbation and acceptance processes at the current temperature T, where each perturbation and acceptance process includes: at the current temperature T, randomly select the relationship groups in the current solution X according to the perturbation ratio μ, re-randomly allocate the selected relationship groups to the target devices, and calculate the total time consumption E new of performing the model operation in the new allocation method. If E new is less than E0, accept the new allocation method; if E new is greater than or equal to E0, accept the new allocation method with probability p; where p = exp(-(E new - E0) / T).
[0053] Step 3: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ.
[0054] Step 4: Determine whether the current temperature T is less than the termination temperature T min . If so, take the current allocation method as the final solution and end; otherwise, jump to Step 2 and continue to execute.
[0055] In one implementation, for the above-described relationship group clustering method that divides relationship groups according to predetermined fields in names and clusters according to aggregation nodes, the step of allocating all relationship groups within all sets of parallel-executable tasks to multiple target devices to minimize the total parallel operation time of all sets of parallel-executable tasks includes:
[0056] Using the simulated annealing algorithm to allocate all serial-running branches within all sets of parallel-executable tasks to multiple target devices to minimize the total parallel operation time of all sets of parallel-executable tasks.
[0057] In one implementation, for the above-described relationship group clustering method that divides relationship groups according to predetermined fields in names and clusters according to aggregation nodes, the step of using the simulated annealing algorithm to allocate all serial-running branches within all sets of parallel-executable tasks to multiple target devices to minimize the total parallel operation time of all sets of parallel-executable tasks includes:
[0058] Step 1: Initialization: Set the initial temperature T0, the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ for each solution update; at the initial temperature T0, randomly generate an initial solution X0 and calculate the initial time consumption E0; where the initial solution X0 refers to randomly allocating all serial-running branches to target devices according to the initial allocation method; the initial time consumption E0 refers to the total time consumption after performing model operations under the initial allocation method.
[0059] Step 2: Perform K perturbation and acceptance processes at the current temperature T, where each perturbation and acceptance process includes: at the current temperature T, randomly select serial-running branches in the current solution X according to the perturbation ratio μ, re-randomly allocate the selected serial-running branches to target devices, and calculate the total time consumption E new of performing model operations under the new allocation method. If E new is less than E0, accept the new allocation method; if E new is greater than or equal to E0, accept the new allocation method with probability p; where p = exp(-(E new - E0) / T).
[0060] Step 3: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ.
[0061] Step 4: Determine whether the current temperature T is less than the termination temperature T min . If so, use the current allocation method as the final solution and end; otherwise, jump to Step 2 and continue execution.
[0062] Embodiment 2
[0063] As Figure 2As shown in the figure, an embodiment of the present invention provides a device for parallel processing of a deep learning model, including:
[0064] A relationship group division module 201, configured to determine the dependency relationship between computing nodes in the model, and divide relationship groups according to the dependency relationship;
[0065] A set division module 202, configured to cluster relationship groups according to a predetermined rule to generate a set of parallel-executable sets; wherein, the relationship groups within each set of parallel-executable sets can run in parallel;
[0066] A device allocation module 203, configured to allocate the relationship groups within all the sets of parallel-executable sets to multiple target devices so that the total parallel operation time of all the sets of parallel-executable sets is the shortest;
[0067] In one embodiment, the relationship group division module is configured to determine the dependency relationship between computing nodes in the model in the following manner, and divide relationship groups according to the dependency relationship:
[0068] Determine the upstream node and downstream node of each computing node, and the attribute of the computing node;
[0069] Divide relationship groups according to at least one of the following methods;
[0070] Method 1: Divide computing nodes with the same attribute into the same relationship group;
[0071] Method 2: Divide a computing node with only one downstream node and no upstream node, and the downstream node of the computing node into the same relationship group.
[0072] In one embodiment, the set division module is configured to cluster relationship groups according to a predetermined rule in the following manner to generate a set of parallel-executable sets:
[0073] Statistically calculate the running time of all relationship groups on a single device, and sort all relationship groups from more to less according to the running time;
[0074] Select multiple relationship groups with the top-ranked running time, and search for a set of parallel-executable sets among the selected relationship groups;
[0075] Wherein, the set of parallel-executable sets satisfies the following conditions: there is no upstream and downstream relationship within n levels between any two relationship groups in the set of parallel-executable sets; and there is a common upstream node or a common downstream node within n levels between any two relationship groups in the set of parallel-executable sets; wherein, n is a preset value.
[0076] In one embodiment, the set division module is configured to select multiple relationship groups with the top-ranked running time in the following manner:
[0077] Select the top a relationship groups in terms of time consumption;
[0078] where a is the smallest integer such that the ratio of the total time consumption of the a relationship groups to the total time consumption of all relationship groups is greater than or equal to a predetermined ratio value.
[0079] In one implementation, a relationship group division module is configured to determine the dependency relationship between calculations in the model in the following manner and divide relationship groups according to the dependency relationship:
[0080] Divide relationship groups for calculation nodes according to a predetermined field in the name of the calculation node, and calculation nodes with the same predetermined field belong to the same relationship group;
[0081] where the predetermined field includes: a name scope field; when the name scope field includes nested levels, the predetermined field is the outermost name scope field.
[0082] In one implementation, a set division module is configured to cluster relationship groups according to a predetermined rule in the following manner to generate a set of parallelizable executions:
[0083] Traverse all relationship groups and search for convergence nodes with multiple inputs or multiple outputs;
[0084] Starting from a convergence node with multiple inputs, traverse all input nodes of the convergence node upstream until another convergence node is encountered, and generate a set of parallelizable executions from all serial running branches between the two convergence nodes; or
[0085] Starting from a convergence node with multiple outputs, traverse all output nodes of the convergence node downstream until another convergence node is encountered, and generate a set of parallelizable executions from all serial running branches between the two convergence nodes;
[0086] where the serial running branch is a set of relationship groups with an upstream and downstream relationship between two convergence nodes.
[0087] In one implementation, for the above relationship group clustering method that divides relationship groups according to Method 1 and / or Method 2 and sorts relationship groups according to time consumption on a single device, a device allocation module is configured to allocate relationship groups within all sets of parallelizable executions to multiple target devices in the following manner to minimize the total parallel operation time consumption of all sets of parallelizable executions:
[0088] Use the simulated annealing algorithm to allocate relationship groups within all sets of parallelizable executions to multiple target devices to minimize the total parallel operation time consumption of all sets of parallelizable executions.
[0089] In one implementation, for the above-mentioned relationship group clustering method that divides relationship groups according to Method 1 and / or Method 2 and sorts them according to the time consumption on a single device, the device allocation module is used to adopt the following method to allocate the relationship groups within all the parallel-executable sets to multiple target devices by using the simulated annealing algorithm, so that the total parallel operation time consumption of all the parallel-executable sets is the shortest:
[0090] Step 1: Initialization: Set the initial temperature T0, the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ when updating the solution each time; at the initial temperature T0, randomly generate the initial solution X0 and calculate the initial time consumption E0; where, the initial solution X0 refers to randomly allocating all the relationship groups to the target devices according to the initial allocation method; the initial time consumption E0 refers to the total time consumption after performing the model operation under the initial allocation method.
[0091] Step 2: Perform K perturbation and acceptance processes at the current temperature T, where each perturbation and acceptance process includes: at the current temperature T, randomly select the relationship groups in the current solution X according to the perturbation ratio μ, re-randomly allocate the selected relationship groups to the target devices, and calculate the total time consumption E new of performing the model operation under the new allocation method. If E new is less than E0, then accept the new allocation method. If E new is greater than or equal to E0, then accept the new allocation method with a probability p; where, p = exp(-(E new - E0) / T).
[0092] Step 3: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ.
[0093] Step 4: Determine whether the current temperature T is less than the termination temperature T min . If so, take the current allocation method as the final solution and end; otherwise, jump to Step 2 to continue execution.
[0094] In one implementation, for the above-mentioned relationship group clustering method that divides relationship groups according to the predetermined fields in the name and clusters them according to the aggregation nodes, the device allocation module is used to adopt the following method to allocate the relationship groups within all the parallel-executable sets to multiple target devices so that the total parallel operation time consumption of all the parallel-executable sets is the shortest:
[0095] Adopt the simulated annealing algorithm to allocate the serial running branches within all the parallel-executable sets to multiple target devices so that the total parallel operation time consumption of all the parallel-executable sets is the shortest.
[0096] In one embodiment, for the above-mentioned relationship group clustering method that divides relationship groups according to predetermined fields in names and clusters according to aggregation nodes, the device allocation module is used to adopt the following method to allocate the serial running branches in all parallelizable execution sets to multiple target devices by using the simulated annealing algorithm to minimize the total parallel operation time of all parallelizable execution sets:
[0097] Step 1: Initialization: Set the initial temperature T0, the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ when updating the solution each time; at the initial temperature T0, randomly generate the initial solution X0 and calculate the initial time consumption E0; where, the initial solution X0 refers to randomly allocating all serial running branches to target devices according to the initial allocation method; the initial time consumption E0 refers to the total time consumption after performing model operations in the initial allocation method.
[0098] Step 2: Perform K perturbation and acceptance processes at the current temperature T, where each perturbation and acceptance process includes: at the current temperature T, randomly select serial running branches in the current solution X according to the perturbation ratio μ, re-randomly allocate the selected serial running branches to target devices, and calculate the total time consumption E after performing model operations in the new allocation method new , if E new is less than E0, then accept the new allocation method; if E new is greater than or equal to E0, then accept the new allocation method with probability p; where, p = exp(-(E new - E0) / T);
[0099] Step 3: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ;
[0100] Step 4: Determine whether the current temperature T is less than the termination temperature T min , if so, take the current allocation method as the final solution and end; otherwise, jump to Step 2 and continue to execute.
[0101] Example 3
[0102] An embodiment of the present invention provides a device for parallel processing of a deep learning model, including:
[0103] A memory, a processor, and a program for parallel processing of a deep learning model stored on the memory and executable on the processor. When the program for parallel processing of the deep learning model is executed by the processor, it implements the steps of the method for parallel processing of the deep learning model described in the above Example 1.
[0104] Example 4
[0105] An embodiment of the present invention provides a computer-readable storage medium, on which a program for parallel processing of a deep learning model is stored. When the program for parallel processing of the deep learning model is executed by a processor, the steps of the method for parallel processing of the deep learning model described in Embodiment 1 above are implemented.
[0106] Example 1
[0107] This example provides a method for parallel processing of a deep learning model. In TensorFlow, each deep learning model corresponds to a computational graph, also called a data flow graph, which is a directed graph composed of nodes and edges to describe mathematical operations. Each computation (OP) is a node on the computational graph, and the edges between the nodes describe the dependencies between the computations. Data (Tensor) flows along the edges between the nodes. The computational graph of a complex deep learning model often contains tens of thousands of OPs. For example, in this example, the computational graph of Inception-V3 is used, and the computational graph contains more than 30,000 OPs. Model parallelism is to group these OPs and then train them on different devices.
[0108] In this example, the method for parallel processing of a deep learning model may include the following steps:
[0109] 1) Determine the upstream OP and downstream OP of each OP. As Figure 3 shown, it is a part of the model computational graph.
[0110] 2) Divide the relationship groups, that is, classify the OPs with close relationships into the same relationship group. Mainly based on two principles: one is to put the OPs with the same colocation attribute into the same relationship group; the other is to divide the OP with only one downstream node and no upstream node and its downstream node into the same relationship group. As Figure 3 shown, the OPs at the bottom on both sides (represented by dotted lines) can be divided into a relationship group with their upstream OPs respectively.
[0111] 3) Statistically calculate the time consumed by all relationship groups when running on a single device, and sort all relationship groups from most time-consuming to least time-consuming;
[0112] 4) Select multiple relationship groups with the top time-consuming rankings, and search for a set of parallelizable executions in the selected relationship groups. Among them, the set of parallelizable executions satisfies the following conditions: there is no upstream and downstream relationship within n levels between any two relationship groups in the set of parallelizable executions; and there is a common upstream node or a common downstream node within n levels between any two relationship groups in the set of parallelizable executions; where n is a preset value.
[0113] As Figure 4As shown in the figure, the three relationship groups indicated by dashed lines in the figure are the top three relationship groups in terms of time consumption. Among them, there is an upstream and downstream relationship within two levels between relationship groups a and b, so they cannot be placed in the same parallel execution set. However, a and c fully meet the above conditions, so they can be placed in the same parallel execution set.
[0114] 5) Use the simulated annealing algorithm to allocate the relationship groups within all the parallel execution sets to multiple target devices to minimize the time consumption of parallel computing.
[0115] Example 2
[0116] This example provides a method for parallel processing of deep learning models. In TensorFlow, each deep learning model corresponds to a computational graph, also called a data flow graph, which is a directed graph composed of nodes and edges to describe mathematical operations. Each computation (OP) is a node on the computational graph, and the edges between the nodes describe the dependencies between the computations. Data (Tensor) flows along the edges between the nodes. The computational graph of a complex deep learning model often contains tens of thousands of OPs. For example, in this example, the computational graph of Inception-V3 is adopted, and the computational graph contains more than 30,000 OPs. Model parallelism is to group these OPs and then train them on different devices.
[0117] In this example, the method for parallel processing of deep learning models may include the following steps:
[0118] 1) Divide the OPs into relationship groups according to the name scope field in the OP name. The OPs with the same name scope belong to the same relationship group. When the name scope includes nested levels, they are divided according to the outermost name scope. As Figure 5 shown, the name scope fields of three OPs are "a", and the name scope fields of the other four OPs are "b", and they are respectively divided into the first relationship group surrounded by the dashed box on the left and the second relationship group surrounded by the dashed box on the right.
[0119] 2) Traverse all the relationship groups and search for convergence nodes with multiple inputs or multiple outputs. Among them, Figure 6 shows four types of convergence nodes, from left to right: single input single output, single input multiple outputs, multiple inputs single output, multiple inputs multiple outputs.
[0120] 3) As Figure 7As shown, starting from a convergence node with multiple outputs, traverse all the output nodes of the convergence node downstream until another convergence node is encountered, and take all the serial running branches between the two convergence nodes as a set that can be executed in parallel. Among them, the serial running branch is a set of relationship groups with an upstream-downstream relationship between two convergence nodes;
[0121] 4) Use the simulated annealing algorithm to allocate the serial running branches within all the sets that can be executed in parallel to multiple target devices to minimize the time-consuming of parallel operations.
[0122] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or be implemented as hardware, or be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disc (DVD), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, the communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanisms, and can include any information delivery medium.
[0123] It should be noted that the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for parallel processing of a deep learning model, comprising: Determining the dependency relationships between computing nodes in the model and dividing relationship groups according to the dependency relationships; Clustering the relationship groups according to a predetermined rule to generate a set of parallelizable executions; wherein, the relationship groups within each set of parallelizable executions can run in parallel; Allocating the relationship groups within all sets of parallelizable executions to multiple target devices such that the total parallel computing time of all sets of parallelizable executions is minimized; The determining the dependency relationships between computing nodes in the model and dividing relationship groups according to the dependency relationships includes: Determining the upstream nodes, downstream nodes, and the attributes of each computing node; Dividing the relationship groups in at least one of the following ways; Way 1: Dividing computing nodes with the same attributes into the same relationship group; Way 2: Dividing a computing node that has only one downstream node and no upstream node, and the downstream node of this computing node into the same relationship group.
2. The method according to claim 1, wherein: The clustering the relationship groups according to a predetermined rule to generate a set of parallelizable executions includes: Statistical the running time of all relationship groups on a single device, and sorting all relationship groups from most time-consuming to least time-consuming; Selecting multiple relationship groups with the top-ranked running times, and searching for sets of parallelizable executions among the selected relationship groups; Wherein, the set of parallelizable executions satisfies the following conditions: there is no upstream and downstream relationship within n levels between any two relationship groups in the set of parallelizable executions; and there is a common upstream node or a common downstream node within n levels between any two relationship groups in the set of parallelizable executions; wherein, n is a preset value.
3. The method according to claim 2, wherein: The selecting multiple relationship groups with the top-ranked running times includes: Selecting a relationship groups with the top-ranked running times; Wherein, a is the smallest integer such that the ratio of the total running time of a relationship groups to the total running time of all relationship groups is greater than or equal to a predetermined ratio value.
4. The method according to claim 1, wherein: The determining the dependency relationships between computing nodes in the model and dividing relationship groups according to the dependency relationships includes: Dividing the computing nodes into relationship groups according to a predetermined field in the name of the computing node, and computing nodes with the same predetermined field belong to the same relationship group; Wherein, the predetermined field includes: a name scope field; when the name scope field includes nested levels, the predetermined field is the outermost name scope field.
5. The method according to claim 4, wherein: The clustering the relationship groups according to a predetermined rule to generate a set of parallelizable executions includes: Traversing all relationship groups and searching for convergence nodes with multiple inputs or multiple outputs; Starting from a convergence node with multiple inputs, traversing all input nodes of the convergence node upstream until another convergence node is encountered, and generating a set of parallelizable executions from all serial running branches between the two convergence nodes; or Starting from a convergence node with multiple outputs, traverse all the output nodes of the convergence node downstream until another convergence node is encountered, and generate a set of parallelizable executions from all the serial running branches between the two convergence nodes; wherein, the serial running branches are a set of relationship groups with an upstream and downstream relationship between two convergence nodes.
6. The method according to claim 2, characterized in that: The step of allocating the relationship groups within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions includes: Using a simulated annealing algorithm to allocate the relationship groups within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions.
7. The method according to claim 6, characterized in that: The step of using a simulated annealing algorithm to allocate the relationship groups within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions includes: Step 1: Initialization: Set the initial temperature T0, the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ for each solution update; at the initial temperature T0, randomly generate the initial solution X0 and calculate the initial elapsed time E0; among them, the initial solution X0 refers to randomly allocating all relation groups to the target devices according to the initial allocation method; the initial elapsed time E0 refers to the total elapsed time after performing the model operation under the initial allocation method; Step 2: Conduct K perturbation and acceptance processes at the current temperature T, where each perturbation and acceptance process includes: at the current temperature T, randomly select a relationship group in the current solution X according to the perturbation ratio μ, re-randomly assign the target device to the selected relationship group, and calculate the total time consumption E after performing the model operation under the new allocation method new , if E new is less than E0, accept the new allocation method, if E new is greater than or equal to E0, accept the new allocation method with probability p; where, p = exp(-(E new - E0) / T); Step three: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ; Step 4: Determine whether the current temperature T is less than the termination temperature T min . If yes, take the current allocation method as the final solution and end; otherwise, jump to Step 2 and continue execution.
8. The method according to claim 5, characterized in that: The step of allocating the relationship groups within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions includes: Using a simulated annealing algorithm to allocate the serial running branches within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions.
9. The method according to claim 8, characterized in that: The step of using a simulated annealing algorithm to allocate the serial running branches within all the sets of parallelizable executions to multiple target devices to minimize the total parallel operation time of all the sets of parallelizable executions includes: Step 1: Initialization: Set the initial temperature T0 and the termination temperature T min , the number of iterations K within each temperature, the cooling rate α, and the perturbation ratio μ when updating the solution each time; at the initial temperature T0, randomly generate the initial solution X0 and calculate the initial elapsed time E0; among them, the initial solution X0 refers to randomly allocating all serial operation branches to the target devices according to the initial allocation method; the initial elapsed time E0 refers to the total elapsed time after performing the model operation under the initial allocation method; Step 2: Conduct K perturbation and acceptance processes at the current temperature T. Each perturbation and acceptance process includes: at the current temperature T, randomly select the serial running branches in the current solution X according to the perturbation ratio μ, re-randomly allocate the target devices for the selected serial running branches, and calculate the total elapsed time E after performing the model operation under the new allocation method. new , if E new is less than E0, accept the new allocation method. If E new is greater than or equal to E0, accept the new allocation method with probability p; where p = exp(-(E new - E0) / T); Step three: Update the current temperature T and the perturbation ratio μ: T = αT, μ = αμ; Step 4: Determine whether the current temperature T is less than the termination temperature T min ; if yes, take the current allocation method as the final solution and end; otherwise, jump to Step 2 and continue to execute.
10. An apparatus for parallel processing of a deep learning model, comprising: A memory, a processor, and a program for parallel processing of a deep learning model stored on the memory and executable on the processor, wherein when the program for parallel processing of a deep learning model is executed by the processor, the steps of the method for parallel processing of a deep learning model according to any one of claims 1-9 above are implemented.
11. A computer-readable storage medium, on which a program for parallel processing of a deep learning model is stored, wherein when the program for parallel processing of a deep learning model is executed by a processor, the steps of the method for parallel processing of a deep learning model according to any one of claims 1-9 above are implemented.
Citation Information
Patent Citations
A method for task scheduling with a simulated annealing-based approach in the cloud computing
US20220019463A1