Data Processing Method, Device, Electronic Device, and Storage Medium
By obtaining tensors in the deep neural network model, determining the tensor slicing results and generating an initial parallel strategy, using the cost model optimization strategy, the problem of low accuracy of parallel strategies in the existing technology is solved, and more efficient deep learning training is achieved.
Patent Information
- Application Number
- CN202210283654.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-22
AI Technical Summary
The parallel strategy of the automatically generated deep neural network models in the prior art is low, resulting in poor training efficiency and effectiveness.
By obtaining the tensor of each operator in the deep neural network model, determining the tensor slicing result of each operator based on the tensor, generating an initial parallel strategy, and using a complete cost model to train the initial parallel strategy, obtaining the cost evaluation result, and finally generating the target parallel strategy.
The accuracy and efficiency of the parallel strategy of deep neural network models are improved, the problem of low accuracy of automatic generation of parallel strategies is solved, and more efficient data processing and training process is realized.
Smart Images

Figure CN114611675B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to data processing methods, devices, electronic devices and storage media. Background Art
[0002] Deploying deep neural network models on multiple computing devices is a way to train large-scale complex models. Data parallelism is the most widely used parallel strategy, but as data sets and models become larger and larger, single card memory is limited, and the number of training devices continues to increase, resulting in an increase in communication overhead, data parallelism encounters a bottleneck, and mixed data and model parallelism is needed.
[0003] In related technologies, the parallel strategy of deep neural network models can generally be implemented through dynamic programming search algorithms. However, the dynamic programming search algorithm of AccPar is only applicable to deep neural network models whose computational graphs are linear structures, and the operators at the same layer in the results of AccPar cannot use the hybrid parallel method of data parallelism plus model parallelism, resulting in low accuracy of the automatically generated parallel strategy.
[0004] Currently, no effective solution has been proposed for the problem of low accuracy of parallel strategies automatically generated in related technologies. Summary of the invention
[0005] Embodiments of the present application provide a data processing method, device, electronic device, and storage medium to at least solve the problem of low accuracy of automatically generated parallel strategies in related technologies.
[0006] In a first aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0007] Obtain a deep neural network model and a tensor of each operator in the deep neural network model;
[0008] Determine a tensor segmentation result of each operator based on the tensor, so as to generate an initial parallel strategy at least according to the tensor segmentation result, and train the initial parallel strategy using a complete cost model to obtain a cost evaluation result;
[0009] A target parallel strategy is generated based at least on the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network model to perform data processing operations.
[0010] In some embodiments, generating an initial parallel strategy at least according to the tensor splitting result includes:
[0011] Get the preset device topology model;
[0012] Determine operator parallel tasks corresponding to all the operators according to the degree of parallelism in the tensor splitting result, and determine device deployment results according to the device topology model and the operator parallel tasks;
[0013] The initial parallel strategy is generated according to the tensor segmentation result and the device deployment result.
[0014] In some embodiments, generating the initial parallel strategy according to the tensor segmentation result and the device deployment result includes:
[0015] Get the preset parallel strategy rules;
[0016] Adjusting the tensor splitting result according to the parallel strategy rule to obtain a tensor splitting adjustment result, and adjusting the device deployment result according to the parallel strategy rule to obtain a device deployment adjustment result;
[0017] The initial parallel strategy is generated according to the tensor splitting adjustment result and the device deployment adjustment result.
[0018] In some embodiments, the training process of the initial parallel strategy using the well-built cost model to obtain the cost evaluation result includes:
[0019] Get the preset device topology model;
[0020] Establishing a task graph according to the initial parallel strategy, and acquiring static parallel tasks according to the task graph;
[0021] Using the cost model, the static parallel tasks are traversed and run on the hardware devices indicated by the device topology model, and the actual memory cost and actual time cost corresponding to each of the static parallel tasks are calculated. The cost evaluation result of the initial parallel strategy is calculated based on the actual memory cost and the actual time cost.
[0022] In some embodiments, generating a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result includes:
[0023] Using a preset strategy search algorithm, the cost evaluation result is converted to obtain a probability distribution feature;
[0024] Randomly determine an operator from all the operators, replace the initial parallel strategy corresponding to the operator with a new proposed parallel strategy to determine multiple parallel strategy samples including the proposed parallel strategy, and calculate the proposed cost evaluation results corresponding to the multiple parallel strategy samples using the cost model;
[0025] According to the probability distribution characteristics and the proposed cost evaluation result, the parallel strategy samples are searched and sampled to obtain search results, so as to generate the target parallel strategy based on the search results.
[0026] In some of the embodiments, the policy search algorithm is a Multiple Proposal Markov Chain Monte Carlo (MP-MCMC) policy search algorithm.
[0027] In some embodiments, the tensor includes a data tensor and a weight tensor, and determining the tensor segmentation result of each operator based on the tensor includes:
[0028] Obtain the sample batch dimension of the data tensor, and obtain the input channel dimension and output channel dimension corresponding to the weight tensor;
[0029] The tensor splitting result is determined according to the sample batch dimension, the input channel dimension, and the output channel dimension.
[0030] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising: an acquisition module, a segmentation module and a generation module;
[0031] The acquisition module is used to acquire a deep neural network model and a tensor of each operator in the deep neural network model;
[0032] The segmentation module is used to determine the tensor segmentation result of each operator based on the tensor, so as to generate an initial parallel strategy at least according to the tensor segmentation result, and train the initial parallel strategy using a complete cost model to obtain a cost evaluation result;
[0033] The generation module is used to generate a target parallel strategy based on at least the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network to perform data processing operations.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data processing method as described in the first aspect above is implemented.
[0035] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the data processing method described in the first aspect above is implemented.
[0036] Compared with the related art, the data processing method, device, electronic device and storage medium provided in the embodiments of the present application obtain a deep neural network model and a tensor of each operator in the deep neural network model; determine the tensor segmentation result of each operator based on the tensor, so as to generate an initial parallel strategy at least according to the tensor segmentation result, and train the initial parallel strategy with a complete cost model to obtain a cost evaluation result; generate a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network model to perform data processing operations, thereby solving the problem of low accuracy of automatically generated parallel strategies.
[0037] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0039] Figure 1 is an application environment diagram of a data processing method according to an embodiment of the present application;
[0040] Figure 2 is a flow chart of a data processing method according to an embodiment of the present application;
[0041] Figure 3 is a flow chart of another data processing method according to an embodiment of the present application;
[0042] Figure 4 is a schematic diagram of the architecture of a data processing method according to a preferred embodiment of the present application;
[0043] Figure 5 is a structural block diagram of a data processing device according to an embodiment of the present application;
[0044] Figure 6 It is a structural diagram of the inside of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means, and should not be understood as insufficient contents disclosed in the present application.
[0046] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0047] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to greater than or equal to two. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The terms "first", "second", "third" and the like involved in the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.
[0048] The data processing method provided in this application can be applied to Figure 1In the application environment shown. Among them, the terminal 12 communicates with multiple servers 14 through the network. The terminal 12 obtains the deep neural network model, calculation graph, and device topology model between each server 14 input and set by the user; the terminal 12 transmits each information to the server 14. The server 14 determines the tensor of each operator in the deep neural network model based on the calculation graph, determines the tensor segmentation result of each operator based on the tensor, and generates an initial parallel strategy according to the tensor segmentation result and the device topology model; the server 14 uses the constructed complete cost model to train the initial parallel strategy to obtain a cost evaluation result, and generates a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result; the target parallel strategy is used to indicate the deployment of the deep neural network model on each server 14, so that each server 14 can relay the training of the deep neural network model to perform data processing operations. The terminal 12 can be, but is not limited to, various smart phones, personal computers, laptops, and tablet computers, and the server 14 can be implemented with an independent server or a server cluster consisting of multiple servers.
[0049] This embodiment provides a data processing method. Figure 2 is a flow chart of a data processing method according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:
[0050] Step S220, obtaining a deep neural network model and a tensor of each operator in the deep neural network model.
[0051] Among them, the above-mentioned operators refer to the attribute operations of tensors in the deep neural network model, such as matrix multiplication, tensor addition, convolution, etc.; in this embodiment, they can be equivalent to the layers in the neural network. The above-mentioned tensor refers to an n-dimensional array, which is an n-dimensional generalization of scalars, 1-dimensional vectors and 2-dimensional matrices. In the model training of machine learning, both training data and intermediate calculation results can be regarded as tensors. It can be understood that each operator in the above-mentioned deep neural network model can be determined based on the calculation graph of the deep network model.
[0052] Step S240, determining the tensor segmentation result of each operator based on the tensor, generating an initial parallel strategy at least according to the tensor segmentation result, and training the initial parallel strategy using a complete cost model to obtain a cost evaluation result.
[0053] The tensor splitting result refers to the result of splitting the tasks of each operator based on the tensor using the tensor splitting strategy preset by the user to obtain multiple parallel tasks. After the tensor splitting result indicates that each operator is split into multiple operator parallel tasks, each operator parallel task can constitute the initial parallel strategy. The cost model refers to the cost impact of the initial parallel strategy on computing, communication, memory, etc. during the training process. The cost model obtains the execution cost of the initial parallel strategy by establishing a task graph based on the parallel strategy, and then running a simulation algorithm on the task graph to use the simulation algorithm to obtain the time cost and memory cost of using the initial parallel strategy to execute a parallel training iteration of the deep neural network model; the task graph will include the forward calculation and reverse calculation tasks of all operators in the model parallel training process, the communication tasks between operators, and the update tasks of the operator weights; the simulation algorithm will process the tasks in sequence according to the dependencies between the tasks, simulating the parallel training process in this way; finally, the simulation algorithm is based on the task graph, topologically traverses each parallel task in the initial parallel strategy, calculates the time cost and memory cost of a model parallel training iteration, and combines the time cost and memory cost into an execution cost estimate of the initial parallel strategy using a linear function, thereby obtaining the cost evaluation result corresponding to the above initial parallel strategy.
[0054] Step S260, generating a target parallel strategy based at least on the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network model to perform data processing operations.
[0055] After the initial parallel strategy and the cost evaluation result are determined through the above steps S220 to S240, the cost evaluation result can be used to optimize the initial parallel strategy to finally obtain the above target parallel strategy.
[0056] Through the above steps S220 to S260, by determining the tensor splitting result of each operator in the deep neural network model, generating an initial parallel strategy based on the tensor splitting result, and optimizing the initial parallel strategy through the cost model, it is possible to efficiently obtain the cost feedback result of the initial parallel strategy, thereby realizing the automatic generation of the parallel strategy of the deep neural network model, solving the problem of low accuracy of the automatically generated parallel strategy, and realizing an accurate and efficient automatic generation method of parallel strategy.
[0057] In some embodiments, the tensor includes a data tensor and a weight tensor, and determining the tensor splitting result of each operator based on the tensor also includes the following steps: obtaining the sample batch dimension of the data tensor, and obtaining the input channel dimension and output channel dimension corresponding to the weight tensor; determining the tensor splitting result according to the sample batch dimension, the input channel dimension and the output channel dimension.
[0058] The data tensor mentioned above refers to the tensor of the parameter-free operator in the deep neural network model. The weight tensor mentioned above refers to the tensor of the operator with parameters, such as the convolution operator. Specifically, the tensor segmentation strategy of the operator in this embodiment can be expressed as The tensor splitting strategy is a three-dimensional positive integer vector, and the three vectors are used to represent the operator θ i The parallelism of the sample batch dimension of the input tensor in the data tensor, the parallelism of the input channel dimension of the weight tensor, and the parallelism of the output channel dimension of the weight tensor. For example, if the join operator uses the data parallel and row-based model parallel splitting methods, you can pre-set the tensor splitting strategy to The input tensor used to represent the fully connected operator is evenly split into two parts in the sample batch dimension, and the weight tensor is evenly split into two parts in the input channel dimension, and no split is made in the output channel dimension. In this embodiment, the tensor splitting strategy of the operator will choose to use an equal splitting method in each dimension, so that the operator parallel tasks generated after the tensor splitting can be calculated in a balanced manner. It can be understood that after determining the tensor splitting result based on the tensor splitting strategy set by the user, the tensor splitting result can be further optimized through subsequent cost models and strategy search algorithms to ultimately achieve data parallel and model parallel splitting methods.
[0059] Through the above embodiment, the data tensor and weight parameter are split through the tensor splitting strategy of three-dimensional positive integer vectors, so that each operator can realize the hybrid parallel method of data parallelism and model parallelism, achieve better deep learning parallel training performance improvement effect, and effectively improve the accuracy of the automatic generation method of parallel strategy.
[0060] In some embodiments, a data processing method is provided. Figure 3 is a flow chart of another data processing method according to an embodiment of the present application. Figure 3 As shown, the process includes Figure 3 The steps S220 and S260 shown in the figure further include the following steps:
[0061] Step S320, determining the tensor segmentation result of each operator based on the tensor; and obtaining a preset device topology model.
[0062] Among them, the above-mentioned device topology model refers to a model used to represent the connection relationship between various hardware devices that need to deploy parallel tasks of deep neural network models; the device topology model can be determined in advance based on the hardware device ID information and bandwidth information input by the user.
[0063] Step S340, determining all operator parallel tasks corresponding to the operator according to the parallelism in the tensor splitting result, and determining the device deployment result according to the device topology model and the operator parallel tasks.
[0064] The above equipment deployment results are related to the operator's tensor splitting strategy. For example, in this embodiment, for The product of the parallelism in three dimensions. According to the tensor splitting strategy of the operator The execution of an operator can be divided into independent tasks; these parallel tasks of the operator can be expressed as At the same time, the device deployment strategy of the operator will determine the allocation device of each operator parallel task. Specifically, in this embodiment, in Indicates the device accelerator number to which the operator parallel task is assigned.
[0065] Step S360, generating the initial parallel strategy according to the tensor segmentation result and the device deployment result, and training the initial parallel strategy using the complete cost model to obtain a cost evaluation result.
[0066] Specifically, based on the analysis of operator tensor segmentation and the demand for operator heterogeneous computing, this application proposes a four-dimensional strategy space, expressed as (Batch, InChannel, OutChannel, Placement); among them, Batch, InChannel, and OutChannel correspond to the three dimensions of operator tensor segmentation: the segmentation of the sample batch dimension of the operator input tensor, and the segmentation of the input channel dimension and the output channel dimension of the operator weight tensor. Placement represents the device where the operator parallel task needs to be deployed. For common data parallel methods, all operators need to be deployed to all devices for execution; however, when the model parameters are so large that they cannot be distributed to each device, the sub-training tasks generated after the operator segmentation need to be deployed on different device accelerators, and the parallel training of the model is completed through "relay" execution. Based on the above analysis, the above tensor segmentation results and device deployment results can together constitute a four-dimensional initial parallel strategy. Initial parallel strategy Including each operator o i Tensor partitioning strategy Device deployment strategy for parallel tasks with operators For a deep neural network model with a total number of operators N, in this embodiment, in According to the initial parallel strategy It can be clarified how to perform model segmentation and task deployment during parallel training of the model, thereby realizing the execution process of parallel training of the model. In addition, engineers can also independently select operators o i Corresponding parallel strategy Replace it with any new feasible strategy It should be noted that this process will not affect the parallel strategies of other operators, but may affect the overall parallel training process of the model.
[0067] Through the above steps S320 to S360, a four-dimensional strategy space is realized through the tensor splitting strategy and the device deployment strategy, so that the sub-training tasks generated after the operator splitting can be deployed on different device accelerators, providing a more efficient and accurate parallel strategy automatic generation method.
[0068] In some embodiments, generating the initial parallel strategy according to the tensor segmentation result and the device deployment result further includes the following steps:
[0069] Step S361, obtaining preset parallel strategy rules.
[0070] The above parallel strategy rules refer to user-defined splitting and deployment strategy rules for parallel tasks; the parallel strategy rules can be pre-set by the user based on actual conditions. For example, the user can input parameters for indicating that the input channel dimension of the weight tensor is split into two parts through the above terminal, or input parameters for indicating that several operator parallel tasks are deployed on the server hardware device with ID information 2, etc.
[0071] Step S362: adjust the tensor splitting result according to the parallel strategy rule to obtain a tensor splitting adjustment result, and adjust the device deployment result according to the parallel strategy rule to obtain a device deployment adjustment result; generate the initial parallel strategy according to the tensor splitting adjustment result and the device deployment adjustment result.
[0072] Among them, considering that users may have the need to customize parallel training, such as the need to limit the parallel training method of some model subgraphs, or the need to allow certain specific operators to execute on special hardware accelerators, etc., the rule engine is introduced in this embodiment to facilitate users to control the results of automatic generation of parallel strategies. Specifically, the user-defined parallel strategy rules will affect the tensor splitting method and device allocation method of some operators in the network model through the rule engine, that is, the user-defined rule set is processed by the rule engine to generate parallel strategy rule operations, which will limit the parallel strategy selection of certain operators; for example, if the parallel strategy rule indicates that the weight tensor is split into two parts in the input channel dimension, and the three operator parallel tasks are allocated to the server hardware device deployment with ID information 2, then the tensor splitting result generated The second dimension of will be adjusted to 2, and the device deployment result will be partially adjusted to comply with the parallel strategy rule. It can be understood that the above parallel strategy rule can also only indicate the adjustment of the tensor segmentation result, or only indicate the adjustment of the device deployment result, which will not be repeated here.
[0073] Through the above steps S361 to S362, the tensor splitting results and / or device deployment results are adjusted based on the obtained parallel strategy rules to generate an initial parallel strategy, thereby providing the user with an optional custom rule interface, allowing the user to adjust the parallel strategy based on actual conditions, thereby improving the user experience.
[0074] In some embodiments, the above-mentioned training process of the initial parallel strategy using the constructed complete cost model to obtain the cost evaluation result further includes the following steps:
[0075] Step S241, obtaining a preset device topology model; establishing a task graph according to the initial parallel strategy, and obtaining static parallel tasks according to the task graph.
[0076] Specifically, the above task graph can be based on the computation graph Device topology model and initial parallel strategy The task graph can be expressed as T = (N, E); where N represents the number of all parallel tasks and E represents the connection relationship between each parallel task. The above-mentioned static parallel tasks can be screened out from the task graph; the static parallel tasks refer to tasks whose calculation time of some operators is independent of the content.
[0077] Step S242, using the cost model, traverse and run the static parallel tasks on the hardware devices indicated by the device topology model, calculate the actual memory cost and actual time cost corresponding to each of the static parallel tasks, and calculate the cost evaluation result of the initial parallel strategy based on the actual memory cost and the actual time cost.
[0078] Specifically, Table 1 shows a pseudo code of a cost model simulation algorithm in an embodiment of the present application, as shown in Table 1:
[0079] Table 1 Cost model simulation algorithm
[0080]
[0081]
[0082] The above simulation algorithm first completes the construction of the task graph T, including the creation of task nodes, the device allocation of task nodes, the creation of communication tasks when adjacent computing tasks are in different devices, and the initialization of task nodes. and Other attributes will be initialized and updated when the simulation algorithm is executed. The data attributes related to the tasks, equipment, and cost involved in the above simulation algorithm can refer to the description in Table 2, as shown in Table 2:
[0083] Table 2 Description of data structure attributes of cost model simulation algorithm
[0084]
[0085]
[0086] Based on the content in the above table, the above simulation algorithm will obtain it through the MeasureOpCost method. Specifically, in the above simulation algorithm, the input tensor shape size and weight tensor shape size of each static parallel task t in the above task graph can be obtained, as well as the data type used to store the tensor. Therefore, the above cost model can randomly initialize the input tensor and weight tensor with the same data type and shape size, and perform multiple simulation calculations on the hardware device where the static parallel task is located, that is, the hardware device indicated by the above device topology model, so as to obtain the parallel execution time of the static parallel task, and the average memory occupancy as the cost estimate of the static parallel task, that is, the above actual time cost and the above actual memory cost. The simulation algorithm can obtain the shape size and data type of the tensor that the communication task needs to transmit, and can also calculate the time cost of the communication transmission. The time cost estimate of the communication task is equal to the time s / b for transmitting a tensor with s bytes between devices with a connection bandwidth of b (unit Byte / s). In order to maximize the efficiency of algorithm operation, the cost model caches the cost estimates of static parallel tasks to avoid repeated calculations; when the simulation algorithm encounters a new task with an operator type, tensor shape, data type, and device type that has appeared before, the cost model can directly return the cached cost data.
[0087] Since the embodiment of the present application is to find a parallel strategy that minimizes the execution time cost of parallel training while satisfying the memory constraints of each device in the device cluster, the simulation algorithm used in the above cost model calculates the memory cost and time cost of the computing tasks assigned to each device during the training process, and then returns the sum of the overall execution time cost of one iteration of parallel training and the memory overflow on all devices. The simulation algorithm of the cost model will finally use a linear function f:R→R to calculate the execution cost of the parallel strategy. The linear function will combine the execution time and memory cost of the parallel strategy into the total cost of the parallel strategy, that is, the above cost evaluation result; the time cost c time Defined as the end time of the latest task in the task graph, the memory cost c mem It is defined as the sum of device memory overflow. The scalar a∈R and scalar b∈R are the time cost coefficient and memory cost coefficient respectively, which are related to the specific hardware and network topology.
[0088] Through the above steps S241 to S242, the static parallel tasks are traversed and run on the hardware device through the cost model, and the actual memory cost and actual time cost corresponding to each static parallel task are calculated to obtain the cost evaluation result of the initial parallel strategy, thereby realizing an efficient acquisition method for the calculation, communication, and memory feedback of the parallel strategy through the cost model, and realizing a cost estimation method for obtaining the actual execution cost overhead by running the calculation of the operator multiple times in the actual physical device environment, avoiding the drawbacks of only considering the execution time cost in the related technology, and further improving the accuracy of automatic generation of parallel strategies.
[0089] In some embodiments, the step of generating a target parallel strategy according to the initial parallel strategy and the cost evaluation result further includes the following steps:
[0090] Step S261 , using a preset strategy search algorithm, converting the cost evaluation result to obtain a probability distribution feature.
[0091] Among them, the above-mentioned strategy search algorithm can adopt MP-MCMC, Metropolis-Hastings algorithm (MH), or MCMC and other strategy search algorithms. It should be supplemented that the MP-MCMC algorithm is a generalization method of the MH algorithm, which improves the calculation speed and statistical efficiency of the existing MH method through parallel calculation. Therefore, the strategy search algorithm based on MP-MCMC can improve the search speed and optimization efficiency of parallel strategies, thereby optimizing the autonomous generation method of parallel strategies, and is suitable for the discrete high-dimensional characteristics of the parallel strategy search space in this embodiment. Specifically, using the MP-MCMC strategy search algorithm, a Markov chain is constructed to make its stable distribution the probability distribution p(x) of the parameter to be estimated, and a sample of the probability distribution is generated through this Markov chain, starting from the initial state x 0 Starting from the conditional probability distribution κ(x t ,·)(t=0,1,…) generates samples, and determines whether to transfer through the transfer conditions. After m updates, it reaches stability, and the samples generated thereafter are p(x). Simply put, MCMC samples high-probability samples from the probability distribution more often than it samples low-probability samples. In addition, MCMC methods are often used in practice in scenarios where the probability distribution of unknown parameters is high-dimensional, complex and uncommon. Similarly, the parallel strategy search space used in this study also has the characteristics of discrete high dimensions, which is very consistent with the MCMC method. Formula 1 represents a common method for converting the cost function cost(S) into the probability distribution feature p(S):
[0092] p(S)∝exp(-β·cost(S)) Formula 1
[0093] Among them, β is an adjustable hyperparameter, and the cost function cost(S) is a method to obtain the cost evaluation result of the above initial parallel strategy through the above cost model. Therefore, the policy search algorithm can use the parallel strategy execution cost obtained by the cost model to convert the problem of minimizing the parallel strategy execution cost into the problem of sampling from the probability distribution of parallel strategies. The policy search algorithm can start from the initial parallel strategy and obtain many feasible parallel strategy samples through sampling. Finally, the policy search algorithm will select the parallel strategy sample with the lowest cost as the final output of the policy optimizer.
[0094] Step S262, randomly determine an operator from all the operators, replace the initial parallel strategy corresponding to the operator with the new proposed parallel strategy to determine multiple parallel strategy samples including the proposed parallel strategy, and use the cost model to calculate the proposed cost evaluation results corresponding to the multiple parallel strategy samples.
[0095] In order to satisfy the symmetry condition of the proposed kernel function in the MH method, a simple method of proposing a new parallel strategy is adopted in the strategy search algorithm in the embodiment of the present application. The proposed method for obtaining a new parallel strategy is to randomly select an operator in the current model calculation graph. i , and randomly replace the parallel strategy of the operator with the new proposed parallel strategy To determine multiple parallel strategy samples including the proposed parallel strategy. This proposal method can satisfy the symmetry condition of the proposed kernel function in the MH method. This is because the parallel strategy of any operator is updated randomly with the same probability, so the proposed kernel function can satisfy the symmetry condition:
[0096]
[0097] Step S263 , searching and sampling the parallel strategy sample according to the probability distribution feature and the proposed cost evaluation result to obtain search results, so as to generate the target parallel strategy based on the search results.
[0098] Among them, when the proposed kernel function satisfies the symmetry when using the above strategy search algorithm, the conversion formula of the probability distribution characteristics described in the above formula 1 is substituted to obtain the constraint conditions for determining whether to accept the above new proposed parallel strategy, as shown in the following formula 3:
[0099]
[0100] It is necessary to add that represents the initial parallel strategy, represents the new proposed parallel strategy, represents a method for obtaining the cost evaluation result of the initial parallel strategy through the above cost model, It represents the method of obtaining the proposed cost estimation result of the proposed parallel strategy through the cost model. From formula 3, we can see that if the new proposed parallel strategy The execution cost is lower than the initial parallel strategy So the new strategy will always be accepted; if the new parallel strategy The loss is higher than the initial parallel strategy It is possible that the transfer will be rejected and the initial parallel strategy will still be used. Through the above strategy search algorithm, starting from the initial parallel strategy, many feasible proposed parallel strategy samples are obtained through sampling, and finally the parallel strategy sample with the lowest cost is searched out as the final output search result, which is the above target parallel strategy. Specifically, the process of obtaining a new parallel strategy may include the following steps: In each iteration of the algorithm, first obtain the current initial parallel strategy S 0 As input, a new proposal parallel strategy is proposed through the proposal kernel function, that is, the conditional distribution Finally, the algorithm will determine whether to transfer based on the transfer conditions. If the transfer conditions are met, the proposed parallel strategy will be accepted. At this time, you can set is the current strategy; if the transfer fails, the strategy search algorithm will continue based on S 0 Propose a new parallel strategy. The above process can be performed continuously until the algorithm has not found a new and better parallel strategy for a long time, or the algorithm reaches the time budget, then the search process can be terminated. Through the above steps, the sampling of parallel strategies can tend to move towards low-cost strategies, which also helps to escape the local optimal point.
[0101] Specifically, Table 3 shows a pseudo code of an MP-MCMC strategy search algorithm in an embodiment of the present application, as shown in Table 3:
[0102] Table 3 MP-MCMC strategy search algorithm
[0103]
[0104]
[0105] Each iteration in the above strategy search algorithm can be divided into two steps. First, step 1 is based on the current initial parallel strategy state For example, if I = 0, N new strategies are proposed in parallel by proposing kernel functions. Then, step 2 proposes a set of strategies It is regarded as a Markov chain with a transition probability of A(i,j)(i,j∈[0,N]) and samples from it to set N new strategy samples. The last sample in each round of MCMC iteration will become the initial state of the next MCMC iteration. Therefore, in this embodiment, a policy search algorithm based on MP-MCMC is used to improve the search speed and optimization efficiency of parallel strategies; through the policy search algorithm, multiple new proposed parallel strategies are proposed and calculated in parallel in each round of search optimization, thereby improving the algorithm's exploration efficiency in the parallel strategy search space, which is conducive to the algorithm finding a parallel strategy that can further improve the parallel training performance within a limited time.
[0106] In the related art, the search space of the FlexFlow parallel strategy automatic generation method is not complete, and does not support the complete tensor segmentation model parallel method. In addition, the search algorithm of FlexFlow uses the ordinary MH random search method, and the search efficiency is low. AccPar's dynamic programming search algorithm is only applicable to deep neural network models with linear structures in the calculation graph, and the granularity of the search algorithm is not a finer operator level, but a coarser level. In the results of AccPar, the operators at the same layer cannot use the hybrid parallel method of data parallelism plus model parallelism, so the performance of the parallel strategy obtained in the end is also suboptimal.
[0107] The embodiment of the present application avoids the drawback of the lack of operator parallel methods in the parallel strategy search space by using the complete operator-level parallel strategy search space through the above steps S261 to S263, so that each operator can implement data parallelism and model parallel methods of row and column splitting, so as to find previously unknown parallel strategy results and achieve better deep learning parallel training performance improvement effects. In addition, by adopting the new MP-MCMC search algorithm, feasible new strategies can be explored in the parallel strategy search space more efficiently. The MP-MCMC search algorithm will find new strategy results that can improve parallel training performance in the search space through parallel calculation and sampling based on the parallel strategy performance cost feedback of the cost model. The MP-MCMC search algorithm is more conducive to exploration in high-dimensional discrete strategy search space, which improves the optimization efficiency of parallel strategies.
[0108] The embodiments of the present application are described and illustrated below through preferred embodiments. Figure 4 is a schematic diagram of the architecture of a data processing method according to a preferred embodiment of the present application, such as Figure 4As shown, the architecture is mainly composed of a policy optimizer. Specifically, in this embodiment, the computational graph model corresponding to the deep neural network model and the device topology model used for parallel training are used as the main inputs of the policy optimizer, and the user-defined parallel strategy rules are input into the policy optimizer through the rule engine, so that the initial parallel strategy is generated based on the input parameters by the policy optimizer. The policy optimizer is mainly composed of a cost model and an MP-MCMC search algorithm. The cost model uses a simulation algorithm to efficiently estimate the cost impact of the parallel strategy on the calculation, communication, and memory during the training process, which is used to evaluate the impact of the parallel strategy on the parallel training performance and feed back the loss to the MP-MCMC search algorithm. Finally, the MP-MCMC search algorithm will propose multiple new candidate parallel strategies, i.e., the above-mentioned parallel strategy samples, based on the parallel strategy cost feedback of the cost model through parallel computing, so as to continuously search for new parallel strategies with lower costs. When the search time budget is exhausted, the policy optimizer will return the parallel strategy with the lowest cost as the result. The policy optimizer uses the MP-MCMC method of parallel computing to improve the search optimization efficiency of parallel strategies, which is conducive to exploring effective parallel strategy results in high-dimensional discrete strategy search space. Compared with the parallel strategy random search algorithm used by FlexFlow in related technologies, the MP-MCMC search algorithm can search for parallel strategies with better performance improvement effects within the same time budget.
[0109] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0110] This embodiment also provides a data processing device, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the terms "module", "unit", "subunit" and the like can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0111] Figure 5 is a structural block diagram of a data processing device according to an embodiment of the present application, such as Figure 5As shown, the device includes: an acquisition module 52, a segmentation module 54 and a generation module 56. The acquisition module 52 is used to acquire a deep neural network model and a tensor of each operator in the deep neural network model; the segmentation module 54 is used to determine the tensor segmentation result of each operator based on the tensor, so as to generate an initial parallel strategy at least according to the tensor segmentation result, and train the initial parallel strategy using a complete cost model to obtain a cost evaluation result; the generation module 56 is used to generate a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network to perform data processing operations.
[0112] Through the above embodiment, the tensor segmentation result of each operator in the deep neural network model is determined by the segmentation module 54, an initial parallel strategy is generated based on the tensor segmentation result, and the initial parallel strategy is optimized through the cost model, so that the cost feedback result of the initial parallel strategy can be efficiently obtained, thereby realizing the automatic generation of the parallel strategy of the deep neural network model, solving the problem of low accuracy of the automatically generated parallel strategy, and realizing an accurate and efficient parallel strategy automatic generation device.
[0113] In some embodiments, the above-mentioned splitting module 54 is also used to obtain a preset device topology model; the splitting module determines all operator parallel tasks corresponding to the operator according to the parallelism in the tensor splitting result, and determines the device deployment result according to the device topology model and the operator parallel task; the splitting module 54 generates the initial parallel strategy according to the tensor splitting result and the device deployment result.
[0114] In some embodiments, the above-mentioned splitting module 54 is also used to obtain preset parallel strategy rules; the splitting module 54 adjusts the tensor splitting result according to the parallel strategy rule to obtain a tensor splitting adjustment result, and adjusts the device deployment result according to the parallel strategy rule to obtain a device deployment adjustment result; the splitting module 54 generates the initial parallel strategy based on the tensor splitting adjustment result and the device deployment adjustment result.
[0115] In some embodiments, the above-mentioned segmentation module 54 is also used to obtain a preset device topology model; the segmentation module 54 establishes a task graph according to the initial parallel strategy, and obtains static parallel tasks according to the task graph; the segmentation module 54 uses the cost model to traverse and run the static parallel tasks on the hardware devices indicated by the device topology model, calculates the actual memory cost and the actual time cost corresponding to each of the static parallel tasks, and calculates the cost evaluation result of the initial parallel strategy based on the actual memory cost and the actual time cost.
[0116] In some of the embodiments, the generation module 56 is also used to use a preset strategy search algorithm to transform the cost evaluation result to obtain a probability distribution feature; the generation module 56 randomly determines an operator from all the operators, replaces the initial parallel strategy corresponding to the operator with a new proposed parallel strategy to determine a plurality of parallel strategy samples including the proposed parallel strategy, and uses the cost model to calculate the proposed cost evaluation results corresponding to the plurality of parallel strategy samples; the generation module 56 searches and samples the parallel strategy samples according to the probability distribution feature and the proposed cost evaluation result to obtain a search result, so as to generate the target parallel strategy based on the search result.
[0117] In some of the embodiments, the above-mentioned policy search algorithm is a MP-MCMC policy search algorithm.
[0118] In some embodiments, the tensor includes a data tensor and a weight tensor; the splitting module 54 is also used to obtain the sample batch dimension of the data tensor, and to obtain the input channel dimension and the output channel dimension corresponding to the weight tensor; the splitting module 54 determines the tensor splitting result according to the sample batch dimension, the input channel dimension and the output channel dimension.
[0119] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0120] In some embodiments, a computer device is provided, which may be a server. Figure 6 is a structural diagram of the inside of a computer device according to an embodiment of the present application, such as Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store target parallel strategies. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data processing method is implemented.
[0121] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0122] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0123] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0124] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0125] S1, obtain the deep neural network model and the tensor of each operator in the deep neural network model.
[0126] S2, determining the tensor segmentation result of each operator based on the tensor, so as to generate an initial parallel strategy at least according to the tensor segmentation result, and training and processing the initial parallel strategy using a complete cost model to obtain a cost evaluation result.
[0127] S3, generating a target parallel strategy based at least on the initial parallel strategy and the cost evaluation result; wherein the target parallel strategy is used to instruct the deep neural network model to perform data processing operations.
[0128] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0129] In addition, in combination with the data processing method in the above embodiments, the present application embodiment may provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the data processing methods in the above embodiments is implemented.
[0130] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0131] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0132] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A data processing method, characterized in that, the method includes: obtaining a deep neural network model and tensors of each operator in the deep neural network model; determining a tensor segmentation result of each operator based on the tensors, and generating an initial parallel strategy at least according to the tensor segmentation result, including: obtaining a preset device topology model; determining operator parallel tasks corresponding to all the operators according to the parallel degree in the tensor segmentation result, and determining a device deployment result according to the device topology model and the operator parallel tasks; generating the initial parallel strategy according to the tensor segmentation result and the device deployment result; performing training processing on the initial parallel strategy by using a constructed complete cost model to obtain a cost evaluation result, including: establishing a task graph according to the initial parallel strategy, and obtaining static parallel tasks according to the task graph; using the cost model to traverse and run the static parallel tasks on the hardware devices indicated by the device topology model, calculating the actual memory cost and the actual time cost corresponding to each static parallel task, and calculating the cost evaluation result of the initial parallel strategy according to the actual memory cost and the actual time cost; generating a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result; wherein, the target parallel strategy is used to instruct the deep neural network model to perform data processing operations.
2. The data processing method according to claim 1, characterized in that, the generating the initial parallel strategy according to the tensor segmentation result and the device deployment result includes: obtaining a preset parallel strategy rule; performing adjustment processing on the tensor segmentation result according to the parallel strategy rule to obtain a tensor segmentation adjustment result, and performing adjustment processing on the device deployment result according to the parallel strategy rule to obtain a device deployment adjustment result; generating the initial parallel strategy according to the tensor segmentation adjustment result and the device deployment adjustment result.
3. The data processing method according to claim 1, characterized in that, the generating a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result includes: using a preset policy search algorithm to perform conversion processing on the cost evaluation result to obtain a probability distribution feature; randomly determining an operator from all the operators, and replacing the initial parallel strategy corresponding to the operator with a new proposed parallel strategy to determine a plurality of parallel strategy samples including the proposed parallel strategy, and calculating the proposed cost evaluation results corresponding to the plurality of parallel strategy samples by using the cost model; performing search sampling on the parallel strategy samples according to the probability distribution feature and the proposed cost evaluation results to obtain a search result, and generating the target parallel strategy based on the search result.
4. The data processing method according to claim 3, characterized in that, the policy search algorithm is an MP-MCMC policy search algorithm.
5. The data processing method according to any one of claims 1 to 4, characterized in that, The tensor includes a data tensor and a weight tensor. Determining the tensor splitting result of each operator based on the tensor includes: Obtaining the sample batch dimension of the data tensor, and obtaining the input channel dimension and the output channel dimension corresponding to the weight tensor; Determining the tensor splitting result according to the sample batch dimension, the input channel dimension, and the output channel dimension.
6. A data processing device, characterized in that the device includes: an acquisition module, a splitting module, and a generation module; the acquisition module is configured to acquire a deep neural network model, and tensors of each operator in the deep neural network model; the splitting module is configured to determine the tensor splitting result of each operator based on the tensor, generate an initial parallel strategy at least according to the tensor splitting result, and perform training processing on the initial parallel strategy by using a constructed complete cost model to obtain a cost evaluation result; the splitting module is further configured to obtain a preset device topology model; determine operator parallel tasks corresponding to all the operators according to the degree of parallelism in the tensor splitting result, and determine a device deployment result according to the device topology model and the operator parallel tasks; generate the initial parallel strategy according to the tensor splitting result and the device deployment result; the splitting module is further configured to establish a task graph according to the initial parallel strategy, and obtain static parallel tasks according to the task graph; use the cost model to traverse and run the static parallel tasks on the hardware device indicated by the device topology model, calculate the actual memory cost and the actual time cost corresponding to each static parallel task, and calculate the cost evaluation result of the initial parallel strategy according to the actual memory cost and the actual time cost; the generation module is configured to generate a target parallel strategy at least according to the initial parallel strategy and the cost evaluation result; wherein, the target parallel strategy is used to instruct the deep neural network to perform data processing operations.
7. An electronic device, including a memory and a processor, characterized in that a computer program is stored in the memory, and the processor is configured to run the computer program to execute the data processing method according to any one of claims 1 to 5.
8. A storage medium, characterized in that a computer program is stored in the storage medium, wherein the computer program is configured to execute the data processing method according to any one of claims 1 to 5 when running.
Citation Information
Patent Citations
Underwater target recognition method based on GRU and one-dimensional CNN neural network fusion
CN110807365A
Self-supervised three-dimensional reconstruction method and system based on collaborative segmentation and data enhancement
CN112767468A