Parallel strategy design method, electronic equipment, storage medium and computer program product

Through a modular system, the parallel strategy of the preset model is visualized, which solves the problem that parallel strategy cannot be designed efficiently in the existing technology, and achieves the effect of improving the model training efficiency.

CN120066775APending Publication Date: 2025-05-30CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125636.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing distributed machine learning framework lacks auxiliary solutions and visualization solutions when designing parallel strategies, resulting in the inability to efficiently design parallel strategies, thereby reducing the training efficiency of the model.

Method used

A parallel strategy design method is provided, the parallel strategy of the preset model is visualized through a modular system, and the second parallel strategy is determined based on the visualization results, and the model is trained through the training data set.

Benefits of technology

Through visual processing, you can intuitively understand the parallel strategy type of the model, and determine efficient parallel strategy based on this, which improves the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066775A_ABST
    Figure CN120066775A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a parallel strategy design method, electronic equipment, a storage medium and a computer program product, the method is applied to a parallel strategy design system, the parallel strategy design system comprises a first module and a second module, and the method comprises the following steps: performing visualization processing on a first parallel strategy of a preset model through the first module; wherein the first parallel strategy comprises a data parallel strategy, and / or, a tensor parallel strategy, and / or, an assembly line parallel strategy, and the first module is used for visualizing a segmentation mode or resource occupation information of the preset model; after the second parallel strategy is determined, a preset model is trained through the second module, the second parallel strategy and the training data set; wherein the second parallel strategy is a part of or all parallel strategies in the first parallel strategy, namely, according to the embodiment of the invention, the first parallel strategy is visualized, so that the design of the parallel strategy is efficiently carried out, and the training efficiency of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of cloud computing and big data technologies, and in particular, to a parallel strategy design method, an electronic device, a storage medium, and a computer program product. Background Art

[0002] Deep learning has been widely applied in various fields, and the emerging machine learning software ecosystem has been continuously enriched. In recent years, the scale and complexity of deep learning models have been increasing. For example, the Generative Pre-trained Transformer 3 (GPT-3) has 175 billion parameters, which is nearly a thousand times larger than the early Bidirectional Encoder Representations from Transformers (BERT) model. At this time, a single Graphics Processing Unit (GPU) or even multiple GPUs on a server can no longer accommodate such a large model. Only with a large number of GPUs can there be sufficient computing power to complete the training of the model within an acceptable time. Therefore, it is necessary to distribute the workload across multiple devices. This cross-server distributed machine learning is particularly important for training large-scale models.

[0003] However, many current distributed machine learning frameworks analyze the model structure manually, then partition the model, and maximize the communication efficiency through continuous experimental tuning. They lack auxiliary and visualization solutions for partitioning the model during distributed training, so it is impossible to intuitively understand the bottlenecks during model training and design the parallel strategy efficiently, resulting in a decline in the training efficiency of the model. Summary of the Invention

[0004] Embodiments of this application provide a parallel strategy design method, an electronic device, a storage medium, and a computer program product, which can improve the training efficiency of the model.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a parallel strategy design method, which is applied to a parallel strategy design system. The parallel strategy design system includes a first module and a second module. The method includes:

[0007] Visualize the first parallel strategy of the preset model through the first module; wherein, the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visualize the segmentation method or resource occupancy information of the preset model;

[0008] After determining the second parallel strategy, train the preset model through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy is part or all of the parallel strategies in the first parallel strategy.

[0009] In a second aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory; wherein,

[0010] The memory is used to store a computer program that can run on the processor;

[0011] The processor is used to execute the parallel strategy design method as described above when running the computer program.

[0012] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program code is stored, and when the computer program code is executed by a computer, the parallel strategy design method as described above is implemented.

[0013] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the parallel strategy design method as described above is implemented.

[0014] An embodiment of the present application provides a parallel strategy design method, an electronic device, a storage medium, and a computer program product. The method is applied to a parallel strategy design system, which includes a first module and a second module. The method includes: visually processing a first parallel strategy of a preset model through the first module; where the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visually process the splitting method or resource occupancy information of the preset model; after determining a second parallel strategy, training the preset model through the second module, the second parallel strategy, and a training data set; where the second parallel strategy is part or all of the parallel strategies in the first parallel strategy. Thus, it can be seen that the first parallel strategy of the preset model can be visually processed through the first module, that is, an embodiment of the present application can visually process the data parallel strategy, and / or the tensor parallel strategy, and / or the pipeline parallel strategy of the preset model through the first module. After determining the second parallel strategy, the preset model can be trained through the second module, the second parallel strategy, and the training data set. That is, an embodiment of the present application visualizes the first parallel strategy, so that the type of the parallel strategy of the preset model can be intuitively understood, and the second parallel strategy can be determined based on different parallel strategies, realizing efficient design of the parallel strategy. Furthermore, the preset model can be trained through the second module, the second parallel strategy, and the training data set, thereby improving the training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Schematic diagram of the parallel strategy design method proposed by an embodiment of the present application Figure 1 ;

[0016] Figure 2 Schematic diagram of the parallel strategy design method proposed by an embodiment of the present application Figure 2 ;

[0017] Figure 3 Schematic diagram of the directed graph structure proposed by an embodiment of the present application;

[0018] Figure 4 Schematic diagram of the first preset color proposed by an embodiment of the present application;

[0019] Figure 5 Schematic diagram of the second preset color proposed by an embodiment of the present application;

[0020] Figure 6 Schematic diagram of the tensor dimension representation proposed by an embodiment of the present application;

[0021] Figure 7 Schematic diagram of the parallel strategy design method proposed by an embodiment of the present application Figure 3 ;

[0022] Figure 8 Schematic diagram of the visual interface layout proposed in the embodiment of the present application;

[0023] Figure 9 Schematic diagram of the model structure proposed in the embodiment of the present application;

[0024] Figure 10 Schematic diagram of the composition structure of the electronic device proposed in the embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the relevant application, rather than limiting the application. Additionally, it should be noted that for the sake of description, only the parts related to the relevant application are shown in the drawings.

[0026] Deep learning emerged approximately a decade ago, becoming a hot topic in the academic and Internet industries and being widely applied to various fields. Subsequently, the emerging machine learning software ecosystem has been continuously enriched, and mainstream machine learning frameworks such as Pytorch, TensorFlow, and MXNet. In recent years, the scale and complexity of deep learning models have been continuously increasing. For example, GPT-3 has 175 billion parameters, which is nearly a thousand times larger than the early BERT model. At this time, a single GPU or even multiple GPUs on a server can no longer accommodate such a large model, and only a large number of GPUs can provide sufficient computing power to complete the training of the model within an acceptable time. Therefore, it is necessary to distribute the workload across multiple devices. This cross-server distributed machine learning is particularly important for training large-scale models.

[0027] However, currently, many distributed machine learning frameworks analyze the model structure manually, then partition the model, and maximize the communication efficiency through continuous experimental tuning. They lack auxiliary and visualization solutions for partitioning the model during distributed training, so it is impossible to intuitively understand the bottlenecks during model training and design parallel strategies efficiently, resulting in a decline in the training efficiency of the model.

[0028] To solve the problem that the current parallel strategy design cannot be carried out efficiently, resulting in a decline in the training efficiency of the model, the embodiments of the present application provide a parallel strategy design method, an electronic device, a storage medium, and a computer program product. The method is applied to a parallel strategy design system, which includes a first module and a second module. The method includes: visually processing the first parallel strategy of a preset model through the first module; where the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visualize the splitting method or resource occupancy information of the preset model; after determining the second parallel strategy, training the preset model through the second module, the second parallel strategy, and the training data set; where the second parallel strategy is part or all of the parallel strategies in the first parallel strategy. It can be seen that the first parallel strategy of the preset model can be visually processed through the first module, that is, the embodiments of the present application can visually process the data parallel strategy, and / or the tensor parallel strategy, and / or the pipeline parallel strategy of the preset model through the first module. After determining the second parallel strategy, the preset model can be trained through the second module, the second parallel strategy, and the training data set. That is, the embodiments of the present application visually process the first parallel strategy, so that the type of the parallel strategy of the preset model can be intuitively understood, and the second parallel strategy can be determined based on different parallel strategies, realizing efficient design of the parallel strategy. Furthermore, the preset model can be trained through the second module, the second parallel strategy, and the training data set, thereby improving the training efficiency of the model.

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application.

[0030] The embodiments of the present application provide a parallel strategy design method, which is applied to a parallel strategy design system. The parallel strategy design system includes a first module and a second module. Figure 1 Schematic diagram of the parallel strategy design method proposed in the embodiments of the present application Figure 1 As Figure 1 shown, the parallel strategy design method may include the following steps:

[0031] Step 101, visually process the first parallel strategy of a preset model through the first module; where the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visualize the splitting method or resource occupancy information of the preset model.

[0032] In the embodiments of the present application, the first module in the parallel strategy design system can visually process the first parallel strategy of the preset model.

[0033] It should be noted that, in the embodiments of the present application, the parallel policy design system may include a first module and a second module, or may include other modules. The present application does not specifically limit the number and types of modules included in the parallel policy design system.

[0034] It should be noted that, in the embodiments of the present application, the first module may be a model display mode module, which can be used to visualize the splitting method or resource occupancy information of a preset model. The present application does not specifically limit the type of the first module.

[0035] It should be noted that, in the embodiments of the present application, the second module may be a monitoring model training module, which can be used to start the training of a preset model and monitor the metric data generated during the training process of the preset model. The present application does not specifically limit the type of the second module.

[0036] It should be noted that, in the embodiments of the present application, the preset model may include deep learning frameworks such as Pytorch and TensorFlow. The present application does not specifically limit the number and types of models included in the preset model.

[0037] It should be noted that, in the embodiments of the present application, the first parallel policy includes a data parallel policy, and / or a tensor parallel policy, and / or a pipeline parallel policy. The present application does not specifically limit the number and types of parallel policies included in the first parallel policy.

[0038] It should be noted that, in the embodiments of the present application, the parallel policy system may further include a display canvas. The present application does not specifically limit the number and types of modules included in the parallel policy system.

[0039] Optionally, in the embodiments of the present application, Figure 2 is a schematic diagram of the parallel policy design method proposed in the embodiments of the present application Figure 2 , such as Figure 2 shown, before visualizing the first parallel policy of the preset model through the first module, that is, before step 101, the following steps may further be included:

[0040] Step 103: Visually present the structural information of the preset model through the display canvas.

[0041] It should be noted that, in the embodiments of the present application, the parallel policy system may further include a fourth module.

[0042] It should be noted that, in the embodiments of the present application, the fourth module may be a connection model training module, which can be used to obtain the preset model structure and extract the preset model structure information. The present application does not specifically limit the type of the fourth module.

[0043] Optionally, in an embodiment of the present application, when the structural information of a preset model is visualized through a display canvas, a fourth module can be used to obtain the preset model from a preset database, and extract model parameter information from the preset model; wherein the model parameter information includes one or more operators, input tensor information and output tensor information corresponding to the operators; then a directed graph can be determined based on the model parameter information; wherein the nodes in the directed graph are represented by operators, and the edges in the directed graph are represented by tensors; and then the directed graph can be visualized through the display canvas to present the structural information of the preset model.

[0044] It should be noted that, in the embodiments of the present application, the preset database may be a machine learning library or other databases, and the present application does not specifically limit the type of the preset database.

[0045] Exemplarily, in an embodiment of the present application, the fourth module is used to derive an application programming interface (API) from the intermediate representation (IR) of the model provided by the machine learning library to obtain a preset model structure.

[0046] It should be noted that, in the embodiments of the present application, the model parameter information may include one or more operators, input tensor information and output tensor information corresponding to the operators, and may also include other information. The present application does not specifically limit the amount and type of information included in the model parameter information.

[0047] It should be noted that in the embodiments of the present application, the operator is the basic computing unit of the deep learning algorithm. In the deep learning model, the operator corresponds to the computing logic in the network layer; for example, the convolution layer is an operator, and the present application does not specifically limit the type and number of operators.

[0048] It should be noted that in an embodiment of the present application, after obtaining a preset model from a preset database using the fourth module and extracting model parameter information from the preset model, a directed graph can be determined based on the model parameter information; and then the directed graph can be visualized through a display canvas to present the structural information of the preset model.

[0049] For example, in the embodiments of the present application, Figure 3 A schematic diagram of a directed graph structure proposed in an embodiment of the present application is shown in FIG. Figure 3As shown, the directed graph is visualized through a display canvas. In the canvas, the directed graph can be arranged from bottom to top. The bottom node represents the model input. Each operator is represented by a rectangle, and each tensor is represented by an oval. For the sake of simplicity, no connection lines are drawn between the operators and tensors. Only the identity number (id) of the input and output tensors is displayed when the user clicks on the operator rectangle.

[0050] That is to say, in the embodiments of the present application, the display canvas in the parallel strategy system can visualize the structural information of the preset model, so that the user can intuitively see the structure of the preset model, and then select a suitable parallel strategy for model training processing.

[0051] Step 104: Determine a first parallel strategy based on the structural information and hardware information of the preset model; wherein, the hardware information includes one or more first nodes, and each first node includes N GPUs, and N is a positive integer.

[0052] In the embodiments of the present application, after visualizing the structural information of the preset model through the display canvas, a first parallel strategy can be determined based on the structural information and hardware information of the preset model.

[0053] Optionally, in the embodiments of the present application, when the parallel strategy system determines the first parallel strategy based on the structural information and hardware information of the preset model, the N GPUs included in one or more first nodes can be converted to obtain a GPU grid of M×N; where the GPUs in the same column of the GPU grid belong to the same first node; then a first parallel strategy can be determined based on the GPU grid and the training data set.

[0054] It should be noted that, in the embodiments of the present application, the first node can be a server node, and the present application does not make specific limitations on the type of the first node.

[0055] Exemplarily, in the embodiments of the present application, assuming that the number of GPUs included in each first node is 16, then the 16 GPUs can be regarded as GPU grids of 2×8, 1×16, 4×4, 8×2, 16×1. The GPUs in the same column of the GPU grid belong to the same first node, that is, the GPUs in the same column belong to the GPUs in the same server. The present application does not make specific limitations on the size of the GPU grid.

[0056] Optionally, in the embodiments of the present application, when the parallel policy system determines the first parallel policy based on the GPU grid and the training dataset, for the first tensor of each data in the training dataset, the first target dimension of the first tensor can be determined based on the first identifier corresponding to the first tensor, and the tensor corresponding to the first target dimension can be copied to the target GPU in the GPU grid to generate a tensor parallel policy; alternatively, the second target dimension of the GPU grid can be determined based on the second identifier corresponding to the first tensor, and the first tensor can be divided based on the second target dimension to generate a tensor parallel policy.

[0057] Exemplarily, in the embodiments of the present application, each data in the training dataset can be represented in the form of a first tensor. For example, text data can be represented as a two-dimensional tensor, where rows represent words and columns represent features. The present application does not make specific limitations on the data types and the number of data included in the training dataset.

[0058] It should be noted that, in the embodiments of the present application, the first identifier is used to determine the target dimension for copying the first tensor. For example, assume that the first tensor is an N-dimensional tensor, denoted as X 0 X 1 , ……, X n-1 , X i ∈{S, R}, S represents dividing the i-th dimension, and R represents not dividing the i-th dimension. Assume that the first tensor is a three-dimensional tensor and the first identifier is SSS, which means copying all three dimensions. The present application does not make specific limitations on the type of the first identifier.

[0059] It should be noted that, in the embodiments of the present application, the second identifier can be used to determine the target dimension of the GPU grid. For example, if the target dimension of the GPU grid is the first dimension, then the tensor can be copied or divided in the first dimension of the GPU grid. The first dimension can represent the GPUs in the same column of the GPU grid. The present application does not make specific limitations on the type of the second identifier.

[0060] Exemplarily, in the embodiments of the present application, when the parallel policy system determines the first parallel policy based on the GPU grid and the training dataset, for the first tensor of each data in the training dataset, if the first identifier of the first tensor determines that the first target dimension is the first dimension and the second identifier determines that the second target dimension of the GPU grid is the first dimension, then the first dimension of the first tensor can be copied to the target GPU in the GPU grid, for example, copied to the first dimension of the GPU grid, that is, copied to the GPUs in the same column of the GPU grid, to generate a tensor parallel policy; or, the first dimension of the first tensor can be divided respectively in the first dimension of the GPU grid to generate a tensor parallel policy.

[0061] That is to say, in the embodiments of the present application, the parallel strategy system can copy or divide tensors based on the first identifier and the second identifier of the first tensor, so as to generate a tensor parallel strategy.

[0062] Optionally, in the embodiments of the present application, when the first tensor is the tensor corresponding to the input data, a data parallel strategy is generated; when the first tensor is the weight tensor and the output tensor, that is, a tensor parallel strategy. The above division method can represent both the tensor parallel strategy and the data parallel strategy at the same time.

[0063] Optionally, in the embodiments of the present application, a first flag can be added to each layer in the preset model to generate a pipeline parallel strategy; wherein, the first flag is used to indicate whether to perform layer-by-layer division on the preset model.

[0064] It should be noted that, in the embodiments of the present application, after the parallel strategy system generates the first parallel strategy, the first parallel strategy of the preset model can be visualized through the first module.

[0065] Optionally, in the embodiments of the present application, when visualizing the first parallel strategy of the preset model through the first module, when the display mode corresponding to the first module is the split mode, the first dimension of the GPU grid can be represented by the first preset color, and the second dimension of the GPU grid can be represented by the second preset color; wherein, the first dimension represents the GPUs in the same column of the GPU grid, and the second dimension represents the GPUs in different columns of the GPU grid; for the first tensor of each data in the training dataset, the first dimension of the first tensor can be represented by the first line style, and the second dimension of the first tensor can be represented by the second line style.

[0066] Exemplarily, in the embodiments of the present application, Figure 4 is the schematic diagram of the first preset color proposed by the embodiments of the present application. As Figure 4 shown, the first preset color can include red, yellow, and blue. Figure 4 The grid, dots, and slashes in it can represent the red, yellow, and blue colors respectively. The present application does not make specific limitations on the color types and the number of colors included in the first preset color.

[0067] Exemplarily, in the embodiments of the present application, Figure 5 is the schematic diagram of the second preset color proposed by the embodiments of the present application. As Figure 5 shown, the second preset color can include green, white, and purple. Figure 4 The gray, blank, and slashes in it can represent the green, white, and purple colors respectively. The present application does not make specific limitations on the color types and the number of colors included in the second preset color.

[0068] Exemplarily, in the embodiments of the present application, the first line style may be a horizontal stripe or other line styles, and the present application does not specifically limit the type of the first line style.

[0069] Exemplarily, in the embodiments of the present application, the second line style may be a vertical stripe, and the type of the second line style is different from that of the first line style. The present application does not specifically limit the type of the second line style.

[0070] That is to say, in the embodiments of the present application, the first dimension of the GPU grid can be characterized by a first preset color, that is, the GPUs in the same column of the GPU grid are characterized by the first preset color, which means that the tensors of the GPUs within the same server are divided; the second dimension of the GPU grid is characterized by a second preset color, that is, the GPUs in different columns of the GPU grid are characterized by the second preset color, which means that the tensors of the GPUs in different servers are divided. In this way, it can be intuitively known in which dimension of the GPU grid the tensor division or replication is performed through the type of the preset color.

[0071] Optionally, in the embodiments of the present application, for the first tensor of each data in the training dataset, the first dimension of the first tensor is characterized by a first line style, and the second dimension of the first tensor is characterized by a second line style. For example, Figure 6 is a schematic diagram of tensor dimension characterization proposed in the embodiments of the present application. As Figure 6 shown, the first dimension of the first tensor can be characterized by a horizontal stripe, and the second dimension of the first tensor can be characterized by a vertical stripe.

[0072] That is to say, in the embodiments of the present application, the dimensions of the GPU grid can be characterized by a first preset color and a second preset color, which can facilitate the user to intuitively judge in which dimension of the GPU grid the tensor division or replication is performed. The dimensions of the first tensor can also be characterized by a first line style and a second line style, which is convenient for the user to know which dimension of the first tensor is divided or replicated. For example, if it is a horizontal stripe, it can indicate that the division or replication is along the first dimension of the first tensor. If it is the first preset color, it can indicate that when dividing or replicating the first dimension of the first tensor, it is divided or replicated along the first dimension of the GPU grid. That is, the embodiments of the present application can visualize the first parallel strategy, enabling the user to intuitively perceive the tensor division strategy, so that an appropriate parallel strategy can be selected, greatly improving the design efficiency of the parallel strategy.

[0073] Optionally, in an embodiment of the present application, when visualizing the first parallel strategy of the preset model through the first module, the preset model can also be divided by layer, and the divided layers of the preset model are represented by the third line style.

[0074] It should be noted that, in an embodiment of the present application, the third line style can be a dashed line, and the present application does not specifically limit the type of the third line style.

[0075] Step 102: After determining the second parallel strategy, train the preset model through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy includes some or all of the parallel strategies in the first parallel strategy.

[0076] In an embodiment of the present application, the first module in the parallel strategy design system visualizes the first parallel strategy of the preset model. After determining the second parallel strategy, the preset model is trained through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy includes some or all of the parallel strategies in the first parallel strategy.

[0077] It should be noted that, in an embodiment of the present application, the second module can be a monitoring model training module, which can be used to start the training of the preset model and monitor the metric data generated during the training process of the preset model. The present application does not specifically limit the type of the second module.

[0078] Optionally, in an embodiment of the present application, when determining the second parallel strategy, it can be determined based on the first parallel strategy, that is, after the first parallel strategy is visualized, the user can select a suitable parallel strategy as the second parallel strategy. For example, the tensor parallel strategy in the first parallel strategy is used as the second parallel strategy. The second parallel strategy can include some or all of the parallel strategies in the first parallel strategy. The present application does not specifically limit the type and quantity of the parallel strategies included in the second parallel strategy.

[0079] Optionally, in an embodiment of the present application, when training the preset model through the second module, the second parallel strategy, and the training data set, the data information during the training process can be monitored; wherein, the data information includes one or more of the communication time of each layer in the preset model, the communication time of each operator in each layer, and the total time of each training iteration; and then the resource occupancy information can be generated based on the data information.

[0080] It should be noted that in the embodiments of the present application, the data information includes one or more of the communication time of each layer in the preset model, the communication time of each operator in each layer, and the total time of each training iteration, and may also include other information. The present application does not make specific limitations on the number and type of information included in the data information.

[0081] It should be noted that in the embodiments of the present application, the parallel strategy design system may further include a third module. The present application does not make specific limitations on the number and type of modules included in the parallel strategy design system.

[0082] It should be noted that in the embodiments of the present application, the third module may be an adjustment parallel division module, which can be used to perform adjustment processing on the parallel strategy. The present application does not make specific limitations on the type of the third module.

[0083] Optionally, in the embodiments of the present application, after the parallel strategy system trains the preset model based on the second parallel strategy and the training data set, that is, after generating the resource occupancy information, the first module can visualize the resource occupancy information, and the third module can update the second parallel strategy to obtain an updated third parallel strategy; wherein, the resource occupancy information is obtained by training the preset model with the second parallel strategy and the training data set; and then the preset model can be retrained with the third parallel strategy and the training data set.

[0084] Optionally, in the embodiments of the present application, when visualizing the resource occupancy information through the first module, when the display mode corresponding to the first module is the resource occupancy mode, the communication time of each operator can be screened to obtain the target communication time; wherein, the target communication time includes the maximum value in the communication time of the operator; then the communication time of each operator can be subjected to a first preset operation with the target communication time to obtain a first operation value; and then the magnitude of each first operation value can be represented by the saturation of a third preset color, that is, the magnitudes of different first operation values can be represented by different saturations of the third preset color.

[0085] It should be noted that in the embodiments of the present application, the third preset color may be green. The present application does not make specific limitations on the type of the third preset color.

[0086] Exemplarily, in the embodiments of the present application, when performing a first preset operation on the communication time of each operator and the target communication time to obtain a first operation value, the first operation value can be obtained through the following formula (1).

[0087]

[0088] wherein, T represents the first operation value, t iRepresents the communication time of the i-th operator, t max Represents the maximum value among the communication times of the operators.

[0089] Optionally, in the embodiments of the present application, when visualizing the resource occupancy information through the first module, the first metric value corresponding to each first tensor in the training dataset can also be determined; wherein, the first metric value is used to characterize the influence value of each first tensor on the communication duration of the operator; then, each first metric value can be screened to obtain the first target metric value; wherein, the first target metric value includes the maximum value among the first metric values; furthermore, each first metric value and the first target metric value can be subjected to a second preset operation to obtain a second operation value; finally, the magnitude of each second operation value can be represented by the saturation of a fourth preset color.

[0090] It should be noted that, in the embodiments of the present application, the fourth preset color can be red, and the present application does not make a specific limitation on the type of the fourth preset color.

[0091] It should be noted that, in the embodiments of the present application, the larger the first metric value, the greater the influence on the operator communication. For example, changing the partitioning method of a tensor can optimize the communication of one operator, but may deteriorate the communication of another operator. Therefore, it is not possible to make all the first metric values corresponding to the first tensors take the minimum value, and it can be subjectively judged by the user.

[0092] Exemplarily, in the embodiments of the present application, when each first metric value and the first target metric value are subjected to a second preset operation to obtain a second operation value, the second operation value can be obtained through the following formula (2).

[0093]

[0094] Wherein, V represents the second operation value, v i Represents the first metric value corresponding to the i-th first tensor, v max Represents the first target metric value.

[0095] It should be noted that, in the embodiments of the present application, after visualizing the resource occupancy information through the first module, it can be determined whether to update the second parallel strategy based on the saturation of the third preset color and / or the saturation of the fourth preset color.

[0096] Optionally, in the embodiments of the present application, when the parallel strategy system updates the second parallel strategy through the third module, the second parallel strategy can be updated through the third module, the saturation of the third preset color, and / or the saturation of the fourth preset color.

[0097] Exemplarily, in an embodiment of the present application, for example, it is possible to determine whether to update the second parallel strategy based on the saturation of the third preset color and / or the saturation of the fourth preset color. In the case where the second parallel strategy needs to be adjusted, the third module can be used to initiate the adjustment of the second parallel strategy. For example, dragging a pipeline (i.e., the third line style) in the display canvas can change the division of pipeline parallelism, or after clicking on any tensor, its corresponding tensor division method can be selected.

[0098] That is to say, in an embodiment of the present application, by visualizing the resource occupancy information, the user can determine whether to update the second parallel strategy based on the saturation of the third preset color and / or the saturation of the fourth preset color. That is, the embodiments of the present application enable the user to intuitively feel the impact of the model division method on distributed training, so as to intuitively understand the bottleneck during model training and timely improve the parallel strategy, achieving the purpose of assisting in optimizing the parallel strategy.

[0099] It should be noted that, in an embodiment of the present application, the parallel strategy system may further include a fifth module.

[0100] It should be noted that, in an embodiment of the present application, the fifth module may be an exported model division module, which can be used to perform an export process on the updated third parallel strategy, so that subsequently, the preset model can be retrained based on the third parallel strategy and the training data set.

[0101] In summary, the first module in the parallel strategy design system can visualize the first parallel strategy of the preset model. For example, the dimensions of the GPU grid can be represented by the first preset color and the second preset color, which can facilitate users to intuitively judge in which dimension of the GPU grid the tensor is divided or copied. The dimensions of the first tensor can also be represented by the first line style and the second line style, which can help users know in which dimension of the first tensor the division or copying is performed. For example, if it is a horizontal stripe, it can indicate that the division or copying is performed along the first dimension of the first tensor. If it is the first preset color, it can indicate that when the division or copying is performed along the first dimension of the first tensor, it is along the first dimension of the GPU grid. That is, the embodiments of the present application can visualize the first parallel strategy, enabling users to intuitively perceive the tensor division strategy, so that appropriate parallel strategies can be selected, greatly improving the design efficiency of parallel strategies. The visualization of resource occupancy information can also be performed, enabling users to determine whether the second parallel strategy needs to be updated based on the saturation of the third preset color and / or the saturation of the fourth preset color. That is, the embodiments of the present application can enable users to intuitively feel the impact of the model division method on distributed training, so that the bottleneck during model training can be intuitively understood and the parallel strategy can be improved in a timely manner, achieving the purpose of assisting in optimizing the parallel strategy.

[0102] The embodiments of the present application provide a parallel strategy design method, which is applied to a parallel strategy design system. The parallel strategy design system includes a first module and a second module. The method includes: visualizing the first parallel strategy of the preset model through the first module; where the first parallel strategy includes a data parallel strategy, and / or, a tensor parallel strategy, and / or, a pipeline parallel strategy, and the first module is used to visualize the splitting method or resource occupancy information of the preset model; after determining the second parallel strategy, training the preset model through the second module, the second parallel strategy, and the training data set; where the second parallel strategy is part or all of the parallel strategies in the first parallel strategy. It can be seen that the first parallel strategy of the preset model can be visualized through the first module. That is, the embodiments of the present application can visualize the data parallel strategy, and / or, the tensor parallel strategy, and / or, the pipeline parallel strategy of the preset model through the first module. After determining the second parallel strategy, the preset model can be trained through the second module, the second parallel strategy, and the training data set. That is, the embodiments of the present application can visually understand the type of parallel strategy of the preset model by visualizing the first parallel strategy, and determine the second parallel strategy based on different parallel strategies, realizing efficient design of parallel strategies. Furthermore, the preset model can be trained through the second module, the second parallel strategy, and the training data set, thereby improving the training efficiency of the model.

[0103] Based on the above embodiments, another embodiment of the present application provides a parallel strategy design method. This method can assist users in designing parallel strategies for distributed machine learning by combining a new optimization algorithm with a graphical interface. The embodiment of the present application can provide an interactive graphical interface for users to display the architecture diagram of a neural network model (i.e., a preset model) and the model parallel partitioning strategy. In addition, an interface for receiving model training information is provided to read the detailed time consumption of each part of the current model (i.e., resource occupancy information). After calculation by the algorithm, a reasonable metric is obtained for each part of the computational graph to show the impact of the partial parallel strategy (i.e., the first parallel strategy) on the parallel efficiency. At the same time, an algorithm is designed to give a solution for improving the parallel partitioning strategy, prompting the user about the parts of the model parallel partitioning that can be improved and the improvement methods. The user can manually change the parallel strategy (i.e., the second parallel strategy) through these prompt messages, and the program backend interface applies the updated strategy (i.e., the third parallel strategy) to model training.

[0104] It should be noted that in the embodiments of the present application, Figure 7 Schematic diagram of the parallel strategy design method proposed in the embodiment of the present application Figure 3 As Figure 7 shown, the main contents included in the parallel strategy design method are as follows: (1) Visualization of the parallel strategy; (2) Calculation and visualization of communication resource occupancy metrics; (3) Algorithm for designing a solution to improve the parallel partitioning strategy.

[0105] It should be noted that in the embodiments of the present application, the acquisition and visualization of the parallel strategy can include the following steps: 1.1. Visualization interface layout, Figure 8 Schematic diagram of the visualization interface layout proposed in the embodiment of the present application. As Figure 8 shown, the embodiment of the present application can use javascript and html to write the front-end interface program. As Figure 8As shown in the figure, there are four buttons, an option box, and a model display canvas in the figure; Button 1 is for connecting to model training (i.e., the fourth module); Button 2 is for monitoring model training (i.e., the second module); Button 3 is for adjusting parallel partitioning (i.e., the third module); Button 4 is for exporting model partitioning (i.e., the fifth module), and the drop-down option box is for the model display mode (i.e., the first module), including options "Segmentation Method" and "Resource Occupancy". The model display canvas is used to display the model structure and resource occupancy. According to the different options of the model display mode (i.e., the first module), the model segmentation situation and the training resource occupancy of each operator of the model are respectively displayed; 1.2. Obtain hardware information: the number of server nodes (i.e., the first node), the number of GPUs per server, and the video memory size of each GPU; 1.3. Obtain the model IR and visualize it; when the model training server starts training, a listening port needs to be established. After the user presses Button 1 to connect to model training (i.e., the fourth module) on the visualization interface, the following information is obtained from the model training script (at this time, the model has not started training), (1) Obtain the model structure through the model IR export API provided by the machine learning library (i.e., the preset database); mainstream frameworks such as Pytorch and Tensorflow support the export of IR in the Open Neural Network Exchange (ONNX) format. In the embodiments of this application, it is assumed that this format is used. If other formats of computational graphs are required, corresponding APIs need to be provided at the backend of the visualization program to process these formats; (2) Extract useful information from the model IR. For example, the required information can be extracted after importing the IR in onnx in python and converted into a custom format convenient for subsequent visualization engines to use; (3) Layer the model; (4) Visualize the above model information.

[0106] It should be noted that in the embodiments of this application, when extracting useful information from the model IR, for example, it can be implemented through the following python code:

[0107] import onnx

[0108] onnx_model = onnx.load(“model_file.onnx”)

[0109] print(onnx_model.graph.input)

[0110] print(onnx_model.graph.node)

[0111] Among them, onnx_model.graph.node contains information (dimensions, IDs) of all operators in the model and the tensors of the inputs and outputs of each operator (i.e., model parameter information). The directed graph formed by these operators and tensors can be denoted as E is the set of edges. The i-th operator is denoted as a node in the graph v i ∈V, where V is the set of nodes. The i-th tensor is denoted as an edge in the graph e i ∈E.

[0112] It should be noted that in the embodiments of the present application, when the model is hierarchically divided, after obtaining the above information (i.e., model parameter information), the model needs to be simply hierarchically divided, and subsequent pipeline parallelism can be achieved by splitting these layers. First, note that the obtained graph is a directed graph and has an initial input tensor. The embodiments of the present application can start iterating from the initial input tensor, as Figure 3 shown. Since current machine learning models such as GPT have a large number of continuous and repetitive substructures (layers), the embodiments of the present application record these repetitive parts and allow users to fold / unfold these structures during visualization. For models based on Transformer such as GPT, etc., Figure 9 is the schematic diagram of the model structure proposed by the embodiments of the present application. As Figure 9 shown, some Encoder layers in the middle will output data to all subsequent layers called Decoder layers. These connections are not considered during hierarchical division, but are considered during subsequent communication.

[0113] It should be noted that in the embodiments of the present application, visualizing the above model information may include the following content. The model display canvas in the visualization layout displays the model structure. In the display canvas, this directed graph can be arranged from bottom to top. The bottom node represents the model input. Each operator is represented by a rectangle, and the content of the rectangle is the operator id generated by onnx, such as " / conv1 / Conv2". Each tensor is represented by an ellipse, and the content of the ellipse is the tensor id generated by onnx. For the sake of simplicity, no connection is made between the operator and the tensor, and only the ids of the input and output tensors are displayed when the user clicks on the operator rectangle. In addition, large rectangles can be used to enclose the operators and tensors belonging to the same layer. The model layer information is provided in the data of the previous step, and there is a certain spacing between the large rectangles. For the foldable blocks formed by repetitive substructures (layers), in the expanded state, the embodiments of the present application can make the large rectangles of these layers closely connected, and a button for performing the folding operation is displayed in the topmost large rectangle. In the folded state, only one layer of the repetitive substructure is displayed, and a button for performing the unfolding operation is displayed in the large rectangle.

[0114] That is to say, in the embodiments of the present application, the display canvas in the parallel strategy system can visualize the structural information of the preset model, so that the user can intuitively see the structure of the preset model, and then select an appropriate parallel strategy for model training processing.

[0115] It should be noted that in the embodiments of the present application, the acquisition and visualization of the parallel strategy may further include the following steps: 1.4. Obtain the model parallel strategy: The visualization module in the embodiments of the present application requires that the model training program can export and import the model partitioning scheme, and uniformly represent the partitioning schemes of data parallelism, tensor parallelism, and pipeline parallelism in the following data structure: (1) Two-dimensional GPU grid layout, regarding multiple GPUs as a two-dimensional grid. For example, for 2 first nodes, with 16 GPUs per node, it can be regarded as grids of 2×8, 1×16, 4×4, 8×2, 16×1. Among them, the GPUs in the same column are the GPUs within the same server. When the number of GPUs within a fixed machine and the total number of GPUs are fixed, the length and width of the two-dimensional grid view are also fixed; (2) Partition tensors. For each tensor (i.e., the first tensor), after removing the dimension of the batch size (the amount of data in a batch of training) of the same batch of data, there are generally 1 to 3 dimensions. For example, a vector is one-dimensional, a matrix is two-dimensional, and some tensors are matrices plus a channel dimension, which is three-dimensional. For the partitioning of such an N-dimensional tensor, it is denoted as X 0 X 1 ,..., X n-1 X i ∈{S, R}, where S represents partitioning the i-th dimension, and R represents not partitioning the i-th dimension (i.e., the first identifier), that is, copying the tensor of this dimension on each GPU; in addition, adding a superscript (i.e., the second identifier) to the GPU partitioning of these tensors indicates copying or partitioning only along a certain dimension of the GPU grid. For example, S 0 represents partitioning only along the first dimension of the GPU grid, and S 1 represents partitioning only along the second dimension of the GPU grid. When partitioning, it is necessary to ensure that the tensor is partitioned along at most one dimension.

[0116] Optionally, in the embodiments of the present application, the parallel strategy system can, for the first tensor of each data in the training dataset, determine the first target dimension of the first tensor based on the first identifier corresponding to the first tensor, and copy the tensor corresponding to the first target dimension to the target GPU in the GPU grid to generate a tensor parallel strategy; or, it can determine the second target dimension of the GPU grid based on the second identifier corresponding to the first tensor, and perform partitioning processing on the first tensor based on the second target dimension to generate a tensor parallel strategy.

[0117] Exemplarily, in the embodiments of the present application, each data in the training dataset can be represented in the form of a first tensor. For example, text data can be represented as a two-dimensional tensor, where rows represent words and columns represent features. The present application does not make specific limitations on the data types and the number of data included in the training dataset.

[0118] It should be noted that, in the embodiments of the present application, the first identifier is used to determine the target dimension for copying the first tensor. For example, assume that the first tensor is an N-dimensional tensor, denoted as X 0 X 1 , ……, X n-1 , X i ∈{S, R}, where S represents dividing the i-th dimension, and R represents not dividing the i-th dimension. Assume that the first tensor is a three-dimensional tensor and the first identifier is SSS, which means that all three dimensions are copied. The present application does not make specific limitations on the type of the first identifier.

[0119] It should be noted that, in the embodiments of the present application, the second identifier can be used to determine the target dimension of the GPU grid. For example, if the target dimension of the GPU grid is the first dimension, then tensor copying or partitioning can be performed in the first dimension of the GPU grid. The first dimension can represent the GPUs in the same column of the GPU grid. The present application does not make specific limitations on the type of the second identifier.

[0120] Exemplarily, in the embodiments of the present application, when the parallel policy system determines the first parallel policy based on the GPU grid and the training dataset, for the first tensor of each data in the training dataset, if the first identifier of the first tensor determines that the first target dimension is the first dimension, and the second identifier determines that the second target dimension of the GPU grid is the first dimension, then the first dimension of the first tensor can be copied to the target GPU in the GPU grid, for example, copied to the first dimension of the GPU grid, that is, copied to the GPUs in the same column of the GPU grid, to generate a tensor parallel policy; or, the first dimension of the first tensor can be divided respectively in the first dimension of the GPU grid to generate a tensor parallel policy.

[0121] That is to say, in the embodiments of the present application, the parallel policy system can perform tensor copying or partitioning based on the first identifier and the second identifier of the first tensor, so as to generate a tensor parallel policy.

[0122] Optionally, in the embodiments of the present application, when the first tensor is the tensor corresponding to the input data, a data parallel policy is generated; when the first tensor is the weight tensor and the output tensor, that is, a tensor parallel policy. The above partitioning method can represent both the tensor parallel policy and the data parallel policy at the same time.

[0123] Optionally, in the embodiments of the present application, a first flag may be added to each layer in the preset model to generate a pipeline parallelism strategy; wherein, the first flag is used to indicate whether to perform layer-by-layer partitioning on the preset model.

[0124] It should be noted that, in the embodiments of the present application, for each operator, since the partitioning methods of the input and output tensors and the GPU grid are already given, the parallel communication strategy of the operator can also be uniquely determined, such as S 0 The identity operator from S 0 to RR communicates on the first dimension of the GPU grid using all-gather (a collective communication primitive), and RS 1 The matrix multiplication from R 0 to RS is an all-reduce (a collective communication primitive) communication across the entire GPU grid.

[0125] It should be noted that, in the embodiments of the present application, the acquisition and visualization of the parallelism strategy may further include the following steps: 1.5. Visualize the model parallelism strategy: When the drop-down option box of the model display mode (i.e., the first module) is "partitioning method", that is, the tensor and pipeline partitioning methods are displayed. The embodiments of the present application can visualize the above-mentioned tensor and pipeline partitioning; specifically, it includes the following contents: (1) Visualize the tensor partitioning. Use horizontal stripes (i.e., the first line style) or vertical stripes (i.e., the second line style) to represent the GPU tensor partitioning, where the internal GPU tensor partitioning is represented by red, yellow, and blue (i.e., the first preset color), the inter-GPU tensor partitioning is represented by green, white, and purple (i.e., the second preset color), and the replicated tensors on each GPU are represented by gray squares; for the input tensor and the activation tensor, the partitioning can be mainly along the first dimension of the matrix (i.e., the first dimension of the first tensor) or the second dimension of the matrix (i.e., the second dimension of the first tensor). For example, horizontal stripes can be used to represent the partitioning of the first dimension of the first tensor, and vertical stripes can be used to represent the partitioning of the second dimension of the first tensor; for the weight tensor, horizontal stripes can be used to represent the partitioning along the first dimension of the matrix, and vertical stripes can be used to represent the second dimension of the matrix; (2) Visualize the model partitioning strategy of pipeline parallelism. Pipeline parallelism divides a model into multiple parts by layer, and a dashed line (i.e., the third line style) is used to separate these pipeline stages during visualization.

[0126] That is to say, in the embodiments of the present application, when visualizing the first parallel strategy of the preset model through the first module, the first dimension of the GPU grid can be characterized by the first preset color, and the second dimension of the GPU grid can be characterized by the second preset color; wherein, the first dimension represents the GPUs in the same column of the GPU grid, and the second dimension represents the GPUs in different columns of the GPU grid; for the first tensor of each data in the training dataset, the first dimension of the first tensor can be characterized by the first line style, and the second dimension of the first tensor can be characterized by the second line style.

[0127] Exemplarily, in the embodiments of the present application, as Figure 4 shown, the first preset color may include red, yellow, and blue, Figure 4 wherein the grid, dots, and diagonal lines in

[0128] Exemplarily, in the embodiments of the present application, as Figure 5 shown, the second preset color may include green, white, and purple, Figure 4 wherein the gray, blank, and diagonal lines in

[0129] Exemplarily, in the embodiments of the present application, the first line style may be horizontal stripes, or other line styles, and the present application does not specifically limit the type of the first line style.

[0130] Exemplarily, in the embodiments of the present application, the second line style may be vertical stripes, and the second line style is different from the type of the first line style. The present application does not specifically limit the type of the second line style.

[0131] That is to say, in the embodiments of the present application, the first dimension of the GPU grid can be characterized by the first preset color, that is, the GPUs in the same column of the GPU grid are characterized by the first preset color, that is, it represents the division of tensors by the GPUs within the same server; the second dimension of the GPU grid is characterized by the second preset color, that is, the GPUs in different columns of the GPU grid are characterized by the second preset color, that is, it represents the division of tensors by the GPUs in different servers. In this way, it is possible to intuitively know in which dimension of the GPU grid the tensor is divided or copied through the type of the preset color.

[0132] It should be noted that in the embodiments of the present application, the calculation and visualization of communication resource occupancy metrics may include the following steps: 2.1. Monitor the calculation time and communication time: After the user presses button 2 to monitor model training (i.e., the second module), the model starts training and exports the monitoring information to the backend of the visualization program. The following monitoring data assumes that the computational graph corresponding to the neural network is static and does not change with the training input. This assumption is reasonable because current mainstream neural networks, including various large models, are all such static models. Monitoring the calculation time and communication time (i.e., data information) may include the following: (1) Select the API of the evaluation tool of the deep learning framework to track the calculation metrics. For example, pytorch profiler can monitor the time used for each layer, and MXNet can monitor the time used for each operator; (2) Monitor the following during neural network training. The following monitored values can be taken as the average value after multiple batches of model training: The total calculation time T of each training iteration n , where n is the iteration number, the calculation time C of each layer k , where k is the layer number, and the communication time t of each operator i , where i is the operator number; (3) Store the calculation metrics (i.e., data information) in a structured log or database, index them by layer and operator, and output the data to the visualization program (i.e., the front-end interface program) after the monitoring is completed; 2.2. Calculate the communication resource occupancy metrics: Some unreasonable matrix partitions will greatly increase the communication volume, resulting in a large amount of computing power waste during the waiting for communication to arrive on the GPU. However, the quality of communication strategies cannot be directly compared by comparing the communication volumes at different operators because the communication volume of some operators is naturally higher than that of other operators. The quality of a tensor partitioning pattern needs to consider the communication volumes brought by all surrounding tensors under different partitions, rather than simply comparing the different partitions of a single tensor. In fact, the matrix partitioning method is equivalent to an element in a multi-dimensional space, and each dimension has only 4 discrete values S 0 、S 1 、R 0 、R 1 . An n-order tensor occupies n dimensions. In the embodiments of the present application, when evaluating the communication resource occupancy metrics of a tensor partition, the following factors can be comprehensively considered: (1) The ratio of the communication volume of the tensor under the current partition to the minimum of the communication volumes under other partitions; (2) The impact of the tensor partition on other operators. It is impossible to globally evaluate the computational volume of these two metrics. It can be intuitively seen that the farther tensor partitions have less impact on the communication of operators. Therefore, in the embodiments of the present application, the global problem can be truncated. Suppose there are I tensors in total. For the i-th tensor, the following local optimization problem is defined. Given an integer k, let split i,k be the partitioning method of the k tensors before and after the i-th tensor, J(I, split i,k)Take the partitioning of these 2k + 1 tensors as split i,k , when the partitioning of other tensors remains unchanged, the communication consumption related to these 2k + 1 tensors, the optimization problem is the following formula (3).

[0133]

[0134] where J(I, split i,k ) represents the partitioning method of 2k + 1 tensors, split i,k represents the partitioning method of the k tensors before and after the i-th tensor, and J * (i, k) represents the minimum time consumption of the communication consumption related to 2k + 1 tensors.

[0135] It should be noted that in the embodiments of the present application, the above formula (3) is a small-scale integer programming problem, and the entire solution space can be directly traversed by exhaustive search for solution. k can be taken as a suitable value according to the computing power, for example, k = 8.

[0136] It should be noted that in the embodiments of the present application, for the i-th tensor, the embodiments of the present application denote v i as the communication strategy index determination value (i.e., the first index value) of this tensor. The larger the value, the greater the impact on communication. For each tensor tensor i , denote O(tensor i ) as the set of all operators whose input or output contains tensor i . Additionally, let the function t(op, split i ) be the communication volume when the operator op changes to the split i partitioning method for the i-th tensor and the partitioning methods of other tensors remain unchanged. The embodiments of the present application can estimate the communication resource occupancy using the following algorithm.

[0137]

[0138] It should be noted that in the embodiments of the present application, all v i cannot be simultaneously set to the minimum value v min in the above algorithm because the partitioning of each tensor affects multiple operators. Changing the tensor to optimize the communication of one operator may deteriorate the communication of another operator.

[0139] It should be noted that in the embodiments of the present application, the data information includes one or more of the communication time of each layer in the preset model, the communication time of each operator in each layer, and the total time of each training iteration, and may also include other information. The present application does not make specific limitations on the quantity and type of information included in the data information.

[0140] It should be noted that in the embodiments of the present application, the calculation and visualization of communication resource occupancy metrics may further include the following steps: 2.3. Visualize the communication resource occupancy: When the drop-down option box of the model display mode (i.e., the first module) is "resource occupancy", it represents the communication resource occupancy. For each communication policy determination value v of the tensors obtained above i (i.e., the first metric), let the maximum value among them be v max (i.e., the first target metric value), and take v i / v max as the saturation of the color of the graph corresponding to the i-th tensor (i.e., the fourth preset color), and the hue of all tensors is taken as red. For the communication time of each operator obtained by monitoring, let the maximum value among them be t max , and take t i / t max as the saturation of the color of the graph corresponding to the i-th operator (i.e., the third preset color), and the hue of all operators is taken as green.

[0141] That is to say, in the embodiments of the present application, when visualizing the resource occupancy information through the first module, the communication time of each operator can be screened to obtain the target communication time; wherein, the target communication time includes the maximum value among the communication times of the operators; then the communication time of each operator can be subjected to a first preset operation with the target communication time to obtain a first operation value; furthermore, the magnitude of each first operation value can be characterized by the saturation of the third preset color, that is, the magnitudes of different first operation values can be represented by different saturations of the third preset color.

[0142] It should be noted that in the embodiments of the present application, the third preset color may be green, and the present application does not make specific limitations on the type of the third preset color.

[0143] Exemplarily, in the embodiments of the present application, when the communication time of each operator is subjected to a first preset operation with the target communication time to obtain a first operation value, the first operation value can be obtained through the above formula (1).

[0144] Optionally, in the embodiments of the present application, when visualizing the resource occupancy information through the first module, the first metric value corresponding to each first tensor in the training dataset can also be determined; wherein, the first metric value is used to characterize the influence value of each first tensor on the communication duration of the operator; then the first metric values can be screened to obtain the first target metric value; wherein, the first target metric value includes the maximum value among the first metric values; furthermore, the first metric value of each can be subjected to a second preset operation with the first target metric value to obtain a second operation value; finally, the magnitude of each second operation value can be characterized by the saturation of the fourth preset color.

[0145] It should be noted that in the embodiments of the present application, the fourth preset color may be red, and the present application does not specifically limit the type of the fourth preset color.

[0146] It should be noted that in the embodiments of the present application, the larger the first index value, the greater the impact on operator communication. For example, changing the partitioning method of a tensor can optimize the communication of one operator, but may deteriorate the communication of another operator. Therefore, it is not possible to make the first index values corresponding to all the first tensors take the minimum value, and it can be subjectively judged by the user.

[0147] Exemplarily, in the embodiments of the present application, when performing a second preset operation on each first index value and a first target index value to obtain a second operation value, the second operation value can be obtained through the above formula (2).

[0148] It should be noted that in the embodiments of the present application, the design algorithm for improving the parallel partitioning strategy scheme may include the following: 3.1. Algorithm idea: The algorithm idea is similar to the evaluation of resource occupancy indicators. For a fixed integer k (such as 8), by calculating the solution of the optimization problem to find the optimization strategy. The embodiments of the present application can use the greedy strategy, and each time find the split with the greatest impact i,k Fix it as the new partitioning strategy for these tensors, and then optimize the remaining tensors. Such a cycle can specifically include the following steps: (1) Solve the optimization problem for the i-th tensor that needs to be optimized (The fixed part is not included in the split i,k ); (2) Take the tensor with the largest ratio J(I, split i,k ) / J * (I, split i,k ), and make the partitioning of these tensors no longer participate in the optimization in the future; (3) Repeat steps 1 and 2 until all tensor partitions are completed. The user can consider choosing to use the partitioning results obtained by some algorithms based on the observation of the model.

[0149] It should be noted that in the embodiments of the present application, the design algorithm for improving the parallel partitioning strategy solution may further include the following: 3.2 User interaction and adjustment of the model parallel strategy: After the user clicks the button 3 to adjust the parallel partitioning (i.e., the third module), the model parallel strategy (i.e., the second parallel strategy) can be adjusted. The user can perform the following operations: (1) After clicking any tensor, the corresponding tensor partitioning method can be selected; (2) Dragging the pipeline parallel partitioning line can change the partitioning of the pipeline parallel. If the video memory occupancy of a certain pipeline stage exceeds the maximum value after dragging, the dragging is prohibited; 3.3 Output the updated parallel strategy through the API / API design: After the user interacts and adjusts the model parallel strategy, clicking the export model partitioning (i.e., the fifth module) can export the parallel strategy; 3.4 Apply the updated parallel strategy: After receiving the updated parallel strategy (i.e., the third parallel strategy), the model training script can re-perform the training task by dynamically importing the parallel strategy.

[0150] Exemplarily, in the embodiments of the present application, for example, it can be determined whether it is necessary to update the second parallel strategy based on the saturation of the third preset color and / or the saturation of the fourth preset color. In the case where it is necessary to adjust the second parallel strategy, the adjustment of the second parallel strategy can be enabled through the third module. For example, dragging the pipeline (i.e., the third line style) in the display canvas can change the partitioning of the pipeline parallel, or after clicking any tensor, the corresponding tensor partitioning method can be selected.

[0151] That is to say, in the embodiments of the present application, by visualizing the resource occupancy information, the user can determine whether it is necessary to update the second parallel strategy based on the saturation of the third preset color and / or the saturation of the fourth preset color. That is, the embodiments of the present application can enable the user to intuitively feel the impact of the model partitioning method on distributed training, so as to intuitively understand the bottleneck during model training and timely improve the parallel strategy, achieving the purpose of assisting in optimizing the parallel strategy.

[0152] In summary, the first module in the parallel strategy design system can visualize the first parallel strategy of the preset model. For example, the dimensions of the GPU grid can be represented by the first preset color and the second preset color, which can facilitate users to intuitively judge in which dimension of the GPU grid the tensor is divided or copied. The dimensions of the first tensor can also be represented by the first line style and the second line style, which can help users know in which dimension of the first tensor the division or copy is performed. For example, if it is a horizontal stripe, it can indicate that the division or copy is performed along the first dimension of the first tensor. If it is the first preset color, it can indicate that when dividing or copying along the first dimension of the first tensor, the division or copy is performed along the first dimension of the GPU grid. That is, the embodiments of the present application can visualize the first parallel strategy and allow users to intuitively perceive the tensor division strategy, so that appropriate parallel strategies can be selected, greatly improving the design efficiency of parallel strategies. The visualization of resource occupancy information can also be performed, enabling users to determine whether the second parallel strategy needs to be updated based on the saturation of the third preset color and / or the saturation of the fourth preset color. That is, the embodiments of the present application can allow users to intuitively feel the impact of the model division method on distributed training, so as to intuitively understand the bottleneck during model training and promptly improve the parallel strategy, achieving the purpose of assisting in optimizing the parallel strategy.

[0153] The embodiments of the present application provide a parallel strategy design method, which is applied to a parallel strategy design system. The parallel strategy design system includes a first module and a second module. The method includes: visualizing the first parallel strategy of the preset model through the first module; where the first parallel strategy includes a data parallel strategy, and / or, a tensor parallel strategy, and / or, a pipeline parallel strategy, and the first module is used to visualize the splitting method or resource occupancy information of the preset model; after determining the second parallel strategy, training the preset model through the second module, the second parallel strategy, and the training data set; where the second parallel strategy is part or all of the parallel strategies in the first parallel strategy. It can be seen that the first parallel strategy of the preset model can be visualized through the first module, that is, the embodiments of the present application can visualize the data parallel strategy, and / or, the tensor parallel strategy, and / or, the pipeline parallel strategy of the preset model through the first module. After determining the second parallel strategy, the preset model can be trained through the second module, the second parallel strategy, and the training data set. That is, the embodiments of the present application can visually understand the type of the parallel strategy of the preset model by visualizing the first parallel strategy, and determine the second parallel strategy based on different parallel strategies, realizing the efficient design of parallel strategies. Furthermore, the preset model can be trained through the second module, the second parallel strategy, and the training data set, thereby improving the training efficiency of the model.

[0154] Based on the above embodiments, an embodiment of the present application provides a parallel strategy design system. Figure 10 It is a schematic diagram of the composition structure of an electronic device, as Figure 10 shown, the electronic device 10 proposed in the embodiment of the present application may further include a processor 11, a memory 12 storing executable instructions of the processor 11. Further, the electronic device 10 may further include a communication interface 13, and a bus 14 for connecting the processor 11, the memory 12, and the communication interface 13.

[0155] In the embodiment of the present application, the above-mentioned processor 11 may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above processor functions may be others, and the embodiments of the present application do not make specific limitations. The electronic device 10 may further include a memory 12, and the memory 12 may be connected to the processor 11. Among them, the memory 12 is used to store executable program codes, and the program codes include computer operation instructions. The memory 12 may include a high-speed RAM memory, and may also include non-volatile memory, for example, at least two disk memories.

[0156] In the embodiment of the present application, the bus 14 is used to connect the communication interface 13, the processor 11, and the memory 12 and for mutual communication between these devices.

[0157] In the embodiment of the present application, the memory 12 is used to store instructions and data.

[0158] Further, in the embodiments of the present application, the above-mentioned processor 11 is configured to visualize the first parallel strategy of the preset model through the first module; wherein, the first parallel strategy includes a data parallel strategy, and / or, a tensor parallel strategy, and / or, a pipeline parallel strategy, and the first module is configured to visualize the splitting method or resource occupancy information of the preset model; after determining the second parallel strategy, the preset model is trained through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy includes some or all of the parallel strategies in the first parallel strategy.

[0159] In practical applications, the above-mentioned memory 12 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or a combination of the above types of memories, and provides instructions and data to the processor 11.

[0160] The embodiments of the present application provide a parallel strategy design system, which includes a first module and a second module. The first parallel strategy of the preset model is visualized through the first module; wherein, the first parallel strategy includes a data parallel strategy, and / or, a tensor parallel strategy, and / or, a pipeline parallel strategy, and the first module is configured to visualize the splitting method or resource occupancy information of the preset model; after determining the second parallel strategy, the preset model is trained through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy is some or all of the parallel strategies in the first parallel strategy. It can be seen that the first parallel strategy of the preset model can be visualized through the first module, that is, the embodiments of the present application can visualize the data parallel strategy, and / or, the tensor parallel strategy, and / or, the pipeline parallel strategy of the preset model through the first module. After determining the second parallel strategy, the preset model can be trained through the second module, the second parallel strategy, and the training data set. That is, the embodiments of the present application visualize the first parallel strategy, so that the type of the parallel strategy of the preset model can be intuitively understood, and the second parallel strategy can be determined based on different parallel strategies, realizing the efficient design of the parallel strategy. Furthermore, the preset model can be trained through the second module, the second parallel strategy, and the training data set, thereby improving the training efficiency of the model.

[0161] An embodiment of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the parallel strategy design method described above is implemented.

[0162] Specifically, the program instructions corresponding to a parallel strategy design method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to a parallel strategy design method in the storage medium are read or executed by an electronic device, the following steps are included:

[0163] Visualize the first parallel strategy of the preset model through the first module; wherein, the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visualize the segmentation method or resource occupancy information of the preset model;

[0164] After determining the second parallel strategy, train the preset model through the second module, the second parallel strategy, and the training data set; wherein, the second parallel strategy includes some or all of the parallel strategies in the first parallel strategy.

[0165] An embodiment of the present application also provides a computer program product, including a computer program, and the computer program can be executed by the processor 11 of the electronic device 10 to complete the steps of any of the foregoing methods.

[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.

[0167] The present application is described with reference to the schematic flowcharts and / or block diagrams of the implementation processes of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the schematic flowcharts and / or block diagrams, and the combination of processes and / or blocks in the schematic flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the specified functions in one or more of the following processes or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the flowcharts and / or boxes Figure 1 of the flowchart or flowcharts and / or boxes Figure 1 specified.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the flowcharts and / or boxes Figure 1 of the flowchart or flowcharts and / or boxes Figure 1 specified.

[0170] As described above, only the preferred embodiments of this application are provided and are not intended to limit the scope of protection of this application.

Claims

1. A parallel strategy design method, characterized in that: The method is applied to a parallel strategy design system, the parallel strategy design system includes a first module and a second module, and the method includes: The first parallel strategy of the preset model is visualized through the first module; wherein the first parallel strategy includes a data parallel strategy, and / or a tensor parallel strategy, and / or a pipeline parallel strategy, and the first module is used to visualize the segmentation method or resource occupancy information of the preset model; After determining the second parallel strategy, the preset model is trained through the second module, the second parallel strategy and the training data set; wherein the second parallel strategy includes part or all of the parallel strategies in the first parallel strategy.

2. The method according to claim 1, characterized in that The parallel strategy design system further includes a third module. After the preset model is trained based on the second parallel strategy and the training data set, the method further includes: The resource occupancy information is visualized through the first module, and the second parallel strategy is updated through the third module to obtain an updated third parallel strategy; wherein the resource occupancy information is obtained by training the preset model through the second parallel strategy and the training data set; The preset model is retrained using the third parallel strategy and the training data set.

3. The method according to claim 1, characterized in that The parallel strategy system further includes a display canvas. Before visualizing the first parallel strategy of the preset model through the first module, the method further includes: Visually presenting the structural information of the preset model through the display canvas; The first parallel strategy is determined based on the structural information and hardware information of the preset model; wherein the hardware information includes one or more first nodes, each of which includes N graphics processors GPU, and N is a positive integer.

4. The method according to claim 3, characterized in that The parallel strategy system further includes a fourth module, wherein the visual presentation of the structural information of the preset model through the display canvas includes: The fourth module is used to obtain the preset model from a preset database, and to extract model parameter information from the preset model; wherein the model parameter information includes one or more operators, input tensor information corresponding to the operators, and output tensor information; Determining a directed graph based on the model parameter information; wherein the nodes in the directed graph are represented by the operator, and the edges in the directed graph are represented by tensors; The directed graph is visualized through the display canvas to present structural information of the preset model.

5. The method according to claim 3, characterized in that: The determining the first parallel strategy based on the structure information and hardware information of the preset model includes: Convert the N GPUs included in the one or more first nodes to obtain an M×N GPU grid; wherein the GPUs in the same column of the GPU grid belong to the same first node; The first parallel strategy is determined based on the GPU grid and the training data set.

6. The method according to claim 5, characterized in that The determining the first parallel strategy based on the GPU grid and the training data set includes: For a first tensor of each data in the training data set, determine a first target dimension of the first tensor based on a first identifier corresponding to the first tensor, and copy the tensor corresponding to the first target dimension to a target GPU in the GPU grid to generate the tensor parallel strategy; or, A second target dimension of the GPU grid is determined based on a second identifier corresponding to the first tensor, and the first tensor is partitioned based on the second target dimension to generate the tensor parallel strategy.

7. The method according to claim 5, characterized in that The method further comprises: A first flag is added to each layer in the preset model to generate the pipeline parallel strategy; wherein the first flag is used to indicate whether the preset model is divided by layer.

8. The method according to claim 1, characterized in that: The visualizing the first parallel strategy of the preset model by the first module includes: Characterizing a first dimension of a GPU grid by a first preset color, and characterizing a second dimension of the GPU grid by a second preset color; wherein the first dimension represents GPUs in the same column of the GPU grid, and the second dimension represents GPUs in different columns of the GPU grid; For a first tensor of each data in the training data set, a first dimension of the first tensor is characterized by a first line style, and a second dimension of the first tensor is characterized by a second line style.

9. The method according to claim 2, characterized in that: The method further comprises: When the preset model is trained by the second module, the second parallel strategy and the training data set, data information in the training process is monitored; wherein the data information includes one or more of the communication time of each layer in the preset model, the communication time of each operator in each layer and the total time of each training iteration; The resource occupancy information is generated based on the data information.

10. The method according to claim 9, characterized in that The visualizing the resource occupancy information by the first module includes: The communication time of each operator is screened to obtain a target communication time; wherein the target communication time includes a maximum value of the communication time of the operators; Perform a first preset operation on the communication time of each operator and the target communication time to obtain a first operation value; The size of each of the first operation values ​​is represented by the saturation of the third preset color.

11. The method according to claim 2, characterized in that The visualizing the resource occupancy information by the first module includes: Determine a first index value corresponding to each first tensor in the training data set; wherein the first index value is used to characterize the impact value of each first tensor on the communication duration of the operator; Performing screening processing on each of the first indicator values ​​to obtain a first target indicator value; wherein the first target indicator value includes a maximum value among the first indicator values; Perform a second preset operation on each of the first indicator values ​​and the first target indicator value to obtain a second operation value; The size of each of the second operation values ​​is represented by the saturation of a fourth preset color.

12. The method according to claim 10 or 11, characterized in that: The updating process of the second parallel strategy by the third module includes: The second parallel strategy is updated by the third module, the saturation of the third preset color, and / or the saturation of the fourth preset color.

13. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; wherein, The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the method according to any one of claims 1 to 12 when running the computer program.

14. A computer-readable storage medium, characterized in that: The storage medium stores computer program codes, and when the computer program codes are executed by a computer, the method according to any one of claims 1 to 12 is executed.

15. A computer program product comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 12 when executed by a processor.