Model segmentation method and device and storage medium

By generating and filtering the segmentation parameters, and determining the target segmentation strategy in combination with simulation and real-machine evaluation, the problem of model segmentation relying on manual experience is solved, and the model segmentation effect and performance are optimized.

CN120353565APending Publication Date: 2025-07-22HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410081475.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the selection of model slicing parameters depends on manual experience, resulting in poor slicing effect, affecting model performance and increasing training/inference costs.

Method used

By generating slicing parameters under different slicing dimensions and filtering according to preset filtering rules, we find candidate slicing strategies that meet the verification conditions, and determine the target slicing strategies based on simulation and actual machine evaluation to reduce manual experience dependence.

Benefits of technology

Effectively improve model segmentation effect, optimize distributed deployment performance, reduce costs, and improve model performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353565A_ABST
    Figure CN120353565A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model segmentation method and device and a storage medium. According to the embodiment of the invention, the segmentation parameters under different segmentation dimensions can be generated for the target model, and the generated segmentation parameters are filtered according to the preset filtering rule, so that the residual segmentation parameters under each segmentation dimension are obviously reduced, and the number of segmentation strategies generated by mutual combination is also obviously reduced. In addition, candidate segmentation strategies conforming to verification conditions are searched from the combined segmentation strategies. Therefore, through multi-level filtering operation, low-quality segmentation strategies can be abandoned reasonably, so that the number of the reserved candidate segmentation strategies is more simplified. Performance evaluation is carried out on the target model under each candidate segmentation strategy, and an appropriate target segmentation strategy can be finally determined. Therefore, the appropriate target segmentation strategy can be efficiently determined for the target model, the appropriate target segmentation strategy is used as a basis for distributed deployment of the target model, manual experience does not need to be relied on any more, and therefore the model segmentation effect can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model splitting method, device, and storage medium. Background Art

[0002] With the development of large language models (LLMs), the scale of large language models has been continuously increasing. Therefore, the demand for distributed computing is also growing. The computing clusters used to host large language models often need to include thousands of GPUs to support the distributed training and distributed inference processes of large language models.

[0003] During the process of distributed deployment, the selection of splitting parameters is very important and will affect the performance of the training process or the inference process. Currently, it is necessary to rely on algorithm engineers to determine the splitting parameters based on experience, which is not only time-consuming and laborious, but also may lead to poor splitting effects due to unreasonable splitting parameters, thereby affecting the model performance. Summary of the Invention

[0004] Multiple aspects of this application provide a model splitting method, device, and storage medium to optimize the model splitting effect.

[0005] An embodiment of this application provides a model splitting method, and the method includes:

[0006] Filter the splitting parameters corresponding to the target model in each splitting dimension according to a preset filtering rule.

[0007] Find candidate splitting strategies that meet the verification conditions from the splitting strategies combined from the remaining splitting parameters in each splitting dimension.

[0008] Perform performance evaluation on the target model under each candidate splitting strategy to determine the target splitting strategy as the basis for distributed deployment of the target model.

[0009] Further, the method may further include:

[0010] Obtain the model parameters corresponding to the target model.

[0011] Query the hardware information corresponding to the target cluster, where the target cluster is used to host the target model.

[0012] Generate splitting parameters for the target model in each splitting dimension based on the model parameters and the hardware information.

[0013] Further, the splitting dimension includes a tensor parallel dimension, the splitting parameter under the tensor parallel dimension adopts a tensor parallelism degree, and the preset filtering rule under the tensor parallel dimension includes:

[0014] The tensor parallelism is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster;

[0015] The tensor parallelism can be divided evenly by the total number of attention heads corresponding to the target model; and / or

[0016] The tensor parallelism does not exceed the number of GPUs installed on a single node in the target cluster;

[0017] Wherein, the target cluster is used to host the target model.

[0018] Further, the partitioning dimension includes a pipeline parallelism dimension, and the partitioning parameter under the pipeline parallelism dimension adopts the pipeline parallelism. The preset filtering rules under the pipeline parallelism dimension include:

[0019] The pipeline parallelism is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; and / or

[0020] The pipeline parallelism can be divided evenly by the total number of model layers corresponding to the target model;

[0021] Wherein, the target cluster is used to host the target model.

[0022] Further, the partitioning dimension includes a data parallelism dimension, and the partitioning parameter under the data parallelism dimension adopts the data parallelism. The preset filtering rules under the data parallelism dimension include: the data parallelism does not exceed the ratio between the total memory of the target cluster and the memory occupied by the target model; and / or,

[0023] The partitioning dimension includes a single batch specification dimension, and the partitioning parameter under the single batch specification dimension adopts the number of samples in a single batch. The preset filtering rules under the single batch specification dimension include: the number of samples in a single batch does not exceed the empirical value supported by the target cluster and does not exceed the number of samples supported by the remaining memory. The remaining memory is the estimated memory remaining in the target cluster after deploying the target model;

[0024] Wherein, the target cluster is used to host the target model.

[0025] Further, the candidate partitioning strategies include tensor parallelism, pipeline parallelism, data parallelism, and / or the number of samples in a single batch. Finding candidate partitioning strategies that meet the verification conditions includes:

[0026] If the product of the tensor parallelism, pipeline parallelism, and data parallelism included in the candidate partitioning strategy is equal to the total number of GPUs that the target cluster can provide, it is determined that the verification conditions are met; and / or,

[0027] If the data parallelism included in the candidate segmentation strategy is divisible by the number of samples in a single batch, it is determined that the verification condition is met;

[0028] Among them, the target cluster is used to host the target model.

[0029] Furthermore, if the number of candidate segmentation strategies is N, the performance of the target model is evaluated under each candidate segmentation strategy to determine the target segmentation strategy, including:

[0030] Under the N candidate segmentation strategies, calculate the corresponding simulation evaluation results of the target model on the target cluster respectively;

[0031] According to the simulation evaluation results, screen K candidate segmentation strategies to be verified from the N candidate segmentation strategies, where K < N;

[0032] Deploy the target model on the target cluster according to the K candidate segmentation strategies to be verified to generate real machine evaluation results;

[0033] According to the real machine evaluation results, determine the target segmentation strategy from the K candidate segmentation strategies to be verified.

[0034] Furthermore, under the N candidate segmentation strategies, calculate the corresponding simulation evaluation results of the target model on the target cluster respectively, including:

[0035] Under the first candidate segmentation strategy, simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster as the simulation evaluation result corresponding to the target model under the first candidate segmentation strategy;

[0036] Among them, the performance overhead includes computing overhead, communication overhead, and / or memory overhead, and the first candidate segmentation strategy is any candidate segmentation strategy.

[0037] Furthermore, simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster, including:

[0038] Taking the unit sample received by a single model layer hosted on the target process as the simulation evaluation unit, simulate and evaluate the corresponding unit performance overhead of the target process, where the target process is any process on the simulated target cluster used to host the target model;

[0039] According to each segmentation parameter in the first candidate segmentation strategy, determine the simulation evaluation coefficient corresponding to the target process;

[0040] Calculate the product of the simulation evaluation coefficient and the unit performance overhead as the performance overhead corresponding to the target process;

[0041] Estimate the performance overhead corresponding to the backbone network according to the performance overhead calculated for each process used to carry the backbone network.

[0042] Further, simulate and evaluate the unit performance overhead corresponding to the target process, including:

[0043] For the inference stage, simulate and evaluate the performance overhead corresponding to the single-round end-to-end propagation of the unit sample as the unit performance overhead corresponding to the target process;

[0044] For the training stage, simulate and evaluate the total performance overhead corresponding to the multi-round end-to-end propagation and other training-required links of the unit sample during training as the unit performance overhead corresponding to the target process.

[0045] Further, estimate the performance overhead corresponding to the backbone network according to the performance overhead calculated for each process used to carry the backbone network, including:

[0046] Select representative processes according to the process time consumed in the performance overhead corresponding to each process;

[0047] Take the performance overhead corresponding to the representative process as the performance overhead corresponding to the backbone network.

[0048] Further, according to the simulation evaluation results, select K candidate partitioning strategies to be verified from the N candidate partitioning strategies, including:

[0049] Filter out the candidate partitioning strategies estimated to have memory overflow problems from the N candidate partitioning strategies according to the memory overhead in the simulation evaluation results;

[0050] Based on the simulation evaluation results, calculate the evaluation index corresponding to each of the remaining candidate partitioning strategies;

[0051] Sort the remaining candidate partitioning strategies according to the evaluation index to select the K candidate partitioning strategies to be verified;

[0052] Among them, the evaluation index includes the sampling rate.

[0053] Further, perform distributed deployment of the target model on the target cluster according to the K candidate partitioning strategies to be verified to generate on-machine evaluation results, including:

[0054] Call the model execution framework and input the partitioning parameters in the K candidate partitioning strategies into the model execution framework respectively;

[0055] Use the model execution framework to distribute and deploy the target model to the target cluster according to the K candidate partitioning strategies respectively;

[0056] Perform on-machine evaluation operations respectively under the K to-be-verified segmentation strategies to generate on-machine evaluation results.

[0057] Further, the method may further include:

[0058] Extract, from the invoked model execution framework, the configuration file generated when the target model is distributedly deployed to the target cluster according to the target segmentation strategy;

[0059] Save the configuration file.

[0060] An embodiment of the present application further provides a computing device, including a memory, a processor, and a communication component;

[0061] The memory is used to store one or more computer instructions;

[0062] The processor is coupled with the memory and the communication component and is used to execute one or more computer instructions to execute the foregoing model segmentation method.

[0063] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing model segmentation method.

[0064] In an embodiment of the present application, before distributedly deploying a target model, segmentation parameters under different segmentation dimensions may be generated for the target model, and the generated segmentation parameters are filtered according to a preset filtering rule, which can effectively reduce the number of segmentation parameters under each segmentation dimension, and the number of segmentation strategies generated by combining the remaining segmentation parameters under each segmentation dimension will also be significantly reduced. Moreover, candidate segmentation strategies that meet the verification conditions are searched from the combined segmentation strategies. In this way, through multi-level filtering operations, low-quality segmentation strategies can be reasonably discarded, so that the number of remaining candidate segmentation strategies is more streamlined. On this basis, the performance of the target model can be evaluated under each candidate segmentation strategy to finally determine a suitable target segmentation strategy. Accordingly, a suitable target segmentation strategy can be efficiently determined for the target model and used as the basis for distributedly deploying the target model, without relying on manual experience, thereby effectively improving the model segmentation effect. Description of the Drawings

[0065] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0066] Figure 1Schematic flowchart of a model slicing method provided by an exemplary embodiment of the present application;

[0067] Figure 2 Logical schematic diagram of a model slicing method provided by an exemplary embodiment of the present application;

[0068] Figure 3 Schematic flowchart of another model slicing method provided by an exemplary embodiment of the present application;

[0069] Figure 4 Schematic diagram of an exemplary model slicing principle provided by an exemplary embodiment of the present application;

[0070] Figure 5 Schematic diagram of the structure of a computing device provided by another exemplary embodiment of the present application. Detailed implementation manners

[0071] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0072] Before starting to elaborate on the technical solutions provided by the embodiments of the present application, several technical concepts involved in the present application are briefly explained as follows.

[0073] Large model: It can refer to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc. It should be noted that when a large model is actually applied, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, large models can be widely applied in the fields of natural language processing (NLP), computer vision, speech processing, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc.

[0074] Distributed deployment: It can be understood that after the model is sliced, each sliced part of the model is carried by a computing cluster. A typical cluster for distributed deployment is a GPU cluster. After completing the distributed deployment, when receiving an inference or training task, the task can be split into parallel subtasks to process the task in a distributed manner, thereby improving the task processing efficiency.

[0075] Model execution framework: It can be understood as a tool for deploying and executing models. The model execution framework usually includes a distributed initialization link. In this link, the model execution framework will slice the model according to the input slicing parameters and deploy it to the parallel processes in the cluster in a distributed manner. There are many model execution frameworks in the current field, such as Megatron, etc., and no more product examples will be given here.

[0076] As introduced in the background technology, currently, professional personnel such as algorithm engineers need to input slicing parameters into the model execution framework according to experience, and the model execution framework will slice and deploy the model accordingly. However, this model slicing method relying on manual experience results in unsatisfactory model slicing effects, leading to poor model performance and high model training / inference costs.

[0077] To this end, the embodiments of the present application provide an innovative model splitting method, which proposes to automatically determine a suitable splitting strategy for the model before distributed deployment of the model, as the basis for distributed deployment of the model execution framework, without relying on manual experience any more.

[0078] The following will describe in detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.

[0079] Figure 1 FIG. is a schematic flowchart of a model splitting method provided by an exemplary embodiment of the present application. This method can be executed by a data processing device, which can be implemented as software, hardware, or a combination of software and hardware, and the data processing device can be integrated in a computing device. Refer to Figure 1 and the method may include:

[0080] Step 100: Filter the splitting parameters corresponding to the target model in each splitting dimension according to a preset filtering rule respectively;

[0081] Step 101: Search for candidate splitting strategies that meet the verification conditions from the splitting strategies combined from the remaining splitting parameters in each splitting dimension;

[0082] Step 102: Perform performance evaluation on the target model under each candidate splitting strategy to determine the target splitting strategy as the basis for distributed deployment of the target model.

[0083] The target model in this embodiment may be the large model mentioned above, and of course it may also be other models that need to be distributedly deployed. This embodiment does not limit the scale and function and other attributes of the target model.

[0084] Refer to Figure 1 In step 100, splitting parameters corresponding to the target model in each splitting dimension can be generated. In this embodiment, a hybrid parallel method can be supported to split the target model. Therefore, there can be multiple splitting dimensions in this embodiment. The splitting dimensions in this embodiment may include but are not limited to tensor parallel dimension, pipeline parallel dimension, data parallel dimension, and sample dimension in a single batch, etc. In this embodiment, the splitting parameters adopted in different splitting dimensions can be: tensor parallel degree in the tensor parallel dimension, pipeline parallel degree in the pipeline parallel dimension, data parallel degree in the data parallel dimension, and the number of samples in a single batch in the sample dimension of a single batch. The following will explain the above several exemplary splitting dimensions as follows:

[0085] Tensor Parallel (TP) dimension can be understood as splitting a single model layer in a model and deploying them to different processes respectively. Among them, the process runs on a node in the cluster.

[0086] The pipeline parallel dimension (PP) can be understood as dividing multiple model layers in a model into multiple groups of model layers, and different groups are deployed to different processes;

[0087] The data parallel dimension (DP) can be understood as multiple processes carrying the same part of the model, processing the same data in parallel;

[0088] The single batch size dimension (BS) can be understood as the number of samples included in a single batch input to the model. When BS is different, the amount of computation that the parallel processes need to bear will be different.

[0089] It should be understood that the above several partitioning dimensions are only exemplary, and this embodiment is not limited thereto.

[0090] In this embodiment, multiple implementation manners can be adopted to generate the partitioning parameters corresponding to each partitioning dimension for the target model. Figure 2 It is a logical schematic diagram of a model partitioning method provided for an exemplary embodiment of the present application. Refer to Figure 2 , in a preferred implementation manner: the model parameters corresponding to the target model can be obtained; the hardware information corresponding to the target cluster, which is used to carry the target model, can be queried; based on the model parameters and the hardware information, the partitioning parameters of the target model in each partitioning dimension are generated.

[0091] Among them, in this preferred implementation manner, the user can input the model parameters of the target model. The model parameters in this embodiment may include but are not limited to the number of model layers, the latent variable dimension, the number of multi-heads, the dimension of each head, the vocabulary size, the data type, the sequence length, and the maximum BS, etc. It should be understood that these model parameters are only exemplary, and this embodiment is not limited thereto, and no exhaustive listing is made here.

[0092] Among them, refer to Figure 2, in this preferred implementation, the hardware information corresponding to the target cluster can be obtained through hardware perception. The dimensions of hardware perception can include but are not limited to storage space, communication connection, computing power, etc. In practical applications, a management and control program is usually deployed in the target cluster. Based on this, the management and control program in the target cluster can be used to automatically and real-time monitor the hardware information in the target cluster. In this preferred implementation, the hardware information of the target cluster can be queried from the management and control program in the target cluster. By obtaining the hardware information corresponding to the target cluster through hardware perception, the accuracy of the hardware information can be effectively improved, and thus the relevant segmentation parameters can be calculated more accurately for the target model. The hardware information in this embodiment can include but are not limited to hardware type, topology, network communication bandwidth, number of computing nodes, number of GPU cards on the nodes, video card memory, computing power, communication bandwidth, etc. It should be understood that these hardware information are only exemplary, and this embodiment is not limited thereto, and no exhaustive list is made here.

[0093] On this basis, according to the model parameters corresponding to the target model and the hardware information corresponding to the target cluster, segmentation parameters for the target model can be generated in each segmentation dimension. The inventor found during the research process that under the two aspects of model parameters and hardware information, the number of segmentation parameters that can be calculated in each segmentation dimension is limited, rather than infinite. Therefore, in this embodiment, all the segmentation parameters supported by the target cluster in each segmentation dimension can be calculated for the target model by means of enumeration. In this embodiment, the calculation rules during the enumeration of segmentation parameters according to model parameters and hardware information are not limited, and the calculation rules can be designed as needed. For example, if there are 8 GPUs in the target cluster and the number of layers of the target model is 16, the values of 1, 2, 3, 4, 5, 6, 7, and 8 can be enumerated for the tensor parallelism in the tensor parallel dimension. Of course, this is only an exemplary calculation rule, and this embodiment is not limited thereto.

[0094] It can be seen that the segmentation parameters generated for the target model in each segmentation dimension according to the model parameters corresponding to the target model and the hardware information corresponding to the target cluster are very comprehensive.

[0095] In this embodiment, in step 100, it is further proposed that the segmentation parameters corresponding to the target model in each segmentation dimension can be filtered respectively according to a preset filtering rule. Here, in this embodiment, by setting a reasonable preset filtering rule, parameter pruning can be performed on the target model in each segmentation dimension to filter out the unreasonable segmentation parameters in each segmentation dimension.

[0096] In this embodiment, preset filtering rules can be configured for different segmentation dimensions. When filtering the segmentation parameters according to the corresponding preset filtering rules under a single segmentation dimension, the segmentation parameters under other segmentation dimensions can be default set to 1 to avoid over-filtering. Several exemplary preset filtering rules will be provided below for several exemplary segmentation dimensions in the foregoing text.

[0097] The preset filtering rules under the tensor parallelism dimension TP may include, but are not limited to: the tensor parallelism is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; the tensor parallelism can be divisible by the total number of attention heads corresponding to the target model; or, the tensor parallelism does not exceed the number of GPUs installed on a single node in the target cluster.

[0098] For example, if the total number of GPUs that can be provided in the target cluster is 8, the possible values of the tensor parallelism selected may include 1, 2, 4, 6, 8, and other enumerated tensor parallelisms will be filtered out. If the total number of attention heads (included in the model parameters) in the target model is 32, since 6 does not satisfy the divisibility rule, therefore, 6 selected previously will be filtered out, and the remaining possible values of the tensor parallelism may include 1, 2, 4, 8. If the number of GPUs installed on a single node in the target cluster is 4, then since 8 is greater than 4, 8 will be filtered out. The filtering in this regard can effectively ensure that the processes of tensor parallelism do not cross nodes, so as to avoid the need for cross-node data transmission between the processes of tensor parallelism. So far, the remaining possible values of the tensor parallelism may include 1, 2, 4.

[0099] The preset filtering rules under the pipeline parallelism dimension PP include, but are not limited to: the pipeline parallelism is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; or, the pipeline parallelism can be divisible by the total number of model layers corresponding to the target model.

[0100] For example, if the total number of model layers of the target model is 32 layers, the possible values of the pipeline parallelism selected may include 1, 2, 4, 8, and other enumerated pipeline parallelisms will be filtered out. The filtering in this regard can make the number of model layers carried on each process of pipeline parallelism consistent, so that the workload between processes is more balanced.

[0101] The preset filtering rules under the data parallelism dimension DP may include: the data parallelism does not exceed the ratio between the total memory of the target cluster and the memory occupied by the target model.

[0102] For example, if the memory required by the target model itself is 26GB, and there are 8 GPUs provided in the target cluster, with the memory capacity of each GPU being 14GB, then the total memory of the target cluster is 14 * 8 = 112GB. Since 112 / 26 ≥ 4, therefore, the possible values of the data parallelism can include 1, 2, 3, and 4. That is, at most 4 copies of the target model can be deployed in the target cluster.

[0103] The preset filtering rules under the single batch specification dimension BS can include: the number of samples in a single batch does not exceed the empirical value supported by the target cluster and does not exceed the number of samples supported by the remaining memory. The remaining memory is the memory remaining in the target cluster after deploying the target model as estimated.

[0104] For example, if the model of the GPU in the target cluster is A10, according to experience, it can be determined that the BS experience it supports is 512. And if the GPU model in the target cluster is A20, whose computing power is twice that of A10, then it can be determined that the BS experience value it supports is 1024. Continuing with the example above where the total memory of the target cluster is 112GB, since DP is set to 1 during filtering, the remaining memory after deploying the target model in the target cluster will be 112 - 26 = 84GB. And by estimation, the memory required for a single sample in a single batch is 32MB, then 84GB / 34MB = 2752 can be calculated. Then the minimum value can be taken from 2752 and 1024, which is 1024, and the BS values exceeding 1024 enumerated will be filtered out. In addition, during the enumeration process, an enumeration step greater than 1 can be used, for example, an enumeration step of 16 can be used, and some BS values can also be filtered out.

[0105] It should be understood that the above preset filtering rules are only exemplary, and this embodiment is not limited thereto, and no more examples are given here.

[0106] It can be understood that in step 100, in this embodiment, by designing reasonable filtering rules, the segmentation parameters under each segmentation dimension are boldly filtered once. After the filtering operation, the number of remaining segmentation parameters of the target model under each segmentation dimension is already relatively small. This not only reduces the number of alternative segmentation parameters, but also ensures the accuracy of the remaining segmentation parameters.

[0107] On this basis, continue to refer to Figure 1, in step 101, a partitioning strategy can be generated by combining the remaining partitioning parameters under each partitioning dimension. As mentioned above, in this embodiment, model partitioning is supported in a hybrid parallel manner. Therefore, the combination here can be understood as taking one partitioning parameter from each partitioning dimension to combine a partitioning strategy. For example, if 4 partitioning dimensions are used, a partitioning strategy combined in step 101 can be understood as a quadruple, taking 1 partitioning parameter from each of the 4 partitioning dimensions to obtain the 4 elements in the quadruple.

[0108] Based on the first-level filtering operation in the foregoing step 100, the remaining partitioning parameters under each partitioning dimension are already relatively few. Therefore, in step 101, an enumeration method can be used to enumerate all the partitioning strategies that can be combined among the remaining partitioning parameters under each partitioning dimension.

[0109] In this embodiment, in step 101, it is further proposed to search for the partitioning strategies that meet the verification conditions from the combined partitioning strategies as candidate partitioning strategies.

[0110] In an exemplary search scheme:

[0111] If the product of the tensor parallelism, pipeline parallelism, and data parallelism included in the candidate partitioning strategy is equal to the total number of GPUs that the target cluster can provide, it is determined that the verification condition is met. This can ensure that all GPUs in the target cluster can be fully utilized.

[0112] And / or, if the data parallelism included in the candidate partitioning strategy and the number of samples in a single batch are divisible, it is determined that the verification condition is met. This can ensure that the number of samples in a single batch can be evenly distributed to multiple parallel model replicas, making the workload among processes more balanced.

[0113] In this exemplary search scheme, two aspects of exemplary verification conditions are provided. It should be noted that the verification conditions in this embodiment are not limited to this.

[0114] It can be understood that in this embodiment, in step 101, another level of filtering is performed. After the filtering operation, various unreasonable partitioning strategies can be screened out, and the number of remaining candidate partitioning strategies will be more refined. This not only reduces the number of candidate partitioning strategies but also ensures the accuracy of the candidate partitioning strategies.

[0115] After that, in this embodiment, in step 102, the performance of the target model can be evaluated under each candidate partitioning strategy to determine the target partitioning strategy as the basis for distributed deployment of the target model.

[0116] Among them, performance evaluation can be understood as evaluating the model performance that the target model can achieve when it is distributedly deployed on the target cluster according to a certain candidate segmentation strategy. In practical applications, the optimal one can be selected from the candidate segmentation strategies according to the performance evaluation results to determine the target segmentation strategy. In this way, when the target model is distributedly deployed according to the target segmentation strategy subsequently, the model performance of the target model can be effectively guaranteed.

[0117] In summary, in this embodiment, before the target model is distributedly deployed, segmentation parameters in different segmentation dimensions can be generated for the target model, and the generated segmentation parameters can be filtered according to a preset filtering rule, which effectively reduces the number of segmentation parameters in each segmentation dimension, and the number of segmentation strategies generated by the combination of the remaining segmentation parameters in each segmentation dimension will also be significantly reduced. Moreover, candidate segmentation strategies that meet the verification conditions are searched from the combined segmentation strategies. In this way, through multi-level filtering operations, low-quality segmentation strategies can be reasonably discarded, so that the number of remaining candidate segmentation strategies is more refined. On this basis, the target model can be performance-evaluated under each candidate segmentation strategy to finally determine a suitable target segmentation strategy. Accordingly, a suitable target segmentation strategy for the target model is efficiently determined and used as the basis for the distributed deployment of the target model, without relying on manual experience, thus effectively improving the model segmentation effect.

[0118] Figure 3 It is a schematic flowchart of another model segmentation method provided by an exemplary embodiment of the present application. Refer to Figure 3 , this method may include:

[0119] Step 300: Filter the segmentation parameters corresponding to the target model in each segmentation dimension respectively according to a preset filtering rule;

[0120] Step 301: Search for N candidate segmentation strategies that meet the verification conditions from the segmentation strategies combined from the remaining segmentation parameters in each segmentation dimension;

[0121] Step 302: Calculate the corresponding simulated evaluation results of the target model on the target cluster under the N candidate segmentation strategies respectively;

[0122] Step 303: Screen K candidate segmentation strategies from the N candidate segmentation strategies according to the simulated evaluation results, where K < N, and K and N are positive integers;

[0123] Step 304: Distributively deploy the target model on the target cluster according to the K candidate segmentation strategies to generate actual machine evaluation results;

[0124] Step 305: Determine the target segmentation strategy from the K candidate segmentation strategies according to the actual machine evaluation results.

[0125] Among them, steps 300-301 can refer to the relevant descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, an optional implementation manner for determining a target segmentation strategy from candidate segmentation strategies can be provided based on steps 302-305.

[0126] Refer to Figure 3 , in step 302, under N candidate segmentation strategies, the performance of the target model can be simulated and evaluated on the target cluster respectively to obtain corresponding simulation evaluation results. As mentioned above, in this embodiment, after the two-layer filtering performed by steps 300 and 301, the number of candidate segmentation strategies has been relatively small. Therefore, in step 302, only the filtered candidate segmentation strategies need to be simulated and evaluated for performance, and the computational efficiency of the performance simulation evaluation operation in step 302 will be greatly improved.

[0127] It should be understood that the simulation evaluation results generated by the performance model evaluation operation performed in step 302 are a theoretical value. Although its accuracy cannot reach complete accuracy, it can basically reflect the model segmentation effects that different candidate segmentation strategies can roughly achieve. Therefore, in this embodiment, the simulation evaluation results are used to perform another layer of filtering on the candidate segmentation strategies.

[0128] In addition, the specific implementation manner of the performance simulation evaluation operation is not limited in this embodiment and will be described exemplarily later, and will not be elaborated here for the time being.

[0129] However, in this embodiment, the final target segmentation strategy is not directly determined based on the simulation evaluation results. Refer to Figure 2 , but an innovative mechanism of combining theoretical search with actual machine verification and optimization is proposed.

[0130] According to this mechanism, in step 303, K candidate segmentation strategies to be verified can be screened from N candidate segmentation strategies according to the simulation evaluation results.

[0131] In an exemplary screening scheme:

[0132] First, according to the memory overhead in the simulation evaluation results, candidate segmentation strategies estimated to have memory overflow problems can be screened out from N candidate segmentation strategies. The inventor found during the research process that the simulation evaluation results usually include memory overhead. Therefore, in this exemplary screening scheme, this information is fully utilized to estimate candidate segmentation strategies with memory overflow problems and screen them out. In practical applications, the memory overhead records the memory amount required for a single batch of input, while the hardware information includes the actual memory amount. If the memory amount required for a single batch of input exceeds the remaining memory amount after the target model is simulated and deployed on the target cluster, it can be determined that there is a memory overflow problem. This is essentially a layer of filtering. After this filtering operation, some defective candidate segmentation strategies can be screened out.

[0133] Then, based on the simulation evaluation results, the evaluation indices corresponding to the remaining candidate segmentation strategies can be calculated. That is, only the evaluation indices corresponding to the candidate segmentation strategies remaining in the previous step are needed. Among them, the evaluation index can include the sampling rate, etc. The sampling rate can be understood as the amount of data that the target model can process per unit time. Exemplarily, an exemplary scheme for calculating the sampling rate can be: determining the number of samples in a single batch from a remaining candidate segmentation strategy; determining the time taken for the target model to perform calculations according to the number of samples from the simulation evaluation results corresponding to this candidate segmentation strategy; taking the ratio between the number of samples and the time taken as the sampling rate. Of course, in addition to the sampling rate, other types of evaluation indices can also be used, and it is not limited to this.

[0134] After that, the remaining candidate segmentation strategies can be sorted according to the evaluation index to screen out K candidate segmentation strategies to be verified. Referring to Figure 2 , in practical applications, the K candidate segmentation strategies with the highest evaluation indices can be selected as the candidate segmentation strategies to be verified.

[0135] After this layer of filtering, the number of remaining segmentation strategies has reached K, which is already very small.

[0136] Continuing to refer to Figure 3 , as proposed in step 304, the target model can be distributedly deployed on the target cluster according to the K candidate segmentation strategies to be verified to generate on-machine evaluation results. It should be understood that in step 304, on-machine verification tests are performed according to each candidate segmentation strategy to be verified. An exemplary on-machine evaluation scheme can be: calling the model execution framework and inputting the segmentation parameters in the K candidate segmentation strategies to be verified into the model execution framework respectively; using the model execution framework to distribute and deploy the target model to the target cluster according to the K candidate segmentation strategies respectively; performing on-machine evaluation operations respectively under the K candidate segmentation strategies to generate on-machine evaluation results.

[0137] In practical applications, a batch processing program can be adopted to call a specific model execution framework (such as Megatron or TensorRT, etc.), and complete the aforementioned distributed initialization process for the target model on the target cluster according to each segmentation strategy to be verified, and conduct on-machine evaluation. Here, the model execution framework called can preferably be the model execution framework that may be adopted when the target model is deployed online in the future. Moreover, the inventor found during the research process that in the process of distributed initialization by the model execution framework, a large amount of internal processing logic needs to be carried out based on the input segmentation parameters to complete the distributed deployment, and a configuration file will be generated during this process. In this embodiment, it is further proposed that the configuration file can be extracted from the called model execution framework; and after determining the target segmentation strategy, the configuration file corresponding to the target segmentation strategy is retained. In this way, when the target model is deployed online in the future, the configuration file can be directly input into the corresponding model execution framework, which can eliminate a large amount of repetitive work in the distributed initialization process of the model execution framework and improve the efficiency of distributed initialization.

[0138] In this embodiment, when conducting the on-machine evaluation operation, the evaluation mechanism adopted can include but is not limited to: evaluating the processing time corresponding to a fixed BS, or evaluating the BS that can be supported by a fixed processing time. In addition, in this embodiment, the same evaluation metrics as those in the aforementioned simulation evaluation can be adopted during the on-machine evaluation. For example, the evaluation metric adopted during the on-machine evaluation can also be the sampling rate. Continuing with this exemplary evaluation metric, evaluating the processing time corresponding to a fixed BS can be understood as: detecting the processing time consumed by the target model when processing a fixed BS, and calculating the ratio between the fixed BS and the consumed processing time as the sampling rate of the target model. And evaluating the BS that can be supported by a fixed processing time can be understood as: detecting the actual BS processed by the target model within a fixed processing time, and calculating the ratio between this BS and the fixed processing time as the sampling rate of the target model. The essence of both mechanisms is to calculate the workload that the target model can process per unit time as the on-machine evaluation result of the target model.

[0139] In this way, in step 304, the on-machine evaluation results corresponding to the target model under K segmentation strategies to be verified can be generated.

[0140] On this basis, referring to Figure 2 , in step 305, the target segmentation strategy can be determined from the K segmentation strategies to be verified according to the on-machine evaluation results. In practical applications, the on-machine evaluation results generated by the K segmentation strategies to be verified can be sorted, and the target segmentation strategy can be found out from them.

[0141] In summary, in this embodiment, a mechanism of combining theoretical search with actual machine verification and optimization is adopted to optimize the final target segmentation strategy from N candidate segmentation strategies, rather than relying entirely on theoretical calculations. Moreover, during the process of actual machine verification and optimization, the model execution framework expected to be used during the subsequent online deployment of the target model can be adopted, and the configuration file precipitated when the model execution framework performs actual machine deployment on the target model can be saved, so that it can be directly used by the model execution framework during subsequent online deployment, saving a lot of repetitive work, and thus improving the online deployment efficiency of the target model.

[0142] In the above or following embodiments, the specific implementation scheme of the simulation evaluation is not limited. The following provides a preferred implementation scheme. For ease of description, taking the first candidate segmentation strategy among the above-mentioned N candidate segmentation strategies as an example, this preferred implementation scheme will be elaborated. It should be understood that the same simulation evaluation logic can be adopted under other candidate segmentation strategies to obtain corresponding simulation evaluation results.

[0143] In this preferred implementation scheme: Under the first candidate segmentation strategy, the performance overhead corresponding to the backbone network in the target model on the target cluster can be simulated and evaluated as the simulation evaluation result corresponding to the target model under the first candidate segmentation strategy. The inventor found during the research process that the model may include multiple parts, such as: the vectorized embedding part, the backbone network, and the prediction head projection head, etc. Among them, the most significant performance overhead is the backbone network. Therefore, in this preferred implementation scheme, it is proposed that during the simulation evaluation, the performance overhead corresponding to the backbone network in the target model can be focused on and used as the simulation evaluation result corresponding to the target model under the first candidate segmentation strategy.

[0144] On this basis, an exemplary simulation evaluation scheme for the performance overhead of the backbone network in the target model can be:

[0145] Taking the unit sample received by a single model layer carried on the target process as the simulation evaluation unit, the unit performance overhead corresponding to the target process is simulated and evaluated, where the target process is any process used to carry the target model on the simulated target cluster;

[0146] According to each segmentation parameter in the first candidate segmentation strategy, the simulation evaluation coefficient corresponding to the target process is determined;

[0147] The product of the simulation evaluation coefficient and the unit performance overhead is calculated as the performance overhead corresponding to the target process;

[0148] According to the performance overhead calculated for each process used to carry the backbone network, the performance overhead corresponding to the backbone network is estimated.

[0149] First, in this exemplary simulation evaluation scheme, the target model is simulatedly partitioned according to the first candidate partitioning strategy to generate multiple model parts, and parallel processes used to host the target model on the target cluster are also simulated. It should be understood that the number of parallel processes simulated on the target cluster can be determined based on the hardware information corresponding to the target cluster.

[0150] Figure 4 This is a schematic diagram of an exemplary model partitioning principle provided by an exemplary embodiment of the present application. Refer to Figure 4 , for example, if there are two nodes in the target cluster and 16 GPUs can be provided, then one process can be started on each GPU ( Figure 4 g0 - g15 in ). And as mentioned above, in any of the candidate partitioning strategies in this embodiment, the total number of model parts to be partitioned should be adapted to the number of parallel processes that can be provided in the target cluster. Therefore, when the target model is simulatedly partitioned, 16 model parts can be simulatedly partitioned. Of course, under different candidate partitioning strategies, the parallel relationships among the 16 partitioned model parts are different. Figure 4 The corresponding candidate partitioning strategy is [TP = 2, PP = 4, DP = 2, BS = 2]. In this way, the 16 partitioned model parts will be hosted by the 16 processes.

[0151] In this way, from the overall perspective of the target model, it will receive a batch of inputs. Figure 4 In, BS = 2, and the single - batch input will be distributed according to multiple partitioning dimensions to each process. From the perspective of the process, in the case of a single - batch input, the workloads that different processes need to handle may be different. Therefore, the performance overheads required by different processes may also be different. In this exemplary simulation evaluation scheme, it is proposed to use the unit samples received by a single model layer hosted on the target process as the simulation evaluation unit to simulate and evaluate the unit performance overhead required by the target process. It should be understood that there is no limitation on the specific rules adopted in the process of simulating and evaluating the unit performance overhead. Since the unit performance overhead is affected by model parameters and hardware parameters, there must be a natural influence law between the model parameters, hardware parameters and the unit performance overhead. In practical applications, this influence law can be extracted through machine learning or manual summarization as the rule in the process of simulating and evaluating the unit performance overhead, and this influence law is not exhaustively listed here.

[0152] Then, in this exemplary simulation evaluation scheme, the simulation evaluation coefficient corresponding to the target process is also determined according to each segmentation parameter in the first candidate segmentation strategy. Among them, the simulation evaluation coefficient is used to characterize the multiple corresponding to the unit performance overhead required on the target process. The simulation evaluation coefficient is usually equal to (the number of model layers carried on the target process * BS) / DP. Specifically, it can be understood that: after a batch of samples (for example, BS = 2) are input into the target model, they are allocated once based on DP (for example, DP = 2). In this way, the BS allocated to each model layer becomes 1. If there are two model layers loaded on the target process, the BS allocated to the target process becomes 2 again. In this way, the simulation evaluation coefficient on the target process can be calculated as 2.

[0153] After that, in this exemplary simulation evaluation scheme, the product of the simulation evaluation coefficient and the unit performance overhead is calculated as the performance overhead corresponding to the target process. Continuing with the above example, the performance overhead corresponding to the target process is 2 (simulation evaluation coefficient) * unit performance overhead.

[0154] Finally, in this exemplary simulation evaluation scheme, the performance overhead corresponding to the backbone network can be estimated according to the performance overhead calculated for each process used to carry the backbone network. The inventor found during the research process that since the model will ultimately be distributed for computing, in this link, the representative process can be selected according to the process time consumption included in the performance overhead corresponding to each process; the performance overhead corresponding to the representative process is used as the performance overhead corresponding to the backbone network. In practical applications, the process with the longest time consumption can be selected as the representative process.

[0155] In this way, through this exemplary simulation evaluation scheme, the performance overhead corresponding to the backbone network in the target model on the target cluster can be simulated and evaluated.

[0156] In addition, the inventor found during the research process that the performance overhead required in the model inference stage and the model training stage is different, which is mainly due to the different number of rounds of the end-to-end propagation process required in the inference stage and the training stage. In the inference stage, usually a single round of the end-to-end propagation process is sufficient to obtain the model output result. However, in the training stage, because there are two propagation directions, forward propagation and backward propagation, multiple rounds of the end-to-end propagation process may be required to complete one training. Therefore, in this exemplary simulation evaluation scheme, when simulating and evaluating the unit performance overhead corresponding to the target process, the inference stage and the training stage can be simulated and evaluated separately.

[0157] Specifically: for the inference stage, simulate and evaluate the performance overhead corresponding to the end-to-end propagation process of a single unit sample, which is used as the unit performance overhead corresponding to the target process; for the training stage, simulate and evaluate the total performance overhead corresponding to the multi-round end-to-end propagation process required for training and other links required for training of a single unit sample, which is used as the unit performance overhead corresponding to the target process. In this way, based on the unit performance overheads respectively simulated and evaluated for the inference stage and the training stage, the performance overheads required for the target model in the inference stage and the training stage can be calculated.

[0158] In the process of screening K candidate splitting strategies described in the foregoing embodiments, the performance overheads simulated and evaluated for the inference stage and the training stage can both be used as the simulation evaluation results, so as to more reasonably screen out K candidate splitting strategies.

[0159] Furthermore, various other events that may occur in the end-to-end propagation process can also be simulated and calculated, and the time consumption that these events may cause can be simulated and added to the foregoing unit performance overhead. Among them, these events that may occur can include, but are not limited to, pipeline bubble events and the preparation work that the subsequent process needs to do for the previous process, etc. The pipeline bubble event can be understood as that the processes in pipeline parallelism may wait for each other due to inconsistent work rhythms before and after. In this way, after simulating the time consumption caused by these other events, the accuracy of the simulated and evaluated unit performance overhead can be further improved, and then the accuracy of the performance overhead simulated and evaluated for the backbone network in the target model under the first candidate splitting strategy can be improved.

[0160] It should be understood that in this embodiment, other exemplary solutions can also be used to simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster. For example, the BS value, the memory amount required, the communication volume required, and the communication overhead that each process can support under the condition of fixed time consumption can be simulated and evaluated. And it is not limited to the above exemplary simulation and evaluation solutions. No more detailed examples are given here.

[0161] In this embodiment, the performance overheads concerned in the simulation and evaluation process can include, but are not limited to, computing overhead, communication overhead, and / or memory overhead, etc. Among them, the computing overhead and the communication overhead can be used as the basis for calculating the evaluation indicators in the foregoing embodiments, while the memory overhead can be used as the basis for evaluating the memory overflow problem in the foregoing.

[0162] In summary, in this embodiment, by simulating and evaluating the performance overhead required by the backbone network in the target model under different candidate splitting strategies, the simulation evaluation results for the target model can be generated more accurately and efficiently, so that K candidate splitting strategies can be more reasonably screened out from N candidate splitting strategies, and further the selected target splitting strategy can be made more reasonable.

[0163] It should be noted that in some of the processes described in the above embodiments and accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The operation numbers such as 301 and 302 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different candidate segmentation strategies, etc., do not represent a sequence, and do not limit that "first" and "second" are different types.

[0164] Figure 5 FIG. is a schematic structural diagram of a computing device provided by another exemplary embodiment of the present application. As Figure 5 shown, the computing device includes: a memory 50, a processor 51, and a communication component 52.

[0165] The processor 51 is coupled to the memory 50 and the communication component 52 and is configured to execute a computer program in the memory 50 for:

[0166] Filter the segmentation parameters corresponding to the target model in each segmentation dimension according to a preset filtering rule;

[0167] Find candidate segmentation strategies that meet the verification conditions from the segmentation strategies combined from the remaining segmentation parameters in each segmentation dimension;

[0168] Perform performance evaluation on the target model under each candidate segmentation strategy to determine the target segmentation strategy as the basis for distributed deployment of the target model.

[0169] In an optional embodiment, the processor 51 may further be configured to:

[0170] Obtain the model parameters corresponding to the target model;

[0171] Query the hardware information corresponding to the target cluster, where the target cluster is used to host the target model;

[0172] Generate segmentation parameters for the target model in each segmentation dimension based on the model parameters and the hardware information.

[0173] In an optional embodiment, the segmentation dimension includes a tensor parallel dimension, and the segmentation parameter under the tensor parallel dimension adopts a tensor parallelism degree. The preset filtering rule under the tensor parallel dimension includes:

[0174] The tensor parallelism degree is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster;

[0175] The tensor parallelism can be evenly divided by the total number of attention heads corresponding to the target model; and / or

[0176] The tensor parallelism does not exceed the number of GPUs installed on a single node in the target cluster;

[0177] wherein, the target cluster is used to host the target model.

[0178] In an alternative embodiment, the partitioning dimension includes a pipeline parallelism dimension, the partitioning parameter under the pipeline parallelism dimension adopts the pipeline parallelism, and the preset filtering rules under the pipeline parallelism dimension include:

[0179] The pipeline parallelism is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; and / or

[0180] The pipeline parallelism can be evenly divided by the total number of model layers corresponding to the target model;

[0181] wherein, the target cluster is used to host the target model.

[0182] In an alternative embodiment, the partitioning dimension includes a data parallelism dimension, the partitioning parameter under the data parallelism dimension adopts the data parallelism, and the preset filtering rules under the data parallelism dimension include: the data parallelism does not exceed the ratio between the total memory of the target cluster and the memory occupied by the target model; and / or,

[0183] The partitioning dimension includes a single batch specification dimension, the partitioning parameter under the single batch specification dimension adopts the number of samples in a single batch, and the preset filtering rules under the single batch specification dimension include: the number of samples in a single batch does not exceed the empirical value supported by the target cluster and does not exceed the number of samples supported by the remaining memory, where the remaining memory is the estimated remaining memory in the target cluster after deploying the target model;

[0184] wherein, the target cluster is used to host the target model.

[0185] In an alternative embodiment, the candidate partitioning strategy includes tensor parallelism, pipeline parallelism, data parallelism, and / or the number of samples in a single batch. When the processor 51 searches for a candidate partitioning strategy that meets the verification conditions, it can specifically be used for:

[0186] If the product of the tensor parallelism, pipeline parallelism, and data parallelism included in the candidate partitioning strategy is equal to the total number of GPUs that the target cluster can provide, it is determined that the verification conditions are met; and / or,

[0187] If the data parallelism included in the candidate segmentation strategy is divisible by the number of samples in a single batch, it is determined that the verification condition is met;

[0188] Among them, the target cluster is used to host the target model.

[0189] In an optional embodiment, if the number of candidate segmentation strategies is N, when the processor 51 performs performance evaluation on the target model under each candidate segmentation strategy to determine the target segmentation strategy, it can specifically be used for:

[0190] Under the N candidate segmentation strategies, respectively calculate the corresponding simulated evaluation results of the target model on the target cluster;

[0191] According to the simulated evaluation results, screen K candidate segmentation strategies to be verified from the N candidate segmentation strategies, where K < N;

[0192] Deploy the target model on the target cluster according to the K candidate segmentation strategies to be verified to generate actual machine evaluation results;

[0193] According to the actual machine evaluation results, determine the target segmentation strategy from the K candidate segmentation strategies to be verified.

[0194] In an optional embodiment, when the processor 51 calculates the corresponding simulated evaluation results of the target model on the target cluster under the N candidate segmentation strategies, it can specifically be used for:

[0195] Under the first candidate segmentation strategy, simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster as the simulated evaluation result corresponding to the target model under the first candidate segmentation strategy;

[0196] Among them, the performance overhead includes computing overhead, communication overhead, and / or memory overhead, and the first candidate segmentation strategy is any candidate segmentation strategy.

[0197] In an optional embodiment, when the processor 51 simulates and evaluates the performance overhead corresponding to the backbone network in the target model on the target cluster, it can specifically be used for:

[0198] Taking the unit sample received by a single model layer hosted on the target process as the simulation evaluation unit, simulate and evaluate the unit performance overhead corresponding to the target process, where the target process is any process used to host the target model on the simulated target cluster;

[0199] According to each segmentation parameter in the first candidate segmentation strategy, determine the simulation evaluation coefficient corresponding to the target process;

[0200] Calculate the product of the simulated evaluation coefficient and the unit performance overhead as the performance overhead corresponding to the target process;

[0201] Estimate the performance overhead corresponding to the backbone network based on the performance overhead calculated for each process used to carry the backbone network.

[0202] In an alternative embodiment, when the processor 51 simulates and evaluates the unit performance overhead corresponding to the target process, it may specifically be used for:

[0203] For the inference phase, simulate and evaluate the performance overhead corresponding to the single-round end-to-end propagation of the unit sample as the unit performance overhead corresponding to the target process;

[0204] For the training phase, simulate and evaluate the total performance overhead corresponding to the multi-round end-to-end propagation of the unit sample during training and other training-required links as the unit performance overhead corresponding to the target process.

[0205] In an alternative embodiment, when the processor 51 estimates the performance overhead corresponding to the backbone network based on the performance overhead calculated for each process used to carry the backbone network, it may specifically be used for:

[0206] Filter representative processes based on the process time consumed included in the performance overhead corresponding to each process;

[0207] Use the performance overhead corresponding to the representative process as the performance overhead corresponding to the backbone network.

[0208] In an alternative embodiment, when the processor 51 filters K candidate segmentation strategies to be verified from the N candidate segmentation strategies according to the simulation evaluation results, it may specifically be used for:

[0209] Filter out the candidate segmentation strategies estimated to have memory overflow problems from the N candidate segmentation strategies according to the memory overhead in the simulation evaluation results;

[0210] Based on the simulation evaluation results, calculate the evaluation index corresponding to each of the remaining candidate segmentation strategies;

[0211] Sort the remaining candidate segmentation strategies according to the evaluation index to filter out the K candidate segmentation strategies to be verified;

[0212] Among them, the evaluation index includes the sampling rate.

[0213] In an alternative embodiment, when the processor 51 performs distributed deployment of the target model on the target cluster according to the K candidate segmentation strategies to be verified to generate real-machine evaluation results, it may specifically be used for:

[0214] Invoke the model execution framework and input the segmentation parameters in the K segmentation strategies to be verified into the model execution framework respectively;

[0215] Use the model execution framework to deploy the target model to the target cluster in a distributed manner according to the K segmentation strategies to be verified respectively;

[0216] Perform physical machine evaluation operations respectively under the K segmentation strategies to be verified to generate physical machine evaluation results.

[0217] In an alternative embodiment, the processor 51 can also be used for:

[0218] Extract the configuration file generated when the target model is deployed to the target cluster in a distributed manner according to the target segmentation strategy from the invoked model execution framework;

[0219] Save the configuration file.

[0220] Furthermore, as Figure 5 shown, the computing device further includes: a power supply component 53 and other components. Figure 5 Only some components are schematically shown in Figure 5 and it does not mean that the computing device only includes

[0221] shown components.

[0222] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executed in the above method embodiments.

[0223] The above Figure 5 The memory is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of these data include instructions for any application or method for operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0224] The above Figure 5The communication component therein is configured to facilitate communication, either wired or wireless, between the device where the communication component is located and other devices. The device where the communication component is located can access a communication standard-based wireless network, such as a WiFi, 2G, 3G, 4G / LTE, 5G, etc. mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0225] The above Figure 5 The power supply component therein provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0226] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0227] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0228] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.

[0229] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one or more boxes.

[0230] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.

[0231] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0232] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A model segmentation method, characterized in that, Including: Filter the segmentation parameters corresponding to the target model in each segmentation dimension respectively according to the preset filtering rules; Search for candidate segmentation strategies that meet the verification conditions from the segmentation strategies combined from the remaining segmentation parameters in each segmentation dimension; Perform performance evaluation on the target model under each candidate segmentation strategy to determine the target segmentation strategy as the basis for distributed deployment of the target model.

2. The method according to claim 1, wherein It also includes: Obtain the model parameters corresponding to the target model; Query the hardware information corresponding to the target cluster, where the target cluster is used to host the target model; Based on the model parameters and the hardware information, generate segmentation parameters for the target model in each segmentation dimension.

3. The method according to claim 1 or 2, characterized in that, The segmentation dimension includes a tensor parallel dimension. The segmentation parameter under the tensor parallel dimension adopts a tensor parallelism degree. The preset filtering rules under the tensor parallel dimension include: The tensor parallelism degree is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; The tensor parallelism degree can be divided evenly by the total number of attention heads corresponding to the target model; and / or The tensor parallelism degree does not exceed the number of GPUs installed on a single node in the target cluster; Wherein, the target cluster is used to host the target model.

4. The method according to claim 1 or 2, characterized in that, The segmentation dimension includes a pipeline parallel dimension. The segmentation parameter under the pipeline parallel dimension adopts a pipeline parallelism degree. The preset filtering rules under the pipeline parallel dimension include: The pipeline parallelism degree is 1 or an even number greater than 1 and does not exceed the total number of GPUs in the target cluster; and / or The pipeline parallelism degree can be divided evenly by the total number of model layers corresponding to the target model; Wherein, the target cluster is used to host the target model.

5. The method according to claim 1 or 2, characterized in that, The segmentation dimension includes a data parallel dimension. The segmentation parameter under the data parallel dimension adopts a data parallelism degree. The preset filtering rules under the data parallel dimension include: the data parallelism degree does not exceed the ratio between the total memory of the target cluster and the memory occupied by the target model; and / or, The segmentation dimension includes a single batch specification dimension. The segmentation parameter under the single batch specification dimension adopts the number of samples in a single batch. The preset filtering rules under the single batch specification dimension include: the number of samples in a single batch does not exceed the empirical value supported by the target cluster and does not exceed the number of samples supported by the remaining memory. The remaining memory is the estimated remaining memory after deploying the target model in the target cluster; Wherein, the target cluster is used to host the target model.

6. The method according to claim 1 or 2, characterized in that, The candidate segmentation strategy includes a tensor parallelism degree, a pipeline parallelism degree, a data parallelism degree, and / or the number of samples in a single batch. Searching for candidate segmentation strategies that meet the verification conditions includes: If the product of the tensor parallelism degree, the pipeline parallelism degree, and the data parallelism degree included in the candidate segmentation strategy is equal to the total number of GPUs that the target cluster can provide, it is determined that the verification conditions are met; and / or, If the data parallelism degree included in the candidate segmentation strategy and the number of samples in a single batch can be divided evenly, it is determined that the verification conditions are met; Wherein, the target cluster is used to host the target model.

7. The method according to claim 2, wherein If the number of candidate segmentation strategies is N, then the performance of the target model is evaluated under each candidate segmentation strategy to determine the target segmentation strategy, including: Under the N candidate segmentation strategies, calculate the corresponding simulation evaluation results of the target model on the target cluster respectively; According to the simulation evaluation results, screen K candidate segmentation strategies to be verified from the N candidate segmentation strategies, where K < N; Deploy the target model on the target cluster according to the K candidate segmentation strategies to be verified to generate the real machine evaluation results; Determine the target segmentation strategy from the K candidate segmentation strategies to be verified according to the real machine evaluation results.

8. The method according to claim 7, wherein Under the N candidate segmentation strategies, calculate the corresponding simulation evaluation results of the target model on the target cluster respectively, including: Under the first candidate segmentation strategy, simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster as the simulation evaluation result corresponding to the target model under the first candidate segmentation strategy; Among them, the performance overhead includes computing overhead, communication overhead, and / or memory overhead, and the first candidate segmentation strategy is any candidate segmentation strategy.

9. The method according to claim 8, wherein Simulate and evaluate the performance overhead corresponding to the backbone network in the target model on the target cluster, including: Taking the unit sample received by a single model layer carried on the target process as the simulation evaluation unit, simulate and evaluate the unit performance overhead corresponding to the target process, where the target process is any process used to carry the target model on the simulated target cluster; Determine the simulation evaluation coefficient corresponding to the target process according to each segmentation parameter in the first candidate segmentation strategy; Calculate the product of the simulation evaluation coefficient and the unit performance overhead as the performance overhead corresponding to the target process; Estimate the performance overhead corresponding to the backbone network according to the performance overhead calculated for each process used to carry the backbone network.

10. The method according to claim 9, wherein Simulate and evaluate the unit performance overhead corresponding to the target process, including: For the inference stage, simulate and evaluate the performance overhead corresponding to the single-round end-to-end propagation process of the unit sample as the unit performance overhead corresponding to the target process; For the training stage, simulate and evaluate the total performance overhead corresponding to the multi-round end-to-end propagation process required for training the unit sample and other training required links as the unit performance overhead corresponding to the target process.

11. The method according to claim 9, wherein Estimate the performance overhead corresponding to the backbone network according to the performance overhead calculated for each process used to carry the backbone network, including: Screen the representative processes according to the process time consumed included in the performance overhead corresponding to each process; Take the performance overhead corresponding to the representative process as the performance overhead corresponding to the backbone network.

12. The method according to claim 7, characterized in that According to the simulation evaluation results, screen K candidate segmentation strategies to be verified from the N candidate segmentation strategies, including: According to the memory overhead in the simulation evaluation results, screen out the candidate segmentation strategies estimated to have memory overflow problems from the N candidate segmentation strategies; Based on the simulation evaluation results, calculate the evaluation index corresponding to each of the remaining candidate segmentation strategies; Sort the remaining candidate segmentation strategies according to the evaluation index to screen out the K candidate segmentation strategies to be verified; Among them, the evaluation index includes the sampling rate.

13. The method according to claim 7, wherein Perform distributed deployment of the target model on the target cluster according to the K candidate segmentation strategies to be verified to generate an in-machine evaluation result, including: Invoke the model execution framework and input the segmentation parameters in the K candidate segmentation strategies into the model execution framework respectively; Use the model execution framework to distribute and deploy the target model to the target cluster according to the K candidate segmentation strategies respectively; Perform in-machine evaluation operations respectively under the K candidate segmentation strategies to generate in-machine evaluation results.

14. The method according to claim 13, wherein It also includes: Extract the configuration file generated when the target model is distributed and deployed to the target cluster according to the target segmentation strategy from the invoked model execution framework; Save the configuration file.

15. A computing device, characterized in that, It includes a memory, a processor, and a communication component; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute the one or more computer instructions to perform the model segmentation method according to any one of claims 1-14.

16. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the model segmentation method according to any one of claims 1-14.

Citation Information

Cited By

  • Model parameter processing method, system on chip and model parameter processing system

    CN121543651A

  • Tensor slice implementation method on TPU (thermoplastic polyurethane) and TPU

    CN122332132A