Model performance estimation method and device and computer equipment

By constructing a distributed overhead calculation graph and quantifying the basic and additional overhead of operation operators, the problem of inaccurate performance estimation of multi-card deployment in the pre-silicon stage of artificial intelligence chip design is solved, and accurate performance evaluation and optimization are achieved under conditions of limited hardware resources.

CN120806199AActive Publication Date: 2025-10-17SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202511299908.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

During the AI ​​chip design phase, existing technologies make it difficult to accurately evaluate the performance ceiling of large models in multi-card distributed deployments in the pre-silicon stage. This is especially true due to the limited resources of hardware simulators, which cannot effectively simulate the multi-card interconnection mode, resulting in inaccurate performance estimates.

Method used

By determining the model configuration data and candidate distributed strategies of the target model, constructing a distributed cost calculation graph, quantifying the basic cost of the operation operator and the additional cost of connecting operators, and realizing the performance estimation of the candidate distributed strategies, including the summary of computing cost, communication cost and memory cost.

Benefits of technology

Accurately evaluate the model's performance ceiling in the pre-silicon stage, optimize hardware design, and provide a multi-dimensional performance optimization basis. This overcomes the limitation of traditional technology that only focuses on a single dimension of overhead and achieves more accurate performance estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806199A_ABST
    Figure CN120806199A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence chips, and discloses a model performance estimation method and device and computer equipment, and the method comprises the steps: determining model configuration data and candidate distributed strategies of a target model; converting the original calculation graph corresponding to the single-card deployment state based on the model configuration data and the strategy configuration data of the candidate distributed strategies to obtain a distributed overhead calculation graph corresponding to the multi-card deployment state; and performing performance estimation on the candidate distributed strategy according to the basic overhead and the additional overhead in the distributed overhead calculation graph to obtain strategy performance data of the target model under the candidate distributed strategy, thereby realizing conversion of a single-card model into a multi-card model. And based on the multi-card model, simulation calculation of the multi-card interconnection mode is realized on the premise of limited hardware resources, so that the influence of a communication operator and a topological structure corresponding to the multi-card interconnection mode in the overall operation of the model is reflected, and the upper limit of the model performance can be accurately evaluated in a simulator verification stage before silicon is applied.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence chip, and in particular to a model performance estimation method and device and computer equipment. BACKGROUND

[0002] A large model is a main application scenario of the current artificial intelligence chip. Therefore, how to obtain an upper limit of the performance of the large model by theoretical calculation in the design stage of the artificial intelligence chip is an important content for examining the design and verification of the artificial intelligence chip.

[0003] In the design and verification stage of the artificial intelligence chip, performance evaluation is one of the core links. In the related art, a hardware simulator is usually directly used to run a large model to evaluate the actual performance of the chip.

[0004] Since it is difficult to directly simulate a multi-card interconnection mode in the pre-silicon stage, an effective model performance estimation method is needed to predict the performance upper limit of the large model in the multi-card distributed deployment. SUMMARY

[0005] The present application provides a model performance estimation method, device and computer equipment, which solves the technical problem that the pre-silicon verification in the related art needs to rely on a hardware simulator to reproduce a large-scale distributed scenario, and provides a reasonable and effective performance estimation scheme for the design of the artificial intelligence chip.

[0006] In order to achieve the above purpose, the main technical scheme adopted by the present application includes: In a first aspect, the present application provides a model performance estimation method, which includes: determining model configuration data of a target model and a candidate distributed strategy; wherein the target model corresponds to a single-card deployment state and a multi-card deployment state in the simulator verification stage before silicon; and the single-card deployment state corresponds to an original computation graph; performing data splitting based on the model configuration data and strategy configuration data of the candidate distributed strategy, to construct distributed nodes for multi-card parallel execution based on single-card nodes in the original computation graph, to obtain a distributed overhead computation graph corresponding to the multi-card deployment state; wherein the nodes of the distributed overhead computation graph are used to represent the basic overhead of operation operators, and the edges of the distributed overhead computation graph are used to represent the additional overhead between connected operation operators; performing performance estimation on the candidate distributed strategy according to the distributed overhead computation graph, to obtain strategy performance data of the target model under the candidate distributed strategy.

[0007] Optionally, the step of constructing distributed nodes for multi-card parallel execution based on single-card nodes in the original computation graph to obtain a distributed overhead computation graph corresponding to the multi-card deployment state includes: transforming single-card nodes in the original computation graph into distributed nodes for multi-card parallel execution, and inserting a communication operator node in the original computation graph based on the strategy configuration data, to obtain the distributed overhead computation graph.

[0008] Optionally, the basic overheads include a computation overhead, a communication overhead, and a memory overhead, and the additional overheads include a re-sharding overhead. The performance estimation of the candidate distributed strategy according to the distributed overhead computation graph includes: The performance data of the target model under the candidate distributed strategy is obtained by aggregating the computation overhead, the communication overhead, the memory overhead, and the re-sharding overhead.

[0009] Optionally, when the candidate distributed strategy is a data parallel strategy, the distributed nodes for multi-card parallel execution are constructed based on the single-card nodes in the original computation graph, to obtain the distributed overhead computation graph corresponding to the multi-card deployment state, including: Each single-device node corresponding to an operation operator in the original computation graph is converted into a data parallel node; wherein, input tensors are split along a batch processing dimension to simulate the allocation of each shard data to a corresponding data parallel node; model weights are split in a full-shard data parallel manner to simulate the storage of a corresponding first shard weight by the data parallel node. A full-aggregation operator node is inserted in the original computation graph according to a data parallel configuration parameter, to obtain the distributed overhead computation graph.

[0010] Optionally, when the candidate distributed strategy is a tensor parallel strategy, the distributed nodes for multi-card parallel execution are constructed based on the single-card nodes in the original computation graph, to obtain the distributed overhead computation graph corresponding to the multi-card deployment state, including: Each single-device node corresponding to an operation operator in the original computation graph is converted into a tensor parallel node; wherein, the size of an input tensor remains unchanged, and model weights are split along a tensor parallel dimension to simulate the storage of a corresponding second shard weight by the tensor parallel node. A full-aggregation operator node and a full-reduction operator node are inserted in the original computation graph according to a tensor parallel configuration parameter, to obtain the distributed overhead computation graph.

[0011] Optionally, the candidate distributed strategy includes at least one of a data parallel strategy, a tensor parallel strategy, a pipeline parallel strategy, and a hybrid parallel strategy.

[0012] Optionally, the candidate distributed strategy is determined in any one of the following ways: In response to a policy configuration operation, the candidate distributed policy is generated; In response to a policy selection operation, the candidate distributed policy is generated; In response to a modification operation on an initial distributed policy, the candidate distributed policy is obtained; Hardware constraint data is input into a policy generator for generation of a distributed policy, and the candidate distributed policy is obtained.

[0013] Optionally, the method further comprises: A repeated module in the target model is determined; wherein the repeated module has marking information; Based on the marking information of the repeated module, the repeated module is split from the target model; A partial calculation subgraph of the repeated module is constructed; According to the number of the repeated modules and sub-computing performance data corresponding to the partial calculation subgraph, overall computing performance data of the target model is calculated.

[0014] Optionally, the repeated module is identified according to model configuration data of the target model.

[0015] Optionally, calculating the overall computing performance data of the target model according to the number of the repeated modules and the sub-computing performance data corresponding to the partial calculation subgraph comprises: Based on the number of the repeated modules, a performance calculation repetition number is configured; wherein the performance calculation repetition number is equal to the number of the repeated modules minus 1; According to the performance calculation repetition number, the sub-computing performance data is repeated to restore the overall computing performance data of the target model.

[0016] Optionally, the sub-computing performance data corresponding to the partial calculation subgraph is determined by: Operator performance data of each computing operator in the partial calculation subgraph is calculated; According to the operator performance data of each computing operator, the sub-computing performance data is obtained by summarizing.

[0017] In a second aspect, an embodiment of the present application provides a model performance estimation device, the device comprising: A distributed policy determination module is configured to determine model configuration data of a target model and a candidate distributed policy; wherein the target model corresponds to a single-card deployment state and a multi-card deployment state in a pre-silicon simulator verification stage; and the single-card deployment state corresponds to an original calculation graph. An overhead computation graph generation module is configured to perform data splitting based on the model configuration data and policy configuration data of the candidate distributed strategy, to construct distributed nodes for multi-card parallel execution based on single-card nodes in the original computation graph, and to obtain a distributed overhead computation graph corresponding to the multi-card deployment state; wherein a node of the distributed overhead computation graph is configured to represent basic overhead of an operation operator, and an edge of the distributed overhead computation graph is configured to represent additional overhead between connected operation operators. A distributed strategy estimation module is configured to perform performance estimation on the candidate distributed strategy based on the distributed overhead computation graph, to obtain policy performance data of the target model under the candidate distributed strategy.

[0018] In a third aspect, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored on the memory and executable on the artificial intelligence chip, and the artificial intelligence chip implements the method described in any of the above aspects when executing the computer program.

[0019] In the embodiments of the present application, the model configuration data of the target model and the candidate distributed strategy are determined; the original computation graph corresponding to the single-card deployment state is converted based on the model configuration data and the policy configuration data of the candidate distributed strategy, to obtain a distributed overhead computation graph corresponding to the multi-card deployment state; the performance of the candidate distributed strategy is estimated based on the basic overhead and the additional overhead in the distributed overhead computation graph, to obtain the policy performance data of the target model under the candidate distributed strategy, to realize the conversion of the single-card model to the multi-card model, and to realize the simulation calculation of the multi-card interconnection mode under the premise of limited hardware resources based on the multi-card model, to reflect the influence of the communication operator and the topology structure corresponding to the multi-card interconnection mode in the overall operation of the model, so as to accurately evaluate the upper limit of the model performance in the pre-silicon simulator verification stage. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the description of the embodiments or the prior art. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0021] Figure 1 The schematic structural diagram of the general-purpose graphics processor provided by the embodiments of the present application is shown in the figure; Figure 2a The flowchart of the model performance estimation method provided by the embodiments of the present application is shown in the figure; Figure 2b The schematic diagram of the original computation graph provided by the embodiments of the present application is shown in the figure; Figure 3 A flow chart of the model performance estimation method provided in the embodiment of the present application; Figure 4a A flow chart of the model performance estimation method provided in the embodiment of the present application; Figure 4b A schematic diagram of a distributed overhead calculation graph corresponding to the data parallel strategy provided in an embodiment of the present application; Figure 5a A flow chart of the model performance estimation method provided in the embodiment of the present application; Figure 5b A schematic diagram of a distributed overhead calculation graph corresponding to the tensor parallel strategy provided in an embodiment of the present application; Figure 6 A flow chart of the model performance estimation method provided in the embodiment of the present application; Figure 7 A schematic diagram of marking information of a repeating module is provided for an embodiment of the present application; Figure 8 A schematic diagram of the framework of the model performance estimation device provided in an embodiment of the present application; Figure 9 A schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0023] The model performance method used in this application is based on artificial intelligence (AI). AI is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence.

[0024] The artificial intelligence processor involved in this application can be any of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), and a GPGPU (General-Purpose Graphics Processing Unit). A general-purpose graphics processor (GPGPU) is used as an example for illustration.

[0025] Figure 1 The following is a schematic diagram of the structure of a general-purpose graphics processing unit (GPGPU). Figure 1 ,A general graphics processor is actually an array of programmable multiprocessors. For example, the programmable multiprocessor can be a streaming processor cluster (SPC), including Figure 1 The stream processor clusters shown are 1, M-1, and M, where M is a positive integer greater than 1. In general-purpose graphics processors, one stream processor cluster processes one computational task, or multiple stream processor clusters can process one computational task. Taking stream processor cluster 1 as an example, a stream processor cluster includes multiple execution units (EUs). Each processing unit includes an arithmetic logic unit (ALU) and a floating-point unit (FPU), and is responsible for executing specific computational tasks. Each processing unit includes thread local registers (TLRs) for storing source and destination data related to the computational task. The global memory buffer (GMB) connects the EUs to high-bandwidth memory (HBM) for data transfer, improving data transfer efficiency between the EUs and HBM. The HBM provides large-capacity, low-latency storage to support the high-performance computing requirements of GPUs. Instructions enter the SPC from the outside and are dispatched to each EU for execution. Input data is read from the HBM, transferred to the SPC via the GMB, and then distributed to the EUs for processing. The results (output data) processed by EU are returned to GMB through SPC and finally written to HBM.

[0026] Because a large-scale model can have hundreds of billions of parameters, a single machine or GPU is no longer sufficient for training such models. Therefore, distributed training is widely used for large-scale model training in artificial intelligence. Distributed training utilizes multiple machines to collaborate to accelerate the training of large deep learning models. In distributed training, the original model, dataset, and training process are broken down and distributed across multiple machines for simultaneous processing, effectively utilizing more computing resources and shortening training time.

[0027] For large models, communication computing time often becomes the main factor affecting model time consumption, but it is difficult for the simulator to deploy a complete 4-card interconnection all-in-one machine or 8-card interconnection all-in-one machine, and it can only deploy at most 2 cards, and the number of execution units is limited.

[0028] Further, in the training process of a large model, different distributed algorithms have a great influence on the performance of the model, and it is necessary to measure the performance of different algorithms by a certain method, analyze the performance bottleneck of different large models under different distributed algorithms, and give guidance and suggestions for the performance optimization of large models. The mainstream distributed algorithms include data parallel algorithm DP, tensor parallel algorithm TP and pipeline parallel algorithm PP.

[0029] The data parallel algorithm refers to that a training data set of a model is divided into multiple sub-data sets, the multiple sub-data sets after division are allocated to multiple devices, each device holds a complete model copy, so as to independently train the sub-data set allocated thereto, for example, each device can independently complete forward propagation and back propagation calculation of the sub-data set. For example, a training data set is divided into 3 parts and allocated to device 0, device 1 and device 2, each of device 0, device 1 and device 2 has a complete model. After each device completes back propagation, the gradients calculated by each device need to be aggregated and the global model parameters are updated.

[0030] The tensor parallel mode refers to that a layer (or operator) of a model is split, and part of the weights of the operator is placed on different devices, so as to reduce the memory occupation of each device (for example, GPU). A neural network model includes multiple layers, each layer can be understood as a function, which can perform specific mathematical operations on input tensors. For example, a fully connected layer can perform linear transformation operation on a tensor, a convolution layer can perform convolution operation on a tensor, and a pooling layer can perform pooling operation on a tensor, etc., to extract or convert features, these operations can change the content and shape of the tensor, and realize the step-by-step conversion from the original data to the high-level abstract representation. In the tensor parallel mode, the tensor of a layer is divided into multiple parts, and the multiple part tensors after division are allocated to multiple devices, for example, device 0, device 1 and device 2, each device is only responsible for part of the calculation of the model, and exchanges the necessary intermediate results through communication.

[0031] Pipeline parallel mode divides the different layers of a model into multiple stages, with each stage assigned to multiple devices to form a pipeline. Activation values ​​and gradients are passed sequentially between devices. For example, during forward propagation, each device passes activation values ​​to the device in the next pipeline stage. During backward propagation, each device passes gradients back to the device in the previous pipeline stage. This allows computations to be performed on multiple devices simultaneously, improving computational efficiency. For example, if a model consists of four layers, the four layers are divided into three stages, with each stage assigned to a device. For example, the first layer is assigned to device 0, the second layer to device 1, and the third and fourth layers to device 2.

[0032] As model size increases, using a single parallel mode often cannot simultaneously meet device memory limitations and high computational efficiency requirements. Therefore, for large-scale models, it is usually necessary to combine multiple parallel technologies such as data parallelism, tensor parallelism, and pipeline parallelism for distributed training.

[0033] Furthermore, Table 1 exemplifies the characteristics, advantages, and disadvantages of different distributed algorithms: Table 1: Comparison of various distributed algorithms In related technologies, training and inference of large models typically rely on multiple GPUs or distributed computing frameworks. During the chip design phase (pre-silicon), limited hardware simulators (such as FPGAs and Zebu) make it difficult to deploy a complete multi-GPU interconnected environment (such as a 4- or 8-GPU cluster). Furthermore, the performance of the simulator's multi-GPU cluster interconnection can differ significantly from that of actual machines, making it difficult to accurately assess the impact of communication overhead on performance. As shown in Table 1, distributed algorithms (such as data parallelism, tensor parallelism, and pipeline parallelism) each have their own advantages and disadvantages. Related technologies lack a systematic approach to simulate the communication overhead and performance bottlenecks of different distributed strategies in the pre-silicon phase.

[0034] Therefore, the embodiment of the present application proposes a model performance estimation method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here. Figure 2a , Figure 2a The figure shows a flow chart of the model performance estimation method in the embodiment of the present application. The model performance estimation method includes the following steps: S110: Determine model configuration data and candidate distribution strategies of the target model.

[0035] The target model corresponds to a single-card deployment state and a multi-card deployment state in the simulator verification stage before silicon. The single-card deployment state corresponds to an original computation graph. The candidate distributed strategy includes at least one of a data parallel strategy, a tensor parallel strategy, a pipeline parallel strategy, and a hybrid parallel strategy.

[0036] In this embodiment, the target model refers to an artificial intelligence model to be evaluated for performance, which includes executable computing logic and associated parameters. Specifically, the target model can be a large language model based on a Transformer architecture (such as Llama, GPT series), or a convolutional neural network model (such as ResNet). The model configuration data of the target model refers to description data describing the model structure and parameters, including but not limited to the number of model layers, the number of parameters per layer, the type of activation function, the memory occupancy, etc.

[0037] In this embodiment, the target model has two deployment forms. The single-card deployment state refers to a state in which the model runs on a single computing device (such as a GPU), and the multi-card deployment state refers to a state in which the model runs on multiple computing devices through a distributed strategy. For example, in the single-card deployment state, the target model is completely loaded on a GPU; in the multi-card deployment state, the target model is split into 4 GPUs through a distributed strategy.

[0038] In this embodiment, the candidate distributed strategy is a set of parallel computing schemes designed to achieve multi-card deployment. The number of candidate distributed strategies can be one or multiple. Specifically, the candidate distributed strategy includes basic strategies such as data parallel strategy (DP), tensor parallel strategy (TP), pipeline parallel strategy (PP), and expert parallel strategy (EP), or a hybrid parallel strategy thereof. The hybrid parallel strategy can be a 3D parallel strategy, i.e., a hybrid strategy of DP, TP, and PP.

[0039] In this embodiment, the model configuration data of the target model and the candidate distributed strategy are determined to provide input parameters for subsequent distributed overhead computation graph conversion. For example, the model configuration data is used to quantify the basic overhead of the operator, and the candidate distributed strategy is used to define the structure of the distributed overhead computation graph.

[0040] In this embodiment, the candidate distributed strategy is determined in any of the following ways: In response to a strategy configuration operation, the candidate distributed strategy is generated.

[0041] In response to a strategy selection operation, the candidate distributed strategy is generated.

[0042] In response to a modification operation on the initial distributed strategy, the candidate distributed strategy is obtained.

[0043] The hardware constraint data is input into the strategy generator for generation of a distributed strategy, to obtain a candidate distributed strategy.

[0044] The above-described determination manner of the candidate distributed strategy can support various different machine topology states and modification of the distributed strategy on the model calculation graph, and allow automatic selection or manual configuration of the distributed strategy.

[0045] S120, based on the model configuration data and the strategy configuration data of the candidate distributed strategy, converting the original calculation graph corresponding to the single-card deployment state to obtain a distributed overhead calculation graph corresponding to the multi-card deployment state.

[0046] The node of the distributed overhead calculation graph is used to represent the basic overhead of the operation operator, and the edge of the distributed overhead calculation graph is used to represent the additional overhead between the connected operation operators.

[0047] In the embodiment, the strategy configuration data is a set of technical parameters for implementing a specific candidate distributed strategy. The original calculation graph refers to an abstract representation of the calculation process of the model in the single-card deployment state. For example, refer to FIG. 1, which shows a schematic diagram of the original calculation graph. The original calculation graph can include operators (such as matrix multiplication, convolution) and their input-output relationships. Figure 2b , Figure 2b The original calculation graph can include operators (such as matrix multiplication, convolution) and their input-output relationships.

[0048] In the embodiment, the distributed overhead calculation graph refers to a calculation process representation obtained by converting the original calculation graph according to the distributed strategy. The node and the edge of the distributed overhead calculation graph represent the basic overhead and the additional overhead in the distributed deployment, respectively. For example, the node of the distributed overhead calculation graph can correspond to the calculation time of the operator (such as the MAC times of matrix multiplication), and the edge can correspond to the communication time across devices (such as the delay of the All-Reduce operation). It should be noted that the distributed overhead calculation graph adopts a data form of a graph structure, and the distributed overhead calculation graph is composed of two basic elements, namely the node and the edge, which work together to form the basis of the mathematical model of the pre-silicon performance estimation. The distributed overhead calculation graph can reflect the calculation and data flow logic of the target model in the multi-card deployment state in the pre-silicon simulator verification phase. The distributed overhead calculation graph is embedded with complete overhead information, which facilitates the pre-silicon quantitative simulation of the performance of the distributed execution. The node in the distributed overhead calculation graph corresponds to a basic calculation operator or a communication operator, and is marked with calculation overhead, communication overhead, and memory overhead. The edge in the distributed overhead calculation graph represents the data dependency relationship between the nodes, and is marked with a type of additional overhead, such as the re-sharding cost (Re-sharding Cost) caused by the mismatch of the strategy.

[0049] It can be understood that the technical problem that the multi-card model cannot be actually run in the pre-silicon stage is solved by expressing the candidate distributed strategy as a concrete and quantifiable graph structure through the distributed overhead calculation graph. The process of determining the distributed overhead calculation graph can be regarded as constructing a performance simulation model corresponding to the candidate distributed strategy. Through the edges and nodes of the model, the performance influencing factors in the multi-card distributed environment can be completely modeled.

[0050] In the embodiment, the basic overhead of the operation operator refers to the calculation time, communication time and memory condition of the operation operator on the multi-card. The additional overhead refers to the additional time consumption introduced by cross-device communication or synchronization in the distributed deployment.

[0051] In some embodiments, the original calculation graph corresponding to the single-card deployment state is converted based on the model configuration data and the strategy configuration data of the candidate distributed strategy to obtain the distributed overhead calculation graph corresponding to the multi-card deployment state, including step S122. In step S122, data splitting is performed based on the model configuration data and the strategy configuration data of the candidate distributed strategy to construct distributed nodes for multi-card parallel execution based on the single-card nodes in the original calculation graph, and the distributed overhead calculation graph corresponding to the multi-card deployment state is obtained.

[0052] Specifically, the operation operator corresponding to the single-card deployment state is determined based on the model configuration data, and the operation operator is represented as a node to construct the original calculation graph. In order to realize the distributed allocation of the calculation load, the tensor data and / or weight data of the single-card node are split and allocated to multiple computing devices to construct distributed nodes for multi-card parallel execution based on the single-card nodes in the original calculation graph. For example, the distributed nodes for multi-card parallel execution can not only be the distributed nodes obtained by converting or transforming the single-card nodes in the original calculation graph, but also the corresponding distributed nodes newly added on the basis of the single-card nodes in the original calculation graph. It can be understood that whether the distributed nodes are the single-card nodes transformed or the newly added distributed nodes, they are all nodes corresponding to the operation operator.

[0053] Further, a basic overhead or a base overhead is set for a node corresponding to an operation operator, and an additional overhead or an extra overhead caused by cross-device deployment is set for an edge connecting the operation operators, to obtain a distributed overhead computation graph corresponding to the multi-card deployment state. By converting the original computation graph into the distributed overhead computation graph, the conversion from the single-card model to the multi-card model is realized, and the performance upper limit of the target model under the multi-card distributed strategy can be predicted without actual hardware deployment, which is beneficial to evaluating the influence of different interconnection architectures on the model performance in the pre-silicon stage, thereby optimizing the hardware design. Illustratively, the strategy configuration data of the candidate distributed strategy can be parsed to determine the communication operation corresponding to the candidate distributed strategy, and the corresponding operator is added to the original computation graph based on the communication operation corresponding to the candidate distributed strategy, for example, a node representing a communication operator is added to the original computation graph.

[0054] In S130, performance of the candidate distributed strategy is estimated according to the distributed overhead computation graph, to obtain strategy performance data of the target model under the candidate distributed strategy.

[0055] In this embodiment, the performance estimation refers to the process of quantitatively evaluating the running efficiency of the candidate distributed strategy based on the distributed overhead computation graph. The strategy performance data refers to the result of the performance estimation, including but not limited to total time consumption, throughput, speedup ratio or video memory occupancy. Specifically, after the original computation graph is converted into the distributed overhead computation graph, the total overhead of the target model under the candidate distributed strategy is calculated, including the sum of the basic overheads of the nodes and the sum of the additional overheads of the edges.

[0056] Further, the basic overhead can include a calculation overhead, a communication overhead and a video memory overhead, and the additional overhead can include a re-sharding overhead. Specifically, the calculation overhead can refer to the equivalent single-card calculation overhead after the calculation of a single device is deployed to multiple devices. For example, under the distributed strategy, each device (such as a GPU) processes part of the calculation task, and the calculation overhead can quantify the calculation burden of each device. Illustratively, the calculation overhead can be the pure calculation time consumption or calculation time of the corresponding operation operator executed on the calculation device marked in the node in the distributed overhead computation graph. The calculation overhead can be distinguished according to different operator types. The calculation bottleneck types of different operation operators are identified and modeled to determine the calculation overhead. Different calculation bottleneck types: 1) vcore instruction emission bottleneck operator, such as a dwc operator; 2) tcore calculation bottleneck, such as a general matrix multiplication (gemm) operator; 3) memory access bottleneck, such as a pointwise operator; 4) for some complex operators, such as self-attention operators (attention), if there is a data dependency, the operator needs to be segmented for processing according to the algorithm. According to the calculation and memory access overheads of different operators, the slowest position is taken, and the calculation overhead can be calculated.

[0057] The communication overhead relates to the communication between devices, such as the All-Reduce operation. The communication overhead refers to the communication time consumed by the synchronization of data or the aggregation of data between devices after the target model is distributed to each device according to the distributed strategy in the multi-card deployment state (such as the All-Reduce and All-Gather operations). Exemplarily, the communication overhead can be the time consumed by the communication operation (such as the All-Reduce and All-Gather operations) corresponding to part of the nodes in the distributed overhead computation graph in the corresponding topology state. There are various protocols for low latency communication or Simple, and there are ring or tree algorithms according to different topologies. For a certain communication algorithm, the performance is dominated by delay or bandwidth according to the size of the data block to be transmitted. If it is bandwidth dominated, the time consumption is calculated according to the bandwidth of the topology, and if it is delay dominated, the delay time is directly calculated.

[0058] The memory overhead can refer to the memory occupation (such as parameter storage and intermediate activation value storage) of the target model in the multi-card deployment state, and the memory occupation of each device after being distributed to each device according to the distributed strategy. Exemplarily, the memory overhead corresponds to the space size mainly depending on the weight parameter and the calculation data size after the distributed strategy is cut in the operation, including the optimizer space occupation in the training. In addition, part of the buffer space for distributed communication needs to be left.

[0059] The reshuffling overhead can refer to the additional time consumption caused by the data distribution method (such as sharding and replication) or the calculation order adjustment (such as pipeline stage division) in the distributed strategy. For example, in the tensor parallel strategy, the input data needs to be split and rearranged on different devices, which introduces additional memory copying and synchronization operations, and the overhead is the reshuffling overhead. The reshuffling overhead is used to quantify the additional requirements of data preprocessing and post-processing in the distributed strategy. Exemplarily, the reshuffling overhead can be understood as an implicit communication operator. If the data cutting methods of multiple operators are different in the distribution, the send / receive operation in the distribution needs to be performed when the interaction occurs.

[0060] Correspondingly, the performance of the target model under the candidate distributed strategy is estimated according to the distributed overhead computation graph, and the performance data of the target model under the candidate distributed strategy is obtained, including: the performance data of the target model under the candidate distributed strategy is obtained by aggregating the calculation overhead, the communication overhead, the memory overhead and the reshuffling overhead.

[0061] Specifically, the computing overhead, the communication overhead, the memory overhead and the re-sharding overhead are integrated into a quantifiable comprehensive performance index as performance data of the target model under the candidate distributed strategy. The performance data reflects the influence degree of different overheads on the performance of the model. The performance of the candidate distributed strategy is comprehensively evaluated, thereby providing multi-dimensional optimization basis for the chip design stage, such as selecting the candidate distributed strategy with the minimum overhead. By subdividing and quantifying various overheads in the distributed strategy, the limitation of the traditional technology of only focusing on the single-dimensional overhead of computing is solved. By summarizing various basic overheads and additional overheads, data basis is provided for the dynamic adjustment of the performance evaluation model, thereby realizing more accurate performance estimation.

[0062] In the above embodiments, the model configuration data of the target model and the candidate distributed strategy are determined; the original computation graph corresponding to the single-card deployment state is converted based on the model configuration data and the strategy configuration data of the candidate distributed strategy to obtain a distributed overhead computation graph corresponding to the multi-card deployment state; the performance of the candidate distributed strategy is estimated according to the basic overhead and the additional overhead in the distributed overhead computation graph to obtain the strategy performance data of the target model under the candidate distributed strategy, thereby realizing the conversion of the single-card model into the multi-card model and the simulation calculation of the multi-card interconnection mode based on the multi-card model under the premise of limited hardware resources, so as to reflect the influence of the communication operator and the topology structure corresponding to the multi-card interconnection mode in the overall operation of the model, thereby accurately evaluating the upper limit of the model performance in the pre-silicon simulator verification stage.

[0063] In some embodiments, the distributed nodes for multi-card parallel execution are constructed based on the single-card nodes in the original computation graph to obtain the distributed overhead computation graph corresponding to the multi-card deployment state, including: converting the single-card nodes in the original computation graph into the distributed nodes for multi-card parallel execution, and inserting the communication operator nodes in the original computation graph based on the strategy configuration data to obtain the distributed overhead computation graph. Further, please refer to Figure 3 The original computation graph corresponding to the single-card deployment state is converted based on the model configuration data and the strategy configuration data of the candidate distributed strategy to obtain the distributed overhead computation graph corresponding to the multi-card deployment state, which can include: S310, data slicing is performed based on the model configuration data and the strategy configuration data to convert the single-card nodes in the original computation graph into the distributed nodes for multi-card parallel execution.

[0064] S320, the communication operator nodes are inserted in the original computation graph based on the strategy configuration data to obtain the distributed overhead computation graph.

[0065] The single-card node refers to an operation operator node in the original computation graph designed to run on a single computing device. The distributed node for multi-card parallel execution can be an operation operator node in a distributed overhead computation graph designed to run on multiple computing devices. The communication operator node can be a communication operation unit inserted for coordinating the execution of the distributed node. Specifically, the type of the communication operator node is defined by the strategy configuration data. Exemplarily, the communication operator node can be a collective communication node.

[0066] Specifically, to realize the distributed allocation of the computing load, the single-card node in the original computation graph is converted into the distributed node for multi-card parallel execution by splitting the tensor data and / or weight data of the single-card node and distributing them to multiple computing devices, to generate a computation flow representation in a distributed deployment scenario. Further, a communication operation is added at a specific position of the computation graph, for example, a TP strategy: an All-Gather node is inserted after a matrix multiplication node. In this way, the graph structure is reconstructed, and the data flow connection relationship is changed. For example, the original computation graph is reconstructed from "computation node A -> computation node B" to "computation node A -> All-Reduce node -> computation node B". Further, overheads corresponding to the computation operator nodes and the communication operator nodes are added or labeled, to obtain a distributed overhead computation graph. In this embodiment, the completeness of the distributed computation is realized by inserting the communication operator node, and the data coordination process between devices is explicitly expressed. Therefore, the communication overhead can be quantified, and the performance bottleneck can be revealed.

[0067] It should be noted that, for the PP strategy, the operators in the original computation graph are divided onto different devices, without involving changes to the computation graph.

[0068] In this embodiment, when the candidate distributed strategy is the data parallel strategy, refer to Figure 4a , the distributed node for multi-card parallel execution is constructed based on the single-card node in the original computation graph, to obtain a distributed overhead computation graph corresponding to the multi-card deployment state, including: S410, converting each single-device node corresponding to an operation operator in the original computation graph into a data parallel node.

[0069] S420, inserting an all-gather operator node into the original computation graph according to the data parallel configuration parameter, to obtain a distributed overhead computation graph.

[0070] Specifically, refer to

[0071] Specifically, refer to Figure 4b , Figure 4bA schematic diagram of the distributed overhead computation graph is shown. First, the computation distribution under the data parallel strategy is realized by converting the single-device nodes of the original computation graph into data parallel nodes and splitting the input tensor along the batch dimension; second, the model weights are split by the full-sharded data parallel method to ensure that each device only stores and calculates the corresponding shard parameters; finally, the parameter synchronization process in distributed training is simulated by inserting an all-gather operator node. The generated distributed overhead computation graph can accurately represent the computation time, communication delay and memory occupancy under the data parallel strategy, and can more comprehensively evaluate the comprehensive demand of the data parallel strategy for hardware resources, thereby providing a theoretical basis for performance optimization in the chip design stage.

[0072] In this embodiment, when the candidate distributed strategy is the tensor parallel strategy, refer to Figure 5a , the distributed nodes for multi-card parallel execution are constructed based on the single-card nodes in the original computation graph, and the distributed overhead computation graph corresponding to the multi-card deployment state is obtained, including: S510, converting each operation operator node in the original computation graph into a tensor parallel node.

[0073] S520, inserting all-gather operator nodes and all-reduce operator nodes in the original computation graph according to the tensor parallel configuration parameters, to obtain the distributed overhead computation graph.

[0074] Among them, the size of the input tensor remains unchanged, and the model weights are split along the tensor parallel dimension to simulate the storage of the corresponding second shard weights by the tensor parallel node. The all-gather operator node (All-Gather) is a collective communication node used to collect scattered data, and its function is to splice the shard data of each device into a complete tensor. For example, in tensor parallel, the partial output tensors of multiple devices are spliced into a complete result. The all-reduce operator node (All-Reduce) is a collective communication node used to aggregate data across devices, and its function is to perform a summation operation on the input tensors of each device and broadcast the result, for example, in the gradient synchronization scenario, the gradient tensors of each device are summed.

[0075] Specifically, refer to Figure 5b , Figure 5bAn illustrative diagram of a distributed overhead computation graph is shown. First, by converting the single-device nodes of the original computation graph into tensor parallel nodes, the size of the input tensor remains unchanged, and the model weights are split along the tensor parallel dimension to simulate the corresponding second split weights of the tensor parallel node. During the forward propagation phase, the partial results calculated by each device need to be aggregated through an all-gather operator (such as All-Gather) to ensure that the output results on all devices are consistent; during the backward propagation phase, the gradients calculated by each device need to be synchronized through an all-reduce operator (such as All-Reduce) to ensure the global consistency of the model parameters. Finally, the tensor parallel nodes and the communication operator nodes are integrated to form a complete distributed computation flow representation. For example, the nodes of the distributed overhead computation graph represent the basic overhead of the operation operator, and the edges represent the additional overhead (such as the delay of All-Gather) of cross-device communication. The technical effect of this step is that through theoretical modeling and algorithm simulation, the performance upper limit of the target model under the multi-card tensor parallel strategy can be predicted without actual hardware deployment.

[0076] In some embodiments, referring to Figure 6 the method further comprises the following steps: S610, determining a repeated module in the target model.

[0077] The repeated module refers to a functional unit in the target model that is used repeatedly. The repeated module is identified according to the model configuration data of the target model, and is used to distinguish the position and attributes of different repeated modules. Referring to Figure 7 , the repeated module has a marker information 702; the marker information (scope) is used to uniquely identify the repeated module. The marker information is used to uniquely identify the position and attributes of the repeated module, ensuring the accuracy of the splitting operation. Specifically, the functional unit that can be reused in the model structure is identified, thereby providing a basis for subsequent performance estimation. For example, by parsing the configuration file of the target model, it can be determined that the 61 layers of Transformer are repeated modules, and each layer contains the same operator combination (such as linear layer, attention mechanism).

[0078] S620, based on the marker information of the repeated module, splitting the repeated module from the target model.

[0079] Specifically, based on the marker information, the repeated module in the target model is extracted from the overall structure to form an independent functional unit. The splitting operation is implemented by parsing the configuration file of the target model, ensuring the integrity and independence of each repeated module. For example, in this embodiment, if the target model contains 61 layers of Transformer, then the __module.layers.1 to __module.layers.61 can be split layer by layer through the marker information to generate 61 independent repeated modules.

[0080] S630: Construct a partial computation subgraph of the repeated module.

[0081] In this embodiment, a partial computation subgraph refers to a small-scale computation graph consisting of a single repeating module and its internal operators. It is used to quantify the module's performance metrics (such as computational effort, memory usage, and communication overhead). The construction process of a partial computation subgraph involves extracting the operator structure of the repeating module, defining the input and output dimensions, and generating the corresponding computation flow.

[0082] S640: Perform performance estimation based on the number of repeated modules and the sub-computation performance data corresponding to some of the computational subgraphs to determine the overall computational performance data of the target model.

[0083] Among them, sub-computing performance data refers to the summary results of the performance indicators of each operator in some computing subgraphs, including computing power (FLOPs), memory usage (video memory requirements), and communication overhead. Overall computing performance data refers to the theoretical performance upper limit of the target model, including total computing power (TFLOPs), peak video memory usage (GB), and communication overhead (GB / s) in distributed training. Overall computing performance data is used to evaluate the adaptability of artificial intelligence chips or clusters. Specifically, sub-computing performance data is obtained by traversing the operators in some computing subgraphs. Through linear superposition or nonlinear correction algorithms, performance is summarized and extrapolated based on the number of repeated modules and the sub-computing performance data corresponding to some computing subgraphs to obtain the overall computing performance data of the target model.

[0084] In the above-mentioned embodiment, the resource consumption of the target model performance estimation can be significantly reduced (such as reducing the memory usage from TB level to the video memory required for several layers of calculation graphs), while supporting the rapid testing of the impact of different architectural changes (such as adjusting the number of layers, parameter scale) on performance. In this embodiment, the operation of performing performance inference based on the sub-computational performance data of the repeated module is simple to implement and applicable to all types of current large models. It can very conveniently test the impact of different architectural changes on the performance of large models. Moreover, when a complete calculation graph is not required, it is easy to replace the input batch, sequence, etc., and a new small calculation graph can be constructed in real time, and then the performance of the complete calculation graph can be calculated. For large models that are difficult to fully deploy to the server, deployment can be completed, and for large models that have already been deployed, efficiency can be improved.

[0085] In some embodiments, performance is extrapolated based on the number of repeated modules and the sub-computing performance data corresponding to some computing subgraphs to determine the overall computing performance data of the target model, including: configuring the number of performance calculation repetitions based on the number of repeated modules; repeating the sub-computing performance data according to the number of performance calculation repetitions to restore the overall computing performance data of the target model.

[0086] The performance calculation repetition number is equal to the number of repeated modules minus 1. The number of repeated modules refers to the total number of repeated modules with the same structure and function in the target model. The performance calculation repetition number refers to a parameter configured based on the number of repeated modules in the target model, used for calculating the repetition number when calculating the overall calculation performance data. The parameter is obtained by subtracting 1 from the number of repeated modules, and the purpose is to avoid additional counting of the first repeated module, so as to more accurately reflect the influence of module repetition on the overall performance. For example, in the embodiment, if the target model contains 61 repeated modules, the performance calculation repetition number is 61 minus 1, which is equal to 60, indicating that the performance of the subsequent 60 repeated modules needs to be superimposed on the performance of the first module by repeated calculation.

[0087] Specifically, the overall calculation performance data of the target model is restored by linear superposition. This step simulates the performance superposition effect of multiple repeated modules by multiplying the sub-computation performance data by the performance calculation repetition number. Through the combination of the performance calculation repetition number and the sub-computation performance data, the theoretical performance upper limit of the target model is generated, including the total calculation amount, the peak memory occupation, and the communication overhead in distributed training.

[0088] In some embodiments, the sub-computation performance data corresponding to the partial computation subgraph is determined by: calculating the operator performance data of each computation operator in the partial computation subgraph; and aggregating the operator performance data of each computation operator to obtain the sub-computation performance data.

[0089] It can be understood that in the above embodiments, on the one hand, since the complete computation graph is not required, resource consumption is reduced, and memory demand is reduced by multiple orders of magnitude. On the other hand, input parameters (such as Batch Size and Sequence Length) can be adjusted in real time, a new computation graph can be quickly generated, and calculation efficiency is improved.

[0090] It should be noted that for different distributed strategies, this method of splitting computation and restoration has no additional impact: 1) Impact on single-machine single-card performance prediction: no impact on single-machine single-card 2) Impact on DP performance prediction: under the premise of data parallelism, the performance of each single-card machine repeated layer is the same, so there is no impact.

[0091] 3) Impact on TP performance prediction: TP cuts each layer horizontally and distributes it to multiple cards. For hidden layers, the TP configuration is generally the same, and it can be assumed that the performance of each hidden layer is also the same, so there is no impact.

[0092] 4) Impact on PP performance prediction: When PP, the layers of the full model are assigned to different machines, i.e. each layer is computed on a full machine, the performance of the same layer is artificially the same, so there is no impact.

[0093] It can be seen that the above splitting, computing and restoring manner can be compatible with different distributed strategies (DP, TP and PP) without additional adaptation.

[0094] The embodiment of the present application provides a model performance estimation device, Figure 8 FIG. 8 is a schematic diagram of a framework of the model performance estimation device provided by the embodiment of the present application. The model performance estimation device 800 can include a distributed strategy determination module 810, an overhead computation graph generation module 820 and a distributed strategy estimation module 830.

[0095] The distributed strategy determination module 810 is configured to determine model configuration data of a target model and a candidate distributed strategy; wherein the target model corresponds to a single-card deployment state and a multi-card deployment state in a pre-silicon simulator verification stage; the single-card deployment state corresponds to an original computation graph; The overhead computation graph generation module 820 is configured to perform data splitting based on the model configuration data and strategy configuration data of the candidate distributed strategy, to construct distributed nodes for multi-card parallel execution based on single-card nodes in the original computation graph, and to obtain a distributed overhead computation graph corresponding to the multi-card deployment state; wherein a node of the distributed overhead computation graph is used to represent basic overhead of an operation operator, and an edge of the distributed overhead computation graph is used to represent additional overhead between connected operation operators. The distributed strategy estimation module 830 is configured to perform performance estimation on the candidate distributed strategy according to the distributed overhead computation graph, to obtain strategy performance data of the target model under the candidate distributed strategy.

[0096] For the convenience of description, the above device is described in various modules respectively. Of course, the functions of the modules can be implemented in the same or multiple software and / or hardware in the implementation of the present application.

[0097] The embodiment of the present application provides a non-transitory computer readable storage medium. The storage medium can be a non-transitory computer readable storage medium, and one or more computer readable instructions can be non-transitorily stored on the storage medium. For example, when the computer readable instructions are executed by a processor, one or more steps of the above model performance estimation method can be performed. The storage medium can be applied to an electronic device, for example, the storage medium can include a storage device in the electronic device.

[0098] The storage device can include any combination of one or more computer program products, which can include various forms of computer-readable storage media, such as volatile and / or non-volatile computer-readable media. For example, volatile computer-readable media can include random access memory (RAM), and / or cache memory, etc. Non-volatile computer-readable media can include read only memory (ROM), hard disks, erasable programmable read only memory (EPROM), portable compact disc read only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored on the computer-readable storage media, and the processor can execute the computer-readable instructions to implement various functions of the processor. Various application programs and various data, etc. can also be stored in the storage media.

[0099] The storage medium can include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), portable compact disc read only memory (CD-ROM), flash memory, or any combination of the above storage media, or other applicable storage medium.

[0100] The embodiment of the present application provides a computer device, as shown in the accompanying drawings, which comprises at least one artificial intelligence chip 1510 and a memory 1520 connected with the at least one artificial intelligence chip 1510. In the embodiment of the present application, the specific connection medium between the artificial intelligence chip 1510 and the memory 1520 is not limited. Figure 9 In the embodiment of the present application, the memory 1520 stores instructions executable by the at least one artificial intelligence chip 1510, and the at least one artificial intelligence chip 1510 can execute the steps of the model performance estimation method by executing the instructions stored in the memory 1520. Figure 9

[0101] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0102] ​The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable

[0103] It should be noted that the term "comprising" or "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element preceded by "comprises a" or "comprises" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0104] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0105] The above merely describes the embodiments of the present application, and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.

[0106] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes shall fall within the scope defined by the appended claims.

Claims

1. A model performance estimation method, characterized in that: The method comprises: Determining model configuration data and candidate distributed strategies for a target model; wherein the target model corresponds to a single-card deployment state and a multi-card deployment state during a pre-silicon simulator verification phase; and the single-card deployment state corresponds to an original computation graph; Data segmentation is performed based on the model configuration data and the policy configuration data of the candidate distributed policy to construct a distributed node for multi-card parallel execution based on the single-card node in the original computation graph, thereby obtaining a distributed cost computation graph corresponding to the multi-card deployment state; wherein the nodes of the distributed cost computation graph are used to represent the basic cost of the operation operator, and the edges of the distributed cost computation graph are used to represent the additional cost between connected operation operators; The performance of the candidate distributed strategy is estimated according to the distributed cost calculation graph to obtain strategy performance data of the target model under the candidate distributed strategy.

2. The method according to claim 1, characterized in that The step of constructing a distributed node for multi-card parallel execution based on a single-card node in the original calculation graph to obtain a distributed cost calculation graph corresponding to the multi-card deployment state includes: The single-card node in the original calculation graph is converted into a distributed node for multi-card parallel execution, and the communication operator node is inserted into the original calculation graph based on the policy configuration data to obtain the distributed cost calculation graph.

3. The method according to claim 1, characterized in that The basic overhead includes computing overhead, communication overhead and video memory overhead, and the additional overhead includes resharding overhead caused by policy mismatch; The performing performance estimation on the candidate distributed strategy according to the distributed cost calculation graph to obtain performance data of the target model under the candidate distributed strategy includes: Performance data of the target model under the candidate distributed strategy is obtained by summarizing the computational overhead, the communication overhead, the video memory overhead, and the resharding overhead.

4. The method according to claim 1, wherein When the candidate distributed strategy is a data parallel strategy, constructing a distributed node for multi-card parallel execution based on a single-card node in the original calculation graph to obtain a distributed cost calculation graph corresponding to the multi-card deployment state includes: Converting the single device node corresponding to each operator in the original computation graph into a data parallel node; wherein the input tensor is split along the batch dimension to simulate the distribution of each shard data to the corresponding data parallel node; the model weight is split using a full-shard data parallel method to simulate the data parallel node storing the corresponding first shard weight; Insert the full aggregation operator node into the original calculation graph according to the data parallel configuration parameters to obtain the distributed cost calculation graph.

5. The method according to claim 1, wherein When the candidate distributed strategy is the tensor parallel strategy, constructing a distributed node for multi-card parallel execution based on the single-card node in the original calculation graph to obtain a distributed cost calculation graph corresponding to the multi-card deployment state includes: Converting the single device node corresponding to each operator in the original computation graph into a tensor parallel node; wherein the size of the input tensor remains unchanged, and the model weight is split along the tensor parallel dimension to simulate the tensor parallel node storing the corresponding second shard weight; According to the tensor parallel configuration parameters, a full aggregation operator node and a full reduction operator node are inserted into the original calculation graph to obtain the distributed cost calculation graph.

6. The method according to any one of claims 1 to 5, characterized in that The candidate distribution strategy is determined by any one of the following methods: generating the candidate distributed policy in response to a policy configuration operation; generating the candidate distributed policies in response to a policy selection operation; In response to a modification operation on the initial distributed strategy, obtaining the candidate distributed strategy; The hardware constraint data is input into a policy generator to generate a distributed policy, thereby obtaining the candidate distributed policy.

7. The method according to claim 1, characterized in that The method further comprises: Determining a repeated module in the target model; wherein the repeated module has tag information; splitting the repeated modules from the target model based on the tag information of the repeated modules; Constructing a partial computational subgraph of the repeated module; The overall computing performance data of the target model is determined by performing performance estimation based on the number of the repeated modules and the sub-computation performance data corresponding to the partial computing subgraphs.

8. The method according to claim 7, characterized in that The repeating module is identified based on the model configuration data of the target model.

9. The method according to claim 7, characterized in that The performing performance estimation based on the number of the repeated modules and the sub-computation performance data corresponding to the partial computation subgraphs to determine the overall computation performance data of the target model includes: Configuring the number of performance calculation repetitions based on the number of the repetitive modules; wherein the number of performance calculation repetitions is equal to the number of the repetitive modules minus 1; The sub-computation performance data is repeated according to the number of performance calculation repetitions to restore the overall calculation performance data of the target model.

10. The method according to claim 7, characterized in that The sub-computation performance data corresponding to the partial computation subgraph is determined in the following manner: Calculating operator performance data of each computing operator in the partial computing subgraph; The sub-computing performance data is obtained by summarizing the operator performance data of each computing operator.

11. A model performance estimation device, characterized in that: The device comprises: A distributed strategy determination module is configured to determine model configuration data and candidate distributed strategies for a target model; wherein the target model corresponds to a single-card deployment state and a multi-card deployment state during a pre-silicon simulator verification phase; and the single-card deployment state corresponds to an original computation graph; A cost calculation graph generation module is configured to perform data segmentation based on the model configuration data and the policy configuration data of the candidate distributed policy, so as to construct a distributed node for multi-card parallel execution based on the single-card node in the original calculation graph, and obtain a distributed cost calculation graph corresponding to the multi-card deployment state; wherein the nodes of the distributed cost calculation graph are used to represent the basic cost of the operation operator, and the edges of the distributed cost calculation graph are used to represent the additional cost between the connected operation operators; A distributed strategy estimation module is used to perform performance estimation on the candidate distributed strategy according to the distributed cost calculation graph, and obtain strategy performance data of the target model under the candidate distributed strategy.

12. A computer device, characterized in that: The method comprises a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip, wherein the artificial intelligence chip implements the method according to any one of claims 1 to 10 when executing the computer program.

Citation Information

Patent Citations

  • Distributed system modeling method, device, equipment, medium and program product

    CN119046122A

  • Distributed system cost evaluation method and device, equipment, medium and product

    CN119046124A

  • Method for evaluating model training performance, computing device, storage medium and computer program product

    CN119201651A

  • Machine learning model hybrid parallel strategy automatic search method and system

    CN119250232A

  • Program Module Applicability Analyzer for Software Development and Testing for Multi-Processor Environments

    US20140007043A1

Cited By

  • Performance analysis method and device for GPGPU

    CN121233412A

  • Model parameter processing method, system on chip and model parameter processing system

    CN121543651A

  • Construction method of reasoning simulation model, data processing method and related products

    CN121880035A

  • Strategy optimization method and device, equipment, storage medium and computer program product

    CN121900978A