Model simulation

By constructing virtual model simulation conditions for training, the problems of high cost and low efficiency in machine learning model training are solved, and the accuracy of model simulation and resource consumption are optimized.

WO2025253251A1PCT designated stage Publication Date: 2025-12-11CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055584
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2025-05-30
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing machine learning models suffer from high training costs and low training efficiency, necessitating a realistic and accurate model simulation technique to optimize model parameters and hardware aspects.

Method used

By determining the model simulation parameters and environmental simulation parameters of the target model, virtual model simulation conditions are constructed for simulation training. The target model is then trained using simulated behavior to determine computational and communication behaviors. A workload execution graph is then constructed for simulation to obtain the simulation results of the target model.

Benefits of technology

This reduces the resource and time consumption of model simulation, ensures the accuracy and realism of the simulation, and optimizes the performance of model parameters and hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055584_11122025_PF_FP_ABST
    Figure IB2025055584_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a model simulation method and system. The model simulation method comprises: determining model simulation parameters of a target model and environment simulation parameters for simulating the target model; on the basis of the model simulation parameters and the environment simulation parameters, determining a simulation behavior for performing model simulation training on the target model; using the simulation behavior to perform model simulation training on the target model, so as to determine a computation behavior and a communication behavior during the model simulation training process; on the basis of the computation behavior, the communication behavior and a dependency relationship between the computation behavior and / or the communication behavior, constructing a workload execution graph; and on the basis of the workload execution graph, performing simulation to obtain a simulation result of the target model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Model simulation

[0002]

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to model simulation. BACKGROUND

[0003]

[0002] At present, the application of machine learning models in many fields has brought revolutionary progress to artificial intelligence. This progress is behind high training costs and urgent demand for high training efficiency. Optimization in the model parameter level and hardware level (such as GPU selection, network topology design) is crucial to improve training performance and reduce costs.

[0004]

[0003] Generally, model simulation can achieve optimization in the model parameter level and hardware level. Therefore, there is an urgent need for a real and accurate model simulation technical solution to solve the above technical problems. SUMMARY

[0005]

[0004] In view of this, embodiments of the present disclosure provide a model simulation method. One or more embodiments of the present disclosure also relate to a model simulation system, a computing device, a computer readable storage medium, and a computer program product to solve the technical defects in the related art.

[0006]

[0005] According to a first aspect of embodiments of the present disclosure, a model simulation method is provided, including: determining model simulation parameters of a target model and environment simulation parameters for simulating the target model; determining a simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters; performing model simulation training on the target model using the simulation behavior, to determine a computing behavior and a communication behavior in the model simulation training process; constructing a workload execution graph according to the computing behavior, the communication behavior, and a dependency relationship between the computing behavior and / or the communication behavior; and performing simulation according to the workload execution graph to obtain a simulation result of the target model.

[0007]

[0006] According to a second aspect of an embodiment of the present disclosure, a model simulation system is provided, comprising: a parameter determination unit configured to determine model simulation parameters of a target model and environment simulation parameters for simulating the target model; a modeling unit configured to determine a simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters, perform model simulation training on the target model by using the simulation behavior, determine a computing behavior and a communication behavior in the model simulation training process, and construct a workload execution graph according to the computing behavior, the communication behavior, and a dependency relationship between the computing behavior and / or the communication behavior; and a simulation unit configured to perform simulation according to the workload execution graph and obtain a simulation result of the target model.

[0008]

[0007] According to a third aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which realize the steps of the above model simulation method when executed by the processor.

[0009]

[0008] According to a fourth aspect of an embodiment of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, which realize the steps of the above model simulation method when executed by a processor.

[0010]

[0009] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer programs / instructions, which realize the steps of the above model simulation method when executed by a processor.

[0011]

[0010] An embodiment of the present disclosure provides a model simulation method, comprising: determining model simulation parameters of a target model and environment simulation parameters for simulating the target model; determining a simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters; performing model simulation training on the target model by using the simulation behavior, determining a computing behavior and a communication behavior in the model simulation training process; constructing a workload execution graph according to the computing behavior, the communication behavior, and a dependency relationship between the computing behavior and / or the communication behavior; and performing simulation according to the workload execution graph and obtaining a simulation result of the target model.

[0012]

[0011] Specifically, the method constructs a virtual model simulation condition by using the model simulation parameter and the environment simulation parameter, i.e., a simulation behavior of model simulation training on the target model, uses the simulation behavior to perform model simulation training on the target model, obtains the calculation behavior and the communication behavior in the model simulation training process, ensures the consistency of the calculation behavior and the communication behavior with the actual training process of the target model, realizes the model simulation without real training, reduces the resource consumption and the time consumption of the target model training process modeling, and ensures the accuracy of the target model training. Then, a workload execution graph is constructed according to the calculation behavior, the communication behavior, and the dependency relationship between the calculation behavior and / or the communication behavior, the authenticity and the accuracy of the workload execution graph are improved, thereby the authenticity and the accuracy of the simulation result obtained by simulation according to the workload execution graph are improved. Furthermore, the model parameter layer and the hardware layer are optimized by accurate and real model simulation, and the optimization accuracy and the optimization effect of the optimized model parameter layer and the hardware layer are improved. BRIEF DESCRIPTION OF DRAWINGS

[0013]

[0012] FIG. 1 is a specific application scene diagram of a model simulation method provided by one embodiment of the present disclosure;

[0014]

[0013] FIG. 2 is a flowchart of a model simulation method provided by one embodiment of the present disclosure;

[0015]

[0014] FIG. 3 is a structural block diagram of a model simulation method provided by one embodiment of the present disclosure;

[0016]

[0015] FIG. 4 is a modeling process flowchart of a model simulation method provided by one embodiment of the present disclosure;

[0017]

[0016] FIG. 5 is a structural schematic diagram of a model simulation system provided by one embodiment of the present disclosure;

[0018]

[0017] FIG. 6 is a structural block diagram of a computing device provided by one embodiment of the present disclosure. DETAILED DESCRIPTION

[0019]

[0018] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details that are set forth in the following description, and it is understood that persons having ordinary skill in the art can make and use modifications to the present disclosure without departing from the scope of the present disclosure, and it is intended that the scope of the present disclosure encompass such modifications.

[0020]

[0020] It should be understood that although the terms first, second, etc. can be employed in this disclosure one or more embodiments to describe various information, these information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information. For example, without departing from the scope of the disclosure one or more embodiments, first can also be referred to as second, and similarly, second can also be referred to as first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to ascertaining".

[0021]

[0021] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the disclosure one or more embodiments are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0022]

[0022] In the disclosure one or more embodiments, the large model refers to a deep learning model with large-scale model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, tens of billions or even more than one hundred billion model parameters. The large model can also be called foundation model, through large-scale unlabeled corpus for pre-training of large model, outputting pre-training model with more than one hundred million parameters, such model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as large-scale language model (Large Language Model, LLM), multi-modal pre-training model

[0023] (multi-modal pre-training model) and the like.

[0024]

[0023] In practical applications, a large model can be applied to different tasks by fine-tuning a pre-trained model with a small amount of samples. Large models can be widely applied in natural language processing (NLP) and computer vision, and can be applied to tasks such as visual question answering (VQA), image captioning (IC), image generation, text-based sentiment classification, text summarization, machine translation, and other natural language processing tasks. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, and the like.

[0025]

[0024] First, the technical terms related to one or more embodiments of the present disclosure are explained.

[0026]

[0025] LLM: Large Language Model, large language model.

[0027]

[0026] GPU: Graphics Processing Unit, graphics processing unit.

[0028]

[0027] CPU: Central Processing Unit, central processing unit.

[0029]

[0028] TP: Tensor Parallel, model parallelism, in model parallelism, different parts of a neural network (such as different layers or parameter groups) are distributed to different GPUs. This method is suitable for cases where a single model is too large to fit into a single GPU memory. Each GPU only processes a part of the model, and coordination between different GPUs is needed to complete the forward and backward propagation of the entire model.

[0030]

[0029] DP: Data Parallel, data parallelism, training data is divided into multiple small batches, each small batch is processed independently on different GPUs, all GPUs run the same model copy, each GPU completes the forward and backward propagation of its own batch data, and then updates the model parameters across all GPUs.

[0031]

[0030] PP: Pipeline Parallel. Pipeline parallelism assigns different layers of the model to different GPUs and runs different stages of the model in a pipeline manner. When the forward propagation of a layer is completed, its output can be immediately passed to the next layer, while the previous layer can start processing the next batch of data.

[0032]

[0031] Map: A way of storing data using a key-value data structure.

[0033]

[0032] NCCL (NVIDIA Collective Communications Library): is a function library for efficient data collection and distribution in multi-GPU and multi-data processing node systems; NCCL provides a set of primitives

[0034] (Primitives) are used to transfer data in GPU memory, supporting various types of collection operations, such as all-reduce, broadcast, reduction, and scatter.

[0035]

[0033] NCCL Object Initialization (NCCL Communicator Init): Initialization of objects in the NCCL library used to define communication relationships between processes. Communicator is an object.

[0036]

[0034] channel: communication path.

[0037]

[0035] Mock: Mock refers to a fake object used to simulate the behavior of a real object.

[0038]

[0036] Bootstrap: The steps to establish the initial communication connection between data processing nodes.

[0039]

[0037] workload: workload.

[0040]

[0038] Network Interface Card: A hardware device in a computer used to connect a computer to a computer network, also known as a network card.

[0041]

[0039] torchrun: A command used for distributed training and other parallel computing tasks.

[0042]

[0040] In the present disclosure, a model simulation method is provided, and the present disclosure also relates to a model simulation system, a computing device, and a computer-readable storage medium, which are described in detail in the following embodiments.

[0043]

[0041] Referring to FIG. 1, FIG. 1 shows a specific application scenario diagram of a model simulation method according to an embodiment of the present disclosure.

[0044]

[0042] As shown in FIG. 1, the model simulation system shown in FIG. 1 includes a client 102 and a model simulation server 104, wherein the client 102 includes but is not limited to a mobile phone, a tablet, a notebook computer, a tablet computer, a desktop computer, etc., and the client 102 communicates with the model simulation server 104 through a network or a wired connection; the model simulation server 104 includes but is not limited to a physical server, a cloud server, etc.

[0045]

[0043] In a specific implementation, the client 102 sends model simulation parameters and environment simulation parameters for simulating a target model to the model simulation server 104 through a network or a wired data transmission form, the model simulation server 104 determines the model simulation parameters and the environment simulation parameters, and determines a simulation behavior of the target model for model simulation training; then, the model simulation server 104 performs model simulation training on the target model according to the model simulation parameters and the environment simulation parameters, and determines a computing behavior and a communication behavior in the model simulation training process; further, the model simulation server 104 constructs a workload execution graph according to the computing behavior, the communication behavior, and a dependency relationship between the computing behavior and / or the communication behavior, simulates the workload execution graph, and obtains a simulation result of the target model; after obtaining the simulation result of the target model, the model simulation server 104 returns the simulation result to the client 102, and the client 102 can evaluate the target model according to the simulation result, and implement optimization of model parameters and environment parameters of a subsequent target model.

[0046]

[0044] The model simulation method provided by one or more embodiments of the present disclosure reduces the optimization cost of the model simulation parameters and the environment simulation parameters of the target model by simulating the training process of the target model without actually performing the complete simulation process of the target model, and reduces the resource consumption of the simulation due to the fact that the simulation does not need to be performed on an actual simulation scene. Furthermore, the simulation flow is clearly described by using the workload execution graph constructed by the computing behavior and the communication behavior, so that the simulation performed according to the workload execution graph ensures the authenticity and accuracy of the model simulation without consuming a large amount of resources.

[0047]

[0045] Referring to FIG. 2, FIG. 2 shows a flowchart of a model simulation method according to one embodiment of the present disclosure, which specifically includes the following steps.

[0048]

[0046] Step 202: Determine the model simulation parameters of the target model and the environment simulation parameters used for simulating the target model.

[0049]

[0047] The target model can be understood as a machine learning model to be simulated, including but not limited to a neural network model (for example, the neural network model can be understood as a multi-modal pre-training model, a large language model (LLM), etc., for example, the large language model can be understood as GPT3-175B, LLaMA2-70B, etc.), a linear model, a decision tree model, etc.

[0050]

[0048] The model simulation parameters can be understood as the model parameters of the target model to be simulated, for example, in the case where the target model is understood as an LLM, the model simulation parameters can be understood as the number of model layers (num-layers), the scale of the hidden layer (hidden-size), etc.

[0051] The model simulation parameters can be understood as the parameters of the target model, such as the number of hidden layers (num-layers), the number of hidden units in each hidden layer (hidden-size), the number of multi-head attention (num-attention-heads), the sequence length (seq-length), the size of the micro-batch (micro-batch-size), the size of the tensor model parallel (tensor-model-parallel-size), the size of the pipeline model parallel (pipeline-model-parallel-size), the global batch size (global-batch-size), the training iteration (train-iters), and the like. Or, in the case where the target model is understood as a decision tree model, the model simulation parameters can be understood as the division criterion (Criterion), the Gini coefficient (Gini), the information gain (Entropy), the node division strategy (Splitter), the class weight (Class_weight), and the like.

[0052]

[0049] The environment simulation parameters can be understood as the parameters of the environment for model simulation, for example, the parameters of the distributed environment for model simulation, wherein the parameters of the distributed environment for model simulation include but are not limited to the machine category (such as servers, GPUs, CPUs, and the like) for implementing simulation, the number of machines, the model of the machine, the number of network cards, the model of the network card, the number ratio of the machine and the network card, and the like.

[0053]

[0050] In specific implementation, the model simulation parameters of the target model and the environment simulation parameters for simulating the target model can be determined by receiving the model simulation data sent by the client. The model simulation data can be obtained by the client through the interaction behavior of the user in the user interaction interface in the client. The client analyzes the model simulation data, obtains and sends the model simulation parameters of the target model and the environment simulation parameters for simulating the target model to the model simulation server. The model simulation server receives and confirms the model simulation parameters of the target model and the environment simulation parameters for simulating the target model, for subsequent model simulation.

[0054]

[0051] Step 204: determining the simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters.

[0055]

[0052] The simulation behavior can be understood as an operation behavior simulating the training process of the target model, and the model simulation training is achieved by the simulation behavior, so that the model simulation does not need to be based on a real distributed environment.

[0056]

[0053] In a specific implementation, after the model simulation server determines the model simulation parameter of the target model and the environment simulation parameter for simulating the target model, the model simulation server can establish a series of simulation behaviors for model simulation training of the target model according to the model simulation parameter of the target model and the environment simulation parameter for simulating the target model.

[0057]

[0054] In practical applications, the simulation behavior can be determined by the environment simulation parameter and the model simulation parameter, and the specific implementation is as follows: the simulation behavior for model simulation training of the target model is determined according to the model simulation parameter and the environment simulation parameter, including: determining at least two data processing nodes and node information of each data processing node according to the environment simulation parameter; determining at least two data processing node groups and simulation behaviors of each data processing node group for model simulation training of the target model according to the at least two data processing nodes, the node information of each data processing node, and the model simulation parameter.

[0058]

[0055] The data processing node can be understood as a virtual machine of any one machine in the cluster machine, such as a virtual GPU, a virtual CPU, a virtual single server, etc., and the node information of the data processing node includes identification information (such as a number) of the virtual data processing node and hardware configuration information (such as a model, configuration parameters, and an interface quantity, etc.) of the virtual data processing node.

[0059]

[0056] Specifically, the node information of the data processing node is saved in a single server (i.e., the model simulation server) in the form of a Map, so that the single model simulation server can simulate a cluster environment containing multiple data processing nodes.

[0060]

[0057] The simulation behavior of each data processing node group for model simulation training of the target model can be understood as an operation behavior of each data processing node group for training of the simulation model.

[0061]

[0058] For example, model behavior can be understood as certain data processing nodes performing forward computation of the simulated model, certain data processing nodes performing backward computation of the simulated model, certain data processing nodes performing simulated aggregate communication, each data processing node (i.e., virtual machine) independently executing single-machine hardware topology awareness, each data processing node (i.e., virtual machine) independently performing computation, each data processing node (i.e., virtual machine) establishing a channel graph on a single machine, and all data processing nodes...

[0062] The channel graphs (i.e., virtual machines) are summarized, and a global communication channel graph is established, etc.

[0063]

[0059] The model simulation method provided in one or more embodiments of this disclosure, by grouping data processing nodes, makes the cluster containing each data processing node closer to the actual distributed environment, and by using the grouped data processing nodes to perform the subsequent simulation model training process, improves the accuracy and efficiency of the entire model simulation.

[0064]

[0060] In practical applications, the specific implementation method for grouping at least two data processing nodes is as follows: Determining the grouping of at least two data processing nodes based on the at least two data processing nodes, the node information of each data processing node, and the model simulation parameters includes: determining a first partition number based on the training data information for the target model and the node information of each data processing node; grouping the at least two data processing nodes based on the first partition number and the node information of each data processing node to obtain a first node group equal to the first partition number; determining a second partition number for the target model based on the model structure parameters in the model simulation parameters, and grouping the data processing nodes in the first node group based on the second partition number to obtain a second node group equal to the second partition number; determining a third partition number based on the number of data processing nodes in the second node group, and grouping the data processing nodes in the second node group based on the third partition number to obtain a third node group equal to the third partition number; and determining the at least two data processing node groups based on the first node group, the second node group, and the third node group.

[0065]

[0061] The training data information of the target model can be understood as the amount of data of the model to be trained, such as 100,000, 1 million, etc.

[0066]

[0062] The model structure parameter in the model simulation parameter includes but is not limited to the number of model layers, the parameter quantity of each layer of the model, etc.

[0067]

[0063] Specifically, after the at least two data processing nodes and the node information of each data processing node are determined, the training data of the target model can be divided according to the training data information of the target model, and the number of the divided data groups is obtained, which can be determined as the first division number. Then, according to the first division number, the at least two data processing nodes need to be grouped, and the number of the first node groups obtained.

[0068]

[0064] In specific implementation, the data processing nodes can be grouped according to the GPU model, the network card model and other information contained in the node information of each data processing node, and the parameter quantity of the training data to be divided, so as to improve the matching degree between each first node group and the parameter quantity of each training data to be divided. For example, the first node group processing the training data with a large parameter quantity is allocated with a data processing node containing a high-configuration GPU, network card and other hardware resources.

[0069]

[0065] It should be noted that, in order to facilitate understanding, the model simulation method provided by one or more embodiments of the present disclosure is exemplarily described by taking the same hardware resources of each first node group, the equal division of the training data of the target model, the division of the model according to the number of layers, and the consistent hardware resources required by each layer as examples. In actual application, the matching degree between the hardware resources of the data processing nodes and the data to be processed can be set according to actual needs, and the model simulation method provided by one or more embodiments of the present disclosure is referred to for simulation. The selection strategy of the specific data processing node and other actual setting rules are not limited by one or more embodiments of the present disclosure.

[0070]

[0066] For example, the number of the at least two data processing nodes is 24 (for example, the data processing nodes are numbered as 1-24), the first division number is 2, then the number of the first node groups obtained by grouping the at least two data processing nodes is 2, then each first node group contains 12 nodes, that is, one first node group contains data processing nodes (1-12), and the other first node group contains data processing nodes (13-24).

[0071]

[0067] Further, the model can be divided into blocks according to the model structure parameter in the model simulation parameters, and each data processing node can be divided, so that each part of the data processing node after division can process one block of the model.

[0072]

[0068] Specifically, the second division number that the target model can be divided into, i.e., the block number of the target model, can be determined according to the model structure parameter in the model simulation parameters, and the number of the second node group that groups the data processing nodes in the first node group can be determined according to the second division number.

[0073]

[0069] For example, the model structure parameter is understood as the number of model layers, the target model has 3 layers, and the second division number is determined to be 3 according to the number of model layers, i.e., the number of the second node group is 3 groups. The data processing nodes in the first node group are divided into 3 groups, for example, the first node group containing data processing nodes (1-12) is divided into 3 groups, and three second node groups (1-4), (5-8), and (9-12) are obtained.

[0074] The first node group containing data processing nodes (13-24) is divided into (13-16), (17-20), and (21-24) three second node groups.

[0075]

[0070] Optionally, the second node groups processing the same model layer in each first node group can be combined to form combined second node groups. In the above example, three combined second node groups (1-4, 13-16), (5-8, 17-20), and (9-12, 21-24) are obtained.

[0076]

[0077]

[0071] Further, the data processing nodes can be grouped according to the pipeline of model training, i.e., the number of pipelines that can be divided, i.e., the third division number, is determined according to the number of data processing nodes contained in the second node group, and the data processing nodes in the second node group are grouped according to the third division number, so that the output of the model layer data processing can be transmitted to the next layer more quickly after the model layer data processing is completed, without the need to select the data processing nodes of the next layer to be transmitted, and the layer can start processing the next data batch more quickly, thereby improving the communication efficiency in the model training process, improving the efficiency of model training, and reducing the consumption of computer resources.​

[0078]

[0072] Continuing with the above example, for example, the second node groups are data processing nodes (1-4), (5-8), (9-12), (13-16), (17-20), (21-24), each of which contains four nodes, and the number of pipelines that can be divided is 1, 2, 3, 4, that is, the third division number can be 1, 2, 3, 4, and taking the third division number 2 as an example, each second node group is divided to obtain a third node small group, that is, data processing nodes (1-4) are divided into two groups, each of which contains two data processing nodes, that is, (1, 2), (3, 4), wherein (1, 2) is a data processing node of a certain layer of a processing model in a pipeline, and (3, 4) is a data processing node of a certain layer of a model in another pipeline, which is the same as (1, 2).

[0079]

[0073] Then, one third node small group is taken out from each layer of the model corresponding third node small group, and three third node small groups corresponding to a pipeline are combined, for example, (1, 2), (5, 6), (9, 10), and the three third node small groups are combined, that is, a third node group contains data processing nodes (1, 2, 5, 6, 9, 10), and similarly, third node groups (3, 4, 7, 8, 11, 12), (13, 14, 17, 18, 21, 22), (15, 16, 19, 20, 23, 24) are obtained.

[0080]

[0074] Further, after obtaining the above first node group, second node group, and third node group, at least two data processing node groups can be obtained according to the first node group, second node group, and third node group.

[0081]

[0075] Continuing with the above example, at least two data processing node groups are obtained according to the above first node group, second node group, and third node group, that is, (1-12), (13-24), (1-4), (5-8), (9-12), (13-16), (17-20), (21-24).

[0082] (17-20), (21-24), (1, 2, 5, 6, 9, 10), (3, 4, 7, 8, 11, 12), (13, 14, 17, 18, 21, 22), (15, 16, 19, 20, 23, 24).

[0083]

[0076] The model simulation method provided by one or more embodiments of the present disclosure determines the first node group after grouping at least two data processing nodes through the training data information of the target model, so that each training data with a smaller parameter quantity can be simulated for training on different data processing nodes, model copies of the same target model are simulated and run on different data processing nodes, and the synchronization update of the model parameters on the data processing nodes is realized through subsequent collective communication, which reduces the data processing pressure and hardware resource consumption of each data processing node and guarantees the integrity and accuracy of the simulation model training process. Further, by dividing the target model into blocks, different layers or different parameter groups of the target model can be distributed on different data processing nodes to avoid the case that a single target model cannot be deployed on a single data processing node due to a large number of model parameters. In addition, since each data processing node processes a part of the target model, the hardware resource consumption of the target model is further reduced, and the communication between the data processing nodes processing different layers or parameter groups of the target model is guaranteed through subsequent collective communication, thereby guaranteeing the model training accuracy of the entire target model. Furthermore, after dividing the pipeline and assigning different layers or parameter groups of the target model to different data processing nodes, the data processing nodes in different layers or parameter groups are grouped again according to the pipeline, so that the data processing nodes processing different layers or parameter groups of the target model can perform point-to-point data interaction. That is, after the forward data processing of a layer is completed, the output thereof can be quickly transmitted to the next layer of the target model, and the layer can quickly process the next data batch, thereby improving the communication efficiency between the data processing nodes processing different layers or parameter groups of the target model, improving the model training efficiency of the entire target model, and reducing the network resource consumption of the entire target model. In addition, the above three grouping methods simulate the grouping of data processing nodes in the target model training process, so that the grouping of data processing nodes is closer to the actual distributed environment, thereby improving the authenticity and accuracy of model simulation.

[0084]

[0077] In actual applications, the at least two data processing nodes are grouped according to the at least two data processing nodes, the node information of each data processing node, and the model simulation parameter. The grouping can also be implemented by any one of the grouping manners or any combination of multiple grouping manners, such as the first grouping manner (that is, grouping the at least two data processing nodes according to the first division quantity and the node information of each data processing node), the second grouping manner (that is, determining the second division quantity of the target model according to the model structure parameter in the model simulation parameter, and grouping the data processing nodes in the first node grouping according to the second division quantity), and the third grouping manner (that is, determining the third division quantity according to the number of data processing nodes in the second node grouping, and grouping the data processing nodes in the second node grouping according to the third division quantity). For specific implementation manners, reference can be made to the above description of the embodiments.

[0085]

[0078] In actual applications, the specific implementation manner of determining the simulation behavior of each data processing node grouping for model simulation training of the target model is as follows: The simulation behavior of each data processing node grouping for model simulation training of the target model is determined according to the node information of each data processing node in the data processing node grouping and the model simulation parameter corresponding to each data processing node.

[0086]

[0079] The model simulation parameter corresponding to each data processing node can be understood as the simulation parameter of the model block of the target model that needs to be simulated by the data processing node, for example, the simulation parameter of a layer of the target model, the simulation parameter of a parameter group of the target model, and the like.

[0087]

[0080] In specific implementation, the simulation operation behavior of each data processing node in each data processing node grouping is determined according to the simulation parameter of the model block that needs to be simulated by each data processing node in each data processing node grouping, and the operation behavior that needs to be simulated by each data processing node grouping is determined according to the simulation operation behavior of each data processing node in each data processing node grouping.

[0088]

[0081] The model simulation method provided by one or more embodiments of the present disclosure improves the accuracy of the simulation behavior of each data processing node grouping by accurately determining the simulation behavior corresponding to each data processing node, so as to determine the simulation behavior of each data processing node grouping corresponding to each data processing node according to the simulation behavior corresponding to each data processing node.

[0089]

[0082] Step 206: model simulation training of the target model using the simulation behavior, determining the computing behavior and the communication behavior in the model simulation training process.

[0090]

[0083] The computing behavior can be understood as the behavior of the computing operation of the simulation data processing node in the computing process of the model simulation training; the communication behavior can be understood as the behavior of the communication operation in the data interaction process between the simulation data processing nodes in the model simulation training process.

[0091]

[0084] In practical applications, according to the model simulation parameters and the environment simulation parameters, the model simulation training of the target model using the simulation behavior determines the specific implementation of the computing behavior and the communication behavior in the model simulation training process as follows: the model simulation training of the target model using the simulation behavior determines the computing behavior and the communication behavior in the model simulation training process, including: using the simulation behavior of the model simulation training of the target model by the data processing node group, training the target model to determine the node information of the data processing node in the data processing node group and the communication relationship between the data processing nodes in the data processing node group; according to the node information of the data processing node in the data processing node group and the communication relationship between the data processing nodes in the data processing node group, constructing a target channel graph, wherein the target channel graph is constructed according to the node information of the data processing node in the data processing node group and the communication relationship between the data processing nodes in the data processing node group; according to the target channel graph and the simulation behavior, determining the running parameters of the data processing node in the data processing node group in the model simulation training process of the target model; according to the running parameters, determining the computing behavior and the communication behavior in the model simulation training process.

[0092]

[0085] The target channel graph can be understood as a communication graph of the data processing node in the data processing node group, and the target channel graph contains a plurality of single machine channel graphs, each single machine channel graph corresponding to a data processing node group, and the single machine channel graph of a data processing node group can be understood as a directed graph with the data processing node in the data processing node group as the node and the communication relationship between the data processing nodes in the data processing node group as the edge.

[0093]

[0086] In a specific implementation, after obtaining the simulation behaviors of the data processing node groups in simulating and training the target model, the communication information between the data processing node groups and the communication information between the data processing nodes in each data processing node group can be determined, and then the communication relationship between the data processing nodes in each data processing node group can be determined according to the communication information between the data processing node groups and the communication information between the data processing nodes in each data processing node group.

[0094]

[0087] Further, after the communication relationship between the data processing nodes in each data processing node group is determined, the target channel graph can be constructed according to the node information of the data processing nodes in each data processing node group and the communication relationship between the data processing nodes in each data processing node group, which can represent the simulation of the communication relationship between the data processing nodes in each data processing node group and the hardware resource status in the distributed environment.

[0095]

[0088] Then, according to the target channel graph and the simulation behaviors, the simulation training process of the target model can be implemented. During the simulation training process of the target model, the running parameters of each data processing node in simulating and training the target model are collected, including but not limited to operation type, input parameter, output parameter, and hardware resource of the executed data processing node. According to the running parameters, the calculation operation information of each data processing node and the communication information between each data processing node and other data processing nodes can be obtained, so that the calculation behaviors corresponding to each calculation operation step and the communication behaviors corresponding to each communication step can be determined according to the calculation operation information of each data processing node and the communication information between each data processing node and other data processing nodes. The calculation step can be understood as the calculation operation step involved in the simulation training process of the target model, and the communication step can be understood as the communication operation step involved in the simulation training process of the target model.

[0089] The model simulation method provided by one or more embodiments of the present disclosure simplifies the complexity of the data processing process by establishing the target channel graph between the data processing nodes. Further, since the target channel graph is consistent with the actual multi-data processing node cluster environment, the running data in the simulation training process of the target model can be collected according to the target channel graph and the simulation behaviors, and the calculation behaviors and the communication behaviors can be determined according to the running data, so that the calculation behaviors and the communication behaviors are more accurate.

[0096]

[0090] In practical applications, the target channel graph can be obtained by combining the single-machine channel graphs of each data processing node to reduce the risk of computer failure that may result from constructing the target channel graph. The specific implementation is as follows: The simulation behavior of using the grouped data processing nodes to perform model simulation training on the target model, and constructing the target channel graph, includes: Simulating the behavior of using the grouped data processing nodes to perform model simulation training on the target model, performing model simulation training on the target model, determining the single-machine topology information of the data processing nodes in each grouped data processing nodes, wherein the single-machine topology information includes the data interaction relationship between at least two simulated hardware components of the data processing node; determining the node information of the data processing nodes in each grouped data processing nodes, and the communication relationship between the data processing nodes in each grouped data processing nodes, based on the single-machine topology information of the data processing nodes in each grouped data processing nodes; Constructing the target channel graph based on the node information of the data processing nodes in each grouped data processing nodes, and the communication relationship between the data processing nodes in each grouped data processing nodes, includes: Based on the node information of the data processing nodes in each data processing node group and the communication relationships between the data processing nodes in each data processing node group, determine the single-machine channel diagram of the data processing nodes in each data processing node group; and construct the target channel diagram based on the single-machine channel diagram of the data processing nodes in each data processing node group.

[0097]

[0091] Among them, the single-machine topology information can be understood as the topology information between the hardware of the data processing nodes that have a dependency relationship.

[0098]

[0092] Data interaction relationship can be understood as the interaction relationship between multiple hardware devices for data transmission and data return.

[0099]

[0093] A single-machine channel graph can be understood as a directed graph constructed with the hardware of a single data processing node as the node and the data interaction relationship between the hardware of the single data processing node as the edge.

[0100]

[0094] Therefore, the target channel diagram determined according to the single-machine channel diagram of the data processing node in each data processing node group can be understood as a directed graph constructed with the nodes corresponding to the single-machine channel diagram of the data processing node in each data processing node group as nodes and the communication relationship between each single-machine channel diagram as edges.

[0101]

[0095] The model simulation method provided by one or more embodiments of the present disclosure greatly reduces the data pressure of data processing of the model simulation server per unit time by respectively constructing single-machine channel graphs of the data processing nodes in each data processing node group to determine the target channel graph in combination, thereby ensuring the stability of the model simulation server.

[0102]

[0096] In actual application, the specific implementation of the running parameters of the data processing nodes in each data processing node group in the model simulation training process of the target model is determined according to the target channel graph and the simulation behavior as follows: the determination of the running parameters of the data processing nodes in each data processing node group in the model simulation training process of the target model according to the target channel graph and the simulation behavior includes: determining the dependency relationship between the data processing nodes in each data processing node group and the simulation behavior according to the target channel graph and the simulation behavior; and determining the running parameters of the data processing nodes in each data processing node group in the model simulation training process of the target model according to the dependency relationship between the data processing nodes in each data processing node group and the simulation behavior.

[0103]

[0097] The dependency relationship between the data processing nodes in each data processing node group and the simulation behavior can be understood as the dependency relationship between the data processing nodes in each data processing node group and the simulation behavior that needs to be executed, that is, the data processing nodes in each data processing node group have a corresponding relationship with each sub-simulation behavior in the simulation behavior, and the sub-simulation behavior in the simulation behavior is a part of the simulation behavior.

[0104]

[0098] The model simulation method provided by one or more embodiments of the present disclosure ensures the accuracy of the correspondence between the data processing nodes in each data processing node group and the simulation behaviors by determining the running parameters according to the dependency relationship between the data processing nodes in each data processing node group and the simulation behaviors, improves the accuracy of the simulation of the data processing nodes, and thus improves the accuracy of the determined running parameters and the accuracy of the subsequently obtained simulation results.

[0099] In actual applications, in order to improve the accuracy of the communication behaviors, the grouping communication behaviors in the communication behaviors corresponding to each data processing node group can be specifically analyzed, so as to ensure the accuracy of the communication behaviors, and the specific implementation manner is as follows: the calculation behaviors and the communication behaviors in the model simulation training process are determined according to the target channel graph and the simulation behaviors, including: the calculation behaviors of the data processing nodes in each data processing node group in the model simulation training process of the target model are obtained according to the target channel graph and the simulation behaviors; the data receiving nodes and the data sending nodes in each data processing node group and the communication relationship between the data receiving nodes and / or the data sending nodes in each data processing node group are determined according to the communication relationship between the data processing nodes in each data processing node group in the target channel graph and the simulation behaviors; the grouping communication behaviors between the data processing nodes in each data processing node group are determined according to the communication relationship between the data receiving nodes and the data sending nodes in each data processing node group and / or the data receiving nodes and / or the data sending nodes in each data processing node group, and the communication behaviors in the model simulation training process are determined according to the grouping communication behaviors.

[0105]

[0100] Here, the data receiving node can be understood as a data processing node for receiving data, and the data sending node can be understood as a data processing node for sending data.

[0106]

[0101] Specifically, according to the data receiving nodes and the data sending nodes in each data processing node group and the communication relationship between the data receiving nodes and / or the data sending nodes in each data processing node group, the grouping communication behaviors in each data processing node group can be constructed, and the grouping communication behaviors represent the data synchronization behaviors between the data processing nodes in the data processing node group.

[0107]

[0102] For example, the group communication behaviors of a certain data processing node group can be illustrated as follows: data processing node 1, 2, 3 each holds its own data, data processing node 1 holds data x, data processing node 2 holds data y, and data processing node 3 holds data z; for example, in the case of first data synchronization according to the group communication behaviors, data processing node 1 sends data x to data processing node 2, data processing node 2 sends data y to data processing node 3, and data processing node 3 sends data z to data processing node 1, data processing node 1 holds data x and z, data processing node 2 holds data y and x, and data processing node 3 holds data z and y; then, data processing node 1 sends data z to data processing node 2, data processing node 2 sends data x to data processing node 3, and data processing node 3 sends data y to data processing node 1, each data processing node holds data x, y, and z, and data synchronization between data processing node 1, data processing node 2, and data processing node 3 is achieved.

[0108]

[0103] The model simulation method provided by one or more embodiments of the present disclosure determines the communication behaviors in the model simulation training process by utilizing the group communication behaviors of each data processing node group for data synchronization, so that the communication behaviors contain data synchronization information of the data processing node group, improves the authenticity of the simulation training process of the target model and the training effect of the target model, and improves the accuracy of the subsequent simulation results.

[0109]

[0104] Step 208: constructing a workload execution graph according to the computing behaviors, the communication behaviors, and the dependency relationships between the computing behaviors and / or the communication behaviors.

[0110]

[0105] The workload execution graph can be understood as a directed graph containing the process of model training and the communication relationship between data processing nodes, the nodes of which are computing behaviors or communication behaviors, and the edges of which are the dependency relationships between the computing behaviors and / or the communication behaviors.

[0111]

[0106] In actual applications, the specific implementation of constructing the workload execution graph according to the computing behaviors, the communication behaviors, the dependency relationships among the computing behaviors and / or the communication behaviors is as follows: the computing behaviors include at least two groups of computing behaviors, and the communication behaviors include at least two groups of communication behaviors; and the constructing the workload execution graph according to the computing behaviors, the communication behaviors, the dependency relationships among the computing behaviors and / or the communication behaviors includes: constructing a communication graph of a group of data processing nodes corresponding to each group of communication behaviors according to the each group of communication behaviors; and constructing the workload execution graph according to the communication graph of the group of data processing nodes corresponding to the each group of communication behaviors, the computing behaviors, the dependency relationships among the computing behaviors and / or the communication graph.

[0112]

[0107] The communication graph of the group of data processing nodes corresponding to each group of communication behaviors can be understood as a directed graph constructed by taking data receiving nodes / data sending nodes in the group of data processing nodes corresponding to each group of communication behaviors as nodes and taking communication relationships between the data receiving nodes and / or the data sending nodes as edges.

[0113]

[0108] Specifically, after the communication graph of the group of data processing nodes corresponding to each group of communication behaviors is determined, the constructing the workload execution graph according to the communication graph of the group of data processing nodes corresponding to each group of communication behaviors, the computing behaviors, the dependency relationships among the computing behaviors and / or the communication graph can be understood as constructing the workload execution graph by taking the communication graph of the group of data processing nodes corresponding to each group of communication behaviors and the computing behaviors as nodes and taking the dependency relationships among the computing behaviors and / or the communication graph of the group of data processing nodes corresponding to each group of communication behaviors as edges.

[0114]

[0109] The model simulation method provided by one or more embodiments of the present disclosure utilizes a more detailed communication graph corresponding to the communication behaviors, so that the workload execution graph contains more detailed communication information, and the accuracy of the workload execution graph is improved.

[0115]

[0110] Step 210: simulating according to the workload execution graph to obtain a simulation result of the target model.

[0116]

[0111] The simulation result can be understood as a simulation result in the form of a chart, a simulation result in the form of text, a simulation result in the form of a combination of text and graphics, and the like. The simulation result includes efficiency information obtained by simulating the target model, and the efficiency information includes, but is not limited to, simulation time length, calculation time length of each data processing node, communication time length, and the like.

[0117]

[0112] In actual application, the simulation result can be determined by the calculation time length, the communication time length, and the total simulation time length of the target model, so as to improve the accuracy of the simulation result. The specific implementation manner is as follows: the simulation according to the workload execution graph to obtain the simulation result of the target model includes: simulating according to the workload execution graph to obtain the calculation time length, the communication time length, and the total simulation time length of the target model; and obtaining the simulation result of the target model according to the calculation time length, the communication time length, and the total simulation time length of the target model.

[0118]

[0113] The calculation time length can be understood as a time length required for simulating a calculation behavior, the communication time length can be understood as a time length required for simulating a communication behavior, and the total simulation time length is determined by adding the calculation time length and the communication time length.

[0119]

[0114] The model simulation method provided by one or more embodiments of the present disclosure can make the simulation result more targeted by obtaining the total simulation time length, the calculation time length, and the communication time length of the target model, so as to facilitate subsequent targeted optimization of the target model, including but not limited to optimization of a calculation part and / or a communication part.

[0120]

[0115] In actual application, the specific implementation manner of simulating according to the workload execution graph to obtain the calculation time length, the communication time length, and the total simulation time length of the target model is as follows: the simulation according to the workload execution graph to obtain the calculation time length, the communication time length, and the total simulation time length of the target model includes: analyzing a calculation behavior in the workload execution graph according to an association relationship between a calculation behavior and a time length to determine a calculation time length corresponding to the calculation behavior; analyzing a communication behavior in the workload execution graph according to an association relationship between a communication behavior and a time length to determine a communication time length corresponding to the communication behavior; and simulating the target model according to the calculation behavior, the communication behavior, a dependency relationship between the calculation behavior and / or the communication behavior, the calculation time length corresponding to the calculation behavior, and the communication time length corresponding to the communication behavior to determine the total simulation time length of the target model.

[0121]

[0116] The association between the computing behavior and the time length can be understood as a corresponding relationship between the computing behavior and the time length actually required for performing the computing behavior, for example, computing behavior A corresponds to a time length of 30s, computing behavior B corresponds to a time length of 40s, and so on.

[0122]

[0117] The association between the communication behavior and the time length can be understood as a corresponding relationship between the communication behavior and the time length actually required for performing the communication behavior, for example, communication behavior A corresponds to a time length of 2s, communication behavior B corresponds to a time length of 6s, and so on.

[0123]

[0118] Specifically, according to the computing behavior, the communication behavior, the dependency relationship between the computing behavior and the communication behavior, and / or the dependency relationship between the computing behavior and the communication behavior contained in the obtained work load execution graph, the order of the computing behavior and / or the communication behavior in the simulation process of the target model can be determined, and according to the computing time length of the computing behavior and the communication time length of the communication behavior, the total simulation time required for model simulation according to the work load execution graph can be determined.

[0124]

[0119] The model simulation method provided by one or more embodiments of the present disclosure improves the accuracy of the computing time length and the communication time length by predicting the computing time length corresponding to the computing behavior and the communication time length corresponding to the communication behavior according to the time length actually required by the computing behavior and the communication behavior, and improves the accuracy of the total simulation time length stabilized according to the computing time length and the communication time length.

[0125]

[0120] In one or more embodiments of the present disclosure, after the simulation according to the work load execution graph is performed to obtain the simulation result of the target model, the method further includes: evaluating the target model according to the simulation result, and adjusting the model parameters of the target model and / or the environment parameters for running the target model according to the evaluation result.

[0126]

[0121] The model parameters of the target model can be understood as the model parameters corresponding to the actual target model, and the environment parameters of the target model can be understood as the environment parameters of the environment in which the target model is actually trained / applied.

[0127]

[0122] Evaluating the target model according to the simulation result can be understood as evaluating the computing efficiency, the communication efficiency, and the overall training efficiency of the target model according to the computing time length, the communication time length, and the total simulation time length contained in the simulation result, so as to adjust the model parameters of the target model and / or the environment parameters for running the target model.

[0128]

[0123] For example, in the case of a long calculation duration, such as a calculation duration exceeding an expected calculation duration, the model parameters of the target model and / or the environment parameters for running the target model can be adjusted by increasing the number of GPUs, reducing the number of layers of the target model, increasing the scale of the GPU, and the like. For example, in the case of a long communication duration, such as a communication duration exceeding an expected communication duration, the model parameters of the target model and / or the environment parameters for running the target model can be adjusted by increasing the number of network cards, reducing the number of data processing nodes, and the like.

[0129]

[0124] The model simulation method provided by one or more embodiments of the present disclosure adjusts the model parameters of the target model and / or the environment parameters for running the target model according to the simulation result after obtaining the simulation result, so that the efficiency of the actual training / application of the target model is higher, and the consumption of computing resources and network resources is less.

[0130]

[0125] The model simulation method provided by one or more embodiments of the present disclosure constructs a simulated model simulation condition by using the model simulation parameters and the environment simulation parameters, i.e., a simulated behavior of model simulation training of the target model, uses the simulated behavior to perform model simulation training on the target model, obtains the calculation behavior and the communication behavior in the model simulation training process, ensures the consistency of the calculation behavior and the communication behavior with the actual training process of the target model, realizes the model simulation without real training, reduces the resource consumption and time consumption of the modeling of the target model training process, and ensures the accuracy of the target model training. Then, a workload execution graph is constructed according to the calculation behavior, the communication behavior, and the dependency relationship between the calculation behavior and / or the communication behavior, the authenticity and accuracy of the workload execution graph are improved, the authenticity and accuracy of the simulation result obtained by simulating according to the workload execution graph are improved, and further, the optimization accuracy and optimization effect of the optimized model parameter level and hardware level are improved through accurate and real model simulation optimization.

[0131]

[0126] The above is a schematic scheme of a model simulation method of the present embodiment. It should be noted that the technical scheme of the model simulation method belongs to the same concept as the technical scheme of the model simulation method described above, and the details of the technical scheme of the model simulation method that are not described in detail can be referred to the description of the technical scheme of the model simulation method.

[0132]

[0127] The following description, in conjunction with Figure 3, takes the application of the model simulation method provided in this disclosure in an LLM model as an example to further illustrate the model simulation method. It should be noted that the embodiments of this disclosure are not limited to model simulation of LLM models, but can also be applied to other machine learning models. For details, please refer to the specific implementation of the model simulation method provided in the embodiments of this disclosure and adjust the corresponding model parameters, environmental parameters, etc. The application of other machine learning models in the embodiments of this disclosure will not be elaborated further.

[0133]

[0128] FIG3 shows a structural block diagram of a model simulation method provided in an embodiment of the present disclosure, specifically including a parameter receiving unit 302, a parsing unit 304, a modeling unit 306, and a simulation unit 308, the specific implementation of which is described below.

[0134]

[0129] The parameter receiving unit 302 of the LLM model simulation system (i.e., the model simulation server mentioned above) receives the model parameters (such as model name, model version number), training framework parameters (such as model parallel parameters, training rounds, etc.) and cluster environment information (GPU model, network card model, GPU / NIC ratio, cluster size, etc.) input by the user, and sends the model parameters, training framework parameters and cluster environment information input by the user to the parsing unit 304; the parsing unit 304 determines whether the model parameters, training framework parameters and cluster environment information input by the user are all data in a preset format. If so, it performs parameter parsing on the data in the model parameters, training framework parameters and cluster environment information input by the user that do not conform to the preset format, and obtains the model parameters, training framework parameters and cluster environment information that conform to the preset format.

[0135]

[0130] Then, the parsing unit 304 inputs the model parameters, training framework parameters and cluster environment information that conform to the preset format to the modeling unit 306. The modeling unit 306 processes the data to obtain the workload execution graph (i.e. the above-mentioned workload execution graph).

[0136]

[0131] Specifically, the modeling unit 306 performs data processing on the model parameters, training framework parameters, and cluster environment information that conform to the preset format to obtain the specific implementation of the workload execution graph, as shown in Figure 4.

[0137]

[0132] Referring to Figure 4, Figure 4 shows a flowchart of the modeling process of a model simulation method provided in an embodiment of the present disclosure.

[0138]

[0133] Among them, steps 402-step 424 show an actual running process of model simulation modeling based on actual distributed environment, and the specific implementation is as follows.

[0139]

[0134] Step 402: input user parameters.

[0140]

[0135] Specifically, the user parameters can be understood as the model parameters and the training framework parameters.

[0141]

[0136] Step 404: input environment parameters.

[0142]

[0137] Specifically, the environment parameters can be understood as the cluster environment information.

[0143]

[0138] Step 406: initialize the distributed environment.

[0144]

[0139] Step 408: divide the communication group.

[0145]

[0140] Step 410: initialize the communication group.

[0146]

[0141] Step 412: perform forward / backward calculation.

[0147]

[0142] Specifically, the forward / backward calculation can be understood as model forward calculation / model backward calculation.

[0148]

[0143] Step 414: generate the calculation node.

[0149]

[0144] Step 416: generate the set communication node.

[0150]

[0145] Step 418: initialize the object.

[0151]

[0146] Step 420: initialize the Bootstrap.

[0152]

[0147] Step 422: collect the cluster topology information.

[0153]

[0148] Step 424: build the group multi-machine communication graph.

[0154]

[0149] Step 426: flow the set communication.

[0155]

[0150] In implementation, the model simulation modeling process of steps 402-426 needs a real distributed environment to be implemented, which results in high network resources and computer resources required and low modeling efficiency. Therefore, one or more embodiments of the present disclosure provide a technical solution for simulating the above steps to implement modeling. See steps 428-436.

[0156]

[0151] Step 428: Simulate the initialization of the distributed environment.

[0157]

[0152] Specifically, according to the cluster size in the above cluster environment information, a simulated distributed environment is constructed. For example, by intercepting the torchrun command for starting the distributed running environment, the nproc_per_node (processes per node), nnodes (node number), and node_rank (node rank) of the distributed environment are obtained, i.e., the above cluster information.

[0158]

[0153] According to the nproc_per_node (processes per node) > nnodes (node number), node_rank (node rank) of the distributed environment, the init initialization function of the distributed running environment is modified, virtual computing nodes corresponding to the nproc_per_node (processes per node), nnodes (node number), and node_rank (node rank) of the distributed environment are virtually generated, and the virtual computing nodes are saved in a single process in the form of Map mapping, and the virtual computing nodes are configured as a distributed group.

[0159]

[0154] Step 430: Simulate the division of the communication group.

[0160]

[0155] Specifically, according to the original grouping rules of Tensor Parallel (TP) > Data Parallel (DP) > Pipeline Parallel (PP), a virtual multi-machine communication group is created using a simulation framework, so that the simulated Group (i.e., the virtual multi-machine communication group, Group is the communication group) is highly consistent with the actual distributed environment.

[0161]

[0156] For example, as shown in FIG. 4, group 0 includes nodes 0-7, group n includes nodes 0-8, and so on.

[0162]

[0157] Step 432: Virtual multi-machine environment.

[0163]

[0158] Specifically, in each multi-machine communication group, the initialization of the communication group is completed by simulating the NCCL Communicator Init related functions using the preset Mock scheme. The initialization process includes collecting the topology information of the above-mentioned computing nodes in advance, and converting the topology information into intra-single-process information transmission (for example, using data structures, function calls, etc., which can be understood as a communication operator between nodes).

[0164]

[0159] After that, according to the intra-single-process information transmission information, a distributed channel communication graph corresponding to each multi-machine communication group is established, which includes the configuration information of the virtual computing nodes in the corresponding multi-machine communication group (the virtual computing nodes can implement the simulated behavior in the Mock scheme) itself and the hardware dependency relationship, the dependency relationship between the virtual computing nodes, etc.

[0165]

[0160] Step 434: Build a computation graph.

[0166]

[0161] Specifically, the execution importance of the simulated behavior on the virtual computing node is obtained, the critical position in the virtual computing node is determined according to the execution importance of the simulated behavior on the virtual computing node, and the critical position is instrumented to embed the Trace logic to capture the information of the simulated behavior execution on the virtual computing node at each critical position, such as operation type, input and output parameters, and hardware resources consumed by execution, and according to the captured information, a directed graph (as shown in FIG. 4 a, b, c, d) is constructed, that is, a behavior graph is trained, wherein the communication is described as a computing node, that is, a collective communication node (i.e., b).

[0167]

[0162] Step 436: Build a collective communication splitting graph.

[0168]

[0163] Specifically, the communication Mock is preset, the logic of data interaction in each multi-machine communication group is captured, and the parameters of data interaction in each multi-machine communication group are recorded, including data volume, sender (Send) and receiver (Recv) of data interaction, communication channel, path (direction of data interaction), and according to these parameters, a traffic model graph of the collective communication of each multi-machine communication group is established.

[0169]

[0170]

[0164] For example, a certain multi-machine communication group set communication is shown in FIG. 4. In the case where set communication nodes 1, 2, and 3 do not perform set communication, each set communication node holds its own data, that is, node 1 holds data x, node 2 holds data y, and node 3 holds data z. In the case where the first set communication is performed (denoted as T = 1), node 1 sends data x to node 2, node 2 sends data y to node 3, and node 3 sends data z to node 1. After the first set communication is completed, node 1 holds data x and z, node 2 holds data y and x, and node 3 holds data z and y. In the case where the second set communication is performed (denoted as T = 2), node 1 sends data z to node 2, node 2 sends data x to node 3, and node 3 sends data y to node 1. After the second set communication is completed, that is, in the case where the third set communication is performed (denoted as T = 3), each node holds data x, y, and z, and data synchronization between node 1, node 2, and node 3 is achieved.

[0171]

[0165] Then, the traffic model graph of the set communication operation obtained in step 436 is implanted into the training behavior graph obtained in step 434, specifically, the traffic model graph is replaced with the set communication nodes in the training behavior graph, to generate a workload execution graph that can accurately describe the entire cluster AI training process.

[0172]

[0166] Further, after the workload is generated by the modeling unit 306, the workload is input to the simulation unit 308, and the simulation unit 308 analyzes the workload to obtain the computing nodes, communication nodes (specifically, each sending / receiving node), dependency relationships between the nodes, interaction data, paths, and the like, executes simulation scheduling logic, and obtains simulation results.

[0173]

[0167] In one or more embodiments of the present disclosure, after the simulation unit 308 obtains the simulation results, the simulation unit 308 can send the simulation results to the visualization processing unit, the visualization processing unit performs visualization processing on the simulation results, draws pictures, and sends the pictures to the client corresponding to the LLM model simulation system, to display them to the developer or other user. Subsequently, the developer or other user can perform model evaluation, model optimization, model environment optimization, and the like according to the pictures.

[0174]

[0168] The model simulation method provided by one or more embodiments of the present disclosure can simulate a distributed cluster environment of multiple machines by modeling the modeling process of a model, i.e., performing fine semantic modification on key components and positions, so that a single machine can simulate the distributed cluster environment of multiple machines, resource consumption and time consumption due to communication between multiple machines are reduced, and the authenticity and accuracy of the distributed cluster environment are improved due to modeling of the distributed cluster environment from the training process. Furthermore, by establishing a simulated behavior of the training process, the training process of the simulation model is simulated, so that the machine does not need to perform real simulation execution, but only needs to model from the communication level and the calculation level according to the behavior of the model training, which ensures the integrity of the modeling process, reduces resource consumption and time consumption of the model modeling process, and reduces resource consumption and time consumption of the entire simulation process. Furthermore, the simulation module in the model simulation process can be customized according to actual needs, which improves the universality of the modeling module in the embodiments of the present disclosure, and improves the universality of the model simulation method provided by one or more embodiments of the present disclosure.

[0175]

[0169] The above is a schematic scheme of the model simulation method of the present embodiment. It should be noted that the technical scheme of the model simulation method belongs to the same concept as the technical scheme of the model simulation method described above, and the details of the technical scheme of the model simulation method that are not described in detail can be referred to the description of the technical scheme of the model simulation method.

[0176]

[0170] Corresponding to the method embodiments described above, the present disclosure also provides a model simulation system embodiment. FIG. 5 shows a structural schematic diagram of a model simulation system according to one embodiment of the present disclosure. As shown in FIG. 5, the system includes: a parameter determination unit 502 configured to determine model simulation parameters of a target model and environment simulation parameters for simulating the target model; a modeling unit 504 configured to determine a simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters, perform model simulation training on the target model by using the simulation behavior, determine a calculation behavior and a communication behavior in the model simulation training process, and construct a workload execution graph according to the calculation behavior, the communication behavior, the calculation behavior, and / or the dependency relationship between the communication behaviors; and a simulation unit 506 configured to perform simulation according to the workload execution graph to obtain a simulation result of the target model.

[0177]

[0171] Specifically, the parameter determination unit 502 can be understood as including the parameter receiving unit 302 and the parameters in the embodiments described above.

[0178]

[0172] Optionally, the modeling unit 504 is further configured to: determine at least two data processing nodes and node information of each data processing node according to the environment simulation parameter; and determine at least two data processing node groups and determine simulation behaviors of each data processing node group for model simulation training of the target model according to the at least two data processing nodes, the node information of each data processing node, and the model simulation parameter.

[0179]

[0173] Optionally, the modeling unit 504 is further configured to: determine a first division number according to training data information for the target model and the node information of each data processing node; group the at least two data processing nodes according to the first division number and the node information of each data processing node to obtain a first node group equal to the first division number; determine a second division number of the target model according to a model structure parameter in the model simulation parameter, and group data processing nodes in the first node group according to the second division number to obtain a second node group equal to the second division number; determine a third division number according to a number of data processing nodes in the second node group, and group data processing nodes in the second node group according to the third division number to obtain a third node group equal to the third division number; and determine the at least two data processing node groups according to the first node group, the second node group, and the third node group.

[0180]

[0174] Optionally, the modeling unit 504 is further configured to: determine simulation behaviors of each data processing node group for model simulation training of the target model according to node information of each data processing node in each data processing node group and corresponding model simulation parameters of each data processing node.

[0181]

[0175] Optionally, the modeling unit 504 is further configured to: perform model simulation training on the target model by using the simulation behaviors of the data processing node groups on the target model; determine single-machine topology information of data processing nodes in the data processing node groups, wherein the single-machine topology information comprises a data interaction relationship between at least two simulation hardware of the data processing nodes; determine node information of the data processing nodes in the data processing node groups and a communication relationship between the data processing nodes in the data processing node groups according to the single-machine topology information of the data processing nodes in the data processing node groups; determine running parameters of the data processing nodes in the data processing node groups in a model simulation training process of the target model according to the target channel graph and the simulation behaviors; and determine a calculation behavior and a communication behavior in the model simulation training process according to the running parameters.

[0182]

[0176] Optionally, the modeling unit 504 is further configured to: perform model simulation training on the target model by using the simulation behaviors of the data processing node groups on the target model; determine single-machine topology information of data processing nodes in the data processing node groups, wherein the single-machine topology information comprises a data interaction relationship between at least two simulation hardware of the data processing nodes; determine node information of the data processing nodes in the data processing node groups and a communication relationship between the data processing nodes in the data processing node groups according to the single-machine topology information of the data processing nodes in the data processing node groups; and the modeling unit 504 is further configured to: determine a single-machine channel graph of the data processing nodes in the data processing node groups according to the node information of the data processing nodes in the data processing node groups and the communication relationship between the data processing nodes in the data processing node groups; and construct a target channel graph according to the single-machine channel graph of the data processing nodes in the data processing node groups.

[0183]

[0177] Optionally, the modeling unit 504 is further configured to: determine a dependency relationship between the data processing nodes in the data processing node groups and the simulation behaviors according to the target channel graph and the simulation behaviors; and determine running parameters of the data processing nodes in the data processing node groups in a model simulation training process of the target model according to the dependency relationship between the data processing nodes in the data processing node groups and the simulation behaviors.

[0184]

[0178] Optionally, the modeling unit 504 is further configured to: obtain, according to the target channel graph and the simulation behavior, a calculation behavior of data processing nodes in each data processing node group in a model simulation training process of the target model; determine, according to a communication relationship between data processing nodes in each data processing node group in the target channel graph and the simulation behavior, a data receiving node and a data sending node in each data processing node group and a communication relationship between the data receiving node and / or the data sending node in each data processing node group; and determine, according to the data receiving node and the data sending node in each data processing node group and the communication relationship between the data receiving node and / or the data sending node in each data processing node group, a grouping communication behavior between data processing nodes in each data processing node group, and determine a communication behavior in the model simulation training process according to each grouping communication behavior.

[0185]

[0179] Optionally, the calculation behavior includes at least two grouping calculation behaviors, and the communication behavior includes at least two grouping communication behaviors; the modeling unit 504 is further configured to: construct, according to each grouping communication behavior in the communication behavior, a communication graph of a data processing node group corresponding to the grouping communication behavior; and construct a workload execution graph according to the communication graph of the data processing node group corresponding to each grouping communication behavior, the calculation behavior, a dependency relationship between the calculation behavior and / or the communication graph.

[0186]

[0180] Optionally, the simulation unit 506 is further configured to: simulate according to the workload execution graph to obtain a calculation time length, a communication time length, and a total simulation time length of the target model; and obtain a simulation result of the target model according to the calculation time length, the communication time length, and the total simulation time length of the target model.

[0187]

[0181] Optionally, the simulation unit 506 is further configured to: analyze, according to an association relationship between a calculation behavior and a time length, the calculation behavior in the workload execution graph to determine a calculation time length corresponding to the calculation behavior; analyze, according to an association relationship between a communication behavior and a time length, the communication behavior in the workload execution graph to determine a communication time length corresponding to the communication behavior; and simulate the target model according to the calculation behavior, the communication behavior, a dependency relationship between the calculation behavior and / or the communication behavior, the calculation time length corresponding to the calculation behavior, and the communication time length corresponding to the communication behavior contained in the workload execution graph to determine a total simulation time length of the target model.

[0188]

[0182] Optionally, the system further comprises a model adjustment module configured to: evaluate the target model according to the simulation result, and adjust model parameters of the target model and / or environment parameters for running the target model according to the evaluation result.

[0189]

[0183] The model simulation system provided by one or more embodiments of the present disclosure, when performing model simulation, constructs a virtual model simulation condition by using model simulation parameters and environment simulation parameters, and performs model simulation training on the target model under the virtual model simulation condition through the simulated behavior of the model simulation training of the target model, so as to ensure the consistency of the calculation behavior and the communication behavior with the training process of the model, and to reduce the resource consumption and time consumption of the modeling of the training process of the target model, and then construct a workload execution graph according to the calculation behavior, the communication behavior and the dependency relationship between the calculation behavior and / or the communication behavior, so as to improve the authenticity and accuracy of the workload execution graph, thereby improving the authenticity and accuracy of the simulation result obtained by simulating according to the workload execution graph, and further improving the optimization accuracy and optimization effect of the optimized model parameter level and hardware level through accurate and real model simulation optimization.

[0190]

[0184] The present embodiment is a schematic scheme of a model simulation system. It should be noted that the technical scheme of the model simulation system belongs to the same concept as the technical scheme of the model simulation method described above, and the details of the technical scheme of the model simulation system that are not described in detail can be referred to the description of the technical scheme of the model simulation method.

[0191]

[0185] Referring to FIG. 6, FIG. 6 shows a structural block diagram of a computing device 600 according to an embodiment of the present disclosure. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to save data.

[0192]

[0186] Computing device 600 also includes access device 640 that enables computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. Access device 640 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, or the like.

[0193]

[0187] In one embodiment of the disclosure, the above-described components of computing device 600, as well as other components not shown in FIG. 6, can also be connected to each other by a bus, for example. It should be understood that the structure block diagram of the computing device shown in FIG. 6 is merely for the purpose of example, and is not a limitation on the scope of the disclosure. Other components can be added or replaced as needed by those skilled in the art.

[0194]

[0188] Computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smartwatch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 600 can also be a mobile or stationary server.

[0195]

[0189] The memory 610 is configured to store a computer program / instruction, and the processor 620 is configured to execute the computer program / instruction stored in the memory 610. The computer program / instruction, when executed by the processor, implements the steps of the model simulation method.

[0196]

[0190] The embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments. In particular, the computing device embodiment is basically similar to the model simulation method embodiment, and thus the description is relatively simple, and the relevant part can be referred to the description of the model simulation method embodiment.

[0197]

[0192] The embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments. In particular, the computer readable storage medium embodiment is basically similar to the model simulation method embodiment, and thus the description is relatively simple, and the relevant part can be referred to the description of the model simulation method embodiment.

[0198]

[0193] An embodiment of the present disclosure further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the model simulation method.

[0199]

[0194] The embodiment is described as a schematic scheme of a computer program product. It should be noted that the technical scheme of the computer program product and the technical scheme of the model simulation method belong to the same concept, and the details of the technical scheme of the computer program product which are not described in detail can be referred to the description of the technical scheme of the model simulation method.

[0200]

[0195] The specific embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Also, the process depicted in the figures does not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0201]

[0196] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice. For example, according to the patent practice in some regions, the computer readable medium does not include electric carrier signals and telecommunication signals.

[0202]

[0197] It should be noted that, for the foregoing method embodiments, in order to facilitate description, each is described as a combination of a series of actions, but those skilled in the art should know that the disclosed embodiments are not limited to the order of the actions described, because according to the disclosed embodiments, certain steps can be performed in other orders or at the same time. In addition, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the disclosed embodiments.

[0203]

[0198] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0204]

[0199] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The optional embodiments do not describe all the details and limit the invention to the specific embodiments described. Obviously, according to the content of the disclosed embodiments, many modifications and changes can be made. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the disclosed embodiments, so that those skilled in the art can well understand and use the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

CLAIM 1. A model simulation method, comprising: determining model simulation parameters of a target model and environment simulation parameters for simulating the target model; determining simulation behaviors of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters; determining computing behaviors and communication behaviors in the model simulation training of the target model by using the simulation behaviors; constructing a workload execution graph according to the computing behaviors, the communication behaviors, and dependency relationships between the computing behaviors and the communication behaviors; obtaining simulation results of the target model by simulating according to the workload execution graph. 2.The model simulation method of claim 1, wherein determining, according to the model simulation parameter and the environment simulation parameter, a simulation behavior of the target model for model simulation training comprises: determining at least two data processing nodes and node information of each data processing node according to the environment simulation parameters; determining simulation behaviors of each data processing node group for model simulation training of the target model according to the at least two data processing node groups and the node information of each data processing node.

3. The model simulation method of claim 2, wherein the determining the at least two data processing node groups according to the at least two data processing nodes, the node information of each data processing node, and the model simulation parameters comprises: determining a first division number according to training data information of the target model and the node information of each data processing node; grouping the at least two data processing nodes according to the first division number and the node information of each data processing node to obtain a first node group equal in number to the first division number; determining a second division number of the target model according to model structure parameters in the model simulation parameters, and grouping data processing nodes in the first node group according to the second division number to obtain a second node group equal in number to the second division number; determining a third division number according to the number of data processing nodes in the second node group, and grouping data processing nodes in the second node group according to the third division number to obtain a third node group equal in number to the third division number; and determining the at least two data processing node groups according to the first node group, the second node group, and the third node group. 4.The model simulation method of claim 2, wherein the determining the simulation behavior of each data processing node group for model simulation training of the target model comprises: determining simulation behaviors of each data processing node group for model simulation training of the target model according to the node information of each data processing node in each data processing node group and the model simulation parameters corresponding to each data processing node.

5. The model simulation method according to claim 2, wherein the step of using the simulation behavior to perform model simulation training on the target model, and determining the computational and communication behaviors during the model simulation training process, includes: According to the simulation behavior of the data processing node groups on the target model, the target model is trained, node information of the data processing nodes in the data processing node groups and communication relationships between the data processing nodes in the data processing node groups are determined; according to the node information of the data processing nodes in the data processing node groups and the communication relationships between the data processing nodes in the data processing node groups, a target channel graph is constructed, wherein the target channel graph is constructed according to the node information of the data processing nodes in the data processing node groups and the communication relationships between the data processing nodes in the data processing node groups; according to the target channel graph and the simulation behavior, running parameters of the data processing nodes in the data processing node groups in a model simulation training process of the target model are determined; and according to the running parameters, calculation behavior and communication behavior in the model simulation training process are determined.

6. The model simulation method according to claim 5, wherein the simulating behavior of the model simulation training of the target model by the data processing node groups comprises: training the model simulation of the target model, determining node information of data processing nodes in the data processing node groups and a communication relationship between data processing nodes in the data processing node groups. According to the simulation behavior of the data processing node groups on the target model, the target model is trained, single-machine topology information of the data processing nodes in the data processing node groups is determined, wherein the single-machine topology information comprises data interaction relationships between at least two simulation hardware of the data processing nodes; according to the single-machine topology information of the data processing nodes in the data processing node groups, node information of the data processing nodes in the data processing node groups and communication relationships between the data processing nodes in the data processing node groups are determined; and the target channel graph is constructed according to the node information of the data processing nodes in the data processing node groups and the communication relationships between the data processing nodes in the data processing node groups, comprising: according to the node information of the data processing nodes in the data processing node groups and the communication relationships between the data processing nodes in the data processing node groups, single-machine channel graphs of the data processing nodes in the data processing node groups are determined; and according to the single-machine channel graphs of the data processing nodes in the data processing node groups, the target channel graph is constructed 7. The model simulation method according to claim 5, wherein the determining of the running parameters of the data processing nodes in the data processing node groups in the model simulation training of the target model according to the target channel graph and the simulation behavior comprises: According to the target channel graph and the simulation behavior, a dependency relationship between the data processing nodes in the data processing node groups and the simulation behavior is determined. According to the dependency relationship between the data processing nodes in the data processing node groups and the simulation behavior, running parameters of the data processing nodes in the data processing node groups in a model simulation training process of the target model are determined.

8. The model simulation method according to claim 5, wherein the determining the computing behavior and the communication behavior in the model simulation training process according to the target channel graph and the simulated behavior comprises: According to the target channel graph and the simulation behavior, calculation behavior of the data processing nodes in the data processing node groups in a model simulation training process of the target model is obtained. determine the data receiving nodes and the data sending nodes in the data processing node groups and the communication relationship between the data receiving nodes and / or the data sending nodes in the data processing node groups according to the communication relationship between the data processing nodes in the data processing node groups in the target channel graph and the simulation behavior; determine the group communication behaviors between the data processing nodes in the data processing node groups according to the data receiving nodes and the data sending nodes in the data processing node groups and the communication relationship between the data receiving nodes and / or the data sending nodes in the data processing node groups, and determine the communication behaviors in the model simulation training process according to the group communication behaviors.

9. The model simulation method according to claim 8, wherein the computing behavior includes at least two packet computing behaviors, and the communicating behavior includes at least two packet communicating behaviors. construct a workload execution graph according to the calculation behaviors, the communication behaviors, the dependency relationship between the calculation behaviors and / or the communication behaviors, including: construct a communication graph of the data processing node group corresponding to each group communication behavior according to the group communication behaviors in the communication behaviors; construct a workload execution graph according to the communication graph of the data processing node group corresponding to each group communication behavior, the calculation behaviors, the dependency relationship between the calculation behaviors and / or the communication graph.

10. The model simulation method of claim 1, wherein the simulating according to the workload execution graph to obtain a simulation result of the target model comprises: simulate according to the workload execution graph to obtain the calculation time length, the communication time length, and the total simulation time length of the target model; obtain the simulation result of the target model according to the calculation time length, the communication time length, and the total simulation time length of the target model. analyze the calculation behaviors in the workload execution graph according to the association relationship between the calculation behaviors and the time lengths to determine the calculation time lengths corresponding to the calculation behaviors; 11.The model simulation method of claim 10, wherein the simulating according to the workload execution graph comprises: obtaining a computation duration, a communication duration, and a total simulation duration of the target model. analyze the communication behaviors in the workload execution graph according to the association relationship between the communication behaviors and the time lengths to determine the communication time lengths corresponding to the communication behaviors; simulate the target model according to the calculation behaviors, the communication behaviors, the dependency relationship between the calculation behaviors and / or the communication behaviors, the calculation time lengths corresponding to the calculation behaviors, and the communication time lengths corresponding to the communication behaviors contained in the workload execution graph to determine the total simulation time length of the target model. evaluate the target model according to the simulation result, and adjust the model parameters of the target model and / or the environment parameters used for running the target model according to the evaluation result.

12. The model simulation method of claim 1, wherein the simulation is performed according to the workload execution graph, and after the simulation result of the target model is obtained, further comprising: a parameter determination unit configured to determine model simulation parameters of a target model and environment simulation parameters used for simulating the target model; 13. A model simulation system, comprising: ​ The modeling unit is configured to determine a simulation behavior of the target model for model simulation training according to the model simulation parameters and the environment simulation parameters, perform model simulation training on the target model by using the simulation behavior, determine a computing behavior and a communication behavior in the model simulation training process, and construct a workload execution graph according to the computing behavior, the communication behavior, and a dependency relationship between the computing behavior and / or the communication behavior. The simulation unit is configured to perform simulation according to the workload execution graph to obtain a simulation result of the target model.

14. A computing device comprising: A storage and a processor; The storage is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, so as to implement the steps of the model simulation method in any one of claims 1 to 12.

15. A computer readable storage medium, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the model simulation method in any one of claims 1 to 12.

16. A computer program product, which comprises computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the model simulation method in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Simulation system, method and device

    CN116151137A

  • Method and apparatus for co-optimizing neural network and neural network-specific hardware

    CN117787389A

  • Installation and operation of different processes of an an engine adapted to different configurations of hardware located on-premises and in hybrid environments

    US20180307945A1