Model simulation training method, electronic device, storage medium and program product

CN122693409APending Publication Date: 2026-09-04CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510255184.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

目前,模型的仿真训练需要在真实的物理集群下追踪训练日志,也就是说,依然需要消耗大量真实的计算资源

Benefits of technology

[0008]The simulation training method for the model provided in this application involves transmitting training data determined according to training instructions to a target graphics processing unit (GPU). The training data includes data for at least one neural network layer of the model as indicated by the training instructions. For any one of the at least one neural network layer, the training time for the target GPU to train that neural network layer is obtained. Based on the training time of the at least one neural network layer, simulation training results for a large model are generated. Therefore, simulation training can be performed using only the target GPU, meaning that only a single real GPU is needed to achieve model simulation training, reducing the consumption of real computing resources by the simulation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122693409A_ABST
    Figure CN122693409A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model simulation training method, an electronic device, a storage medium and a program product, relating to the technical field of artificial intelligence, the method comprising: transmitting training data determined according to a training instruction to a target graphics processing unit, the training data comprising data of at least one neural network layer of a model indicated by the training instruction; for any neural network layer in the at least one neural network layer, obtaining a training duration of the target graphics processing unit training the neural network layer; and generating a simulation training result of the model according to the training duration of the at least one neural network layer. In the technical solution of the embodiments of the present application, the consumption of real computing resources by simulation training is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a simulation training method for a model, an electronic device, a storage medium, and a program product. Background Technology

[0002] Training large models typically involves a large number of training parameters and training samples, and requires significant computational resources, such as multiple Graphics Processing Units (GPUs), Network Interface Cards (NICs), and storage devices. In other words, model training usually consumes a considerable amount of time and computational resources. Therefore, it is essential to perform simulation training on large models and evaluate their performance based on the simulation results. Currently, model simulation training requires tracking training logs on real physical clusters, meaning it still consumes substantial real-world computational resources. Summary of the Invention

[0003] This application provides a simulation training method for a model, an electronic device, a storage medium, and a program product to alleviate or solve one or more technical problems existing in the prior art.

[0004] In a first aspect, embodiments of this application provide a simulation training method for a model, the method comprising: transmitting training data determined according to a training instruction to a target graphics processing unit, the training data including data of at least one neural network layer of the model indicated by the training instruction; obtaining, for any one of the at least one neural network layer, the training duration of the target graphics processing unit training the neural network layer; and generating a simulation training result of the large model based on the training duration of the at least one neural network layer.

[0005] Secondly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods of embodiments of this application when executing the computer program.

[0006] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this application.

[0007] Fourthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, implements any of the methods described in the embodiments of this application.

[0008] The simulation training method for the model provided in this application involves transmitting training data determined according to training instructions to a target graphics processing unit (GPU). The training data includes data for at least one neural network layer of the model as indicated by the training instructions. For any one of the at least one neural network layer, the training time for the target GPU to train that neural network layer is obtained. Based on the training time of the at least one neural network layer, simulation training results for a large model are generated. Therefore, simulation training can be performed using only the target GPU, meaning that only a single real GPU is needed to achieve model simulation training, reducing the consumption of real computing resources by the simulation training.

[0009] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0010] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this application and should not be construed as limiting the scope of this application.

[0011] Figure 1 This illustration shows an application scenario diagram of the simulation training method for the model according to an embodiment of this application;

[0012] Figure 2 A first flowchart of a simulation training method for a model according to an embodiment of this application is shown;

[0013] Figure 3 A second flowchart of the simulation training method for the model according to an embodiment of this application is shown;

[0014] Figure 4 A schematic diagram of the simulation training method for the model according to an embodiment of this application is shown;

[0015] Figure 5 A block diagram of a simulation training apparatus for a model according to an embodiment of this application is shown;

[0016] Figure 6 A block diagram of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0017] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0018] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.

[0019] The following terms will be used in the following text:

[0020] Large Language Model (LLM): A type of natural language processing model based on deep learning, typically consisting of hundreds of millions to trillions of parameters. LLMs are trained on large-scale text data and are capable of performing various natural language processing tasks, such as text generation, translation, question answering, and dialogue systems.

[0021] Graphics Processing Unit (GPU): Also known as a graphics processor, it is widely used in fields such as graphics processing, parallel computing, and high-performance computing.

[0022] Bus bandwidth (Busbw): A metric used to evaluate hardware utilization efficiency. Since algorithm bandwidth (algbw) cannot directly reflect the true performance of hardware in multi-node communication, Busbw is introduced. Busbw adjusts the algorithm bandwidth based on different communication modes to reflect the communication speed between nodes.

[0023] Parallel strategy: A framework for handling deep learning tasks. Training deep learning models typically requires a lot of computing resources, while parallel strategies allow computation to be performed simultaneously on multiple processing units (such as multiple GPUs or even clusters), thereby shortening training time.

[0024] Data parallelism (DP) is a parallel strategy that divides the training data into multiple mini-batches and performs forward and backward propagation simultaneously on different processing units. The model parameters are identical within each processing unit. Afterward, the gradients calculated by each processing unit are aggregated and used to update the overall model parameters.

[0025] Tensor Parallelism (TP): A parallel strategy suitable for models that are so large that they cannot be run on a single processing unit. This method distributes different parts of the model across different computational units. Each unit computes its assigned part, and intermediate results are then passed between the units.

[0026] Pipeline parallelism (PP): Pipeline parallelism is a special case of model parallelism, which distributes different layers of the model across multiple processing units and forms a pipeline between different computational units. Each processing unit can run in parallel on different batches when more inputs arrive.

[0027] Expert Parallelism (EP) is a parallel strategy that expands model complexity by introducing expert models, typically applicable to sparse activation (i.e., not all model parameters are activated simultaneously). In this approach, the entire model is divided into multiple "experts," each corresponding to a part of the model. When processing a task, only a subset of experts is activated for computation, allowing for the training of larger models without the corresponding huge computational overhead, thus improving model performance.

[0028] Collective Communication (CCM) is a common communication pattern used in parallel computing and distributed systems, involving data exchange and synchronization between multiple processes (or threads). In contrast to point-to-point communication, CCM allows a group of processes to exchange data in a coordinated manner, rather than simply transferring data between two processes.

[0029] Broadcast: A collective communication mode in which one process sends data to all other processes. For example, one process broadcasts a set of model parameters to all other worker processes.

[0030] Gather: A communication mode that aggregates data from all processes into one process; it is an operation that combines multiple data fragments into one.

[0031] All-Reduce: A communication mode that aggregates data from multiple processes into a single result through a certain operation (such as summation, finding the maximum value, etc.), and all processes will receive the reduced result.

[0032] All-Gather: A communication mode where each process collects data from all other processes, ultimately providing each process with the data set from all processes. For example, in model training, processes can share local gradients to form a global view.

[0033] Reduce Scatter: A collective communication mode that combines the functions of reduce and scatter. It is used to perform reduction operations (such as summation, averaging, etc.) on data from multiple processes or nodes, and then distribute the reduced results to various processes or nodes.

[0034] All-to-All: A collective communication mode in which each process sends data to all other processes, forming a complete communication network.

[0035] The rise of large language models has brought revolutionary changes to fields such as natural language processing, image recognition, and augmented reality. However, these massive models also place higher demands on training resources. Training large language models typically involves billions or even hundreds of billions of parameters, with enormous datasets, and requires highly optimized hardware architectures, including high-performance graphics processing units, network interface cards (NICs), storage devices, and other computing resources. The configuration of these hardware resources is crucial to model training efficiency, model performance, and cost reduction. Traditional model training often relies on manual parameter tuning and repeated trial and error, which not only wastes a lot of time and computing resources but also easily leads to inaccurate training results. Therefore, it is necessary to conduct simulation training of models and evaluate model performance based on simulation training results. ASTRA-Sim (Distributed Machine Learning System Simulator) is one of the main simulation training methods in related technologies, but it requires tracking training logs on a real physical cluster, meaning it still consumes a lot of real computing resources. In addition, real physical clusters, or even those that have not yet been built, are difficult to obtain for simulation, thus greatly reducing their practicality. Furthermore, ASTRA-Sim struggles to meet the demands of various new model architectures, and its simulation accuracy is relatively low in large-scale clusters. Another simulation method in related technologies is real-world streaming simulation, which offers higher accuracy but often requires more than 6 hours of simulation time, resulting in long training times and low efficiency.

[0036] Based on this, this application provides a simulation training method for a model, an electronic device, a storage medium, and a program product to improve simulation training efficiency. Figure 1 This is a schematic diagram illustrating an application scenario for simulation training of a model provided in an embodiment of this application, such as... Figure 1As shown, the scenario includes: simulation training equipment, which can be a terminal device or a server. The terminal device can be a mobile phone, tablet, desktop computer, laptop, etc., and the server can be a physical server or a cloud server, etc.

[0037] Figure 1 The simulation training equipment used in this example is a desktop computer. It should be understood that... Figure 1 The illustration is merely a schematic representation of the application scenario of the simulation training method for the model involved in this application and does not constitute a limitation on the technical solution of this application. In other embodiments, the application scenario of the simulation training method for the model involved in this application may include more or fewer components.

[0038] In one embodiment, the simulation training device may include a display screen and a target graphics processing unit (i.e., a target GPU). The display screen is configured to display the user interface (UI) of the simulation training device and provide an interface for interaction and information exchange between the simulation training device and the user. The UI involved in this application embodiment can be configured as a medium interface for user operation and interaction with the simulation training device. As an interaction interface with the user, the UI can convert the computer language of the simulation training device into a form that the user can accept and recognize, including displayed images, text, buttons, etc. A common form of UI is a graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the display screen of the simulation training device. The control can include visual interface elements such as icons, buttons, menus, tabs, and text boxes. The control can be implemented as a visual functional interface of the simulation training device. When the control receives a corresponding trigger operation from the user, the corresponding functional interface of the simulation training device receives the corresponding processing instruction, thereby enabling the simulation training device to respond to the processing instruction and perform functional processing. For example, each user action (e.g., click, double-click, edit, etc.) inputs a corresponding instruction to the simulation training device through the corresponding GUI control, thereby triggering the simulation training device to execute relevant processing and update the GUI based on the processed results. In one embodiment, the user can operate the corresponding GUI (e.g., training controls, etc.) to send training instructions for the model to the simulation training device. Correspondingly, the simulation training device, in response to the training instruction, transmits the training data determined according to the training instruction to the target graphics processing unit. The training data includes data for at least one neural network layer of the model indicated by the training instruction. Furthermore, for any one of the at least one neural network layer, the training duration of the neural network layer trained by the target graphics processing unit is obtained; based on the training duration of the at least one neural network layer, the simulation training result of the model is generated. In other words, the target graphics processing unit is used to train at least one neural network layer based on the training data received. The model can be a large model or a regular model; the type of model is not specifically limited in this application.

[0039] In another implementation, the scenario may also include: a client ( Figure 1(Not shown in the image), the client can communicate with the simulation training device. The client may have the aforementioned display screen, and the simulation training device may have the aforementioned target graphics processor. Accordingly, the user can operate the client to input training commands. The client sends the received training commands to the simulation training device. In response to the training commands, the simulation training device transmits the training data determined according to the training commands to the target graphics processing unit. The training data includes data of at least one neural network layer of the model indicated by the training commands. Furthermore, for any one of the at least one neural network layer, the training time of the target graphics processing unit for training the neural network layer is obtained; based on the training time of the at least one neural network layer, the simulation training result of the model is generated.

[0040] As can be seen, in the simulation training process of the above model, only the target graphics processing unit is needed for simulation training. That is, only one real graphics processing unit is required to realize the simulation training of the model, which reduces the consumption of real computing resources by the simulation training. Furthermore, since the simulation training results of the model can be obtained by training only at least one neural network layer of the model, without the need to train the complete model, the simulation training time is reduced and the simulation training efficiency is improved.

[0041] It should be noted that the application scenarios or examples provided in the embodiments of this application are for ease of understanding, and the embodiments of this application do not specifically limit the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0042] The technical solution of this application and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0043] Figure 2 A flowchart illustrating the simulation training method of the model according to an embodiment of this application is shown. Figure 2 The method in can be derived from Figure 1 The simulation training equipment in the middle is used for execution. For example... Figure 2 As shown, the method may include steps S201 to S203.

[0044] Step S201: The training data determined according to the training instructions is transmitted to the target graphics processing unit. The training data includes data of at least one neural network layer of the model indicated by the training instructions.

[0045] Step S202: For any neural network layer in at least one neural network layer, obtain the training time of the target graphics processing unit training neural network layer.

[0046] Step S203: Generate the simulation training results of the model based on the training time of at least one neural network layer.

[0047] Considering that the training duration of a model can often characterize its performance—for example, it can be used to analyze whether the model architecture is reasonable (generally, the longer the training duration, the lower the reasonableness of the model architecture), and to evaluate whether the parallel strategy used in training is reasonable (generally, the longer the training duration, the lower the reasonableness of the parallel strategy), etc.—in some embodiments of this application, during the simulation training process, the training duration of the target graphics processor for any one of the at least one neural network layer is obtained, and the simulation training results of the model are generated based on the training duration of the at least one neural network layer.

[0048] Specifically, in some embodiments, the simulation training device includes not only the aforementioned display screen and target graphics processing unit, but also a processor (CPU) and a simulation training apparatus. Accordingly, the user can operate the corresponding GUI on the display screen to send training instructions to the processor. In response to these training instructions, the processor sends a training start instruction to the target graphics processing unit. This training start instruction is generated based on training parameters pre-determined according to the user's training needs. Alternatively, the user can operate the corresponding GUI on the display screen to edit training parameters and send training instructions to the processor based on these parameters. The processor sends the training start instruction to the target graphics processing unit according to the training parameters in the training instruction. When the simulation training apparatus detects the training start instruction, it intercepts the instruction, determines the training data based on the training parameters in the instruction, and sends the training data to the target graphics processing unit. The target graphics processing unit trains at least one neural network layer based on the received training data. The simulation training apparatus can obtain the training time for any of the at least one neural network layer trained by the target graphics processing unit; the training times of the at least one neural network layer are then summed to obtain the simulation training result of the model.

[0049] The training instructions can specify the simulation training of the model in a virtual distributed training system. This means the target graphics processor (GPU) is a real GPU within the virtual distributed training system, and this real GPU simulates the distributed training of the model. Training parameters can include environment parameters and model parameters. Environment parameters can include the number of hosts (physical machines, such as servers and terminal devices) used for parallel simulation training (nnodes), the number of GPUs per host (nproc_per_node), the communication speed of the GPUs, the number of network interface cards (NICs) in the host, and the communication speed of the NICs. Model parameters can include the parallelism strategy, the number of neural network layers in the model, the communication mode and weights corresponding to each neural network layer, the batch size, and the number of experts. Parallelism strategies can include Tensor Parallelism (TP), Pipeline Parallelism (PP), and Expert Parallelism (EP), among others. The communication mode of any neural network layer can be any of the following: broadcast, gather, all-reduce, all-gather, or all-to-all. The weights corresponding to a neural network layer are the parameters used in model training, i.e., the parameters involved in the data processing of that layer. The data for at least one neural network layer of the model can include data from the training instructions or data determined based on those instructions.

[0050] In some implementations, after obtaining the simulation training results, the model's performance can be analyzed based on these results to obtain analysis results. For example, the simulation training results can be used to analyze whether the model architecture is reasonable, or to evaluate the rationality of the model's parallel strategy.

[0051] The simulation training method for the model provided in this application involves transmitting training data determined according to training instructions to a target graphics processing unit (GPU). The training data includes data for at least one neural network layer of the model as indicated by the training instructions. For any one of the at least one neural network layer, the training time for the target GPU to train that neural network layer is obtained. Based on the training time of the at least one neural network layer, simulation training results for a large model are generated. Therefore, simulation training can be performed using only the target GPU, meaning that only a single real GPU is needed to achieve model simulation training, reducing the consumption of real computing resources by the simulation training.

[0052] To facilitate training of the target graphics processing unit, in some embodiments, before transmitting the training data determined according to the training instructions to the target graphics processing unit in step S201, the following may be included:

[0053] Based on the parallel strategy included in the training instructions, the total number of graphics processing units for parallel simulation training is determined; based on the total number and the number of neural network layers in the model included in the training instructions, at least one neural network layer is determined, the number of which is based on the parallel strategy and the number of neural network layers; based on the input data attributes of the first neural network layer in the at least one neural network layer, the first input data of the first neural network layer is generated; the total number, the first input data, the communication mode and weights corresponding to the at least one neural network layer are determined as the training data for the at least one neural network layer.

[0054] Specifically, the user's training requirements can be obtained in advance. Based on these requirements, the input data attributes of each layer in the model can be analyzed, and a correspondence between neural network layers and input data attributes can be established. When training instructions are obtained, the tensor model parallelization (TP) is obtained from the parallelization strategy included in the training instructions, and the value of TP is determined as the total number of graphics processing units (GPUs) for parallel simulation training. If the parallelization strategy only includes TP, the model is vertically divided into an average number of parts (e.g., 8 parts, each including all neural network layers of the model) according to the value of TP (e.g., TP = 8). The neural network layers included in any part are determined as the network layers for simulation training of the target GPU. If the parallelization strategy includes both TP and PP, the model can first be vertically divided into an average number of parts according to the value of TP, and then horizontally divided according to the value of PP for any part. That is, the number of neural network layers in the model included in the training instructions is divided by the value of PP to obtain the number of at least one neural network layer. This number of neural network layers is then selected from the model as the network layers for simulation training of the target GPU. The method involves obtaining the input data attributes of the first neural network layer from the correspondence between neural network layers and input data attributes, and generating the first input data of the first neural network layer based on the obtained input data attributes. It also involves obtaining the communication mode and weights of each neural network layer from the model parameters included in the training instructions. The determined total quantity, the generated first input data, and the obtained communication mode and weights corresponding to the least one neural network layer are used to define the data for at least one neural network layer. Training data is generated based on the data of the at least one neural network layer. That is, the training data may also include other data, which is not specifically limited in this application. The communication modes of different neural network layers can be the same or different. When a neural network layer does not require data communication between graphics processing units undergoing parallel simulation training, the communication mode can be empty.

[0055] Furthermore, the number of neural network layers selected from the model as the network layers for simulation training of the target graphics processor can be chosen randomly or sequentially from front to back. These neural network layers can be continuous or discontinuous.

[0056] For example, TP = 8, PP = 4, and the number of neural network layers in the model is 32. The total number of graphics processing units (GPUs) used for parallel simulation training can be determined to be 8, meaning the virtual distributed training system includes 8 GPUs. The number of at least one neural network layer is 32 / 4 = 8. Eight consecutive neural network layers can be randomly selected as the target GPU simulation training layers, or the first to eighth neural network layers of the model can be determined as the target GPU simulation training layers, or the third to tenth neural network layers can be determined as the target GPU simulation training layers, etc. Furthermore, if the input data attribute of the first neural network layer among the acquired 8 neural network layers is an 8*8 image, then an 8*8 image is generated and used as the first input data of that first neural network layer.

[0057] By using the total number of graphics processing units (GPUs) trained in parallel simulation as data for at least one neural network layer, the target GPU can determine the validity of the processing result based on this number when it determines that the communication mode of any of the at least one neural network layer is not empty. By using the first input data as data for at least one neural network layer, the target GPU can perform simulated training on the at least one network layer based on the first input data. By using the communication mode and weights corresponding to at least one neural network layer as data for at least one neural network layer, the target GPU can determine whether it needs to communicate with other GPUs in the virtual distributed training system (i.e., GPUs other than the target GPU) based on the communication mode, and process the first input data of the corresponding neural network layer according to the weights. Furthermore, as described above, the target GPU only needs to train at least one neural network layer of the model, without needing to train the entire model, thus reducing the simulation training time and improving the simulation training efficiency.

[0058] It should be noted that the virtual distributed training system includes only one real graphics processing unit (GPU), namely the aforementioned target GPU; other GPUs are not present in the virtual distributed training system. For the target GPU, when the communication mode is not empty, the target GPU will consider that it needs to interact with other GPUs. At this time, the simulation training device can simulate communication to make the target GPU believe that it has engaged in data interaction with other GPUs in the corresponding communication mode, thereby continuing the simulation training. The specific process of simulated communication will be described later; repetitions will not be repeated here.

[0059] To effectively evaluate model performance, such as assessing the rationality of the model architecture and the parallelization strategy, some implementations include... Figure 3 As shown, step S202 may include steps S2021 to S2023:

[0060] Step S2021: For any neural network layer in at least one neural network layer, obtain the processing time of the neural network layer for the first input data. The processing time is the time from the input of the first input data into the neural network layer to the output of the processing result by the network layer.

[0061] Specifically, for any one of the at least one neural network layers, the input time of the first input data into the neural network layer and the output time of the neural network layer are determined, and the duration between the input time and the output time is determined as the processing time of the neural network layer for the first input data.

[0062] In some implementations, any neural network layer may include at least one sub-module. Accordingly, obtaining the processing time of the neural network layer for the first input data in step S2021 may include: for any sub-module, recording a first moment when second input data is detected input to the sub-module; wherein, if the sub-module is the first sub-module, the second input data is the first input data, and if the sub-module is not the first sub-module, the second input data is the data determined based on the processing result output by the target sub-module, the target sub-module being the sub-module preceding the first sub-module; recording a second moment when the processing result of the second input data is detected output by the sub-module; determining the duration from the first moment to the second moment as the sub-processing time of the sub-module for the second input data; and determining the processing time based on the sub-processing time of at least one sub-module.

[0063] Specifically, for any submodule, if it is the first submodule in a neural network layer, the first input data of that neural network layer is determined as the second input data of that submodule; if it is not the first submodule in a neural network layer, the second input data is determined based on the processing result output by the previous submodule (this can be the processing result output by the previous submodule, or data generated based on the processing result output by the previous submodule). When the second input data is detected entering the submodule, a point-marking operation is performed to record the corresponding first time point; and when the processing result of the second input data is detected, a point-marking operation is performed to record the corresponding second time point. The duration from the first time point to the second time point is determined as the sub-processing time of the submodule on the second input data. If the parallel strategy includes only at least one of TP, DP, and EP, the sub-processing times of at least one submodule are summed, and the sum is determined as the processing time of the corresponding neural network layer on the first input data. When the parallel strategy includes PP, since PP affects the number of neural network layers in the simulated training, i.e. the number of neural network layers in the simulated training will be reduced, the sub-processing times of at least one sub-module can be added together, and the sum can be divided by the value of PP. The result of the division is determined as the processing time of the corresponding neural network layer for the first input data.

[0064] It is understandable that when any neural network layer includes a submodule, that neural network layer is that submodule, meaning that the neural network layer is not subdivided, and the above dotting operation is performed based on that neural network layer.

[0065] It should be noted that when the structure of a submodule can be further subdivided, this submodule can be called a first-level submodule. Further subdivision of a first-level submodule yields a second-level submodule. When a second-level submodule can be further subdivided, a third-level module can be obtained, and so on. Correspondingly, the above-mentioned dotting operation can be performed on the final-level submodule (e.g., a second-level submodule, i.e., a second-level submodule that is no longer subdivided) or on the first-level submodule; this application does not specifically limit this. For any neural network layer, this neural network layer can be divided into multiple first-level submodules. Some first-level submodules can be further differentiated into second-level submodules, while some first-level submodules may not require further subdivision.

[0066] Taking the Transformer model as an example, since both the encoder and decoder parts of the Transformer model are composed of multiple identical layers stacked together—that is, each encoder and decoder layer has the same structure, but they play different roles and functions in the model—any neural network layer in the Transformer model can include multiple first-level sub-modules. Some first-level sub-modules can include multiple second-level sub-modules, and some first-level sub-modules cannot be further subdivided. The structural division of any neural network layer in the Transformer model is illustrated below:

[0067]

[0068] It should be noted that in the above example, the second-level submodule is -- indicating that the corresponding first-level submodule has not been further divided.

[0069] Therefore, by recording the first and second moments, the sub-processing time of the corresponding sub-module on the second input data can be accurately determined. Based on this sub-processing time, the processing time of the corresponding neural network layer on the first input data can be accurately determined. This provides a guarantee for improving the accuracy of simulation training results when determining simulation training results based on processing time.

[0070] To improve the accuracy of simulation training results, in some implementations, the communication parameters of each network layer in at least one neural network layer can be determined according to the training instructions. That is, the communication parameters include the sub-communication parameters corresponding to at least one sub-module in the neural network layer. Correspondingly, the above-mentioned processing result of detecting the output of the second input data from a sub-module can further include: if the sub-module is not the last sub-module of at least one neural network layer and the sub-communication parameters corresponding to the sub-module are not empty, simulating a target number of processing results based on the processing result of the second input data, where the target number refers to the total number of other graphics processing units included in the virtual distributed training system; concatenating the processing result of the second input data and the processing result of the target number to obtain the second input data of the next sub-module of the sub-module; and transmitting the second input data of the next sub-module of the sub-module to the target graphics processing unit.

[0071] Specifically, considering that in practical applications, the processing results of some sub-modules may require collective communication between the target graphics processing unit (GPU) and other graphics processing units (GPUs) included in the virtual distributed training system, and since only the target GPU is a real GPU during simulation training, while other GPUs in the virtual distributed training system are not present, in order to ensure the effective execution of simulation training, in some implementations, when the processing result of the second input data of a sub-module is detected, if it is determined that the sub-module is not the last sub-module of at least one neural network layer and collective communication is required, then a simulated communication operation is performed. That is, when the processing result of the second input data of a sub-module is detected, if it is determined that the sub-module is not the last sub-module of at least one neural network layer and the sub-communication parameter corresponding to the sub-module is not empty, then the processing result of the second input data is simulated based on the processing result of the second input data, and the processing result of the second input data and the processing result of the target number of processing results are concatenated to obtain the second input data of the next sub-module of the sub-module; and the second input data of the next sub-module is transmitted to the target GPU, so that the target GPU determines that the collective communication is complete based on the second input data and inputs it into the next sub-module of the sub-module.

[0072] The simulation of the target quantity processing result based on the processing result of the second input data can be achieved by either copying the processing result of the second input data to obtain the target quantity processing result, or by transforming the processing result of the second input data according to a preset transformation rule. The target quantity is the difference between the value of TP and 1. The preset transformation rule could be, for example, replacing data at a preset position in the processing result of the second input data with default data. The simulation method for the target quantity processing result can be set as needed in practical applications, and this application does not impose specific limitations on it.

[0073] Therefore, when the processing result of the second input data output by the submodule is detected and the sub-communication parameter corresponding to the submodule is not empty, the processing result of the target number is simulated, and the processing result of the second input data and the processing result of the target number are concatenated. In the case that the submodule is not the last submodule of at least one neural network layer, the concatenated data is transmitted to the target graphics processing unit, thereby erasing the real communication and realizing simulated communication. This allows the target graphics processing unit to determine the completion of set communication based on the result of simulated communication, i.e., the concatenated data, and to train the next submodule, ensuring the effective implementation of simulated training.

[0074] It should be noted that when it is determined that the submodule is the last submodule of at least one neural network layer, since the target graphics processing unit does not need to continue to perform training operations, the aforementioned simulated communication operation can be performed or not.

[0075] Furthermore, the aforementioned determination of communication parameters for each network layer in at least one neural network layer based on training instructions may include: obtaining the number of hosts for parallel simulation training from the training instructions; determining the communication type as intra-host communication when the number of hosts for parallel simulation training is one; determining the communication type as inter-host communication when the number of hosts for parallel simulation training is not one and each host contains one graphics processor; and determining the communication type as including both intra-host and inter-host communication when the number of hosts for parallel simulation training is not one and at least one host contains not one graphics processor. Additionally, for any submodule in a neural network layer, if it is determined that the processing result output by the submodule requires aggregate communication based on its module type (e.g., embedded submodule), then the corresponding communication mode (e.g., broadcast) and parallel strategy for the submodule are obtained from the training instructions, as well as the communication speed corresponding to the determined communication type are obtained from the training instructions, and the communication type, communication mode, parallel strategy, and communication speed are determined as the sub-communication parameters of the submodule.

[0076] The process of obtaining the communication speed corresponding to a specific communication type from the training instructions can include: when the communication type is intra-machine communication, obtaining the communication speed of the graphics processor from the training instructions and using it as the communication speed corresponding to the communication type (i.e., intra-machine communication speed); when the communication type is inter-machine communication, obtaining the communication speed of the network card in the host and the number of network cards in the host from the training instructions, and multiplying the communication speed of the network card by the number of network cards to obtain the communication speed corresponding to the communication type (i.e., inter-machine communication speed); when the communication type includes both intra-machine communication and inter-machine communication, determining the intra-machine communication speed and inter-machine communication speed as the communication speed corresponding to the communication type in the aforementioned manner.

[0077] Step S2022: Determine the communication duration of the target graphics processing unit based on the communication parameters corresponding to the neural network layer. The communication duration is used to characterize the communication duration between the target graphics processing unit and other graphics processing units in the scenario of the virtual distributed training system. The communication parameters are determined according to the training instructions.

[0078] As mentioned earlier, in cases where the target graphics processing unit (GPU) needs to perform ensemble communication with other GPUs, the actual communication is omitted. However, in the actual training process, ensemble communication takes time, and this time is part of the training time. Therefore, this application calculates the corresponding communication duration based on the relevant communication parameters for the locations where ensemble communication is required.

[0079] In some implementations, determining the communication duration of the target graphics processing unit based on the communication parameters corresponding to the neural network layer in step S2022 may include: for any submodule, if the sub-communication parameters corresponding to the submodule are empty, determining that the sub-communication duration of the submodule is zero, where the sub-communication duration is the communication duration between the target graphics processing unit and other graphics processing units for the processing result output by the submodule; if the sub-communication parameters corresponding to the submodule are not empty, determining the sub-communication duration of the submodule based on the sub-communication parameters and the data length of the processing result output by the submodule; and determining the communication duration based on the sub-communication duration of at least one submodule.

[0080] It is understandable that when the sub-communication parameter corresponding to the sub-module is empty, it indicates that the target graphics processing unit does not need to perform collective communication with other graphics processing units, so the sub-communication duration corresponding to the sub-module is zero. When the sub-communication parameter corresponding to the sub-module is not empty, it indicates that the target graphics processing unit needs to perform collective communication with other graphics processing units. Therefore, the sub-communication duration corresponding to the sub-module is determined based on the sub-communication parameter and the data length of the processing result output by the sub-module. The sub-communication durations corresponding to at least one sub-module included in the neural network layer are added together, and the sum is determined as the communication duration corresponding to the target graphics processing unit in the corresponding neural network layer.

[0081] Therefore, based on the communication parameters corresponding to the neural network layers, the communication duration corresponding to the target graphics processing unit is determined, ensuring the accuracy of the communication duration, thereby ensuring the accuracy of the training duration determined based on the communication duration, and further ensuring the accuracy of the simulation training results determined based on the training duration.

[0082] During the distributed training of the model, multiple graphics processing units (GPUs) trained in parallel may reside on the same host (i.e., a physical device, such as a server or terminal device) or on different hosts. Some GPUs may reside on the same host while others reside on different hosts. Correspondingly, communication types can include intra-machine communication and inter-machine communication, and different communication types correspond to different initial communication speeds. Based on this, in some implementations, sub-communication parameters may include communication types. Accordingly, the aforementioned determination of the initial communication speed based on sub-communication parameters may include: when the communication type is intra-machine communication, obtaining the communication mode and intra-machine communication speed from the sub-communication parameters; determining a first communication speed based on the speed calculation method corresponding to the communication mode and the intra-machine communication speed; and determining the first communication speed as the initial communication speed. Here, intra-machine communication indicates that the target GPU and other GPUs reside on the same host. When the communication type is inter-machine communication, the communication mode and inter-machine communication speed are obtained from the sub-communication parameters. A second communication speed is determined based on the speed calculation method corresponding to the communication mode and the inter-machine communication speed, and this second communication speed is set as the initial communication speed. Here, inter-machine communication indicates that the target graphics processing unit and other graphics processing units are located in different hosts. When the communication type includes both intra-machine communication and inter-machine communication, the first communication speed and the second communication speed are determined respectively according to the aforementioned method, and both are set as the initial communication speed.

[0083] The determination of the first communication speed based on the speed calculation method corresponding to the communication mode and the internal communication speed can include: in the case of All-Reduce, determining the first communication speed based on the parallel strategy in the sub-communication parameters and the following formula 1 based on the internal communication speed; in the case of ReduceScatter, determining the first communication speed based on the parallel strategy in the sub-communication parameters and the following formula 2 based on the internal communication speed; in the case of All-Gather, determining the first communication speed based on the parallel strategy in the sub-communication parameters and the following formula 3 based on the internal communication speed; and in the case of Reduce, determining the first communication speed based on the following formula 4 based on the internal communication speed.

[0084] Formula 1: v = algbw * (2 * (n-1) / n);

[0085] Formula 2: v = algbw * ((n-1) / n);

[0086] Formula 3: v = algbw * ((n-1) / n);

[0087] Formula 4: v = algbw * 1.

[0088] In Formulas 1 to 4 above, v represents the first communication speed, which can also be called bus bandwidth; algbw represents the internal communication speed, which can also be called algorithm bandwidth; n is related to the parallel strategy. When the parallel strategy is TP, n is the value of TP; when the parallel strategy is DP, n is the value of DP.

[0089] The process of determining the second communication speed based on the speed calculation method corresponding to the communication mode and the inter-machine communication speed is similar to the process of determining the first communication speed, as described above. This will not be repeated here. The difference lies in the fact that in Formulas 1 to 4, 'v' represents the second communication speed, and 'algbw' represents the inter-machine communication speed.

[0090] Therefore, by determining the corresponding initial communication speed based on different communication types and modes, the accuracy of the initial communication speed is ensured, which in turn ensures the accuracy of the sub-communication duration determined based on the initial communication speed, thus providing a guarantee for improving the accuracy of simulation training results.

[0091] In some implementations, determining the sub-communication duration corresponding to the sub-module based on the sub-communication parameters and the data length of the processing result output by the sub-module may include: determining an initial communication speed based on the sub-communication parameters; obtaining a target first ratio corresponding to the data length of the processing result from the correspondence between data length and a first ratio; determining a target communication speed based on the initial communication speed and the target first ratio; and determining the sub-communication duration corresponding to the sub-module based on the data size of the processing result and the target communication speed.

[0092] Specifically, considering that in practical applications, communication speed is related to data length, and the larger the data length, the smaller the speed attenuation, in order to accurately determine the sub-communication duration, in some implementations, a correspondence between data length and a first ratio can be preset. During the determination of the sub-communication duration, the target first ratio corresponding to the processed data length is obtained from this correspondence; the determined initial communication speed is multiplied by the target first ratio to obtain the target communication speed; and the processed data length is divided by the target communication speed to obtain the sub-communication duration corresponding to the sub-module, i.e., sub-communication duration = data length / (initial communication speed * target first ratio). Here, the first ratio is used to characterize the degree of communication speed attenuation. The first ratio can be a decimal within a preset range, such as a decimal between 0 and 1. The larger the target first ratio, the smaller the degree of communication speed attenuation, i.e., the larger the target communication speed.

[0093] Therefore, by determining the target first ratio corresponding to the data length of the processing result, and determining the target communication speed based on the initial communication speed and the target first ratio; and by determining the sub-communication duration corresponding to the sub-module based on the data length of the processing result and the target communication speed, the accuracy of the determined sub-communication duration is improved because the impact of data size on communication speed is taken into account in this process.

[0094] Furthermore, when the communication type includes intra-machine communication and inter-machine communication, the target communication speed may include a third communication speed corresponding to intra-machine communication (i.e., the product of the intra-machine communication speed and the target attenuation rate) and a fourth communication speed corresponding to inter-machine communication (i.e., the product of the initial communication speed of inter-machine communication and the target attenuation rate). Accordingly, the aforementioned determination of the sub-communication duration corresponding to the sub-module based on the data size of the processing result and the target communication speed includes: determining a first candidate duration based on the data size of the processing result and the third communication speed; determining a second candidate duration based on the data size of the processing result and the fourth communication speed; and determining the larger candidate duration among the first and second candidate durations as the sub-communication duration.

[0095] Specifically, when the communication type includes intra-machine communication and inter-machine communication, the data length of the processed result is divided by the third communication speed to obtain the first candidate duration. The data length of the processed result is divided by the fourth communication speed to obtain the second candidate duration. The first candidate duration and the second candidate duration are compared to obtain the maximum candidate duration, and the maximum candidate duration is determined as the sub-communication duration.

[0096] Therefore, when the communication types include intra-machine communication and inter-machine communication, by calculating the first candidate duration and the second candidate duration, and determining the maximum candidate duration between the first candidate duration and the second candidate duration as the sub-communication duration, the time consumption of both intra-machine communication and inter-machine communication is taken into account, ensuring the accuracy of the determined sub-communication duration, and thus providing a guarantee for improving the accuracy of simulation training results.

[0097] In some implementations, to improve the accuracy of sub-communication duration, the corresponding sub-communication duration can be determined based on the ratio corresponding to the communication mode. That is, the aforementioned method of dividing the data size of the processing result by the target communication speed to obtain the sub-communication duration corresponding to the sub-module can include: dividing the data size of the processing result by the target communication speed to obtain the third candidate duration corresponding to the sub-module; obtaining the target ratio associated with the communication mode of the sub-module from the correlation between communication mode and ratio; and multiplying the third candidate duration by the target ratio to obtain the sub-communication duration corresponding to the sub-module.

[0098] Step S2023: Determine the training duration of the neural network layer based on the processing duration and communication duration.

[0099] In some implementations, the processing time and communication time can be added together to obtain the training time of the corresponding neural network layer.

[0100] In other implementations, considering that there may be partial overlap between communication duration and processing duration, the processing duration and communication duration can be added together, and the result can be multiplied by a preset second ratio to determine the training duration of the corresponding neural network layer.

[0101] Therefore, for any neural network layer, the training duration is determined by acquiring the processing time of the neural network layer for the first input data and the communication duration of the target graphics processing unit corresponding to that neural network layer. This approach considers both the data processing process and the communication process, thus improving the accuracy of the determined training duration and consequently enhancing the accuracy of the simulation training results determined based on the training duration.

[0102] Furthermore, to facilitate the management of various parameters during model simulation training, in some implementations, these parameters can be saved through configuration files. Specifically, for example... Figure 4 As shown, after determining the communication parameters of each network layer in at least one neural network layer according to the training instructions, the method may further include: saving the determined communication parameters, model parameters and environment parameters contained in the training instructions to a configuration file.

[0103] Correspondingly, before determining the communication duration of the target graphics processing unit based on the communication parameters of the neural network layer, the process may further include: obtaining the communication parameters of the neural network layer from the data pre-stored in the configuration file, wherein the data pre-stored in the configuration file is determined according to the training instructions, and the data pre-stored in the configuration file includes communication parameters, model parameters, and environment parameters, etc.

[0104] Furthermore, after obtaining the processing time of the neural network layer for the first input data, the process may further include saving the processing time to a configuration file.

[0105] In some implementations, the configuration file may include relevant data of each sub-module contained in at least one neural network layer of the model specified in the training instructions. Accordingly, saving the determined communication parameters, model parameters included in the training instructions, and environment parameters to the configuration file may include: saving the environment parameters to the storage location corresponding to the environment parameter field in the configuration file; for each sub-module, obtaining the sub-communication parameters corresponding to that sub-module from the communication parameters, obtaining the weights corresponding to that sub-module from the model parameters, and saving the sub-communication parameters and weights to the storage location corresponding to the parameter field of that sub-module in the configuration file; and saving all parameters in the model parameters except the weights to the storage location corresponding to the model parameter field in the configuration file. Correspondingly, obtaining the communication parameters corresponding to the neural network layer from the data pre-stored in the configuration file may include: for each sub-module, obtaining the sub-communication parameters corresponding to that sub-module from the data pre-stored in the configuration file, and determining the sub-communication parameters of at least one sub-module included in the neural network layer as the communication parameters corresponding to the network layer.

[0106] Furthermore, the aforementioned method of saving the processing time to the configuration file may include: for each submodule, upon obtaining the sub-processing time of that submodule for the second input data, saving the sub-processing time to the storage location corresponding to the processing time field of that submodule in the configuration file. Correspondingly, the aforementioned method of determining the processing time based on the sub-processing time of at least one submodule may include: if it is determined that the sub-processing time of the last submodule of the corresponding neural network layer has been stored in the configuration file, then obtaining the sub-processing times corresponding to each submodule in that neural network layer from the configuration file; summing the obtained sub-processing times, and determining the sum as the processing time of that neural network layer for the first input data.

[0107] It should be noted that, based on the configuration file, the processing time of each neural network layer for the first input data and the communication time of the target graphics processing unit (GPU) during the training of each neural network layer can be determined, using the neural network layer as the dimension. Specifically, for each neural network layer, the processing time for the first input data is first determined according to the configuration file in the manner described above. Then, the communication parameters corresponding to the neural network layer are obtained from the configuration file in the manner described above, and the communication time of the target GPU is determined based on these communication parameters. Alternatively, the communication parameters corresponding to the neural network layer are first obtained from the configuration file in the manner described above, and the communication time of the target GPU is determined based on these communication parameters. Then, the processing time of the neural network layer for the first input data is determined in the manner described above. Alternatively, the processing time of the neural network layer for the first input data and the communication time of the target GPU can be determined simultaneously based on the configuration file in the manner described above.

[0108] Alternatively, the processing time of each neural network layer for the first input data and the communication time of the target graphics processing unit (GPU) during the training of each neural network layer can be determined using at least one neural network layer as a dimension. That is, based on the configuration file, the processing time of each neural network layer for the first input data is first determined, and then the communication time of the target GPU in each neural network layer is determined. Alternatively, based on the configuration file, the communication time of the target GPU in each neural network layer is first determined, and then the processing time of each neural network layer for the first input data is determined. Or, based on the configuration file, the processing time of each neural network layer for the first input data and the communication time of the target GPU in each neural network layer are determined simultaneously. It can also be done from the perspective of sub-modules.

[0109] Alternatively, based on the configuration file, the sub-processing time of each sub-module and the sub-communication time of the target graphics processing unit can be determined by sub-module. The total processing time is obtained by adding the sub-processing times and the total communication time. Finally, the simulation training results are obtained by adding the total processing time and the total communication time.

[0110] The specific method for determining the processing and communication duration based on the configuration file is not specifically limited in this application, and can be set as needed in practical applications.

[0111] Therefore, by saving the parameters of each model to a configuration file and retrieving the communication parameters from the configuration file before determining the communication duration, not only can the efficiency of model parameter management be improved, but the accuracy of the acquired communication parameters can also be guaranteed. By saving the processing duration to a configuration file, all relevant model data is stored in the same location, which not only improves the convenience of managing model-related data, but also avoids data confusion caused by simultaneous simulation training of multiple models, thus ensuring the accuracy of simulation training results.

[0112] To facilitate user understanding of the simulation training results, in some implementations, step S203 may further include: displaying the simulation training results. It should be noted that not only the total processing time and total communication time of at least one network can be displayed, but also the processing time and communication time of each neural network layer can be displayed, and some or all model parameters can be displayed (e.g., showing...). In one implementation, such as... Figure 4 As shown, the processing time and communication time of each sub-process can be obtained from the configuration file corresponding to the model, and the visualization results can be displayed based on the obtained data.

[0113] It is understood that in the simulation training method of the model provided in this application, one simulation training can correspond to one network architecture and one model weight. To facilitate users in selecting appropriate network architectures from different network architectures and appropriate model weights from different model weights, in some embodiments, users can also select multiple simulation training results to be compared after multiple simulation training sessions and submit a comparison request. Correspondingly, the method also includes: responding to the user's comparison request, obtaining the multiple simulation training results selected by the user, and displaying the multiple simulation training results.

[0114] For example, since the current model training process involves multiple all-to-all ensemble communications, and these communications are quite complex, to determine whether eliminating all-to-all ensemble communications helps improve the overall model performance, simulation training can be performed based on different training parameters. For instance, simulation training can be performed using the following three training parameters:

[0115]

[0116] Here, "model" indicates simulation training for Model 1; "GPUs" indicates the total number of graphics processors used for parallel simulation training is 1024; "TP" indicates tensor model parallelism, "DP" indicates data parallelism, "PP" indicates pipeline parallelism, and "RP" indicates expert parallelism; "Seq_len" indicates the sequence length; "Mbs" indicates the mini-batch size, which refers to the number of samples used for training in each iteration; "Hidden_size" indicates the dimension of the hidden layer; "GBS" (Global Batch Size) indicates the total number of samples processed by all graphics processors in a training step during distributed training; and "GA" (Gradient Accumulation) indicates the accumulated number of gradients in the mini-batch, at which point the model parameters are updated. In addition, other parameters can be included among the three training parameters, such as 64 experts and 4 topk (selecting the top k items from a set of candidates, where k is greater than or equal to 1).

[0117] Accordingly, in response to the user's comparison request, the simulation training results of the three training parameters selected by the user are obtained and displayed as follows:

[0118]

[0119] Wherein, DP comm represents the communication time involved when using data parallelism DP; Exposed TP&EP comm represents the communication time involved when using tensor parallelism TP and expert parallelism EP; Compute represents the total processing time, which is the sum of the processing times corresponding to each neural network layer; Total represents the simulation training result, which is the sum of the communication times and the total processing time.

[0120] See the simulation training results for TP=2, EP=8 and TP=16, EP=1. When EP=1, all-to-all set communication is eliminated. When EP=8, all-to-all set communication is required. Analysis of the simulation training results shows that, since 1300 is less than 4439 and 3307 is less than 6038, eliminating all-to-all set communication, while reducing inter-machine communication of EP, actually leads to greater computational overhead and more TP communication; that is, the model's performance is not improved.

[0121] As can be seen, by displaying multiple simulation training results, the corresponding model structure, model weights, and other training parameters can be analyzed based on these results, thereby determining the appropriate model architecture, appropriate model weights, and appropriate other training parameters.

[0122] It is important to emphasize that the simulation training method for the model provided in this application achieves single-machine simulation of multiple machines by erasing the underlying real communication and performing simulated communication. That is, only one real graphics processor is needed to complete the model simulation training, without requiring a real distributed physical environment, thus greatly reducing the consumption of real material resources. Furthermore, the simulation training can be completed within one minute, significantly shortening the simulation training time and improving its efficiency.

[0123] Corresponding to the application scenarios and methods provided in the embodiments of this application, the embodiments of this application also provide a simulation training device for a model, which can be applied to... Figure 1 The simulation training equipment shown, such as Figure 5 As shown, the device includes:

[0124] The first transmission module 501 is used to transmit training data determined according to the training instructions to the target graphics processing unit, wherein the training data includes data of at least one neural network layer of the model indicated by the training instructions.

[0125] The first acquisition module 502 is used to acquire the training time of the target graphics processing unit training the neural network layer for any one of the at least one neural network layers.

[0126] The generation module 503 is used to generate the simulation training results of the model based on the training time of the at least one neural network layer.

[0127] In some implementations, the first acquisition module 502 is specifically used for:

[0128] The processing time of the neural network layer on the first input data is obtained, wherein the processing time is the time from when the first input data is input into the neural network layer to when the network layer outputs the processing result.

[0129] Based on the communication parameters corresponding to the neural network layer, the communication duration corresponding to the target graphics processing unit is determined. The communication duration is used to characterize the communication duration between the target graphics processing unit and other graphics processing units in the scenario of the virtual distributed training system. The communication parameters are determined according to the training instructions.

[0130] The training duration is determined based on the processing duration and the communication duration.

[0131] In some embodiments, the neural network layer includes at least one sub-module, and the first acquisition module 502 is further specifically used for:

[0132] For any submodule, when the second input data is detected to be input into the submodule, a first moment is recorded. If the submodule is the first submodule, then the second input data is the first input data. If the submodule is not the first submodule, then the second input data is the data determined according to the processing result output by the target submodule, which is the submodule preceding the first submodule.

[0133] If the processing result of the second input data output by the submodule is detected, the second moment is recorded.

[0134] The duration from the first moment to the second moment is determined as the sub-processing duration of the sub-module for the second input data.

[0135] The processing time is determined based on the sub-processing time of the at least one sub-module.

[0136] In some embodiments, the communication parameters include sub-communication parameters corresponding to each of the at least one sub-module, and the device further includes an analog module, a splicing module, and a second transmission module.

[0137] The simulation module is used to simulate the processing result of a target number based on the processing result of the second input data, provided that the submodule is not the last submodule of the at least one neural network layer and the sub-communication parameter corresponding to the submodule is not empty. The target number refers to the total number of other graphics processing units included in the virtual distributed training system.

[0138] The splicing module is used to splice the processing result of the second input data and the processing result of the target quantity to obtain the second input data of the next submodule of the submodule.

[0139] The second transmission module is used to transmit the second input data of the next sub-module to the target graphics processing unit.

[0140] In some embodiments, the neural network layer includes at least one sub-module, and the communication parameters include sub-communication parameters corresponding to each of the at least one sub-module. The first acquisition module 502 is further specifically used for:

[0141] For any submodule, if the sub-communication parameters corresponding to the submodule are empty, the sub-communication duration corresponding to the submodule is determined to be zero. The sub-communication duration refers to the communication duration between the target graphics processing unit and the other graphics processing units for the processing result output by the submodule.

[0142] If the sub-communication parameters corresponding to the sub-module are not empty, the sub-communication duration corresponding to the sub-module is determined based on the sub-communication parameters and the data length of the processing result output by the sub-module.

[0143] The communication duration is determined based on the sub-communication duration corresponding to the at least one sub-module.

[0144] In some implementations, the first acquisition module 502 is further specifically used for:

[0145] The initial communication speed is determined based on the sub-communication parameters; the target first ratio corresponding to the data length of the processing result is obtained from the correspondence between data length and the first ratio; the target communication speed is determined based on the initial communication speed and the target first ratio; the sub-communication duration corresponding to the sub-module is determined based on the data size of the processing result and the target communication speed, wherein the first ratio is used to characterize the degree of communication speed attenuation.

[0146] In some implementations, the sub-communication parameters include the communication type, and the first acquisition module 502 is further specifically used for:

[0147] When the communication type is intra-machine communication, the communication mode and intra-machine communication speed are obtained from the sub-communication parameters. A first communication speed is determined according to the speed calculation method corresponding to the communication mode and the intra-machine communication speed. The first communication speed is determined as the initial communication speed. The intra-machine communication indicates that the target graphics processing unit and the other graphics processing units are located in the same host.

[0148] When the communication type is inter-machine communication, the communication mode and inter-machine communication speed are obtained from the sub-communication parameters. The second communication speed is determined according to the speed calculation method corresponding to the communication mode and the inter-machine communication speed. The second communication speed is determined as the initial communication speed. The inter-machine communication indicates that the target graphics processing unit and the other graphics processing units are located in different hosts.

[0149] When the communication type includes intra-machine communication and inter-machine communication, the first communication speed and the second communication speed are determined as the initial communication speed.

[0150] In some embodiments, when the communication type includes intra-machine communication and inter-machine communication, the target communication speed includes a third communication speed corresponding to the intra-machine communication and a fourth communication speed corresponding to the inter-machine communication. Accordingly, the first acquisition module is further specifically used for:

[0151] Based on the data size of the processing result and the third communication speed, a first candidate duration is determined; based on the data size of the processing result and the fourth communication speed, a second candidate duration is determined; the maximum candidate duration between the first candidate duration and the second candidate duration is determined as the sub-communication duration.

[0152] In some embodiments, the apparatus further includes a second acquisition module, configured to acquire the communication parameters corresponding to the neural network layer from data pre-stored in a configuration file before the first acquisition module determines the communication duration corresponding to the target graphics processing unit based on the communication parameters corresponding to the neural network layer, wherein the data pre-stored in the configuration file is determined according to the training instructions.

[0153] In some embodiments, the training instructions include a parallel strategy, the number of neural network layers in the model, the communication mode and weights corresponding to each neural network layer of the model, and the device further includes: a determining module, configured to: determine the total number of graphics processing units to be trained in parallel simulation according to the parallel strategy before the first transmission module transmits the training data determined according to the training instructions to the target graphics processing unit; determine the at least one neural network layer according to the parallel strategy and the number of neural network layers, wherein the number of the at least one neural network layer is determined based on the parallel strategy and the number of neural network layers; generate first input data for the first neural network layer according to the input data attributes of the first neural network layer in the at least one neural network layer that have been predetermined; and determine the total number, the first input data, the communication mode and weights corresponding to the at least one neural network layer as the training data for the at least one neural network layer.

[0154] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0155] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components illustrated as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0156] Figure 6 This is a block diagram of an electronic device used to implement embodiments of this application. Figure 6 As shown, the electronic device includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the computer program, it implements the method described in the above embodiments. The number of memories 601 and processors 602 can be one or more. In a specific implementation, the electronic device may also include a communication interface 603 for communicating with external devices and exchanging data.

[0157] In practical implementation, if the memory 601, processor 602, and communication interface 603 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0158] Optionally, in a specific implementation, if the memory 601, processor 602 and communication interface 603 are integrated on a single chip, the memory 601, processor 602 and communication interface 603 can communicate with each other through an internal interface.

[0159] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this application.

[0160] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method provided in this application.

[0161] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.

[0162] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0163] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0164] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0165] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0166] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0167] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0168] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0169] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0170] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0171] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0172] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A simulation training method for a model, characterized in that, The method includes: The training data determined according to the training instructions is transmitted to the target graphics processing unit, the training data including data of at least one neural network layer of the model indicated by the training instructions; For any one of the at least one neural network layers, obtain the training time for the target graphics processing unit to train the neural network layer; Based on the training time of the at least one neural network layer, the simulation training results of the model are generated.

2. The method according to claim 1, characterized in that, The step of obtaining the training time for the target graphics processing unit to train the neural network layer includes: The processing time of the neural network layer on the first input data is obtained, wherein the processing time is the time from when the first input data is input into the neural network layer to when the neural network layer outputs the processing result; Based on the communication parameters corresponding to the neural network layer, the communication duration corresponding to the target graphics processing unit is determined. The communication duration is used to characterize the communication duration between the target graphics processing unit and other graphics processing units in the scenario of the virtual distributed training system. The communication parameters are determined according to the training instructions. The training duration is determined based on the processing duration and the communication duration.

3. The method according to claim 2, characterized in that, The neural network layer includes at least one sub-module, and obtaining the processing time of the neural network layer for the first input data includes: For any submodule, when the second input data is detected to be input into the submodule, a first moment is recorded. If the submodule is the first submodule, then the second input data is the first input data. If the submodule is not the first submodule, then the second input data is the data determined according to the processing result output by the target submodule, where the target submodule is the submodule preceding the first submodule. If the processing result of the second input data output by the submodule is detected, the second moment is recorded; The duration from the first moment to the second moment is determined as the sub-processing duration of the sub-module for the second input data; The processing time is determined based on the sub-processing time of the at least one sub-module.

4. The method according to claim 3, characterized in that, The communication parameters include sub-communication parameters corresponding to each of the at least one sub-module, and further include, when the processing result of the second input data output by the sub-module is detected: If the submodule is not the last submodule of the at least one neural network layer and the sub-communication parameter corresponding to the submodule is not empty, the processing result of the target number is simulated based on the processing result of the second input data. The target number refers to the total number of other graphics processing units included in the virtual distributed training system. The processing result of the second input data and the processing result of the target quantity are concatenated to obtain the second input data of the next submodule of the submodule; The second input data of the next submodule is transmitted to the target graphics processing unit.

5. The method according to claim 2, characterized in that, The neural network layer includes at least one sub-module, and the communication parameters include sub-communication parameters corresponding to each of the at least one sub-module. Determining the communication duration corresponding to the target graphics processing unit based on the communication parameters corresponding to the neural network layer includes: For any submodule, if the sub-communication parameters corresponding to the submodule are empty, the sub-communication duration corresponding to the submodule is determined to be zero. The sub-communication duration is the communication duration between the target graphics processing unit and the other graphics processing units for the processing result output by the submodule. If the sub-communication parameters corresponding to the sub-module are not empty, the sub-communication duration corresponding to the sub-module is determined based on the sub-communication parameters and the data length of the processing result output by the sub-module. The communication duration is determined based on the sub-communication duration corresponding to the at least one sub-module.

6. The method according to claim 5, characterized in that, The step of determining the sub-communication duration corresponding to the sub-module based on the sub-communication parameters and the data length of the processing result output by the sub-module includes: The initial communication speed is determined based on the sub-communication parameters; From the correspondence between data length and first ratio, obtain the target first ratio corresponding to the data length of the processing result, whereby the first ratio is used to characterize the degree of attenuation of communication speed. The target communication speed is determined based on the initial communication speed and the first target ratio; Based on the data size of the processing result and the target communication speed, the sub-communication duration corresponding to the sub-module is determined.

7. The method according to claim 6, characterized in that, The sub-communication parameters include the communication type, and determining the initial communication speed based on the sub-communication parameters includes: When the communication type is intra-machine communication, the communication mode and intra-machine communication speed are obtained from the sub-communication parameters. A first communication speed is determined according to the speed calculation method corresponding to the communication mode and the intra-machine communication speed. The first communication speed is determined as the initial communication speed. The intra-machine communication indicates that the target graphics processing unit and the other graphics processing units are located in the same host. When the communication type is inter-machine communication, the communication mode and inter-machine communication speed are obtained from the sub-communication parameters. The second communication speed is determined according to the speed calculation method corresponding to the communication mode and the inter-machine communication speed. The second communication speed is determined as the initial communication speed. The inter-machine communication indicates that the target graphics processing unit and the other graphics processing units are located in different hosts. When the communication type includes intra-machine communication and inter-machine communication, the first communication speed and the second communication speed are determined as the initial communication speed.

8. The method according to claim 7, characterized in that, When the communication type includes intra-machine communication and inter-machine communication, the target communication speed includes a third communication speed corresponding to intra-machine communication and a fourth communication speed corresponding to inter-machine communication. Determining the sub-communication duration corresponding to the sub-module based on the data size of the processing result and the target communication speed includes: Based on the data size of the processing result and the third communication speed, a first candidate duration is determined; The second candidate duration is determined based on the data size of the processing result and the fourth communication speed; The maximum candidate duration between the first candidate duration and the second candidate duration is determined as the sub-communication duration.

9. The method according to any one of claims 2-8, characterized in that, Before determining the communication duration corresponding to the target graphics processing unit based on the communication parameters corresponding to the neural network layer, the method further includes: The communication parameters corresponding to the neural network layer are obtained from the data pre-stored in the configuration file, which is determined according to the training instructions.

10. The method according to any one of claims 1 to 8, characterized in that, The training instructions include a parallel strategy, the number of neural network layers in the model, the communication mode and weights corresponding to each neural network layer of the model, and before transmitting the training data determined according to the training instructions to the target graphics processing unit, the method further includes: Based on the parallel strategy, determine the total number of graphics processing units for parallel simulation training; The at least one neural network layer is determined based on the parallel strategy and the number of neural network layers; Based on the predetermined input data attributes of the first neural network layer in the at least one neural network layer, the first input data of the first neural network layer is generated; The total number, the first input data, the communication mode and weights corresponding to the at least one neural network layer are determined as the training data for the at least one neural network layer.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 10.

12. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 10.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10.