Data Parallelism in Distributed Training of Artificial Intelligence Models

Through the multi-level parallel parameter reduction method, the problem of inefficient training of deep learning models on memory-constrained devices is solved, efficient AI model training is achieved, and the rapid execution of large AI models on memory-constrained devices is supported.

CN114127740BActive Publication Date: 2025-07-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080051343.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-30
Filing Date
2020-06-09
Publication Date
2025-07-11
Estimated Expiration
2040-06-09

AI Technical Summary

Technical Problem

Deep learning models face the problems of memory capacity limitation and inefficient computing efficiency during training. Especially when using multiple GPUs, it is difficult to achieve efficient data parallelism and model parallelism, resulting in long training time and large resource consumption.

Method used

The multi-level parallel parameter reduction method is adopted to work together by the coordinated work of the parameter server and the target device, and the AI model is decomposed and executed on the target device with limited memory. The combination of micro batches and mini batches is used to optimize communication overhead and achieve efficient data parallelism and model parallelism.

Benefits of technology

Efficient AI model training is implemented on memory-constrained devices, reducing training time and resource consumption, improving computing efficiency, and supporting the rapid execution of large AI models on memory-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114127740B_ABST
    Figure CN114127740B_ABST
Patent Text Reader

Abstract

Methods, systems, apparatuses, and computer program products are described herein for enabling the execution of large AI models on memory-constrained target devices communicatively connected to a parameter server that stores a master copy of the AI model. The AI model can be decomposed into smaller parts (e.g., layers or sublayers), and each part can be executed as efficiently as possible on the target device. After the execution of a part of the AI model is completed, another part of the AI model can be downloaded and executed at the target device. To improve efficiency, input samples can be divided into micro-batches, and multiple micro-batches executed in sequence can form a mini-batch. The size of a set of micro-batches or mini-batches can be adjusted to reduce communication overhead. Multilevel parallel parameter reduction can be performed at the parameter server and the target device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Artificial intelligence has had a huge impact on many aspects of modern society. Machine learning (a subset of artificial intelligence that uses mathematical algorithms to process large datasets) has become increasingly popular in commercial applications and is increasingly appearing in consumer products. Deep learning is a branch of machine learning that is based on algorithms for modeling high-level abstractions in data. Many applications of artificial intelligence are driven by deep learning, such as natural language processing, speech recognition, and image analysis.

[0002] However, there are many challenges that hinder the widespread adoption of deep learning. These challenges include the complexity of managing large datasets and the significant amount of time and resources required to train deep learning networks. For example, a speech recognition program may require data from multiple dialects and demographics, which may include several terabytes of data for a single language. The complexity of a deep neural network (DNN) can be represented by the number of parameters, such that the more parameters there are, the more complex the DNN is. Additionally, optimizing hyperparameters (parameters defined before the learning process of an artificial intelligence (AI) model begins) can greatly affect the performance of the AI model. Furthermore, a large amount of computing power is required to process the large amount of data used to train such an AI model.

[0003] In deep learning, certain types of AI models may require the processing power of a GPU (graphics processing unit) with a high memory capacity. To increase throughput, multiple GPUs can be run in data parallelism, which typically requires synchronizing hundreds of millions to billions of parameters stored separately in different GPUs. This method may be limited by the memory capacity of the GPUs and may not achieve the maximum computational efficiency of the GPUs. Summary of the Invention

[0004] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0005] Methods, systems, devices, and computer program products are described herein for enabling the execution of arbitrarily large AI models on memory-constrained target devices communicatively connected to a parameter server. In an example embodiment, multi-level parallel parameter reduction can be performed by both the target device and the parameter server.

[0006] Further described herein are methods, systems, apparatuses, and computer program products that include a parameter server communicatively coupled to a target device. The parameter server includes: a data manager configured to store a master copy of an AI model; a transmitter configured to transmit a portion of the AI model to the target device; a batch manager configured to determine a micro-batch size suitable for the target device; and simultaneously, execute a collection of micro-batches of a training dataset on a first sub-portion of the transmitted portion of the AI model at the target device to generate gradients, a weight updater configured to perform parameter reduction on a second sub-portion of the transmitted portion of the AI model, and the transmitter configured to transmit weights of a third sub-portion of the transmitted portion of the AI model to the target device.

[0007] Other features and advantages, as well as the structure and operation of various examples, are described in detail below with reference to the drawings. Note that the concepts and techniques are not limited to the specific examples described herein. Such examples are presented herein for illustrative purposes only. Additional examples will be apparent to those skilled in the relevant art based on the teachings contained herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The drawings incorporated herein and forming a part of the specification illustrate embodiments of the present application and, together with the description, further serve to explain the principles of the embodiments and enable those skilled in the relevant art to make and use the embodiments.

[0009] Figure 1 is a block diagram of a system enabling execution of an arbitrarily large AI model on a memory-constrained target device according to an example embodiment.

[0010] Figure 2 illustrates a flowchart of a process for providing to run an AI model on a memory-constrained device during a forward pass according to an example embodiment.

[0011] Figure 3 illustrates a flowchart of a process for providing to run an AI model on a memory-constrained device during a backward pass according to an example embodiment.

[0012] Figure 4 illustrates a table representing a forward pass through a machine learning model according to an example embodiment.

[0013] Figure 5 illustrates a table representing a backward pass through a machine learning model according to an example embodiment.

[0014] Figure 6 illustrates a flowchart of a process for providing to run an AI model on a memory-constrained device at a parameter server according to an example embodiment.

[0015] Figure 7Shows a flowchart of a process for generating activations during a forward pass at a parameter server according to an example embodiment.

[0016] Figure 8 Shows a flowchart of a process for updating an AI model at a parameter server according to an example embodiment.

[0017] Figure 9 Shows a block diagram of parameter reduction with multi - level parallelism in a system according to an example embodiment.

[0018] Figure 10 Shows a timing diagram of parameter reduction with multi - level parallelism according to an example embodiment.

[0019] Figure 11 Shows a flowchart of a process for providing parallel parameter reduction in a system according to an example embodiment.

[0020] Figure 12 Shows a flowchart of a process for providing mixed - precision training of an AI model according to an example embodiment.

[0021] Figure 13 Shows a flowchart of a process for providing training of an AI model using multiple target devices according to an example embodiment.

[0022] Figure 14 Shows a flowchart of a process for providing dynamic execution for AI modeling according to an example embodiment.

[0023] Figure 15 Shows a flowchart of a process for providing determination of computational precision for dynamic execution for AI modeling according to an example embodiment.

[0024] Figure 16 Shows a flowchart of a process for providing determination of whether to stop or continue execution of an AI model based on the accuracy of the AI model according to an example embodiment.

[0025] Figure 17 Is a block diagram of an example computer system in which embodiments can be implemented.

[0026] The features and advantages of the embodiments will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which like reference numerals throughout the drawings identify corresponding elements. In the drawings, like reference numerals generally denote identical, functionally similar, and / or structurally similar elements. The drawing in which an element first appears is indicated by the left - most (one or more) digits in the corresponding reference numeral. Detailed Description

[0027] I. Introduction

[0028] The following specific embodiments disclose many examples. The scope of this patent application is not limited to the disclosed embodiments, but also includes combinations of the disclosed embodiments and modifications to the disclosed embodiments.

[0029] References to "an embodiment", "embodiment", "exemplary embodiment", etc. in the specification indicate that the described embodiment may include a specific feature, structure, or characteristic, but each embodiment may not necessarily include the specific feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. In addition, when a feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not).

[0030] In the discussion, unless otherwise stated, adjectives such as "substantially", "approximately", and "about" that modify a condition or relationship characteristic of one or more features of the embodiments of the present disclosure are understood to mean that the condition or characteristic is defined to be within an acceptable tolerance for the operation of the embodiments for their intended applications.

[0031] Many exemplary embodiments are described as follows. Note that any section / subsection headings provided herein are not intended to be restrictive. Embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. In addition, the embodiments disclosed in any section / subsection may be combined with any other embodiments described in the same section / subsection and / or different section / subsections in any manner.

[0032] II. Exemplary Embodiments

[0033] The exemplary embodiments described herein are provided for illustrative purposes and are not restrictive. The examples described herein may be applicable to any type of target crawling system. Through the teachings herein, additional structural and operational embodiments (including modifications / changes) will become apparent to those skilled in the relevant art.

[0034] Deep learning has many applications, including natural language processing, speech recognition, image analysis, machine translation, object classification and detection in photos, automatic handwriting generation, automatic gaming, generative model chatbots. Deep learning models are widely used in various tasks because of their ability to simulate the human brain.

[0035] In deep learning, large AI models (e.g., trained for natural language processing or image analysis) may require multiple GPUs with high memory capacity to perform their training. To increase speed, these GPUs can have high-speed interfaces, such as high-bandwidth memory (HBM) interfaces. However, even with high-quality hardware, there are ways to improve the inference and training processes of large AI models. For example, there are two ways to parallelize the training of an AI model to increase throughput: model parallelism and data parallelism.

[0036] Model parallelism involves dividing a learning model into multiple parts and placing these parts on different computing nodes (e.g., placing the first half of a layer on a first GPU and the second half of the layer on a second GPU, or splitting intermediate layers and distributing them to separate GPUs). For example, a typical large AI model with 24 layers can be run on GPUs in the following way. The forward pass is performed layer by layer on the same mini-batch, such as by starting with the mini-batch on layer 1, then layer 2, and so on until layer 24. After each layer, the activations of that layer (also referred to in this article as hidden activations, hidden states, or intermediate results) can be saved (e.g., on-chip or off-chip) for the backward pass, which can be performed on the same mini-batch in a similar layer-by-layer manner (in reverse order). For example, the mini-batch can be performed on layer 24, then layer 23, and so on until layer 1, after which the AI model is updated. Sometimes, as a trade-off between computational cost and efficient memory usage, the hidden activations can be recomputed during the backward pass. In some types of AI models (e.g., natural language processing), there may be a large number of parameters, but the mini-batch size may be small (e.g., a few kilobytes). In other types of models (such as dense networks or computer vision models), the number of parameters may be relatively small, but the number of hidden activations may be large. Generally, these types of models may not be able to run on devices without global memory, such as application-specific integrated circuit (ASIC) devices. Therefore, the available technique used is model parallelism, where the model is split across multiple devices. However, model parallelism is inefficient due to long dormant memory and computational times.

[0037] In addition, a GPU may have certain data structures mapped to its global memory, which is off-chip and connected to a high-speed memory interface (e.g., HBM). For example, input and output activations may reside on-chip, and sometimes gradients as well, while the main copies of weights and hidden activations may be stored off-chip. There are several residency issues with these data structures. Weights may be loaded before actual use, thus occupying precious memory. After a forward pass is completed, hidden activations may be generated, but may not be needed until a backward pass. In addition, global memory data is moved in and out of the chip via loads and stores, resulting in memory access amplification even with limited-time buffering through caches and registers. Therefore, when running large AI models (e.g., using model parallelism or unified memory addressing techniques) in such a GPU or GPU cluster, the AI model size depends on the number of devices, and the performance penalty due to communication overhead cannot be mitigated or hidden.

[0038] Data parallelism is where the input data is split across computing devices, and each device holds a complete copy of the learning model, referred to as replicas or workers. Each replica computes the gradients for its portion of the data, and these gradients are combined to update the model parameters. In asynchronous distributed stochastic gradient descent (SGD), each replica accesses a shared memory space where the global parameters are stored. After replicating the parameters in its local memory, the replica can compute the gradients and updated weights with respect to its current weights, and then apply the updated weights to the global parameters in the shared memory space. The advantage of this configuration is that replicas can work at their own pace without waiting for other replicas to finish computing their gradients. However, there is no way to ensure that when one replica is computing the gradients for a set of parameters, another replica is not updating the global parameters, resulting in the global parameters being updated with stale gradients. In synchronous distributed SGD, each GPU may run a mini-batch of the input data (or, samples), and then stop execution to synchronize all model parameters by exchanging gradients, which are computed as adjustments through backpropagation of the loss through the AI model. This method is highly limited by the GPU memory capacity. In cases where the AI model requires more memory than a single GPU, model compilation on that GPU may fail due to out-of-memory errors. Data parallelism typically requires synchronizing hundreds of millions to billions of parameters stored separately in different GPUs. Therefore, this method may not achieve the maximum computational efficiency of GPUs as they need to pause for a long time during computation to complete synchronization.

[0039] The embodiments described herein overcome such difficulties to support the operation of AI models on devices with large on-chip memory but without global memory. The embodiments described herein can execute any arbitrarily sized AI model in a fast and efficient manner in memory-constrained devices such as GPUs, ASICs, or FPGAs (field programmable gate arrays). In an example embodiment, any arbitrarily sized AI model can be executed on an ASIC that has no global memory but can execute the AI model faster than a GPU. Thus, the embodiments described herein support the execution of large AI models on memory-constrained devices.

[0040] An example embodiment can be implemented in a system having at least one parameter server and one target device. A master copy of the AI model can reside in the parameter server. The AI model can be decomposed into smaller parts or chunks (e.g., individual layers), and each part or layer can be executed as efficiently as possible on the target device. After one layer is completed, the next layer is executed. To improve balance and efficiency, the technique iterates over the same layer for a large number of input samples until either (a) the next layer is loaded onto the target device, thus fully hiding its latency, or (b) the next layer is loaded after the current layer is completed, exposing its latency, but minimizing the overhead through the long computation cycle of the current layer. To make the current computation cycle longer, the input samples can be partitioned into micro-batches. A set of micro-batches forms a mini-batch, which is a term for the number of samples used for each update (for training) or provided for each inference cycle (for inference). By using the size of a set of micro-batches and / or mini-batches as a knob that can be adjusted manually or automatically (e.g., using software of an AI framework), the communication overhead can be minimized or even reduced to zero.

[0041] If an AI model can be optimized with a larger batch size, such as in the case of natural language processing models, vision models, or models with a high weight / activation ratio, the embodiments described herein will allow these models to run on one or more memory-constrained devices with peak performance. Thus, according to an example embodiment, a large AI model can be executed on a target device whose memory is less than the memory required to efficiently run the large AI model. In other words, the AI model can be executed with a minimum device batch size, at which point peak efficiency of speed (i.e., efficient TFLOP) can be achieved. For example, the performance of an AI model depends only on the efficiency of the computational throughput of the library running on the target device, i.e., TFLOP (teraFLOPS). Floating-point operations per second (FLOPS) is a measure of computer performance, e.g., for measuring the ability of an algorithm or computer hardware to compute one trillion floating-point operations per second. In other example embodiments, a large AI model can be executed on multiple target devices whose combined memory (e.g., global memory) may be less than the memory required to efficiently run the large AI model.

[0042] A. Executing a Large Artificial Intelligence Model on a Memory-Constrained Target Device

[0043] Supporting the execution of a large AI model on a memory-constrained device can be achieved in a variety of ways. For example, Figure 1 is a block diagram of a system that supports the execution of any large AI model on a memory-constrained target device according to an example embodiment. As Figure 1 shown, system 100 includes a parameter server 102 and target devices 134a - 134k. Although Figure 1 only one parameter server 102 is shown, system 100 can include multiple parameter servers. Similarly, although Figure 1 target devices 134a - 134k are depicted, system 100 can include fewer or more target devices. Based on the following discussion of system 100 as depicted in Figure 1 the other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0044] The parameter server 102 can include any type of computing device (mobile or fixed). The parameter server 102 can provide functionality to other programs or devices, such as sharing data or resources or performing computations. The parameter server 102 can include a memory 104 configured to store data (e.g., data sets, software programs, AI models) and a processor 132 configured to perform programming functions. The parameter server 102 can include off-the-shelf components and / or custom components and can be a stand-alone device or part of another computing device. The parameter server 102 can include Figure 1Other components not shown, such as peripheral interfaces, communication interfaces, integrated devices, multi-processors, and different types of memories. In an embodiment, the parameter server 102 may be implemented as one or more of a CPU, FPGA, or ASIC. For example, as a CPU, the parameter server 102 may include electronic circuitry within a computing device that executes instructions of a computer program by performing operations based on the instructions (e.g., mathematical, logical, control, or input / output). As an FPGA, the parameter server 102 may include an array of programmable logic blocks configured to perform complex combinatorial functions or other operations. As an ASIC or system-on-chip, the parameter server 102 may include a custom integrated circuit configured to perform operations based on computer program instructions.

[0045] The parameter server 102 may be configured to store the AI model 106 in the memory 104. The AI model 106 may include weights 108, and during the execution of the AI model 106, activations 112 and gradients 110 may be stored in the memory 104. The parameter server 102 may also store a dataset 114, which may be a training or test dataset. The parameter server 102 may also include computer program logic (e.g., computer program code or instructions) for performing operations. For example, the parameter server 102 may include an AI model manager 116 configured to manage the AI model 106 during inference or training of the AI model 106. The AI model manager 116 includes computer program logic for managing the AI model 106, such as a data manager 118, a batch manager 120, a transmitter 122, and an output data manager 124. The output data manager 124 is configured to receive and manage output data from the target devices 134a - 134k and other data for managing the AI model 106. The output data manager 124 includes a weight updater 126 configured to update the weights 108 of the AI model 106, a precision formatter 128 configured to manage precision (e.g., mixed precision training, precision conversion, etc.) formats, and a model evaluator 132 configured to evaluate the AI model 106 and manage the execution of the AI model 106 accordingly. In an example embodiment, the AI model manager 116 may include fewer or more components than Figure 1 shown. In other embodiments, the functions of the components of the AI model manager 116 may overlap.

[0046] The parameter server 102 can serve one or more target devices 134a - 134k. The parameter server 102 can be communicatively connected to the target devices 134a - 134k via a suitable interface (such as, Peripheral Component Interconnect (PCI) or PCI express (PCIe)) and / or a network (e.g., for cloud computing or edge computing). In an example embodiment, the parameter server 102 and one or more target devices 134a - 134k can reside on the same chip or can reside on different chips or different devices. In an example embodiment, the parameter server 102 and the target devices 134a - 134k can include software for communicating with each other. For example, the parameter server 102 can include driver software specifically designed for communication, such as sending commands (e.g., initiating a function call to the target devices 134a - 134k) and receiving responses.

[0047] Each of the target devices 134a - 134k can include Figure 1 an instance of the features shown for target device 134a in. Target device 134a includes a data interface 136, a processor 140, and a memory 142. Each of the target devices 134a - 134k can include multiple processors and different types of interfaces and memories, even Figure 1A single data interface is depicted for target device 134a. In an example embodiment, target devices 134a - 134k have the same hardware specifications, such as the same memory size. In other example embodiments, target devices 134a - 134k may have different hardware specifications. Target devices 134a - 134k can be specifically designed to perform compute - intensive operations in an accelerated manner. For example, the target device can include a high - compute - density, fixed - function processor (e.g., processor 140), as well as other general - purpose capabilities. Target devices 134a - 134k can be managed by parameter server 102 to run specific computations. For example, parameter server 102 can execute a main program that prepares input data for processing at target devices 134a - 134k, invokes parallel routines (e.g., kernels) at target devices 134a - 134k, and receives the results after the routines terminate. Parameter server 102 can also utilize a high - level computing language, computing platform, or framework that includes an integrated library of acceleration algorithms and data structures to make it easier to accelerate computations on target devices 134a - 134k. In an example embodiment, target devices 134a - 134k can be implemented as GPUs (e.g., dedicated or general - purpose GPUs), ASICs, FPGAs, or edge devices (e.g., stand - alone devices or microcontrollers residing at the end or edge of a network connection and possibly having a small memory footprint). In an embodiment, target devices 134a - 134k can be memory - constrained devices, although not necessarily in all cases. For example, target devices 134a - 134k, individually or in one or more groups, may be memory - constrained such that they are unable to efficiently run large AI models due to insufficient memory. For illustrative purposes, Figure 1 the characteristics of target device 134a are described as follows as representative of each of the target devices 134a - 134k.

[0048] The data interface 136 can be configured to interface the target device 134a with the parameter server 102 and other devices including other target devices. For example, the data interface 136 can include PCI, PCIe, and / or HBM. The processor 140 is configured to execute operations requested by the parameter server 102 and operations specific to the target device 134a. The memory 142 is configured to store data and computer program logic. For example, the memory 142 includes an accelerator 144, which is configured to execute functions and / or accelerate certain operations, such as as indicated by the parameter server 102. The accelerator 144 includes a data downloader 146 configured to download data (e.g., a model and / or its data, such as weights, activations, and data sets), a data manager 148 configured to store or otherwise manage the downloaded data, a layer executor 150 configured to execute an AI model or a portion thereof (i.e., execute a data set on the AI model or a portion thereof), and an output manager 152 configured to manage output data (e.g., gradients and activations) generated according to the model execution, such as by saving, sending, or restoring the output data. In an example embodiment, the accelerator 144 can include fewer or more components than Figure 1 shown. In other embodiments, the functionality of the components of the accelerator 144 can overlap.

[0049] In conjunction with Figure 2 the following describes additional operational aspects of the parameter server 102 and the target device 134a, Figure 2 FIG. 200 is a flow chart showing a process for running an AI model on a memory-constrained device during a forward pass according to an example embodiment. Although described with reference to Figure 1 the system 100 of Figure 2 the process is not limited to that system. Based on Figure 2 the flow chart 200 of Figure 1 and the following discussion of the system 100 of

[0050] The flow chart 200 begins at step 202. At step 202, a portion of an artificial intelligence (AI) model is downloaded from a parameter server to the memory of a target device, where the parameter server stores the master copy of the AI model. For example, Figure 1The target device 134a, specifically the data downloader 146, can be configured to download a portion of the AI model 106 from the parameter server 102 to the memory 142 of the target device 134a. In an example embodiment, one or more target devices 134a - 134k can form a group that is configured to run an instance of the AI model. For example, one group can include the target device 134a, while another group can include the target devices 134b - 134k. The memory required for the AI model has both invariant and variable requirements. For example, the size of the weights or parameters may be invariant for a particular precision type (e.g., 32 bits), and the variable requirements may depend on the target device batch size. In an example embodiment, the target device 134a (or, a group of target devices 134a - 134k) can be configured to run a large AI model, and the memory of the target device 134a (or, the combined memory or global memory of the group of target devices 134a - 134k) may be smaller than the memory required to efficiently run the AI model. Efficiently running the AI means that the AI model runs with the minimum device batch size to achieve peak efficiency in speed (i.e., effective TLOP). In other words, the AI model does not benefit from a larger device batch size (no more efficient TFLOP). For example, the target device 134a can have a memory with a size smaller than the overall size of the AI model's specific optimal batch size. The optimal batch size of a target device is the batch size that achieves the desired accuracy (e.g., 85%) with the highest possible throughput (e.g., the rate of processing a data set). The optimal group batch size is the global batch size (e.g., the batch size of a group of target devices) divided by the number of groups of target devices communicatively connected to the parameter server 102.

[0051] The AI model 106 can include any type of machine learning model that can have various applications in many fields, such as natural language processing, autonomous vehicles, image processing, deep learning robots, automatic machine translation, and automatic handwriting generation. The AI model 106 can have any type of deep learning architecture, such as deep neural networks, recurrent neural networks, and convolutional neural networks.

[0052] A simple neural network can include several layers, one for receiving input signals and another for sending output signals. One or more hidden or processing layers can be located between the input layer and the output layer. In a DNN constructed to generate one or more inferences, there may be many hidden layers composed of artificial neurons. Such neurons can include activation functions, constant inputs, other inputs, and outputs. The neuron can produce an output by performing an activation function on a weighted version of the inputs. The inputs to the activation function are weighted according to their corresponding weights. For example, the inputs can include normalized data. The activation function can be configured to accept a single number (e.g., a linear combination of weighted inputs) based on all inputs and perform a fixed operation, such as sigmoid, tanh, or rectified linear unit options. The constant input can be a constant value.

[0053] A single neuron by itself may not accomplish much work, and useful AI models typically involve the combined computational work of a large number of neurons working together. For example, a DNN can include multiple neurons assembled in layers and connected in a cascaded manner. These layers can include an input layer, an output layer, and some hidden layers in between. The output of each layer of neurons can be weighted according to certain weights and then used as the input to the neurons in the next layer. Other interconnection strategies known in the art can be employed. The neurons in the input layer can be configured to accept normalized or otherwise feature-engineered or processed data corresponding to user data. The output of each neuron in the input layer or hidden layer can be weighted according to the weights of its corresponding output edges and then applied as an input at each neuron in the next layer. The output(s) of the output layer include the output of the DNN or AI model. In an inference context, such output can be the inference(s) or prediction(s). Constructing such a DNN is just the beginning of generating a useful machine learning or AI model. The accuracy of the inferences generated by such an AI model requires the selection of appropriate activation functions, and then, each weight of the entire model is adjusted to provide an accurate output. The process of adjusting such weights is called "training". Training a DNN or other type of network requires a set of training data with known characteristics. For example, if the DNN is designed to predict the probability that an input image of an animal is a cat, the training data will include many different images of cats, and usually not only images of cats, but also images of other similar animals. Training requires preprocessing the image data corresponding to each image according to normalization and / or feature extraction techniques known in the art to produce input features for the DNN, and then, these features are provided as inputs to the network, e.g., as inputs to the neurons in the input layer.

[0054] Afterwards, each neuron of a layer performs its corresponding activation operation, and the output of the activation operation is weighted and fed forward to the next layer in the forward pass until the output(s) of the DNN are generated by the output layer. The output(s) of the DNN can be compared with the known or expected value of the output, and the difference can be fed back in reverse in the backward pass through the DNN to adjust the weights contained therein according to the backpropagation algorithm known in the art. With the AI model including the updated weights, the image features can be input into the model again and new outputs are generated. Training includes iterating the AI model over a training dataset and updating the weights at each iteration. Once the AI model reaches sufficient accuracy, or its output has converged and the weight changes have little effect, the AI model is said to have been trained. Then, the trained model can be used to evaluate any input data, the nature of which is unknown beforehand and has not been considered by the model before (e.g., a new picture of an animal), and output the desired inference (e.g., the probability that the image is an image of a cat).

[0055] Gradient descent is an algorithm often used in training AI models. Gradient descent involves an objective function (e.g., a loss function or a cost function) (there may be many of them), and the goal is to minimize this function. The objective function is used to monitor the error in the predictions of the AI model. Thus, by minimizing this function, the lowest error value can be found, thereby improving the accuracy of the AI model. Stochastic gradient descent (SGD) is a variant of the gradient descent algorithm that calculates the error and updates the model for each sample in the training dataset. SGD updates frequently and has a faster learning rate, but has a high computational cost and may take longer to train on large datasets. Batch SGD is another variant that calculates the error for each sample of the training dataset but updates the AI model only after the entire dataset has been processed (i.e., at the end of the training phase). Batch SGD updates less frequently and is more computationally efficient than SGD. The separation of the prediction error calculation and model update in batch SGD makes this algorithm suitable for parallel processing-based implementations, but the update at the end of the training epoch requires additional complexity to accumulate the prediction errors of the entire dataset and is typically implemented in a way that requires the entire training dataset in memory and is available to the algorithm. Mini-batch SGD is yet another variant of SGD that splits the training dataset into small batches, which are used to calculate the model error and update the parameters. The implementation can sum the gradients over the mini-batches, thereby further reducing the variance of the gradients. Thus, mini-batch SGD strikes a balance between SGD and batch SGD. Mini-batch SGD requires an additional "mini-batch size" hyperparameter to be configured for the learning algorithm. Error information can accumulate in the training examples of the mini-batch. The mini-batch size can be configured as a function of the computational architecture on which the AI model is being executed, e.g., a power of 2 that fits the memory requirements of the target device or accelerator hardware, such as 32, 64, 128, 256, etc. The batch size can serve as a tuning of the learning process, where smaller values give a learning process that converges quickly at the expense of noise in the training process, while larger values give a learning process that converges slowly with an accurate estimate of the error gradient.

[0056] Return reference Figure 2In step 202, the downloaded portion of AI model 106 can include any part of AI model 106. AI model 106 can be broken down into layers or composite layers or composite fractional layers (e.g., broken down into 1.5x layers). The AI model is neatly segmented at the layers, so these models can be divided as a whole. In one example embodiment, the downloaded portion can include one or more layers of AI model 106. However, in another example embodiment, depending on the AI model and other factors, fractional division is possible. One reason fractional division may be needed is that such fractional parts may fit inside the target device while the whole layer may not. This will support an implementation where any number of layers can be run without a memory shortage problem, rather than hitting a memory shortage error after a specific number of layers. In an example embodiment, the portion of AI model 106 downloaded to target device 134a can include any part of AI model 106 up to the entire AI model 106. Target device 134a can download the portion of AI model 106 in various ways. For example, target device 134a can download the next portion of AI model 106 into one or more memory buffers while executing the current portion of AI model 106. This method can use slightly more memory and special libraries, but can result in a higher performance AI model 106. In another example, target device 134a can execute the current sub-portion, synchronize, and then download the next sub-portion. This method may be a bit slower, but does not require buffering. Flowchart 200 continues to step 204.

[0057] In step 204, a micro-batch collection of the dataset is stored in the memory of the target device. For example, as Figure 1As shown, a collection of micro-batches of the dataset 114 can be downloaded from the parameter server 102 via the data downloader 146. Then, the data manager 148 can store the collection of micro-batches in the memory 142, buffer, or any other known memory structure of the target device 134a. The dataset 114 can be user input data, such as a training dataset for training, a test dataset for testing purposes, or any input data for inference. The collection of micro-batches includes a plurality of micro-batches that are configured to be executed sequentially at the target device 134a. The collection of micro-batches forms a mini-batch that includes a number of samples for training the AI model 106 for each update or a number of samples for inference provided in each inference cycle. Each micro-batch in the collection of micro-batches can have a micro-batch size that can be configured automatically or manually. In an embodiment, the micro-batch size can be selected based on the execution rate of the plurality of micro-batches and the communication rate between the target device 134a and the parameter server 102. For example, the micro-batch size can be initially selected for the target device 134a based on its hardware specifications, and then the micro-batch size can be adjusted during the iteration process as needed to sufficiently hide the communication latency. In an embodiment, the optimal micro-batch size can be a trade-off between the required memory and the percentage of communication overhead that can be hidden. Since more communication overhead is hidden, more memory may be required for the computation. Therefore, the micro-batch size may be large enough to fully utilize the execution of the layers in the target device, but small enough to fit into the memory of the target device.

[0058] The flowchart 200 proceeds to step 206, which performs a micro-batch aggregation on a first sub-part of the download portion of the AI model to generate activations. For example, the layer executor 150 can perform the micro-batch aggregation on the first sub-part of the portion of the AI model 106 downloaded by the data downloader 146 at the target device 134a. In an example embodiment in which the download portion of the AI model 106 includes one or more layers, the micro-batch aggregation can be performed on one or more download layers of the AI model 106 one layer at a time to generate activations. An activation can be a value that is an intermediate result, e.g., the output of each micro-batch execution. An activation can be internal data needed in the backpropagation to determine how the weights 108 of the AI model 106 should be adjusted. After each micro-batch is executed for a sub-part (e.g., layer) of the AI model 106, the activations can be saved on the target device 134a, sent to the parameter server 102 to save memory, or discarded to save memory and then recomputed. For example, if the AI model 106 has 12 layers and each mini-batch has 8 micro-batches, the activations can be stored 96 times during the forward pass and restored 96 times during the backpropagation. If not all activations are saved during the forward pass, the activations can be recomputed during the backpropagation. In an example embodiment, the storage of the activations of the micro-batches (whether at the target device 134a or at the parameter 102) can occur while the target device 134a is executing different micro-batches. In an example embodiment, the restoration of the activations or the recomputation of the activations can occur before the execution of the sub-part or during the execution of the sub-part as needed, e.g., the restoration / recomputation of the activations for the next micro-batch can occur in parallel with the execution of the current micro-batch.

[0059] The flowchart 200 ends at step 208. In step 208, the weights of the second sub - part of the download portion of the AI model are downloaded from the parameter server to the memory of the target device. For example, if the download portion of the AI model includes multiple layers, the weights of the second layer can be downloaded to the memory 142 of the target device 134a via the data downloader 146. In an example embodiment, the download of the weights of the next layer can occur while the current layer is being executed. For example, the target device 134a can be configured to simultaneously execute a micro - batch set of the dataset on the second sub - part using the downloaded weights of the second sub - part, and download the weights of the third sub - part of the download portion of the AI model 106 from the parameter server 102 to the memory 142 of the target device 134a. For example, the layer executor 150 can execute a micro - batch set on a layer using the weights that have been downloaded for that layer, while the data downloader 146 downloads the weights of the next layer of the AI model 106. Alternatively, the target device 134a can be configured to serially execute a micro - batch set on the second sub - part using the downloaded weights of the second sub - part, and download the weights of the third sub - part of the download portion of the AI model 106 from the parameter server 102 to the memory 142 of the target device 134a. For example, the layer executor 150 can execute a micro - batch set on a layer using the weights that have been downloaded for that layer, and after the execution of that layer, the data downloader 146 can download the weights of the next layer of the AI model 106.

[0060] Thus, the execution of the AI model 106 continues at the target device 134a one sub - part at a time as described above, while other sub - parts of the AI model 106 can also be executed at other target devices. For example, in the forward pass, a set of micro - batches or mini - batches are executed on the first layer, then the second layer, and so on until the last layer.

[0061] Once the forward pass of the AI model 106 is complete, the backward pass can be executed. For example, Figure 3 FIG. 300 is a flowchart showing a process for running an AI model on a memory - constrained device during a backward pass according to an example embodiment. Although described with reference to Figure 1 system 100, the Figure 3 process is not limited to this system. Based on Figure 3 the following discussion of flowchart 300 and Figure 1 system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0062] Flowchart 300 begins at step 302 where a microbatch aggregation is performed on the third sub-part of the download section of the AI model to generate gradients. For example, the microbatch aggregation can be performed by layer executor 150 on the third sub-part of the download section of AI model 106 to generate gradients for the third sub-part. If AI model 106 has 24 layers, the microbatch aggregation can be performed on layer 24 to generate gradients for that layer to initiate backpropagation.

[0063] Flowchart 300 continues to step 304. In step 304, the weights and activations of the fourth sub-part of the download section of the AI model are downloaded. For example, the weights and activations from the fourth sub-part of the download section of AI model 106 can be downloaded from parameter server 102 to target device 134a by data downloader 146. For example, if AI model 106 has 24 layers, the weights and activations from layer 23 can be downloaded from parameter server 102 to target device 134a.

[0064] In step 306, simultaneously, a microbatch aggregation is performed on the fourth sub-part using the downloaded weights and output activations, the weights and output activations of the fifth sub-part of the download section of the AI model are downloaded from the parameter server, and the gradients for the third sub-part are sent to the parameter. For example, simultaneously or substantially simultaneously in a parallel manner, layer executor 150 can perform a microbatch aggregation on the fourth sub-part using the downloaded weights and output activations for that sub-part, data downloader 146 can download the weights and output activations of the fifth sub-part of the download section of AI model 106 from parameter server 102, and output manager 152 can send the gradients for the third sub-part of AI model 106 to parameter server 102. In an example embodiment where model 106 has 24 layers, target device 134a can be configured to perform multiple steps in parallel or simultaneously. In this embodiment, target device 134a can be configured to simultaneously perform layer 23 using the downloaded weights and output activations of layer 23, download the weights and output activations of layer 22 from parameter 102, and send the gradients 110 generated for layer 24 to parameter server 102.

[0065] Target device 134a is configured to continue the above steps of flowchart 300 to complete the execution of the entire dataset 114 in microbatches on AI model 106 one sub-part (e.g., layer) at a time in reverse order (i.e., layer 24, layer 23,..., and layer 1) for backpropagation.

[0066] As Figure 2 and Figure 3 described, the forward and backward passes can be visualized as Figure 4 and Figure 5 depicted. For example, Figure 4Table 400 shows a forward pass through a machine learning model with 24 layers according to an example embodiment. Table 400 is for a target device, which can be implemented as Figure 1 target device 134a. Table 400 has three rows. Row 410 shows the execution of a set of micro-batches sequentially on each layer of the AI model, 10 of which form a mini-batch here. Row 412 shows a set of actions of the target device (e.g., receiving weights from a parameter server), and row 414 shows another set of actions that the target device can take (e.g., sending activations to the parameter server). Data exchange at the target device can be done via an interface, such as Figure 4 PCI as shown in. Although the AI model has 24 layers, only layer 1, layer 2, and layer 24 are shown in detail in Table 400 because the execution of the AI model on each layer is similar. For example, column 402 of Table 400 depicts the execution of a set of 10 micro-batches on layer 1, and the 10 micro-batches form a mini-batch. During this execution, the target device receives the weights of layer 2 (the next layer to be executed). Since each micro-batch is executed on layer 1, the activation of the micro-batch can be saved (e.g., at the target device or the parameter server) if memory and / or other resources permit. Then, as shown in column 406 of Table 400, the same set of 10 micro-batches is executed on layer 2 while receiving the weights of layer 3 and saving the activation of each micro-batch. This process continues for all layers of the AI model until the last layer, i.e., layer 24, which can be referred to as the "decoding layer" (DL) or "embedding layer" or "output layer". When the set of micro-batches is executed on the last layer (layer 24), its weights and activations are determined at the target device and sent to the parameter server, as shown in column 408 of Table 400.

[0067] Figure 5 Table 500 shows a backward pass through a machine learning model with 24 layers according to an example embodiment. Table 500 is for a target device, which can be implemented as Figure 1 target device 134a. Table 500 has four rows. Row 510 shows the execution of a set of micro-batches sequentially on each layer of the AI model, 16 of which form a mini-batch here. Row 512 shows a set of actions of the target device (e.g., loading weights and activations from a parameter server), row 514 shows another set of actions (e.g., sending gradients to the parameter server), and row 516 shows another set of actions that the target device can take (e.g., parameter reduction). Data exchange at the target device can be done via an interface, such as Figure 5The PCI shown. Although the AI model has 24 layers, only the 24th, 23rd, 22nd, and 1st layers are shown in detail in Table 500 because the execution of the AI model on each layer is similar. For example, column 502 of Table 500 depicts the execution of a set of 16 micro-batches on layer 24, and the 16 micro-batches form a mini-batch. During this execution, the target device loads the weights and activations of layer 23 (the next layer to be executed in the backward pass). Then, in column 504, the set of 16 micro-batches is executed on layer 23 using the loaded weights and activations. In parallel or simultaneously (or, substantially simultaneously), the target device is configured to load the weights and activations of layer 22 (the next layer to be executed); send the gradients of the most recently executed layer 24 to the parameter server; and reduce the parameters of the AI model. In column 506, the same set of 16 micro-batches is executed on layer 22 using the loaded weights and activations of that layer. During the execution of layer 22, the weights and activations of layer 21 are loaded, the gradients of layer 23 are sent to the parameter server, and the parameters are reduced at the target device. In column 508, the same set of 16 micro-batches is executed for layer 1. Simultaneously with this execution, the gradients of layer 2 are sent to the parameter server, and the parameters are reduced at the target device.

[0068] By running many micro-batches on the same layer, there is enough time to hide or cover the latency of preparing the next layer. Thus, the total memory complexity of the target device can be two layers plus one layer of hidden activations and one layer of output activations.

[0069] In the above description, for example, in combination with Figures 2 - 5 the target device serves as a support component for executing a large AI model on a memory-constrained device. In the following description, in combination with Figures 6 - 8 the parameter server can serve as a support component for executing a large AI model on a memory-constrained device. For example, Figure 6 shows a flowchart 600 of a process for running an AI model on a memory-constrained device at a parameter server according to an example embodiment. Although described with reference to Figure 1 system 100, the Figure 6 process is not limited to this system. Based on the following discussion of Figure 6 flowchart 600 and Figure 1 system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0070] Flowchart 600 begins at step 602. In step 602, a main copy of the artificial intelligence model is stored at the parameter server, which is communicatively connected to the target device. For example, as Figure 1As shown, the main copy of the AI model 106 can be stored by the data manager 118 in the memory 104 of the parameter server 102. The parameter server 102 can communicate with the target devices 134a - 134k via suitable means (such as, PCI and PCIe interfaces or other network interfaces). In an example embodiment, the parameter server 102 stores a complete copy of the AI model 106, while the target devices 134a - 134k can store a portion of the AI model 106 rather than the entire copy of the AI model 106.

[0071] In step 604, a micro - batch size suitable for the target device is determined. For example, as Figure 1 shown, the target device 134a can be a memory - constrained device, the memory size of which is less than the entirety of the AI model 106 for a particular optimal batch size stored at the parameter server 102. The batch manager 120 is configured to determine the micro - batch size suitable for the target device 134a by, for example, considering the memory size of the target device 134a and / or other hardware specifications. In one embodiment, the batch manager 120 configures the micro - batch size for load balancing such that the ratio of execution time to communication time is maximized. For example, the micro - batch size can depend on the computation time of the target device (C), the size of the sub - part (S) to be transmitted, and the communication bandwidth (B) of the target device and the parameter server system. In this example, the micro - batch size can be determined as Minimum_numMicroBatches = S / B / C. This equation can be static, but in some cases (such as, neural architecture search), the micro - batch size can be determined dynamically. Thus, when the ratio of execution time and communication time can be manipulated, it allows the parameter server to have more time to perform complex data parallelism or background tasks. In an example embodiment, the micro - batch size can be dynamically configured at certain times or boundary points during the training or inference process. For example, the micro - batch size can be dynamically configured at the end of an iteration, but must be constant for mini - batch iterations in the forward and backward passes.

[0072] Return Figure 6, the flow chart 600 ends at step 606. In step 606, a portion of the AI model is transmitted to the target device. For example, the transmitter 122 can transmit a portion of the AI model 106 from the parameter server 102 to the target device 134a. The AI model 106 can be divided into different portions in any number of ways. For example, the portion can be a layer of the AI model 106, a combination of layers, or a combination of fractional layers. The portion size can be determined based on the memory available on the target device 134a such that the portion has an optimal size for the target device 134a. For example, the AI model manager 116 can consider the hardware specifications of the target device 134a when determining the size of the portion to be sent to the target device 134a. The transmitter 122 can transmit a portion of the AI model 106 to the target device 134a while the target device 134a is executing another portion, thereby requiring the target device 134a to buffer the portion. Alternatively, the transmitter 122 can transmit a portion of the AI model 106 to the target device 134a after the target device 134a has completed the current portion to avoid the need to buffer the portion. In this alternative example, the target device 134a can perform a synchronization after the execution of the current portion before receiving the portion.

[0073] The parameter server 102 or specifically the AI model manager 116 can perform additional steps to improve the throughput of distributed training and inference of the AI model on memory-constrained devices. For example, Figure 7 A flow chart 700 is shown that provides a process at the parameter server for generating activations during the forward pass according to an example embodiment. Although described with reference to Figure 1 the system 100, the Figure 7 process is not limited to this system. Based on Figure 7 the following discussion of the flow chart 700 and Figure 1 the system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0074] Figure 7Beginning at step 702, activations are received from the target devices after each micro-batch is executed. For example, after each target device executes a micro-batch, the output data manager 124 may receive activations from the target devices 134a - 134k. For example, the activations may include hidden activations, or intermediate results of executing the micro-batch, or the output of executing the micro-batch at the target device 134a. In an example embodiment, activations are received from the target device 134a after each micro-batch. In this embodiment, saving and / or storing the activations after each micro-batch may provide the best efficiency for executing the AI model 106. Thus, when the target device 134a is executing a mini-batch that includes multiple micro-batches, the activations for each of the multiple micro-batches may be saved after each of the multiple micro-batches is executed. In another example embodiment, the activations are saved at the target device 134a. In yet another example embodiment, not all activations are saved during the forward pass, only the data needed to recompute the activations during the backward pass. Such data may include the input state of a particular sub-part of the AI model, for example, because the input state may require less memory space than the output state. Thus, the memory space of the target device 134a may be saved by not saving all activations during the forward pass.

[0075] Flowchart 700 ends at step 704, where output activations are generated for a sub-part of the downloaded portion of the AI model based on the received activations. For example, the weight updater 126 may generate output activations for a sub-part of the downloaded portion of the AI model 106 based on the activations received from the target devices 134a - 134k. In an example embodiment, the generated output activations may be saved as activations 112 in the memory 104 of the parameter server 102. In an example where the sub-part includes a layer, the output activations of the layer may be generated by the weight updater 126 based on the hidden activations that are received after each micro-batch is executed at the target devices 134a - 134k.

[0076] The parameter server 102 may perform additional steps to improve the throughput of distributed training and inference of AI models on memory-constrained devices. For example, Figure 8 FIG. 800 shows a flowchart 800 of a process for updating an AI model at a parameter server according to an example embodiment. Although described with reference to Figure 1 system 100, the Figure 8 process is not limited to this system. Based on Figure 8 the following discussion of flowchart 800 and Figure 1 system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0077] Flowchart 800 begins at step 802, where gradients are received from the target devices. For example,Figure 1 The output data manager 124 can receive gradients from the target devices 134a - 134k. A gradient is an adjustment calculated by backpropagating a prediction error through an AI model. Thus, a gradient is a value representing the difference between where the model weights are and where they should be. Gradients can be placed in a data structure, such as a matrix. In an example embodiment, gradients can be received after the execution of each micro - batch, and the output data manager 124 and / or the weight updater 126 are configured to accumulate the received gradients until a specific number of micro - batches have been executed, and then use the received gradients to perform additional calculations and / or update the AI model 106. In another example embodiment, gradients can be accumulated at the target devices 134a - 134k for each micro - batch and then sent to the parameter server 102 after each micro - batch is completed.

[0078] In step 804, the weights of the AI model are updated based on the received gradients. For example, Figure 1 the weight updater 126 can use the gradients received from the target devices 134a - 134k to update the weights 108 of the AI model 106. The received gradients can be further processed (e.g., averaged) before using the processed gradients to update the AI model 106. In an example embodiment, the output data manager 124 receives gradients after each micro - batch and accumulates the gradients over mini - batches before the weight updater 126 updates the AI model 106 by updating the weights 108. In another embodiment, the output data manager 124 receives gradients after each micro - batch and at this time the weight updater 126 updates the AI model 106. For example, for an image analysis model, the size of a mini - batch can be set to 512 images. Thus, after the execution of 512 images, the gradients of the mini - batch can be provided to the parameter server to update the model. However, if each target device can only accommodate a micro - batch of 16 images, the gradients can be accumulated at the target devices after the execution of each micro - batch, and the gradients will only be applied to the model after 32 micro - batches. Thus, the micro - batch approach is mathematically equivalent to the mini - batch approach. That is, executing 512 images in one mini - batch and then applying the gradients of that mini - batch to the model is mathematically the same as executing 16 images in micro - batches, accumulating the gradients of each micro - batch until 512 images have been executed in 32 micro - batches or one mini - batch, and then applying the accumulated gradients to the model.

[0079] B. Data Parallelism in Distributed Training of an Artificial Intelligence Model

[0080] When training a distributed deep learning model in a large-scale environment, one challenge in deep learning is communication between target devices. For example, the latency of exchanging gradients across all target devices (e.g., in an implementation without a parameter server) is a time-consuming process. Typically, in synchronous data parallel distributed deep learning, the main computational steps include computing gradients using mini-batches on GPUs, computing the mean of the gradients via inter-GPU communication, and then updating the model. To compute the mean of the gradients, a communication operation (e.g., AllReduce) can be used to reduce the target arrays in all GPUs to a single array and return the single array to all GPUs. Even in a scenario where a parameter server is used, it may be necessary for GPUs to cache all layers of the AI model at high speed.

[0081] In an example embodiment, performing a dataset in micro-batches on one sub-part of an AI model at a time provides several advantages, particularly for distributed training of such an AI model in a data parallel manner. For example, the technique enables one or more parameter servers to reduce (e.g., optimize, average, and update) all parameters of the AI model while reducing the parameters that occur in the target devices. Thus, parameter reduction can occur simultaneously at different levels (e.g., target device level and parameter server level). The benefit of this technique is zero or near-zero communication overhead in large-scale data parallelism.

[0082] For example, Figure 9 FIG. shows a diagram of parameter reduction with multi-level parallelism in a system 900 according to an example embodiment. The system 900 includes a parameter server 902 and target devices 906a - 906n, and this system can be implemented as a system 100 with a parameter server 102 and target devices 134a - 134k. As Figure 9 shown, while the target devices 906a - 906n are performing parameter reduction of a specific sub-part (e.g., the current layer) of the AI model at the target device level 908, the parameter server 902 can also perform parameter reduction of its parameters of the AI model at the parameter server level 904 for another sub-part (e.g., the previous layer) of the AI model. Thus, the parameter server 902 can be responsible for reducing parameters, such as averaging the gradients and / or otherwise optimizing them, and then performing subsequent weight updates of the AI model outside the target devices in parallel with the computations at the target devices, thereby accelerating the overall computation. More parameter servers can be added to the system 900, and this multi-level parallel parameter reduction technique scales well with the addition of parameter servers to reduce communication overhead even at commodity network speeds. For example, the parameter server can perform parameter server-level parameter reduction in parallel with the target devices that perform parameter reduction at the target device level.

[0083] Figure 10FIG. 1000 is a timing diagram showing parameter reduction in multi - level parallelism in a system according to an example embodiment. For example, FIG. 1000 depicts parameter reduction in multi - level parallelism of an AI model executed in a system (such as, Figure 9 the system 900 shown). FIG. 1000 shows a timeline 1004 with different time periods 1014, 1016, and 1018. During each time period, tasks 1002 related to the training of the AI model executed at the target device (e.g., Figure 9 the target devices 906a - 906n) and the parameter server (e.g., Figure 9 the parameter server 902) can be executed in parallel to improve the computing speed.

[0084] For example, during the first time period 1014, the target device can execute task 1008, which is the computation of the current layer N, and at the same time, also execute task 1010, which is a full reduction operation between the target devices of the previous layer N + 1. The result 1024 of the full reduction operation of the previous layer N + 1 is sent to the parameter server. Additionally, during the first time period 1014, the parameter server executes task 1006 and task 1012. Task 1006 is the preparation for the next layer N - 1, and task 1012 is the parameter reduction of the layer N + 2 before the previous layer. The preparation for the next layer N - 1 includes sending the necessary data 1020 (e.g., weights and activations of the AI model) to the target device.

[0085] During the second time period 1016, the target device can execute task 1008, which is the computation of layer N - 1, based on the received data 1020, and at the same time, also execute task 1010, which is a full reduction operation between the target devices of layer N. The result 1026 of the full reduction operation of layer N is sent to the parameter server. Additionally, during the second time period 1014, the parameter server executes task 1006 and task 1012. Task 1006 is the preparation for layer N - 2, and task 1012 is the parameter reduction of layer N + 1. The preparation for layer N - 2 includes sending the necessary data 1022 to the target device.

[0086] During the second time period 1016, the target device can execute task 1008, which is the computation of layer N - 1, based on the received data 1020, and at the same time, also execute task 1010, which is a full reduction operation between the target devices of layer N. The result 1026 of the full reduction operation of layer N is sent to the parameter server. Additionally, during the second time period 1016, the parameter server executes task 1006 and task 1012. Task 1006 is the preparation for layer N - 2, and task 1012 is the parameter reduction of layer N + 1. The preparation for layer N - 2 includes sending the necessary data 1022 to the target device.

[0087] The multi-level reduction process continues in a similar manner at the parameter server and the target device for each time period until the training of the AI model is completed. For example, during the third time period 1018, the target device may perform task 1008 based on the received data 1022, where task 1008 is the calculation of layer N-2, and at the same time perform task 1010, where task 1010 is the full reduction operation between the target devices of layer N-1. Additionally, during the third time period 1018, the parameter server performs task 1006 and task 1012, where task 1006 is the preparation of layer N-3 and task 1012 is the parameter reduction of layer N.

[0088] The multi-level parallel parameter reduction process can be implemented in various ways. For example, Figures 11 - 13 illustrates the process for distributed training of an AI model. More specifically, Figure 11 FIG. 1100 is a flowchart showing a process for parallel parameter reduction in a provided system according to an example embodiment. For example, the parallel parameter reduction can be performed by Figure 9 the parameter server 902 and the target devices 906a-906n of the system 900 shown and / or Figure 1 the parameter server 102 and the target devices 134a-134k of the system 100 shown.

[0089] Flowchart 1100 begins at step 1102, where a main copy of the AI model is stored. For example, as Figure 1 shown in, when the AI model 106 is trained using the dataset 114, for example, the data manager 118 can store the main copy of the AI model 106 together with its associated weights 108, activations 112, and gradients 110 at the parameter server 102.

[0090] In step 1104, a portion of the AI model is transmitted to the target device. For example, as Figure 1 shown in, the transmitter 122 transmits a portion of the AI model 106 from the parameter server 102 to the target device 134a, which may be a memory-constrained device. Thus, in the example embodiment, the target device 134a may not have enough memory to efficiently execute the AI model 106. In another example embodiment, the target device 134a may have a large enough memory to store the entire AI model 106, but it may be more efficient to download and store only a portion of the AI model 106 rather than storing the entire instance of the AI model 106 according to the execution requirements.

[0091] In step 1106, a micro-batch size suitable for the target device is determined. As referred to above with reference to Figure 2 and Figure 6As described, the micro-batch size can be automatically or manually configured by the batch manager 120 at discrete points during the training of the AI model 106 based on the communication rate between the target device 134a and the parameter server 102. In one embodiment, the batch manager 116 can initially select the micro-batch size based on the hardware specifications of the target device 134a and then iteratively adjust it to the optimal micro-batch size, e.g., based on the computation time of the target device 134a, the size of the sub-part of the AI model 106 to be transmitted, and / or the communication bandwidth of the system 100.

[0092] Flowchart 1100 ends at step 1108. In step 1108, simultaneously, a micro-batch set of the training dataset is executed on a first sub-part of the transmission part of the AI model at the target device to generate gradients, parameter reduction is performed on a second sub-part of the transmission part of the AI model, and weights of a third sub-part of the transmission part of the AI model are sent to the target device. For example, when the target device 134a executes a micro-batch set of the dataset 114 on a first sub-part (e.g., the current layer) of the AI model 106, the weight updater 126 can perform parameter reduction on a second sub-part (e.g., the layer before the previous layer) of the AI model 106, and simultaneously (or, substantially simultaneously), the transmitter 112 can send the weights of a third sub-part (e.g., the next layer) of the AI model 106 to the target device 134a. For example, the parameter server 102 can perform these tasks according to Figure 10 Figure 1000 as shown.

[0093] In an example embodiment, the weight updater 126 is configured to perform parameter reduction using the gradients received from the target device 134a, which are generated by the target device 134a executing a micro-batch set of the dataset 114 on a second sub-part (e.g., the layer before the layer before the current layer) of the AI model 106. The weight updater 126 is also configured to generate an average value of the received gradients in any manner known in the art. For example, the weight updater 126 can generate an average value of the received gradients by using operations and libraries provided in the AI framework. The weight updater 126 can also perform other operations on the received gradients and / or optimize them in other ways. The weight updater 126 is also configured to update the AI model 106 by using the average value of the received gradients to update the weights 108.

[0094] In an example embodiment, the target devices 134a - 134k are configured to perform parameter reduction on the gradients generated by the target devices 134a - 134k in a manner similar to the parameter server 102. For example, the output manager 154 can generate an average value of the gradients generated by the target device 134a. The output manager 154 can also perform other operations on the gradients and / or optimize them in other ways.

[0095] In addition to performing the above-described processes depicted in the flowchart 1100, the parameter server 102 may perform additional processes. Training of AI models requires computational and memory resources, and for larger AI models, more computational and memory resources are needed. Deep learning systems may use single-precision (i.e., 32-bit) format (which is a common floating-point format), double-precision (i.e., 64-bit) format, or half-precision (i.e., 16-bit) format for computational workloads, such as storage and update of data such as weights, activations, and gradients. Mixed-precision methods combine the use of different numerical formats within one computational workload. By using mixed-precision training, the memory bandwidth requirements can be reduced because fewer bits can be used to store the same number of values. The computational time on the processor can also be improved, which can provide higher throughput for lower-precision mathematics. Additionally, some devices and AI frameworks may include automatic support for mixed-precision methods. For example, Figure 12 FIG. 1200 is a flowchart showing a process for providing mixed-precision training of an AI model according to an example embodiment. Based on Figure 12 the flowchart 1200 and Figure 1 the following discussion of the system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0096] The flowchart 1200 begins at step 1202, where the weights of the fourth sub-part of the transmission portion of the AI model are converted to a first precision format before being sent to the target device. For example, as Figure 1 shown in, before the transmitter 122 sends the converted weights to the target device 134a, the precision formatter 128 may convert the weights 108 of the AI model 106 to a first precision (e.g., half-precision) format. For example, using a lower-precision format, the computational time can be less. In an example embodiment, any precision format may be used as needed to optimize the performance of the AI model 106.

[0097] In step 1204, the gradients received from the target device are converted to a second precision format. For example, as Figure 1 shown in, the precision formatter 128 is configured to convert the gradients received from the target device 134a to a second precision (e.g., single-precision) format. In an example embodiment, the conversion of the gradients to the second precision format may be performed before or after certain operations (e.g., summation, averaging, etc.) on the received gradients. In other embodiments, the received gradients may simply be converted to the second precision format before being stored as gradients 110 in the memory 104.

[0098] In step 1206, the converted gradients are used to update the weights. For example, as Figure 1As shown in , the weights 108 of the AI model 106 can be updated by the weight updater 126 using the transformed gradients.

[0099] In embodiments, the flowchart 1200 can be performed with fewer or more steps or different steps than those shown. For example, different mixed-precision methods can be utilized with different precisions. For example, for a training iteration of a sub-part (e.g., a layer) of the AI model 106, the weights 108 can be transformed into a half-precision format for the forward pass, and the generated activations can also be kept in the half-precision format. In the backward pass, the weights 108 can be kept in the half-precision format together with the generated gradients. Once the average gradient is computed, the average gradient can be transformed into a single-precision format before updating the weights 108 of the AI model 106. For various reasons, many other operational embodiments can be implemented using the system 100. For example, the weight update (e.g., the weight gradient multiplied by the learning rate) may become too small to be represented in half-precision to maintain model accuracy. A single-precision or double-precision format may result in more computational time and / or resources for training the model.

[0100] The parameter server 102 can perform additional processes to manage the target devices 134a - 134k. For example, Figure 13 A flowchart 1300 showing a process for providing training of an AI model using multiple target devices according to an example embodiment is shown. Based on Figure 13 the flowchart 1300 and Figure 1 the following discussion of the system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0101] The flowchart 1300 begins at step 1302, where another portion of the AI model is transmitted to another target device. For example, the transmitter 122 can transmit another portion of the AI model 106 to another target device (such as, Figure 9 the target device 906n shown in ). In an example embodiment, multiple target devices can be used to accelerate the training time of the AI model. The system 900 can include any number of target devices, from one to multiple, and each target device is communicatively connected to the parameter server 902 via one or more suitable interfaces (e.g., PCI or PCIe).

[0102] In step 1304, gradients are received from another target device to perform parameter reduction for another portion of the AI model. Continuing the example of step 1302, the target device 906n can send gradients for the output data manager 124 at the parameter server 902 to receive for the portion of the AI model 106 that the target device 906n receives and executes on.

[0103] C. Dynamic Multi-Layer Execution of Artificial Intelligence Modeling

[0104] Another significant advantage of the execution paradigm above (i.e., executing the dataset in micro-batches on one sub-part of the AI model at a time) is that it only requires statically defining a sub-part or a portion thereof (e.g., a layer or a sub-layer), rather than the entire model computational graph as conventionally required. Thus, the number of layers within the AI model can be dynamically modified based on any number of factors, such as performance, alternative datasets, or other statistical observations.

[0105] New models based on neural architecture search (NAS) and its probabilistic counterparts are emerging, and the frictionless approach to dynamic execution provides improved modeling techniques that are currently very challenging to develop. NAS is a technique or algorithm for searching for the best neural network architecture based on a set of defined building blocks that can be used for the neural network to be constructed. These building blocks can be sampled and pieced together to construct a network similar to other known networks in the art, but can include different combinations and configurations of the building blocks. The network constructed by NAS can be trained and tested, and based on the test results, the building blocks can be adjusted. The network constructed by NAS can be improved by operations such as adding layers, removing layers, or otherwise changing the layers.

[0106] Therefore, techniques that allow the number of layers within the AI model to be dynamically modified based on any number of factors are very beneficial in NAS and other application areas. For example, Figure 14 FIG. 1400 is a flow chart showing a process for providing dynamic execution for AI modeling according to an example embodiment. Although described with reference to Figure 1 system 100, the Figure 14 method is not limited to this system. Based on the following discussion of Figure 1 system 100, other structural and operational embodiments will be apparent to those skilled in the relevant art. Flow chart 1400 may include the steps already described above with reference to, for example, Figure 1 , Figure 2 and Figure 6 and will not be described in detail below for the sake of brevity.

[0107] Flow chart 1400 begins at step 1402, where a master copy of the AI model is stored in the parameter server. For example, as shown in Figure 1 , the data manager is configured to store the master copy of the AI model 106 in the memory 104 of the parameter server 102.

[0108] In step 1404, a micro-batch size suitable for the target device is determined. For example, the batch manager 120 may be configured to determine a micro-batch size suitable for the target device 134a. In an example embodiment, the target device 134a may be a memory-constrained device such that the memory of the target device 134a may be insufficient to efficiently execute the AI model 106. In an alternative embodiment, the target device 134a may be able to accommodate the AI model 106 in its entirety. However, in this embodiment, it may be more efficient or desirable to download and store only a portion of the AI model 106 at a given time rather than the entire instance of the AI model 106.

[0109] In step 1406, a portion of the AI model is transmitted to the target device. For example, the transmitter 122 may be configured to transmit a portion of the AI model 106 to the target device 120b.

[0110] In step 1408, output data can be received from the target device, where the output data is generated by executing a micro-batch set of the dataset on a sub-portion of the transmitted portion of the AI model at the target device. For example, the output data manager 124 may be configured to receive output from Figure 1 the target device 134a. The output data can be generated by executing a micro-batch set of the dataset (e.g., dataset 114) on a sub-portion (e.g., layer or sub-layer) of the transmitted portion (e.g., layer or sub-layer) of the AI model 106 at the target device 134a. The output data can be, for example, activations and gradients generated for inference or for training of the AI model 106 on the forward pass and the backward pass, respectively. In an example embodiment, the target device 134a may send a signal instead of the output data, where the signal indicates that the micro-batch set has been executed at the target device 134a. The parameter server 102 can then be configured to act (e.g., perform subsequent steps) based on the signal rather than based on the output data.

[0111] The flowchart 1400 ends with step 1410. In step 1410, the AI model is evaluated based on one or more metrics to determine if any changes to the execution of the AI model are needed. For example, the model evaluator 130 may be configured to evaluate the AI model 106 based on one or more metrics to determine if any changes to the execution of the AI model 106 are needed, such as dynamically increasing or decreasing the number of layers to be executed.

[0112] The one or more metrics can be based on any number of factors, such as current performance, alternative datasets, or other statistical observations. In an example embodiment, the one or more metrics include precision statistics of the gradients and weights of a sub-portion (e.g., layer or sub-layer) of the transmitted portion of the AI model 106. For example, Figure 15FIG. 1500 shows a flowchart providing a process for determining a computational precision for dynamic execution for AI modeling. Although described with reference to Figure 1 system 100 of Figure 15 , the method is not limited to this system. Based on the following discussion of Figure 1 system 100 of

[0113] , other structural and operational embodiments will be apparent to those skilled in the relevant art.

[0114] In another embodiment, one or more metrics include an accuracy measurement of the AI model. For example, Figure 16 FIG. 1600 shows a flowchart providing a process for determining whether to stop or continue the execution of an AI model based on the accuracy of the AI model. Although described with reference to Figure 1 system 100 of Figure 16 , the method is not limited to this system. Based on the following discussion of Figure 1 system 100 of

[0115] , other structural and operational embodiments will be apparent to those skilled in the relevant art. Figure 1The model accessor 130 is configured to determine whether to stop or continue the execution of the AI model based on the accuracy of the AI model. That is, the model accessor 130 can stop the execution of the AI model when the accuracy of the AI model exceeds a predetermined threshold, or can continue the execution of the AI model when the accuracy of the AI model does not exceed the predetermined threshold. For example, the accuracy measurement of the AI model 106 can be the classification accuracy, which is the ratio of the number of correct predictions to the total number of input samples. In an exemplary embodiment, when the accuracy measurement of the AI model 106 exceeds a predetermined threshold, the execution of the AI model 106 can be paused, and the predetermined threshold can be any predefined value, such as 95%. According to this exemplary embodiment, when the accuracy measurement of the AI model 106 does not exceed the predetermined threshold of 95%, for example, when the accuracy is 80%, the execution of the AI model 106 can be continued. In an exemplary embodiment, for example, the continued execution of the AI model 106 can be determined dynamically for some number of layers or until the next evaluation of the AI model 106. In an exemplary embodiment, the AI model 106 can be executed one sub - part (e.g., layer or sub - layer) at a time, and the AI model 106 can be evaluated after each sub - part is executed.

[0116] Other metrics can be used to evaluate the AI model 106. For example, log loss, metrics derived from a confusion matrix, area under the curve, F1 score, mean absolute error, mean squared error. When other metrics are used, an appropriate threshold for each metric can be determined and applied to the evaluation of the AI model 106. Other factors such as a new dataset being used can cause the AI model 106 to be evaluated and / or its execution to change.

[0117] In the foregoing discussion of the flowcharts 200, 300, 600 - 800, and 1100 - 1600, it should be understood that, sometimes, these steps can be performed in a different order or even simultaneously with other steps. Other operational embodiments will be apparent to those skilled in the relevant art. It is also noted that the foregoing general description of the operation of the systems 100 and 900 is for illustrative purposes only, and embodiments of the systems 100 and 900 can include different hardware and / or software and can operate in a manner different from that described above.

[0118] III. Exemplary Computer System Implementations

[0119] Parameter server 102, target devices 134a - 134k, parameter server 904, and target devices 906a - 906n, as well as each of flowcharts 200, 300, 600 - 800, and / or 1100 - 1600, can be implemented in hardware or in hardware in combination with software and / or firmware. For example, parameter server 102, target devices 134a - 134k, parameter server 904, and target devices 906a - 906n, as well as flowcharts 200, 300, 600 - 800, and / or 1100 - 1600 can be implemented as computer program code / instructions configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, parameter server 102, target devices 134a - 134k, parameter server 904, and target devices 906a - 906n, as well as flowcharts 200, 300, 600 - 800, and / or 1100 - 1600 can be implemented as hardware logic / circuitry.

[0120] For example, in one embodiment, one or more of parameter server 102, target devices 134a - 134k, parameter server 904, and target devices 906a - 906n, as well as flowcharts 200, 300, 600 - 800, and / or 1100 - 1600 can be implemented together in a SoC in any combination. The SoC can include an integrated circuit chip that includes one or more of the following: a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or other circuitry, and can optionally execute the received program code and / or include embedded firmware to perform functions.

[0121] Figure 17 An exemplary implementation of a computing device 1700 in which embodiments can be implemented is depicted. For example, parameter server 102, target devices 134a - 134k, parameter server 904, and target devices 906a - 906n can each be implemented in one or more computing devices similar to computing device 1700 in a fixed or mobile computer embodiment, including one or more features and / or alternative features of computing device 1700. The description of computing device 1700 provided herein is for illustrative purposes and is not intended to be limiting. Embodiments can be implemented in other types of computer systems, as known to those skilled in the relevant art.

[0122] As Figure 17As shown in the figure, the computing device 1700 includes one or more processors (referred to as processor circuitry 1702), system memory 1704, and a bus 1706 that couples various system components, including system memory 1704, to processor circuitry 1702. Processor circuitry 1702 is an electrical and / or optical circuit implemented in one or more physical hardware circuit devices and / or integrated circuit devices (semiconductor material chips or dies) as a central processing unit (CPU), microcontroller, microprocessor, and / or other physical hardware processor circuitry. Processor circuitry 1702 can execute program code stored in a computer-readable medium, such as program code of an operating system 1730, application programs 1732, other programs 1734, and the like. Bus 1706 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of various bus architectures. System memory 1704 includes read-only memory (ROM) 1708 and random access memory (RAM) 1710. The basic input / output system 1712 (BIOS) is stored in ROM 1708.

[0123] The computing device 1700 also has one or more of the following drives: a hard disk drive 1714 for reading from and writing to a hard disk, a disk drive 1716 for reading from or writing to a removable disk 1718, and an optical disk drive 1720 for reading from or writing to a removable optical disk 1722 such as a CD ROM, DVD ROM, or other optical medium. The hard disk drive 1714, disk drive 1716, and optical disk drive 1720 are connected to bus 1706 via a hard disk drive interface 1724, a disk drive interface 1726, and an optical disk drive interface 1728, respectively. The drives and their associated computer-readable media provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer. Although hard disks, removable disks, and removable optical disks are described, other types of hardware-based computer-readable storage media can be used to store data, such as flash memory cards, digital video disks, RAM, ROM, and other hardware storage media.

[0124] Multiple program modules can be stored on a hard disk, magnetic disk, optical disk, ROM, or RAM. These programs include an operating system 1730, one or more application programs 1732, other programs 1734, and program data 1736. The application programs 1732 or other programs 1734 can include, for example, computer program logic (e.g., computer program code or instructions) for implementing the parameter server 102, the target devices 134a - 134k, the parameter server 904, and the target devices 906a - 906n, as well as the flowcharts 200, 300, 600 - 800, and / or 1100 - 1600 (including any suitable steps of the flowcharts 200, 300, 600, and / or 1100 - 1600), and / or other embodiments described herein.

[0125] A user can input commands and information into the computing device 1700 through input devices such as a keyboard 1738 and a pointing device 1740. Other input devices (not shown) can include a microphone, a joystick, a gamepad, a satellite antenna, a scanner, a touch screen and / or a touchpad, a voice recognition system for receiving voice input, a gesture recognition system for receiving gesture input, etc. These and other input devices are typically connected to the processor circuit 1702 through a serial port interface 1742, and the serial port interface 1742 is coupled to the bus 1706, but can also be connected through other interfaces, such as a parallel port, a game port, or a Universal Serial Bus (USB).

[0126] The display screen 1744 is also connected to the bus 1706 through an interface (such as a video adapter 1746). The display screen 1744 can be external to the computing device 1700 or incorporated into the computing device 1700. The display screen 1744 can display information, as well as a user interface for receiving user commands and / or other information (e.g., through touch, finger gestures, a virtual keyboard, etc.). In addition to the display screen 1744, the computing device 1700 can include other peripheral output devices (not shown), such as speakers and printers.

[0127] The computing device 1700 is connected to a network 1748 (e.g., the Internet) through an adapter or network interface 1750, a modem 1752, or other devices for establishing communication on the network. The modem 1752 (which can be internal or external) can be connected to the bus 1706 through the serial port interface 1742, as Figure 17 shown, or can be connected to the bus 1706 using another interface type (including a parallel interface).

[0128] As used herein, the terms "computer program medium", "computer-readable medium", and "computer-readable storage medium" are used to refer to physical hardware media such as hard disks associated with hard disk drive 1714, removable disks 1718, removable optical disks 1722, and other physical hardware media such as RAM, ROM, flash memory cards, digital video disks, zip disks, MEM, nanotechnology-based storage devices, and other types of physical / tangible hardware storage media. Such computer-readable storage media are distinct from and do not overlap with communication media (excluding communication media). Communication media include computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave. The term "modulated data signal" refers to a signal having one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media include wireless media such as acoustic, RF, infrared, and other wireless media, as well as wired media. Embodiments also relate to such communication media that are separate from and do not overlap with embodiments involving computer-readable storage media.

[0129] As described above, computer programs and modules (including application programs 1732 and other programs 1734) may be stored on hard disks, disks, optical disks, ROM, RAM, or other hardware storage media. Such computer programs may also be received via network interface 1750, serial port interface 1742, or any other interface type. Such computer programs, when executed or loaded by an application, enable computing device 1700 to implement the features of the embodiments described herein. Thus, such computer programs represent the controller of computing device 1700.

[0130] Embodiments also relate to computer program products that include computer code or instructions stored on any computer-readable medium. Such computer program products include hard disk drives, optical disk drives, storage device packages, portable storage sticks, memory cards, and other types of physical storage hardware.

[0131] IV. Additional Example Embodiments

[0132] A system is described herein. In one embodiment, the system includes: a parameter server communicatively coupled to a target device, the parameter server including: a data manager configured to store a master copy of an artificial intelligence (AI) model; a transmitter configured to transmit a portion of the AI model to the target device; a batch manager configured to determine a micro-batch size suitable for the target device; and simultaneously, execute a set of micro-batches of a training dataset on a first sub-portion of the transmitted portion of the AI model at the target device to generate gradients, a weight updater configured to perform parameter reduction for a second sub-portion of the transmitted portion of the AI model, and the transmitter further configured to send weights of a third sub-portion of the transmitted portion of the AI model to the target device.

[0133] In one embodiment of the foregoing system, the weight updater is configured to perform parameter reduction by: receiving gradients from the target device, the gradients being generated by the target device executing a set of micro-batches of a training dataset on the second sub-portion; and generating an average of the received gradients.

[0134] In another embodiment of the foregoing system, the weight updater is further configured to update the AI model using the average of the received gradients.

[0135] In an additional embodiment of the foregoing system, the set of micro-batches includes a plurality of micro-batches configured to be executed in sequence, the set of micro-batches forming a mini-batch that includes a number of samples for training the AI model for each update.

[0136] In yet another embodiment of the foregoing system, the micro-batch size is configurable based on the execution rate of the set of micro-batches at the target device and the communication rate between the target device and the parameter server.

[0137] In still another embodiment of the foregoing system, the parameter server further includes a precision formatter configured to: convert the weights of a fourth sub-portion of the transmitted portion of the AI model to a first precision format before sending the weights to the target device; convert the gradients received from the target device to a second precision format; and update the weights using the converted gradients.

[0138] In another embodiment of the foregoing system, the transmitter is further configured to transmit another portion of the AI model to another target device; and the weight updater is further configured to receive gradients from the another target device to perform parameter reduction for the another portion of the AI model.

[0139] A method implemented in a parameter server is described herein. The method includes: storing a main copy of an artificial intelligence (AI) model; transmitting a portion of the AI model to a target device; determining a micro-batch size suitable for the target device; and simultaneously, executing a set of micro-batches of a training dataset on a first sub-portion of the transmitted portion of the AI model at the target device to generate gradients, performing parameter reduction on a second sub-portion of the transmitted portion of the AI model, and sending weights of a third sub-portion of the transmitted portion of the AI model to the target device.

[0140] In another embodiment of the foregoing method, performing parameter reduction includes: receiving gradients from the target device, the gradients being generated by executing a set of micro-batches of a training dataset on the second sub-portion at the target device; and generating an average value of the received gradients.

[0141] An embodiment of the foregoing method further includes updating the AI model using the average value of the received gradients.

[0142] In another embodiment of the foregoing method, the set of micro-batches includes a plurality of micro-batches configured to be executed in sequence, and the set of micro-batches forms a mini-batch that includes a number of samples for training per update or a number of samples for inference per inference cycle.

[0143] In an additional embodiment of the foregoing method, the micro-batch size is configurable based on the execution rate of the set of micro-batches at the target device and the communication rate between the target device and the parameter server.

[0144] Another embodiment of the foregoing method further includes: before sending the weights to the target device, converting the weights of a fourth sub-portion of the transmitted portion of the AI model to a first precision format; converting the gradients received from the target device to a second precision format; and updating the weights using the converted gradients.

[0145] Yet another embodiment of the foregoing method further includes: transmitting another portion of the AI model to another target device; and receiving gradients from the another target device to perform parameter reduction on the another portion of the AI model.

[0146] A computer program product is also described herein. The computer program product includes a computer-readable storage device having recorded thereon computer program logic that, when executed by a processor-based computer system, causes the processor-based system to perform a method that includes: storing a primary copy of an artificial intelligence (AI) model at a parameter server; transmitting a portion of the AI model to a target device; determining a microbatch size suitable for the target device; and simultaneously, at the target device, performing a set of microbatches of a training dataset on a first sub-portion of the transmitted portion of the AI model to generate gradients, performing parameter reduction on a second sub-portion of the transmitted portion of the AI model, and sending weights of a third sub-portion of the transmitted portion of the AI model to the target device.

[0147] In an embodiment of the foregoing computer program product, performing parameter reduction includes: receiving gradients from the target device, the gradients being generated by performing a set of microbatches of a training dataset on the second sub-portion at the target device; and generating an average of the received gradients.

[0148] In an embodiment of the foregoing computer program product, the method further includes updating the AI model using the average of the received gradients.

[0149] In another embodiment of the foregoing computer program product, the set of microbatches includes a plurality of microbatches configured to be executed in sequence, and the set of microbatches forms a minibatch that includes a number of samples for training the AI model for each update.

[0150] In yet another embodiment of the foregoing computer program product, the microbatch size is configurable based on the execution rate of the set of microbatches at the target device and the communication rate between the target device and the parameter server.

[0151] In an embodiment of the foregoing computer program product, the method further includes transmitting another portion of the AI model to another target device; and receiving gradients from the another target device to perform parameter reduction on the another portion of the AI model.

[0152] V. Conclusion

[0153] Although various embodiments of the disclosed subject matter have been described above, it should be understood that they are presented by way of example only and not limitation. Those skilled in the relevant art will understand that various changes may be made in form and detail without departing from the spirit and scope of the embodiments as defined in the appended claims. Accordingly, the breadth and scope of the disclosed subject matter should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.

Claims

1. A system, comprising: A parameter server communicatively connected to a target device, the parameter server comprising: A data manager configured to store a main copy of an artificial intelligence (AI) model; A transmitter configured to transmit a portion of the AI model to the target device, the target device including an integrated circuit chip having on-chip memory smaller than the size of the overall AI model, and the size of the portion being based on the available size of the on-chip memory and the size of one or more layers of the AI model; A batch manager configured to determine a micro-batch size suitable for the target device; and Simultaneously with the integrated circuit chip in the target device executing a set of micro-batches of a training dataset on a first sub-portion of the transmitted portion of the AI model stored in the on-chip memory to generate gradients, A weight updater configured to perform parameter reduction for a second sub-portion of the transmitted portion of the AI model, and The transmitter is further configured to send weights for a third sub-portion of the transmitted portion of the AI model to the target device.

2. The system according to claim 1, wherein the weight updater is configured to perform parameter reduction by: Receiving gradients from the target device, the gradients being generated by the target device executing the set of micro-batches of the training dataset on the second sub-portion at the target device; and Generating an average value of the received gradients.

3. The system according to claim 2, wherein the weight updater is further configured to update the AI model using the average value of the received gradients.

4. The system according to claim 2, wherein the set of micro-batches includes a plurality of micro-batches configured to be executed in sequence, the set of micro-batches forming a mini-batch, the mini-batch including a number of samples for training the AI model for each update.

5. The system according to claim 2, wherein the micro-batch size is configurable based on the execution rate of the set of micro-batches at the target device and the communication rate between the target device and the parameter server.

6. The system according to claim 1, wherein the transmitter is further configured to transmit another portion of the AI model to another target device; and the weight updater is further configured to receive gradients from the another target device to perform parameter reduction for the another portion of the AI model.

7. A system, comprising: A parameter server communicatively connected to a target device, the parameter server comprising: A data manager configured to store a main copy of an artificial intelligence (AI) model; A transmitter configured to transmit a portion of the AI model to the target device; A batch manager configured to determine a micro-batch size suitable for the target device; and Simultaneously with executing a set of micro-batches of a training dataset on a first sub-portion of the transmitted portion of the AI model at the target device to generate gradients, The weight updater is configured to perform parameter reduction on a second sub - part of the transmitted part of the AI model, and the transmitter is further configured to send weights of a third sub - part of the transmitted part of the AI model to the target device, wherein the parameter server further includes a precision formatter configured to: before sending the weights to the target device, convert the weights of a fourth sub - part of the transmitted part of the AI model to a first precision format; convert the gradients received from the target device to a second precision format; and use the converted gradients to update the weights.

8. A method implemented in a parameter server, comprising: storing a main copy of an artificial intelligence (AI) model; transmitting a part of the AI model to a target device, the target device including an integrated circuit chip having on - chip memory smaller than the size of the whole AI model, and the size of the part being based on the available size of the on - chip memory and the size of one or more layers of the AI model; determining a micro - batch size suitable for the target device; and simultaneously with the integrated circuit chip in the target device executing a set of micro - batches of a training dataset on a first sub - part of the transmitted part of the AI model stored in the on - chip memory to generate gradients, performing parameter reduction on a second sub - part of the transmitted part of the AI model and sending weights of a third sub - part of the transmitted part of the AI model to the target device.

9. The method according to claim 8, wherein performing the parameter reduction includes: receiving gradients from the target device, the gradients being generated by executing the set of micro - batches of the training dataset on the second sub - part at the target device; and generating an average value of the received gradients.

10. The method according to claim 9, further comprising: updating the AI model using the average value of the received gradients.

11. The method according to claim 9, wherein the set of micro - batches includes a plurality of micro - batches configured to be executed in sequence, the set of micro - batches forms a mini - batch, and the mini - batch includes a number of samples for training the AI model for each update.

12. The method according to claim 9, wherein the micro - batch size is configurable based on the execution rate of the set of micro - batches at the target device and the communication rate between the target device and the parameter server.

13. The method according to claim 8, further comprising: transmitting another part of the AI model to another target device; and receiving gradients from the another target device to perform parameter reduction on the another part of the AI model.

14. A method implemented in a parameter server, comprising: storing a main copy of an artificial intelligence (AI) model; transmitting a part of the AI model to a target device; determining a micro - batch size suitable for the target device; Simultaneously with performing a set of micro-batches of a training dataset on a first sub-part of the transmitted portion of the AI model at the target device to generate gradients, perform parameter reduction on a second sub-part of the transmitted portion of the AI model and send weights for a third sub-part of the transmitted portion of the AI model to the target device; Before sending the weights to the target device, convert the weights for a fourth sub-part of the transmitted portion of the AI model to a first precision format; Convert the gradients received from the target device to a second precision format; and Use the converted gradients to update the weights.

15. A computer program product comprising a computer-readable storage device having computer program logic recorded thereon, the computer program logic causing a processor-based computer system to perform a method when executed by the processor-based computer system, the method comprising: Store a main copy of an artificial intelligence (AI) model at a parameter server; Transmit a portion of the AI model to a target device, the target device including an integrated circuit chip having on-chip memory that is smaller than the size of the overall AI model, and the size of the portion being based on the available size of the on-chip memory and the size of one or more layers of the AI model; Determine a micro-batch size suitable for the target device; And Simultaneously with the integrated circuit chip at the target device performing a set of micro-batches of a training dataset on a first sub-part of the transmitted portion of the AI model stored in the on-chip memory to generate gradients, perform parameter reduction on a second sub-part of the transmitted portion of the AI model and send weights for a third sub-part of the transmitted portion of the AI model to the target device.

16. The computer program product according to claim 15, wherein performing the parameter reduction comprises: Receive gradients from the target device, the gradients being generated by performing the set of micro-batches of the training dataset on the second sub-part at the target device; And Generate an average value of the received gradients.

17. The computer program product according to claim 16, wherein the method further comprises: Update the AI model using the average value of the received gradients.

18. The computer program product according to claim 16, wherein the set of micro-batches includes a plurality of micro-batches configured to be executed in sequence, the set of micro-batches forming a mini-batch, the mini-batch including a number of samples for training the AI model for each update.

19. The computer program product according to claim 15, wherein the micro-batch size is configurable based on the execution rate of the set of micro-batches at the target device and the communication rate between the target device and the parameter server.

20. The computer program product according to claim 15, wherein the method further comprises: Transmit another portion of the AI model to another target device; and receiving a gradient from the other target device to perform parameter reduction for the other part of the AI model.

Citation Information

Patent Citations

  • Communication optimizations for distributed machine learning

    US20190205745A1