Training a machine learning model using an accelerated pipeline having

By postponing the update of model parameters in the accelerator pipeline, especially for the embedding set of recommended models, the problems of inefficient resource utilization and different results during the training process are solved, and more efficient resource usage and training efficiency are achieved.

CN120584348APending Publication Date: 2025-09-02MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480008853.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-24
Filing Date
2024-03-11
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Training large and/or complex machine learning models is computationally time consuming and resource-intensive, and existing accelerator pipeline methods may introduce differences in training results and resource utilization is not efficient enough.

Method used

Using an accelerator pipeline with delayed updates, by storing frequent access values ​​in GPU memory, infrequent access values ​​in main memory, and delaying update model parameters, especially for the embedded set of recommended models, GPUs are scheduling small batch inputs to reduce latency and improve resource utilization efficiency.

Benefits of technology

Use resources more efficiently during training operations, reduce or eliminate differences in training results, and improve training efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120584348A_ABST
    Figure CN120584348A_ABST
Patent Text Reader

Abstract

Innovations are described for training machine learning models using an accelerator pipeline with deferred updates to model parameters. The accelerator identifies one or more first class micro-batches ("MBs") and second class MBs of the working set. The first class of MBs contain, as input, frequent access embedding stored in graphics processing unit ("GPU") memory. The accelerator schedules the first class MB (s) for training using one or more GPUs. During training, the accelerator obtains a second type of MB, which contains as input the infrequent access embedding stored in the main memory. At least some updates to the model parameters from training with the first class MB (s) are delayed until after training with the second class MB. The accelerator schedules a second type of MBs for training. Finally, after training with the second class of MBs, the accelerator updates non-frequent access values for the second class of MBs. At this time, the GPU (s) also update the model parameters, applying delayed update.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In computer systems, machine learning uses statistical techniques to extract features from a training data set. The extracted features can then be applied when classifying new data. Machine learning techniques can be useful in a wide range of use cases, such as recognizing images and speech, analyzing and classifying information, and performing various other classification tasks. For example, features can be extracted by training an artificial neural network, which can be a deep neural network ("DNN") with multiple layers of hidden nodes. After the DNN is trained, new data can be classified based on the trained DNN.

[0002] Recommendation models (also known as recommendation systems, platforms, or engines) typically provide suggestions to users for items that are interesting, promising, or otherwise relevant to the user. These items can be books, television programs, movies, news stories, products, songs, people, other entities, restaurants, other locations, services, or any other type of item. In many implementations, a recommendation model is trained using machine learning (e.g., using a DNN) based on a set of user ratings of item selections. After training, the recommendation model can be used to identify new items relevant to a given user, or the recommendation model can be used to identify new users relevant to a given item.

[0003] Machine learning tools are typically executed on a general-purpose processor in a computer system, such as a central processing unit ("CPU") or a general-purpose graphics processing unit ("GPU"). However, training operations can be computationally expensive, making training impractical for large and / or complex machine learning models. Even when training is performed using specialized computer hardware, it can be time-consuming and resource-intensive.

[0004] In particular, recommendation models targeting a large number of users and a wide range of items can require a large amount of memory and computing resources in a computer system for training operations. Some methods use accelerator pipelines for training operations to more efficiently use resources during training operations, which can potentially save time and reduce energy consumption. However, in some cases, this approach may introduce differences in the results of the training process compared to the training results without accelerator pipelines. Therefore, there is room for improvement in methods for training machine learning models. Summary of the Invention

[0005] In summary, the detailed description presents innovations for training machine learning models ("MLMs") (such as recommendation models) using an accelerator pipeline with deferred updates to model parameters. With these innovations, in many usage scenarios, resources can be used more efficiently during training operations due to the accelerator pipeline, while reducing or even eliminating variance in training process results, compared to MLM training results without the accelerator pipeline.

[0006] According to the first technology and tool set described herein, a computer system includes at least one graphics processing unit ("GPU"), GPU memory, main memory associated with at least one central processing unit ("CPU"), and an accelerator. The GPU(s) are configured to train an MLM, such as a recommendation model. The MLM has model parameters, which may include network parameters (such as weight values ​​for nodes, bias values ​​for nodes, and values ​​for activation functions in a neural network) and values ​​in an embedding set. The embeddings may be organized as an embedding table. The GPU memory is configured to store frequently accessed values ​​of the embedding set. On the other hand, the main memory is configured to store infrequently accessed values ​​of the embedding set. The accelerator is configured to perform accelerator operations using deferred updates to the model parameters.

[0007] In operation, the accelerator identifies one or more first-class micro-batches ("MBs") and second-class MBs for a given work set. For example, the accelerator identifies a mini-batch input associated with the given work set and classifies the mini-batch input into (multiple) first-class MBs and second-class MBs. Each of the (multiple) first-class MBs contains only inputs from frequently accessed values ​​stored in GPU memory. The second-class MBs contain inputs from infrequently accessed values ​​stored in main memory. (The second-class MBs may also contain inputs from frequently accessed values ​​stored in GPU memory). The accelerator schedules (multiple) first-class MBs for training the MLM using (multiple) GPUs. During at least part of the training of the MLM using (multiple) first-class MBs, the accelerator obtains the second-class MBs. In this way, (multiple) GPUs can perform training operations on (multiple) first-class MBs while the accelerator obtains the second-class MBs, which can avoid introducing additional delays from collecting inputs for the second-class MBs. In terms of the update timing of the model parameters, at least some updates to the model parameters from the training using (multiple) first-class MBs are postponed until after the MLM is trained using the second-class MBs. By deferring updates, training results can closely track or even match the results of training the MLM with a small batch of inputs in training iterations without using the accelerator pipeline. The accelerator uses (multiple) GPUs to schedule the second type of MBs for training the MLM. Finally, after training the MLM with the second type of MBs, the accelerator updates the infrequently accessed values ​​for the second type of MBs, which are stored in main memory. At this time, the (multiple) GPUs also update model parameters, such as network parameters and frequently accessed embeddings. In doing so, the (multiple) GPUs apply the deferred updates.

[0008] According to the second technology and tool set described herein, an accelerator performs accelerator operations to train an MLM, such as a recommendation model. The accelerator identifies (multiple) first-class MBs and second-class MBs for a given working set. The accelerator uses (multiple) GPUs to schedule (multiple) first-class MBs for training the MLM. During at least part of the time of training the MLM using (multiple) first-class MBs, the accelerator obtains second-class MBs. In terms of update timing of model parameters, at least some updates to model parameters from training using (multiple) first-class MBs are postponed until after training the MLM using the second-class MBs. The accelerator uses (multiple) GPUs to schedule second-class MBs for training the MLM. Finally, after training the MLM using the second-class MBs, the accelerator updates the infrequently accessed values ​​for the second-class MBs, which are stored in main memory.

[0009] According to the third technology and tool set described herein, (multiple) GPUs perform GPU operations to train MLMs, such as recommendation models. In response to scheduling (multiple) first-class MBs of a given working set for training MLMs, the accelerator continuously performs training operations using (multiple) first-class MBs. (Multiple) GPUs postpone at least some updates to model parameters from training using (multiple) first-class MBs until after training MLMs using second-class MBs of a given working set. In response to scheduling second-class MBs for training MLMs, (multiple) GPUs perform training operations using second-class MBs. After training MLMs using second-class MBs, (multiple) GPUs provide updates for infrequently accessed values ​​for the second-class MBs, which are stored in main memory. At this time, (multiple) GPUs also update network parameters and frequently accessed values ​​stored in GPU memory. In doing so, (multiple) GPUs apply at least some of the postponed updates.

[0010] The innovations described herein may be implemented as part of a method, as part of a computer system configured to perform the method, or as part of a tangible computer-readable medium storing computer-executable instructions for causing one or more processors to perform the method when programmed thereby. The various innovations may be used in combination or individually. The innovations described herein include, but are not limited to, the innovations covered by the claims. This summary is provided to introduce, in simplified form, a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. The foregoing and other objects, features, and advantages of the present invention will become more apparent from the following detailed description, which continues with reference to the accompanying drawings and illustrates a number of examples. The examples may also have other and different applications, and some details may be modified in various aspects without departing from the spirit and scope of the disclosed innovations. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The following drawings illustrate some features of the disclosed innovation.

[0012] Figure 1 is a diagram illustrating an example computer system in which some described embodiments may be implemented.

[0013] Figure 2 is a diagram illustrating an example architecture for training an MLM using an accelerator with deferred updates to model parameters.

[0014] Figure 3a is a diagram illustrating an example set of embeddings for a machine learning model (“MLM”), and Figure 3b is a diagram illustrating an example deep neural network for MLM.

[0015] Figure 4a and Figure 4b is a diagram illustrating operations in an accelerator pipeline with deferred updates to model parameters when training an MLM.

[0016] Figure 5 is a flow chart illustrating a general technique for training an MLM using an accelerator pipeline with deferred updates to model parameters from the perspective of the accelerator.

[0017] Figure 6 is a flow chart illustrating a general technique for training an MLM using an accelerator pipeline with deferred updates to model parameters from the perspective of one or more GPUs. DETAILED DESCRIPTION

[0018] The detailed description presents innovations for training machine learning models ("MLMs") (such as recommendation models) using an accelerator pipeline with deferred updates to model parameters. Utilizing these innovations, resources can be used more efficiently during training operations due to the accelerator pipeline in many use cases. By deferring updates, the training process can minimize or even completely eliminate differences in training results compared to MLM training results without the accelerator pipeline. The innovations include, but are not limited to, the features of the claims.

[0019] In the examples described herein, the same reference numerals in different figures indicate the same components, modules or operations. More generally, various alternatives to the examples described herein are possible. For example, some methods described herein may be replaced by changing the order of the described method actions by splitting, repeating or omitting certain method actions, etc. Various aspects of the disclosed technology may be used in combination or alone. Some of the innovations described herein solve one or more of the problems mentioned in the background. Generally, a given technology / tool ​​cannot solve all such problems. It is to be understood that other examples may be used, and that structural, logical, software, hardware and electrical changes may be made without departing from the scope of this disclosure. Therefore, the following description should not be considered restrictive.

[0020] I. Example Computer System.

[0021] Figure 1 A general example of a suitable computer system (100) is illustrated in which several of the described innovations may be implemented. The innovations described herein relate to training a machine learning model ("MLM") using an accelerator pipeline with deferred updates to model parameters. The computer system (100) is not intended to suggest any limitation as to the scope of use or functionality, as these innovations may be implemented in a variety of computer systems, including special-purpose computer systems.

[0022] refer to Figure 1 , a computer system (100) includes a central processing unit ("CPU") or one or more processing cores (110...11x) of multiple CPUs and local memory (118). The processing core(s) (110...11x) are, for example, single on-chip processing cores and execute computer-executable instructions. The number of processing core(s) (110...11x) depends on the implementation and may be, for example, 4 or 8. The local memory (118) may be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two, accessible by the corresponding processing core(s) (110...11x). Alternatively, the processing cores (110...11x) may be part of a system on a chip ("SoC"), an application-specific integrated circuit ("ASIC"), or other integrated circuit.

[0023] The local memory (118) may store software (180) in the form of computer-executable instructions that implement various aspects of the innovation for training an MLM using an accelerator pipeline with deferred updates to model parameters, for operations performed by the corresponding processing core(s) (110 . . . 11x). For example, the instructions may be for operations to read / write infrequently accessed values ​​of an embedding set in the main memory (120), for operations to identify mini-batches of inputs for training by sampling, or for other operations. Figure 1 In the embodiment, the local memory (118) is an on-chip memory, such as one or more caches, for which access operations, transfer operations, etc. are fast using the (multiple) processing cores (110...11x).

[0024] The computer system (100) further includes a graphics processing unit ("GPU") or processing cores (130...13x) of multiple GPUs and local memory (138). The number of processing cores (130...13x) of the GPU depends on the implementation. For example, the processing cores (130...13x) are part of a single instruction multiple data ("SIMD") unit of the GPU. The SIMD width n, which depends on the implementation, indicates the number of elements (sometimes called lanes) of the SIMD unit. For example, for an ultra-wide SIMD architecture, the number of elements (lanes) of the SIMD unit can be 16, 32, 64, or 128. The GPU memory (138) can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two, which is accessible by the corresponding processing cores (130...130x). In some examples described herein, the GPU memory (138) stores frequently accessed values ​​of the embedded set.

[0025] The GPU memory (138) may store software (180) in the form of computer-executable instructions (such as shader code) that implement various aspects of the innovation for training an MLM using an accelerator pipeline with deferred updates to model parameters for operations performed by the corresponding processing cores (130...13x). For example, these instructions are for operations to read / write frequently accessed values ​​in the GPU memory (138), for training operations, or for other operations. Figure 1 In the embodiment, if there are multiple GPUs, the GPU memory (138) is a high-bandwidth memory with paired interconnections between the GPUs. For the GPU memory (138), access operations, transfer operations, etc. are very fast using the processing cores (130...13x).

[0026] The computer system (100) includes a main memory (120), which can be volatile memory (e.g., RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two, accessible by the processing core(s) (110 ... 11x, 130 ... 13x). In some examples described herein, the main memory (120) stores infrequently accessed values ​​of the embedding set. The main memory (120) can also store software (180) in the form of computer-executable instructions that implements various aspects of the innovation for training an MLM using an accelerator pipeline with deferred updates to model parameters. Figure 1 In the present invention, the main memory (120) is an off-chip memory. For the off-chip memory, access operations and transfer operations using the processing cores (110...11x, 130...13x) are slow. The computer system (100) may include a memory access engine (not shown) to facilitate access to the main memory (120).

[0027] The computer system (100) further includes an accelerator (190). The accelerator (190) includes logic (191) and a buffer (198). The logic (191) implements various aspects of the innovation for training an MLM using an accelerator pipeline with deferred updates to model parameters in the form of computer-executable instructions or other forms. For example, the logic implements operations for classifying and reordering small batches of inputs for a working set into micro-batches ("MBs") of the working set, scheduling MBs for execution on (multiple) GPUs (130), obtaining inputs for MBs with infrequently accessed values, reading / writing infrequently accessed values ​​to main memory (120), or other operations. The buffer (198) can, for example, store inputs classified as frequently accessed or infrequently accessed, store values ​​from an embedding set (retrieved from main memory (120) and / or GPU memory (138)) to be used as inputs to MBs in MLM training, or store other values.

[0028] More generally, the term "processor" may refer to any device that can process computer-executable instructions, and may include microprocessors, microcontrollers, programmable logic devices, digital signal processors, and / or other computing devices. A processor may be a processing core of a CPU, other general-purpose unit, or GPU. A processor may also be a specialized processor implemented using, for example, an ASIC or a field programmable gate array ("FPGA").

[0029] The term "control logic" may refer to a controller or more generally one or more processors operable to process computer-executable instructions, determine results, and generate outputs. Depending on the implementation, the control logic may be implemented by software executable on a CPU, software controlling dedicated hardware (e.g., a GPU or other graphics hardware), or dedicated hardware (e.g., in an ASIC).

[0030] The computer system (100) includes one or more network interface devices (140). The network interface device(s) (140) enable communication to another computing entity (e.g., a server, other computer system) over a network. The network interface device(s) (140) can support wired and / or wireless connections for use with a wide area network, a local area network, a personal area network, or other networks. For example, the network interface device(s) can include one or more Wi-Fi transceivers, Ethernet ports, cellular transceivers, and / or another type of network interface device and associated drivers, software, etc. The network interface device(s) (140) communicate information, such as computer executable instructions, audio or video input or output, or other data, in a modulated data signal over the network connection(s). A modulated data signal is a signal that has one or more of its characteristics set or changed in such a way as to encode the information in the signal. By way of example and not limitation, the network connection can use electrical, optical, RF, or other carrier waves.

[0031] The computer system (100) optionally includes other components such as: a motion sensor / tracker input (142) for a motion sensor / tracker; a game controller input (144) that receives control signals from one or more game controllers via a wired or wireless connection; a media player (146); a video source (148); and an audio source (150).

[0032] The computer system (100) optionally includes a video output (160) that provides video output to a display device. The video output (160) may be an HDMI output or other type of output. An optional audio output (160) provides audio output to one or more speakers.

[0033] The storage device (170) may be removable or non-removable and includes magnetic media (such as disks, tapes, or cartridges), optical media, and / or any other media that can be used to store information and that can be accessed within the computer system (100). The storage device (170) stores instructions for software (180) that implements various aspects of the innovation for training an MLM using an accelerator pipeline with deferred updates to model parameters.

[0034] The computer system (100) may have additional features. For example, the computer system (100) may include one or more other input devices and / or one or more other output devices. The other input device(s) may be a touch input device such as a keyboard, mouse, pen, or trackball, a scanning device, or another device that provides input to the computer system (100). The other output device(s) (160) may be a printer, a CD burner, or another device that provides output from the computer system (100).

[0035] An interconnection mechanism (not shown) such as a bus, controller, or network interconnects the components of the computer system (100). Typically, operating system software (not shown) provides an operating environment for other software executing in the computer system (100) and coordinates the activities of the components of the computer system (100).

[0036] Figure 1 The computer system (100) is a physical computer system. A virtual machine may include Figure 1 Parts of the tissue shown.

[0037] The term "application" or "program" may refer to software, such as any user mode instructions for providing functionality. The software of an application (or program) may also include instructions for an operating system and / or device drivers. The software may be stored in an associated memory. The software may be, for example, firmware. Although it is contemplated that a suitably programmed general-purpose computer or computing device may be used to execute such software, it is also contemplated that hard-wired circuitry or custom hardware (e.g., an ASIC) may be used in place of, or in combination with, software instructions. Therefore, the examples described herein are not limited to any specific combination of hardware and software.

[0038] The term "computer-readable medium" refers to any medium that participates in providing data (e.g., instructions) that can be read by a processor and accessed within a computing environment. Computer-readable media can take many forms, including but not limited to non-volatile media and volatile media. Non-volatile media include, for example, optical or magnetic disks and other persistent memories. Volatile media include dynamic random access memory ("DRAM"). Common forms of computer-readable media include, for example, solid-state drives, flash drives, hard disks, any other magnetic media, CD-ROMs, DVDs, any other optical media, RAM, programmable read-only memory ("PROM"), erasable programmable read-only memory ("EPROM"), USB memory sticks, any other memory chip or cassette tape, or any other medium from which a computer can read. The term "non-transient computer-readable medium" specifically excludes transient propagating signals, carrier waves and waveforms, or other intangible or transient media that can still be readable by a computer. The term "carrier wave" can refer to an electromagnetic wave that is modulated in amplitude or frequency to convey a signal.

[0039] Innovation can be described in the general context of computer executable instructions, which are executed in a computer system on a target real or virtual processor. Computer executable instructions may include instructions executable on the processing core of a general-purpose processor to provide the functionality described herein, instructions executable to control a GPU or dedicated hardware to provide the functionality described herein, instructions executable on the processing core of a GPU to provide the functionality described herein, and / or instructions executable on the processing core of a dedicated processor to provide the functionality described herein. In some implementations, computer executable instructions can be organized in program modules. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or split between program modules as desired. Computer executable instructions for program modules can be executed within a local or distributed computer system.

[0040] The terms "system" and "device" are used interchangeably herein. Unless the context clearly indicates otherwise, neither term implies any limitation on the type of computer system or device. In general, a computer system or device may be local or distributed and may include dedicated hardware and / or any combination of hardware and software that implements the functionality described herein.

[0041] Many examples are described in this disclosure and are presented for illustrative purposes only. The examples described are not restrictive in any sense and are not intended to be restrictive. As is readily apparent from this disclosure, the innovations currently disclosed are widely applicable in many contexts. Those of ordinary skill in the art will recognize that the disclosed innovations can be practiced with various modifications and alterations, such as structural, logical, software, and electrical modifications. Although specific features of the disclosed innovations may be described with reference to one or more specific examples, it will be understood that, unless expressly specified otherwise, such features are not limited to their usage in the one or more specific examples to which they are described. This disclosure is neither a literal description of all examples nor a list of features of the invention that must be present in all examples.

[0042] When a sequence number (such as "first," "second," "third," etc.) is used as an adjective before a term, the sequence number is used only (unless otherwise expressly specified) to indicate a particular characteristic, such as to distinguish the particular characteristic from another characteristic described by the same or similar term. The mere use of the sequence numbers "first," "second," "third," etc. does not indicate any physical order or position, any temporal order, or any ranking of importance, quality, or otherwise. In addition, the mere use of a sequence number does not define a numerical limitation on the characteristic identified by the sequence number.

[0043] When introducing an element, the articles "a," "an," "the," and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0044] When single device, parts, module or structure are described, multiple devices, parts, modules or structures (no matter whether they collaborate) can replace single device, parts, modules or structures and use instead.Described as the functionality owned by a single device can be owned by multiple devices instead, no matter whether they collaborate.Similarly, when multiple devices, parts, modules or structures are described in this article, no matter whether they collaborate, single device, parts, modules or structures can replace multiple devices, parts, modules or structures and use instead.Described as the functionality owned by multiple devices can be owned by single device instead.Usually, computer system or device can be local or distributed, and can include any combination of dedicated hardware and / or hardware and software realizing the functionality described herein.

[0045] Further, the techniques and tools described herein are not limited to the specific examples described herein. Rather, the corresponding techniques and tools can be utilized independently and separately from other techniques and tools described herein.

[0046] Devices, components, modules, or structures that communicate with each other do not need to communicate with each other continuously unless otherwise explicitly specified. Instead, such devices, components, modules, or structures only need to communicate with each other when necessary or desired, and can actually avoid exchanging data most of the time. For example, a device that communicates with another device via the Internet can go weeks without transmitting data to the other device. In addition, devices, components, modules, or structures that communicate with each other can communicate directly or indirectly through one or more intermediaries.

[0047] As used herein, the term "send" represents any way in which information is communicated from one device, component, module or structure to another device, component, module or structure. The term "receive" represents any way in which information is obtained from another device, component, module or structure at a device, component, module or structure. Devices, components, modules or structures can be parts of the same computer system or different computer systems. Information can be passed by value (e.g., as parameters of a message or function call), or passed by reference (e.g., in a buffer). Depending on the context, information can be directly transmitted or communicated through one or more intermediate devices, components, modules or structures. As used herein, the term "connect" represents an operable communication link between a device, component, module or structure, which can be part of the same computer system or different computer systems. An operable communication link can be a wired or wireless network connection, which can be direct or pass through one or more intermediaries (e.g., a network).

[0048] The description of an example with several features does not imply that all or even any such features are required. Instead, various optional features are described to illustrate various possible examples of the innovations described herein. Unless expressly specified otherwise, no feature is essential or required.

[0049] Furthermore, although process steps and phases can be described in sequential order, such a process can be configured to work in a different order. The description of a specific sequence or order does not necessarily indicate that the steps / phases are required to be performed in that order. The steps or phases can be performed in any actual order. Furthermore, although described or implied as occurring non-simultaneously, some steps or phases can be performed simultaneously. Describing a process as including multiple steps or phases does not imply that all or even any steps or phases are necessary or required. Various other examples can omit some or all of the steps or phases described. Unless otherwise expressly specified, no step or phase is necessary or required. Similarly, although a product can be described as including multiple aspects, qualities, or characteristics, this does not mean that all of them are necessary or required. Various other examples can omit some or all aspects, qualities, or characteristics.

[0050] An enumerated listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise. Likewise, an enumerated listing of items does not imply that any or all of the items are comprehensive of any category, unless expressly specified otherwise.

[0051] For presentation purposes, the detailed description uses terms like "determine" and "select" to describe computer operations in a computer system. These terms represent operations performed by one or more processors or other components in a computer system and should not be confused with actions performed by a human. The actual computer operations corresponding to these terms vary depending on the implementation.

[0052] II. Accelerate the training of recommendation models.

[0053] A recommendation model can provide suggestions for interesting, promising, or relevant items for a user. The term item is general. These items can be books, television programs, movies, news stories, products, songs, people, other entities, restaurants, other locations, services, or any other type of item. In many implementations, a recommendation model is trained using machine learning based on ratings for a larger group of users for a larger selection of items. After training, the recommendation model can be used to identify new items relevant to a given user, or the recommendation model can be used to identify new users relevant to a given item. A recommendation model can also be referred to as a recommendation (or recommender) system, platform, or engine.

[0054] Typically, recommendation models use continuous features as well as categorical features. Continuous features are processed by a neural network layer, which can be a multilayer perceptron, a deep neural network ("DNN") layer, or another type of neural network layer. Continuous features can include, for example, weight values ​​for nodes, bias values ​​for nodes, and values ​​for activation functions in the DNN. Categorical features for recommendation models are part of an embedding set, which can be organized into embedding tables. For example, an embedding set includes categorical values ​​for users and items rated by users. The embedding set can be quite large, even if the actual ratings only sparsely populate the embedding set. Suppose the embedding set tracks the ratings of 100,000 users and 25,000 items. The number of ratings in the user × item matrix can be as high as 100,000 × 25,000 = 2,500,000,000, even though, in reality, the matrix is ​​only sparsely populated with actual rating values, which motivates the use of recommendation models. The embedding set can be decomposed into multiple embedding tables, such as one or more embedding tables storing feature vectors for users and one or more embedding tables storing feature vectors for items. The embedding table can be quite large, and its size increases as more users and items interact.

[0055] Training a recommendation model can be both computationally intensive (due to the neural network used for training operations) and memory intensive (due to the need to buffer large embedding tables). In some previous approaches, recommendation models are trained in a hybrid CPU-GPU mode. In the hybrid CPU-GPU mode, the CPU (with associated main memory) provides high memory capacity for the embedding set, and the GPU provides high throughput, data-parallel execution of the training operations for the neural network. In the hybrid CPU-GPU approach, the latency due to transfers between the CPU and GPU can be significant.

[0056] Accelerators can reduce this latency. In some approaches, accelerators take advantage of the property that in a typical training scenario only a few entries in the embedding set are frequently accessed (popular). For details on one such approach, see Adnan et al., “Heterogeneous Acceleration Pipeline for Recommendation System Training,” arXiv:2204.05436v1, published on April 11, 2022. In this approach, the accelerator uses (multiple) GPUs for all training operations and uses CPU-based main memory to store most embeddings, just like in the hybrid CPU-GPU approach. However, unlike the hybrid CPU-GPU approach, some embeddings are stored in GPU memory. In particular, GPU memory stores frequently accessed embeddings.

[0057] More specifically, the accelerator uses an access frequency-aware memory layout for the category features (embeddings) of the recommendation model. The accelerator exploits the insight that recommendation models are often trained on inputs with a very high access frequency bias. Frequently accessed entries tend to have a smaller memory footprint and are very important for the training process. During the learning phase, the accelerator identifies frequently accessed embeddings. For example, the accelerator takes small batches in its first epoch and uses these inputs to identify frequently accessed embeddings. For most datasets, training a small proportion of mini-batches can identify a large percentage of frequently accessed embeddings.

[0058] Subsequently, frequently accessed embedding entries are stored in GPU memory. Less frequently accessed embeddings are stored in CPU main memory. The accelerator can re-evaluate the access patterns of the embeddings from time to time to ensure that the classification follows the current trends in the training data. The accelerator itself only stores the indices of the frequently accessed embeddings, which the accelerator can use to classify small batches of input into micro-batches (“MBs”).

[0059] In the acceleration phase, the accelerator pipes MBs to (multiple) GPUs. Specifically, the accelerator operates on multiple MBs in the working set within a training iteration. The accelerator classifies and reorders the mini-batch input into different categories of MBs for the working set based on access frequency.

[0060] MBs in the first category (“popular” MBs) contain only inputs that are frequently accessed embeddings. The accelerator can schedule popular MBs directly on (multiple) GPUs. Popular MBs contain indices to the inputs; frequently accessed embeddings are already stored in GPU memory.

[0061] MBs in the second category ("unpopular" MBs) contain inputs that are not frequently accessed embeddings, and may also contain inputs that are frequently accessed embeddings. For unpopular MBs, the accelerator collects the required operating parameters from the CPU main memory (and from the GPU memory for any frequently accessed values ​​that are inputs to unpopular MBs) while executing the training operation using the popular MBs on the GPUs. Since the accelerator operation of collecting embeddings for unpopular MBs is performed concurrently with the GPU operation of using the popular MBs for training, the latency of the accelerator operation for collecting embeddings is effectively hidden (i.e., the additional latency for the operation of collecting embeddings is avoided).

[0062] The accelerator then schedules the unpopular MB for training on the GPU(s) and provides the input for the unpopular MB to the GPU(s). Thus, the GPU(s) perform training operations for the forward pass, backward pass, and optimization, regardless of whether the MB is a popular MB or an unpopular MB. While the GPU(s) perform training operations for the unpopular MB, the accelerator can begin identifying inputs for the next working set of MBs for the next mini-batch.

[0063] In some implementations, a given entry from the embedding set may appear as a frequently accessed value in multiple MBs in the same working set. In this case, if the frequently accessed value is updated after training with a given MB, subsequent training of a different MB (in the same working set) that also includes the frequently accessed value may be affected. This may introduce differences in the results of training with an accelerator compared to the training results in the "baseline" approach (without the accelerator pipeline, which splits the mini-batch input between the MBs in the working set and then pipelines the training operations for the corresponding MBs).

[0064] Furthermore, the same network parameters can be used for training different MBs in a working set. If one of the network parameters is updated after training with a given MB, subsequent training of another MB in the same working set may be affected. Again, this may introduce differences in the results of training with an accelerator compared to the training results in the "baseline" approach (without the accelerator pipeline, which splits the mini-batch input between the MBs in the working set and then pipelines the training operations for the corresponding MBs).

[0065] III. Training machine learning models using an accelerator pipeline with deferred updates to model parameters.

[0066] This section describes an innovation for training MLMs (such as recommendation models) using an accelerator pipeline with deferred updates to model parameters. In many use cases, resources can be used more efficiently during training operations due to the accelerator pipeline. By deferring updates to the model parameters, the differences in the results of the training process can be reduced or even eliminated compared to the results of training MLMs using a baseline approach (an accelerator pipeline that does not split mini-batches of inputs between micro-batches ("MBs") in a working set and then pipelines the training operations for the corresponding MBs). For example, by deferring updates to the network parameters, the training of (multiple) later MBs in the working set is not affected by changes to the network parameters from the training of (multiple) earlier MBs in the working set. As another example, assume that multiple MBs in a given working set can include the same frequently accessed value. By deferring updates to the frequently accessed value, even if multiple MBs in a given working set include the same frequently accessed value, the training of (multiple) later MBs in the working set is not affected by changes to the frequently accessed value from the training of earlier MBs in the working set.

[0067] In some examples described herein, the MLM is a recommendation model. Alternatively, the MLM can be a model trained for another use case, such as image recognition, speech recognition, image classification, object detection, facial recognition or other biometric recognition, emotion detection, question-answering ("chatbots"), natural language processing, automated language translation, query processing in search engines, automatic content selection, analysis of email and other electronic documents, relationship management, biomedical informatics, identification or screening of candidate biomolecules, generative adversarial networks, or other classification tasks.

[0068] An MLM has model parameters, such as network parameters and embedding sets. Typically, network parameters are parameters of the neural network or other model used to train the MLM. For example, network parameters include weight values ​​for hidden layer nodes of the DNN, bias values ​​for hidden layer nodes of the DNN, values ​​of activation functions for the DNN, and / or other parameters of the DNN that are adjusted during the training process. Typically, the embedding set provides input to the neural network or other model used to train the MLM, but the values ​​of the embedding set are also adjusted during the training process. Examples of network parameters and embeddings are described in the next section.

[0069] A. Sample architecture.

[0070] Figure 2 An example architecture (200) for training an MLM using an accelerator with deferred updates to model parameters is shown. The example architecture (200) includes one or more CPUs (210), a main memory (220) associated with the CPU(s) (210), a memory access engine (224), one or more GPUs (230a...230x) including GPU memories (232a...232x), a high-bandwidth switch matrix (238), a bus (240), and an accelerator (290). In general, a computer system having the example architecture (200) can train an MLM using operations in an acceleration pipeline with deferred updates to the MLM's model parameters.

[0071] In the example architecture (200), values ​​for an embedding set of MLMs are stored across main memory (220) and GPU memory (232a . . . 232x). Main memory (220) can store more values ​​for the embedding set, but accessing values ​​stored in main memory (220) is relatively slow. Accessing values ​​stored in main memory (220) via a memory access engine (224) is faster than accessing such values ​​via a CPU (s), but still relatively slow. CPU (s) (210), GPU (s) (230a . . . 230x), and accelerator (290) can transfer information across a bus (240), but transfers across the bus (240) are relatively slow.

[0072] In contrast, accessing values ​​stored in GPU memories (232a...232x) is relatively fast. GPU(s) (230a...230x) can access values ​​stored in their respective GPU memories (232a...232x). Moreover, GPU(s) (230a...230x) can access values ​​stored in other GPU memories (232a...232x) via the high bandwidth switch matrix (238), and such access is relatively fast. GPU memories (232a...232x) tend to be smaller and more expensive than main memory (220). Therefore, as Figure 2As shown, GPU memory (232a...232x) stores frequently accessed values ​​of the embedding set, while main memory (220) stores less frequently accessed values ​​of the embedding set. Frequently accessed values ​​of the embedding set can be distributed across GPU memories (232a...232x), with memories for different GPUs storing different portions (e.g., tables) of frequently accessed values. Even so, retrieving values ​​for training through the switch matrix (238) is very fast.

[0073] The memory access engine (224) implements operations for reading values ​​from the main memory (220) and writing values ​​to the main memory (220). The (multiple) CPUs (210) also implement operations for reading values ​​from the main memory (220) and writing values ​​to the main memory (220) as an alternative, slower path. The accelerators (290) and (multiple) GPUs (230a...230x) can update values ​​in the main memory (220) through the memory access engine (224) or the (multiple) CPUs (210). With respect to the training process, the (multiple) CPUs (210) can implement operations for selecting inputs for small batches from the embedding set for the working set. In doing so, the (multiple) CPUs (210) can apply various sampling strategies. Typically, the (multiple) CPUs apply a sampling strategy that is biased towards selecting certain values ​​that are more relevant to the training process, which has the effect of selecting values ​​stored in the GPU memory more frequently. The (multiple) CPUs (210) can also implement any other operations. (As explained below, the accelerator (290) sorts the mini-batch input into micro-batches ("MBs") of the working set.)

[0074] The example architecture (200) may include a single GPU, two GPUs, four GPUs, eight GPUs, or some other number of GPUs, depending on the implementation. The GPU(s) (230a...230x) implement operations for training the MLM. For example, for training the MLM using a given MB, the GPU(s) (230a...230x) implement forward operations in a neural network with inputs in the given MB, implement backpropagation operations for the inputs in the given MB, implement backpropagation operations for the neural network parameters, and implement operations for determining updates to the model parameters (e.g., optimization operations for determining gradients for the inputs in the given MB and determining gradients for the neural network parameters). The GPU(s) (230a...230x) are configured to perform training operations for different MBs in a continuous training of the MBs (one MB after another MB).

[0075] The GPU(s) (230a...230x) also implement operations for updating model parameters of the MLM. Thus, the GPU(s) (230a...230x) implement operations for updating network parameters and for updating values ​​of embedding sets stored in the GPU memory (232a...232x). In the example architecture (200), some updates to model parameters are deferred within the working set. For example, updates to network parameters are deferred until after the MLM is trained with the last MB of the working set. Or, as another example, updates to frequently accessed embeddings are deferred until after the MLM is trained with the last MB of the working set, or at least until these frequently accessed embeddings are no longer read for any later MBs of the working set. The deferred updates can be buffered in the GPU memory (232a...232x), thereby avoiding transferring the deferred updates to / from main memory.

[0076] When updates to model parameters are deferred within a working set, the GPU(s) (230a...230x) implement operations to aggregate any updates that affect the same model parameter. For example, for a given network parameter, the GPU(s) (230a...230x) may aggregate any updates to the given network parameter from training MLMs using different MBs in the working set. Or, as another example, for a given frequently accessed value of an embedding set, the GPU(s) (230a...230x) may aggregate any updates to the given frequently accessed value from training MLMs using different MBs in the working set.

[0077] The GPU(s) (230a...230x) further implement operations to provide updates to the infrequently accessed values ​​of the embedding set to the accelerator (290) after training is completed for the MB containing the infrequently accessed values ​​as input. Alternatively, the GPU(s) (230a...230x) may directly update the values ​​in the main memory (220).

[0078] The accelerator (290) implements an operation for identifying MBs of a working set. Specifically, the accelerator (290) implements an operation for sorting mini-batches of inputs for the working set into corresponding MBs. A mini-batch contains multiple entries from the embedding set, but less than all entries of the embedding set. A MB contains a subset of entries in the mini-batch. The number of MBs in the working set depends on the implementation. For example, the number of MBs in the working set is 2, 4, 8, or some other number of MBs greater than 2.

[0079] The MB may be a first type MB that includes only frequently accessed values ​​of the embedding set as input, and these frequently accessed values ​​are stored in the GPU memory (232a...232x). Alternatively, the MB may be a second type MB that includes infrequently accessed values ​​of the embedding set as input, and these infrequently accessed values ​​are stored in the main memory (220), and may also include frequently accessed values ​​of the embedding set as input, and these frequently accessed values ​​are stored in the GPU memory (232a...232x).

[0080] In a typical usage scenario, a working set includes one or more first-class MBs and one second-class MB. Thus, for example, a working set may include one first-class MB, three first-class MBs, seven first-class MBs, or some other number of first-class MBs. The number of first-class MBs in a working set depends on the number of inputs in the working set and also on how the inputs selected for the working set are classified.

[0081] In some example implementations, to hide the latency associated with identifying MBs of a working set, the accelerator ( 290 ) may identify MBs during training of the MLM using the last MB of a previous working set.

[0082] The accelerator (290) also implements operations for scheduling (multiple) first-class MBs for use in training MLMs using (multiple) GPUs (230a...230x). The accelerator (290) provides (multiple) first-class MBs to the (multiple) GPUs (230a...230x). In doing so, frequently accessed values ​​in the (multiple) first-class MBs are already in the GPU memory (232a...232x). The inputs to the (multiple) first-class MBs are represented as indices in the (multiple) first-class MBs scheduled by the accelerator (290), where each index in the indices refers to an entry in an embedding set stored in the GPU memory (232a...232x).

[0083] The accelerator (290) also implements operations for obtaining second-class MBs during at least a portion of the training of the MLM using (multiple) first-class MBs. For example, the accelerator (290) implements read operations to request and receive inputs from frequently accessed values ​​stored in the GPU memory (232a...232x) for the second-class MBs via the GPU (230a...230x), and the accelerator (280) implements read operations to request and receive inputs from infrequently accessed values ​​stored in the main memory (220) via the memory access engine (224) or (multiple) CPUs (210). The read operations for the second-class MBs can be performed during at least some of the training operations for training the MLM using (multiple) first-class MBs to hide the latency associated with the read operations and the operations of combining the inputs to obtain the second-class MBs. (The accelerator (290) also implements operations for combining the inputs received from the read operations.)

[0084] The accelerator (290) also implements operations for scheduling second-class MBs for training MLMs using GPUs (multiple). The accelerator (290) provides the second-class MBs to the GPUs (230a...230x). In doing so, the accelerator (290) passes actual values ​​for the inputs. Each of the inputs to the second-class MBs is an entry from an embedding set retrieved from GPU memory (232a...232x) or main memory (220).

[0085] After training with the second class MBs is completed, the model parameters of the MLM are updated. As explained above, at least some updates to the model parameters from training with the working set of (multiple) first class MBs are postponed until after the MLM is trained with the working set of second class MBs. This may include updates to the frequently accessed values ​​of the network parameters and embedding set by the (multiple) GPUs (230a...230x). After training the MLM with the second class MBs, the accelerator (290) may implement operations to update the less frequently accessed values ​​stored in the main memory (220) for the second class MBs. For example, the accelerator (290) implements a write operation to write the values ​​stored in the main memory (220) through the memory access engine (224) or the (multiple) CPUs (210). Alternatively, the (multiple) GPUs (230a...230x) may directly update the values ​​in the main memory (220).

[0086] B. Example machine learning model.

[0087] Figure 3a An example embedding set (300) for MLM is shown. Figure 3a In the example of MLM, the MLM is a recommendation model. The embedding set (300) for the MLM includes values ​​organized along a first dimension for m users and a second dimension for n items. The number of users, m, depends on the usage scenario and can be hundreds, thousands, or even millions of users. The number of items, n, also depends on the usage scenario and can be hundreds, thousands, or even millions of items. These items can be books, TV shows, movies, news stories, products, songs, people, other entities, restaurants, other locations, services, or any other type of item.

[0088] exist Figure 3a, the embedding set (300) is organized into a plurality of embedding tables (301, 302). The first embedding table (301) stores a vector of k weights for each user. Thus, the first embedding table (301) is an m×k table. The second embedding table (302) stores a vector of k weights for each item. Thus, the second embedding table (302) is an n×k table. Typically, the product of the transpose of the first embedding table (301) and the second embedding table (302) indicates the (estimated) rating matrix for the corresponding items for the corresponding users according to the trained MLM: (m×k)(k×n)=m×n, where each entry of m×n indicates an estimated rating for a given user and a given item.

[0089] The number k of weights (also called features) depends on the implementation and affects the complexity of the MLM. For smaller values ​​of k, the MLM is simpler. For larger values ​​of k, the MLM is more complex. For example, k can be 16, 32, 64, 128, or some other number of weights per vector.

[0090] In practice, embedding tables can be extremely large, totaling billions or even terabytes. Large embedding tables can be split into multiple smaller tables. For example, the large example of embedding table (301) can be split into multiple embedding tables for users, while the large example of embedding table (302) can be split into multiple embedding tables for projects.

[0091] Figure 3b The topology of an example deep neural network ("DNN") (350) that can be used when training an MLM is shown. Typically, a DNN operates in at least two different modes. The DNN is trained in training mode and then used as a classifier in inference mode. During training, examples in a training dataset (here, the embedding set (300)) are applied as input to the DNN, and various network parameters of the DNN are adjusted so that when training is complete, the DNN can be used as an effective classifier.

[0092] Training is performed in iterations, with each iteration using multiple examples (in mini-batches) of the embedding set (300). For iterations in conventional methods using mini-batch sampling, training typically includes performing a forward propagation of the input, computing a loss (e.g., determining the difference between the output of the DNN and the expected output for a given input), and performing a backpropagation through the DNN to adjust the network parameters of the DNN (e.g., weights and biases) and the values ​​of the embedding set used as input. In the methods described herein, where a mini-batch is split into multiple MBs of a working set for training, training for a given MB includes forward propagation of the input of the given MB, computing a loss, and determining updates to the network parameters and the values ​​of the embedding set used as input, but the updates are postponed until after training has been performed using the last MB of the working set. When the parameters of the DNN are suitable for classifying the training data, the parameters converge, and the training process can be complete. After training, the DNN can be used in an inference mode, where one or more examples are applied to the DNN as input and forward propagated through the DNN so that the DNN can classify the example(s).

[0093] exist Figure 3b In the example topology (350) shown, a first set of nodes (360) forms the input layer. A second set of nodes (370) forms the first hidden layer. A second hidden layer is formed from a third set of nodes (380), and an output layer is formed from a fourth set of nodes (390). More generally, the topology of a DNN can have more hidden layers. (DNNs conventionally have at least two hidden layers.) The input layer, hidden layers, and output layer can have more or fewer nodes than in the example topology (350), and different layers can have the same node count or different node counts. Hyperparameters (such as the number of hidden layers and the node counts for the respective layers) can define the overall topology.

[0094] Nodes in a given layer may provide input to each node in one or more nodes in a later layer or to fewer than all nodes in a later layer. In the example topology (350), each node in a given layer is fully interconnected to nodes in each adjacent layer in the adjacent layer(s). For example, each node in the first node set (360) (input layer) is connected to and provides input to each node in the second node set (370) (first hidden layer). Each node in the second node set (370) (first hidden layer) is connected to and provides input to each node in the third node set (380) (second hidden layer). Finally, each node in the third node set (380) (second hidden layer) is connected to and provides input to each node in the fourth node set (390) (output layer). Thus, a layer may include nodes that have common inputs with other nodes in the layer and / or provide outputs to a common destination with other nodes in the layer. More generally, a layer may include nodes that have a common subset of inputs with other nodes in the layer and / or provide outputs to a common subset of destinations with other nodes in the layer. Therefore, nodes at a given layer do not need to be interconnected to every node at an adjacent layer.

[0095] Typically, during forward propagation, a node produces an output by applying a weight to each input from the previous layer of nodes and collecting the weighted input values ​​to produce an output value. Each individual node may have an activation function and / or an applied bias. For example, for node n in the hidden layer, the forward function f() may produce an output that is mathematically expressed as:

[0096]

[0097] where variable E is the count of connections (edges) providing input to the node, variable b is the bias value for node n, function σ() represents the activation function for node n, and variable x i and w i are the input value and weight value of one of the connections from the previous layer node. For each connection (edge) that provides input to node n, the input value x i Multiply by the weight w i The products are added together, a bias value b is added to the sum of the products, and the resulting sum is input to an activation function σ(). In some implementations, the activation function σ() produces a continuous value (represented as a floating point number) between 0 and 1. For example, the activation function is a sigmoid function. Alternatively, the activation function σ() produces a binary 1 or 0 value depending on whether the sum is above or below a threshold.

[0098] A neural network can be trained and retrained by adjusting the parameters that make up the output function f(n). For example, by adjusting the weights w for the corresponding nodesi And bias value b, adjust the behavior of the neural network. During backpropagation, a cost function C(w,b) can be used to find appropriate weights and biases for the network, where the cost function can be mathematically described as:

[0099]

[0100] Where the variables w and b represent weights and biases, the variable m is the number of training inputs, and the variable a is a vector of output values ​​from the neural network for the input vector y(x) of expected outputs (labels) for examples from the training data. By adjusting the network weights and biases, the cost function C can be driven to a target value (e.g., to zero) using various search techniques (such as stochastic gradient descent). When the cost function C is driven to the target value, the neural network is said to have converged.

[0101] although Figure 3b The example topology (350) is for a non-recursive DNN, but the tools described herein can be used for other types of neural networks, including recursive neural networks or other artificial neural networks.

[0102] C. Example operations in the pipeline.

[0103] Figure 4a and 4b The general timing of operations (401, 402) in an accelerator pipeline with deferred updates to model parameters when training an MLM is shown. The operations (401, 402) are split between the accelerator, one or more GPUs, and a memory access engine, as generally indicated by the dashed lines splitting the operations (401, 402).

[0104] exist Figure 4a In the illustrated operation (401), the accelerator identifies (410) a working set of MBs, for example, by classifying a mini-batch of inputs into the working set of MBs. The working set includes a first type of MB and a second type of MB. The accelerator schedules the first type of MB using (multiple) GPUs.

[0105] The GPU(s) train (420) the MLM using the first class MBs. For training, the GPU(s) read (422) frequently accessed values ​​from GPU memory, which are inputs for the first class MBs. However, the GPU(s) defer updates from training using the first class MBs. Instead of applying the updates, the GPU(s) cache the deferred updates in GPU memory.

[0106] Concurrently, the accelerator obtains (426) second-class MBs. To do so, the accelerator reads (424) any frequently accessed values ​​that are input to the second-class MBs from the GPU memory using a read operation on the GPU(s), and the accelerator reads (428) infrequently accessed values ​​that are input to the second-class MBs from the main memory using a read operation on the memory access engine. The accelerator combines the retrieved inputs for the second-class MBs. Although retrieving and combining the inputs for the second-class MBs can be time-consuming, by overlapping the training (420) of the GPU(s), the delay for such operations is at least partially eliminated. The accelerator then schedules the second-class MBs on the GPU(s), thereby providing the second-class MBs to the GPU(s).

[0107] The GPU(s) train (480) the MLM using the second type of MBs. For training, the GPU(s) use values ​​provided by the second type of MBs. During the training (480) using the second type of MBs, deferred updates from the training (420) using the first type of MBs are buffered in GPU memory.

[0108] After training is complete, the GPU(s) update (490) the model parameters for the MLM. To do this, the GPU(s) write values ​​(492) to the GPU memory. For example, the GPU(s) update network parameters, which may be buffered in the GPU memory, and the GPU(s) update frequently accessed values ​​stored in the GPU memory. In doing so, the GPU(s) apply updates from training (480) using the second type of MBs and also apply deferred updates from training (420) using the first type of MBs.

[0109] The GPU(s) also provide updates to the accelerator for infrequently accessed values ​​that are inputs to the second type of MBs. For example, the GPU(s) write the updates to GPU memory, and the accelerator then reads the updated values ​​from GPU memory using a read operation on the GPU(s). Alternatively, the GPU(s) provide updates to infrequently accessed values ​​in some other manner.

[0110] The accelerator obtains (496) updates for the infrequently accessed values ​​and applies the updates to the infrequently accessed values ​​stored in the main memory. To this end, the accelerator writes (498) the updated infrequently accessed values ​​for the inputs in the second type MB to the main memory using a write operation to the memory access engine.

[0111] Figure 4b Many of the operations shown are similar to Figure 4a The corresponding operations shown are the same. Figure 4aThe training and update operations are shown for a working set with one first class MB, however, Figure 4b The working set in has multiple first-class MBs.

[0112] exist Figure 4b In the illustrated operation (402), the accelerator identifies (410) different working sets of MBs, for example, by classifying a mini-batch of inputs into different working sets of MBs. The working set includes a plurality of first-class MBs and second-class MBs. The accelerator schedules the first-class MBs using (multiple) GPUs.

[0113] The GPU(s) train (420) an MLM using a first MB (MB1) in a first class of MBs. For training, the GPU(s) read (422) frequently accessed values ​​from GPU memory, which are inputs for the first MB (MB1). However, the GPU(s) defer updates from training using the first MB (MB1). Instead of applying the updates, the GPU(s) cache the deferred updates in GPU memory.

[0114] The GPU(s) then train (440) the MLM using a second MB (MB2) from the first class of MBs. For training, the GPU(s) read (442) frequently accessed values ​​from GPU memory, which are inputs for the second MB (MB2). However, the GPU(s) defer updates from training using the second MB (MB2). Instead of applying the updates, the GPU(s) cache the deferred updates in GPU memory.

[0115] The GPU(s) may similarly perform additional training (460) for other MBs in the first class of MBs while continuing to train the other MBs on a per-MB basis. For such training, the GPU(s) may perform additional read (462) operations from the GPU memory to obtain frequently accessed values ​​that are inputs for the other first class of MBs.

[0116] Concurrently with at least a portion of the training (420, 440, 460) of the GPU(s), the accelerator obtains (426) second-class MBs. To do so, the accelerator reads (424) any frequently accessed values ​​as inputs for the second-class MBs from GPU memory using a read operation on the GPU(s), and the accelerator reads (428) less frequently accessed values ​​as inputs for the second-class MBs from main memory using a read operation on the memory access engine. The accelerator combines the retrieved inputs for the second-class MBs. By overlapping the training (420, 440, 460) of the GPU(s), the latency of such operations is at least partially negated. The accelerator then schedules the second-class MBs on the GPU(s), thereby providing the second-class MBs to the GPU(s).

[0117] The GPU(s) train (480) the MLM using the second class of MBs. For training, the GPU(s) use values ​​provided using the second class of MBs. During other later training sessions for the working set, deferred updates from training (420, 440, 460) using the first class of MBs are buffered in GPU memory. For example, during training (440, 460, 480) using later MBs in the working set, updates from training (420) using the first MB (MB1) of the first class of MBs are buffered in GPU memory. Similarly, during training (460, 480) using later MBs in the working set, updates from training (440) using the second MB (MB2) in the first class of MBs are buffered in GPU memory.

[0118] After training is complete, the GPU(s) update (490) the model parameters for the MLM. To do this, the GPU(s) write values ​​(492) to the GPU memory. For example, the GPU(s) update network parameters, which may be buffered in the GPU memory, and the GPU(s) update frequently accessed values ​​stored in the GPU memory. In doing so, the GPU(s) apply updates from training (480) using the second type of MBs and also apply deferred updates from training (420, 440, 460) using the first type of MBs.

[0119] The GPU(s) also provide updates to the accelerator for infrequently accessed values ​​that are inputs to the second type of MBs. For example, the GPU(s) write the updates to GPU memory, and the accelerator then reads the updated values ​​from GPU memory using a read operation on the GPU(s). Alternatively, the GPU(s) provide updates to the infrequently accessed values ​​in some other manner.

[0120] The accelerator obtains (496) updates for the infrequently accessed values ​​and applies the updates to the infrequently accessed values ​​stored in the main memory. To this end, the accelerator writes (498) the updated infrequently accessed values ​​for the inputs in the second type MB to the main memory using a write operation to the memory access engine.

[0121] D. Example accelerator operations and GPU operations.

[0122] Figure 5 A general technique (500) for training an MLM using an accelerator pipeline with deferred updates to model parameters is shown from the accelerator's perspective. Figure 2 Or otherwise described, a computer system implementing an accelerator for training MLM using one or more GPUs can perform the technique (500). Specifically, the accelerator of such a computer system performs Figure 5 Accelerator operation shown.

[0123] on the contrary, Figure 6 A general technique (600) for training an MLM using an accelerator pipeline with deferred updates to model parameters is shown from the perspective of one or more GPUs. Figure 2 Or otherwise described, a computer system implementing an accelerator for training MLM using one or more GPUs can perform the technique (600). Specifically, the GPU(s) of such a computer system perform Figure 6 The GPU operations shown.

[0124] The MLM has model parameters. For example, the model parameters include network parameters and embedding sets. The network parameters can be neural network parameters, such as weight values ​​for nodes, bias values ​​for nodes, or values ​​of activation functions for neural networks. The embedding set includes frequently accessed values ​​stored in GPU memory and infrequently accessed values ​​stored in main memory associated with the CPU. For example, for a recommendation model, the embedding set includes values ​​organized along a first dimension (for users) and a second dimension (for items). The embedding set can be organized into multiple embedding tables, such as an embedding table with feature vectors for users and an embedding table with feature vectors for items, where the values ​​(feature vectors) for some users and items are frequently accessed as input for the MB, while the values ​​(feature vectors) for other users and items are not frequently accessed as input for the MB.

[0125] The number of GPUs used for training in a computer system depends on the implementation. For example, the computer may use a single GPU, two GPUs, four GPUs, eight GPUs, or some other number of GPUs.

[0126] Figure 5 and 6 The operations performed on the working set of MBs are shown. The mini-batch input is split among the MBs in the working set. The working set includes one or more first-class MBs (which only contain frequently accessed values ​​as input) and one or more second-class MBs (which contain infrequently accessed values ​​as input, but may also contain frequently accessed values ​​as input). Typically, a mini-batch includes multiple entries from the embedding set, but less than all embedding sets. An MB (whether a first-class MB or a second-class MB) contains a subset of the entries in the mini-batch. For a recommendation model, the working set of MBs typically includes a small portion of the embeddings in the embedding set.

[0127] The number of MBs in a working set depends on the implementation. For example, a working set includes a single first-class MB, three first-class MBs, seven first-class MBs, or some other number of first-class MBs. Typically, a working set includes a single second-class MB, but a working set may alternatively include multiple second-class MBs.

[0128] This can be repeated for one or more subsequent working sets of MB for different mini-batches in different training iterations. Figure 5 and Figure 6 For example, the accelerator and GPU(s) may repeat the operations for subsequent working sets of the MB until the epoch for the embedding set is complete, at which point each entry of the embedding set has been processed in training. The accelerator and GPU(s) may further repeat the operations in the working set of the MB for one or more additional epochs for the embedding set, e.g., until the model parameters converge.

[0129] refer to Figure 5 The accelerator identifies (510) one or more first-class MBs and second-class MBs for a given working set. Each of the first-class MBs in the plurality of first-class MBs contains only inputs of frequently accessed values ​​from the embedding set. The frequently accessed values ​​are stored in GPU memory. The second-class MBs contain inputs of infrequently accessed values ​​from the embedding set. The second-class MBs may also contain inputs from frequently accessed values ​​stored in GPU memory. The infrequently accessed values ​​are stored in main memory associated with at least one CPU.

[0130] For example, the accelerator identifies (e.g., receives or otherwise determines) a set of inputs for a mini-batch of the working set. The accelerator classifies the corresponding inputs of the mini-batch as frequently accessed values ​​or infrequently accessed values. The accelerator reorders the mini-batch inputs into MBs such that (multiple) first-class MBs contain only frequently accessed values ​​as inputs, and the second-class MBs contain the remaining inputs (including at least some infrequently accessed values). Typically, the mini-batch inputs for the working set are selected from the embedding set according to a sampling strategy that is biased towards selecting certain values ​​(frequently accessed values), which is why such values ​​are stored in GPU memory.

[0131] The accelerator can pipeline the operation of identifying MBs of a working set. For example, while training an MLM using the second-class MBs of a previous working set, the accelerator can identify the first-class MBs and the second-class MBs of the current working set (e.g., identifying mini-batches of inputs and classifying the mini-batches of inputs into MBs of the current working set).

[0132] The accelerator schedules (520) the first class MBs for training the MLM using the GPUs. The inputs to the first class MBs are represented as indices into the first class MBs scheduled by the accelerator. That is, the accelerator does not pass actual values ​​to the GPUs. Instead, each index in the index refers to an entry in the embedding set.

[0133] refer to Figure 6 In response to scheduling one or more first-class MBs of a given working set for training of the MLM, the GPU(s) sequentially (on a MB-by-MB basis) execute (610) training operations using the first-class MB(s). Each of the first-class MB(s) contains only inputs of frequently accessed values ​​from the embedding set, which are stored in GPU memory for the GPU(s). The GPU(s) can use indexes to the inputs in the first-class MB(s) to retrieve the actual frequently accessed values ​​from the GPU memory.

[0134] For example, as part of training the MLM using a given MB in the first category of MBs, the GPU(s) perform a forward operation in the neural network using input from the given MB, perform a backpropagation operation on the input (embedded value) in the given MB, perform a backpropagation operation on the neural network parameters, and determine (but not yet apply) updates to the model parameters, including updates to the input in the given MB and updates to the neural network parameters. When the first category of MBs includes a plurality of first category MBs, in continuously training the plurality of first category MBs on an MB-by-MB basis, the GPU(s) perform training operations on the plurality of first category MBs, respectively.

[0135] The GPU(s) defer at least some updates to model parameters from training with the first class MB(s) until after the MLM is trained with the second class MB(s) of the given working set. For example, the GPU(s) defer updates to network parameters from training with the first class MB(s). In this manner, training of later MBs in the first class MB(s) and the second class MB(s) is not affected by network parameter changes from training of earlier MBs in the first class MB(s) in the same working set, which could introduce discrepancies (compared to a baseline training approach without an accelerator pipeline that splits mini-batches of inputs between MBs in the working set).

[0136] As another example, the GPU(s) may defer updates to frequently accessed values ​​stored in GPU memory from training of the first class MB(s). In some example implementations, a given entry from the embedding set may appear as a frequently accessed value in multiple MBs of the same working set. By deferring updates to the frequently accessed values, training of later MBs in the first class MB(s) and the second class MB(s) is not affected by changes in frequently accessed values ​​from training of earlier MBs in the first class MB(s) in the same working set, which may introduce differences (compared to a baseline training approach without an accelerator pipeline that splits mini-batches of inputs between MBs in the working set). On the other hand, if the inputs in the MBs of a given working set are guaranteed to be disjoint (i.e., no entry from the embedding set can appear in different MBs of a given working set), the GPU(s) may immediately apply updates to the frequently accessed values ​​stored in GPU memory without affecting later training using the MBs of the given working set.

[0137] Or, as another example, the GPU(s) may defer updates to network parameters and updates to frequently accessed values ​​stored in GPU memory.

[0138] During at least part of the time of training the MLM with the first class MB(s), the accelerator obtains (530) the second class MBs. In terms of timing, as described above, at least some updates to the model parameters from training with the first class MBs are deferred until after the MLM is trained with the second class MBs.

[0139] For example, the accelerator requests any input for the second type of MB from the frequently accessed values ​​stored in the GPU memory from the GPU(s), and receives such input from the GPU(s). The accelerator also requests any input for the second type of MB from the infrequently accessed values ​​stored in the main memory from the CPU or main memory engine, and receives such input. The accelerator performs the operation of requesting / receiving input for the second type of MB during at least some operations of training the MLM using the first type of MB(s), which effectively avoids additional latency that would otherwise be associated with the operation of requesting / receiving input for the second type of MB. The accelerator combines the input received from the frequently accessed values ​​stored in the GPU memory (if any) with the input received from the infrequently accessed values ​​stored in the main memory.

[0140] The accelerator schedules (540) the second type of MBs for training the MLM using the GPU(s). The inputs of the second type of MBs are represented as actual values ​​in the second type of MBs scheduled by the accelerator. That is, each of the inputs of the second type of MBs is an entry from an embedding set retrieved from the GPU memory or the main memory.

[0141] In response to scheduling the second type of MB for use in training the MLM, the GPU(s) perform (630) a training operation using the second type of MB. For example, as part of training the MLM using the second type of MB as a given MB, the GPU(s) perform a forward operation in the neural network using inputs from the given MB, perform a backpropagation operation on the inputs (embedded values) from the given MB, perform a backpropagation operation on the neural network parameters, and determine updates to the model parameters, including updates to the inputs from the given MB and updates to the neural network parameters.

[0142] After training the MLM using the second class of MBs, the GPU(s) provide (640) updates to the infrequently accessed values ​​stored in main memory for the second class of MBs and also update (650) network parameters and frequently accessed values ​​stored in GPU memory. In doing so, the GPU(s) apply at least some of the deferred updates.

[0143] For example, the GPU(s) update network parameters (such as neural network parameters) in the model parameters. When updates to the network parameters from training with the first type of MBs are deferred, the GPU(s) apply the deferred updates to the network parameters. In the event that a given one of the network parameters is affected by multiple updates, the GPU(s) may aggregate any updates from MLM training with the first type of MBs and MLM training with the second type of MBs for the given network parameter.

[0144] As another example, the GPU(s) update frequently accessed values ​​in the model parameters (used as input to the MBs) stored in the GPU memory. When updates to the frequently accessed values ​​stored in the GPU memory from training with the first type of MBs are deferred, the GPU(s) apply the deferred updates to the frequently accessed values ​​stored in the GPU memory. In the event that a given one of the frequently accessed values ​​is affected by multiple updates, the GPU(s) may aggregate any updates from the MLM training with the first type of MBs and the MLM training with the second type of MBs for the given frequently accessed value.

[0145] After training the MLM with the second type of MBs, the accelerator updates ( 550 ) the infrequently accessed values ​​stored in the main memory for the second type of MBs.

[0146] In view of the many possible embodiments to which the principles of the disclosed invention may be applied, it should be recognized that the illustrated embodiments are merely preferred examples of the invention and should not be construed as limiting the scope of the invention. Rather, the scope of the invention is defined by the following claims. Therefore, all that falls within the scope and spirit of these claims is hereby claimed as the present invention.

Claims

1. A computer system (200), comprising: at least one graphics processing unit ("GPU") (230a...230x) configured to train a machine learning model ("MLM") having model parameters, the model parameters including network parameters and an embedding set; GPU memory (232a...232x) configured to store frequently accessed values ​​of the embedding set; a main memory (220) associated with at least one central processing unit ("CPU") (210), said main memory (220) being configured to store infrequently accessed values ​​of said embedding set; as well as An accelerator (290) is configured to perform an accelerator operation, the accelerator operation comprising: identifying (510) one or more first-class micro-batches ("MBs") and second-class MBs for a given working set, wherein each of the one or more first-class MBs contains input only from the frequently accessed values ​​stored in the GPU memory, and wherein the second-class MBs contain input from the infrequently accessed values ​​stored in the main memory; scheduling (520) the one or more first-category MBs for training of the MLM using the at least one GPU; obtaining (530) MBs of the second class during at least part of the training of the MLM with the one or more MBs of the first class, at least some updates to the model parameters from the training with the one or more MBs of the first class being deferred until after the training of the MLM with the MBs of the second class; and The second type of MBs are scheduled (540) using the at least one GPU for training of the MLM.

2. The computer system of claim 1 , wherein the at least one GPU is configured to update the network parameters, and wherein With respect to the at least some updates to the model parameters being deferred, updates to the network parameters are deferred until after the training of the MLM with the second class of MBs.

3. The computer system according to claim 2, wherein: For a given one of the network parameters, the at least one GPU is configured to aggregate any updates from the training of the MLM with the one or more first-type MBs and the training of the MLM with the second-type MBs.

4. The computer system of claim 1 , wherein the at least one GPU is configured to update the frequently accessed values ​​stored in the GPU memory, and wherein, For the deferred at least some updates to the model parameters, updating of the frequently accessed values ​​stored in the GPU memory for the one or more first class MBs is deferred until after the training of the MLM with the second class MBs.

5. The computer system according to claim 4, wherein: For a given one of the frequently accessed values ​​stored in the GPU memory, the at least one GPU is configured to aggregate any updates from the training of the MLM using the one or more first-class MBs and the training of the MLM using the second-class MBs.

6. The computer system according to any one of claims 1 to 5, wherein the accelerator operation further comprises: After the training of the MLM with the second class of MBs, the infrequently accessed values ​​stored in the main memory for the second class of MBs are updated.

7. The computer system of any one of claims 1 to 6, wherein the second type of MB further includes input from the frequently accessed values ​​stored in the GPU memory, and wherein obtaining the second type of MB comprises: For the second type of MB, requesting and receiving the input from the frequently accessed values ​​stored in the GPU memory; requesting and receiving the input from the infrequently accessed values ​​stored in the main memory for the second class of MBs, wherein the requesting and receiving the input from the infrequently accessed values ​​stored in the main memory for the second class of MBs occurs during at least some training operations of the training of the MLM utilizing the one or more first class of MBs; as well as The input received from the frequently accessed values ​​stored in the GPU memory is combined with the input received from the infrequently accessed values ​​stored in the main memory.

8. The computer system of any one of claims 1 to 7, wherein identifying the one or more first-category MBs and the second-category MBs of the given working set comprises during the training of the MLM utilizing second-category MBs of a previous working set: identifying mini-batches of inputs for the given working set; and The inputs in the mini-batch are classified into the one or more first-category MBs and the second-category MBs.

9. The computer system according to claim 1 , wherein the network parameters are neural network parameters including weight values ​​for nodes, bias values ​​for nodes, and / or values ​​for activation functions, and wherein the training of the MLM comprises training a given MB among the one or more first-category MBs and the second-category MBs: performing a forward operation in a neural network using the input in the given MB; performing a back-propagation operation on the input in the given MB; performing a backpropagation operation on the neural network parameters; and The update to the model parameters is determined, the update to the model parameters comprising an update to the input in the given MB and an update to the neural network parameters.

10. The computer system of claim 1 , wherein the one or more first-class MBs consist of a single first-class MB, three first-class MBs, or seven first-class MBs, and wherein each of the one or more first-class MBs and the second-class MBs includes a plurality of entries from a mini-batch, the mini-batch containing a plurality of entries from the embedding set but less than all entries in the embedding set.

11. A computer system according to any one of claims 1 to 7, wherein the MLM is a recommendation model, wherein the embedding set includes values ​​organized along a first dimension and a second dimension, wherein the first dimension is users, wherein the second dimension is items, and wherein the embedding set is organized into a plurality of embedding tables.

12. The computer system of any one of claims 1 to 7, wherein identifying the one or more first-category MBs and the second-category MBs of the given working set comprises: identifying a mini-batch of inputs for the given working set, the inputs for the mini-batch being selected from the embedding set according to a sampling strategy that biases selection of the frequently accessed values ​​stored in the GPU memory; as well as The input of the mini-batch is classified into the one or more first-category MBs and the second-category MBs.

13. A method of performing accelerator operations, in a computer system implementing an accelerator for training a machine learning model ("MLM"), the method comprising: identifying (510) one or more first-class mini-batches ("MBs") and second-class MBs for a given working set, wherein the MLM has model parameters including network parameters and an embedding set, wherein each of the one or more first-class MBs contains inputs only of frequently accessed values ​​from the embedding set, the frequently accessed values ​​being stored in a graphics processing unit ("GPU") memory, and wherein the second-class MBs contain inputs of infrequently accessed values ​​from the embedding set, the infrequently accessed values ​​being stored in a main memory associated with at least one central processing unit ("CPU"); scheduling (520) the one or more first-category MBs for training of the MLM using the at least one GPU; obtaining (530) MBs of the second class during at least part of the training of the MLM with the one or more MBs of the first class, at least some updates to the model parameters from the training with the one or more MBs of the first class being deferred until after the training of the MLM with the MBs of the second class; scheduling (540) the second type of MBs for training of the MLM using the at least one GPU; as well as After the training of the MLM with the second class of MBs, the infrequently accessed values ​​stored in the main memory for the second class of MBs are updated (550).

14. One or more computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions being operable, when programmed, to cause at least one graphics processing unit ("GPU") to perform GPU operations, the GPU operations comprising: In response to scheduling one or more first-class micro-batches ("MBs") of a given work set for training a machine learning model ("MLM"), continuously performing (610) training operations using the one or more first-class MBs, wherein the MLM has model parameters including network parameters and an embedding set, each of the one or more first-class MBs containing inputs only of frequently accessed values ​​from the embedding set, the frequently accessed values ​​being stored in GPU memory; deferring at least some updates to the model parameters from the training with the one or more first class MBs until after training of the MLM with the second class MBs of the given working set; In response to scheduling the second class of MBs for training of the MLM, performing (630) a training operation using the second class of MBs, the second class of MBs containing inputs of infrequently accessed values ​​from the embedding set, the infrequently accessed values ​​stored in a main memory associated with at least one central processing unit ("CPU"); as well as After the training of the MLM with the second type of MBs: providing (640) an update of the infrequently accessed value stored in the main memory for the second class of MBs; as well as The network parameters and the frequently accessed values ​​stored in the GPU memory are updated (650), including applying the deferred at least some of the updates.

15. The one or more computer-readable media of claim 14, wherein the at least one GPU is configured to: Update the network parameters, wherein, for said at least some updates being deferred, updating of said network parameters is deferred until after said training of said MLM with said second class of MBs; and / or Updating the frequently accessed values ​​stored in the GPU memory, wherein, for the at least some of the deferred updates, updating the frequently accessed values ​​for the one or more first class MBs is deferred until after the training of the MLM with the second class MBs.