Method and system for embedding via intelligent memory pool
Through the intelligent memory pooling system, multiple interconnected electronic cards are used for parallel processing and distributed embedding operations, solving the memory capacity and bandwidth challenges of the embedding layer in deep learning applications, and achieving high-performance and high-throughput deep learning computing.
Patent Information
- Application Number
- CN202080103779.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-10-21
AI Technical Summary
The prior art faces challenges in memory capacity and bandwidth when dealing with embedded layers in deep learning applications, resulting in system-level performance bottlenecks.
An intelligent memory pool system was designed to utilize multiple interconnected electronic cards through a new hardware architecture and accompanying software tools, each electronic card containing programmable devices and memory cards to achieve parallel processing and distributed embedding operations.
The system can effectively expand memory capacity and bandwidth, reduce the computing workload of the host, reduce data exchange bandwidth, and improve the performance and throughput of deep learning jobs.
Smart Images

Figure CN116171437B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to machine learning and system architecture. More specifically, the present disclosure relates to methods and systems for embedding within a smart memory pool. Background Art
[0002] In recent years, machine learning algorithms (e.g., deep learning (DL)) based on deep neural networks (DNNs) have been rapidly expanding. To meet the computational demands of DNNs, general-purpose graphics processing units (GPGPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and neural processing units (NPUs) have been widely studied and deployed to accelerate machine learning workloads. However, recent studies on multiple superscalars have shown that system-level challenges are imminent in emerging DL applications, which are very memory-intensive (in terms of capacity and bandwidth). These studies further position the embedding layer, which provides input data to DNNs and pre-processes the input data, as one of the most memory-intensive layers in DL applications. Due to the exponential data growth in superscalars, the embedding layer typically needs to process tens to hundreds of terabytes of data.
[0003] Some existing technologies have proposed the idea of aggregating a pool of local memory modules within a device-side interconnect (e.g., “device” refers to a GPU, TPU, or NPU that implements a neural network), which are decoupled from the host (e.g., CPU side) interface and serve as a tool for transparent memory capacity expansion. However, the device-side interconnect is device-dependent and incompatible with different types of GPUs and NPUs. More importantly, pure memory capacity expansion still leaves the embedded layer processing to the host (e.g., CPU) without direct access to the expanded memory. Therefore, a more intelligent, scalable, and compatible storage architecture needs to be designed to release the embedded processing workload from both the host and device sides. Summary of the invention
[0004] Various embodiments of the present specification may include systems, methods, and non-transitory computer-readable media for handling large-scale embedding within a smart memory pool. The system consists of a novel hardware architecture and accompanying software tools to utilize the novel hardware architecture. This novel embedding layer solution provides various technical advantages for memory capacity challenges and memory bandwidth challenges.
[0005] According to one aspect, a method for embedding features is implemented on a computing system including multiple electronic cards, each electronic card storing at least a portion of one or more embedding tables, the method comprising: obtaining a machine learning job for a machine learning model, the machine learning job comprising multiple features, wherein the multiple features include multiple sparse features; based on the one or more embedding tables, executing multiple embedding tasks through multiple electronic cards in parallel to obtain multiple embedding results, wherein each of the multiple embedding tasks includes one or more embedding operations for one or more sparse features among the multiple sparse features, and the multiple embedding results include multiple dense features; and aggregating the multiple embedding results obtained to obtain an input to the machine learning model, thereby processing the machine learning job.
[0006] In some embodiments, each of the plurality of electronic cards includes a near memory coprocessor (NMC), a system on chip (SoC) processor, and one or more memory cards.
[0007] In some embodiments, the method may further include partitioning the one or more embedded tables and loading at least a portion of the partitioned one or more embedded tables from a non-volatile storage medium into one or more memory cards of each of the multiple electronic cards.
[0008] In some embodiments, the one or more embedding operations performed by one of the electronic cards include performing one or more embedding lookups to obtain one or more embeddings for one or more of the plurality of sparse features.
[0009] In some embodiments, performing one or more embedding lookups includes: determining, through one of the multiple electronic cards, whether one or more embeddings corresponding to one or more of the sparse features are stored in one or more memory cards of the electronic card; upon determining that the one or more embeddings are stored in the one or more memory cards, reading the one or more embeddings from the memory card of the electronic card; and upon determining that the one or more embeddings are not stored in the one or more memory cards, redirecting the embedding lookup to a different electronic card and receiving one or more embeddings returned from the different electronic card.
[0010] In some embodiments, performing one or more embedding lookups includes: performing one or more local embedding lookups in the memory card to obtain one or more first embeddings for one or more sparse features in the plurality of sparse features; receiving one or more second embeddings, the one or more second embeddings being generated by one or more remote embedding lookups in another electronic card; and performing one or more embedding operations based on the first embedding and the second embedding to obtain one or more embedding results.
[0011] In some embodiments, executing multiple embedded tasks includes: executing one or more embedded operations by one electronic card among the multiple electronic cards based on one or more operators defined in the NMC of the one electronic card and one or more instructions from the SoC processor of the one electronic card.
[0012] In some embodiments, the method may further include: generating one or more executable files by compiling the machine learning job; and deploying the executable files to a plurality of electronic cards for execution.
[0013] In some embodiments, the method may further include programming the NMCs in the plurality of electronic cards, respectively, to add one or more operators executable by the NMCs.
[0014] In some embodiments, the method may further include updating the plurality of embeddings based on back-propagation from the machine learning model.
[0015] In some embodiments, the plurality of electronic cards are organized into one or more memories, and one or more of the plurality of electronic cards in the same memory are interconnected by a low-latency link with adaptive congestion control.
[0016] In some embodiments, each memory includes a peripheral component interconnect express (PCI-e) interface, and the method further includes: feeding the input to the machine learning model through the PCI-e interface of the memory.
[0017] In some embodiments, the one or more embedding operations include one or more of: lookup, summation, mean, normalization, multiplication, concatenation, reduction, slicing, or hashing.
[0018] According to other embodiments, a system includes one or more processors and one or more computer-readable memories, which are coupled to the one or more processors and store instructions that can be executed by the one or more processors to perform the method of any of the aforementioned embodiments.
[0019] According to some other embodiments, a non-transitory computer-readable storage medium is configured with instructions, and the instructions can be executed by one or more processors to cause the one or more processors to perform the method of any of the foregoing embodiments.
[0020] According to another aspect, a system for embedding features may include: multiple electronic cards, each electronic card including: one or more memory cards, a programmable device (e.g., a system on chip (SoC) field programmable gate array (FPGA) device), a first connection interface and a second connection interface, each of the multiple electronic cards being interconnected with one or more other electronic cards via the first connection interface, wherein each of the multiple electronic cards stores at least a portion of one or more embedding tables in its memory card; and wherein the multiple electronic cards are configured to perform multiple embedding operations based on the one or more embedding tables and utilizing the SoC FPGA device to obtain multiple embedding results, wherein the embedding results are transmitted to one or more processors via the second connection interface.
[0021] In some embodiments, the programmable device includes a near memory coprocessor (NMC) and a SoC processor; the SoC processor is configured to send an embedding instruction to the NMC; and the NMC is configured to perform one or more embedding operations among a plurality of embedding operations according to the embedding instruction.
[0022] In some embodiments, the first connection interface supports a low latency link with congestion control; and the second connection interface comprises a PCI-e interface.
[0023] In some embodiments, the system further comprises an Ethernet interface connected to a pool of host devices, the host devices configured to schedule the plurality of embedded operations.
[0024] In some embodiments, the plurality of electronic cards form a memory, and each of the plurality of electronic cards in the memory is configured to exchange data with an electronic card in a different memory through the second connection interface.
[0025] In some embodiments, the system further includes a compiler configured to compile the machine learning job and generate one or more executable files for multiple electronic cards to perform multiple embedded operations.
[0026] In some embodiments, the multiple electronic cards are further configured to: receive one or more sparse features; perform one or more of the multiple embedding operations on the one or more sparse features; and generate one or more dense features, wherein the multiple embedding results include the generated one or more dense features.
[0027] The embodiments disclosed in the specification have one or more technical effects. In some embodiments, the described smart memory pool is easy to expand to meet the storage requirements for storing a large number of embeddings and for memory-based embedded lookups. Scalability is achieved by building the smart pool as an independent device independent of the host side (e.g., CPU) or device (e.g., GPU or NPU). For example, the smart pool can be configured with a desired number of memories to respond to memory capacity and bandwidth challenges, wherein each memory can include multiple interconnected electronic cards with local memory. Therefore, the smart memory pool can be expanded by stacking multiple memories and / or increasing the number of electronic cards in the memory. In some embodiments, the described smart memory pool is highly compatible with existing system architectures. As an independent device, the smart memory pool can be configured with the most commonly used connection interfaces to exchange data / instructions with the host and device. For example, the smart memory pool can have a peripheral component interconnect express (PCI-e) interface connected to the device side, and an Ethernet interface connected to the host side. These two interfaces are compatible with most hosts and devices on the market. In some embodiments, each electronic card in the smart memory pool can have the computing power to process embedded data (e.g., lookups and other operations) stored in its local memory card and communicate with other electronic cards to exchange embedded data. This feature enables distributed embedded processing to achieve optimal performance. In addition, the processing power of the smart memory pool allows near memory operations, which not only reduces the computing workload of the host, but also avoids continuous data exchange with the host to save bandwidth and time. This provides the benefit of reduced latency. In some embodiments, the electronic card is equipped with configurable (programmable) components, such as a near memory coprocessor (NMC) and a system-on-chip (SoC) processor, which provides fast and easy deployment of new embedded operations, as well as the logic required for new deep learning jobs or models. For example, the described smart memory pool can be automatically configured by using a compiler to compile deep learning jobs and automatically generate executable files for hosts, devices, and smart memory pools. The executable file can automatically coordinate collaboration between electronic cards for distributed embedded processing.
[0028] The above and other features of the systems, methods, and non-transitory computer-readable media disclosed herein, as well as the methods of operation and functions of the related structural elements and the economy of combination and manufacture of the parts will become more apparent after considering the following description and the appended claims with reference to the accompanying drawings, all of which constitute a part of this specification, wherein like reference numerals in the various drawings represent corresponding parts. It should be expressly understood, however, that the drawings are for illustration and description purposes only and are not intended as a definition of the limits of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 An example diagram of a system architecture for processing machine learning tasks according to some embodiments is shown.
[0030] Figure 2 An example diagram illustrating an architecture for handling embedded memory smart cards in accordance with some embodiments.
[0031] Figure 3 An example diagram of a coprocessor for processing an embedded near memory is shown in accordance with some embodiments.
[0032] Figure 4 An example diagram of a process for processing a machine learning job according to some embodiments is shown.
[0033] Figure 5 An exemplary flow chart for compiling a machine learning job according to some embodiments is shown.
[0034] Figure 6 An exemplary method for processing embeddings according to some embodiments is shown.
[0035] Figure 7 An exemplary block diagram of a computer system is shown upon which any of the embodiments described herein may be implemented. DETAILED DESCRIPTION
[0036] This description is intended to enable those skilled in the art to make and use the embodiments, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the description. Therefore, the description is not limited to the embodiments shown, but is in accordance with the widest scope consistent with the principles and features disclosed herein.
[0037] Machine learning (ML) algorithms (also known as deep learning (DL) algorithms) based on deep neural networks (DNNs) are expanding rapidly. Emerging DL-based applications are typically memory-intensive (in terms of capacity and bandwidth), with the embedding layer often referred to as the most memory-intensive stage. A key reason why the embedding layer requires a large amount of memory capacity and bandwidth is that the total number of embeddings scales proportionally with the number of users / items / objects involved in the DL application. The embedding layer is typically configured to embed the input features before feeding them to the DNN. In this application, the term "embedding" in different contexts may refer to a relatively low-dimensional space into which a vector in a high-dimensional space may be transformed, or a vector in a low-dimensional space, or the action of transforming a vector in a high-dimensional space into a vector in a low-dimensional space.
[0038] In some embodiments, the input of the DNN can be constructed as a combination of dense features and sparse features. Dense features can generally be represented as real vectors. Sparse features can be represented as an index of a one-hot encoded vector. These one-hot indexes are used to query multiple embedding lookup tables to map sparse indexes into dense vector dimensional space. The content stored in these lookup tables is called embedding, which is trained to extract deep learning features (for example, extracting information about pages that a particular user likes, which is used to recommend related content or posts to the user). The embedding readout of the lookup table can be combined with other dense embeddings (embedding results of dense features or other sparse features) to generate an output tensor, which is then sent to the DNN for processing. That is, the output tensor of the embedding layer becomes the input tensor of the DNN.
[0039] In this description, an embedding layer solution is described to address the memory capacity and bandwidth challenges in feature embedding tasks. The solution relates to a system that includes a new hardware architecture and accompanying software tools to fully utilize the hardware architecture. The hardware architecture is not only highly scalable to provide extended memory capacity and bandwidth for feature embedding tasks, but also has sufficient flexibility (e.g., programmable) to optimally utilize the extended memory capabilities and bandwidth to release embedding workloads from hosts and devices. Herein, "host" may refer to a host CPU, and "device" may refer to a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a general-purpose GPU (GPGPU), other suitable processing units for neural network machine learning, or any combination of the above processing units. For simplicity, the above-mentioned various processing units for neural network machine learning may be collectively referred to as xPUs in the following description.
[0040] Figure 1 An example diagram of a system architecture for processing machine learning tasks according to some embodiments is shown. As shown, the system architecture may include an xPU pool, a server pool 103, a memory pool 102, and various hardware and interconnects. Figure 1 The types of hardware and interconnects shown are for illustrative purposes and may be replaced by other suitable hardware and interconnects serving the same or similar purposes.
[0041] In some embodiments, the xPU pool 101 may include multiple xPU machines, each of which may include one or more xPU chips 104. As described above, the xPU chip 104 may refer to a dedicated circuit that implements the necessary control and arithmetic logic to execute machine learning algorithms, such as training and reasoning. Exemplarily, the xPU chip 104 includes a GPU, a GPGPU, a TPU, an NPU, etc. In some embodiments, each xPU chip 104 is connected to a PCI-e switch 107 via a PCI-e connection 108 to obtain instructions and exchange data with external components.
[0042] In some embodiments, the smart memory pool 102 may include a plurality of interconnected electronic cards, each of which includes a programmable device (e.g., a system on chip (SoC) field programmable gate array (FPGA) device), a first connection interface, a second connection interface, and one or more memory cards. Such an electronic card may be referred to as a memory intelligent card (MIC) 111. In various embodiments, the electronic card may be implemented as other suitable devices, including, for example, some components of the MIC 111 or components configured to perform some functions of the MIC 111. The electronic card may be represented by other suitable names. The description of the structure, function, and operation of the MIC 111 herein is applicable to all suitable embodiments of the electronic card. Each of the plurality of MICs 111 may be interconnected with one or more other MICs 111 via a first connection interface. Each of the plurality of MICs 111 stores at least a portion of one or more embedding tables in its memory card. The plurality of MICs are configured to perform a plurality of embedding operations based on one or more embedding tables using the SoC FPGA device to obtain a plurality of embedding results. Each of the multiple MICs can exchange data with the deep learning processor pool through the second connection interface, for example, sending multiple embedding results to the deep learning processor pool. For illustrative purposes, Figure 1 Four MICs 111 are shown in FIG. 1 , and the number of MICs 111 can be configured based on application requirements. In some embodiments, multiple MICs can be organized into one or more memories 120, each memory including multiple MICs 111. In some embodiments, the interconnection between the MICs 111 within the memory 120 can be implemented as a mesh network with low-latency links 110, which are referred to as LL links 110 in the following description.
[0043] In the mesh network, each pair of MICs 111 may be connected and exchange data via an LL link 110. The LL link 110 may provide a high-speed and low-latency connection between MICs 111 in the same memory 120, which allows one MIC 111 to read data from other MICs 111 in the same memory machine 120, and the data reading speed is as fast as accessing the local memory. In some embodiments, the LL link 110 may be provided with congestion control, and the congestion control has a time difference congestion notification. For example, based on a first pair of timestamps for sending a first data packet and a second pair of timestamps for sending a second data packet, the LL link 110 may determine the degree of congestion by calculating a one-way time difference from a sender to a receiver, wherein each pair of timestamps includes a timestamp generated by the sender when sending a data packet and a timestamp generated by the receiver when receiving a data packet. In response to a low congestion degree, the controller of the LL link 110 may increase the packet transmission rate at the sender. In response to a high congestion degree, the controller of the LL link 110 may reduce the packet transmission rate at the sender. In some embodiments, LL link 110 may be replaced by other suitable connection means, such as PCI-e structure, Or other suitable link configured to connect to the MIC.
[0044] In some embodiments, these interconnected MICs 111 can provide storage capacity of about tens to hundreds of terabytes. For example, each MIC 111 can include a local memory card array, such as a double data rate (DDR) memory card 106, a SoC FPGA device 105, a first connection interface for connecting to the LL link 110, and a second connection interface (e.g., a PCI-e interface for connecting to the PCI-e connection 108) connected to the machine learning processor pool 101. The SoC FPGA device 105 integrates both the SoC processor and the FPGA architecture into a single device, wherein both the SoC processor and the FPGA architecture are programmable. The programmability of the SoC FPGA device 105 allows the user to configure the smart memory pool 102 to perform the desired embedded processing near the memory card 106, such as embedded lookup and embedded operations, and release such workloads from the host (e.g., the server pool 103) and the machine learning processor (e.g., the xPU pool 101). In some embodiments, the SoC FPGA device 105 refers to a SoC class FPGA, and can be replaced by other suitable programmable devices, such as an application specific integrated circuit (ASIC), an application specific standard component (ASSP), a system on chip (SoC), or a field programmable gate array (FPGA). For simplicity, the following description uses SoC FPGA device 105 for illustration purposes.
[0045] In some embodiments, the second connection interface of each MIC 111 in the smart memory pool 102 may include a PCI-e port, which is connected to the PCI-e switch 107 through the PCI-e connection 108 to exchange data with the xPU pool 101. Through the PCI-e switch 107, the xPU 104 can access any storage location in the smart memory pool 102. In this way, the application running on the xPU 104 is no longer constrained by the limited on-chip memory or board memory of the xPU 104.
[0046] In some embodiments, each MIC 111 may also include an Ethernet interface that is connected to the server pool 103 via the Ethernet 109 structure and the Ethernet switch 112. The server pool 103 may include a data center server having a CPU and a memory. In some embodiments, the server pool 103 may coordinate the processing of machine learning jobs. For example, upon receiving a user request for training or reasoning, the server pool 103 may load (e.g., partition and store in a distributed manner) the relevant embedding table into the smart memory pool 102. Subsequently, the server pool 103 may receive the results via the Ethernet 109 structure and present the results to the user.
[0047] In some embodiments, the PCI-e switch 107 may also be connected to a non-volatile storage medium 113 storing embeddings. When used for machine learning jobs, the embeddings may be loaded from the storage medium 113 through the PCI-e switch 107 and stored in the MIC 111 in the smart memory pool 102. In some embodiments, the embeddings are in the form of an embedding table and may be divided into embedding segments before being loaded into the smart memory pool 102. In this way, the embedding segments may be evenly distributed in the MICs 111 of the smart memory pool 102 to achieve load balancing.
[0048] Figure 2 An example diagram of an architecture 200 for handling an embedded memory smart card (MIC) is shown in accordance with some embodiments. Figure 2 The figure shown can be referred to Figure 1 The architecture of the MIC 111 in FIG. Depending on the implementation, the architecture of the MIC 111 may include more, fewer, or alternative components.
[0049] like Figure 2As shown, the MIC 111 may include a PCI-e interface 207 connected to a PCI-e switch 209, one or more LL-link interfaces 208 for corresponding LL-links 210, a network interface 205 connected to a host CPU, and a SoCFPGA device 105, wherein the LL-link 210 is connected to other MICs in the same memory. In some embodiments, the SoCFPGA device 105 may include a SoC 204 and a near memory coprocessor (NMC) 203. The SoC 204 may be configured to perform general and scheduling / flow control tasks. The NMC 203 may be configured to perform vector-oriented embedding layer operations.
[0050] In some embodiments, for data located in the MIC 111, the SoC 204 may instruct the NMC 203 to obtain the data from the local memory (e.g., DDR 201) of the MIC 111 through the memory controller 202. Then, the obtained data may be sent to the NMC 203 for further processing according to the instructions from the SoC 204. In some embodiments, the processing result may be sent to the SoC 204, to the xPU 104 requesting the result, stored in the local DDR 201, or sent to different MICs 203 through the PCI-e interface 207 (e.g., different MICs are located in different memories) or the LL link interface 208 (e.g., different MICs are located on the same memory). For example, if the processing result is required by one xPU 104, the processing result may be sent to the PCI-e switch 209 through the PCI-e interface 207, and the PCI-e switch 209 routes the result to the xPU 104.
[0051] In some embodiments, for data not in MIC 111 (marked as the first MIC, but the data is in the second MIC), SoC 204 can send instructions to the second MIC through LL link interface 208 and LL link 201. In some embodiments, the instructions can cause the second MIC to perform one or more local memory lookups and embedding operations in the second MIC to generate a result. The result can be sent to the corresponding xPU 104, stored in the local memory in the second MIC, or returned to the NMC 203 of the first MIC 111 for further processing.
[0052] In some embodiments, the MIC 111 may further include a crossbar switch 206, which is used to handle interconnection and data communication between the SoC 204, the NMC 203, the LL link interface 208, and the PCI-e interface 207. In some embodiments, the crossbar switch 206 may be configured to determine whether the requested embedding is stored in the memory card (DDR 201) of the MIC 111; and when it is determined that the embedding is not stored in the memory card of the MIC 111, redirect the embedding lookup to a different MIC storing the requested embedding. In some embodiments, the embedding stored in the MIC 111 is globally addressed, and each MIC 111 may determine whether the requested embedding is stored locally or remotely. In some embodiments, the determination may be performed by the SoC 204 and / or the crossbar switch 206 in the MIC 111.
[0053] Figure 3 An example diagram of a near memory coprocessor (NMC) 300 for processing embedded memory according to some embodiments is shown. The example NMC 300 may be referenced Figure 2 Depending on the implementation, the NMC 300 may include more, fewer, or alternative components.
[0054] In some embodiments, within the MIC, the SoC 204 may send instructions 301 to the NMC 300 to cause the NMC 300 to perform embedded operations. The instructions from the SoC 204 may include control instructions and / or data. The NMC 300 may include a plurality of defined operators that are arranged and perform various embedded operations. In some embodiments, the embedded operations may be performed based on the operators defined in the NMC and the instructions from the SoC processor.
[0055] In some embodiments, the NMC 300 may include a decoder 302 that decodes the instruction to generate one or more of the following: a memory operation, an embedded computation instruction, optional data, or other suitable information. The memory operation may be routed to a MUX 305 for memory access. The embedded computation instruction and optional data may be sent to an Ops / Data block 303 that instructs an embedded computation unit (ECU) 304 to perform the embedded computation accordingly. In some embodiments, the ECU 304 may include multiple operators as atomic operations that may be triggered based on the embedded computation instruction and optional data from the Ops / Data block 303.
[0056] In some embodiments, MUX 305 may refer to a multiplexer as a data selector, and the MUX 305 is used to select an input signal from a plurality of input signals. Figure 3As shown, the input signal of MUX 305 may include signals from the crossbar switch 206, the decoder 302, and the ECU 304. The signal selected by MUX 305 may be referred to as memory Rd / Wr306, and the signal selected by MUX 305 is sent to the memory controller ( Figure 2 202) to access the corresponding memory address in the local memory.
[0057] In some embodiments, the data read from the local memory can be processed in the ECU 304. The processing results of the ECU 304 can be written back to the local memory and sent to the xPU or remote MIC that requested the results through the crossbar switch 206. In some embodiments, the ECU 304 can also notify the SoC 204 after the data is processed. In some embodiments, processing the data can involve one or more embedded operations, such as lookup, summation, averaging, normalization, multiplication, concatenation, slicing, hashing, other suitable operations, or any combination of the above operations.
[0058] Figure 4 An exemplary data flow diagram for processing a machine learning job with an embedding layer is shown in accordance with some embodiments. Figure 4 The data flow shown in FIG. 400 relates to a machine learning job (training or inference job) processing a deep neural network (DNN) 408. For example, a DNN in a commercial search or recommendation system may read tens to hundreds of terabytes of embedding tables in addition to tens to hundreds of gigabytes of weights and parameters in the DNN.
[0059] In some embodiments, the machine learning job acquired or received at block 401 may include multiple dense features 402 and multiple sparse features 403. Dense features 402 may include some embeddings (e.g., user profiles and item brands) and parameters learned in a continuous modeling and attention network (e.g., based on the behavior / time situation of multiple items in a recommendation system). Dense features 402 are typically represented as real vectors and can be directly processed in dense feature processing 404. Dense feature processing is typically small-scale and can be processed by an FPGA or a host.
[0060] In contrast, sparse features 403 may cover features such as user ID, item ID, query ID, etc. Sparse features 403 may initially be represented as one-hot encoded vectors. These one-hot encoded vectors are high-dimensional and sparse and are not suitable for being directly fed into DNN 408 for processing. In some embodiments, these one-hot indexes may be used to query an embedding lookup table (also referred to as an embedding lookup) to project the sparse representation into a low-dimensional and dense space.
[0061] In some embodiments, one or more embedded lookup tables may be loaded into a smart memory pool (e.g., Figure 1 102 in ), and distributed across multiple MICs in a smart memory pool (e.g., Figure 1 Each MIC may store at least a portion of one or more embedded cards. For example, the embedded lookup table may initially be stored in a non-volatile storage medium (e.g., Figure 1 The memory is then partitioned and loaded into a memory card (eg, a DDR chip) of the MIC in the smart memory pool for near memory embedding operations.
[0062] In some embodiments, embedding sparse features 403 in a machine learning job may include: within multiple MICs in parallel, executing multiple embedding tasks based on one or more embedding tables to obtain multiple embedding results, wherein each of the multiple embedding tasks includes one or more embedding operations for one or more sparse features among the multiple sparse features, and the multiple embedding results include multiple dense features.
[0063] In some embodiments, the embedding lookup may include a local embedding lookup 405 (e.g., read from the local DDR of each MIC 111) and a remote embedding lookup 406 (e.g., read from the DDR of a different MIC 111). In some embodiments, each MIC may be able to determine whether the embedding requested by the embedding lookup is stored in its local memory card. When it is determined that the embedding or a portion of the embedding table including the embedding is stored locally, the MIC may read the embedding directly from the memory card; and when it is determined that the embedding or a portion of the embedding table is not stored in the memory card, the MIC may redirect the embedding lookup to a different MIC and receive the embedding returned from the different MIC.
[0064] In some embodiments, performing multiple embedding tasks within multiple MICs may include, by each MIC: performing one or more local embedding lookups from a memory card in one MIC to obtain one or more first embeddings for one or more sparse features; receiving one or more second embeddings generated by one or more remote embedding lookups in another MIC; and performing one or more embedding operations based on the first embeddings and the second embeddings to obtain one or more embedding results.
[0065] For example, embeddings obtained from local storage or remote MICs can be directly processed in a near-memory manner: each MIC in the smart memory pool can perform various embedding operations on the obtained embeddings before sending the embedding results to the DNN. Such near-memory embedding operations can significantly reduce the amount of data transmitted in the network and release computational workloads from the host CPU and neural network devices (e.g., xPUs).
[0066] In some embodiments, the output of dense feature processing 404 and the output of sparse feature embedding (e.g., embedding / dense features) can be aggregated through concatenation and / or reduction 407 to obtain an embedding result (which can be referred to as a dense tensor) before being sent to DNN408 for further calculation. The tensor generated after concatenation and / or reduction 407 is dense and therefore suitable for processing in the xPU (where the DNN is located). In existing solutions, feature processing (404, 405, 406, and 407) is processed by the host. By using a smart memory pool, the above steps can be offloaded from the host to the smart memory pool and can be processed in a distributed manner to obtain better performance and throughput. In some embodiments, sparse features 403 can be input into the smart memory pool for embedding, and the output embedded features (dense) can be aggregated to form the input of DNN 408.
[0067] In some embodiments, embeddings (or their weights) stored in the smart memory pool may be updated during the training of the DNN 408, for example, based on back-propagation from the DNN 408. The embedding updates may be local 409 (e.g., updating embeddings within a MIC) or remote 410 (e.g., updating embeddings in a second MIC via a first MIC).
[0068] Figure 5 An example diagram for compiling a machine learning job according to some embodiments is shown. Figure 1 One of the technical features of the smart memory pool 102 in the embodiment is flexibility. In some embodiments, this flexibility comes from the programmability of components within the smart memory pool 102, such as MICs implemented with SoC FPGA devices. For example, these MICs can be programmed according to new machine learning models 501 to be trained or used by compilers 504 designed for the smart memory pool 102.
[0069] In some embodiments, after receiving the machine learning model 501, the front-end compiler 502 can compile the machine learning model 501 to generate an intermediate representation (IR) 503. The machine learning IR is both front-end independent and target independent. The machine learning model 501 may refer to a tensor flow model, a PyTorch model, or other suitable models. In some embodiments, the IR 503, the system architecture information, and the information of the resources 505 can be fed to the back-end compiler 504 to generate three sets of executable files, including a host executable file 506, a MIC executable file 507, and a DNN executable file 508. The information of the system architecture information and the resources 505 may include information of the host / server pool, the smart memory pool (and the MIC therein), and the xPU pool. The three sets of executable files generated can be deployed in the host / server pool, the smart memory pool, and the xPU pool, respectively, to program the corresponding devices, and / or used for the corresponding devices to execute the machine learning model 501. For example, the host executable 506 may include X86 code for a host / server pool, the DNN executable 508 may include GPU code for a GPU, and the MIC executable 507 may be deployed in the MIC to program the NMC (and / or SoC processor) to add one or more operators executable by the NMC.
[0070] Figure 6 An exemplary method for processing embedding according to some embodiments is shown. The method 600 may be Figure 1 The method 600 can be implemented by Figures 1 to 5 The device, apparatus or system shown is used to perform, for example Figure 1 Memory pool 102 shown in . Depending on the implementation, method 600 may include additional steps, reduced steps, or alternative steps performed in various sequences or in parallel.
[0071] In some embodiments, method 600 may be implemented on a computer system including a plurality of interconnected MICs, each MIC storing at least a portion of one or more embedded tables. In some embodiments, each of the plurality of MICs includes one or more memory cards, a near memory coprocessor (NMC), and a system-on-chip (SoC) processor. In some embodiments, the embedded table may be stored in the MIC by partitioning one or more embedded tables stored in a non-volatile storage medium and loading into one or more memory cards of each of the plurality of MICs.
[0072] Block 620 includes obtaining a machine learning job for a machine learning model, the machine learning job including a plurality of features, wherein the plurality of features includes a plurality of sparse features.
[0073] Block 640 includes performing multiple embedding tasks based on one or more embedding tables in multiple MICs in parallel to obtain multiple embedding results, wherein each of the multiple embedding tasks includes one or more embedding operations, the one or more embedding operations are for one or more sparse features in multiple sparse features, and the multiple embedding results include multiple dense features. In some embodiments, the one or more embedding operations performed by one of the MICs include performing one or more embedding lookups to obtain one or more embeddings for one or more sparse features. In some embodiments, performing the one or more embedding lookups includes: determining, by one of the multiple MICs, whether one or more embeddings corresponding to one or more sparse features in the sparse features are stored in one or more memory cards of the MIC; when determining that the one or more embeddings are stored in the one or more memory cards, reading the one or more embeddings from the memory card of the MIC; and when determining that the one or more embeddings are not stored in the one or more memory cards, redirecting the embedding lookup to a different MIC and receiving one or more embeddings returned from the different MIC.
[0074] In some embodiments, performing one of the multiple embedding tasks by one of the MICs includes: performing one or more local embedding lookups in a memory card in one of the MICs to obtain one or more first embeddings for one or more sparse features; receiving one or more second embeddings generated by one or more remote embedding lookups in another MIC; and performing one or more embedding operations based on the first embedding and the second embedding to obtain one or more embedding results. In some embodiments, performing multiple embedding tasks includes: performing one or more embedding operations by one of the MICs based on one or more operators defined in the NMC and one or more instructions from the SoC processor.
[0075] In some embodiments, the embedding operation includes one or more of: lookup, summation, mean, normalization, multiplication, concatenation, slicing, and hashing.
[0076] Block 660 includes aggregating the obtained embedding results to obtain input to the machine learning model to process the machine learning job.
[0077] In some embodiments, method 600 may further include generating one or more executable files by compiling the machine learning job; and programming NMCs in the plurality of MICs by deploying the executable files to add one or more operators executable by the NMCs.
[0078] In some embodiments, method 600 may further include updating the plurality of embeddings based on back-propagation from the machine learning model.
[0079] In some embodiments, multiple MICs are organized into one or more memories, and one or more MICs within the same memory are interconnected by a low-latency link with adaptive congestion control. In some embodiments, each of the memories includes a peripheral component interconnect express (PCI-e) interface, and method 600 may further include: feeding input to the machine learning model through the PCI-e interface of the memory.
[0080] Figure 7 An exemplary block diagram of a computer system that can implement any of the embodiments described herein is shown. A computing device can be used to implement Figures 1 to 6 The computing device 700 may include a bus 702 or other communication mechanism for transmitting information and one or more hardware processors 704 coupled to the bus 702 for processing information. The hardware processor 704 may be, for example, one or more general-purpose microprocessors.
[0081] The computing device 700 may also include a main memory 707 coupled to the bus 702, such as a random access memory (RAM), a cache, and / or other dynamic storage device 710, for storing information and instructions executed by the processor 704. During the execution of the instructions executed by the processor 704, the main memory 707 may also be used to store temporary variables or other intermediate information. When the instructions are stored in a storage medium accessible to the processor 704, the instructions can make the computing device 700 a special machine configured to perform the operations specified in the instructions. The main memory 707 may include non-volatile media and / or volatile media. Non-volatile media may, for example, include optical disks or magnetic disks. Volatile media may include dynamic memory. Common media forms may include, for example, floppy disks, floppy disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, DRAM, PROM and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge memory, or a network version of a storage medium.
[0082] The computing device 700 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which in combination with the computing device may make the computing device 700 a special purpose machine or program the computing device 700 to make it a special purpose machine. According to one embodiment, the techniques herein are performed by the computing device 700 in response to the processor 704 executing one or more sequences of one or more instructions contained in the main memory 707. Such instructions may be read into the main memory 707 from another storage medium such as the storage device 710. The execution of the sequence of instructions contained in the main memory 707 may cause the processor 704 to perform the processing steps described herein. For example, the processes / methods disclosed herein may be implemented by computer program instructions stored in the main memory 707. When these instructions are executed by the processor 704, the processor 704 may perform the steps shown in the corresponding figures and described above. In alternative embodiments, hardwired circuits may be used in place of software instructions or in combination with software instructions.
[0083] The computing device 700 also includes a communication interface 717 coupled to the bus 702. The communication interface 717 can provide bidirectional data communication coupled to one or more network links that are connected to one or more networks. As another example, the communication interface 717 can be a local area network (LAN) card to provide a data communication connection with a compatible LAN (or a WAN component to communicate with a WAN). Wireless links can also be implemented.
[0084] The performance of certain operations can be distributed across the processor, not just resident in a single machine, but deployed across multiple machines. In some exemplary embodiments, the processor or processor-implemented engine can be located in a single geographic location (e.g., in a home environment, an office environment, or a server farm). In other exemplary embodiments, the processor or processor-implemented engine can be distributed across multiple geographic locations.
[0085] Each process, method and algorithm described in the previous sections can be embodied in a code module executed by one or more computer systems or computer processors including computer hardware, and fully or partially automated by the code module. The processes and algorithms can be implemented in part or in whole in a dedicated circuit.
[0086] When the functions disclosed herein are implemented in the form of software functional units and sold or used as independent products, the functions may be stored in a non-volatile computer-readable storage medium executable by a processor. The (all or part) specific technical solutions disclosed herein or aspects that contribute to current technology may be embodied in the form of software products. The software product may be stored in a storage medium, and the software product includes multiple instructions for enabling a computing device (which may be a personal computer, a server, a network device, etc.) to perform all or some steps in the method of an embodiment of the present application. The storage medium may include a flash drive, a portable hard drive, a ROM, a RAM, a disk, an optical disk, another medium operable to store program code, or any combination of the above components.
[0087] Certain embodiments also provide a system including a processor and a non-transitory computer-readable storage medium, wherein the instructions stored in the storage medium can be executed by the processor to cause the system to perform operations corresponding to the steps in any of the methods of the above embodiments. Certain embodiments also provide a non-transitory computer-readable storage medium configured to be executed by one or more processors to cause the one or more processors to perform operations corresponding to the steps in any of the methods of the above embodiments.
[0088] The embodiments disclosed herein may be implemented by a cloud platform, a server or a server group (hereinafter collectively referred to as a "service system") that interacts with a client. The client may be a terminal device or a client registered by a user on the platform, wherein the terminal device may be a mobile terminal, a personal computer (PC), and any device that may be installed with a platform application.
[0089] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. In addition, certain method or process blocks may be omitted in some embodiments. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith may be performed in other appropriate sequences. For example, the blocks or states described may be performed in a sequence different from that specifically disclosed, or multiple blocks or states may be merged into a single block or state. The example blocks or states may be performed in series, in parallel, or in some other manner. Blocks or states may be added to or removed from the disclosed example embodiments. The exemplary systems and components described herein may be different from the described configurations. For example, elements may be added, removed, or rearranged compared to the disclosed example embodiments.
[0090] The various operations of the example methods described herein may be performed at least in part by an algorithm. The algorithm may include a program code or instruction stored in a memory (e.g., the non-transitory computer-readable storage medium described above). The algorithm may include a machine learning algorithm. In some embodiments, the machine learning algorithm may not explicitly program the computer to perform a function, but may learn from training data to establish a predictive model for performing the function.
[0091] The various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute a processor-implemented engine that runs the processor-implemented engine to perform one or more operations or functions described herein.
[0092] Similarly, the methods described herein may be implemented at least in part by a processor, wherein a particular processor or processors are examples of hardware. For example, at least some of the operations of the method may be performed by one or more processors or an engine implemented by a processor. In addition, the one or more processors may be run in a "cloud computing" environment to support the performance of related operations or as "software as a service" (SaaS). For example, at least some operations may be performed by a group of computers (e.g., a machine including a processor), which may be accessed over a network (e.g., the Internet) and through one or more appropriate interfaces (e.g., an application program interface (API)).
[0093] In this specification, multiple instances can implement the components, operations or structures described as single instances. Although the individual operations of one or more methods are shown and described as individual operations, one or more of the individual operations can be performed simultaneously, and the operations do not need to be performed in the order shown. The structure and function presented as individual components in the example configuration can be implemented as a combined structure or component. Similarly, the structure and function presented as a single component can be implemented as an individual component. The above and other changes, modifications, additions and improvements are all within the scope of this paper theme.
[0094] Although the overview of the subject matter has been described with reference to specific exemplary embodiments, various modifications and changes may be made to these embodiments without departing from the scope of the broader embodiments of the present disclosure. For convenience only, these embodiments of the subject matter may be referred to herein individually or collectively by the term "invention", and if more than one embodiment is in fact disclosed, there is no intention to voluntarily limit the scope of the present application to any single disclosure or concept.
[0095] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be derived therefrom and used so that structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Therefore, the detailed description should not be regarded as limiting, and the scope of the various embodiments is limited only by the claims and the full range of equivalents to which these claims are entitled.
[0096] Any process description, element or block in the flowcharts described herein and / or described in the accompanying drawings should be understood to potentially represent a module, segment or portion of code, including one or more executable instructions for implementing specific logical functions or steps in the process. As will be appreciated by those skilled in the art, alternative implementations are included within the scope of the embodiments described herein, wherein elements or functions may be deleted, executed in a different order than shown or discussed, including substantially simultaneously or in reverse order, depending on the functions involved.
[0097] Unless otherwise expressly stated or the context otherwise dictates, the "or" used in this article is inclusive and not exclusive. Therefore, in this article, unless otherwise expressly stated or the context otherwise dictates, "A, B or C" means "A, B, A and B, A and C, B and C, or A, B and C". In addition, unless otherwise expressly stated or the context otherwise dictates, "and" is both joint and separate. Therefore, in this article, unless otherwise expressly stated or the context otherwise dictates, "A and B" means "A and B, joint or each". In addition, multiple instances can be provided for the resources, operations or structures of a single instance described in this article. In addition, the boundaries between various resources, operations, engines and data stores are somewhat arbitrary, so specific operations are described in the context of a specific illustrative configuration. Other allocations of functions can be envisioned, and the allocations are within the scope of various embodiments of the present disclosure. In general, structures and functions presented as separate resources in the example configuration can be implemented as combined structures or resources. Similarly, structures and functions presented as single resources can be implemented as separate resources. The above and other changes, modifications, additions and improvements are within the scope of the embodiments of the present disclosure as represented by the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
[0098] The terms "include" or "comprises" are used to indicate the presence of subsequently stated features, but do not preclude the addition of other features. Unless otherwise specifically stated or otherwise understood in the context of use, conditional language, such as "can", "may", "possibly", or "may", etc., is generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not include the certain features, elements, and / or steps. Therefore, the above conditional language is generally not intended to indicate that the features, elements, and / or steps are in any way required by one or more embodiments, or that one or more embodiments must include logic to determine (with or without user input or prompting) whether the features, elements, and / or steps are included in any particular embodiment or will be performed in any particular embodiment.
Claims
1. A method for embedding a feature, the method being implemented on a computing system comprising a plurality of electronic cards, each electronic card storing at least a portion of one or more embedding tables, the method comprising: Obtaining a machine learning job for a machine learning model, the machine learning job comprising a plurality of features, wherein the plurality of features comprises a plurality of sparse features; Based on the one or more embedding tables, executing a plurality of embedding tasks by the plurality of electronic cards in parallel to obtain a plurality of embedding results, wherein each of the plurality of embedding tasks comprises one or more embedding operations for one or more sparse features among the plurality of sparse features, and the plurality of embedding results comprises a plurality of dense features; and Aggregating the obtained multiple embedding results to obtain an input of the machine learning model, thereby processing the machine learning job; The one or more embedding operations performed by one of the plurality of electronic cards include: performing one or more embedding lookups to obtain one or more embeddings for one or more sparse features of the plurality of sparse features.
2. The method according to claim 1, wherein: Each of the plurality of electronic cards includes a near memory coprocessor NMC, a system-on-chip SoC processor, and one or more memory cards.
3. The method according to claim 2, further comprising: The one or more embedded tables are partitioned, and at least a portion of the partitioned one or more embedded tables are loaded from a non-volatile storage medium into one or more memory cards of each of the plurality of electronic cards.
4. The method according to claim 1, wherein: The performing one or more embedded lookups comprises: determining, by one of the plurality of electronic cards, whether one or more embeddings corresponding to one or more of the sparse features are stored in one or more memory cards of the electronic card; Upon determining that the one or more embeddings are stored in the one or more memory cards, reading the one or more embeddings from the memory card of the electronic card; and Upon determining that the one or more embeddings are not stored in the one or more memory cards, redirecting the embedding lookup to a different electronic card and receiving the one or more embeddings returned from the different electronic card.
5. The method according to claim 1, wherein: Executing the plurality of embedded tasks by one of the plurality of electronic cards comprises: performing one or more local embedding lookups in a memory card of the electronic card to obtain one or more first embeddings for one or more sparse features of the plurality of sparse features; receiving one or more second embeddings resulting from one or more remote embedding lookups in another electronic card; and One or more embedding operations are performed based on the first embedding and the second embedding to obtain one or more of the embedding results.
6. The method according to claim 2, wherein: Executing the plurality of embedding tasks includes: The one or more embedded operations are performed by one electronic card among the plurality of electronic cards based on one or more operators defined in the NMC of the one electronic card and one or more instructions from the SoC processor of the one electronic card.
7. The method according to claim 1, further comprising: Generate one or more executable files by compiling the machine learning job; as well as The executable file is deployed to the electronic card for execution.
8. The method according to claim 2, further comprising: The NMCs are programmed in the plurality of electronic cards respectively to add one or more operators executable by the NMCs.
9. The method according to claim 1, further comprising: The one or more embedding tables are updated based on back-propagation from the machine learning model.
10. The method according to claim 1, wherein: The plurality of electronic cards are organized into one or more memories, and one or more electronic cards in the same memory are interconnected by a low-latency link with adaptive congestion control.
11. The method according to claim 10, wherein: Each of the memories comprises a Peripheral Component Interconnect Express (PCI-e) interface, and the method further comprises: The input is fed to the machine learning model through the PCI-e interface of the memory.
12. An intelligent memory pool system for embedding features, comprising: A plurality of electronic cards, each electronic card comprising: one or more memory cards, a programmable device, a first connection interface and a second connection interface; Each of the plurality of electronic cards is interconnected with one or more other electronic cards via the first connection interface; wherein each of the plurality of electronic cards stores at least a portion of one or more embedded tables in a memory card thereof; and wherein the plurality of electronic cards are configured to perform a plurality of embedding operations based on the one or more embedding tables and using a plurality of programmable devices to obtain a plurality of embedding results, wherein the embedding results are transmitted to the one or more processors via the second connection interface; The multiple embedding operations performed by the multiple electronic cards include: performing multiple embedding searches.
13. The system according to claim 12, wherein: The programmable device includes a system on chip (SoC) field programmable gate array (FPGA) device.
14. The system of claim 12, wherein: The programmable device includes a near memory coprocessor NMC and a SoC processor; The SoC processor is configured to send embedded instructions to the NMC; and The NMC is configured to perform one or more embedding operations of the plurality of embedding operations according to the embedding instruction.
15. The system of claim 12, wherein: The first connection interface supports a low latency link with congestion control; and The second connection interface includes a PCI-e interface.
16. The system of claim 12, further comprising an Ethernet interface connected to a pool of host devices, the host devices configured to schedule the plurality of embedded operations.
17. The system of claim 12, wherein: The plurality of electronic cards form a memory, and each of the plurality of electronic cards in the memory is configured to exchange data with an electronic card in a different memory through the second connection interface.
18. The system of claim 12, further comprising a compiler, wherein: The compiler is configured to compile the machine learning job and generate one or more executable files for the plurality of electronic cards to perform the plurality of embedded operations.
19. The system of claim 12, wherein: The multiple electronic cards are further configured as: receiving one or more sparse features; performing one or more embedding operations of the plurality of embedding operations on the one or more sparse features; as well as One or more dense features are generated, wherein the plurality of embedding results include the generated one or more dense features.
Citation Information
Patent Citations
Neural network data entry system
CN110036399A
Hardware accelerator architecture for processing very-sparse and hyper-sparse matrix data
US10146738B2