Memory Bandwidth Management for Deep Learning Applications

By using FPGA processors that support parallel processing in data centers and using batch and parallel locking steps, the memory bandwidth problem in deep learning applications is solved, and efficient deep neural network evaluation is achieved.

CN113112006BActive Publication Date: 2025-06-27MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110584882.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-06-25
Filing Date
2016-06-23
Publication Date
2025-06-27
Estimated Expiration
2036-06-23

AI Technical Summary

Technical Problem

In deep learning applications, as complexity and scalability increase, memory bandwidth becomes a serious problem, and it is difficult to meet the needs of high computing power and low cost.

Method used

By introducing processors that support parallel processing, such as FPGAs, in the data center, leverage batch processing technology and parallel locking steps, memory bandwidth requirements during neural network weight loading and efficient processing of multiple parallel streams are achieved.

Benefits of technology

This technology significantly reduces memory bandwidth requirements, improves efficiency and throughput for deep learning applications, and enables complex deep neural network evaluations within processor-to-memory bandwidth constraints in the current data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113112006B_ABST
    Figure CN113112006B_ABST
Patent Text Reader

Abstract

In a data center, neural network evaluation can be included for services involving image or speech recognition by using field programmable gate arrays (FPGAs) or other parallel processors. The memory bandwidth limitation of providing a weighted data set from an external memory to the memory of an FPGA (or other parallel processor) can be managed by queuing input data from multiple cores performing services at the FPGA (or other parallel processor) in batches of at least two feature vectors. The at least two feature vectors can be at least two observation vectors from the same data stream or from different data streams. The FPGA (or other parallel processor) can then take action on a batch of data for each load of the weighted data set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application "Memory Bandwidth Management for Deep Learning Applications" with an application date of June 23, 2016 and an application number of 201680034405.4. Background Art

[0002] Artificial intelligence (AI) applications involve machines or software that are made to exhibit intelligent behaviors such as learning, communicating, perceiving, moving and manipulating, and even creating. These machines or software can achieve such intelligent behaviors through various methods such as search and optimization, logic, probabilistic methods, statistical learning, and neural networks. Along these lines, various deep learning architectures, such as deep neural networks (deep NNs) including deep multi-layer perceptrons (MLPs) (often referred to as DNNs), convolutional deep neural networks, deep belief networks, recurrent neural networks (RNNs), and long short-term memory (LSTM) RNNs, have received attention for their application to fields such as computer vision, image processing / recognition, speech processing / recognition, natural language processing, audio recognition, and bioinformatics.

[0003] Deep NNs typically include an input layer, any number of hidden layers, and an output layer. Each layer contains a specific number of units, which can follow a neural model, and each unit corresponds to an element in a feature vector (e.g., an observation vector of an input data set). Each unit typically uses a weighted function (e.g., a logistic function) to map its total input from the lower layer to a scalar state that is sent to the upper layer. The layers of the neural network are trained (usually via unsupervised machine learning) and the units of the layer are assigned weights. Depending on the depth of the neural network layer, the total number of weights used in the system can be huge.

[0004] Many computer vision, image processing / recognition, speech processing / recognition, natural language processing, audio recognition, and bioinformatics are run and managed at data centers that support services available to a large number of consumers and enterprise customers. Data centers are designed to run and operate computer systems (servers, storage devices, and other computers), communication devices, and power systems in a modular and flexible manner. Data center workloads require high computing power, flexibility, power efficiency, and low cost. At least some parts that can accelerate large-scale software services can achieve the desired throughput and enable these data centers to meet the needs of their resource consumers. However, the increasing complexity and scalability of deep learning applications may exacerbate problems regarding memory bandwidth. Summary of the Invention

[0005] Memory bandwidth management techniques and systems for accelerating neural network evaluation are described.

[0006] In a data center, a neural network evaluation accelerator may include a processor that supports parallel processing (“parallel processor”), such as a field programmable gate array (FPGA). This processor is separated from the general purpose computer processing unit (CPU) at the data center and, after at least two observation vectors from the same or different data streams (from the cores of the CPU), it uses a weight data set loaded from an external memory to perform a process. By queuing the input data of at least two streams or at least two observation vectors before applying the weighted data set, the memory bandwidth requirement for neural network weight loading can be reduced by a factor of K, where K is the number of input data sets in a batch. Additionally, by using a processor that supports parallel processing, N simultaneous streams can be processed in parallel locking steps to ensure that the memory bandwidth requirement for N parallel streams remains the same as the memory bandwidth requirement for a single stream. For each load of the weight data set, this achieves a throughput of N*K input data sets.

[0007] Services that benefit from including a deep learning architecture hosted at a data center may include deep neural network (deep NN) evaluation performed on an FPGA, where the method includes loading a first weight data set from off-chip storage, queuing a batch of at least two feature vectors at the input of the FPGA, performing a first layer process of the deep NN evaluation on the batch to generate intermediate results, loading a second weight data set from off-chip storage, and performing a second layer process of the deep NN evaluation on the intermediate results. In some cases, the at least two feature vectors may include from at least two data streams, where the at least two data streams are from corresponding cores. In some cases, the at least two feature vectors may be from the same data stream. In some cases, the at least two feature vectors may include at least two observation vectors from each of at least two data streams.

[0008] The present invention content is provided to introduce in a simplified form a series of concepts that are further described below in the detailed description. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Brief Description of the Drawings

[0009] Figure 1 Illustrates an example operating environment that provides memory bandwidth management for deep learning applications.

[0010] Figure 2 Illustrates an example architecture of at least components for managing and accelerating deep learning applications hosted by resources of a data center.

[0011] Figures 3A to 3D Illustrates a comparison of bandwidth management in neural network evaluation.

[0012] Figure 4 An example implementation of using an FPGA to accelerate the DNN process is illustrated.

[0013] Figure 5 An example computing system on which the described techniques may be performed is illustrated. Detailed Description

[0014] Memory bandwidth management techniques and systems that can accelerate neural network evaluation are described.

[0015] Due to the computational patterns attributed to many neural network evaluations, general-purpose processors and purely software-based techniques tend to be inefficient and, in some cases, unable to meet the performance requirements of the applications of which they form a part. Additionally, these evaluations are often limited by the resources available at the data centers where the computations are performed. By including FPGAs in the data center and leveraging these FPGAs in the manner described herein, complex deep neural network evaluations can be performed within the processor-to-memory bandwidth constraints and machine-to-machine networking bandwidth constraints of current data centers. In some cases, particularly where power efficiency is not a priority or fewer parallel computational streams are needed, graphics processing units (GPUs) can be used to perform neural network evaluations.

[0016] Figure 1 An example operating environment for providing memory bandwidth management for deep learning applications is illustrated. Referring Figure 1 , data center 100 may include a number of resources 101, physical and virtual, on which applications and services may be hosted. One or more routing servers 110 may facilitate directing requests to appropriate resources. One or more routing servers 110 may physically reside within a particular data center 100. One or more services hosted at data center 100 may serve a number of clients, such as client 0 121, client 1 122, client 2 123, client 3 124, client 4 125, client 5 126, and client 6 127, which communicate with the (one or more) services (and access one or more data center resources 101) via the Internet 130.

[0017] Various implementations of the described techniques are suitable for use as part of a process for services that involve computer vision, image processing / recognition, speech processing / recognition, natural language processing, audio recognition, bioinformatics, weather prediction, stock forecasting, control systems, and any other applications to which neural networks may be applied.

[0018] As an example scenario, the operating environment supports a translation service for audio or video calls. The translation service can involve deep learning applications to identify words from the conversation. A client (e.g., 121, 122, 123, …) can enable a user to select to participate in the translation service so that the audio of the user conversation can be sent to the translation service. For example, the audio from a conversation at a device running client 0 121 can be sent to the translation service as input 0, the audio from a conversation at a device running client 1 122 can be sent to the translation service as input 1, the audio from a conversation at a device running client 2 123 can be sent to the translation service as input 2, and the audio from a conversation at a device running client 3 124 can be sent to the translation service as input 3.

[0019] These independent conversations can be processed at the data center 100 and can have the output sent to the same client or a different client (which may or may not participate in sending the audio to the service). For example, the translated conversation from input 0 can be sent to client 4 125 as output 0, the translated conversation from input 1 can be sent to client 5 126 as output 1, the translated conversation from input 2 can be sent to client 6 127 as output 2, and the translated conversation from input 3 can be sent back to client 3 125 as output 3. Accelerating one or more of the processes associated with the translation service can contribute to the real-time functionality of such a service. However, any acceleration or even just the actual computation is at least partially constrained by the physical limitations of the system (i.e., the data center resources 101 at the data center 100).

[0020] Figure 2 Illustrated is an example architecture for managing and accelerating at least components of a deep learning application hosted by resources of a data center. In a data center 200 that houses many servers, switches, and other devices for various services and applications such as those described regarding Figure 1 the data center 100, deep learning applications and in particular neural network evaluations where the weight data sets are on the order of megabytes (tens of megabytes, hundreds of megabytes, or even more) can be accelerated and their memory bandwidth requirements can be reduced by batching the input data.

[0021] Servers at data center 200 may include a processor having two or more cores. Each core tends to process a single thread or a single data stream. According to some embodiments, parallel processor 210 is used to accelerate neural network evaluation. As an example, parallel processor 210 may be a GPU or an FPGA. The FPGA may exhibit improved power savings over the use of a GPU. However, it should be understood that some embodiments may use a GPU or other parallel processors to perform some or all of the methods described herein.

[0022] Parallel processor 210 may include: an input buffer 211 having a specified queue depth (e.g., K = 2 or more); and an output buffer 212 for holding intermediate result outputs or other data before the data is further processed and / or output to another component. In some cases, the parallel processor may include logic 220 for processing data from one or more input buffers 211. Depending on the type of parallel processor 210, logic 220 may be programmable and / or reconfigurable between and / or during operations.

[0023] An "observation vector" refers to an initial data set (or feature vector) that is input into a neural network and used to initiate or trigger an identification or classification process. This may include data representing color, price, sound amplitude, or any other quantifiable value that may have been observed in an object of interest. The observation vector input into parallel processor 210 may be generated by a core of the CPU (discussed in more detail in the example below), by another computing unit such as an FPGA, a CPU, or other logic on parallel processor 210 that is performing neural network evaluation.

[0024] An "intermediate result output" or "intermediate result" refers to an internal state value of a neural network used to track the progress of data through the network for the current evaluation or across multiple evaluations of the network (in the case of an RNN).

[0025] The intermediate result values may be related to the features of the observation vector, but generally they represent some abstract form of the original observation data as the "reason" for the network algorithm given its input. Logically, the intermediate result values represent the network's attempt to classify the data based on a hypersphere decision boundary between competing concepts. Mathematically, the intermediate result values represent the proximity of the observation vector, previous intermediate result values, or a combination of both to the boundary between competing concepts represented by one or more neural network nodes.

[0026] As Figure 2As shown in the figure, a single parallel processor 210 can receive input data from multiple cores, such as core 0 221, core 1 222, core 2 223, and core 3 224, which can be provided in one or more processing units housed in one or more servers. For example, a single server can contain a processing unit with 12 - 24 cores, which is common today. In some applications, the input data from these cores can be loaded into the corresponding buffers of the input buffer 211. In other applications, the input data from one of these cores can be loaded into more than one input buffer.

[0027] In another embodiment, instead of many separate queues (provided by the input buffer 211), one queue per core among the N cores, the parallel processor can have a single queue into which different cores add their data sets when they become available. The parallel processor can poll this queue periodically (after each complete evaluation of depth NN) and read new batches of data sets from this queue for parallel processing. The new batches of data sets will be processed in parallel by the depth NN, and then a decoder on the parallel processor or on a CPU core will send them back to the appropriate cores. Each data set will be tagged by the core that sent it to facilitate this demultiplexing operation. This type of embodiment is suitable for cases where a parallel processor such as an FPGA can handle a computational load such as adding decoding processing.

[0028] As mentioned above, the observation vector can be generated by the core and provided directly to the parallel processor 210 as input data; however, in some cases, the observation vector may not be output by the core. In some of those cases, the parallel processor 210 can generate the observation vector (using separate logic to do so) or other computational units can generate the observation vector based on data output from the core or in a system implemented entirely using other computational units. Thus, when describing the core and data flow herein, the use of other computational units can be considered as other embodiments that can be suitable as processing units for a particular recognition process (or other applications that benefit from deep learning).

[0029] When a parallel processor executes a weighted function for neural network evaluation, the weight dataset is typically too large to be stored on-chip using the processor. Instead, the weight dataset is stored in off-chip storage 230 and loaded onto the parallel processor 210 in partitions small enough for on-chip storage each time a particular weighted function is executed. To achieve maximum efficiency, the weights must be loaded at a rate that the processor can consume them, which requires a significant amount of memory bandwidth. Off-chip storage 230 can include memory modules (e.g., DDR, SDRAM DIMM), hard disk drives (solid state, hard disk, magnetic, optical, etc.), CDs, DVDs, and other removable storage devices. It should be understood that storage 230 does not include propagated signals.

[0030] According to various techniques described herein, the memory bandwidth is managed by processing parallel data streams (e.g., from cores: core 0 221, core 1 222, core 2 223, and core 3 224) in batches at the parallel processor 210, such that the weight dataset input from off-chip storage 230 to the parallel processor 210 for each layer processes at least two feature vectors. Although the techniques described can be useful for two feature vectors (from the same or different data streams), a significant effect on bandwidth and / or power efficiency can be seen when at least four feature vectors are processed in parallel. For example, doubling the number of items processed in parallel approximately halves the memory bandwidth.

[0031] In some cases, the acceleration of deep NN evaluation can be managed by a manager agent 240. The manager agent 240 can be implemented using software executable by any suitable computing system (including physical servers and virtual servers and any combination thereof). The manager agent 240 can be installed on a virtual machine and run in the context of the virtual machine in some cases or can be directly installed on the computing system and executed on the computing system in a non-virtualized implementation. In some cases, the manager agent 240 can be implemented wholly or in part using hardware.

[0032] The manager agent 240, when used, can coordinate the timing of data transfer between various components at the data center (e.g., between the off-chip storage 230 storing the weights for deep NN evaluation and the parallel processor 210). Thus, in certain embodiments, the manager agent 240 and / or the bus / data routing configuration for the data center 200 enables a dataset (e.g., from cores: core 0 221, core 1 222, core 2 223, core 3 224) to be transferred to a single parallel processor 210 for processing in batches.

[0033] Figures 3A to 3D Illustrates a comparison of bandwidth management in neural network evaluation. InFigure 3A In the scenario illustrated in the figure, the deep NN evaluation involves software or hardware evaluation and does not include acceleration or memory management as described herein. Instead, the weight data sets stored in the external memory 300 are separately applied to each data set (e.g., from Core: Core 0 301 and Core 1 302) during the corresponding deep NN evaluations (e.g., DNN evaluation (0) 303 and DNN evaluation (1) 304) respectively. Specifically, the first layer weight set 310 from the external memory 300 and the first feature vector (vector 0) 311 of the first stream (Stream 0) from Core 0 301 are acquired / received (312) for DNN evaluation (0) 303; the first layer process is executed (313), generating an intermediate result 314.

[0034] Subsequently, the second layer weight set 315 from the external memory 300 is acquired / received 316 to perform the second layer process (317) on the intermediate result 314. This evaluation process of acquiring / receiving weights from the external memory 300 continues for each layer until the entire process is completed for a specific input vector. This process can be repeated for each input feature vector (e.g., vector 01 of vector 0). If there are multiple cores running the deep NN evaluation, multiple deep NN evaluations can be executed in parallel, but without the memory management as described herein, each evaluation requires acquiring / receiving the weight data set from the external memory 300 as an independent request and using additional processors / cores.

[0035] For example, the first layer weight set 310 from the external memory 300 and the second feature vector (vector 1) 318 of the second stream (Stream 1) from Core 1 302 are acquired / received (319) for DNN evaluation (0) 304; and the first layer process is executed (313) to generate an intermediate result 320 for the second stream. Subsequently, the second layer weight set 315 from the external memory 300 is acquired / received 316 to perform the second layer process (317) on the intermediate result 320. As described for DNN evaluation (0) 303, DNN evaluation (1) 304 continues for each layer until the entire process is completed for a specific input feature vector, and is repeated for each input feature vector (e.g., vector 11 of Stream 1). As can be seen, this is not an efficient mechanism for performing deep NN evaluations and requires a considerable amount of memory bandwidth to execute, because each evaluation requires acquiring the weighted data set and requires its own processor / (one or more) cores to perform the evaluation.

[0036] In Figure 3BIn the scenario illustrated, the memory - managed, accelerated deep NN process uses the FPGA 330 to perform deep NN evaluations on two feature vectors from the same data stream. Here, two feature vectors, namely the observation vectors (vector 0 and vector 01) 331 can be loaded as a batch from core 0 301 onto the FPGA 330 and evaluated. The first - layer weight set 310 can be received and / or fetched for loading (332) onto the FPGA 330, and the first - layer process is performed (333) at the FPGA 330 on both vector 0 and vector 01 331. The intermediate result 334 from the first - layer process can be available when the first - layer process is completed for the two observation vectors; and the second - layer weight set 315 is loaded (335) onto the FGPA 330.

[0037] The intermediate result 334 can be loaded into a buffer (e.g., Figure 2 input buffer 211) for the next - layer processing and the second - layer process performed (337). The deep NN evaluation at the FPGA 330 continues for each layer until the entire process is completed for the two observation vectors (vector 0 and vector 01) 331. Then the process can be repeated for the next pair of observation vectors (e.g., vector 02 and vector 03). It should be understood that although only two feature vectors are described, the FPGA can be used to evaluate more than two feature vectors / observation vectors in parallel. This will further reduce the required memory bandwidth, but it will be found to have limitations in the latency of buffering batches from a single data stream, because in real - time applications, buffering N observation vectors before starting the computation delays the system's output by N vectors, and in some applications, the acceptable latency is not large enough to allow the effective operation of the parallel processor. Additionally, batching of vectors from the same stream is not suitable for recurrent networks (RNNs) because the vectors within a batch are not independent (the computation at time step t requires the output at time step t - 1).

[0038] In Figure 3C the scenario illustrated, the memory - managed, accelerated deep NN process using the FPGA 330 needs to perform DNN evaluations on two feature vectors from two different data streams. That is, two feature vectors, one observation vector (vector 0) 341 from core 0 301 and one observation vector (vector 1) 342 from core 1 302 are evaluated as a batch. The first - layer weight set 310 can be received / fetched for loading (343) onto the FPGA 330, and the first - layer process is performed (344) at the FPGA 330 on both vector 0 341 and vector 1 342, generating an intermediate result 345. The intermediate result 345 (the data output of the first - layer process) can be loaded into a buffer for the next - layer processing. UnlikeFigure 3B The scenario illustrated in Figure 1 is suitable for evaluating the DNN of the MLP but not for evaluating the deep NN of the RNN. This scenario is suitable for evaluating the RNN (recurrent neural network) because all the observations within the batch are from different streams and are thus independent.

[0039] After the first-layer process is completed for the two observation vectors 341 and 342, the second-layer weight set 315 is loaded (346) onto the FPGA 330 and the second-layer process is executed (347). The deep NN evaluation at the FPGA 330 continues for each layer until the entire process is completed for the two observation vectors (vector 0 341 and vector 1) 342. The process can then be repeated for subsequent observation vectors of the two streams (e.g., vector 01 and vector 11). Of course, although only two streams and cores are shown for simplicity, more than two can be evaluated in parallel at the FPGA 330.

[0040] In Figure 3D In the scenario illustrated in Figure 2, the memory-managed, accelerated deep NN process using the FPGA 330 requires performing deep NN evaluation on four feature vectors, with every two feature vectors (observation vectors) coming from two data streams. That is, two observation vectors 351 (vector 0, vector 01) from core 0 301 and two observation vectors 352 (vector 1, vector 11) from core 1 302 are loaded and evaluated as a batch. The two observation vectors from each stream can be loaded through an input buffer (see, for example, Figure 2 the input buffer 211 of Figure 3) with a queue depth of 2 for the FPGA. Thus, with a single loading (353) of the first-layer weight set 310, two observation vectors 351 (vector 0, vector 01) from core 0 301, and two observation vectors 352 (vector 1, vector 11) from core 1 302, the first-layer process can be executed (354). As described above regarding the scenario illustrated in Figure 3B Figure 1, although this scenario is suitable for evaluating various deep NN architectures, it is not as suitable for the RNN as described regarding Figure 3C Figure 2.

[0041] The intermediate result 355 (data output of the first layer process) can be loaded into the buffer for the next layer processing. The second layer weight set 315 can be fetched / received for loading (356) onto the FPGA 330 and for the second layer process to be performed (357) thereafter. At the FPGA 330, DNN evaluation continues for each layer until the entire processing is completed for at least four observation vectors (vector 0, vector 01, vector 1, and vector 11). Thereafter, the process can be repeated for the next pair of observation vectors of the two streams (e.g., vector 02, vector 03, and vector 12, vector 13). Of course, this scenario can also be extended to additional cores processed in parallel by the FPGA 330.

[0042] As can be seen from the illustrated scenario, Figure 3B and 3C the configuration shown in Figure 3A reduces the memory bandwidth required to evaluate the same amount of data as the configuration shown in Figure 3D Moreover, the configuration shown in Figure 3B and 3D can even further reduce the required memory bandwidth. For the configurations shown in Figure 3B and 3D , there is a time delay cost for queuing multiple observation vectors from a single data stream. Additionally, there may also be some delay cost for evaluating the amount of data twice (or more) for each line of the available parallel processes.

[0043] Therefore, the input data of at least two streams and / or at least two observation vectors can be queued for processing at the FPGA to reduce the memory bandwidth requirement for neural network weight loading by K times, where K is the number of input data sets in a batch (and can also be regarded as the queue depth for the FPGA). To optimize the bandwidth efficiency, processing is performed when a batch of K input data sets has been accumulated in the on-chip FPGA memory. By queuing the inputs in this way, any I / O boundary issues can be handled in the case where the bandwidth of the database (weights) required to process the input data sets is too high. Thus, overall, for the required bandwidth B, the average effective bandwidth required using the queuing method is B / K. N simultaneous streams can be processed using parallel locking steps to ensure that the memory bandwidth requirement for N parallel streams remains the same as that for a single stream.

[0044] Example scenario - Internet translator

[0045] For an Internet translator, a session can arrive at a data center after an input via a microphone at a client and be transformed (e.g., via a fast Fourier transform) to establish power bands based on frequencies from which an observation vector can be obtained. The deep NN evaluation for an example scenario involves a DNN (for an MPL) that performs eight-layer matrix multiplication, adds a bias vector, and applies a non-linear function for all layers except the top layer. The output of the DNN evaluation can establish a probability score indicating how likely the slice being viewed belongs to a speech unit, e.g., how likely the slice being viewed belongs to the middle part of "ah" read out in the left context of "t" and the right context of "sh". Additional processing tasks performed by a processor core can involve identifying words based on the probability score and applying it to certain dictionaries to perform translation between languages.

[0046] Figure 4 An example implementation of using an FPGA to accelerate the deep NN process is illustrated. Refer to Figure 4 , a server blade 400 at the data center can include an external storage 401 that stores a weight data set. Advantageously, a single FPGA 402 may be capable of performing all of the deep NN evaluation for an entire server blade that includes 24 CPU cores 403 (N = 24); while the deep NN is being executed, these cores 403 handle other processing tasks required for 24 simultaneous sessions.

[0047] In an example scenario, the input data set from each of the 24 cores can be loaded onto an input buffer 404 with a queue depth of K = 2. Thus, two observation vectors of the data stream from each session / core can undergo processing through layer logic 405 as a single batch, e.g., matrix multiplication (e.g., for a deep MLP) or multiple parallel matrix multiplications and non-linear matrix multiplications (e.g., for an LSTM RNN). The intermediate result 406 of the layer from layer logic 405 can be routed back (407) when a new weight data set is loaded from storage 401 so as to undergo another process with the new weighting function. This process can be repeated until all layers have been processed, at which point the output is then sent back to the processing core, which can demultiplex (demux) the data at some point.

[0048] Live speech translation requires close attention to latency, power, and accuracy. FPGAs typically have relatively low power requirements (10W) and can still deliver high computational performance. Since using only CPU cores (a pure software approach) for performing speech recognition using deep NN evaluations typically requires at least 3 CPU cores per session, with at least two being used for deep NN evaluations, the FPGA 402 is able to effectively eliminate the need for 2 * 24 = 48 CPU cores, which translates to high power savings. For example, assuming a bloated estimate of 25W for FPGA power consumption and a reasonable average power consumption of 100W / 12 = 8.33W for CPU cores, the net power savings would be approximately 48 * 8.33W - 25W = 375W per server blade. Calculated another way without the FPGA, the power usage would be 3 * 8.33W = 25W per session, while the power per session with the FPGA would be 8.33W + 25W / 24 = 9.37W.

[0049] When scaled to a large number of users, the 3 - fold increase in computational power required by the pure software deep NN approach compared to using just a single CPU core per session (and simply a single FPGA for 24 sessions) makes the pure software approach cost - prohibitive, even though the deep NN provides greater recognition accuracy and thus a better user experience when incorporated into speech recognition and translation services. Thus, the inclusion of the FPGA 402 enables the deep NN to be incorporated into speech recognition and translation services.

[0050] Typically, performing deep NN evaluations on an FPGA will require very high bandwidth requirements. One of the main difficulties with FPGA implementations is the management of memory bandwidth. For an exemplary Internet translator, there are approximately 500 million 16 - bit neural network weights that must be processed for each complete evaluation of the neural network. For each evaluation of the neural network, the FPGA must load 50M * 2 bytes = 100M bytes of data from memory. To meet the performance specifications, the neural network must be evaluated 100 times per second per session. Even for a single session, this means a memory bandwidth requirement of 100 * 100MB = 10GB / second for the FPGA. The absolute peak memory bandwidth of a typical FPGA memory interface is approximately 12.8GB / second, but this is difficult to achieve and assumes perfect operating conditions with no other activities in the system. If one considers that the FPGA's task is to handle N = 24 such sessions simultaneously, the problem seems intractable. However, Figures 3B to 3D illustrated in Figure 4 and reflected in the specific implementation for N = 24 in

[0051] First consider the case of a single session, which can be regarded as Figure 3B As illustrated in Figure 3B , the memory bandwidth requirement can be reduced by batching the input data in groups of K input data sets (of the observation vectors). By delaying processing until K input data sets have been accumulated, and then loading the neural network weight data all at once for all K inputs, the effective memory bandwidth required drops by a factor of K (while delaying the speech recognition output by the duration corresponding to K - 1 vectors, e.g., (K - 1)*10 ms). For example, if the memory bandwidth requirement is 10 GB / second and K = 2, the effective memory bandwidth required is 10 GB / second / 2 = 5 GB / second, which is a more manageable number. Larger values of K result in lower effective memory bandwidth and can be chosen to reduce the memory bandwidth requirement to a manageable number for the application. This comes at the cost of increased computational latency, since the input data sets are delayed until K have been accumulated, but it is a good trade-off in those specific cases where maintaining throughput may be more important than latency.

[0052] In the case of processing N simultaneous sessions, such as Figure 3C and 3D illustrated in Figure 3C and 3D , where N = 2 and K = 1 and K = 2 respectively, each session uses a queue of K input data sets in use simultaneously and N such queues (for the example illustrated in Figure 4 where N = 24 and K = 2). The input data sets can be arranged such that during the layer in layer logic 405 (e.g., matrix multiplication or other weighted processing steps, which can be performed on intermediate result 406, which is then re-queued 407 to utilize new weights for the next layer processing) in a locked step, the exact same weights are used to process all N queues simultaneously across all queues. Thus, the neural network weight data is loaded only once for all N sessions (for each layer of processing), and the memory bandwidth requirement for the FPGA remains the same as when only a single session is being processed. Figure 4 Figure 4 is a block diagram of components of a computing device or system that can be used to perform some of the processes described herein. Referring to Figure 5 , system 500 can include one or more blade server devices, stand-alone server devices, personal computers, routers, hubs, switches, bridges, firewall devices, intrusion detection devices, mainframe computers, network-attached storage devices, and other types of computing devices. The hardware can be configured according to any suitable computer architecture such as a symmetric multiprocessing (SMP) architecture or a non-uniform memory access (NUMA) architecture. Thus, more or fewer elements described with respect to system 500 can be incorporated to implement a particular computing system.

[0053] Figure 5 Figure 5 is a block diagram of components of a computing device or system that can be used to perform some of the processes described herein. Referring to Figure 5 , system 500 can include one or more blade server devices, stand-alone server devices, personal computers, routers, hubs, switches, bridges, firewall devices, intrusion detection devices, mainframe computers, network-attached storage devices, and other types of computing devices. The hardware can be configured according to any suitable computer architecture such as a symmetric multiprocessing (SMP) architecture or a non-uniform memory access (NUMA) architecture. Thus, more or fewer elements described with respect to system 500 can be incorporated to implement a particular computing system. Figure 5 Figure 5 , system 500 can include one or more blade server devices, stand-alone server devices, personal computers, routers, hubs, switches, bridges, firewall devices, intrusion detection devices, mainframe computers, network-attached storage devices, and other types of computing devices. The hardware can be configured according to any suitable computer architecture such as a symmetric multiprocessing (SMP) architecture or a non-uniform memory access (NUMA) architecture. Thus, more or fewer elements described with respect to system 500 can be incorporated to implement a particular computing system.

[0054] System 500 may include a processing system 505, which may include one or more processing devices, such as a central processing unit (CPU) having one or more CPU cores, a microprocessor, or other circuitry that fetches and executes software 510 from a storage system 520. The processing system 505 may be implemented within a single processing device but may also be distributed across multiple processing devices or subsystems that cooperate when executing program instructions.

[0055] One or more processing devices of the processing system 505 may include a multi-processor or multi-core processor and may operate according to one or more suitable instruction sets, including but not limited to a reduced instruction set computing (RISC) instruction set, a complex instruction set computing (CISC) instruction set, or a combination thereof. In certain embodiments, one or more digital signal processors (DSPs) may be included as part of the system's computer hardware in lieu of or in addition to a general-purpose CPU.

[0056] Software 510 may include any computer-readable storage medium that can be read by the processing system 505 and that is capable of storing software 510 including instructions for performing various processes, with neural network evaluation executed on an FPGA forming part of the various processes. Software 510 may also include additional processes, programs, or components, such as operating system software, database management software, or other application software. Software 510 may also include firmware or some other form of machine-readable processing instructions executable by the processing system 505. In addition to storing software 510, the storage system 520 may also store matrix weights and other data sets for performing neural network evaluations. In some cases, the manager agent 240 is at least partially stored on a computer-readable storage medium that forms part of the storage system 520 and that implements virtual memory and / or non-virtual memory.

[0057] The storage system 520 may include volatile media and non-volatile media, removable media and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data.

[0058] Although storage system 520 is shown as a box, storage system 520 represents on-chip storage and external storage available to computing system 500. Storage system 520 can include various storage media, such as random access memory (RAM), read-only memory (ROM), magnetic disks, optical disks, CDs, DVDs, flash memory, solid-state memory, phase change memory, or any other suitable storage medium. Some embodiments can include either or both virtual memory and non-virtual memory. In any case, the storage medium does not include a propagated signal or carrier wave. In addition to the storage medium, in some embodiments, storage system 520 can also include a communication medium through which software and data can communicate internally or externally.

[0059] Storage system 520 can be implemented as a single storage device or can be implemented across multiple storage devices or subsystems that are co-located or distributed relative to each other. In some cases, processing system 505 can access storage system 520 (or a portion of storage system 520) via a system bus. Storage system 520 can include additional elements, such as a controller, that can communicate with processing system 505.

[0060] Computing system 500 also includes FPGA 530 for performing neural network evaluations. Multiple FPGAs can be obtained in a data center. In some cases, multiple FPGAs can be incorporated into a daughter card and housed with a subset of servers. Alternatively, a single FPGA can be housed in a single server, where services that require more than one FPGA can be mapped across FPGAs residing in multiple servers and / or services that require more than one server can access a single FPGA residing at one of the servers in which the FPGA resides. In some cases, one or more FPGAs can be housed separately from the servers. When incorporated in the same server, the (one or more) FPGAs can be coupled to processing system 505 on the same board or on separate boards interfaced using a communication interface technology such as PCIe (PCI Express).

[0061] Communication interface 540 is included to provide a communication connection and device that allows device 500 to communicate with other computing systems (not shown) via a communication network or a collection of networks (not shown) or over the air. Examples of connections and devices that collectively allow inter-system communication can include network interface cards, antennas, power amplifiers, RF circuits, transceivers, and other communication circuits. The connections and devices can communicate via a communication medium to exchange communications with other computing systems or a network of systems, the communication medium such as metal, glass, air, or any other suitable communication medium. The above communication medium, network, connections, and devices are well known and need not be discussed in detail herein.

[0062] It should be noted that many components of the device 500 can be included on a system-on-chip (SoC) device. These components can include, but are not limited to, the processing system 505, the components of the storage system 520, and even the components of the communication interface 540.

[0063] Certain aspects of the present invention provide the following non-limiting embodiments.

[0064] Example 1. A method of performing a neural network process, the method comprising: receiving, at a field programmable gate array (FPGA), a batch of input data for accelerated processing of neural network evaluation, wherein the batch of input data includes at least two feature vectors; loading the FPGA with a first layer weight set for neural network evaluation from an external memory; and applying the first layer weight set to the batch of input data within the FPGA to generate an intermediate result.

[0065] Example 2. The method according to Example 1, wherein the at least two feature vectors include one observation vector from each of at least two data streams.

[0066] Example 3. The method according to Example 2, wherein the neural network evaluation is a recurrent neural network evaluation.

[0067] Example 4. The method according to Example 1, wherein the at least two feature vectors include at least two observation vectors from each of at least two data streams.

[0068] Example 5. The method according to Example 1, wherein the at least two feature vectors include at least two observation vectors from a single data stream.

[0069] Example 6. The method according to any one of Examples 1-5, further comprising: after applying the first layer weight set to the batch, loading the FPGA with a second layer weight set for neural network evaluation from an external memory; and applying the second layer weight set to the intermediate result within the FPGA.

[0070] Example 7. The method according to any one of Examples 1, 2, or 4-6, wherein the neural network evaluation is a deep neural network multi-layer perception evaluation.

[0071] Example 8. The method according to any one of Examples 1-7, wherein the batch of input data is received from at least one core.

[0072] Example 9. The method according to any one of Examples 1-7, wherein a batch of input data is received from other logic on the FPGA.

[0073] Example 10. The method according to any one of Examples 1-7, wherein the batch of input data is received from other processing units.

[0074] Example 11. One or more computer-readable storage media having instructions stored thereon that, when executed by a processing system, direct the processing system to manage memory bandwidth for deep learning applications by: queuing at a field-programmable gate array (FPGA) a batch of at least two observation vectors from at least one core; loading at least one weighted data set on the FPGA, each weighted data set in the at least one weighted data set being loaded once for a batch of at least two observation vectors queued at the FPGA; and directing an evaluation output from the FPGA to at least one core for further processing.

[0075] Example 12. The medium of Example 11, wherein the instructions that queue the batch of at least two observation vectors from at least one core at the FPGA direct one observation vector from each of at least two cores to be queued at the FPGA.

[0076] Example 13. The medium of Example 11, wherein the instructions that queue the batch of at least two observation vectors from at least one core at the FPGA direct at least two observation vectors from each of at least two cores to be queued at the FPGA.

[0077] Example 14. A system comprising: one or more storage media; a plurality of processing cores; a service stored on at least one of the one or more storage media and executed on at least the plurality of processing cores; a parallel processor in communication with the plurality of cores to perform a neural network evaluation on a batch of data for a process of the service; and a weight data set for the neural network evaluation stored on at least one of the one or more storage media.

[0078] Example 15. The system of Example 14, wherein the parallel processor is a field-programmable gate array (FPGA).

[0079] Example 16. The system of Example 14 or 15, wherein the parallel processor receives one observation vector from each of the plurality of cores as the batch of data.

[0080] Example 17. The system of Example 16, wherein the neural network evaluation includes a recurrent neural network evaluation.

[0081] Example 18. The system of any of Examples 14-16, wherein the neural network evaluation includes a deep neural network multi-layer perception evaluation.

[0082] Example 19. The system of any of Examples 14, 15, or 18, wherein the parallel processor receives at least two observation vectors from each of the plurality of cores as the batch of data.

[0083] Example 20. The system according to any one of Examples 14-19, wherein the service includes a speech recognition service.

[0084] Example 21. The system according to any one of Examples 14-20, further comprising: a manager agent, at least partially stored on at least one of one or more storage media, the manager agent, when executed, bootstrapping the system to: queue the batch of data from at least one of a plurality of processing cores at a parallel processor; load at least one weighted data set of a weight data set onto the parallel processor, each weighted data set in the at least one weighted data set being loaded once in batches; and direct an evaluation output from the parallel processor to a plurality of processing cores.

[0085] Example 22. The system according to Example 21, wherein the manager agent bootstraps the system to queue the batch of data at the parallel processor by directing at least one observation vector from each of at least two of the plurality of cores to the parallel processor.

[0086] Example 23. The system according to Example 21, wherein the manager agent bootstraps the system to queue the batch of data at the parallel processor by directing at least two observation vectors from each of at least two of the plurality of cores to the parallel processor.

[0087] It should be understood that the examples and embodiments described herein are for illustrative purposes only and will suggest to those skilled in the art various modifications or changes therefrom and such modifications or changes should be included within the spirit and scope of the present application.

[0088] Although the subject matter has been described in language specific to structural features and / or acts, it is to be understood that the subject matter defined in the claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as examples of implementing the claims, and other equivalent features and acts are intended to be within the scope of the claims.

Claims

1. A method for performing a neural network process, comprising: loading a parallel processor with a first set of weights for a neural network process from an external memory; sequentially applying the first set of weights to at least two input data sets such that when each of the at least two input data sets is separately and concurrently processed by the parallel processor, one of the at least two input data sets is processed after another of the at least two input data sets; queuing intermediate results of the at least two input data sets; loading the parallel processor with a second set of weights for the neural network process from the external memory; sequentially applying the second set of weights to the intermediate results of the at least two input data sets; and repeating the queuing, loading, and applying until the neural network process is completed for the at least two input data sets.

2. The method according to claim 1, wherein, The input data includes data for processing for speech recognition or translation services.

3. The method according to claim 1, wherein, The input data includes data for processing for computer vision applications.

4. The method according to claim 1, wherein The input data includes data for image processing or recognition services.

5. The method according to claim 1, wherein The input data includes data for natural language processing.

6. The method according to claim 1, wherein, The input data includes data for audio recognition.

7. The method according to claim 1, wherein The input data includes data for bioinformatics.

8. The method according to claim 1, wherein, The input data includes data for weather prediction.

9. A computing system, comprising: a parallel processor having a plurality of available parallel processing streams and a buffer with a queue depth of at least 2 for each stream such that, prior to processing on a particular weight data set, one stream of parallel input data from at least two parallel input data sets is stored in the buffer, wherein the at least two parallel input data sets are queued at the buffer and one of the at least two parallel input data sets is processed after another; and a storage unit storing a weight data set including the particular weight data set for neural network evaluation.

10. The computing system according to claim 9, wherein, The parallel processor is a field programmable gate array (FPGA).

11. The computing system according to claim 9, further comprising: a plurality of processing cores operatively coupled to the parallel processor, the plurality of processing cores performing audio, speech, or image processing.

12. The computing system according to claim 9, wherein, The parallel processor is operatively coupled to receive at least one observation vector of a process executed on at least one of the plurality of processing cores and to transmit an evaluation output to a corresponding at least one of the plurality of processing cores.

13. One or more storage media having instructions stored thereon that, when executed, direct a computing system: Load a parallel processor with a first set of weights for a neural network process from an external memory, wherein, The first weight set is sequentially applied to at least two input data sets such that when each of the at least two input data sets is separately and concurrently processed by the parallel processor, one of the at least two input data sets is processed after another one of the at least two input data sets, wherein after the first weight set is applied, intermediate results of the at least two input data sets are queued; and The parallel processor is loaded with a second weight set for the neural network process from the external memory, wherein the second weight set is sequentially applied to the intermediate results of the at least two input data sets.

14. The medium according to claim 13, wherein The input data includes data for processing for speech recognition or translation services.

15. The medium according to claim 13, wherein, The input data includes data for processing for computer vision applications.

16. The medium according to claim 13, wherein The input data includes data for image processing or recognition services.

17. The medium according to claim 13, wherein The input data includes data for natural language processing.

18. The medium according to claim 13, wherein The input data includes data for audio recognition.

19. The medium according to claim 13, wherein, The input data includes data for bioinformatics.

20. The medium according to claim 13, wherein, The input data includes data for weather prediction.

Citation Information

Patent Citations

  • Neural network missing data estimation machine and evaluation method based on FPGA

    CN101246508A

  • System and method for increasing signal real-time mode recognizing processing speed in DSP+FPGA frame

    CN101673343A