Deep learning inference system

The deep learning inference system addresses inefficiencies in memory usage and throughput by employing a shared global memory space and local memory spaces for each client, enabling efficient parallel processing of inferences and improving overall efficiency.

JP7683735B2Active Publication Date: 2025-05-27NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023565720
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-05-27
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Existing deep learning inference systems using multi-layer neural networks face inefficiencies in memory usage and request throughput due to the lack of shared memory space among processors, preventing multi-core and pipeline processing, and resulting in wasted memory space for identical requests.

Method used

A deep learning inference system with a shared global memory space for operation codes and parameters, and local memory spaces for each client, allowing multiple processors to read and write data efficiently, enabling parallel processing of inferences using the same model and improving memory usage and throughput.

Benefits of technology

The proposed system enhances computer efficiency and energy efficiency by allowing parallel execution of inferences with the same model, reducing memory waste, and improving request throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683735000001
    Figure 0007683735000001
  • Figure 0007683735000002
    Figure 0007683735000002
  • Figure 0007683735000003
    Figure 0007683735000003
Patent Text Reader

Abstract

This deep learning inference system comprises a DRAM (100a), and processors (101a-1, 101a-2) for reading an operation code (200) and a parameter (201) from a global memory space (1007) of the DRAM (100a) and performing computation of a neural network. The processors (101a-1, 101a-2) read data to be processed (202-1, 202-2) from local memory spaces (1000, 1001) that correspond to subject clients, perform computation, and then store the results of computation in the local memory spaces (1000, 1001) that correspond to the subject clients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deep learning inference system that performs inference serving using a multi-layer neural network.

Background Art

[0002] In recent years, there have been many services that perform information processing using a multi-layer neural network and utilize the results. Giving an operation of neural network calculation, parameters of the neural network, and data to be processed to an arithmetic unit to obtain processed data is called inference. Inference requires a large number of operations and memory. For this reason, inference may be performed on a server.

[0003] A client sends a request and data to be processed to a server and receives the result of the process as a response. Providing such a service is inference serving. Various methods have been proposed for inference serving (see Non-Patent Document 1).

[0004] When using an FPGA (field-programmable gate array) accelerator as an arithmetic unit for inference serving, a method of constructing a Neumann-type processor on the FPGA accelerator is common (see Non-Patent Document 2). The internal structure of the Neumann-type processor is generalized and shown in FIG. 5.

[0005] The operation code 200 of the neural network calculation, the parameters 201 of the neural network, and the data to be processed are stored in a DRAM (Dynamic Random Access Memory) 100. In the example of FIG. 5, the data to be processed and the data during the calculation are used as input data 202.

[0006] The Instruction Fetch Module 102 reads the operation code 200 from the DRAM 100 and transfers it to the Load Module 103, the Compute Module 104, and the Store Module 105.

[0007] The Load Module 103 reads the input data 202 from the DRAM 100, batches a plurality of input data 202, and transfers it to the Compute Module 104. The Compute Module 104 performs neural network operations using the input data 202 and the parameter 201 according to the operation code 200 transferred from the Instruction Fetch Module 102. The Compute Module 104 is equipped with an ALU (Arithmetic Logic Unit) 1040 and a GEMM (General matrix multiply) circuit 1041. After performing the operation according to the operation code 200, the Compute Module 104 transfers the operation result to the Store Module 105.

[0008] The Store Module 105 stores the operation result by the Compute Module 104 in the DRAM 100. At this time, not only the processed data is stored in the DRAM 100 as the output data 203, but also the data during the operation may be temporarily stored as the output data 203. The data during the operation becomes the input data 202 to the Load Module 103.

[0009] The DRAM 100 is a memory external to the processor. The DRAM 100 stores the operation code 200 for neural network operations, the neural network parameters 201, the input data 202, and the output data 203 as described above. These data exist for each request client. The memory space of the DRAM 100 when a memory area is allocated for each request client is as shown in FIG. 6. In the example of FIG. 6, each row 1000 to 1006 of the memory space schematically represents the memory space for each client.

[0010] When there are multiple Neumann-type processors, since there is no memory space shared by each processor in the memory space of the conventional DRAM 100, there are the following problems. (I) The processor cannot perform multi-core processing on requests from clients. (II) The processor cannot perform pipeline processing on requests from clients. (III) Even when requests from clients are the same requests (for example, when the parameters are the same), the processor needs to read from the memory for each client, resulting in wasted memory space.

Prior Art Documents

Non-Patent Documents

[0011]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0012] The present invention is made to solve the above problems, and an object thereof is to provide a deep learning inference system capable of improving computer efficiency and executing inference services with high energy efficiency.

Means for Solving the Problems

[0013] The deep learning inference system of the present invention includes a memory having a global memory space in which operation codes of neural network operations and parameters of the neural network are stored, and a local memory space secured for each client that transmits a request, and in response to a request from a client, a plurality of processors configured to perform processing of reading the operation codes and the parameters from the global memory space and performing operations of the neural network for each of the plurality of clients, and each processor reads data to be processed from the local memory space corresponding to the target client, performs operations of the neural network, and stores the operation result in the local memory space corresponding to the target client.

[0014] In addition, the deep learning inference system of the present invention includes a memory having a global memory space in which data to be processed by a convolutional neural network is stored and a local memory space secured for each of a plurality of kernels of the convolutional neural network, and a plurality of processors configured to perform a process of reading the data to be processed from the global memory space and performing a convolutional operation for each of the plurality of kernels. Each processor reads a convolutional operation instruction code and kernel parameters of a target kernel from the local memory space corresponding to the target kernel and performs a convolutional operation, and stores the operation result in the local memory space corresponding to the target kernel.

[0015] In addition, the deep learning inference system of the present invention includes a memory having a global memory space in which intermediate data of a multi-layer neural network is stored and a local memory space secured for each layer of the multi-layer neural network, and a plurality of processors configured to perform a process of reading an operation code and parameters of an operation of a target layer from the local memory space corresponding to the target layer of the multi-layer neural network and performing the operation of the target layer for each layer of the multi-layer neural network. Among the processors, a processor targeting an upper layer reads data to be processed from the local memory space corresponding to the target layer and performs an operation of the target layer, and stores the operation result as intermediate data in the global memory space. Among the processors, a processor targeting a lower layer reads the intermediate data to be processed from the global memory space and performs an operation of the target layer, and stores the operation result in the local memory space corresponding to the target layer. In addition, one configuration example of the deep learning inference system of the present invention further includes a plurality of cache memories respectively provided between the memory and the plurality of processors and configured to store data, codes, and parameters read and written between the memory and the plurality of processors.

Effects of the Invention

[0016] According to the present invention, operation codes and parameters are stored in a global memory space shared by a plurality of processors. In the present invention, although the data to be processed is different, for inferences using the same model, a plurality of inferences can be executed in parallel by different processors. As a result, in the present invention, it is possible to save memory space and improve request throughput.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0018] [Principle of the Invention] The present invention provides a shared memory space on the memory space of a deep learning inference system and allows data to be shared by each Neumann type processor.

[0019] [First Embodiment] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing the configuration of an arithmetic unit provided in a server of a deep learning inference system according to a first embodiment of the present invention. The arithmetic unit is composed of a DRAM 100a and a plurality of Neumann type processors 101a-1, 101a-2. Each of the processors 101a-1, 101a-2 is composed of an instruction fetch module 102a, a load module 103a, a compute module 104, and a store module 105a.

[0020] In the memory space of the DRAM 100a of this embodiment, there are a local memory space 1000~1006 secured for each client and a global memory space 1007 secured for sharing by a plurality of processors 101a-1, 101a-2.

[0021] When the server's CPU (Central Processing Unit) 110 receives an inference request from client A via the network, it stores the operation code 200 of the neural network operation corresponding to the inference request and the neural network parameters 201 in the global memory space 1007 of the DRAM 100a. Also, the CPU 110 stores the data to be processed received from client A as input data 202-1 in the local memory space 1000 of the DRAM 100a corresponding to client A.

[0022] Also, when the CPU 110 receives an inference request and data to be processed from client B via the network, it stores the data to be processed as input data 202-2 in the local memory space 1001 of the DRAM 100a corresponding to client B. The inference request specifies which model to use for inference. In this embodiment, it is assumed that the inference requests received from each of clients A and B specify the same model.

[0023] The instruction fetch module 102a of the processor 101a-1 reads the operation code 200 and the parameter 201 from the global memory space 1007 of the DRAM 100a and transfers them to the load module 103a, the compute module 104, and the store module 105a of the processor 101a-1.

[0024] The load module 103a of the processor 101a-1 reads the input data 202-1 from the local memory space 1000 of the DRAM 100a corresponding to the client A that sent the inference request, batches the plurality of input data 202-1, and transfers them to the compute module 104.

[0025] The compute module 104 of the processor 101a-1 performs neural network operations using the input data 202-1 and the parameter 201 according to the operation code 200 transferred from the instruction fetch module 102a. The compute module 104 transfers the operation result to the store module 105a.

[0026] The store module 105a of the processor 101a-1 stores the operation result by the compute module 104 as the output data 203-1 in the local memory space 1000 of the DRAM 100a corresponding to the client A.

[0027] The CPU 110 of the server reads the processed data from the local memory space 1000 of the DRAM 100a and returns this data to the client A as a response to the inference request.

[0028] On the other hand, the instruction fetch module 102a of the processor 101a-2 reads the operation code 200 and the parameter 201 from the global memory space 1007 of the DRAM 100a and transfers them to the load module 103a, the compute module 104, and the store module 105a of the processor 101a-2.

[0029] The load module 103a of the processor 101a-2 reads the input data 202-2 from the local memory space 1001 of the DRAM 100a corresponding to the client B that sent the inference request, batches the plurality of input data 202-2, and transfers it to the compute module 104.

[0030] The compute module 104 of the processor 101a-2 performs neural network operations using the input data 202-2 and the parameter 201 according to the operation code 200 transferred from the instruction fetch module 102a. The compute module 104 transfers the operation result to the store module 105a.

[0031] The store module 105a of the processor 101a-2 stores the operation result by the compute module 104 as the output data 203-2 in the local memory space 1001 of the DRAM 100a corresponding to the client B.

[0032] The CPU 110 of the server reads the processed data from the local memory space 1001 of the DRAM 100a and returns this data to the client B as a response to the inference request.

[0033] As described above, in this embodiment, the operation code 200 and the parameter 201 are stored in the global memory space 1007 shared by the plurality of processors 101a-1, 101a-2. In this embodiment, although the data to be processed is different, for inferences with the same model (inferences with the same operation code 200), a plurality of inferences can be executed in parallel by different processors 101a-1, 101a-2. As a result, in this embodiment, the memory space can be saved and the request throughput can be improved.

[0034] [Second Embodiment] Next, a second embodiment of the present invention will be described. FIG. 2 is a block diagram showing the configuration of an arithmetic unit provided in a server of a deep learning inference system according to the second embodiment of the present invention. The arithmetic unit is composed of a DRAM 100b and a plurality of Neumann type processors 101b-1 and 101b-2. Each of the processors 101b-1 and 101b-2 is composed of an instruction fetch module 102b, a load module 103b, a compute module 104b, and a store module 105b.

[0035] In the case of a convolutional neural network, convolution operations are performed using multiple types of filters (kernels), and a weighted sum of the results of multiple convolution operations is calculated. In the memory space of the DRAM 100b of this embodiment, there are local memory spaces 1000 to 1006 reserved for each of the multiple types of kernels of the convolutional neural network, and a global memory space 1007 reserved for sharing by the multiple processors 101b-1 and 101b-2. The input data 202 for the convolution operation is stored in the global memory space 1007 by the CPU 110 of the server.

[0036] The instruction fetch module 102b of the processor 101b-1 reads the kernel parameter 204-1 and the convolution operation instruction code 205-1 of the kernel from the local memory space 1000 of the DRAM 100b and transfers them to the load module 103b, the compute module 104b, and the store module 105b of the processor 101b-1.

[0037] The load module 103b of the processor 101b-1 reads the input data 202 from the global memory space 1007 of the DRAM 100b and transfers it to the compute module 104b.

[0038] The compute module 104b of the processor 101b-1 performs a convolution operation using the input data 202 and the kernel parameter 204-1 according to the convolution operation instruction code 205-1 transferred from the instruction fetch module 102b. The compute module 104b transfers the operation result to the store module 105b.

[0039] The store module 105b of the processor 101b-1 stores the operation result by the compute module 104b as the output data 203-1 in the local memory space 1000 of the DRAM 100b.

[0040] On the other hand, the instruction fetch module 102b of the processor 101b-2 reads the kernel parameter 204-2 and the convolution operation instruction code 205-2 of the kernel from the local memory space 1001 of the DRAM 100b, and transfers them to the load module 103b, the compute module 104b, and the store module 105b of the processor 101b-2.

[0041] The load module 103b of the processor 101b-2 reads the input data 202 from the global memory space 1007 of the DRAM 100b and transfers it to the compute module 104b.

[0042] The compute module 104b of the processor 101b-2 performs a convolution operation using the input data 202 and the kernel parameter 204-2 according to the convolution operation instruction code 205-2 transferred from the instruction fetch module 102b. The compute module 104b transfers the operation result to the store module 105b.

[0043] The store module 105b of the processor 101b-2 stores the operation result by the compute module 104b as the output data 203-2 in the local memory space 1001 of the DRAM 100b.

[0044] As described above, in this embodiment, the input data 202 is stored in the global memory space 1007 shared by the plurality of processors 101b-1 and 101b-2, and the kernel parameters 204-1 and 204-2 and the convolution operation instruction codes 205-1 and 205-2 are stored in different local memory spaces 1000 to 1006 for each convolution operation. As a result, in this embodiment, a plurality of convolution operations can be executed in parallel by different processors 101b-1 and 101b-2, and the inference throughput can be improved.

[0045] [Third Embodiment] Next, a third embodiment of the present invention will be described. FIG. 3 is a block diagram showing the configuration of an arithmetic unit provided in a server of a deep learning inference system according to the third embodiment of the present invention. The arithmetic unit is composed of a DRAM 100c and a plurality of von Neumann type processors 101c-1 and 101c-2. Each processor 101c-1 and 101c-2 is composed of an instruction fetch module 102c, a load module 103c, a compute module 104c, and a store module 105c.

[0046] In the case of a multi-layer neural network, the upper-layer operations can be performed by the processor 101c-1 and the lower-layer operations can be performed by the processor 101c-2, so that pipeline processing can be performed. In the memory space of the DRAM 100c of this embodiment, there are local memory spaces 1000 to 1006 reserved for each layer of the multi-layer neural network and a global memory space 1007 reserved for sharing by the plurality of processors 101c-1 and 101c-2.

[0047] The CPU 110 of the server stores the operation code 200-1 and the parameter 201-1 of the upper-layer operation of the multi-layer neural network in the local memory space 1000 of the DRAM 100c, and stores the operation code 200-2 and the parameter 201-2 of the lower-layer operation in the local memory space 1001 of the DRAM 100c. Also, the CPU 110 stores the data to be processed received from the client as input data 202 in the local memory space 1000.

[0048] The instruction fetch module 102c of the processor 101c-1 reads the operation code 200-1 and the parameter 201-1 from the local memory space 1000 of the DRAM 100c and transfers them to the load module 103c, the compute module 104c, and the store module 105c of the processor 101c-1.

[0049] The load module 103c of the processor 101c-1 reads the input data 202 from the local memory space 1000 of the DRAM 100c and transfers it to the compute module 104c.

[0050] The compute module 104c of the processor 101c-1 performs the upper-layer operation of the multi-layer neural network using the input data 202 and the parameter 201-1 according to the operation code 200-1 transferred from the instruction fetch module 102c. The compute module 104c transfers the operation result to the store module 105c.

[0051] The store module 105c of the processor 101c-1 stores the operation result by the compute module 104c as intermediate data 206 in the global memory space 1007 of the DRAM 100c.

[0052] Next, the instruction fetch module 102c of the processor 101c-2 reads the operation code 200-2 and the parameter 201-2 from the local memory space 1001 of the DRAM 100c and transfers them to the load module 103c, the compute module 104c, and the store module 105c of the processor 101c-2.

[0053] The load module 103c of the processor 101c-2 reads the intermediate data 206 from the global memory space 1007 of the DRAM 100c and transfers it to the compute module 104c.

[0054] The compute module 104c of the processor 101c-2 performs operations on the lower layer of the multi-layer neural network using the intermediate data 206 and the parameter 201-2 according to the operation code 200-2 transferred from the instruction fetch module 102c. The compute module 104c transfers the operation result to the store module 105c.

[0055] The store module 105c of the processor 101c-2 stores the operation result by the compute module 104c as the output data 203 in the local memory space 1001 of the DRAM 100c.

[0056] As described above, in this embodiment, the intermediate data 206 of the operations of the multi-layer neural network is stored in the global memory space 1007 shared by the plurality of processors 101c-1 and 101c-2. As a result, in this embodiment, pipeline processing of the operations of the multi-layer neural network becomes possible, and the inference throughput can be improved.

[0057] [Fourth Embodiment] Next, a fourth embodiment of the present invention will be described. FIG. 4 is a block diagram showing the configuration of an arithmetic unit provided in a server of a deep learning inference system according to the fourth embodiment of the present invention. The arithmetic unit is composed of a DRAM 100a, a plurality of Neumann type processors 101a-1 and 101a-2, and cache memories 106-1 and 106-2 provided between the DRAM 100a and the respective processors 101a-1 and 101a-2.

[0058] The DRAM 100a is as described in the first embodiment. The CPU 110 of the server stores the operation code 200 and the parameter 201 stored in the global memory space 1007 of the DRAM 100a in the cache memories 106-1 and 106-2. Further, the CPU 110 stores the input data 202-1 stored in the local memory space 1000 of the DRAM 100a in the cache memory 106-1, and stores the input data 202-2 stored in the local memory space 1001 of the DRAM 100a in the cache memory 106-2.

[0059] The instruction fetch module 102a of the processor 101a-1 reads the operation code 200 and the parameter 201 from the cache memory 106-1 and transfers them to the load module 103a, the compute module 104, and the store module 105a of the processor 101a-1.

[0060] The load module 103a of the processor 101a-1 reads the input data 202-1 from the cache memory 106-1, batches a plurality of input data 202-1, and transfers them to the compute module 104.

[0061] The compute module 104 of the processor 101a-1 performs neural network operations using the input data 202-1 and the parameter 201 according to the operation code 200 transferred from the instruction fetch module 102a. The store module 105a of the processor 101a-1 stores the operation result by the compute module 104 in the cache memory 106-1.

[0062] The CPU 110 of the server writes the processed data stored in the cache memory 106-1 to the local memory space 1000 of the DRAM 100a corresponding to the client A, and then reads the processed data from the local memory space 1000 and returns it to the client A.

[0063] On the other hand, the instruction fetch module 102a of the processor 101a-2 reads the operation code 200 and the parameter 201 from the cache memory 106-2 and transfers them to the load module 103a, the compute module 104, and the store module 105a of the processor 101a-2.

[0064] The load module 103a of the processor 101a-2 reads the input data 202-2 from the cache memory 106-2, batches a plurality of input data 202-2, and transfers them to the compute module 104.

[0065] The compute module 104 of the processor 101a-2 performs neural network operations using the input data 202-2 and the parameter 201 according to the operation code 200 transferred from the instruction fetch module 102a. The store module 105a of the processor 101a-2 stores the operation result by the compute module 104 in the cache memory 106-2.

[0066] The CPU 110 of the server writes the processed data stored in the cache memory 106-2 to the local memory space 1001 of the DRAM 100a corresponding to the client B, and then reads the processed data from the local memory space 1001 and returns it to the client B.

[0067] As described above, in this embodiment, by providing cache memories 106-1 and 106-2 between DRAM 100a and respective processors 101a-1 and 101a-2, the memory latency of DRAM 100a can be hidden and the inference latency can be shortened.

[0068] In this embodiment, the example in which cache memories 106-1 and 106-2 are applied to the first embodiment has been described, but it goes without saying that they may also be applied to the second and third embodiments. When applying cache memories 106-1 and 106-2 to the second embodiment, input data 202, kernel parameter 204-1, convolution operation instruction code 205-1, and output data 203-1 may be stored in cache memory 106-1, and input data 202, kernel parameter 204-2, convolution operation instruction code 205-2, and output data 203-2 may be stored in cache memory 106-1.

[0069] When applying cache memories 106-1 and 106-2 to the third embodiment, input data 202, operation code 200-1, and parameter 201-1 may be stored in cache memory 106-1, and operation code 200-2, parameter 201-2, intermediate data 206, and output data 203 may be stored in cache memory 106-2.

[0070] Also, in the first to fourth embodiments, the number of von Neumann type processors and cache memories is two, but it goes without saying that it may be three or more.

Industrial Applicability

[0071] The present invention can be applied to technologies for providing services using neural networks.

Explanation of Signs

[0072] 100a, 100b, 100c… DRAM, 101a-1, 101a-2, 101b-1, 101b-2, 101c-1, 101c-2… Neumann type processors, 102a, 102b, 102c… Instruction fetch modules, 103a, 103b, 103c… Load modules, 104, 104b, 104c… Compute modules, 105a, 105b, 105c… Store modules, 106-1, 106-2… Cache memories, 1000~1006… Local memory spaces, 1007… Global memory space.

Claims

1. A memory having a global memory space storing operation codes for operations of a neural network and parameters of the neural network, and a local memory space secured for each client that sends a request, and a plurality of processors configured to perform, for each of a plurality of clients, a process of reading the operation code and the parameters from the global memory space and performing an operation of the neural network in response to a request from the client. Each processor reads data to be processed from the local memory space corresponding to the target client, performs an operation of the neural network, and stores the operation result in the local memory space corresponding to the target client. A deep learning inference system characterized by that.

2. A memory having a global memory space storing data to be processed by a convolutional neural network and a local memory space secured for each of a plurality of kernels of the convolutional neural network, and a plurality of processors configured to perform, for each of the plurality of kernels, a process of reading the data to be processed from the global memory space and performing a convolution operation. Each processor reads a convolution operation instruction code and kernel parameters of the target kernel from the local memory space corresponding to the target kernel, performs a convolution operation, and stores the operation result in the local memory space corresponding to the target kernel. A deep learning inference system characterized by that.

3. A memory having a global memory space storing intermediate data of a multi-layer neural network and a local memory space secured for each layer of the multi-layer neural network, and a plurality of processors configured to perform, for each layer of the multi-layer neural network, a process of reading an operation code and parameters of an operation of the target layer from the local memory space corresponding to the target layer of the multi-layer neural network and performing an operation of the target layer. Among the processors, the processor targeting the upper layer reads data to be processed from the local memory space corresponding to the target layer, performs an operation of the target layer, and stores the operation result as intermediate data in the global memory space. Among the processors, the processor targeting the lower layer reads the intermediate data to be processed from the global memory space, performs calculations on the target layer, and stores the calculation result in the local memory space corresponding to the target layer. A deep learning inference system characterized by this.

4. In the deep learning inference system according to any one of claims 1 to 3, A deep learning inference system further comprising a plurality of cache memories respectively provided between the memory and the plurality of processors and configured to store data, code, and parameters read and written between the memory and the plurality of processors.

Citation Information

Patent Citations

  • Neural network processing elements incorporating computational and local memory elements

    JP2020515989A

  • Neural network devices and methods of operating the same

    US20180253635A1

  • Neural network processing

    US20200184320A1