Method and system for processing images based on neural network and processing system in memory

Mapping machine learning applications to PIM accelerators through pipelined methods solves the problems of poor computing efficiency and power performance in the existing technology, and realizes efficient machine learning inference.

CN112750068BActive Publication Date: 2025-05-09SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011171408.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-02
Filing Date
2020-10-28
Publication Date
2025-05-09
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

The prior art has difficulty mapping machine learning applications into hardware structures efficiently, resulting in poor computing efficiency and power performance.

Method used

The pipelined method is used to map machine learning applications to accelerator based on in-memory processing (PIM). Computation efficiency and power performance are improved through the combination of inter-layer pipeline, in-layer pipeline and batch pipeline.

Benefits of technology

It realizes the generation of multiple outputs to multiple images in one clock cycle, significantly improving the power performance and processing speed of machine learning inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112750068B_ABST
    Figure CN112750068B_ABST
Patent Text Reader

Abstract

The present application relates to a method and system for processing images based on a neural network and a processing system in a memory. The neural network includes an i-th layer (i is an integer greater than zero) and an (i+1)-th layer, and the method includes: for a first input image, processing the first i-th value of the i-th layer to generate a first (i+1)-th value for the (i+1)-th layer, for the first input image, processing the first (i+1)-th value of the (i+1)-th layer to generate an output value, and while processing the first (i+1)-th value for the first input image, for a second input image, processing the second i-th value of the i-th layer to generate a second (i+1)-th value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of embodiments of the present disclosure are generally related to machine learning. Background Art

[0002] The explosive growth of big data-driven machine learning (ML) applications and the prospect of a slowdown in Moore's Law have prompted the search for alternative application-specific hardware structures. With a focus on bringing computation into memory bit cells, processing-in-memory (PIM) has been proposed to accelerate ML inference applications. ML applications and networks need to be efficiently mapped onto the underlying hardware structure to extract the highest power performance.

[0003] The above information disclosed in this Background section is only for enhancement of understanding of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the invention

[0004] Various aspects of the embodiments of the present disclosure are directed to a system and method for mapping machine learning (ML) applications to an accelerator based on processing in memory (PIM) in a pipelined manner. According to some embodiments, pipelining includes inter-layer pipelining, intra-layer pipelining, and / or a combination of the two. In addition, the pipelining schemes of various embodiments allow multiple outputs to be generated for multiple images simultaneously within one clock cycle.

[0005] Aspects of embodiments of the present disclosure are directed to a mixed-signal PIM-based configurable hardware accelerator, where machine learning problems are mapped onto PIM sub-arrays in a pipelined manner.

[0006] Various aspects of the embodiments of the present disclosure are directed to a configurable hardware accelerator based on a mixed signal PIM that receives multiple images as input for recognition. The input activations of the multiple images are received on the accelerator in a pipelined manner.

[0007] According to some embodiments of the present disclosure, a method for inference of a pipelined neural network is provided, wherein the neural network includes an i-th layer (i is an integer greater than zero) and an (i+1)-th layer, and the method includes: for a first input image, processing a first i-th value of the i-th layer to generate a first (i+1)-th value for the (i+1)-th layer; for the first input image, processing the first (i+1)-th value of the (i+1)-th layer to generate an output value; and in parallel with processing the first (i+1)-th value for the first input image, processing the second i-th value of the i-th layer for a second input image to generate a second (i+1)-th value.

[0008] In some embodiments, processing the second i-th value for the second input image is performed in parallel with processing the first i-th value for the first input image.

[0009] In some embodiments, the first i-th value comprises a pixel value of the first input image, and the second i-th value comprises a pixel value of the second input image.

[0010] In some embodiments, the first i-th value comprises a value of a first feature map generated by a previous layer of the neural network, the first feature map corresponding to the first input image, and the second i-th value comprises a value of a second feature map generated by the previous layer of the neural network, the second feature map corresponding to the second input image.

[0011] In some embodiments, processing the first i-th value of the i-th layer for the first input image includes: applying the i-th filter associated with the i-th layer to the first i-th value of the i-th layer to generate the first (i+1)th value for the (i+1)th layer.

[0012] In some embodiments, processing the second i-th value of the i-th layer for the second input image includes: applying the i-th filter associated with the i-th layer to the second i-th value of the i-th layer to generate the second (i+1)th value for the (i+1)th layer.

[0013] In some embodiments, the i-th filter is a sliding convolution filter in the form of a p×q matrix, where p and q are integers greater than zero.

[0014] In some embodiments, applying the i-th filter comprises performing a matrix multiplication operation between the i-th filter and values ​​of the first i-th value that overlap with the i-th filter.

[0015] In some embodiments, processing the second i-th value of the i-th layer for the second input image is initiated at a time offset after initiation of processing the first i-th value of the i-th layer for the first input image, and wherein the time offset is greater than or equal to the number of clock cycles corresponding to a single step of the i-th filter.

[0016] According to some embodiments of the present disclosure, a system for inference of a pipelined neural network is provided, wherein the neural network includes multiple layers, wherein the multiple layers include an i-th layer (i is an integer greater than zero), an (i+1)-th layer, and an (i+2)-th layer, and the system includes: a processor; and a processor memory local to the processor, wherein the processor memory stores instructions, and when the instructions are executed by the processor, the processor executes: for a first input image, processing a first i-th value of the i-th layer to generate a first (i+1)-th value for the (i+1)-th layer; for the first input image, processing the first (i+1)-th value of the (i+1)-th layer to generate an output value; and in parallel with processing the first (i+1)-th value for the first input image, processing the second i-th value of the i-th layer for a second input image to generate a second (i+1)-th value.

[0017] In some embodiments, processing the second (i)th value for the second input image is performed in parallel with processing the first (i)th value for the first input image.

[0018] In some embodiments, the first i-th value comprises a pixel value of the first input image, and wherein the second i-th value comprises a pixel value of the second input image.

[0019] In some embodiments, the first i-th value comprises a value of a first feature map generated by a previous layer of the neural network, the first feature map corresponding to the first input image, and wherein the second i-th value comprises a value of a second feature map generated by the previous layer of the neural network, the second feature map corresponding to the second input image.

[0020] In some embodiments, processing the first i-th value of the i-th layer for the first input image includes: applying the i-th filter associated with the i-th layer to the first i-th value of the i-th layer to generate the first (i+1)th value for the (i+1)th layer.

[0021] In some embodiments, processing the second i-th value of the i-th layer for the second input image includes: applying the i-th filter associated with the i-th layer to the second i-th value of the i-th layer to generate the second (i+1)th value for the (i+1)th layer.

[0022] In some embodiments, the i-th filter is a sliding convolution filter in the form of a p×q matrix, where p and q are integers greater than zero.

[0023] In some embodiments, applying the i-th filter comprises performing a matrix multiplication operation between the i-th filter and a value in the first i-th value that overlaps with the i-th filter.

[0024] In some embodiments, processing the second i-th value of the i-th layer for the second input image is started at a time offset after starting processing the first i-th value of the i-th layer for the first input image, and the time offset is greater than or equal to the number of clock cycles corresponding to a single step of the i-th filter.

[0025] According to some embodiments of the present disclosure, a configurable processing-in-memory (PIM) system configured to implement a neural network is provided, the system comprising: a first at least one PIM subarray configured to perform a filtering operation of an i-th filter of an i-th layer of the neural network (i is an integer greater than zero); a second at least one PIM subarray configured to perform a filtering operation of an (i+1)-th filter of an (i+1)-th layer of the neural network; and a controller configured to control the first at least one PIM and the second at least one PIM subarray, the controller being configured to perform: filtering the first filter of the i-th layer The i-th value is supplied to the first at least one PIM subarray to generate a first (i+1)th value for the (i+1)th layer, the first i-th value corresponding to a first input image; the first (i+1)th value of the (i+1)th layer is supplied to the second at least one PIM subarray to generate an output value associated with the first input image; and in parallel with supplying the first (i+1)th value corresponding to the first input image, the second i-th value of the i-th layer is supplied to the first at least one PIM subarray to generate a second (i+1)th value, the second i-th value corresponding to a second input image.

[0026] In some embodiments, the PIM subarrays of the first at least one PIM subarray and the second at least one PIM subarray include: a plurality of bit cells for storing a plurality of weights corresponding to a respective one of the i-th filter or the (i+1)-th filter. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings, together with the specification, illustrate example embodiments of the present disclosure, and, together with the description, serve to explain the principles of the present disclosure.

[0028] Figure 1 is a schematic diagram illustrating a configurable PIM system according to some embodiments of the present disclosure.

[0029] Figure 2A is a schematic diagram illustrating blocks of a configurable PIM system according to some embodiments of the present disclosure.

[0030] Figure 2B A PIM sub-array of a tile is shown according to some embodiments of the present disclosure.

[0031] Figure 3A-3C Inter-layer pipelining of inference in a neural network according to some embodiments of the present disclosure is shown.

[0032] Figure 4A-4C A combination of inter-layer pipelining and intra-layer pipelining for inference in a neural network according to some embodiments of the present disclosure is shown.

[0033] Figure 5 Batch pipelining of inference in a neural network according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0034] The specific embodiments set forth below are intended to be descriptions of example embodiments of systems and methods for pipelined machine learning acceleration provided in accordance with the present disclosure, and are not intended to represent the only form in which the present disclosure can be constructed or utilized. This description sets forth the features of the present disclosure in conjunction with the embodiments shown. However, it should be understood that the same or equivalent functions and structures can be implemented by different embodiments, which are also intended to be included within the scope of the present disclosure. As indicated elsewhere herein, similar element numbers are intended to indicate similar elements or features.

[0035] Various aspects of the present disclosure are directed to mapping machine learning applications to a PIM-based accelerator in a pipelined manner. Pipelining may include inter-layer pipelining, intra-layer pipelining, or a combination of these two types. The PIM-based system is configurable, and (multiple) pipeline schemes are use-specific and can be mapped to the PIM-based system one by one to improve (e.g., maximize) the power performance of each application. According to some examples, the (multiple) pipeline schemes of the configurable PIM-based system can provide significant (e.g., orders of magnitude) power performance improvements compared to other digital or PIM-based inference accelerators of the related art. In addition, the (multiple) pipeline schemes can provide stall-free or low-latency operation of network topologies in hardware.

[0036] Figure 1 is a schematic diagram illustrating a configurable PIM system 1 according to some embodiments of the present disclosure.

[0037] refer to Figure 1, once trained, the configurable PIM system 1 performs inference on input data to generate output data, which may be a prediction based on the input data. According to some embodiments, the configurable PIM system 1 includes: a PIM array 10 including a plurality of tiles 100 for performing inference operations, a controller 20 for controlling the operation of the PIM array 10, and a memory 30 (e.g., on-logic-die memory) for storing outputs or intermediate results of each of the tiles 100 of the configurable PIM system 1. In some examples, the memory 30 may be an embedded magnetoresistive random access memory (eMRAM), a static random access memory (SRAM), or the like.

[0038] Figure 2A is a schematic diagram of a block 100 illustrating a configurable PIM system 1 according to some embodiments of the present disclosure. Figure 2B A PIM sub-array 110 is shown according to some embodiments of the present disclosure.

[0039] Reference Figure 2A According to some embodiments, the tile 100 of the PIM array 10 includes a plurality of PIM sub-arrays (XA) 110 that may be organized in a matrix form; an input register (IR) 120, the input register 120 being configured to receive and store an input signal Vin (e.g., an input voltage signal, also referred to as input activation), which may correspond to input data, and the input register 120 may provide the stored input signal to an appropriate one (or more) of the PIM sub-arrays 110; an analog-to-digital converter (ADC) 130, being configured to convert an analog output of the PIM sub-array 110 into a digital signal (e.g., a binary signal); a shift-and-add circuit 140 for storing and adding the output of the ADC 130 within one clock cycle; and an output register 150 for storing the output of the shift-and-add circuit 140 (referred to as output activation) before transmitting the output to the memory 30 for storage. In some examples, controller 20 may control operations of input register 120 and output register (OR) 150 , and may determine to which PIM sub-array 110 data stored in input register 120 is transferred.

[0040] refer to Figure 2BIn some embodiments, each PIM subarray 110 includes a plurality of bit cells 112, which may be analog or digital (single bit or multiple bits). In some examples, the bit cell 112 may be a two or three terminal non-volatile synaptic weight bit cell. The bit cell 112 may be a resistive random access memory (RRAM) that may be used as an analog or digital memory. However, the embodiments of the present disclosure are not limited thereto, and in some examples, the bit cell 112 may be a conductive bridging random access memory (CBRAM), a phase change memory (PCM), a ferroelectric field effect transistor (FerroFET), a spin transfer torque (STT) memory, and the like. The bit cell may also include multiple cells of a memory cell and have more than three terminals. In order to select a single bit cell 112, one or more diodes or field effect transistors (FETs) may be attached in series to the bit cell 112. The PIM subarray 110 also includes peripheral circuits such as a digital-to-analog converter (DAC) 114 for converting digital inputs into analog voltage signals to be applied to one or more of the bit cells 112, and a sample-and-hold circuit 116 for storing the output of the bit cell 112 before passing it to a subsequent block.

[0041] In some embodiments, the PIM subarray 110 acts as a filter (eg, a convolutional filter), and each bit cell 112 stores a learnable weight (w1, w2, w3, etc.) of the filter.

[0042] Reference Figure 2A-2B , a machine learning system may be described by a network such as a convolutional neural network (CNN), a fully connected (FC) neural network, and / or a recurrent neural network (RNN) having multiple interconnected layers. These interconnected layers may be described by a weight matrix mapped and stored in the bit cells 112.

[0043] In some embodiments, a set of PIM sub-arrays 110 is assigned to each layer of the neural network based on the network topology.

[0044] The number of rows of the PIM subarray 110 in the tile 100 can be determined by the filter size of the interconnect layer (e.g., a layer of a CNN). The number of columns of the PIM subarray 110 in the tile 100 can be determined by the number of filters of the layer, the number of bits mapped by each bit cell 112, and the precision of the filter weights. In some embodiments, each row of the PIM subarray 110 represents a filter of the CNN. The input activations from the previous layer can be fed in parallel (e.g., simultaneously) as row inputs to the rows of the PIM subarray 110 in one clock cycle. In some examples, the rows and columns of the tile 100 can be interchanged as long as the logic within the controller 20 is changed accordingly.

[0045] In some embodiments, the PIM subarray 110 may generate an output current I that is conductance-weighted (and may represent a synaptic neuron). OUT , and the output currents of the PIM sub-arrays 110 along the row are summed together (eg, via a common electrical connection coupled to the PIM sub-arrays 110). Thus, the row of the PIM sub-arrays 110 may generate a summed output current I OUT1 to I OUTn . ADC 130 converts the summed output current (which may be an analog signal) into a digital signal (e.g., binary data). Each summed output may represent a calculated partial sum, which is then shifted and added to the partial sum of the next input activation bit by shift-and-add circuit 140. Output register 150 then stores the calculated partial sum. Once controller 20 determines that a final output has been generated, the final output (return) is stored in memory 30, which may then be used as an input to the next layer. How and when this output data is used for processing in the next layer may determine the pipelining scheme being implemented.

[0046] According to some embodiments, the mapping of network layers onto PIM subarray 110 is performed in a pipelined manner, which may result in stall-free (or low-stall) operation. In some embodiments, controller 20 begins processing the next layer of the neural network once sufficient outputs have been generated from the current layer, rather than waiting for one layer of the neural network to complete processing before moving on to the next layer. This may be referred to as inter-layer pipelining (also referred to as inter-layer concurrency).

[0047] Figure 3A-3C Inter-layer pipelining of inference in a neural network according to some embodiments of the present disclosure is shown.

[0048] refer to Figure 3A-3C , the set of values ​​of the i-th layer (i is an integer greater than zero) can form a 2-dimensional matrix 300. In some examples, the 2-dimensional matrix 300 can represent an image, and each element 302 of the matrix 300 can represent a pixel value (e.g., color intensity). The 2-dimensional matrix 300 can also represent a feature map, where each element 302 represents the output value of the previous layer (i.e., the (i-1)th layer). For ease of explanation, Figure 3A-3C The 2-dimensional matrix 300 is a 6×6 matrix; however, as one of ordinary skill in the art will appreciate, embodiments of the present disclosure are not limited thereto, and the matrix 300 may have any suitable size represented as n×m, where n and m are integers greater than 1.

[0049] The i-th filter 304 (also referred to as a kernel) may be a sliding convolution filter that operates on a set of values ​​of the i-th layer to generate values ​​of the next layer (i.e., the (i+1)-th layer). The filter 304 may be represented by a 2-dimensional matrix of size pxq, where p and q are integers greater than zero and less than or equal to n and m, respectively. For ease of illustration, the i-th filter 304 is shown as a 2×2 filter. Each element of the convolution filter may be a learnable weight value.

[0050] The i-th filter 304 shifts / slides / moves the stride length across the set of values ​​of the i-th layer until the entire set of values ​​of the i-th layer is traversed. At each shift, the i-th filter 304 performs a matrix multiplication operation between the filter 304 and the portion of the matrix of values ​​of the i-th layer (on which the filter 304 is operating at that point). Figure 3A-3C The example of shows a stride length of 1 (i.e., a single stride), however, the stride length can be 2, 3, or any suitable value. The output of the convolution operation that forms a set of values ​​for the next layer (i.e., the (i+1)th layer) can be referred to as the convolution feature outputs. These outputs can populate the matrix 310 associated with the (i+1)th layer.

[0051] In related techniques, it may be necessary to fully process one layer of a neural network before processing the values ​​of subsequent layers, which can be slow.

[0052] However, according to some embodiments, the controller 20 monitors the output values ​​generated by the i-th filter 304, and once it is determined that a sufficient number of values ​​are available for processing at the next layer (i.e., the (i+1)th layer), the controller 20 processes the available output values ​​of the (i+1)th layer via the (i+1)th filter 314. Therefore, the controller 20 can apply the (i+1)th filter 314 to the available values ​​of the (i+1)th layer in parallel (e.g., simultaneously) with applying the i-th filter 304 to the values ​​of the i-th layer. This inter-layer pipelining can be performed for any number of layers or all layers of the neural network. In other words, convolution operations can be performed in parallel (e.g., performed simultaneously in two or more layers of the neural network). This can result in a significant improvement in the inference speed of the neural network.

[0053] According to some embodiments, the first at least one PIM subarray 110 is configured to perform a filtering operation of an i-th filter 304, and the second at least one PIM subarray 110 is configured to perform a filtering operation of an (i+1)-th filter 314. The controller 20 may supply a first set of i-th values ​​of an i-th layer to the first at least one PIM subarray 110 to generate an (i+1)-th value for an (i+1)-th layer, and when it is determined that the number of (i+1)-th values ​​is sufficient to be processed by the second at least one PIM subarray 110, the (i+1)-th value may be supplied to the second at least one PIM subarray 110 to generate output values ​​for a subsequent layer of the neural network while supplying a second set of i-th values ​​of the i-th layer to the first at least one PIM subarray 110 for processing. In some examples, the number of available values ​​for the filtering operation is determined to be sufficient when the corresponding layer has data for each unit operated by the corresponding filter.

[0054] In some embodiments, the time (e.g., clock cycles) at which the next layer (i.e., the (i+1)th layer) can begin processing depends on the size of the i-th filter associated with the i-th layer, the size of the (i+1)th filter associated with the (i+1)th layer, the stride length of the (i+1)th filter, and the size of the image or feature map corresponding to the values ​​of the i-th layer. Figure 3A-3C In the example of (in this example, each of the i-th filter 304 and the (i+1)-th filter 314 has a size of 2×2, the stride length of the i-th filter 304 is 1, and the stride length of the (i+1)-th filter 314 is 2), the controller 20 starts processing the (i+1)-th layer 7 cycles after starting processing of the i-th layer. Figure 3C As shown, there are not enough values ​​in the feature map 310 for the (i+1)th filter 314 to perform further operations. According to some embodiments, each convolution operation may be performed in one clock cycle.

[0055] According to some embodiments of the present disclosure, the processing speed gain from inter-layer pipelining may be further improved by additionally employing intra-layer pipelining (also referred to as intra-layer concurrency) to generate more than one output value from a layer per cycle.

[0056] Figure 4A-4C A combination of inter-layer pipelining and intra-layer pipelining for inference in a neural network according to some embodiments of the present disclosure is shown.

[0057] refer to Figure 4A-4C According to some embodiments, at each layer of a neural network, more than one filter operates in parallel (eg, simultaneously) to generate more than one output value at a time.

[0058] In some embodiments, at any given time, the controller 20 applies a first i-th filter (e.g., a first sliding i-th filter) 304 associated with the i-th layer to a first portion / block of values ​​of the i-th layer to generate a first output value for the (i+1)-th layer (e.g., a first portion of the (i+1)-th value), and in parallel (e.g., simultaneously) applies a second i-th filter (e.g., a second sliding i-th filter) 305 associated with the i-th layer to a second portion / block of values ​​of the i-th layer to generate a second output value for the (i+1)-th layer (e.g., a second portion of the (i+1)-th value).

[0059] According to some embodiments, the first i-th filter and the second i-th filter are identical (e.g., contain the same weight map / weight values), but are offset in position by the stride length of the first filter 304. Thus, in effect, the second i-th filter 305 performs the same operation as the first i-th filter 304 will perform in the next clock cycle, but performs the operation in the same clock cycle as the first i-th filter 304. As a result, in one clock cycle, the PIM array 10 can generate two (or more) values ​​for the next layer. Here, each layer can be subdivided and mapped to different tiles with one or more copies of the same weight matrix. For example, the first i-th filter 304 can be implemented with one tile 100, while the second i-th filter 305 can be implemented with a different tile 100, whereby the weight matrices used for the two tiles 100 are the same.

[0060] In some embodiments, the number of concurrent operations performed by the filters at each layer may be equal to the number of filters forming a combined / composite filter (e.g., 303) at that layer. Here, the stride length of the composite filter may be equal to the number of filters that make up the composite filter (and perform concurrent operations) multiplied by the stride length of the constituent filters. Figure 4A-4C In the example of , the stride length of the composite filter 303 including two i-th filters 304 and 305 each having a stride length of 1 is equal to 2. As will be appreciated by those skilled in the art, embodiments of the present invention are not limited to two concurrent operations per layer and may be extended to include any suitable number of concurrent operations.

[0061] like Figure 4A-4C As shown, by utilizing both inter-layer pipelining and intra-layer pipelining (the combination of which may be referred to as “super-pipelining”), the controller 20 may start processing the (i+1)th filter after only four clock cycles, and may continue processing the (i+1)th layer in the next clock cycle (t=5) without any delay. Figure 3A-3CThis marks an improvement when compared to the inter-layer pipelining scheme of the example in which the controller 20 can only start processing the (i+1)th filter after 7 clock cycles, and cannot perform the next filtering operation of the (i+1)th filter at the next clock cycle (t=8) due to insufficient number of values ​​available at the (i+1) layer.

[0062] Although intra-layer pipelining / concurrency has been described above with respect to the i-th layer, according to some embodiments, the PIM array 10 may utilize intra-layer pipelining / concurrency in more than one layer (eg, in all layers) of a neural network.

[0063] In addition to using inter-layer pipelining and intra-layer pipelining to improve (e.g., increase) the processing speed of a single image or feature map, embodiments of the present disclosure also utilize batch pipelining to improve the processing speed of consecutive images / feature maps.

[0064] Figure 5 Batch pipelining of inference in a neural network according to some embodiments of the present disclosure is shown.

[0065] In the related art, each input image / feature map may be processed one by one. As a result, when the first filter associated with the first layer completes processing the first layer, it may remain idle and not process any other information until all other layers of the neural network have completed processing.

[0066] According to some embodiments, the PIM array 10 utilizes batch processing to concurrently (e.g., in parallel / simultaneously) process more than one input image / feature map. In doing so, when a filter associated with a layer completes processing of an image / feature map of that layer, the filter can continue to process the same layer of a subsequent image / feature map.

[0067] According to some embodiments, a neural network for batch processing of multiple images includes a plurality of layers mapped onto different tiles 100 of a PIM array 10. The plurality of layers may include an i-th layer (i is an integer greater than zero) and an (i+1)-th layer. The configurable PIM system 1 processes a first i-th value 402 of the i-th layer for a first input image to generate a first (i+1)-th value 412, which is used as an input to the (i+1)-th layer. Then, the configurable PIM system 1 processes the first (i+1)-th value 412 of the (i+1)-th layer for the first input image to generate an output value for a subsequent layer. According to some embodiments, while (e.g., in parallel with) processing the (i+1)-th value 412 for the first input image, the configurable PIM system 1 processes a second i-th value 422 of the i-th layer for a second input image to generate a second (i+1)-th value 432. In some embodiments, processing the second i-th value 422 for the second input image may be performed in parallel with processing the first i-th value 402 for the first input image. Processing the first i-th value 402 for the first input image may include applying the i-th filter 404 associated with the i-th layer to the first i-th value 402 of the i-th layer to generate the (i+1)th value 412 for the (i+1)th layer. In addition, processing the second i-th value 422 for the second input image may include applying the i-th filter 404 associated with the i-th layer to the second i-th value 422 to generate the second (i+1)th value 432 for the (i+1)th layer. In other words, the same i-th filter 404 may be used to process the first i-th value and the second i-th value in a time-interleaved manner. However, embodiments of the present disclosure are not limited thereto, and in some examples, a filter similar to the i-th filter 404, but having the same size and stride length as the i-th filter 404, may be used to process the second i-th value.

[0068] According to some embodiments, the first at least one PIM subarray 110 is configured to perform a filtering operation of the i-th filter 404, and the second at least one PIM subarray 110 is configured to perform a filtering operation of the (i+1)-th filter of the (i+1)-th layer of the neural network. The controller 20 may supply the first i-th value of the i-th layer to the first at least one PIM subarray 110 to generate a first (i+1)-th value for the (i+1)-th layer, wherein the first i-th value corresponds to the first input image. The controller 20 may also supply the first (i+1)-th value of the (i+1)-th layer to the second at least one PIM subarray 110 to generate an output value associated with the first input image. While supplying the (i+1)-th value corresponding to the first input image, the controller 20 may supply the second i-th value of the i-th layer to the first at least one PIM subarray 110 to generate a second (i+1)-th value, wherein the second i-th value corresponds to the second input image.

[0069] like Figure 5 As shown, the first i-th value 402 corresponding to the first input image can form a 2-dimensional matrix 400, and the second i-th value 422 corresponding to the second input image can form a 2-dimensional matrix 420. In some examples, the first i-th value 402 includes pixel values ​​of the first input image (or a rectangular block of the first input image), and the second i-th value 422 includes pixel values ​​of the second input image (or a rectangular block of the second input image). The first input image and the second input image can have the same size / dimensions. In some examples, the first i-th value 402 includes values ​​of a first feature map generated by a previous layer of the neural network, and the second i-th value 422 includes values ​​of a second feature map generated by a previous layer of the neural network. The first feature map and the second feature map correspond to (e.g., are generated from) the first input image and the second input image, respectively. The i-th filter 404 can be a sliding convolution filter in the form of a pxq matrix, where p and q are integers greater than zero (in Figure 5 , for convenience, the i-th filter 404 is shown as a 2x 2 matrix).

[0070] According to some embodiments, processing the second i-th value for the second input image is initiated at a time offset after starting processing the first i-th value for the first input image. The time offset may be greater than or equal to the number of clock cycles corresponding to a single step of the i-th filter. For example, Figure 5 In the example, the stride length of the filter 404 is 1, which corresponds to a single clock cycle, and the time offset between processing the same layer for the first input image and the second input image may be at least one clock cycle.

[0071] According to some examples, filters 404 operating on the first input image and the second input image may be copies of each other, but may be implemented in hardware via different PIM subarrays 110 .

[0072] According to some embodiments, increasing the number of images that are processed concurrently by the configurable PIM system 1 improves processing time (e.g., improves image recognition time). In some embodiments, the number of images that can be batch processed by the configurable PIM system 1 can be limited to the number of clock cycles it takes to fully process a single image. For example, when processing of a single image takes 100 clock cycles, the configurable PIM system 1 can batch process 100 or fewer images.

[0073] According to some embodiments, inter-layer pipelining and intra-layer pipelining (e.g., see Figure 3A-3C and Figure 4A-4C ) can be used in conjunction with batch pipelining of multiple images to achieve even greater performance gains.

[0074] According to some examples, the neural network mentioned in the present disclosure may be a convolutional neural network (ConvNet / CNN), which can obtain an input image / video, assign importance to various aspects / objects in the image / video (e.g., via learnable weights and biases), and be able to distinguish one from another. However, embodiments of the present disclosure are not limited thereto. For example, the neural network may be a recursive neural network (RNN), a multilayer perceptron (MLP), etc.

[0075] As described herein, the (multiple) pipelining schemes of the reconfigurable PIM system according to some embodiments of the present disclosure provide significant (e.g., orders of magnitude) power performance improvements relative to other digital or PIM-based inference accelerators of the related art. In addition, the (multiple) pipelining schemes can provide low-latency, stall-free operation of neural networks on hardware.

[0076] As will be appreciated by one of ordinary skill in the art, the operations performed by the controller 20 may be performed by a processor. A memory local to the processor may contain instructions that, when executed, cause the processor to perform the operations of the controller.

[0077] It will be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or parts, these elements, components, regions, layers and / or parts should not be limited by these terms. These terms are used to distinguish one element, component, region, layer or part from another element, component, region, layer or part. Therefore, without departing from the scope of the present invention, the first element, component, region, layer or part discussed below may be referred to as a second element, component, region, layer or part.

[0078] The terms used herein are for the purpose of describing a specific embodiment, and are not intended to limit the inventive concept. As used herein, the singular forms "one" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include", "include", "include" and / or "include" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or increase of one or more other features, integers, steps, operations, elements, components and / or their groups. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. When expressions such as "at least one" are after the element list, the entire element list is modified and the single element in the list is not modified. In addition, "may" is used when describing the embodiment of the inventive concept to refer to "one or more embodiments of the inventive concept". Similarly, the term "exemplary" is intended to represent an example or explanation.

[0079] As used herein, the terms "use," "in use," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively.

[0080] The configurable PIM system and / or any other related devices or components (such as controllers and processors) according to the embodiments of the present disclosure described herein can be implemented by utilizing any suitable hardware, firmware (e.g., application specific integrated circuits), software, or any suitable combination of software, firmware, and hardware. For example, the various components of the PIM system can be formed on one integrated circuit (IC) chip or on separate IC chips. In addition, the various components of the PIM system can be implemented on a flexible printed circuit film, a tape carrier package (TCP), a printed circuit board (PCB), or formed on the same substrate. In addition, the various components of the PIM system can be processes or threads that run on one or more processors in one or more computing devices, execute computer program instructions, and interact with other system components to perform the various functions described herein. The computer program instructions are stored in a memory, which can be implemented in a computing device using a standard memory device such as, for example, a random access memory (RAM). The computer program instructions can also be stored in other non-transitory computer-readable media (e.g., CD-ROMs, flash drives, etc.). Moreover, it should be appreciated by those skilled in the art that the functions of various computing devices can be combined or integrated into a single computing device, or the functions of a particular computing device can be distributed on one or more other computing devices without departing from the scope of the exemplary embodiments of the present disclosure.

[0081] Although the present disclosure has been described in detail with specific reference to the illustrative embodiments of the present disclosure, the embodiments described herein are not intended to be exhaustive or to limit the scope of the present disclosure to the precise forms disclosed. Those skilled in the art to which the present disclosure belongs will recognize that changes and modifications can be made to the described structures and methods of assembly and operation without meaningfully departing from the principles and scope of the present disclosure as set forth in the appended claims and their equivalents.

Claims

1. A method for pipelined reasoning of a neural network, the neural network comprising a plurality of layers, the plurality of layers comprising an i-th layer and an (i+1)-th layer, wherein i is an integer greater than zero, the method comprising: for a first input image, processing, by a controller, a first i-th value of the i-th layer using a composite filter including a first i-th filter associated with the i-th layer and a second i-th filter associated with the i-th layer to generate a first (i+1)-th value for the (i+1)-th layer, wherein the first i-th filter and the second i-th filter are positionally offset by a stride of the first i-th filter so that the stride of the composite filter is greater than or equal to the sum of a stride of the first i-th filter and a stride of the second i-th filter, and the stride of the first i-th filter is associated with a movement of the first i-th filter on the i-th layer; For the first input image, processing the first (i+1)th value of the (i+1)th layer to generate an output value; and In parallel with processing the (i+1)th value for the first input image, For a second input image, a second i-th value of the i-th layer is processed to generate a second (i+1)-th value.

2. The method according to claim 1, wherein: Processing the second i-th value for the second input image is performed in parallel with processing the first i-th value for the first input image.

3. The method according to claim 1, wherein: The first i-th value comprises a pixel value of the first input image, and The second i-th value includes a pixel value of the second input image.

4. The method according to claim 1, wherein: The first i-th value comprises a value of a first feature map generated by a previous layer of the neural network, the first feature map corresponding to the first input image, and wherein the second i-th value comprises a value of a second feature map generated by the previous layer of the neural network, the second feature map corresponding to the second input image.

5. The method according to claim 1, wherein: Processing the first i-th value of the i-th layer for the first input image includes: The composite filter is applied to the first i-th value of the i-th layer to generate the first (i+1)th value for the (i+1)th layer.

6. The method according to claim 5, wherein: Processing the second i-th value of the i-th layer for the second input image comprises: The composite filter is applied to the second i-th value of the i-th layer to generate the second (i+1)th value for the (i+1)th layer.

7. The method according to claim 5, wherein: The first i-th filter is a sliding convolution filter in the form of a p×q matrix, where p and q are integers greater than zero.

8. The method according to claim 5, wherein: Applying the composite filter comprises: A matrix multiplication operation is performed between the composite filter and values ​​of the first i-th value that overlap with the composite filter.

9. The method according to claim 1, wherein: processing the second i-th value of the i-th layer for the second input image is initiated at a time offset after initiating processing of the first i-th value of the i-th layer for the first input image, and The time offset is greater than or equal to the number of clock cycles corresponding to a single step of the complex filter.

10. A system for pipelined reasoning of a neural network, the neural network comprising a plurality of layers, the plurality of layers comprising an i-th layer and an (i+1)-th layer, wherein i is an integer greater than zero, the system comprising: processor; and A processor memory local to a processor, wherein the processor memory has stored thereon instructions that, when executed by the processor, cause the processor to perform: for a first input image, processing, by a controller, a first i-th value of the i-th layer using a composite filter including a first i-th filter associated with the i-th layer and a second i-th filter associated with the i-th layer to generate a first (i+1)-th value for the (i+1)-th layer, wherein the first i-th filter and the second i-th filter are positionally offset by a stride of the first i-th filter so that the stride of the composite filter is greater than or equal to the sum of a stride of the first i-th filter and a stride of the second i-th filter, and the stride of the first i-th filter is associated with a movement of the first i-th filter on the i-th layer; For the first input image, processing the first (i+1)th value of the (i+1)th layer to generate an output value; and In parallel with processing the (i+1)th value for the first input image, For a second input image, a second i-th value of the i-th layer is processed to generate a second (i+1)-th value.

11. The system according to claim 10, wherein: Processing the second i-th value for the second input image is performed in parallel with processing the first i-th value for the first input image.

12. The system according to claim 10, wherein: The first i-th value comprises a pixel value of the first input image, and The second i-th value includes a pixel value of the second input image.

13. The system according to claim 10, wherein: The first i-th value comprises a value of a first feature map generated by a previous layer of the neural network, the first feature map corresponding to the first input image, and wherein the second i-th value comprises a value of a second feature map generated by the previous layer of the neural network, the second feature map corresponding to the second input image.

14. The system according to claim 10, wherein: Processing the first i-th value of the i-th layer for the first input image includes: The composite filter is applied to the first i-th value of the i-th layer to generate the first (i+1)th value for the (i+1)th layer.

15. The system of claim 14, wherein: Processing the second i-th value of the i-th layer for the second input image comprises: The composite filter is applied to the second i-th value of the i-th layer to generate the second (i+1)th value for the (i+1)th layer.

16. The system of claim 14, wherein: The first i-th filter is a sliding convolution filter in the form of a p×q matrix, where p and q are integers greater than zero.

17. The system of claim 14, wherein: Applying the composite filter comprises: A matrix multiplication operation is performed between the composite filter and values ​​of the first i-th value that overlap with the composite filter.

18. The system of claim 10, wherein: processing the second i-th value of the i-th layer for the second input image is initiated at a time offset after initiating processing of the first i-th value of the i-th layer for the first input image, and The time offset is greater than or equal to the number of clock cycles corresponding to a single step of the complex filter.

19. A configurable processing-in-memory (PIM) system configured to implement a neural network, the system comprising: at least one first PIM subarray configured to perform a filtering operation of a composite filter including a first i-th filter and a second i-th filter of an i-th layer of the neural network, wherein i is an integer greater than zero, wherein the first i-th filter and the second i-th filter are offset in position by a stride of the first i-th filter, such that the stride of the composite filter is greater than or equal to the sum of the stride of the first i-th filter and the stride of the second i-th filter, and the stride of the first i-th filter is associated with the movement of the first i-th filter on the i-th layer; second at least one PIM subarray, configured to perform a filtering operation of an (i+1)th filter of an (i+1)th layer of the neural network; and a controller configured to control the first at least one PIM sub-array and the second at least one PIM sub-array, the controller configured to perform: supplying a first i-th value of the i-th layer to the first at least one PIM subarray to generate an (i+1)th value for the (i+1)th layer, the first i-th value corresponding to a first input image; supplying the first (i+1)th value of the (i+1)th layer to the second at least one PIM subarray to generate an output value associated with the first input image; and In parallel with supplying the first (i+1)th value corresponding to the first input image, A second i-th value of the i-th layer is supplied to the first at least one PIM subarray to generate a second (i+1)-th value, the second i-th value corresponding to a second input image.

20. The system of claim 19, wherein: The PIM subarrays in the first at least one PIM subarray and the second at least one PIM subarray include: A plurality of bit units are used to store a plurality of weights corresponding to a respective one of the first i-th filter or the second i-th filter or the (i+1)-th filter.

Citation Information

Patent Citations

  • Multi-layer neural network

    US20180053084A1