Computation-in-Memory (CIM) Architecture and Dataflow Supporting Deeply Convolutional Neural Networks (CNN)
Patent Information
- Application Number
- JP2023577150
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-29
- Filing Date
- 2022-06-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-28
AI Technical Summary
Traditional compute-in-memory (CIM) processes for machine learning model data processing, such as depth-separable convolutional neural networks, require additional hardware elements like digital multiply-and-accumulate circuits (DMACs) that increase space, power consumption, and complexity, limiting their efficiency in edge devices.
A CIM architecture with CIM cells configured for different kernels on different rows and columns of a CIM array, allowing parallel operations and analog-to-digital conversion, coupled with nonlinear activation circuits for further processing, to perform depthwise and pointwise convolutions efficiently.
This approach reduces power consumption and processing time by performing calculations directly in memory, enhancing the efficiency of machine learning tasks on edge devices like mobile devices and IoT devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Application No. 17 / 361,784, filed June 29, 2021, which is assigned to the assignee of this application and is incorporated by reference in its entirety into this specification. [Background technology]
[0002] Aspects of the present disclosure relate to performing machine learning tasks, and in particular to in-memory computational architectures and data flows.
[0003]
[0003] Machine learning is generally the process of creating a trained model (e.g., an artificial neural network, tree, or other structure) that represents a generalized fit to a set of training data that is known a priori. By applying the trained model to new data, inferences are generated, which can be used to gain insight into the new data. Sometimes, applying a model to new data is described as "performing inference" on the new data.
[0004]
[0004] As the use of machine learning has proliferated to enable various machine learning (or artificial intelligence) tasks, a need has arisen for more efficient processing of machine learning model data. In some cases, specialized hardware such as machine learning accelerators can be used to enhance the processing system's ability to process machine learning model data. However, such hardware requires space and power, which is not always available on the processing device. For example, "edge processing" devices such as mobile devices, always-on devices, and internet of things (IoT) devices must balance processing power with power and packaging constraints. Furthermore, accelerators may need to move data across a common data bus, which can cause significant power usage and introduce latency to other processes sharing the data bus. Therefore, other aspects of the processing system are considered to process machine learning model data.
[0005]
[0005] Memory devices are an example of another aspect of a processing system that can be leveraged to perform processing of machine learning model data through a so-called computation in memory (CIM) process. Unfortunately, conventional CIM processes may not be able to perform processing of complex model architectures such as depthwise separable convolutional neural networks without additional hardware elements such as digital multiply-and-accumulate circuits (DMACs) and related peripherals. These additional hardware elements use additional space, power, and complexity in their implementation, which tends to reduce the benefits of leveraging memory devices as additional computational resources. Even if auxiliary aspects of a processing system have DMACs available to perform processing that cannot be performed directly in memory, moving data to and from those auxiliary aspects requires time and power, thus reducing the benefits of the CIM process.
[0006]
[0006] Therefore, there is a need for systems and methods for performing in-memory computations of a wider variety of machine learning model architectures, such as deeply separable convolutional neural networks. Summary of the Invention
[0007]
[0007] Some aspects provide an apparatus for signal processing in a neural network. The apparatus generally includes a first set of computation in memory (CIM) cells configured as a first kernel for neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array, and a second set of CIM cells configured as a second kernel for neural network computation, the second set of CIM cells including one or more second columns and a second plurality of rows of the CIM array. In some aspects, the one or more first columns are different from the one or more second columns, and the first plurality of rows are different from the second plurality of rows.
[0008]
[0008] Some aspects provide a method of signal processing in a neural network. The method generally includes loading a first plurality of weight parameters of a first kernel to a first set of CIM cells via one or more first columns, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array, to perform a neural network calculation. The method may also include loading a second plurality of weight parameters of a second kernel to a second set of CIM cells via one or more second columns, the second set of CIM cells including one or more second columns and a second plurality of rows of the CIM array, to perform a neural network calculation. The one or more first columns may be different from the one or more second columns, and the first plurality of rows may be different from the second plurality of rows. The method may also include performing the neural network calculation by applying a first activation input to the first plurality of rows and applying a second activation input to the second plurality of rows.
[0009]
[0009] Some aspects provide a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method of signal processing in a neural network. The method generally includes loading a first plurality of weight parameters of a first kernel to a first set of CIM cells via one or more first columns, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array, for performing a neural network calculation. The method may also include loading a second plurality of weight parameters of a second kernel to a second set of CIM cells via one or more second columns, the second set of CIM cells including one or more second columns and a second plurality of rows of the CIM array, for performing a neural network calculation. The one or more first columns may be different from the one or more second columns, and the first plurality of rows may be different from the second plurality of rows. The method may also include performing a neural network computation by applying a first activation input to the first plurality of rows and applying a second activation input to the second plurality of rows.
[0010]
[0010] Another aspect provides a processing system configured to perform the aforementioned methods and methods further described herein; a non-transitory computer readable medium comprising instructions which, when executed by one or more processors of the processing system, cause the processing system to perform the aforementioned methods and methods further described herein; a computer program product embodied on a computer readable storage medium comprising code for performing the aforementioned methods and methods further described herein; and a processing system comprising means for performing the aforementioned methods and methods further described herein.
[0011] The following description and the related drawings set forth in detail certain illustrative features of the one or more aspects. [Brief description of the drawings]
[0012]
[0012] The accompanying drawings illustrate some aspects of the one or more aspects and therefore should not be considered as limiting the scope of the present disclosure. [Figure 1A]
[0013] 1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1B] 1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1C] 1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1D] 1A-1C are diagrams illustrating examples of various types of neural networks. [Diagram 2]
[0014] FIG. 1 is a diagram illustrating an example of a conventional convolution operation. [Figure 3A]
[0015] FIG. 13 is a diagram illustrating an example of a depthwise separable convolution operation. [Figure 3B] FIG. 13 is a diagram illustrating an example of a depthwise separable convolution operation. [Figure 4]
[0016] FIG. 1 illustrates an exemplary compute-in-memory (CIM) array configured to perform machine learning model computations. [Figure 5A]
[0017] FIG. 5 illustrates additional details of an exemplary bitcell, which may represent the bitccells of FIG. 4. [Figure 5B] FIG. 5 illustrates additional details of an exemplary bitcell, which may represent the bitccells of FIG. 4. [Figure 6]
[0018] FIG. 2 is an example timing diagram of various signals during a compute-in-memory (CIM) array operation. [Figure 7]
[0019] FIG. 1 illustrates an exemplary convolutional layer architecture implemented by a compute-in-memory (CIM) array. [Figure 8]
[0020] FIG. 2 illustrates a CIM architecture including multiple CIM arrays in accordance with some aspects of the present disclosure. [Figure 9]
[0021] FIG. 1 illustrates example operations for signaling processing via a CIM architecture in accordance with certain aspects of the disclosure. [Figure 10]
[0022] FIG. 2 illustrates a CIM array divided into sub-banks to conserve power, in accordance with some aspects of the disclosure. [Figure 11]
[0023] FIG. 1 illustrates a CIM array having diagonally stacked kernels in accordance with some aspects of the present disclosure. [Figure 12]
[0024] FIG. 1 illustrates a CIM array implemented with an iterated kernel, in accordance with some aspects of the present disclosure. [Figure 13]
[0025] FIG. 1 is a flow diagram illustrating example operations for signal processing in a neural network in accordance with some aspects of the present disclosure. [Figure 14]
[0026] FIG. 1 illustrates an example electronic device configured to perform operations for signal processing in a neural network in accordance with some aspects of the present disclosure.
[0013]
[0027] For ease of understanding, wherever possible, like reference numbers have been used to designate like elements common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014]
[0028] Aspects of the present disclosure provide apparatus, methods, processing systems, and computer-readable media for performing computation in memory (CIM) of machine learning models including depthwise separable convolutional neural networks. Some aspects are directed to CIM cells of a CIM array configured for different kernels, where the CIM cells are implemented on different rows and columns of the CIM array to facilitate parallel operation of a first and a second kernel. For example, a first kernel may be implemented on a first row and column of the CIM array, and a second kernel may be implemented on a second row and column of the CIM array, where the first row and column are different from the second row and column. Each of the kernels implemented on different rows and columns can be coupled to an analog-to-digital converter (ADC), enabling parallel depth-wise (DW) computation and analog-to-digital conversion via the kernels. The results of the DW computation may be input to a nonlinear activation circuit for further processing and input to another CIM array for point-wise computation, as described in more detail herein.
[0015]
[0029] CIM-based machine learning (ML) / artificial intelligence (AI) task accelerators can be used for a wide variety of tasks, including image and audio processing. Furthermore, CIMs can be based on various types of memory architectures, such as DRAM, SRAM (e.g., based on SRAM cells as in FIG. 5), MRAM, and ReRAM, and can be attached to various types of processing units, including central processor units (CPUs), digital signal processors (DSPs), graphical processor units (GPUs), field-programmable gate arrays (FPGAs), AI accelerators, and the like. In general, CIMs can advantageously reduce the "memory wall" problem, where moving data in and out of memory consumes more power than computing the data. Thus, significant power savings can be realized by performing computations in memory. This is particularly useful for various types of electronic devices, such as low-power edge processing devices, mobile devices, and the like.
[0016]
[0030] For example, a mobile device may include a memory device configured to store data and in-memory computation operations. The mobile device may be configured to perform ML / AI operations based on data generated by the mobile device, such as image data generated by a camera sensor of the mobile device. Thus, a memory controller unit (MCU) of the mobile device may load weights from another on-board memory (e.g., flash or RAM) into a CIM array of the memory device and allocate input feature buffers and output (e.g., activation) buffers. The processing device may then begin processing the image data, for example, by loading layers in the input buffers and processing the layers with the weights loaded in the CIM array. This process may be repeated for each layer of the image data, and the outputs (e.g., activations) may be stored in an output buffer and then used by the mobile device for ML / AI tasks such as face recognition.
[0017] A brief background on neural networks, deep neural networks, and deep learning
[0031] Neural networks are organized into layers of interconnected nodes. In general, a node (or neuron) is where computations are performed. For example, a node may combine input data with a set of weights (or coefficients) that either amplify or attenuate the input data. Amplification or attenuation of an input signal may thus be viewed as an assignment of relative importance to various inputs with respect to the task the network is trying to learn. In general, input-weight products are added (or accumulated) and then this sum is passed through the node's activation function to determine whether and how far the signal should proceed further through the network.
[0018]
[0032] In its most basic implementation, a neural network may have an input layer, a hidden layer, and an output layer. "Deep" neural networks generally have two or more hidden layers.
[0019]
[0033] Deep learning is a method of training deep neural networks. In general, deep learning is sometimes called a "universal approximator" because it maps inputs to the network to outputs from the network and can therefore learn to approximate an unknown function f(x)=y between any input x and any output y. In other words, deep learning finds the correct f to transform x to y.
[0020]
[0034] More specifically, deep learning trains each layer of nodes on a different set of features, i.e., the output from the previous layer. Thus, with each successive layer of a deep neural network, the features become more complex. Deep learning is therefore powerful because it can progressively extract higher level features from input data by learning to represent the input at successively higher levels of abstraction at each layer, thereby building useful feature representations of the input data, to perform complex tasks such as object recognition.
[0021]
[0035] For example, when presented with visual data, the first layer of a deep neural network can be trained to recognize relatively simple features, such as edges, in the input data. In another example, when presented with auditory data, the first layer of a deep neural network can be trained to recognize the spectral power at a particular frequency in the input data. The second layer of the deep neural network can then be trained to recognize combinations of features, such as simple shapes in the visual data, or combinations of sounds in the auditory data, based on the output of the first layer. The higher layers can then be trained to recognize complex shapes in the visual data or words in the auditory data. The higher layers can then be trained to recognize common visual objects or spoken phrases. Thus, deep learning architectures can perform particularly well when applied to problems with natural hierarchical structures.
[0022] Layer Connectivity in Neural Networks
[0036] Neural networks, such as deep neural networks, can be designed with a variety of connectivity patterns between layers.
[0023]
[0037] 1A illustrates an example of a fully-connected neural network 102, in which a node in a first layer communicates its output to every node in a second layer, such that each node in the second layer receives input from every node in the first layer.
[0024]
[0038] 1B shows an example of a locally connected neural network 104. In the locally connected neural network 104, a node in a first layer may be connected to a limited number of nodes in a second layer. More generally, the locally connected layers of the locally connected neural network 104 may be configured such that each node in a layer has the same or similar connectivity pattern, but with connection strengths (or weights) that can have different values (e.g., 110, 112, 114, and 116). The connectivity patterns of the local connections may result in spatially distinct receptive fields in the upper layers, since the higher layer nodes in a given region may receive inputs that are tuned to the characteristics of a constrained portion of the total inputs to the network through training.
[0025]
[0039] One type of locally connected neural network is a convolutional neural network. Figure 1C shows an example of a convolutional neural network 106. The convolutional neural network 106 can be configured such that the connection strengths associated with the inputs to each node in the second layer are shared (e.g., 108). Convolutional neural networks are well suited to problems where the spatial location of the inputs is meaningful.
[0026]
[0040] One type of convolutional neural network is the deep convolutional network (DCN), which is a network of multiple convolutional layers and can be further configured with, for example, pooling layers and normalization layers.
[0027]
[0041] 1D shows one embodiment of a DCN 100 designed to recognize visual features in an image 126 generated by an image capture device 130. For example, if the image capture device 130 were a vehicle-mounted camera, the DCN 100 could be trained using various supervised learning techniques to identify traffic signs, and even numbers on traffic signs. Similarly, the DCN 100 could be trained for other tasks, such as identifying lane markings, or identifying traffic signals. These are just a few example tasks, and many other tasks are possible.
[0028]
[0042] In this example, the DCN 100 includes a feature extraction section and a classification section. Upon receiving an image 126, the convolutional layer 132 applies a convolution kernel to the image 126 (e.g., as shown and described in FIG. 2) to generate a first set of feature maps (or intermediate activations) 118. In general, a "kernel" or "filter" includes a multi-dimensional array of weights designed to emphasize different aspects of the input data channels. In various examples, "kernel" and "filter" can be used interchangeably to refer to a set of weights applied in a convolutional neural network.
[0029]
[0043] The first set of feature maps 118 may then be subsampled by a pooling layer (e.g., a max pooling layer, not shown) to generate a second set of feature maps 120. The pooling layer can reduce the size of the first set of feature maps 118 while retaining much of the information to improve model performance. For example, the second set of feature maps 120 can be downsampled from 28×28 to 14×14 by the pooling layer.
[0030]
[0044] This process can be repeated through many layers, in other words, the second set of feature maps 120 may be further convolved through one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0031]
[0045] 1D , the second set of feature maps 120 is provided to a fully connected layer 124, which then generates an output feature vector 128. Each feature in the output feature vector 128 may include a number corresponding to a possible feature of the image 126, such as "sign," "60," and "100." In some cases, a softmax function (not shown) may convert the numbers in the output feature vector 128 into probabilities. The output 122 of the DCN 100 is then the probability that the image 126 contains one or more features.
[0032]
[0046] A softmax function (not shown) can convert the individual elements of the output feature vector 128 into probabilities such that the output 122 of the DCN 100 is one or more probabilities that the image 126 contains one or more features, such as a sign having the number "60" on it, as in the input image 126. Thus, in this example, the probabilities in the output 122 for "sign" and "60" should be higher than the probabilities for others of the output 122, such as "30", "40", "50", "70", "80", "90", and "100".
[0033]
[0047] Prior to training the DCN 100, the output 122 produced by the DCN 100 may be inaccurate. Thus, an error may be calculated between the output 122 and a target output known a priori. For example, here the target output is an indication that the image 126 contains a "sign" and the number "60." Then, utilizing the known target output, the weights of the DCN 100 may be adjusted through training such that subsequent outputs 122 of the DCN 100 achieve the target output.
[0034]
[0048] To adjust the weights of the DCN 100, the learning algorithm may calculate a gradient vector for the weights. The gradient may indicate the amount by which the error would increase or decrease if the weights were adjusted in a particular way. The weights may then be adjusted to reduce the error. This method of adjusting the weights is sometimes called "backpropagation" because it involves a "backward pass" through the layers of the DCN 100.
[0035]
[0049] In practice, the error gradient of the weights may be calculated over a small number of examples such that the calculated gradient approximates the true error gradient. This approximation method is sometimes called stochastic gradient descent. Stochastic gradient descent may be iterated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level.
[0036]
[0050] After training, the DCN 100 may be presented with new images and the DCN 100 can generate inferences such as classifications or probabilities that various features are present in the new image.
[0037] Convolution Techniques for Convolutional Neural Networks
[0051] Convolution is commonly used to extract useful features from an input dataset. For example, in convolutional neural networks as described above, convolution allows the extraction of different features using kernels and / or filters whose weights are automatically learned during training. The extracted features are then combined to make inferences.
[0038]
[0052] Activation functions can be applied before and / or after each layer of a convolutional neural network. An activation function is generally a mathematical function (e.g., a formula) that determines the output of a node of a neural network. Thus, an activation function determines whether a node should pass on information based on whether the node's input is relevant to the model's prediction. In one embodiment, if y=conv(x) (i.e., y=convolution of x), then both x and y can generally be considered "activations". However, for a particular convolution operation, x may also be referred to as a "pre-activation" or "input activation" since it exists before the particular convolution, and y may be referred to as an output activation or feature map.
[0039]
[0053] 2 shows an example of conventional convolution where an input image of 12 pixels by 12 pixels by 3 channels is convolved using a 5×5×3 convolution kernel 204 and a stride (or step size) of 1. The resulting feature map 206 is 8 pixels by 8 pixels by 1 channel. As can be seen in this example, conventional convolution can vary the dimensionality of the input data compared to the output data (here, from 12×12 to 8×8 pixels) and the channel dimensionality (here, from 3 to 1 channel).
[0040]
[0054] One way to reduce the computational load (e.g., measured in floating point operations per second (FLOPs)) and number parameters associated with neural networks with convolutional layers is to factorize the convolutional layers. For example, a spatially separable convolution as shown in FIG. 2 may be factorized into two components: (1) a depth-wise convolution (e.g., spatial fusion), where each spatial channel is independently convolved with a depth-wise convolution, and (2) a point-wise convolution (e.g., channel fusion), where all spatial channels are linearly combined. An example of a depth-wise separable convolution is shown in FIG. 3A and FIG. 3B. In general, during spatial fusion, the network learns features from the spatial plane, and during channel fusion, the network learns the relationship between these features across channels.
[0041]
[0055] In one embodiment, separable depthwise convolution can be implemented using a 3×3 kernel for spatial fusion and a 1×1 kernel for channel fusion. Specifically, channel fusion can use a 1×1×d kernel that iterates through every single point in the input image at depth d, where the depth of the kernel d generally matches the number of channels in the input image. Channel fusion with pointwise convolution is useful for dimensionality reduction for efficient computation. Applying a 1×1×d kernel and adding an activation layer after the kernel can give the network additional depth, which can increase its performance.
[0042]
[0056] 3A and 3B show an example of a depthwise separable convolution operation.
[0043]
[0057] Specifically, in FIG. 3A, a 12 pixel by 12 pixel by 3 channel input image 302 is convolved with a filter comprising three separate kernels 304A-C, each having a dimensionality of 5 by 5 by 1, to generate an 8 pixel by 8 pixel by 3 channel feature map 306, with each channel generated by a separate kernel in 304A-C.
[0044]
[0058] The feature map 306 is then further convolved using a point-wise convolution operation where a kernel 308 (e.g., a kernel) has dimensionality 1×1×3 to generate an 8 pixel×8 pixel×1 channel feature map 310. As shown in this example, the feature map 310 has reduced dimensionality (1 channel vs. 3), which allows for more efficient computations using the feature map 310.
[0045]
[0059] The results of the depthwise separable convolution in FIGS. 3A and 3B are substantially similar to the conventional convolution in FIG. 2, but the number of computations is significantly reduced, and therefore the depthwise separable convolution offers significant efficiency gains, when the network design permits.
[0046]
[0060] Although not shown in FIG. 3B, multiple (e.g., m) pointwise convolution kernels 308 (e.g., individual components of a filter) can be used to increase the channel dimensionality of the convolution output. Thus, for example, m=256 1×1×3 kernels 308 can be generated, each of which outputs an 8 pixel×8 pixel×1 channel feature map (e.g., 310), which can be stacked to obtain a resulting 8 pixel×8 pixel×256 channel feature map. The resulting increase in channel dimensionality can provide more parameters for training, thereby improving the ability of the convolutional neural network to identify features (e.g., in the input image 302).
[0047] Exemplary Computation in Memory (CIM) Architecture
[0061] 4 illustrates an exemplary compute-in-memory (CIM) array 400 configured to perform machine learning model computations according to aspects of the present disclosure. In this example, the CIM array 400 is configured to simulate MAC operations using mixed analog / digital operations for artificial neural networks. Thus, as used herein, the terms multiplication and addition may refer to such simulated operations. The CIM array 400 may be used to implement aspects of the processing techniques described herein.
[0048]
[0062] In the illustrated embodiment, the CIM array 400 includes precharge word lines (PCWL) 425a, 425b, and 425c (collectively 425), read word lines (RWL) 427a, 427b, and 427c (collectively 427), analog-to-digital converters (ADC) 410a, 410b, and 410c (collectively 410), digital processing unit 413, bit lines 418a, 418b, and 418c (collectively 418), PMOS transistors 411a-111i (collectively 411), NMOS transistors 413a-413i (collectively 413), and capacitors 423a-423i (collectively 423).
[0049]
[0063] Weights associated with neural network layers may be stored in SRAM cells of CIM array 400. In this example, binary weights are shown in SRAM bit cells 405a-405i of CIM array 400. Input activations (e.g., input values which may be input vectors) are provided on PCWLs 425a-c.
[0050]
[0064] The multiplication is performed in each bit cell 405a-405i of the CIM array 400 associated with a bit line, and the accumulation (sum) of all bit cell multiplication results is performed on the same bit line for one column. The multiplication in each bit cell 405a-405i is a form of operation equivalent to an AND operation of the corresponding activation and weight, and the result is stored as a charge on the corresponding capacitor 423. For example, a product of 1, and therefore a charge on the capacitor 423, is generated only if the activation is 1 (here, PMOS is used, so PCWL is 0 for an activation of 1) and the weight is 1.
[0051]
[0065] For example, in the accumulation phase, RWL 427 is switched high to allow any charge on capacitor 423 (based on the corresponding bit cell (weight) and PCWL (activation) values) to accumulate on the corresponding bit line 418. The voltage value of the accumulated charge is then converted to a digital value by ADC 410 (e.g., the output value may be a binary value indicating whether the total charge is greater than a reference voltage). These digital values (outputs) can be provided as inputs to another aspect of the machine learning model, such as the next layer.
[0052]
[0066] When the activations on precharge word lines (PCWLs) 425a, 425b, and 425c are, for example, 1, 0, 1, the sums of bit lines 418a-c correspond to 0+0+1=1, 1+0+0=1, and 1+0+1=2, respectively. The outputs of ADCs 410a, 410b, and 410c are passed to digital processing unit 413 for further processing. For example, if CIM 100 is processing multi-bit weighted values, the digital outputs of ADCs 110 can be summed to generate a final output.
[0053]
[0067] The exemplary 3×3 CIM circuit 400 can be used, for example, to perform an efficient three-channel convolution for a three-element kernel (or filter), with each kernel weight corresponding to an element in each of the three columns, so that for a given three-element receptive field (or input data patch), the output of each of the three channels is computed in parallel.
[0054]
[0068] In particular, although Figure 4 illustrates an embodiment of a CIM using SRAM cells, other memory types may be used, for example, in other embodiments dynamic random access memory (DRAM), magnetoresistive random-access memory (MRAM), and resistive random-access memory (ReRAM or RRAM) may be used as well.
[0055]
[0069] FIG. 5A shows additional details of an example bitcell 500.
[0056]
[0070] The embodiment of Figure 5A may be illustrative of or otherwise related to the embodiment of Figure 4. Particularly, bit line 521 is similar to bit line 418a, capacitor 523 is similar to capacitor 423 of Figure 4, read word line 527 is similar to read word line 427a of Figure 4, precharge word line 525 is similar to precharge word line 425a of Figure 4, PMOS transistor 511 is similar to PMOS transistor 411a of Figure 1, and NMOS transistor 513 is similar to NMOS transistor 413 of Figure 1.
[0057]
[0071] The bit cell 500 includes a static random access memory (SRAM) cell 501 (which may represent the SRAM bit cell 405a of FIG. 4), as well as a transistor 511 (e.g., a PMOS transistor) and a transistor 513 (e.g., an NMOS transistor) and a capacitor 523 coupled to ground. Although a PMOS transistor is used for the transistor 511, other transistors (e.g., an NMOS transistor) may be used in place of the PMOS transistor, along with corresponding adjustments (e.g., inversions) of their respective control signals. The same applies to the other transistors described herein. The additional transistors 511 and 513 are included to implement a computational array in memory according to aspects of the present disclosure. In one aspect, the SRAM cell 501 is a conventional six transistor (6T) SRAM cell.
[0058]
[0072] Programming the weights in the bit cells may be performed once for many activations. For example, during operation, SRAM cell 501 receives only one bit of information at nodes 517 and 519 via write word line (WWL) 516. For example, during a write (WWL 216 is high), if write bit line (WBL) 229 is high (e.g., "1"), node 217 is set high and node 219 is set low (e.g., "0"), or if WBL 229 is low, node 217 is set low and node 219 is set high. Conversely, during a write (WWL 216 is high), if write bit bar line (WBBL) 231 is high, node 217 is set low and node 219 is set high, or if WBBL 229 is low, node 217 is set high and node 219 is set low.
[0059]
[0073] The programming of the weights may be followed by an activation input to charge the capacitors according to the corresponding products and a multiplication step. For example, transistor 511 is activated by an activation signal (PCWL signal) via a precharge word line (PCWL) 525 of the in-memory computation array to perform the multiplication step. Transistor 513 is then activated by a signal via another word line (e.g., read word line (RWL) 527) of the in-memory computation array to perform an accumulation of the multiplied value from bit cell 500 with other bit cells of the array, such as described above with respect to FIG.
[0060]
[0074] When node 517 is "0" (e.g., when the stored weight value is "0"), if a low PCWL indicates an activation of "1" at the gate of transistor 511, capacitor 523 is not charged. Thus, no charge is provided to bit line 521. However, when node 517 corresponding to the weight value is "1" and PCWL is set low (e.g., when the activation input is high), it turns on PMOS transistor 511, thereby acting as a short and allowing capacitor 523 to be charged. After capacitor 523 is charged, transistor 511 is turned off, so that charge is stored in capacitor 523. To move charge from capacitor 523 to bit line 521, NMOS transistor 513 is turned on by RWL 527, causing NMOS transistor 513 to act as a short.
[0061]
[0075] Table 1 shows an example of an in-memory computational array operation according to an AND operation setting, such as may be implemented by bit cell 500 of FIG. 5A.
[0062] [Table 1]
[0063]
[0076] The first column (Activation) of Table 1 contains the possible values of the input activation signal.
[0064]
[0077] The second column (PCWL) of Table 1 includes PCWL values that activate transistors designed to implement in-memory computation functions according to aspects of the present disclosure. Since transistor 511 in this example is a PMOS transistor, the PCWL value is the inverse of the activation value. For example, the in-memory computation array includes transistor 511 that is activated by an activation signal (PCWL signal) via precharge word line (PCWL) 525.
[0065]
[0078] The third column (Cell Node) of Table 1 includes weight values stored in the SRAM cell node that correspond to weights in a weight tensor, such as may be used in a convolution operation.
[0066]
[0079] The fourth column (Capacitor Node) of Table 1 shows the resulting product that is stored as a charge on a capacitor. For example, the charge can be stored at the node of capacitor 523 or at the node of one of capacitors 423a-423i. The charge from capacitor 523 is transferred to bit line 521 when transistor 513 is activated. For example, with reference to transistor 511, when the weight at cell node 517 is "1" (e.g., high voltage) and the input activation is "1" (so PCWL is "0"), capacitor 523 is charged (e.g., capacitor node is "1"). For all other combinations, the capacitor node has a value of 0.
[0067]
[0080] FIG. 5B shows additional details of another example bitcell 550.
[0068]
[0081] Bitcell 550 differs from bitcell 500 of FIG. 5A primarily based on the inclusion of an additional precharge wordline 552 coupled to an additional transistor 554 .
[0069]
[0082] Table 2 shows an example of an in-memory computational array operation similar to Table 1, except following the XNOR operation setting, such as may be implemented by bit cell 550 of FIG. 5B.
[0070] [Table 2]
[0071]
[0083] The first column (Activation) of Table 2 contains the possible values of the input activation signal.
[0072]
[0084] The second column (PCWL1) of Table 2 includes PCWL1 values that activate transistors designed to implement in-memory computation functions according to aspects of the present disclosure. Here again, transistor 511 is a PMOS transistor, and the PCWL1 value is the reciprocal of the activation value.
[0073]
[0085] The third column (PCWL2) of Table 2 includes PCWL2 values that activate additional transistors designed to implement in-memory computation functions according to aspects of the present disclosure.
[0074]
[0086] The fourth column (Cell Node) of Table 2 includes weight values stored in the SRAM cell node that correspond to weights in a weight tensor, such as may be used in a convolution operation.
[0075]
[0087] The fifth column (Capacitor Node) of Table 2 shows the resulting product that is stored as a charge on a capacitor, such as capacitor 523.
[0076]
[0088] FIG. 6 illustrates an example timing diagram 600 of various signals during a compute-in-memory (CIM) array operation.
[0077]
[0089] In the illustrated embodiment, the first row of the timing diagram 600 shows a precharge word line PCWL (e.g., 425a in FIG. 4 or 525 in FIG. 5A) going low. In this embodiment, a low PCWL indicates an activation of "1". A PMOS transistor turns on when PCWL is low, thereby allowing a capacitor to charge (if the weight is "1"). The second row shows a read word line RWL (e.g., read word line 427a in FIG. 4 or 527 in FIG. 5A). The third row shows a read bit line RBL (e.g., 418 in FIG. 4 or 521 in FIG. 5A), the fourth row shows an analog-to-digital converter (ADC) read signal, and the fifth row shows a reset signal.
[0078]
[0090] For example, referring to transistor 511 in FIG. 5A, charge from capacitor 523 is gradually transferred to the read bit line RBL when the read word line RWL is high.
[0079]
[0091] The summed charge / current / voltage (e.g., 403 in FIG. 4, or summed charge from bit line 521 in FIG. 5A) is passed to a comparator or ADC (e.g., ADC 411 in FIG. 4) where the summed charge is converted to a digital output (e.g., a digital signal / number). The summation of the charges may occur in the accumulation region of timing diagram 600, and the readout from the ADC may be associated with the ADC readout region of timing diagram 600. After the ADC readout is obtained, a reset signal discharges all of the capacitors (e.g., capacitors 423a-423i) in preparation for processing the next set of activation inputs.
[0080]
[0092] The parallel processing techniques of the present disclosure can be useful for any type of edge computing involving artificial neural networks. The techniques have applicability in the inference stage or any other stage of neural network processing. The illustrated example is based on a binary network that can be used when high accuracy is not required, but the same concepts apply to networks that use multi-bit weights.
[0081] Example of convolution in memory
[0093] 7 illustrates an exemplary convolutional layer architecture 700 implemented by a compute-in-memory (CIM) array 708. The convolutional layer architecture 700 may be part of a convolutional neural network (e.g., as described above with respect to FIG. 1D ) and may be designed to process multi-dimensional data, such as tensor data.
[0082]
[0094] In the illustrated example, the input 702 to the convolutional layer architecture 700 has dimensions of 38 (height) x 11 (width) x 1 (depth). The output 704 of the convolutional layer has dimensions of 34 x 10 x 64 and includes 64 output channels corresponding to the 64 kernels of the filter tensor 714 that are applied as part of the convolution process. Furthermore, in this example, each kernel (e.g., the exemplary kernel 712) of the 64 kernels of the filter tensor 714 has dimensions of 5 x 2 x 1 (collectively, the kernels of the filter tensor 714 are equivalent to one 5 x 2 x 64 filter).
[0083]
[0095] During the convolution process, each 5×2×1 kernel is convolved with the input 702 to generate one 34×10×1 layer of output 704. During the convolution, the 640 weights of the filter tensor 714 (5×2×64) can be stored in a computation in memory (CIM) array 708, which in this example includes a column for each kernel (i.e., 64 columns). The activations of each of the 5×2 receptive fields (e.g., receptive field inputs 706) are then input into the CIM array 708 using word lines, e.g., 716, and multiplied by the corresponding weights to generate a 1×1×64 output tensor (e.g., output tensor 710). The output tensor 704 represents the accumulation of the individual 1×1×64 output tensors for all of the receptive fields (e.g., receptive field inputs 706) of the input 702. For simplicity, the CIM array 708 in FIG. 7 shows only a few example lines for the inputs and outputs of the CIM array 708 .
[0084]
[0096] In the illustrated embodiment, the CIM array 708 includes word lines 716, along which the CIM array 708 receives receptive fields (e.g., receptive field input 706), as well as bit lines 718 (corresponding to columns of the CIM array 708). Although not shown, the CIM array 708 may also include precharge word lines (PCWL) and read word lines RWL (as discussed above with respect to Figures 4 and 5).
[0085]
[0097] In this example, the word line 716 is used for the initial weight definition. However, once the initial weight definition is done, the activation input activates a specially designed line in the CIM bit cell to perform the MAC operation. Thus, each intersection of the bit line 718 and the word line 716 represents a filter weight value, which is multiplied by the input activation on the word line 716 to generate a product. The individual products along each bit line 718 are then summed to generate a corresponding output value in the output tensor 710. The sum value may be a charge, a current, or a voltage. In this example, the dimensions of the output tensor 704 after processing the entire input 702 of the convolution layer are 34×10×64, but only 64 filter outputs are generated by the CIM array 708 at tme. Thus, the processing of the entire input 702 can be completed in 34×10 or 340 cycles.
[0086] A CIM Architecture for Depthwise Separable Convolution
[0098] Computation-in-memory (CIM)-based artificial intelligence (AI) hardware (HW) accelerators can be used for a variety of tasks, including image, sensor, and audio processing AI tasks. CIM can help reduce issues associated with power consumption when moving data from memory. In some cases, data movement can consume more power than computation. Using CIM can result in power savings due to the weight-fixed nature of CIM. In other words, weights for neural network computations can be stored in random access memory (RAM), such as static random access memory (SRAM) memory cells, allowing the computations to be performed in memory, resulting in reduced power consumption.
[0087]
[0099] Although vector-matrix multiplication blocks implemented in memory for CIM architectures can generally perform conventional convolutional neural network processing well, they are not efficient for supporting depthwise separable convolutional neural networks found in many state-of-the-art machine learning architectures. For example, existing CIM architectures generally cannot perform depthwise separable convolutional neural network processing in one phase because each multidimensional filter uses different input channels. Thus, filter weights in the same row may not share the same activation inputs for different channels. As a result, a matrix-matrix multiplication (MxM) architecture is generally required to support depthwise separable convolution processing in one phase cycle.
[0088]
[0100] Conventional solutions to address this shortcoming include adding a separate digital MAC block to handle the processing for the depthwise portion of the separable convolution, while the CIM array can handle the pointwise portion of the separable convolution. However, this hybrid approach results in increased data movement, which can offset the memory efficiency advantages of the CIM architecture. Furthermore, the hybrid approach generally involves additional hardware (e.g., digital multiply and accumulate (DMAC) elements), which increases space and power requirements and increases processing latency. Furthermore, the use of DMAC can affect the timing of processing operations, causing model output timing constraints (or other dependencies) to be exceeded. To solve the problem, various compromises may be necessary, such as reducing the frame rate of the input data, increasing the clock rate of the processing system elements (including the CIM array), and reducing the input feature size.
[0089]
[0101] The CIM architecture described herein improves timing performance of processing operations for depthwise separable convolution. These improvements beneficially result in shorter cycle times and higher total operations per second (TOPS) per Watt of processing power, i.e., TOPS / W, for depthwise separable convolution operations, as compared to conventional architectures that use more hardware (e.g., DMAC) and / or more data movement.
[0090]
[0102] FIG. 8 illustrates a CIM system 800 including multiple CIM arrays according to some embodiments of the present disclosure.
[0091]
[0103] As shown, the CIM system 800 includes a CIM array 802 configured for depth-wise (DW) convolution and a CIM array 804 configured for pointwise (PW) convolution. In some aspects, kernels (e.g., a 3×3 kernel) can be implemented on different columns of the CIM array 802 in a diagonal fashion. For example, a kernel 806 can be implemented using CIM cells on columns 810, 812 (e.g., bit lines) and nine rows 814-1, 814-2 through 814-8, and 814-9 (e.g., word lines (WL), collectively referred to as row 814) to implement a 3×3 filter having a 2-bit weight parameter. Another kernel 808 can be implemented on columns 816, 818 and nine rows 820-1 through 820-9 (collectively referred to as row 820) to implement another 3×3 filter. Thus, the kernels 806 and 808 are implemented on different rows and columns to facilitate parallel convolution operations for DW, i.e., activating the rows and columns of one of the kernels 806, 808 does not affect the rows and columns of the other of the kernels 806, 808. Different activation inputs can be provided to each of the kernels 806, 808, and the kernels 806, 808 can be operated in parallel.
[0092]
[0104] The input activation buffer of each kernel may be filled (e.g., stored) with the corresponding output channel patch from the previous layer. For example, a row (e.g., row 814) of kernel 808 may be coupled to activation buffers 830-1, 830-2 through 830-8, and 830-9 (collectively referred to as activation buffers 830), and a row (e.g., row 820) of kernel 806 may be coupled to activation buffers 832-1 through 832-9 (collectively referred to as activation buffers 832).
[0093]
[0105] The outputs of kernel 806 (e.g., at columns 810, 812) may be coupled to an analog-to-digital converter (ADC) 840, and the outputs of kernel 808 (e.g., at columns 816, 818) may be coupled to an ADC 842. For example, each input of ADC 840 may receive the accumulated charges of row 814 from each of columns 810, 812, and each input of ADC 842 may receive the accumulated charges of row 820 from each of columns 816, 818, based on which each of ADCs 840, 842 generates a digital output signal. ADC 840 receives signals from columns 810, 812 as inputs and generates a digital representation of the signal, taking into account that the bits stored in column 812 represent lower importance in their respective weights than the bits stored in column 810. Similarly, ADC 842 receives as input the signals from columns 816, 818 and generates a digital representation of the signals, taking into account that the bits stored in column 818 represent less importance in their respective weights than the bits stored in column 816.
[0094]
[0106] Although the ADCs 840, 842 are implemented to receive signals from two columns to facilitate analog-to-digital conversion for kernels having 2-bit weight parameters, the aspects described herein may be implemented for ADCs configured to receive signals from any number of columns (e.g., three columns to perform analog-to-digital conversion for kernels having 3-bit weight parameters).
[0095]
[0107] The outputs of the ADCs 840, 842 can be coupled to a nonlinear arithmetic circuit 850 (and buffer) to implement nonlinear operations such as rectified linear unit (ReLU) and average pooling (AvePool), to name a few. Nonlinear operations allow for the creation of complex mappings between inputs and outputs, thus enabling learning and modeling of complex data such as images, videos, audio, and data sets that are nonlinear or have high dimensionality. The output of the nonlinear arithmetic circuit 850 can be coupled to an input activation buffer 860 for the CIM array 804 configured for PW convolution. As shown, the output of the CIM array 804 may be coupled to an ADC 870, the output of which may be provided to a nonlinear arithmetic circuit 880. Although a single ADC 870 is shown, multiple ADCs may be implemented for different columns of the CIM array 804.
[0096]
[0108] Although each of the kernels 806, 808 includes two columns that allow for two-bit weights to be stored in each row of the kernel, the kernels 806, 808 may be implemented using any number of suitable columns, such as one column for one-bit binary weights, or two or more columns for multi-bit weights. For example, each of the kernels 806, 808 may be implemented using three columns to facilitate a three-bit weight parameter being stored in each row of the kernel, or a single column to facilitate a one-bit weight being stored in each row of the kernel. Furthermore, while each of the kernels 806, 808 is implemented using nine rows for a 3×3 kernel for ease of understanding, the kernels 806, 808 may be implemented using any number of rows to implement an appropriate kernel size. Furthermore, more than two kernels may be implemented using a subset of the cells of the CIM array. For example, the CIM array 802 may include one or more other kernels, with all of the kernels of the CIM array 802 being implemented on different rows and columns to facilitate parallel convolution operations. For example, kernel 806 may correspond to kernel 304A described with respect to Figure 3A, and kernel 808 may correspond to kernel 304B described with respect to Figure 3A. Another kernel (not shown in Figure 8) corresponding to kernel 304C may also be implemented on different rows and columns than kernels 806, 808.
[0097]
[0109] FIG. 9 illustrates an example operation 900 for signal processing via the CIM system 800 of FIG. 8 according to some aspects of the disclosure. The operation 900 can start with processing of a DW-CNN layer. For example, in block 904, DW convolution weights can be loaded into CIM cells of a CIM array (e.g., for kernels 806, 808) as described herein. For example, in block 904, DW 3×3 kernel weights can be grouped and written to the CIM array 802 of FIG. 8. That is, 2-bit kernel weights can be provided to columns 810, 812, and pass gate switches of memory cells (e.g., memory cells b01 and b11 shown in FIG. 8) can be closed to store the 2-bit kernel weights in the memory cells. Filter weights can be stored in each row of CIM cells for each of the kernels 806, 808 in a similar manner.
[0098]
[0110] Weights that may have been previously stored in memory cells on the same column but on a different row than the active kernel may be zeroed. For example, a logical zero may be stored in memory cells (not shown) in columns 816, 818 and row 820, and in columns 810, 812 and row 814. In some cases, CIM array 802 may be initially zeroed before storing the weights of kernels 806 and 808.
[0099]
[0111] In some implementations, the CIM array can be divided into tiles. For example, tiles on the same column as an active kernel can be configured in tri-state mode. In tri-state mode, the outputs of the memory cells of the tile can be configured to have a relatively high impedance, effectively eliminating the influence of the cell on the output. As described herein, DW convolution kernels in different columns and rows may be stacked. Both the DW convolution weights and the PW convolution weights can be updated for each subsequent layer.
[0100]
[0112] In block 906, the DW convolution activation inputs (e.g., in activation buffers 830, 832) can be applied to each group of rows of the kernels 806, 808 during the same cycle to generate DW convolution outputs in parallel using both kernels.
[0101]
[0113] In block 908, the ADCs 840, 842 may convert the convolution outputs of the kernels 806, 808 (e.g., in columns 810, 812 and columns 816, 818) from the analog domain to the digital domain. Based on the output of the ADCs 840, 842 for DW convolution, a nonlinear operation may be performed via the nonlinear operation circuit 850.
[0102]
[0114] In block 910, the output from the nonlinear operation circuit 850 may be applied to the PW input (e.g., stored in the input activation buffer 860) of the CIM array 804 to perform the PW convolution. In block 912, the ADC 870 may convert the PW convolution output from the CIM array 804 from the analog domain to the digital domain. Based on the output of the ADC 870 for the PW convolution, a nonlinear operation may be performed via the nonlinear operation circuit 880.
[0103]
[0115] By implementing kernels on different rows and columns, convolution operations can be performed in parallel, facilitating faster processing times and lower dynamic power compared to conventional implementations. In other words, performing parallel convolution operations allows multiple filters to be processed in one cycle, as opposed to processing each filter in a different cycle, saving processing time and reducing dynamic power. In some aspects, each kernel can be iterated multiple times to increase row utilization and reduce ADC range compression, as described in more detail herein.
[0104] Techniques for reducing power consumption and increasing CIM array utilization.
[0116] FIG. 10 illustrates a CIM array 1000 divided into tiles (also called sub-banks) to conserve power, according to some aspects of the disclosure. The CIM array 1000 may have, as one example, 1024 rows and 256 columns. Individual tiles (e.g., sub-banks) of the rows and columns may be enabled or disabled. For example, a tile may include 128 rows and 23 columns. As one example, a tile array 1002 (e.g., including multiple tiles, such as tile 1004) may be active for DW-CNN convolutions, while the remaining tiles may be disabled. In other words, the remaining tiles may be configured in a tri-state mode.
[0105]
[0117] In some implementations, row and column filler cells can be implemented in the CIM array 1000. Filler circuits (e.g., buffers or switches) can be used to enable or disable tiles of the CIM array to save power. The column filler cells can be AND gate logic, and the row filler cells can be buffers on the write bit-lines (WBLs) and transmission switches on the read bit-lines (RBLs). The size and type of the transmission switches can be configured based on linearity specifications.
[0106]
[0118] DW convolution can use relatively small kernel dimensions (3x3, 5x5, ...), and insufficient utilization of the CIM array can impact the output signal to noise ratio (SNR) due to range compression (e.g., the output of the neural network is distributed within a small range due to nonlinear activation). Some aspects of the present disclosure are directed to techniques for improving the SNR. For example, a fine-grained tiling design can be used to mitigate the impact on the SNR, as described in more detail herein with respect to FIG. 11.
[0107]
[0119] FIG. 11 illustrates a CIM array 802 having diagonally stacked kernels in accordance with some aspects of the disclosure. A variety of diagonally stacked kernels can be implemented in the CIM array 802. For example, the CIM array 802 can include CIM cells for kernels 806 and 808, and CIM cells for kernels 1108, 1110, 1112, 1114, 1116, each implemented on a different row and column of the CIM array 802, as described with respect to FIG. 8. As shown, the CIM array 802 can be divided into tiles, such as tiles 1104, 1106, etc. Each of the tiles of the CIM array that do not include at least a portion of a kernel (e.g., tile 1106) can be deactivated to conserve power.
[0108]
[0120] In some aspects, a relatively small tile size can be used (e.g., selected during chip design) for a CIM bank that performs DW convolutions (e.g., 16 rows and 32 columns) to increase CIM array utilization and save power. Using a smaller tile size increases utilization of active CIM cells, which are cells that are not part of disabled tiles.
[0109]
[0121] As an example, three CIM cell groups can be designed to perform the inverse bottleneck of a neural network architecture. Inverse bottleneck operation generally refers to an operation used to expand input features, followed by DW convolution and reduction of DW output dimension via PW convolution. The first CIM cell group (CIM1) can be used for the bottleneck operation, the second CIM cell group (CIM2) can be used for the DW convolution operation, and the third CIM cell group (CIM3) can be used for the bottleneck operation. In some aspects, CIM2 for DW convolutions can have a finer tiling configuration (e.g., 16 rows to implement a 3×3 kernel, or 32 rows to implement a 5×5 kernel) to improve CIM array utilization and save power, while CIM1 and CIM3 can have a coarser tiling (e.g., 64 rows or 128 rows) to avoid the impact of filler cells for non-DW convolution operations (e.g., because using smaller tiles for the CIM array results in a larger number of filler cells). In this way, the reusability of the CIM array library can be doubled for DW and non-DW operations.
[0110]
[0122] As an example, the average (e.g., approximate) CIM utilization with coarse-grained tiling (e.g., each tile using 64 rows and 32 columns of a CIM array having 1024 rows) may be 13.08% for a 3×3 kernel and 31.44% for a 5×5 kernel. In other words, only 13.08% of the active memory cells in the CIM array may be utilized for a 3×3 kernel and 31.44% of the active memory cells in the CIM array may be utilized for a 5×5 kernel. Meanwhile, the average CIM utilization with fine-grained tiling using 16 rows and 32 columns per tile and with a CIM array having 1024 rows may be 40.46% for a 3×3 kernel and 47.64% for a 5×5 kernel. The average CIM utilization with fine tiling using 32 rows and 32 columns per tile of a CIM array with 1024 rows can be 24.18% for a 3×3 kernel and 47.64% for a 5×5 kernel. Thus, fine tiling improves CIM array utilization for filters with smaller kernel sizes (e.g., for DW convolutions). Improving the utilization of the CIM array results in a higher percentage of active memory cells being utilized, reducing power loss caused by unused active memory cells.
[0111]
[0123] In some aspects, utilization can be improved by selecting a tiling size closer to the kernel size. For example, as shown in FIG. 11, a tile size of 16 (e.g., as shown for tile 1104) can be used for a kernel size of 9 (e.g., 9 rows as shown for kernel 806). The tile size may be a power of 2 (logarithmic scale) larger than the kernel size to improve flexibility for handling different neural network models. In some aspects, the kernel can be iterated to improve row utilization and improve the ADC SNR, as described in more detail with respect to FIG. 12.
[0112]
[0124] FIG. 12 illustrates a CIM array implemented with an iterated kernel in accordance with some aspects of the present disclosure.
[0113]
[0125] As shown, multiple kernels can be repeated to form a kernel group. For example, multiple kernels can be implemented on the same column, such as kernel 806, 1204, or kernel 808, 1208. The same weight parameters can be stored in the repeated kernels of the kernel group on the same column (e.g., kernel 806, 1204), and the same activation inputs can be provided to the repeated kernels. Thus, the repeated kernels can generate the same output signal that is combined in each column (output), resulting in an increase in the dynamic range at the output of the repeated kernels. For example, using two repeated kernels can result in a doubling of the dynamic range at the output of the repeated kernels provided to the ADC (e.g., ADC 840). Increasing the dynamic range at the output of the kernels facilitates analog-to-digital conversion with higher accuracy, since a wider range of the ADC can be utilized. In other words, using the full range of the ADC input allows the digital output of the ADC to more accurately identify the analog input of the ADC and improve the SNR of the ADC.
[0114]
[0126] In some cases, the number of DW convolution channels that can be implemented in a CIM array may be limited by the dimensions of the CIM array. For example, when implementing a 3×3 filter, 113 channels can be implemented for a CIM array with 1024 rows (e.g., because 113×9 is less than 1024). In other words, the DW kernel for the DW convolution may not fit into one CIM array due to the limitations on the number of rows or columns associated with the CIM array. Therefore, the input activations and DW convolution weights can be configured by a sequencer to allow partial DW convolution channel sums to be calculated.
[0115]
[0127] In some cases, the maximum number of kernels that can be implemented in the CIM array may be less than the total number of kernels for all channels. The maximum number of kernels can be implemented in the CIM array. Then, all corresponding channel inputs can be processed to generate partial channel outputs. Then, the array can be loaded with the next batch of kernels, and the partial outputs can be processed until all kernels are processed. As another example, the DW convolution input batch size can be determined based on the subsequent PW layer dimension information. The kernels can be loaded multiple times to process the input batch size. The partial DW output can then be fed to the next PW convolution layer to generate the partial bottleneck output.
[0116] Exemplary Operations for Performing Neural Network Processing in a CIM Array
[0128] 13 is a flow diagram illustrating example operations 1300 for signal processing in a neural network according to some aspects of the disclosure. The operations 1300 may be performed by a controller, such as the CIM controller 1432, as described with respect to FIG.
[0117]
[0129] The operation 1300 begins in block 1305 by the controller loading a first set of computation in memory (CIM) cells having a first plurality of weight parameters of a first kernel (e.g., kernel 806 of FIG. 8 ) via one or more first columns (e.g., 810, 812 of FIG. 8 ) to perform a neural network calculation (e.g., a DW neural network calculation), the first set of CIM cells having one or more first columns and a first plurality of rows (e.g., row 814 of FIG. 8 ) of a CIM array (e.g., CIM array 802 of FIG. 8 ). At block 1310, the controller loads a second set of CIM cells having a second plurality of weight parameters of a second kernel (e.g., kernel 808 of FIG. 8 ) via one or more second columns (e.g., columns 816, 818 of FIG. 8 ), the second set of CIM cells having one or more second columns and a second plurality of rows (e.g., rows 820 of FIG. 8 ) of the CIM array to perform the neural network calculation. For example, the first set of CIM cells can include a subset of the cells of the CIM array, and the second set of CIM cells includes another subset of the cells of the CIM array. In some aspects, the one or more first columns may be different from the one or more second columns, and the first plurality of rows may be different from the second plurality of rows. At block 1315, the controller can perform the neural network calculation by applying a first activation input to the first plurality of rows and applying a second activation input to the second plurality of rows.
[0118]
[0130] In some aspects, the operations 1300 may also include loading a third plurality of weights of the third kernel into another CIM array (e.g., CIM array 804 of FIG. 8) to perform the pointwise neural network calculation. The controller may also generate input signals to the second CIM array (e.g., provided via input activation buffer 860 of FIG. 8) based on output signals from the depth-wise neural network calculation.
[0119]
[0131] In some aspects, the operations 1300 may also include generating a first digital signal (e.g., via ADC 840 of FIG. 8) by converting the voltages on one or more first columns from the analog domain to the digital domain, and generating a second digital (e.g., ADC 842 of FIG. 8) signal by converting the voltages on one or more second columns from the analog domain to the digital domain. The operations 1300 may also include performing a nonlinear activation operation (e.g., via nonlinear activation circuit 850) based on the first digital signal and the second digital signal.
[0120]
[0132] In some aspects, the kernels can be iterated to improve CIM array utilization and increase input range compression for the ADC. For example, the controller may also load a first plurality of weight parameters of a third kernel (e.g., kernel 1204 of FIG. 12) to a third CIM cell via one or more first columns to perform a neural network calculation. The third CIM cells may be on one or more first columns and a third plurality of rows of the CIM array. The controller may perform the neural network calculation by applying at least a first activation input (e.g., the same activation input provided to the first kernel) to the third plurality of rows. As described herein, each bit of the weight parameter can be stored via a column of the kernel. For example, an amount of one or more first columns may be associated with an amount of one or more bits of each of the first plurality of weight parameters, and an amount of one or more second columns may be associated with an amount of one or more bits of each of the second plurality of weight parameters.
[0121] Exemplary Processing System for Performing Phase-Selective Convolution
[0133] 14 illustrates an example electronic device 1400. The electronic device 1400 can be configured to perform the methods described herein, including the operations 1300 described with respect to FIG.
[0122]
[0134] The electronic device 1400 includes a central processing unit (CPU) 1402, which in some aspects may be a multi-core CPU. Instructions executed by the CPU 1402 may be loaded from a program memory associated with the CPU 1402 or may be loaded from memory 1424, for example.
[0123]
[0135] The electronic device 1400 also includes additional processing blocks tailored to specific functions, such as a graphics processing unit (GPU) 1404, a digital signal processor (DSP) 1406, a neural processing unit (NPU) 1408, a multimedia processing block 1410, and a wireless connectivity processing block 1412. In one implementation, the NPU 1408 is implemented in one or more of the CPU 1402, the GPU 1404, and / or the DSP 1406.
[0124]
[0136] In some aspects, the wireless connection processing block 1412 may include components for, for example, third generation (3G), fourth generation (4G) (e.g., 4G LTE), fifth generation (e.g., 5G or NR), Wi-Fi, Bluetooth, and wireless data transmission standards. The wireless connection processing block 1412 is further coupled to one or more antennas 1414 to facilitate wireless communication.
[0125]
[0137] The electronic device 1400 may also include one or more sensor processors 1416 associated with any type of sensor, one or more image signal processors (ISP) 1418 associated with any type of image sensor, and / or a navigation processor 1420, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.
[0126]
[0138] Electronic device 1400 may also include one or more input and / or output devices 1422, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, a speaker, a microphone, etc. In some aspects, one or more of the processors of electronic device 1400 may be based on the ARM instruction set.
[0127]
[0139] The electronic device 1400 also includes memory 1424, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1424 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the electronic device 1400 or the CIM controller 1432. For example, the electronic device 1400 may include a CIM circuit 1426 including one or more CIM arrays, such as CIM array 802 and CIM array 804, as described herein. The CIM circuit 1426 may be controlled via the CIM controller 1432. For example, in some aspects, the memory 1424 may include code 1424A for loading (e.g., loading weight parameters into CIM cells) and code 1424B for calculating (e.g., performing neural network calculations by applying activation inputs). As shown, the CIM controller 1432 can include circuitry 1428A for loading (e.g., loading weight parameters into CIM cells) and circuitry 1428B for computing (e.g., performing neural network computations by applying activation inputs). The components shown and other components not shown can be configured to perform various aspects of the methods described herein.
[0128]
[0140] In some aspects, such as when the electronic device 1400 is a server device, various aspects, such as one or more of the multimedia components 1410, the wireless connection components 1412, the antenna 1414, the sensors 1416, the ISP 1418, or the navigation 1420, may be omitted from the aspects shown in FIG. 14 .
[0129] Example clause
[0141] Aspect 1. An apparatus for signal processing in a neural network, comprising: first computation in memory (CIM) cells configured as a first kernel for neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array; and a second set of CIM cells configured as a second kernel for neural network computation, the second set of CIM cells including one or more second columns and a second plurality of rows of the CIM array, wherein the one or more first columns are different from the one or more second columns and the first plurality of rows are different from the second plurality of rows.
[0130]
[0142] Aspect 2. The apparatus of aspect 1, wherein the first set of CIM cells comprises a subset of cells of the CIM array, and the second set of CIM cells comprises another subset of cells of the CIM array.
[0131]
[0143] Embodiment 3. The apparatus of embodiment 1 or 2, wherein the neural network calculation comprises a depth-wise (DW) neural network calculation.
[0132]
[0144] Aspect 4. The apparatus of aspect 3, further comprising another CIM array configured as a third kernel for point-wise (PW) neural network calculations, wherein an input signal to the other CIM array is generated based on an output signal from the CIM array.
[0133]
[0145] Embodiment 5. The apparatus of any one of embodiments 1 to 4, further comprising a first analog-to-digital converter (ADC) coupled to the one or more first columns and a second ADC coupled to the one or more second columns.
[0134]
[0146] Example 6. The apparatus of example 5, further comprising a nonlinear activation circuit coupled to the outputs of the first ADC and the second ADC.
[0135]
[0147] Aspect 7. The apparatus of any one of aspects 1 to 6, further comprising a third CIM cell configured as a third kernel for neural network calculations, the third CIM cell being on one or more of the first columns and a third plurality of rows of the CIM array.
[0136]
[0148] Aspect 8. The apparatus of aspect 7, wherein the same weight parameter is configured to be stored in the first set of CIM cells and in the third CIM cell.
[0137]
[0149] Aspect 9. The apparatus of any one of aspects 1 to 8, wherein one or more of the first set of CIM cells on each row of the first plurality of rows are configured to store a first weight parameter, and one or more of the second set of CIM cells on each row of the second plurality of rows are configured to store a second weight parameter.
[0138]
[0150] Aspect 10. The apparatus of aspect 9, wherein a quantity of one or more first columns is associated with a quantity of one or more bits of a first weighting parameter, and a quantity of one or more second columns is associated with a quantity of one or more bits of a second weighting parameter.
[0139]
[0151] Aspect 11. A method of signal processing in a neural network, comprising: loading a first plurality of weight parameters of a first kernel into a first computation in memory (CIM) cell via one or more first columns to perform a neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array; loading a second plurality of weight parameters of a second kernel into a second set of CIM cells via one or more second columns to perform the neural network computation, the second set of CIM cells including one or more second columns and a second plurality of rows of the CIM array, the one or more first columns being different from the one or more second columns and the first plurality of rows being different from the second plurality of rows; and performing the neural network computation by applying a first activation input to the first plurality of rows and applying a second activation input to the second plurality of rows.
[0140]
[0152] Aspect 12. The method of aspect 11, wherein the first set of CIM cells comprises a subset of cells of the CIM array, and the second set of CIM cells comprises another subset of cells of the CIM array.
[0141]
[0153] Aspect 13. The method of aspect 11 or 12, wherein the neural network calculation comprises a depth-wise (DW) neural network calculation.
[0142]
[0154] Aspect 14. The method of aspect 13, further comprising: loading a third plurality of weights of a third kernel into another CIM array to perform a point-wise (PW) neural network calculation; and generating an input signal to the other CIM array based on an output signal from the DW neural network calculation.
[0143]
[0155] Embodiment 15. The method of any one of embodiments 11 to 14, further comprising: generating a first digital signal by converting a voltage on one or more first strings from the analog domain to the digital domain; and generating a second digital signal by converting a voltage on one or more second strings from the analog domain to the digital domain.
[0144]
[0156] Aspect 16. The method of aspect 15, further comprising performing a non-linear activation operation based on the first digital signal and the second digital signal.
[0145]
[0157] Aspect 17. The method of any one of aspects 11 to 16, further comprising loading a first plurality of weight parameters of a third kernel into a third CIM cell via one or more first columns to perform a neural network calculation, the third CIM cell being on one or more of the first columns and a third plurality of rows of the memory, and performing the neural network calculation further comprising applying a first activation input to the third plurality of rows.
[0146]
[0158] Aspect 18. A method according to any one of aspects 11 to 17, wherein the quantity of one or more first columns is associated with the quantity of one or more bits of each of a first plurality of weight parameters, and the quantity of one or more second columns is associated with the quantity of one or more bits of each of a second plurality of weight parameters.
[0147]
[0159] Aspect 19. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method of signal processing in a neural network, the method comprising: loading a first plurality of weight parameters of a first kernel into first computation-in-memory (CIM) cells via one or more first columns to perform a neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array; and loading a first plurality of weight parameters of a first kernel into first computation-in-memory (CIM) cells via one or more first columns to perform a neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array. , loading a second plurality of weight parameters of a second kernel into a second set of CIM cells via one or more second columns, the second set of CIM cells including one or more second columns and a second plurality of rows of a CIM array, the one or more first columns being different from the one or more second columns and the first plurality of rows being different from the second plurality of rows; and performing a neural network computation by applying a first activation input to the first plurality of rows and applying a second activation input to the second plurality of rows.
[0148] Additional Considerations
[0160] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The embodiments described herein are not intended to limit the scope, applicability, or aspects described in the claims. Various modifications of these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of the elements described without departing from the scope of the disclosure. Various embodiments may omit, substitute, or add various procedures or components as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some embodiments may be combined in some other embodiments. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects described herein. Additionally, the scope of the disclosure is intended to encompass such apparatus or methods practiced using other structures, functions, or structures and functions in addition to or other than the various aspects of the disclosure described herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.
[0149]
[0161] As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.
[0150]
[0162] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items and includes single members. By way of example, "at least one of a, b, or c" is intended to encompass a, b, c, ab, ac, bc, and abc, as well as any combination having multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other permutation of a, b, and c).
[0151]
[0163] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, etc. Also, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Also, "determining" can include resolving, selecting, choosing, establishing, etc.
[0152]
[0164] The methods disclosed herein include one or more steps or actions for achieving the method. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Furthermore, various operations of the methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component(s) and / or module(s), including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.
[0153]
[0165] The following claims are not limited to the embodiments set forth herein, but are to be accorded the full scope consistent with the language of the claims. In the claims, reference to an element in the singular is not intended to mean "one and only one," but rather "one or more." Unless otherwise specified, the term "several" refers to one or more. No element of a claim is to be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means of," or, in the case of a method claim, unless the element is recited using the phrase "step of." All structural and functional equivalents of the elements of the various embodiments described throughout this disclosure that are known or that later become known to those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be made public, regardless of whether such disclosure is expressly recited in the claims.
Claims
1. An apparatus comprising: A first set of in-memory computing (CIM) cells configured as a first kernel for depthwise (DW) separable convolution operations of a convolutional neural network, the first set of CIM cells being included in one or more first columns and a first plurality of rows of a first CIM array; A second set of CIM cells configured as a second kernel for the DW convolution operation, the second set of CIM cells being included in one or more second columns and a second plurality of rows of the first CIM array; Wherein The one or more first columns are different from the one or more second columns; The rows among the first plurality of rows are different from the rows among the second plurality of rows; A second CIM array configured as a third kernel for pointwise (PW) convolution operations; Wherein an input signal to the second CIM array is generated based on an output signal from the first CIM array; An apparatus.
2. A first analog-to-digital converter (ADC) coupled to the one or more first columns; A second ADC coupled to the one or more second columns; The apparatus according to claim 1, further comprising.
3. The apparatus according to claim 2, further comprising a non-linear activation circuit coupled to outputs of the first ADC and the second ADC.
4. The apparatus according to claim 1, further comprising a third set of CIM cells configured as a third kernel for the DW convolution operation of the convolutional neural network, the third set of CIM cells being within the one or more first columns and a third plurality of rows of the first CIM array.
5. The apparatus according to claim 4, wherein the apparatus is configured to store the same weight parameters in the first set of CIM cells and the third set of CIM cells.
6. One or more cells of the first set of CIM cells within each row of the first plurality of rows are configured to store a first weight parameter; One or more cells of the second set of CIM cells within each row of the second plurality of rows are configured to store a second weight parameter; The apparatus according to claim 1.
7. The amount of the one or more first columns is associated with the amount of one or more bits of the first weight parameter; The amount of the one or more second columns is associated with the amount of one or more bits of the second weight parameter. The apparatus according to claim 6. **Claim 8** A method of operating the apparatus according to claim 1, comprising: Loading a first plurality of weight parameters of the first kernel into the first set of CIM cells in memory via the one or more first columns to perform the DW convolution operation of the convolutional neural network, wherein the first set of CIM cells is included in the one or more first columns and the first plurality of rows of the first CIM array; Loading a second plurality of weight parameters of the second kernel into the second set of CIM cells via the one or more second columns to perform the DW convolution operation of the convolutional neural network, wherein the second set of CIM cells is included in the one or more second columns and the second plurality of rows of the first CIM array, the one or more first columns are different from the one or more second columns, and the rows among the first plurality of rows are different from the rows among the second plurality of rows; Performing the DW convolution operation of the convolutional neural network by applying a first activation input to the first plurality of rows and a second activation input to the second plurality of rows; Loading a third plurality of weights of the third kernel into the second CIM array to perform the PW convolution operation; Generating an input signal to the second CIM array based on the output signal from the DW convolution operation of the convolutional neural network; A method comprising the above steps. **Claim 9** Generating a first digital signal by converting a voltage in the one or more first columns from an analog region to a digital region; Generating a second digital signal by converting a voltage in the one or more second columns from the analog region to the digital region; The method according to claim 8, further comprising the above steps. **Claim 10** The method according to claim 9, further comprising performing a non-linear activation operation based on the first digital signal and the second digital signal. **Claim 11** To execute the DW convolution operation of the convolutional neural network, further comprising loading, via the one or more first columns, the first plurality of weight parameters of the third kernel into a third set of CIM cells, the third set of CIM cells being in the one or more first columns and a third plurality of rows of the first CIM array, and executing the DW convolution operation of the convolutional neural network further comprises applying the first activation input to the third plurality of rows. The method according to claim 8. Claim 12 The amount of the one or more first columns is associated with the amount of one or more bits of each of the first plurality of weight parameters. The amount of the one or more second columns is associated with the amount of one or more bits of each of the second plurality of weight parameters. The method according to claim 8. Claim 13 A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to execute the method according to any one of claims 8 to 12.