An In-Memory Computation Architecture for Depthwise Convolution

JP2024525333A5Active Publication Date: 2025-06-05QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577151
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-29
Filing Date
2022-06-28
Publication Date
2025-06-05
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Existing machine learning systems face inefficiencies in processing machine learning model data due to the need for additional hardware elements like digital multiply-and-accumulate circuits, which consume space, power, and increase complexity, especially in edge devices and IoT devices, and existing in-memory computation methods struggle to implement depth-separable convolutional neural networks without these components.

Method used

A compute-in-memory (CIM) array architecture that includes CIM cells configured for depthwise and pointwise neural network computations, allowing for efficient implementation of depthwise separable convolutions without additional hardware, by performing multiple convolution operations sequentially on a single CIM array.

Benefits of technology

This approach reduces power consumption and increases computational efficiency by performing calculations directly in memory, enabling more efficient processing of machine learning models, particularly in resource-constrained devices like mobile devices and IoT devices, while maintaining performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Some aspects provide an apparatus for signal processing in a neural network. The apparatus generally includes first computation in memory (CIM) cells configured as a first kernel for neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array. The apparatus also includes a second set of CIM cells configured as a second kernel for neural network computation, the second set of CIM cells including one or more first columns and a second plurality of rows of the CIM array. The first plurality of rows may be different from the second plurality of rows.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001]

[0001] This application claims priority to U.S. Application No. 17 / 361,807, filed June 29, 2021, which is assigned to the assignee of this application and is incorporated by reference in its entirety into this specification. [Technical field]

[0002] Aspects of the present disclosure relate to performing machine learning tasks, and in particular to in-memory computational architectures and data flows for performing depthwise separable convolutions in memory. [Background technology]

[0003]

[0003] Machine learning is generally the process of creating a trained model (e.g., an artificial neural network, tree, or other structure) that represents a generalized fit to a set of training data that is known a priori. By applying the trained model to new data, inferences are generated, which can be used to gain insight into the new data. Sometimes, applying a model to new data is described as "performing inference" on the new data.

[0004]

[0004] As the use of machine learning has proliferated to enable various machine learning (or artificial intelligence) tasks, a need has arisen for more efficient processing of machine learning model data. In some cases, specialized hardware such as machine learning accelerators can be used to enhance the processing system's ability to process machine learning model data. However, such hardware requires space and power, which is not always available on the processing device. For example, "edge processing" devices such as mobile devices, always-on devices, and internet of things (IoT) devices must balance processing power with power and packaging constraints. Furthermore, accelerators may need to move data across a common data bus, which can cause significant power usage and introduce latency to other processes sharing the data bus. Therefore, other aspects of the processing system are considered to process machine learning model data.

[0005]

[0005] Memory devices are an example of another aspect of a processing system that can be leveraged to perform processing of machine learning model data through so-called computation in memory (CIM) processes. Unfortunately, CIM processes may not be able to perform processing of complex model architectures such as depthwise separable convolutional neural networks without additional hardware elements such as digital multiply-and-accumulate circuits (DMACs) and related peripherals. These additional hardware elements use additional space, power, and complexity in their implementation, which tends to reduce the benefits of leveraging memory devices as additional computational resources. Even if auxiliary aspects of a processing system have DMACs available to perform processing that cannot be performed directly in memory, moving data to and from those auxiliary aspects requires time and power, thus reducing the benefits of the CIM process.

[0006]

[0006] Therefore, there is a need for systems and methods for performing in-memory computations of a wider variety of machine learning model architectures, such as deeply separable convolutional neural networks. Summary of the Invention

[0007]

[0007] Some aspects provide an apparatus for signal processing in a neural network. The apparatus generally includes a first computation in memory (CIM) cell configured as a first kernel for depthwise (DW) neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array, and a second set of CIM cells configured as a second kernel for neural network computation, the second set of CIM cells including one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows. The apparatus may also include a third set of CIM cells of the CIM array configured as a third kernel for pointwise (PW) neural network computation.

[0008]

[0008] Some aspects provide a method of signal processing in a neural network. The method generally includes performing a plurality of DW convolution operations via a plurality of kernels implemented using a plurality of CIM cell groups on one or more first columns of a CIM array, and generating an input signal for a PW convolution operation based on an output from the plurality of DW convolution operations. The method also includes performing a PW convolution operation based on the input signal, the PW convolution operation being performed via a kernel implemented using a CIM cell group on one or more second columns of the CIM array.

[0009]

[0009] Some aspects provide a non-transitory computer-readable medium having instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method of signal processing in a neural network. The method generally includes performing a plurality of DW convolution operations via a plurality of kernels implemented using a plurality of CIM cell groups on one or more first columns of a CIM array, and generating an input signal for a PW convolution operation based on an output from the plurality of DW convolution operations. The method also includes performing a PW convolution operation based on the input signal, the PW convolution operation being performed via a kernel implemented using a CIM cell group on one or more second columns of the CIM array.

[0010]

[0010] Another aspect provides a processing system configured to perform the aforementioned methods and methods further described herein; a non-transitory computer readable medium comprising instructions which, when executed by one or more processors of the processing system, cause the processing system to perform the aforementioned methods and methods further described herein; a computer program product embodied on a computer readable storage medium comprising code for performing the aforementioned methods and methods further described herein; and a processing system comprising means for performing the aforementioned methods and methods further described herein.

[0011] The following description and the related drawings set forth in detail certain illustrative features of the one or more aspects. [Brief description of the drawings]

[0012]

[0012] The accompanying drawings illustrate some aspects of the one or more aspects and therefore should not be considered as limiting the scope of the present disclosure. [Figure 1A]

[0013] 1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1B] 1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1C]1A-1C are diagrams illustrating examples of various types of neural networks. [Figure 1D] 1A-1C are diagrams illustrating examples of various types of neural networks. [Diagram 2]

[0014] FIG. 1 is a diagram illustrating an example of a conventional convolution operation. [Figure 3A]

[0015] FIG. 13 is a diagram illustrating an example of a depthwise separable convolution operation. [Figure 3B] FIG. 13 is a diagram illustrating an example of a depthwise separable convolution operation. [Figure 4]

[0016] 1 illustrates an exemplary compute-in-memory (CIM) array configured to perform machine learning model computations. [Figure 5A]

[0017] FIG. 5 illustrates additional details of an exemplary bitcell, which may represent the bitccells of FIG. 4. [Figure 5B] FIG. 5 illustrates additional details of an exemplary bitcell, which may represent the bitccells of FIG. 4. [Figure 6]

[0018] FIG. 13 is an example timing diagram of various signals during CIM array operation. [Figure 7]

[0019] FIG. 1 illustrates an exemplary convolutional layer architecture implemented by a CIM array. [Figure 8A]

[0020] FIG. 2 illustrates a CIM architecture including a CIM array in accordance with some aspects of the present disclosure. [Figure 8B] FIG. 2 illustrates a CIM architecture including a CIM array in accordance with some aspects of the present disclosure. [Figure 9]

[0021] FIG. 8C illustrates an example operation for signal processing via the CIM architecture of FIG. 8B in accordance with certain aspects of the disclosure. [Figure 10]

[0022] FIG. 2 illustrates a CIM array divided into sub-banks to improve processing efficiency, in accordance with some aspects of the present disclosure. [Figure 11]

[0023] FIG. 1 illustrates a CIM array implemented with repeated kernels to improve processing accuracy, in accordance with some aspects of the present disclosure. [Figure 12]

[0024] FIG. 1 is a flow diagram illustrating example operations for signal processing in a neural network in accordance with some aspects of the present disclosure. [Figure 13]

[0025] FIG. 1 illustrates an example electronic device configured to perform operations for signal processing in a neural network in accordance with some aspects of the present disclosure.

[0013]

[0026] For ease of understanding, wherever possible, like reference numbers have been used to designate like elements common to the figures, and it is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014]

[0027] Aspects of the present disclosure provide apparatus, methods, processing systems, and computer-readable media for performing computation in memory (CIM) of machine learning models including depth-wise (DW) separable convolutional neural networks. Some aspects provide a two-phase convolution technique implemented on a CIM array. For example, one of the two phases can include a DW convolution operation using a kernel implemented on the CIM array, and another of the two phases can include a point-wise (PW) convolution operation using a kernel implemented on the CIM array.

[0015]

[0028] For example, some aspects are directed to CIM cells of a CIM array configured for different kernels used for DW convolution, where the kernels are implemented on different rows and the same column of the CIM array. The kernels can be processed using a topological approach, as described herein. The output of the cells implementing the kernels can be coupled to an analog-to-digital converter (ADC). The results of the DW computations can be input to a nonlinear activation circuit for further processing, as described in more detail herein, and may be input back to the same CIM array for point-wise computation. The aspects described herein provide flexibility in configuring any CIM array on demand for DW convolution operations, while increasing the number of kernels that can be implemented on the CIM array, as compared to conventional implementations, as described in more detail herein.

[0016]

[0029] CIM-based machine learning (ML) / artificial intelligence (AI) task accelerators can be used for a wide variety of tasks, including image and audio processing. Furthermore, CIMs can be based on various types of memory architectures, such as dynamic random access memory (DRAM), static random access memory (SRAM) (e.g., based on SRAM cells as in FIG. 5), magnetoresistive random-access memory (MRAM), and resistive random-access memory (ReRAM), and can be attached to various types of processing units, including central processor units (CPUs), digital signal processors (DSPs), graphical processor units (GPUs), field-programmable gate arrays (FPGAs), AI accelerators, and the like. In general, CIMs can advantageously reduce the "memory wall" problem, where moving data in and out of memory consumes more power than computing the data. Thus, significant power savings can be realized by performing computations in memory, which is particularly useful for various types of electronic devices such as low power edge processing devices, mobile devices, etc.

[0017]

[0030] For example, a mobile device may include a memory device configured to store data and in-memory computation operations. The mobile device may be configured to perform ML / AI operations based on data generated by the mobile device, such as image data generated by a camera sensor of the mobile device. Thus, a memory controller unit (MCU) of the mobile device may load weights from another on-board memory (e.g., flash or RAM) into a CIM array of the memory device and allocate input feature buffers and output (e.g., activation) buffers. The processing device may then begin processing the image data, for example, by loading layers in the input buffers and processing the layers with the weights loaded in the CIM array. This process may be repeated for each layer of the image data, and the outputs (e.g., activations) may be stored in an output buffer and then used by the mobile device for ML / AI tasks such as face recognition.

[0018] A brief background on neural networks, deep neural networks, and deep learning

[0031] Neural networks are organized into layers of interconnected nodes. In general, a node (or neuron) is where computations are performed. For example, a node may combine input data with a set of weights (or coefficients) that either amplify or attenuate the input data. Amplification or attenuation of an input signal may thus be viewed as an assignment of relative importance to various inputs with respect to the task the network is trying to learn. In general, input-weight products are added (or accumulated) and then this sum is passed through the node's activation function to determine whether and how far the signal should proceed further through the network.

[0019]

[0032] In its most basic implementation, a neural network may have an input layer, a hidden layer, and an output layer. "Deep" neural networks generally have two or more hidden layers.

[0020]

[0033] Deep learning is a method of training deep neural networks. In general, deep learning is sometimes called a "universal approximator" because it maps inputs to the network to outputs from the network and can therefore learn to approximate an unknown function f(x)=y between any input x and any output y. In other words, deep learning finds the correct f to transform x to y.

[0021]

[0034] More specifically, deep learning trains each layer of nodes on a different set of features, i.e., the output from the previous layer. Thus, with each successive layer of a deep neural network, the features become more complex. Deep learning is therefore powerful because it can progressively extract higher level features from input data by learning to represent the input at successively higher levels of abstraction at each layer, thereby building useful feature representations of the input data, to perform complex tasks such as object recognition.

[0022]

[0035] For example, when presented with visual data, the first layer of a deep neural network can be trained to recognize relatively simple features, such as edges, in the input data. In another example, when presented with auditory data, the first layer of a deep neural network can be trained to recognize the spectral power at a particular frequency in the input data. The second layer of the deep neural network can then be trained to recognize combinations of features, such as simple shapes in the visual data, or combinations of sounds in the auditory data, based on the output of the first layer. The higher layers can then be trained to recognize complex shapes in the visual data or words in the auditory data. The higher layers can then be trained to recognize common visual objects or spoken phrases. Thus, deep learning architectures can perform particularly well when applied to problems with natural hierarchical structures.

[0023] Layer Connectivity in Neural Networks

[0036] Neural networks, such as deep neural networks, can be designed with a variety of connectivity patterns between layers.

[0024]

[0037] 1A illustrates an example of a fully-connected neural network 102, in which a node in a first layer communicates its output to every node in a second layer, such that each node in the second layer receives input from every node in the first layer.

[0025]

[0038] 1B shows an example of a locally connected neural network 104. In the locally connected neural network 104, a node in a first layer may be connected to a limited number of nodes in a second layer. More generally, the locally connected layers of the locally connected neural network 104 may be configured such that each node in a layer has the same or similar connectivity pattern, but with connection strengths (or weights) that can have different values ​​(e.g., 110, 112, 114, and 116). The connectivity patterns of the local connections may result in spatially distinct receptive fields in the upper layers, since the higher layer nodes in a given region may receive inputs that are tuned to the characteristics of a constrained portion of the total inputs to the network through training.

[0026]

[0039] One type of locally connected neural network is a convolutional neural network. Figure 1C shows an example of a convolutional neural network 106. The convolutional neural network 106 can be configured such that the connection strengths associated with the inputs to each node in the second layer are shared (e.g., 108). Convolutional neural networks are well suited to problems where the spatial location of the inputs is meaningful.

[0027]

[0040] One type of convolutional neural network is the deep convolutional network (DCN), which is a network of multiple convolutional layers and can be further configured with, for example, pooling layers and normalization layers.

[0028]

[0041] 1D shows one embodiment of a DCN 100 designed to recognize visual features in an image 126 generated by an image capture device 130. For example, if the image capture device 130 were a vehicle-mounted camera, the DCN 100 could be trained using various supervised learning techniques to identify traffic signs, and even numbers on traffic signs. Similarly, the DCN 100 could be trained for other tasks, such as identifying lane markings, or identifying traffic signals. These are just a few example tasks, and many other tasks are possible.

[0029]

[0042] In this example, the DCN 100 includes a feature extraction section and a classification section. Upon receiving an image 126, the convolutional layer 132 applies a convolution kernel to the image 126 (e.g., as shown and described in FIG. 2) to generate a first set of feature maps (or intermediate activations) 118. In general, a "kernel" or "filter" includes a multi-dimensional array of weights designed to emphasize different aspects of the input data channels. In various examples, "kernel" and "filter" can be used interchangeably to refer to a set of weights applied in a convolutional neural network.

[0030]

[0043] The first set of feature maps 118 may then be subsampled by a pooling layer (e.g., a max pooling layer, not shown) to generate a second set of feature maps 120. The pooling layer can reduce the size of the first set of feature maps 118 while retaining much of the information to improve model performance. For example, the second set of feature maps 120 can be downsampled from 28×28 to 14×14 by the pooling layer.

[0031]

[0044] This process can be repeated through many layers, in other words, the second set of feature maps 120 may be further convolved through one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0032]

[0045] 1D , the second set of feature maps 120 is provided to a fully connected layer 124, which then generates an output feature vector 128. Each feature in the output feature vector 128 may include a number corresponding to a possible feature of the image 126, such as "sign," "60," and "100." In some cases, a softmax function (not shown) may convert the numbers in the output feature vector 128 into probabilities. The output 122 of the DCN 100 is then the probability that the image 126 contains one or more features.

[0033]

[0046] Prior to training the DCN 100, the output 122 produced by the DCN 100 may be inaccurate. Thus, an error may be calculated between the output 122 and a target output known a priori. For example, here the target output is an indication that the image 126 contains a "sign" and the number "60." Then, utilizing the known target output, the weights of the DCN 100 may be adjusted through training such that subsequent outputs 122 of the DCN 100 achieve the target output.

[0034]

[0047] To adjust the weights of the DCN 100, the learning algorithm may calculate a gradient vector for the weights. The gradient may indicate the amount by which the error would increase or decrease if the weights were adjusted in a particular way. The weights may then be adjusted to reduce the error. This method of adjusting the weights is sometimes called "backpropagation" because it involves a "backward pass" through the layers of the DCN 100.

[0035]

[0048] In practice, the error gradient of the weights may be calculated over a small number of examples such that the calculated gradient approximates the true error gradient. This approximation method is sometimes called stochastic gradient descent. Stochastic gradient descent may be iterated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level.

[0036]

[0049] After training, the DCN 100 may be presented with new images and the DCN 100 can generate inferences such as classifications or probabilities that various features are present in the new image.

[0037] Convolution Techniques for Convolutional Neural Networks

[0050] Convolution is commonly used to extract useful features from an input dataset. For example, in convolutional neural networks as described above, convolution allows the extraction of different features using kernels and / or filters whose weights are automatically learned during training. The extracted features are then combined to make inferences.

[0038]

[0051] Activation functions can be applied before and / or after each layer of a convolutional neural network. An activation function is generally a mathematical function (e.g., a formula) that determines the output of a node of a neural network. Thus, an activation function determines whether a node should pass on information based on whether the node's input is relevant to the model's prediction. In one embodiment, if y=conv(x) (i.e., y=convolution of x), then both x and y can generally be considered "activations". However, for a particular convolution operation, x may also be referred to as a "pre-activation" or "input activation" since it exists before the particular convolution, and y may be referred to as an output activation or feature map.

[0039]

[0052] 2 shows an example of conventional convolution where an input image of 12 pixels by 12 pixels by 3 channels is convolved using a 5×5×3 convolution kernel 204 and a stride (or step size) of 1. The resulting feature map 206 is 8 pixels by 8 pixels by 1 channel. As can be seen in this example, conventional convolution can vary the dimensionality of the input data compared to the output data (here, from 12×12 to 8×8 pixels) and the channel dimensionality (here, from 3 to 1 channel).

[0040]

[0053] One way to reduce the computational load (e.g., measured in floating point operations per second (FLOPs)) and number parameters associated with neural networks with convolutional layers is to factorize the convolutional layers. For example, a spatially separable convolution as shown in FIG. 2 may be factorized into two components: (1) a depth-wise convolution (e.g., spatial fusion), where each spatial channel is independently convolved with a depth-wise convolution, and (2) a point-wise convolution (e.g., channel fusion), where all spatial channels are linearly combined. An example of a depth-wise separable convolution is shown in FIG. 3A and FIG. 3B. In general, during spatial fusion, the network learns features from the spatial plane, and during channel fusion, the network learns the relationship between these features across channels.

[0041]

[0054] In one embodiment, separable depthwise convolution can be implemented using a 3×3 kernel for spatial fusion and a 1×1 kernel for channel fusion. Specifically, channel fusion can use a 1×1×d kernel that iterates through every single point in the input image at depth d, where the depth of the kernel d generally matches the number of channels in the input image. Channel fusion with pointwise convolution is useful for dimensionality reduction for efficient computation. Applying a 1×1×d kernel and adding an activation layer after the kernel can give the network additional depth, which can increase its performance.

[0042]

[0055] 3A and 3B show an example of a depthwise separable convolution operation.

[0043]

[0056] Specifically, in FIG. 3A, a 12 pixel by 12 pixel by 3 channel input image 302 is convolved with a filter comprising three separate kernels 304A-C, each having a dimensionality of 5 by 5 by 1, to generate an 8 pixel by 8 pixel by 3 channel feature map 306, with each channel generated by a separate kernel in 304A-C.

[0044]

[0057] The feature map 306 is then further convolved using a point-wise convolution operation where the kernel 308 (e.g., kernel) has dimensionality 1×1×3 to generate an 8 pixel×8 pixel×1 channel feature map 310. As shown in this example, the feature map 310 has reduced dimensionality (1 channel vs. 3), which allows for more efficient computations with the feature map 310. In some aspects of the present disclosure, the kernels 304A-C and the kernel 308 may be implemented using the same compute-in-memory (CIM) array, as described in more detail herein.

[0045]

[0058] The results of the depthwise separable convolution in FIGS. 3A and 3B are substantially similar to the conventional convolution in FIG. 2, but the number of computations is significantly reduced, and therefore the depthwise separable convolution offers significant efficiency gains, when the network design permits.

[0046]

[0059] Although not shown in FIG. 3B, multiple (e.g., m) pointwise convolution kernels 308 (e.g., individual components of a filter) can be used to increase the channel dimensionality of the convolution output. Thus, for example, m=256 1×1×3 kernels 308 can be generated, each of which outputs an 8 pixel×8 pixel×1 channel feature map (e.g., 310), which can be stacked to obtain a resulting 8 pixel×8 pixel×256 channel feature map. The resulting increase in channel dimensionality can provide more parameters for training, thereby improving the ability of the convolutional neural network to identify features (e.g., in the input image 302).

[0047] Exemplary Computation in Memory (CIM) Architecture

[0060] 4 illustrates an exemplary compute-in-memory (CIM) array 400 configured to perform machine learning model computations according to aspects of the present disclosure. In this example, the CIM array 400 is configured to simulate MAC operations using mixed analog / digital operations for artificial neural networks. Thus, as used herein, the terms multiplication and addition may refer to such simulated operations. The CIM array 400 may be used to implement aspects of the processing techniques described herein.

[0048]

[0061] In the illustrated embodiment, the CIM array 400 includes precharge word lines (PCWL) 425a, 425b, and 425c (collectively 425), read word lines (RWL) 427a, 427b, and 427c (collectively 427), analog-to-digital converters (ADC) 410a, 410b, and 410c (collectively 410), digital processing unit 413, bit lines 418a, 418b, and 418c (collectively 418), PMOS transistors 411a-111i (collectively 411), NMOS transistors 413a-413i (collectively 413), and capacitors 423a-423i (collectively 423).

[0049]

[0062] Weights associated with neural network layers may be stored in SRAM cells of CIM array 400. In this example, binary weights are shown in SRAM bit cells 405a-405i of CIM array 400. Input activations (e.g., input values ​​which may be input vectors) are provided on PCWLs 425a-c.

[0050]

[0063] The multiplication is performed in each bit cell 405a-405i of the CIM array 400 associated with a bit line, and the accumulation (sum) of all bit cell multiplication results is performed on the same bit line for one column. The multiplication in each bit cell 405a-405i is a form of operation equivalent to an AND operation of the corresponding activation and weight, and the result is stored as a charge on the corresponding capacitor 423. For example, a product of 1, and therefore a charge on the capacitor 423, is generated only if the activation is 1 (here, PMOS is used, so PCWL is 0 for an activation of 1) and the weight is 1.

[0051]

[0064] For example, in the accumulation phase, RWL 427 is switched high to allow any charge on capacitor 423 (based on the corresponding bit cell (weight) and PCWL (activation) values) to accumulate on the corresponding bit line 418. The voltage value of the accumulated charge is then converted to a digital value by ADC 410 (e.g., the output value may be a binary value indicating whether the total charge is greater than a reference voltage). These digital values ​​(outputs) can be provided as inputs to another aspect of the machine learning model, such as the next layer.

[0052]

[0065] When the activations on precharge word lines (PCWLs) 425a, 425b, and 425c are, for example, 1, 0, 1, the sums of bit lines 418a-c correspond to 0+0+1=1, 1+0+0=1, and 1+0+1=2, respectively. The outputs of ADCs 410a, 410b, and 410c are passed to digital processing unit 413 for further processing. For example, if CIM 100 is processing multi-bit weighted values, the digital outputs of ADCs 110 can be summed to generate a final output.

[0053]

[0066] The exemplary 3×3 CIM circuit 400 can be used, for example, to perform an efficient three-channel convolution for a three-element kernel (or filter), with each kernel weight corresponding to an element in each of the three columns, so that for a given three-element receptive field (or input data patch), the output of each of the three channels is computed in parallel.

[0054]

[0067] In particular, although Figure 4 illustrates an embodiment of a CIM using SRAM cells, other memory types may be used, for example, in other embodiments, dynamic random access memory (DRAM), magnetoresistive random access memory (MRAM), and resistive random access memory (ReRAM or RRAM) may be used as well.

[0055]

[0068] FIG. 5A shows additional details of an example bitcell 500.

[0056]

[0069] The embodiment of Figure 5A may be illustrative of or otherwise related to the embodiment of Figure 4. Particularly, bit line 521 is similar to bit line 418a, capacitor 523 is similar to capacitor 423 of Figure 4, read word line 527 is similar to read word line 427a of Figure 4, precharge word line 525 is similar to precharge word line 425a of Figure 4, PMOS transistor 511 is similar to PMOS transistor 411a of Figure 1, and NMOS transistor 513 is similar to NMOS transistor 413 of Figure 1.

[0057]

[0070] The bit cell 500 includes a static random access memory (SRAM) cell 501 (which may represent the SRAM bit cell 405a of FIG. 4), as well as a transistor 511 (e.g., a PMOS transistor) and a transistor 513 (e.g., an NMOS transistor) and a capacitor 523 coupled to ground. Although a PMOS transistor is used for the transistor 511, other transistors (e.g., an NMOS transistor) can be used in place of the PMOS transistor, along with corresponding adjustments (e.g., inversions) of their respective control signals. The same applies to the other transistors described herein. The additional transistors 511 and 513 are included to implement a computational array in memory according to aspects of the present disclosure. In one aspect, the SRAM cell 501 is a conventional six transistor (6T) SRAM cell.

[0058]

[0071] Programming the weights in the bit cells may be performed once for many activations. For example, during operation, SRAM cell 501 receives only one bit of information at nodes 517 and 519 via write word line (WWL) 516. For example, during a write (WWL 216 is high), if write bit line (WBL) 229 is high (e.g., "1"), node 217 is set high and node 219 is set low (e.g., "0"), or if WBL 229 is low, node 217 is set low and node 219 is set high. Conversely, during a write (WWL 216 is high), if write bit bar line (WBBL) 231 is high, node 217 is set low and node 219 is set high, or if WBBL 229 is low, node 217 is set high and node 219 is set low.

[0059]

[0072] The programming of the weights may be followed by an activation input to charge the capacitors according to the corresponding products and a multiplication step. For example, transistor 511 is activated by an activation signal (PCWL signal) via a precharge word line (PCWL) 525 of the in-memory computation array to perform the multiplication step. Transistor 513 is then activated by a signal via another word line (e.g., read word line (RWL) 527) of the in-memory computation array to perform an accumulation of the multiplied value from bit cell 500 with other bit cells of the array, such as described above with respect to FIG.

[0060]

[0073] When node 517 is "0" (e.g., when the stored weight value is "0"), if a low PCWL indicates an activation of "1" at the gate of transistor 511, capacitor 523 is not charged. Thus, no charge is provided to bit line 521. However, when node 517 corresponding to the weight value is "1" and PCWL is set low (e.g., when the activation input is high), it turns on PMOS transistor 511, thereby acting as a short and allowing capacitor 523 to be charged. After capacitor 523 is charged, transistor 511 is turned off, so that charge is stored in capacitor 523. To move charge from capacitor 523 to bit line 521, NMOS transistor 513 is turned on by RWL 527, causing NMOS transistor 513 to act as a short.

[0061]

[0074] Table 1 shows an example of an in-memory computational array operation according to an AND operation setting, such as may be implemented by bit cell 500 of FIG. 5A.

[0062] [Table 1]

[0063]

[0075] The first column (Activation) of Table 1 contains the possible values ​​of the input activation signal.

[0064]

[0076] The second column (PCWL) of Table 1 includes PCWL values ​​that activate transistors designed to implement in-memory computation functions according to aspects of the present disclosure. Since transistor 511 in this example is a PMOS transistor, the PCWL value is the inverse of the activation value. For example, the in-memory computation array includes transistor 511 that is activated by an activation signal (PCWL signal) via precharge word line (PCWL) 525.

[0065]

[0077] The third column (Cell Node) of Table 1 includes weight values ​​stored in the SRAM cell node that correspond to weights in a weight tensor, such as may be used in a convolution operation.

[0066]

[0078] The fourth column (Capacitor Node) of Table 1 shows the resulting product that is stored as a charge on a capacitor. For example, the charge can be stored at the node of capacitor 523 or at the node of one of capacitors 423a-423i. The charge from capacitor 523 is transferred to bit line 521 when transistor 513 is activated. For example, with reference to transistor 511, when the weight at cell node 517 is "1" (e.g., high voltage) and the input activation is "1" (so PCWL is "0"), capacitor 523 is charged (e.g., capacitor node is "1"). For all other combinations, the capacitor node has a value of 0.

[0067]

[0079] FIG. 5B shows additional details of another example bitcell 550.

[0068]

[0080] Bitcell 550 differs from bitcell 500 of FIG. 5A primarily based on the inclusion of an additional precharge wordline 552 coupled to an additional transistor 554 .

[0069]

[0081] Table 2 shows an example of an in-memory computational array operation similar to Table 1, except following the XNOR operation setting, such as may be implemented by bit cell 550 of FIG. 5B.

[0070] [Table 2]

[0071]

[0082] The first column (Activation) of Table 2 contains the possible values ​​of the input activation signal.

[0072]

[0083] The second column (PCWL1) of Table 2 includes PCWL1 values ​​that activate transistors designed to implement in-memory computation functions according to aspects of the present disclosure. Here again, transistor 511 is a PMOS transistor, and the PCWL1 value is the reciprocal of the activation value.

[0073]

[0084] The third column (PCWL2) of Table 2 includes PCWL2 values ​​that activate additional transistors designed to implement in-memory computation functions according to aspects of the present disclosure.

[0074]

[0085] The fourth column (Cell Node) of Table 2 includes weight values ​​stored in the SRAM cell node that correspond to weights in a weight tensor, such as may be used in a convolution operation.

[0075]

[0086] The fifth column (Capacitor Node) of Table 2 shows the resulting product that is stored as a charge on a capacitor, such as capacitor 523.

[0076]

[0087] FIG. 6 illustrates an example timing diagram 600 of various signals during a compute-in-memory (CIM) array operation.

[0077]

[0088] In the illustrated embodiment, the first row of the timing diagram 600 shows a precharge word line PCWL (e.g., 425a in FIG. 4 or 525 in FIG. 5A) going low. In this embodiment, a low PCWL indicates an activation of "1". A PMOS transistor turns on when PCWL is low, thereby allowing a capacitor to charge (if the weight is "1"). The second row shows a read word line RWL (e.g., read word line 427a in FIG. 4 or 527 in FIG. 5A). The third row shows a read bit line RBL (e.g., 418 in FIG. 4 or 521 in FIG. 5A), the fourth row shows an analog-to-digital converter (ADC) read signal, and the fifth row shows a reset signal.

[0078]

[0089] For example, referring to transistor 511 in FIG. 5A, charge from capacitor 523 is gradually transferred to the read bit line RBL when the read word line RWL is high.

[0079]

[0090] The summed charge / current / voltage (e.g., 403 in FIG. 4, or summed charge from bit line 521 in FIG. 5A) is passed to a comparator or ADC (e.g., ADC 411 in FIG. 4) where the summed charge is converted to a digital output (e.g., a digital signal / number). The summation of the charges may occur in the accumulation region of timing diagram 600, and the readout from the ADC may be associated with the ADC readout region of timing diagram 600. After the ADC readout is obtained, a reset signal discharges all of the capacitors (e.g., capacitors 423a-423i) in preparation for processing the next set of activation inputs.

[0080] Example of convolution in memory

[0091] 7 illustrates an exemplary convolutional layer architecture 700 implemented by a compute-in-memory (CIM) array 708. The convolutional layer architecture 700 may be part of a convolutional neural network (e.g., as described above with respect to FIG. 1D ) and may be designed to process multi-dimensional data, such as tensor data.

[0081]

[0092] In the illustrated example, the input 702 to the convolutional layer architecture 700 has dimensions of 38 (height) x 11 (width) x 1 (depth). The output 704 of the convolutional layer has dimensions of 34 x 10 x 64, which includes 64 output channels corresponding to the 64 kernels of the kernel tensor 714 that are applied as part of the convolution process. Furthermore, in this example, each kernel (e.g., the example kernel 712) of the 64 kernels of the kernel tensor 714 has dimensions of 5 x 2 x 1 (collectively, the kernels of the filter tensor 714 are equivalent to one 5 x 2 x 64 kernel).

[0082]

[0093] During the convolution process, each 5×2×1 kernel is convolved with the input 702 to generate one 34×10×1 layer of output 704. During the convolution, the 640 weights of the kernel tensor 714 (5×2×64) can be stored in a computation in memory (CIM) array 708, which in this example includes a column for each kernel (i.e., 64 columns). The activations of each of the 5×2 receptive fields (e.g., receptive field inputs 706) are then input into the CIM array 708 using word lines, e.g., 716, and multiplied by the corresponding weights to generate a 1×1×64 output tensor (e.g., output tensor 710). The output tensor 704 represents the accumulation of the individual 1×1×64 output tensors for all of the receptive fields (e.g., receptive field inputs 706) of the input 702. For simplicity, the in-memory computational array 708 of FIG. 7 shows only a few example lines for the inputs and outputs of the in-memory computational array 708 .

[0083]

[0094] In the illustrated embodiment, the CIM array 708 includes word lines 716, along which the CIM array 708 receives receptive fields (e.g., receptive field input 706), as well as bit lines 718 (corresponding to columns of the CIM array 708). Although not shown, the CIM array 708 may also include precharge word lines (PCWL) and read word lines RWL (as discussed above with respect to Figures 4 and 5).

[0084]

[0095] In this example, the word line 716 is used for the initial weight definition. However, once the initial weight definition is done, the activation input activates a specially designed line in the CIM bit cell to perform the MAC operation. Thus, each intersection of the bit line 718 and the word line 716 represents a kernel weight value, which is multiplied by the input activation on the word line 716 to generate a product. The individual products along each bit line 718 are then summed to generate a corresponding output value in the output tensor 710. The sum value may be a charge, a current, or a voltage. In this example, the dimensions of the output tensor 704 after processing the entire input 702 of the convolution layer are 34×10×64, but only 64 kernel outputs are generated by the CIM array 708 at tme. Thus, the processing of the entire input 702 can be completed in 34×10 or 340 cycles.

[0085] A CIM Architecture for Depthwise Separable Convolution

[0096] While vector-matrix multiplication blocks implemented in memory for CIM architectures can generally perform traditional convolutional neural network processing well, they are not efficient for supporting depthwise separable convolutional neural networks, which are found in many state-of-the-art machine learning architectures.

[0086]

[0097] Conventional solutions to improve efficiency include adding a separate digital MAC block to handle the processing for the depthwise portion of the separable convolution, while the CIM array can handle the pointwise portion of the separable convolution. However, this hybrid approach results in increased data movement, which can offset the memory efficiency advantages of the CIM architecture. Furthermore, the hybrid approach generally involves additional hardware (e.g., digital multiply and accumulate (DMAC) elements), which increases space and power requirements and increases processing latency. Furthermore, the use of DMAC can affect the timing of processing operations, causing model output timing constraints (or other dependencies) to be exceeded. To solve the problem, various compromises may be made, such as reducing the frame rate of the input data, increasing the clock rate of the processing system elements (including the CIM array), and reducing the input feature size.

[0087]

[0098] The CIM architecture described herein improves timing performance of processing operations for depthwise separable convolution. These improvements beneficially result in shorter cycle times for depthwise separable convolution operations and achieve higher total operations per second (TOPS) per watt of processing power, i.e., TOPS / W, compared to conventional architectures that require more hardware (e.g., DMAC) and / or more data movement.

[0088]

[0099] 8A and 8B illustrate a CIM system 800 including a CIM array 802 according to some aspects of the disclosure. As shown in FIG. 8A, the CIM array 802 can be used to implement kernels 806, 808, 809 for DW convolution operations and kernel 890 for PW convolution operations. For example, as described with respect to FIG. 3A and FIG. 3B, kernels 806, 808, 809 can correspond to kernels 304A, 304B, 304C, respectively, and kernel 890 can correspond to kernel 308. The DW convolution operations can be performed sequentially during a first phase (phase 1). For example, kernel 806 can be processed during phase 1-1, kernel 808 can be processed during phase 1-2, and kernel 809 can be processed during phase 1-3. The output of the DW convolution operations for kernels 806, 808, 809 can be used to generate inputs for kernel 890 to perform a PW convolution operation in a second phase. In this manner, both DW and PW convolution operations can be performed using kernels implemented on a single CIM array. The DW kernels can be implemented on the same column of the CIM array, allowing a larger number of DW kernels to be implemented on the CIM array compared to conventional implementations.

[0089]

[0100] As shown in Figure 8B, the CIM system 800 includes a CIM array 802 configured for DW convolutional neural network computations and point-wise (PW)-CNN computations (e.g., CNN 1x1). Kernels for the DW convolutional and PW convolutional operations can be implemented on different groups of columns and activated separately during different phases, as described with respect to Figure 8A. In some aspects, kernels (e.g., 3x3 kernels) can be implemented on the same columns (also referred to herein as bitlines) of the CIM array 802. For example, a 3×3 kernel 806 of 2-bit weights (i.e., nine 2-bit values ​​including a first 2-bit value b01, b11, a second 2-bit value b02, b12, etc.) can be implemented using CIM cells on columns 810, 812 (e.g., one column for each bit width of the weights) and nine rows 814-1, 814-2 through 814-8, and 814-9 (e.g., also referred to herein as word-lines (WL) and collectively referred to as rows 814, one row for each value in the kernel). Another kernel 808 can be implemented on columns 810, 812 and nine rows 820-1 through 820-9 (collectively referred to as rows 820) to implement another 3×3 filter. Thus, kernels 806 and 808 are implemented on different rows but on the same column. As a result, kernels 806 and 808 can operate sequentially. In other words, activating one row of kernels 806, 808 does not affect the other row of kernels 806, 808. However, activating one column of kernels 806, 808 affects the other column of kernels 806, 808. Thus, kernels 806, 808 may operate sequentially. Although only two kernels 806, 808 are shown, in some aspects more than two kernels may be implemented. For example, kernels 806, 808, 809 shown in FIG. 8A may be implemented in CIM array 802.

[0090]

[0101] In some aspects, the input activation buffer of each kernel is filled (e.g., stored) with the corresponding output from the previous layer. Each kernel can be operated sequentially one by one to generate the DW convolution output. The input of an inactive kernel can be filled with 0 (e.g., logic low) such that the read BL (RBL) output of the inactive kernel is 0 (e.g., as supported in a ternary mode bit cell). In this way, the inactive kernel may not affect the output from the active kernel implemented on the column (BL).

[0091]

[0102] In some aspects, a row (e.g., row 814) of kernel 806 may be coupled to activation buffers 830-1, 830-2 through 830-8, and 830-9 (collectively referred to as activation buffers 830), and a row (e.g., row 820) of kernel 808 may be coupled to activation buffers 832-1 through 832-9 (collectively referred to as activation buffers 832). The outputs of kernel 806 (e.g., at columns 810, 812) may be coupled to an analog-to-digital converter (ADC) 840. ADC 840 receives as inputs the signals from columns 810, 812 and generates a digital representation of the signals, taking into account that bits stored in column 812 represent less importance in their respective weights than bits stored in column 810.

[0092]

[0103] The CIM array 802 may also include PW convolution cells 890 on columns 816, 818 for PW convolution calculations, as shown. The outputs of the PW convolution cells 890 (e.g., at columns 816, 818) may be coupled to an ADC 842. For example, each input of the ADC 840 may receive the accumulated charges of row 814 from each of columns 810, 812, and each input of the ADC 842 may receive the accumulated charges from each of columns 816, 818, based on which each of the ADCs 840, 842 generates a digital output signal. For example, the ADC 842 receives signals from columns 816, 818 as inputs and generates a digital representation of the signal, taking into account that the bits stored in column 818 represent less importance in their respective weights than the bits stored in column 816. Although ADCs 840, 842 are shown as receiving signals from two columns to facilitate analog-to-digital conversion for kernels with 2-bit weight parameters, aspects described herein may be implemented for ADCs configured to receive signals from any number of columns (e.g., three columns to perform analog-to-digital conversion for kernels with 3-bit weight parameters). In some aspects, an ADC such as ADC 840 or 842 may be coupled to eight columns. Additionally, in some aspects, accumulation may be distributed across two or more ADCs.

[0093]

[0104] The output of the ADC 840, 842 can be coupled to a nonlinear arithmetic circuit 850 (and buffer) to implement (e.g., in order) one or more nonlinear operations such as rectified linear unit (ReLU) and average pooling (AvePool), to name a few. Nonlinear operations allow for the generation of complex mappings between inputs and outputs, thus allowing for learning and modeling of complex data such as images, videos, audio, and data sets that are nonlinear or have high dimensionality. The output of the nonlinear arithmetic circuit 850 can be coupled to an activation output buffer circuit 860. The activation output buffer circuit 860 can store outputs from the nonlinear arithmetic circuit 850 to be used as PW convolution inputs for PW convolution calculations via the PW convolution cell 890. For example, the output of the activation output buffer circuit 860 can be provided to the activation buffer 830. The corresponding activation inputs stored in the activation buffer 830 can be provided to the PW convolution cell 890 to perform the PW convolution calculations.

[0094]

[0105] Although each of the kernels 806, 808 includes two columns that allow for two-bit weights to be stored in each row of the kernel, the kernels 806, 808 may be implemented using any number of suitable columns, such as one column for one-bit binary weights, or two or more columns for multi-bit weights. For example, each of the kernels 806, 808 may be implemented using three columns to facilitate a three-bit weight parameter being stored in each row of the kernel, or a single column to facilitate a one-bit weight being stored in each row of the kernel. Furthermore, while each of the kernels 806, 808 is implemented using nine rows for a 3×3 kernel for ease of understanding, the kernels 806, 808 may be implemented using any number of rows to implement a suitable kernel size. Furthermore, more than two kernels may be implemented using a subset of the cells of the CIM array. For example, the CIM array 802 may include one or more other kernels, with the kernels of the CIM array 802 being implemented on different rows and the same columns.

[0095]

[0106] The embodiments described herein provide flexibility in configuring any CIM array on demand for DW convolution operations. For example, the number of rows used to implement each of the kernels 806, 808 can be increased to increase the size of each respective kernel (e.g., implement a 5×5 kernel). Furthermore, some embodiments allow an increase in the number of kernels that can be implemented on a CIM array compared to conventional implementations. In other words, some embodiments of the present disclosure reduce the area on the CIM array consumed for DW convolution operations by implementing kernels for DW convolution on the same column. In this way, the number of kernels for DW convolution that can be implemented on a CIM array can be increased compared to conventional implementations. For example, a total of 113 3×3 filters can be implemented on a CIM array with 1024 rows. Thus, the area consumption for implementing DW convolution operations can be reduced compared to conventional implementations that can use DMAC hardware.

[0096]

[0107] 9 illustrates an example operation 900 for signal processing via the CIM architecture 800 of FIG. 8B in accordance with some aspects of the disclosure. A single CIM array may be used for both DW and PW convolution operations. The kernel for the DW convolution is run in two phases on the same CIM array hardware.

[0097]

[0108] During the first phase for the DW convolution, the columns 810, 812 used by the DW convolution kernels are active. The operation 900 can begin with the processing of the DW convolution layer. For example, in block 904, the DW convolution weights can be loaded into the CIM cells for the kernels. That is, in block 904, the DW 3×3 kernel weights can be grouped into rows and written into the CIM cells for the kernels 806, 808 of the CIM array 802 of FIG. 8. That is, the 2-bit kernel weights can be provided to the columns 810, 812, and the pass gate switches of the memory cells (e.g., memory cells b11 and b01 shown in FIG. 8) can be closed to store the 2-bit kernel weights in the memory cells. Filter weights can be stored in each row of each of the kernels 806, 808. The remaining CIM columns can be used to write the PW convolution weights to the PW convolution cells 890. Both the DW convolution weights and the PW convolution weights are updated for each subsequent layer. In some implementations, the CIM array can be divided into tiles that can be configured in a tri-state mode, as described in more detail herein. In some aspects, tiles on the same column as an active kernel can be configured in a tri-state mode. In the tri-state mode, the outputs of the memory cells of the tile can be configured to have a relatively high impedance, effectively eliminating the influence of the cell on the output.

[0098]

[0109] In block 906, the DW convolution activation inputs (e.g., in the activation buffer 830) may be applied sequentially to each group of rows of the kernels 806, 808 to generate a DW convolution output for each kernel. Only one of the kernels 806, 808 may be active at a time. Inactive filter rows may be placed in a tri-state mode of operation.

[0099]

[0110] In block 908, the ADC 840 may convert the convolution output of the kernels 806, 808 (e.g., in columns 810, 812) from the analog domain to the digital domain. Based on the output of the ADC 840 for DW convolution, a nonlinear operation may be performed via the nonlinear operation circuit 850. The output from the nonlinear operation circuit 850 may be applied to the PW convolution input (stored in the activation output buffer circuit 860) to perform the PW convolution operation. In other words, the PW convolution input may be written to the activation buffer 830 and applied to the PW convolution cells 890 on the rows 814, 820 and columns 816, 818.

[0100]

[0111] The operations 900 may continue to phase 2 by processing a PW convolution operation. For example, in block 912, the CIM array may be loaded with kernels for PW convolution. For example, the PW convolution columns (e.g., columns 816, 818) may be enabled and the DW convolution columns (e.g., columns 810, 812) may be disabled. In block 914, the PW convolution may be performed and the output of the PW convolution may be converted to a digital signal via the ADC 842. In block 916, the ADC 842 may convert the output of the PW convolution cell 890 from the analog domain to the digital domain. Based on the output of the ADC 842 for the PW convolution, a nonlinear activation operation (e.g., ReLU) may be performed via the nonlinear operation circuit 850.

[0101] Techniques for reducing power consumption and improving CIM array utilization.

[0112] FIG. 10 illustrates a CIM array 1000 divided into tiles (also called sub-banks) to conserve power, according to some aspects of the disclosure. The CIM array 1000 may have, as one example, 1024 rows and 256 columns. Individual tiles in the rows and columns may be enabled or disabled. For example, a tile may include 128 rows and 23 columns. As one example, tile 1002 (e.g., including multiple tiles, such as tile 1004) may be active for convolution, while the remaining tiles may be disabled. In other words, the remaining tiles may be configured in a tri-state mode.

[0102]

[0113] In some implementations, row and column filler cells can be implemented in the CIM array 1000. Filler circuits (e.g., buffers or switches) can be used to enable or disable tiles of the CIM array to save power. As an example, column filler cells can be implemented using AND gate logic and row filler cells can be implemented using buffers on the write bit lines (WBL) and transmission switches on the read bit lines (RBL). The size and type of the transmission switches can be configured based on linearity specifications.

[0103]

[0114] DW convolution can use relatively small kernel dimensions (3×3, 5×5, ...), and insufficient utilization of the CIM array can affect the output signal to noise ratio (SNR) due to range compression (e.g., the output of the neural network is distributed within a small range due to nonlinear activation). Some aspects of the present disclosure are directed to techniques for improving the SNR, as described in more detail with respect to FIG. 11.

[0104]

[0115] FIG. 11 illustrates a CIM array implemented with an iterated kernel in accordance with some aspects of the present disclosure.

[0105]

[0116] As shown, each of the kernels 806, 808 may be repeated to form a kernel group. For example, kernels 806, 1104, 1106 form kernel group 1102, where each of kernels 806, 1104, and 1106 includes the same weight. Furthermore, multiple kernel groups, such as kernel groups 1102 and 1104, may be implemented on the same column. Because the repeated kernels 806, 1104, 1106 in group 1102 have the same weight, the same activation input may be provided to each of the repeated kernels in the group. Similarly for group 1104.

[0106]

[0117] The repeated kernels can generate the same output signal that is combined in each column (output), resulting in an increase in the dynamic range at the output of the repeated kernel. For example, using three repeated kernels can triple the dynamic range at the output of the repeated kernel provided to the ADC (e.g., ADC840). Increasing the dynamic range at the output of the kernels facilitates analog-to-digital conversion with higher accuracy since a wider range of the ADC can be utilized. In other words, using the full range of the ADC input allows the digital output of the ADC to more accurately identify the analog input of the ADC and improves the signal-to-noise ratio (SNR) of the ADC.

[0107]

[0118] In some aspects, a relatively small tile size can be used for a CIM bank performing DW convolution (e.g., 16 rows and 32 columns) to allow a larger number of CIM cells to be deactivated to save power. For example, a group of three CIM cells (e.g., with multiple tiles) can be designed to perform the inverse bottleneck of a neural network architecture. Inverse bottleneck operations generally refer to operations used to expand input features, followed by DW convolution and reduction of the DW output dimensionality via PW convolution.

[0108]

[0119] As an example, a first CIM cell group (CIM1) can be used for bottleneck operations, a second CIM cell group (CIM2) can be used for DW convolution operations, and a third CIM cell group (CIM3) can be used for bottleneck operations. In some aspects, CIM2 for DW convolution can have a finer tiling configuration (e.g., 16 rows to implement a 3×3 kernel, or 32 rows to implement a 5×5 kernel) to improve CIM array utilization, while CIM1 and CIM3 can have a coarse-grained tiling (e.g., 64 rows or 128 rows) to avoid the impact of filler cells for non-DW convolution operations. In this way, the reusability of the CIM array library can be doubled for DW and non-DW operations.

[0109]

[0120] The average (e.g., approximate) CIM utilization with coarse-grained tiling (e.g., using 64 rows and 32 columns of a CIM array with each tile having 1024 rows) may be 13.8% for a 3×3 kernel and 31.44% for a 5×5 kernel. In other words, only 13.8% of the active memory cells in the CIM array may be utilized for a 3×3 kernel and 31.44% of the active memory cells in the CIM array may be utilized for a 5×5 kernel. Meanwhile, the average CIM utilization with fine-grained tiling (e.g., using 16 rows and 32 columns per tile and with a CIM array having 1024 rows) may be 40.46% for a 3×3 kernel and 47.64% for a 5×5 kernel. The average CIM utilization with fine-grained tiling (e.g., using 32 rows and 32 columns per tile of a CIM array with 1024 rows) can be 24.18% for a 3×3 kernel and 47.64% for a 5×5 kernel. Thus, fine tiling improves CIM array utilization for filters with smaller kernel sizes (e.g., such as those used in many common DW-CNN architectures). Improving the utilization of the CIM array results in a higher percentage of active memory cells being utilized, reducing power loss caused by unused active memory cells.

[0110]

[0121] In general, utilization can be improved by selecting (e.g., during chip design) a tiling size closer to the kernel size. For example, a tile size of 16 can be used for a kernel size of 9. In some embodiments, the tile size can be determined to be a power of 2 (logarithmic scale) larger than the kernel size to improve flexibility for handling different neural network models.

[0111] Exemplary Operations for Performing Neural Network Processing in a CIM Array

[0122] 12 is a flow diagram illustrating example operations 1200 for signal processing in a neural network according to some aspects of the disclosure. The operations 1200 may be performed by a neural network system, which may include a controller, such as the CIM controller 1332 described with respect to FIG. 13, and a CIM system, such as the CIM system 800.

[0112]

[0123] The operations 1200 begin at block 1205 by the neural network system performing a plurality of depth-wise (DW) convolution operations via a plurality of kernels (e.g., kernels 806, 808, 809) implemented using a plurality of CIM cell groups on one or more first columns (e.g., columns 810, 812) of a computation in memory (CIM) array (e.g., CIM array 802). As an example, performing the plurality of DW convolution operations may include loading a first plurality of weight parameters of a first kernel (e.g., kernel 806) of the plurality of kernels into a first set of CIM cells of the plurality of CIM cell groups, the first set of CIM cells including a first plurality of rows (e.g., row 814) of the CIM array, via the one or more first columns, and performing a first DW convolution operation of the plurality of DW convolution operations via the first kernel including applying a first activation input (e.g., via activation buggers 830) to the first plurality of rows. Performing the plurality of DW convolution operations may also include loading a second plurality of weight parameters of a second kernel (e.g., kernel 808) of the plurality of kernels via one or more first columns into a second set of CIM cells of the plurality of CIM cell groups, the second set of CIM cells including one or more first columns and a second plurality of rows (e.g., rows 820) of the CIM array, the first plurality of rows being different from the second plurality of rows, and performing a second DW convolution operation of the plurality of DW convolution operations via the second kernel including applying a second activation input (e.g., via activation buffer 832) to the second plurality of rows. In some aspects, the first set of CIM cells includes a subset of cells of the CIM array and the second set of CIM cells includes another subset of cells of the CIM array.

[0113]

[0124] In block 1210, the neural network system can generate an input signal for a PW convolution operation (e.g., via the ADC 840 and the nonlinear operation circuit 850) based on the output from the plurality of DW convolution operations. In block 1215, the neural network system can perform a PW convolution operation based on the input signal, the PW convolution operation being performed via a kernel implemented using a CIM cell group on one or more second columns of the CIM array. For example, performing the PW convolution operation can include loading a third plurality of weights into a CIM cell group for the kernel on one or more second columns. In some aspects, the neural network system can generate a digital signal by converting the voltages in the one or more first columns from the analog domain to the digital domain after performing the plurality of DW convolution operations. The input signal to the CIM cell group on one or more second columns can be generated based on the digital signal.

[0114]

[0125] In some aspects, the kernel can be iterated to improve CIM array utilization and improve ADC dynamic range, as described herein. For example, the neural network system can load the first plurality of weight parameters via one or more of the first columns into a third set of CIM cells of the plurality of CIM cell groups, the third set of CIM cells including one or more of the first columns and a third plurality of rows of the CIM array, to perform a first DW convolution operation.

[0115] Exemplary Processing System for Performing Phase-Selective Convolution

[0126] 13 illustrates an example electronic device 1300. The electronic device 1300 can be configured to perform the methods described herein, including the operations 1200 described with respect to FIG.

[0116]

[0127] The electronic device 1300 includes a central processing unit (CPU) 1302, which in some aspects may be a multi-core CPU. Instructions executed by the CPU 1302 may be loaded from a program memory associated with the CPU 1302 or may be loaded from a memory 1324, for example.

[0117]

[0128] The electronic device 1300 also includes additional processing blocks tailored to specific functions, such as a graphics processing unit (GPU) 1304, a digital signal processor (DSP) 1306, a neural processing unit (NPU) 1308, a multimedia processing block 1310, and a wireless connectivity processing block 1312. In one implementation, the NPU 1308 is implemented in one or more of the CPU 1302, the GPU 1304, and / or the DSP 1306.

[0118]

[0129] In some aspects, the wireless connection processing block 1312 may include components for, for example, third generation (3G), fourth generation (4G) (e.g., 4G LTE), fifth generation (e.g., 5G or NR), Wi-Fi, Bluetooth, and wireless data transmission standards. The wireless connection processing block 1312 is further coupled to one or more antennas 1314 to facilitate wireless communication.

[0119]

[0130] The electronic device 1300 may also include one or more sensor processors 1316 associated with any type of sensor, one or more image signal processors (ISP) 1318 associated with any type of image sensor, and / or a navigation processor 1320, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.

[0120]

[0131] Electronic device 1300 may also include one or more input and / or output devices 1322, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, a speaker, a microphone, etc. In some aspects, one or more of the processors of electronic device 1300 may be based on the ARM instruction set.

[0121]

[0132] The electronic device 1300 also includes a memory 1324, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1324 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the electronic device 1300 or the CIM controller 1332. For example, the electronic device 1300 may include a CIM circuit 1326 including one or more CIM arrays, such as the CIM array 802 and the CIM array 804, as described herein. The CIM circuit 1326 may be controlled via the CIM controller 1332. For example, in some aspects, the memory 1324 may include code 1324B for convolution (e.g., performing a DW or PW convolution operation by applying activation inputs). The memory 1324 may also include code 1324C for generating input signals. The memory 1324 may also optionally include code 1324A for loading (e.g., loading weight parameters into a CIM cell). As shown, the CIM controller 1332 may include circuitry 1328B for convolution (e.g., performing a DW or PW convolution operation by applying activation inputs). The CIM controller 1332 may also include circuitry 1328C for generating input signals. The CIM controller 1332 may also optionally include circuitry 1328A for loading (e.g., loading weight parameters into the CIM cells). The components shown and other components not shown may be configured to perform various aspects of the methods described herein.

[0122]

[0133] In some aspects, such as when the electronic device 1300 is a server device, various aspects, such as one or more of the multimedia processing block 1310, the wireless connection component 1312, the antenna 1314, the sensor processor 1316, the ISP 1318, or the navigation 1320, may be omitted from the aspects shown in FIG. 13 .

[0123] Example clause

[0134] Aspect 1. An apparatus for signal processing in a neural network, comprising: a first computation in memory (CIM) cell configured as a first kernel for depth-wise (DW) neural network computation, the first set of CIM cells including one or more first columns and a first plurality of rows of a CIM array; a second set of CIM cells configured as a second kernel for neural network computation, the second set of CIM cells including one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows; and a third set of CIM cells of the CIM array configured as a third kernel for point-wise (PW) neural network computation.

[0124]

[0135] Aspect 2. The apparatus of aspect 1, wherein the first set of CIM cells comprises a subset of cells of the CIM array, and the second set of CIM cells comprises another subset of cells of the CIM array.

[0125]

[0136] Embodiment 3. The apparatus of embodiment 2, wherein the third set of CIM cells is a third subset of cells of the CIM array.

[0126]

[0137] Aspect 4. The apparatus of any one of aspects 1 to 3, wherein the third set of CIM cells includes one or more second columns and a first plurality of rows of the CIM array, and the one or more second columns are different from the one or more first columns.

[0127]

[0138] Embodiment 5. The apparatus of any one of embodiments 1 to 4, further comprising an analog-to-digital converter (ADC) coupled to the one or more first columns.

[0128]

[0139] Example 6. The apparatus of example 5, further comprising a non-linear circuit coupled to the output of the ADC.

[0129]

[0140] Example 7. The apparatus of any one of Examples 1 to 6, further comprising a third set of CIM cells configured as a third kernel for neural network calculations, the third set of CIM cells including one or more of the first columns and a third plurality of rows of the CIM array.

[0130]

[0141] Aspect 8. The apparatus of aspect 7, configured to store the same weight parameters in the first set of CIM cells and the third set of CIM cells when performing a neural network calculation.

[0131]

[0142] Aspect 9. The apparatus of any one of aspects 1 to 8, wherein one or more of the first set of CIM cells on each row of the first plurality of rows are configured to store a first weight parameter, and one or more of the second set of CIM cells on each row of the second plurality of rows are configured to store a second weight parameter.

[0132]

[0143] Aspect 10. The apparatus of aspect 9, wherein a quantity of the one or more first columns is associated with a quantity of one or more bits of the first weight parameter.

[0133]

[0144] Aspect 11. A method of signal processing in a neural network, comprising: performing a plurality of depth-wise (DW) convolution operations via a plurality of kernels implemented using a plurality of CIM cell groups on one or more first columns of a computation in memory (CIM) array; generating an input signal for a point-wise (PW) convolution operation based on output from the plurality of DW convolution operations; and performing a PW convolution operation based on the input signal, the PW convolution operation being performed via a kernel implemented using a CIM cell group on one or more second columns of the CIM array.

[0134]

[0145] Aspect 12. The method of aspect 11, wherein performing the plurality of DW convolution operations includes: performing a first DW convolution operation of the plurality of DW convolution operations via a first kernel, the first DW convolution operation including: loading a first plurality of weight parameters of a first kernel of the plurality of kernels via one or more first columns to a first set of CIM cells of the plurality of CIM cell groups, the first set of CIM cells including a first plurality of rows of the CIM array, and applying a first activation input to the first plurality of rows; and performing a second DW convolution operation of the plurality of DW convolution operations via a second kernel, the second DW convolution operation including: loading a second plurality of weight parameters of a second kernel of the plurality of kernels via one or more first columns to a second set of CIM cells of the plurality of CIM cell groups, the second set of CIM cells including one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows;

[0135]

[0146] Aspect 13. The method of aspect 12, wherein the first set of CIM cells comprises a subset of cells of the CIM array, and the second set of CIM cells comprises another subset of cells of the CIM array.

[0136]

[0147] Aspect 14. The method of aspect 13, wherein performing the PW convolution operation includes loading a third plurality of weights into a CIM cell group for a kernel on one or more of the second columns.

[0137]

[0148] Aspect 15. The method of aspect 14, further comprising generating a digital signal by converting the voltages in one or more first columns from the analog domain to the digital domain after performing multiple DW convolution operations, wherein an input signal to a group of CIM cells on one or more second columns is generated based on the digital signal.

[0138]

[0149] Aspect 16. The method of any one of aspects 12 to 15, further comprising loading the first plurality of weight parameters via one or more first columns into a third set of CIM cells of the plurality of CIM cell groups, the third set of CIM cells including one or more first columns and a third plurality of rows of the CIM array, to perform the first DW convolution operation.

[0139]

[0150] Aspect 17. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method of signal processing in a neural network, the method including: performing a plurality of depth-wise (DW) convolution operations via a plurality of kernels implemented using a plurality of CIM cell groups on one or more first columns of a computation in memory (CIM) array; generating an input signal for a point-wise (PW) convolution operation based on output from the plurality of DW convolution operations; and performing a PW convolution operation based on the input signal, the PW convolution operation being performed via a CIM cell group on one or more second columns of the CIM array.

[0140]

[0151] Aspect 18. The non-transitory computer-readable medium of aspect 17, wherein performing the plurality of DW convolution operations comprises: performing a first DW convolution operation of the plurality of DW convolution operations via a first kernel, the first DW convolution operation comprising: loading a first plurality of weight parameters of a first kernel of the plurality of kernels via one or more first columns to a first set of CIM cells of the plurality of CIM cell groups, the first set of CIM cells comprising a first plurality of rows of the CIM array, and applying a first activation input to the first plurality of rows; performing a second DW convolution operation of the plurality of DW convolution operations via a second kernel, the second DW convolution operation comprising: loading a second plurality of weight parameters of a second kernel of the plurality of kernels via one or more first columns to a second set of CIM cells of the plurality of CIM cell groups, the second set of CIM cells comprising one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows;

[0141]

[0152] Aspect 19. The non-transitory computer-readable medium of aspect 18, wherein the first set of CIM cells comprises a subset of cells of the CIM array, and the second set of CIM cells comprises another subset of cells of the CIM array.

[0142]

[0153] Aspect 20. The non-transitory computer-readable medium of aspect 19, wherein performing the PW convolution operation includes loading a third plurality of weights into a CIM cell group for a third kernel on one or more second columns.

[0143]

[0154] Aspect 21. The non-transitory computer-readable medium of aspect 20, wherein the method further includes generating a digital signal by converting the voltages in the one or more first columns from the analog domain to the digital domain after performing multiple DW convolution operations, and an input signal to the CIM cell group on the one or more second columns is generated based on the digital signal.

[0144]

[0155] Aspect 22. The non-transitory computer-readable medium of any one of aspects 18 to 21, wherein the method further includes loading the first plurality of weight parameters via one or more first columns into a third set of CIM cells of the plurality of CIM cell groups, the third set of CIM cells including one or more first columns and a third plurality of rows of the CIM array, to perform a first DW convolution operation.

[0145] Additional Considerations

[0156] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The embodiments described herein are not intended to limit the scope, applicability, or aspects described in the claims. Various modifications of these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of the elements described without departing from the scope of the disclosure. Various embodiments may omit, substitute, or add various procedures or components as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some embodiments may be combined in some other embodiments. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects described herein. Additionally, the scope of the disclosure is intended to encompass such apparatus or methods practiced using other structures, functions, or structures and functions in addition to or other than the various aspects of the disclosure described herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0146]

[0157] As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0147]

[0158] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items and includes single members. By way of example, "at least one of a, b, or c" is intended to encompass a, b, c, ab, ac, bc, and abc, as well as any combination having multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other permutation of a, b, and c).

[0148]

[0159] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, etc. Also, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Also, "determining" can include resolving, selecting, choosing, establishing, etc.

[0149]

[0160] The methods disclosed herein include one or more steps or actions for achieving the method. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Furthermore, various operations of the methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component(s) and / or module(s), including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

[0150]

[0161] The following claims are not limited to the embodiments set forth herein, but are to be accorded the full scope consistent with the language of the claims. In the claims, reference to an element in the singular is not intended to mean "one and only one," but rather "one or more." Unless otherwise specified, the term "several" refers to one or more. No element of a claim is to be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means of," or, in the case of a method claim, unless the element is recited using the phrase "step of." All structural and functional equivalents of the elements of the various embodiments described throughout this disclosure that are known or that later become known to those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be made public, regardless of whether such disclosure is expressly recited in the claims.

Claims

1. 1. An apparatus for signal processing in a neural network, comprising: a first set of computation-in-memory (CIM) cells configured as a first kernel for depth-wise (DW) neural network computation, the first one of the CIM cells comprising one or more first columns and a first plurality of rows of a CIM array; a second set of CIM cells configured as a second kernel for the neural network computation, the second set of CIM cells including the one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows; a third set of CIM cells of the CIM array configured as a third kernel for point-wise (PW) neural network computation; and An apparatus comprising:

2. 2. The apparatus of claim 1, wherein the first set of CIM cells comprises a subset of cells of the CIM array, the second set of CIM cells comprises another subset of cells of the CIM array, and the third set of CIM cells is a third subset of cells of the CIM array.

3. 2. The apparatus of claim 1, wherein the third set of CIM cells includes one or more second columns and the first plurality of rows of the CIM array, the one or more second columns being different from the one or more first columns.

4. an analog-to-digital converter (ADC) coupled to the one or more first columns; a non-linear circuit coupled to an output of the ADC; The apparatus of claim 1 further comprising:

5. 2. The apparatus of claim 1, further comprising a fourth set of CIM cells configured as a fourth kernel for the neural network calculation, the third set of CIM cells comprising the one or more first columns and a third plurality of rows of the CIM array.

6. 6. The apparatus of claim 5, configured such that when performing the neural network calculations, the same weight parameters are stored in the first set of CIM cells and the fourth set of CIM cells.

7. one or more of the first set of CIM cells on each row of the first plurality of rows are configured to store a first weight parameter; one or more of the second set of CIM cells on each row of the second plurality of rows are configured to store a second weight parameter; 2. The apparatus of claim 1.

8. The apparatus of claim 7 , wherein a quantity of the one or more first columns is associated with a quantity of one or more bits of the first weighting parameter.

9. 1. A method for signal processing in a neural network, comprising: performing a plurality of depth-wise (DW) convolution operations via a plurality of kernels implemented using a plurality of groups of compute-in-memory (CIM) cells on one or more first columns of a CIM array; generating an input signal for a point-wise (PW) convolution operation based on outputs from the plurality of DW convolution operations; performing a PW convolution operation based on the input signal, the PW convolution operation being performed via a kernel implemented using a group of CIM cells on one or more second columns of the CIM array; A method comprising:

10. performing the plurality of DW convolution operations, loading a first plurality of weight parameters of a first kernel of the plurality of kernels into a first set of CIM cells of the plurality of CIM cell groups via the one or more first columns, the first set of CIM cells comprising a first plurality of rows of the CIM array; performing a first DW convolution operation of the plurality of DW convolution operations through the first kernel; and performing the first DW convolution operation comprises applying a first activation input to the first plurality of rows; loading a second plurality of weight parameters of a second kernel of the plurality of kernels into a second set of CIM cells of the plurality of CIM cell groups via the one or more first columns, the second set of CIM cells comprising the one or more first columns and a second plurality of rows of the CIM array, the first plurality of rows being different from the second plurality of rows; performing a second DW convolution operation of the plurality of DW convolution operations through the second kernel; and performing the second DW convolution operation comprises applying second activation inputs to the second plurality of rows; Equipped with 10. The method of claim 9.

11. 11. The method of claim 10, wherein the first set of CIM cells comprises a subset of cells of the CIM array, the second set of CIM cells comprises another subset of cells of the CIM array, and the third set of CIM cells comprises a third subset of cells of the CIM array.

12. 12. The method of claim 11, wherein performing the PW convolution operation comprises loading a third plurality of weights into the CIM cell groups for the kernels on the one or more second columns.

13. generating a digital signal by converting the voltages at the one or more first columns from an analog domain to a digital domain after performing the plurality of DW convolution operations; the input signals to the CIM cell groups on the one or more second columns are generated based on the digital signal. The method of claim 12.

14. loading the first plurality of weight parameters into a third set of CIM cells of the plurality of CIM cell groups via the one or more first columns to perform the first DW convolution operation, the third set of CIM cells comprising the one or more first columns and a third plurality of rows of the CIM array; The method of claim 10 further comprising:

15. A non-transitory computer readable medium comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method of signal processing in a neural network according to any one of claims 9 to 14.