Memory circuit and operating method thereof

By designing a memory circuit that can flexibly store data according to the type of neural network layer, the problem of limited performance when handling different layer types is solved, and higher energy efficiency and throughput are achieved.

CN119988013APending Publication Date: 2025-05-13TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071536.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing AI accelerators have limited performance problems when dealing with neural networks of different layer types because they are customized to support fixed data flows and cannot flexibly adapt to the needs of various layer types.

Method used

A memory circuit is designed, including multiple processing elements, each processing element includes multiple storage units, through a data router and a controller, selectively store weighted data elements or input data elements according to identification of the neural network layer type, and optimize storage and computing efficiency.

Benefits of technology

Flexible processing of neural networks of different layer types is realized, the overall energy efficiency and throughput of AI accelerators is improved, and the dependence on additional buffers and read/write operations is reduced, especially when processing attention layers, which show the advantages of low energy, low latency and small area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988013A_ABST
    Figure CN119988013A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a memory circuit and an operating method thereof. A memory circuit includes: a first buffer configured to store a plurality of first data elements; a second buffer configured to store a plurality of second data elements; a controller configured to generate a control signal based on the layer type; an array comprising a plurality of processing elements (PEs), each PE comprising a plurality of memory cells; and a data router configured to receive the control signal and determine whether to store a respective first data element of the plurality of first data elements or a respective second data element of the plurality of second data elements in the storage unit of each of the PEs based on the control signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of electronic circuits, and more particularly, to semiconductor devices and methods of manufacturing the same. Background Art

[0002] Artificial intelligence (AI) or machine learning (ML) is a powerful tool that can be used to simulate human intelligence in machines programmed to think and act like humans. AI can be used in a variety of applications and industries. AI accelerators are hardware devices used to efficiently process AI workloads such as neural networks. One type of AI accelerator includes a systolic array that can perform operations on inputs through multiplication and accumulation operations. Summary of the invention

[0003] One embodiment of the present invention provides a memory circuit, comprising: a first buffer configured to store a plurality of first data elements; a second buffer configured to store a plurality of second data elements; a controller configured to generate a control signal based on a layer type; an array comprising a plurality of processing elements (PEs), each PE comprising a plurality of storage cells; and a data router configured to receive the control signal and determine, based on the control signal, whether to store a corresponding first data element among the plurality of first data elements or a corresponding second data element among the plurality of second data elements in a storage cell of each PE.

[0004] Another embodiment of the present invention provides a memory circuit, comprising: an array including a plurality of processing elements (PEs); wherein each PE includes a plurality of storage cells; and wherein each PE is configured to selectively store (i) a single first data element of a plurality of first data elements in a storage cell of a corresponding storage cell; or (ii) a plurality of second data elements of a plurality of second data elements in respective corresponding storage cells based on a control signal indicating a layer type.

[0005] Yet another embodiment of the present invention provides a method for operating a memory circuit, comprising: identifying a layer type of a neural network for processing a plurality of input data elements and a plurality of weight data elements; in response to the layer type being a first type, storing a single weight data element among the plurality of weight data elements in one of a plurality of storage units of a corresponding processing element; and in response to the layer type being a second type, storing a plurality of input data elements among the plurality of input data elements in the plurality of storage units of the corresponding processing element, respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various aspects of the present invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, the various components are not drawn to scale. In fact, the sizes of the various components may be arbitrarily increased or reduced for clarity of discussion.

[0007] Figure 1 An example neural network is shown in accordance with some embodiments.

[0008] Figure 2 An example block diagram of an input (A) processed in a convolutional mechanism with weights (W) is shown in accordance with some embodiments.

[0009] Figure 3 shows at least a portion of an input A according to some embodiments Figure 2 Schematic diagram of an example convolution between the weights W of .

[0010] Figure 4 An example schematic diagram of an input (X) processed in an attention mechanism according to some embodiments is shown.

[0011] Figure 5 An example block diagram of a compute-in-memory (CIM) circuit is shown in accordance with some embodiments.

[0012] Figure 6 According to some embodiments Figure 5 Example circuit diagram of a data router of the CIM circuit.

[0013] Figure 7 Another example block diagram of an input (A) processed in a convolutional mechanism with weights (W) is shown in accordance with some embodiments.

[0014] Figure 8 shows how data elements are stored in when the layer type is indicated as a regular convolutional layer (or attention layer) according to some embodiments. Figure 5 A schematic diagram of the processing elements of the CIM circuit.

[0015] Fig. 9 shows how data elements are stored in when the layer type is indicated as a depthwise convolutional layer according to some embodiments. Figure 5 A schematic diagram of the processing elements of the CIM circuit.

[0016] Fig.10 According to some embodiments Figure 5 An example circuit diagram of a column-type write circuit for a CIM circuit.

[0017] Fig.11 An example block diagram of portions of an attention mechanism according to some embodiments is shown.

[0018] Fig.12 It is shown that according to some embodiments, Fig.10 How does the column-based write circuit handle Fig.11 The input tensor X, key weight matrix W K And the operation flow of query matrix Q.

[0019] Fig.13 A method for operating according to some embodiments is shown. Figure 5 An example flow chart of a method of a CIM circuit.

[0020] Fig.14 A method for operating according to some embodiments is shown. Figure 5 An example flow chart of another method of a CIM circuit. DETAILED DESCRIPTION

[0021] The present invention provides many different embodiments or examples for realizing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present invention. Of course, these are merely examples and are not intended to limit the present invention. For example, in the following description, forming a first component above or on a second component may include an embodiment in which the first component and the second component are formed in direct contact, and may also include an embodiment in which an additional component may be formed between the first component and the second component so that the first component and the second component may not be in direct contact. In addition, the present invention may repeat reference numerals and / or characters in various examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or configurations discussed.

[0022] Furthermore, for ease of description, spatially relative terms such as "below," "beneath," "lower," "above," "upper," etc. may be used herein to describe the relationship of one element or component to another (or additional) elements or components as illustrated in the figures. The spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein should likewise be interpreted accordingly.

[0023] Artificial intelligence accelerators are a class of specialized hardware designed to accelerate machine learning workloads processed by deep neural networks (DNNs), which are typically neural networks that involve large amounts of memory access and highly parallel but simple computations. AI accelerators can be based on application-specific integrated circuits (ASICs), which include multiple processing elements (PEs) (or processing circuits) arranged in space or time to perform parts of multiply and accumulate (MAC) operations. MAC operations are performed based on input activation states (sometimes called input data elements) and weights (sometimes also called weight data elements), which are then summed together to provide an output activation state. The input activation state and output activation state are often referred to as the input and output of the PE, respectively.

[0024] Typical AI accelerators, known as fixed dataflow accelerators (FDAs), are tailored to support one fixed dataflow, such as output fixed dataflow, input fixed dataflow, or weight fixed dataflow. However, AI workloads include a variety of layer types / shapes that may favor different dataflows, e.g., one dataflow is suitable for one workload, or one layer may not be the best solution for other layers, thereby limiting performance. For example, the various layer types may include regularized convolutional layers, depthwise convolutional layers, attention layers, fully connected layers, etc. In a typical dataflow architecture, one or more convolutional layers may be followed by a fully connected layer that outputs (or flattens) the previous output into a single vector. However, convolutional layer types are generally more efficient for certain dataflows, while fully connected layer types are generally more efficient for different dataflows. Given the diversity of workloads in terms of layer types, one dataflow that is suitable for one workload or one layer may not be the best solution for other workloads or layers, thereby limiting performance.

[0025] The present disclosure provides various embodiments of an AI accelerator implemented as a memory circuit that can adaptively process various layer types. Based on identifying the layer type of the corresponding neural network, the memory circuit can adjust the configuration or operation of its components to optimize the efficiency of component use. For example, the memory circuit may include a memory array having a plurality of processing elements, and each processing element may include a plurality of storage cells. The memory circuit may include a data router so that the storage cell of each processing element selectively stores at least one of the plurality of weight data elements or stores a plurality of input data elements based on the layer type (e.g., a regular convolution layer, an attention layer, or a deep convolution layer) used to process the weight data elements and the input data elements. With this flexibility, the plurality of storage cells of each processing element can be utilized with higher efficiency, which in turn improves the overall energy efficiency and throughput of the disclosed AI accelerator. In addition, the memory circuit may include a column-based write circuit that can simultaneously read out intermediate results from the processing element row by row (or row-wise) and write these intermediate results column-wise (or column-wise) back to the processing element. Therefore, the memory circuit does not require additional buffers and read / write operations to transpose the matrix, which usually makes it very challenging to process the attention layer of the neural network. Despite the disclosed column-wise write circuit, the disclosed memory circuit can even process the attention layer (except for the regular convolution layer and the depthwise convolution layer) with low energy, low latency and small area.

[0026] Figure 1 An example neural network 100 according to various embodiments is shown. As shown, the neural network 100 includes four layers 110, 120, 130 and 140, wherein layers 110 and 140 are respectively referred to as input layers and output layers, and layers 120 to 130 are all referred to as hidden layers. Each layer can contain multiple neurons. In general, the hidden layers of the neural network 100 can be regarded as neuron layers to a large extent, and each neuron layer receives weighted outputs from neurons of other (e.g., previous layer) neuron layers in the mesh interconnection structure between the layers. The connection weight from the output of a specific previous neuron to the input of another subsequent neuron is set according to the influence or effect of the previous neuron on the subsequent neuron (for simplicity, only one neuron 101 and the weight of the input connection are marked). Here, the output value of the previous neuron is multiplied by the connection weight with the subsequent neuron to determine the specific stimulus presented by the previous neuron to the subsequent neuron.

[0027] The total input stimulus of a neuron corresponds to the combined stimulus of all its weighted input connections. According to various implementations, if the total input stimulus of a neuron exceeds a certain threshold, the neuron is triggered to perform some, e.g. linear or nonlinear, mathematical function on its input stimulus. The output of the mathematical function corresponds to the output of the neuron, which is then multiplied by the corresponding weight of the output connection of the neuron to its subsequent neuron. In general, the more connections between neurons, the more neurons per layer and / or the more layers of neurons, the greater the intelligence that the network is able to achieve. Therefore, neural networks used for practical, real-world artificial intelligence applications are characterized by a large number of neurons and a large number of connections between neurons. Therefore, processing information through neural networks involves a large number of computations (not only for the neuron output functions, but also for the weighted connections).

[0028] The processing performed on the input stimulus is based on the layer type (or mechanism). A neural network can have or implement various layer types (or mechanisms), such as fully connected layers, convolutional layers, deconvolutional layers, recurrent layers, attention layers, etc. In general, convolutional layers (or convolutional mechanisms) are the core building blocks of convolutional neural networks. The parameters of a convolutional layer consist of a set of learnable filters (sometimes called kernels or weights), where each filter has a width and height and is usually square in shape. These filters are small (in terms of their spatial dimensions) but extend throughout the depth of the volume. Depending on the configuration of the convolutional layer, there may be other convolutions: regular convolutions (sometimes called regular convolutional layers) and depthwise convolutions (sometimes also called depthwise convolutional layers). The key difference between regular convolutional layers and depthwise convolutional layers is that depthwise convolutions apply convolutions only along one spatial dimension (sometimes called channels), while regular convolutions are applied to all spatial dimensions / channels at each step. The concept of an attention layer (or attention mechanism) is to improve recurrent neural networks (RNNs) to process longer sequences or sentences. The attention mechanism enhances the information content of the input stimulus embedding by including information about the input context. In other words, the attention mechanism enables the model to weigh the importance of different elements in the input stimulus and dynamically adjust their impact on the output.

[0029] Generally speaking, neural networks compute weights to perform calculations on input data (input stimuli or inputs). Machine learning currently relies on the computation of dot products and vector absolute differences, typically computed by performing multiply-accumulate (MAC) operations on parameters, input data, and weights. The computation of large deep neural networks typically involves so many data elements that it is impractical to store them in the processor cache. Therefore, these data elements are typically stored in memory. As a result, machine learning is very computationally intensive in terms of computing and comparing many different data elements. The computation of operations within the processor is several orders of magnitude faster than the transfer of data elements between the processor and main memory resources. Placing all data elements in cache closer to the processor is prohibitively expensive for the vast majority of practical systems due to the size of memory required to store the data elements. As a result, the transfer of data elements becomes the main bottleneck for AI computations. As data sets grow, the time and power / energy used by the computing system to move data elements can end up being several times the time and power used to actually perform the computation.

[0030] In this regard, a computation-in-memory (CIM) circuit has been proposed to perform such MAC operations. The CIM circuit performs on-site data processing within a suitable memory circuit. The CIM circuit suppresses the latency of data / program reading and output result uploading in the corresponding memory (e.g., memory array), thereby solving the memory (or von Neumann) bottleneck of traditional computers. Another key advantage of the CIM circuit is high computational parallelism, thanks to the specific architecture of the memory array, where computations can be performed along multiple current paths simultaneously. The CIM circuit also benefits from the high density of multiple memory arrays with computing devices, which typically have excellent scalability and 3D integration capabilities. As a non-limiting example, CIM circuits for various machine learning applications can perform MAC operations locally within the memory (i.e., without sending data elements to the main processor) to achieve higher throughput dot products of neuron activations and weight matrices, while still providing higher performance and lower energy compared to the main processor's computation.

[0031] Figure 2 shows an example block diagram of an input "A" processed in a convolutional mechanism with weights "W" in accordance with various embodiments, Figure 3 Shows Figure 2 Schematic diagram of an example convolution between at least a portion of the input A and at least a portion of the weight W in . It should be noted that Figure 3 The schematic diagrams are provided as non-limiting examples only for illustrative purposes and are not intended to limit the scope of the present disclosure. For example, within the scope of the present disclosure, the disclosed memory circuits (e.g. Figure 5 ) can also be implemented to handle any of a variety of other convolutional layer types.

[0032] like Figure 2 As shown, the input (A) of a regular convolutional layer is typically arranged as an input tensor A with "P" planes of input data elements (sometimes called neurons or activations). Each plane has dimensions X×Y of input data elements, often called input channels or channels. A regular convolutional layer is associated with one or more trainable weights, filters, or kernels (W). Each filter W includes multiple weight data elements. For example, for a regular convolutional layer, each filter W has dimensions m×n×P. Therefore, the filters W are shared across multiple planes (channels) of the input tensor A. In other words, the filters W are as deep as the input tensor, allowing the channels to mix freely to generate the output. In another example, for a deep convolutional layer, the channels of the input tensor A are separated, and each channel is convolved with its own filter W. Therefore, multiple filters W with corresponding different dimensions (e.g., m×n×1) are typically utilized.

[0033] To produce an output (e.g., by multiplying the input tensor A by one or more filters W), each filter W is convolved with the input tensor A by sliding the filter W over the input tensor A in the X and Y directions in steps "s" and "t", respectively. The sliding step size in a certain direction is often referred to as the stride size in that direction. At each step, the dot product of the input data element and the weight data element is calculated to produce an output data element (which can be referred to as an output neuron). The input data element applied to the weight data element in any step is often referred to as a convolution window (or window) of the input tensor A. Therefore, each filter W produces an output plane or output tensor "B" of output (e.g., a set of two-dimensional output data elements or output neurons, which can be referred to as an activation map or output channel).

[0034] Typically, a convolution operation produces an output tensor B that is smaller in the X and / or Y direction than the input tensor A. For example, Figure 3 shows a 5×5 input tensor A (with one plane or channel) convolved with a 3×3 filter W with a stride of 2 in the X and Y directions, resulting in a 2×2 output tensor B. Specifically, the input tensor A has 5×5 input data elements, such as A 1,1 , A 1,2 , A 1,3 , A 1,4 , A 1,5 , A 2,1 , A 2,2 , A 2,3 , A 2,4 , A 2,5 、A3,1、A 3,2 , A 3,3 , A 3,4 , A 3,5 , A 4,1 , A4,2 , A 4,3 , A 4,4 , A 4,5 , A 5,1 , A 5,2 , A 5,3 , A 5,4 and A 5,5 ; The filter W has 3×3 weight data elements, such as W 1,1 , W 1,2 , W 1,3 , W 2,1 , W 2,2 , W 2,3 , W 3,1 , W 3,2 and W 3,3 When the first weight data element W 1,1 With input data element A 2i-1,2j-1 When aligned, each output data element of the output tensor B,Bi,j (where "i" represents a row in the output tensor and "j" represents a column in the output tensor) is equal to the dot product of the input data element and the weight data element.

[0035] For example, when the weight data element W 1,1 With input data element A 1,1 When aligned (as shown in 301), output data element B 1,1 is equal to the dot product of the input data element and the weight data element. Specifically, the output element B 1,1 Equal to A 1,1 ×W 1,1 +A 1,2 ×W 1,2 +A 1,3 ×W 1,3 +A 2,1 ×W 2,1 +A 2,2 ×W 2,2 +A 2,3 ×W 2,3 +A 3,1 ×W 3,1 +A 3,2 ×W 3,2 +A 3,3 ×W 3.3 Assuming the stride size is 2, the window is then moved in the X direction (e.g., to the right) by a step size of 2 input data elements, so that the weight data element W 1,1 With input data element A 1,3 Align (as shown in 303). Therefore, the output data element B 1,2 Equal to A 1,3 ×W 1,1 +A 1,4 ×W 1,2 +A1,5 ×W 1,3 +A 2,3 ×W 2,1 +A 2,4 ×W 2,2 +A 2,5 ×W 2,3 +A 3,3 ×W 3,1 +A 3,4 ×W 3,2 +A 3,5 ×W 3,3 According to the same principle, the output data element B can be generated by moving the window in the X direction (to the left) and the Y direction (downward). 2,1 , so that the weight data element W 1,1 With input data element A 3,1 align, and can be done by moving the window in the X direction (to the right) so that the weight data element W 1,1 With input data element A 3,3 Aligned to generate output data element B 2,2 .

[0036] As mentioned above, various types of convolutional layers have been implemented in neural networks, such as regularized convolutional layers and depthwise convolutional layers. Figure 3 In the example diagram of , the input tensor A is shown as having a single plane (or channel), but it should be understood that the principles described should apply to both regular convolutional layers and depthwise convolutional layers. For example, when filter W is implemented as a regular convolutional layer, when the input tensor A has multiple planes (or channels), the same filter W is convolved with all channels. In another example, when filter W is implemented as a depthwise convolutional layer, when the input tensor A has multiple planes (or channels), filter W is convolved with only one of the channels.

[0037] In addition to the convolutional layers discussed above, attention layers have been widely used in Transformer-based models (e.g., large language models) for processing longer sequences or sentences. In general, attention mechanisms mimic cognitive attention by emphasizing important parts of the input and downplaying less important parts of the input. The attention mechanism involves queries, values, and keys, where queries mimic the volitional cues in cognitive attention, values ​​(such as intermediate feature representations) mimic the sensory inputs of cognitive attention, and keys mimic the non-volitional cues of the sensory inputs of cognitive attention. The attention mechanism maps a set of query and key-value pairs to the corresponding output, where the query, key, value, and output are all vectors; the output is calculated as the weighted sum of the values, where the weight assigned to each value is calculated by the compatibility function of the query and the corresponding key. In other words, each query attends to all key-value pairs and generates an attention output.

[0038] Figure 4 An example schematic diagram of an input (tensor) “X” processed in an attention mechanism according to various embodiments is shown. Figure 4 The schematic diagram of gives an overview of the self-attention mechanism flow chart. It should be noted that Figure 4 The schematic diagrams are provided as non-limiting examples only for illustrative purposes and are not intended to limit the scope of the present disclosure. For example, within the scope of the present disclosure, the disclosed memory circuits (e.g., Figure 5 ) can also be implemented to handle any of a variety of other attention layer types (e.g., criss-cross attention or multi-head attention).

[0039] As shown in the figure, the attention mechanism usually includes a transformer (or transformer model) for defining three learnable weight matrices for transformation, including the query weight matrix W Q , key weight matrix W K Sum value weight matrix W V In general, these three weight matrices are operable to project the input tensor X into the query, key, and value components of the sequence, respectively. The input tensor X is first projected onto these weight matrices (e.g., by multiplying the input tensor X by each weight matrix) to generate a query matrix Q (Q = X·W Q ), key matrix K (K = X·W K ) and vector matrix V(V=X·W V ). The transformer then calculates the dot product of all keys and the query, i.e. A = Q·K T , where K TDenotes the transposed key matrix K. The matrix A is then normalized or scaled using the softmax operator to obtain an attention score A', sometimes also referred to as an attention weight A'. Thus, the output Z can be generated as A'·V, where each entity in the output Z is the weighted sum of all entities in the input, when the weight is given by the attention score A'. In some embodiments, Figure 4 The converter shown (eg, component W Q , W K , W V , Q, K, V, Q·K T and A) may sometimes be referred to as an attention layer. In some other embodiments, the attention mechanism may include multiple attention layers as shown, and the multiple attention layers are coupled to a fully connected layer that outputs (or flattens) the previous outputs (Z) of the multiple attention layers into a single vector. This attention mechanism allows the transformer to focus on relevant parts of the input tensor X based on the similarity between the query matrix (or vector) Q and the key matrix (or vector) K, thereby enhancing the ability of the corresponding model to effectively capture dependencies and relationships within the data.

[0040] Figure 5 1 shows an example block diagram of a computing-in-memory (CIM) circuit 500 according to various embodiments. It should be understood that Figure 5 The block diagram has been simplified, and thus, the CIM circuit 500 may include any of a variety of other components while remaining within the scope of the present disclosure.

[0041] As shown, the CIM circuit 500 includes an array 510, a first buffer 520, a second buffer 530, a data router 540, a column-based write circuit 550, a controller 560, and an adder peripheral circuit 570. In short, the CIM circuit 500, as part of an AI accelerator, can adaptively configure its components based on the layer type of a neural network for processing multiple input data elements and weight data elements to achieve high efficiency, low power consumption, and low latency.

[0042] The array 510 may include a plurality of processing elements (PEs) 512 arranged in a plurality of columns (C1, C2 ... CY) and a plurality of rows (R1, R2 ... RX). Each PE 512 is located at the intersection of a corresponding column and a corresponding row. Each PE 512 may include at least a number of registers (or storage units), such as M0, M1 ... M2, M3, M4, M5, M6, M7, M8, M9, M10, M11, M12, M13, M14, M15, M16, M17, M18, M19, M20, M21, M22, M23, M24, M25, M26, M27, M28, M29, M30, M31, M32, M33, M44, M45, M46, M47, M48, M49, M50, M51, M Nand a computing component CP (e.g., a multiplier). A storage unit may be a storage space for a memory unit that is configured to transfer data for immediate use by a central processing unit (CPU) or a graphics processing unit (GPU) for data processing. In some embodiments, each PE 512 may include a plurality of such storage units. The storage units M0 to M1 of each PE 512 may be configured to store data in a plurality of memory units. N The memory cells M0 to M1 of each PE 512 may be configured to selectively store a single element of a plurality of weight data elements or multiple elements of a plurality of input data elements, as will be discussed below. N The PEs may be arranged along a single column, while the memory cells are arranged in corresponding rows. Therefore, such PEs are sometimes referred to as multi-row memory cells. The computing component CP may perform operations on the memory cells M0 to M1. N The output and activation of 512 are multiplied. Each PE 512 (or its computing component CP) can be configured to perform a multiplication operation on a corresponding one of a plurality of first data elements (e.g., input activations or input data elements) and a corresponding one of a plurality of second data elements (e.g., weights or weight data elements), and then perform a summation operation to combine one or more products to generate a partial product. Each PE can provide an output (e.g., a partial product) to the adder peripheral circuit 570 for summation operation. In some embodiments, the adder peripheral circuit 570 may include a plurality of adder trees, a plurality of shifters, or other suitable circuits, each of which is configured to perform a summation operation.

[0043] First buffer 520 may include one or more memories (e.g., registers) that may receive and store input activations (or input data elements) of the neural network. First buffer 520 may sometimes be referred to as activation buffer 520. These input data elements may be received as outputs from, for example, different memory circuits (not shown), global buffers (not shown), or different devices. In some embodiments, input data elements from activation buffer 520 may be provided to data router 540 for selective storage in PE 512 based on control signals 561 provided by controller 560, as will be described in further detail below.

[0044] The second buffer 530 may include one or more memories (e.g., registers) that may receive and store weights (or weight data elements) of the neural network. The second buffer 520 may sometimes be referred to as a weight buffer 530. These weight data elements may be received as outputs from, for example, different memory circuits (not shown), global buffers (not shown), or different devices. In some embodiments, the weight data elements from the weight buffer 530 may be provided to the data router 540 to be selectively stored in the PE 512 based on a control signal 561 provided by the controller 560, which will be described in further detail below.

[0045] The data router 540 is operably coupled to the activation buffer 520 and the weight buffer 530, and can select data elements to be stored in the PE 512 based on the control signal 561 provided by the controller 560. For example, the array 510 may also include at least one write port 514 and an input port 516. In various embodiments of the present disclosure, the write port 514 is configured to receive data elements to be programmed into the PE 512; and the input port 516 is configured to receive data elements to be multiplied with the data elements stored in the PE. The control signal 561 may indicate the layer type of the neural network used to process the input data elements and the weight data elements. For example, the layer type may include at least a regular convolution layer, an attention layer, and a depth convolution layer.

[0046] In one aspect, based on a control signal 561 indicating that the data element to be processed is associated with a regular convolutional layer (mechanism) or an attention layer (mechanism), the data router 540 can select the input data elements received from the activation buffer 520 and forward them to the input port 516, and select the weight data elements received from the weight buffer 530 and forward them to the write port 514. Therefore, the weight data elements are stored in the PE 512 while the input data elements are multiplied by the corresponding stored weight data elements, which is sometimes referred to as "weight stationary (WS) data flow". In addition, in some embodiments, each PE 512 can utilize a single storage unit in its storage unit to store a corresponding one of the weight data elements.

[0047] In another aspect, based on the control signal 561 indicating that the data element to be processed is associated with the deep convolutional layer (mechanism), the data router 540 can select the input data elements received from the activation buffer 520 and forward them to the write port 514, and select the weight data elements received from the weight buffer 530 and forward them to the input port 516. Therefore, the input data elements are stored in the PE 512, and the weight data elements are multiplied by the corresponding stored input data elements, which is sometimes referred to as "input stationary (IS) data flow". In addition, in some embodiments, each PE 512 can use its multiple storage units to store the corresponding input data elements respectively.

[0048] The controller 560 may generate a control signal 561 by identifying the layer type of the neural network. In some embodiments, the controller 560 may be communicatively coupled to another component (e.g., a user interface) that indicates the layer type. In addition to generating the control signal 561 for the data router 540 to select which data elements are to be programmed into the PE 512 of the array 510, the controller 560 may also generate another control signal 563 to selectively configure the column write circuit 550.

[0049] For example, when the layer type is identified as including an attention layer, the controller 560 can generate a control signal 563 to switch between a first logic state and a second logic state, wherein the first logic state enables the column-based write circuit 550 to enable a column-based write-back operation and the second logic state enables the column-based write circuit 550 to disable the column-based write-back operation. When the layer type is identified as including a convolutional layer (e.g., a regular or depth-wise convolutional layer), the controller 560 can generate a control signal 563 in the second logic state to enable the column-based write circuit 550 to disable the column-based write-back operation. When the column-based write-back operation is disabled, the column-based write circuit 550 can perform a row-based write operation. Through the column-based write-back operation, the CIM circuit 500 does not need to include additional circuitry to perform a transposition function, which is typically required to process a neural network with an attention layer. This column-based write-back operation will be discussed in more detail below.

[0050] Figure 6 An example circuit diagram of a data router 540 according to various embodiments is shown. The data router 540 is operably coupled between the buffers 520, 530 and the ports 514, 516, and is configured to select data elements to be forwarded to the write port 514 based on a control signal 561. It should be understood that Figure 6 The circuit diagram has been simplified, and thus, data router 540 may include any of a variety of other components while remaining within the scope of the present disclosure.

[0051] As shown, the data router 540 includes a first multiplexer (MUX) 610 and a second multiplexer (MUX) 620. Figure 6 In the illustrative example of , each of the first MUX 610 and the second MUX 620 is implemented as a 2-to-1 MUX controlled by a corresponding control signal. For example, the first MUX 610 is controlled by a control signal 561, and the second MUX 620 is controlled by another control signal 565 that is logically opposite to the control signal 561. In addition, the first MUX 610 has a first input and a second input, which are configured to receive input data elements and weight data elements from the activation buffer 520 and the weight buffer 530, respectively, and the second MUX 620 has a first input and a second input, which are configured to receive input data elements and weight data elements from the activation buffer 520 and the weight buffer 530, respectively. Based on the logic state of the control signal 561, the first MUX 610 can select one of the data elements received through its first and second inputs as an output forwarded to the write port 514. Similarly, based on the logic state of the control signal 565, the second MUX 620 can select one of the data elements received through its first or second input as an output forwarded to the input port 516.

[0052] Since the control signals 561 and 565 are logically opposite to each other, the data router 540 can determine whether to route the input data element or the weight data element to the write port 514 based on the control signal 561. For example, when the control signal 561 is in a first logic state indicating that the layer type is a regular convolution layer or an attention layer, the first MUX 610 can select the data element (e.g., weight data element) received from its second input and forward it to the write port 514. At the same time, the second MUX 620 can select the data element (e.g., input data element) received from its first input and forward it to the input port 516. When the control signal 561 is in a second logic state indicating that the layer type is a depthwise convolution layer, the first MUX 610 can select the data element (e.g., input data element) received from its first input and forward it to the write port 514. At the same time, the second MUX 620 can select the data element (e.g., weight data element) received from its second input and forward it to the input port 516.

[0053] According to various embodiments of the present disclosure, when the layer type is indicated as a regular convolution layer or an attention layer (i.e., the write port 514 receives the weight data element and the input port 516 receives the input data element), the data router 540 may output a single element of the weight data element to a corresponding one of the PEs 512; and when the layer type is indicated as a depthwise convolution layer (i.e., the write port 514 receives the input data element and the input port 516 receives the weight data element), the data router 540 may output a plurality of weight data elements to a corresponding one of the PEs 512. In addition.

[0054] Figure 7 Another example block diagram showing an input tensor "A" processed in a convolutional mechanism using filters / weights "W", where the input tensor A has a dimension of 4×4 and the filter W has a dimension of 3×3 with a stride of 1. For example, in Figure 7 , the input tensor A has input data elements A 1,1 ,A 1,2 ,A 1,3 ,A 1,4 ,A 2,1 ,A 2,2 ,A 2,3 ,A 2,4 ,A 3,1 ,A 3,2 ,A 3,3 ,A 3,4 ,A 4,1 ,A 4,2 ,A 4,3 , and A 4,4 , which are arranged in 4 columns and 4 rows; and the filter W has weight data elements W arranged in 3 columns and 3 rows 1,1 ,W 1,2 ,W 1,3 ,W 2,1 ,W 2,2 ,W 2,3 ,W 3,1 ,W 3,2 and W 3,3 .

[0055] according to Figures 2 to 3 The convolution principle discussed above generates the first convolution window (with a dimension of 3×3) and transforms the weight data element W 1,1 With input data element A 1,1 Align, the remaining weight data elements W 1,2 W 1,3 ,W 2,1 ,W 2,2 ,W 2,3 ,W 3,1 ,W 3,2 and W 3,3 Respectively with the input data element A1,2 ,A 1,3 ,A 2,1 ,A 2,2 ,A 2,3 ,A 3,1 ,A 3,2 and A 3,3 Thus, the corresponding PE (e.g., 512) can be aligned by aligning the input data elements with the corresponding aligned weight data elements (i.e., A 1,1 ×W 1,1 +A 1,2 ×W 1,2 +A 1,3 ×W 1,3 +A 2,1 ×W 2,1 +A 2,2 ×W 2,2 +A 2,3 ×W 2,3 +A 3,1 ×W 3,1 +A 3,2 ×W 3,2 +A 3,3 ×W 3,3 ) to generate the first partial product. Such a first convolution window is represented by 701. Next, the second, third and fourth convolution windows (on the same dimension of 3×3) are generated, and the weight data element W is multiplied by 1. 1,1 With input data element A 1,2 , A 2,1 and A 2,2 Aligned, as shown at 703, 705 and 707 respectively.

[0056] use Figure 7 As an illustrative example, Figure 8 and Fig. 9 Schematic diagrams showing how data elements are stored in PE 512 when the layer type is indicated as a regular convolution layer (or attention layer) and a depth convolution layer according to various embodiments. Specifically, Figure 8 shows an example of generating partial products based on a weighted stationary (WS) data flow, while Fig. 9 An example of generating partial products from an input stationary (IS) data stream is shown.

[0057] exist Figure 8 In the example, the weight data element W 1,1 The weight data element W is routed by the data router 540 to the write port 514 of the array 510 and then stored in the first PE 512 of the PE 512 of the array 510, where each PE 512 may have 4 rows of storage cells M0, M1, M2 and M3. 1,1can be stored in the storage unit M0 of the first PE 512. As the first PE 512 is activated, the data router 540 transfers the corresponding input data element A 1,4 , A 1,3 , A 1,2 , A 1,1 Routed to input port 516. In some embodiments, activation of first PE 512 (eg, input data element A 1,4 , A 1,3 , A 1,2 , A 1,1 ) can be fed into the array 510 row by row. In addition, the input data element A 1,4 , A 1,3 , A 1,2 , A 1,1 are the weight data elements W in windows 701 to 707. 1,1 The data elements to be multiplied (aligned). According to the same principle, the weight data element W 1,2 , W 1,3 , W 2,1 , W 2,2 , W 2,3 、W3,1、W 3,2 and W 3,3 Each is stored in one of the memory cells of a corresponding PE in the second, third, fourth, fifth, sixth, seventh, eighth, and ninth PEs, while the corresponding activations (input data elements) are fed into array 510. In some embodiments, the nine PEs may be arranged along a single column of array 510, while other configurations are contemplated.

[0058] exist Fig. 9 , enter data element A 1,4 , A 1,3 , A 1,2 , A 1,1 The data element A is routed by the data router 540 to the write port 514 of the array 510 and then stored in the first PE 512 of the PEs 512 of the array 510. Specifically, the PE 512 may 1,1 ,A 1,2 ,A 1,3 , and A 1,4 In some embodiments, the input data element A is stored in the memory cells M0, M1, M2 and M3 respectively. 1,4 ,A 1,3 ,A 1,2 ,A 1,1 can be fed into the array 510 by column. In addition, the input data element A 1,4 ,A 1,3 ,A 1,2 ,A 1,1are the weight data elements W in windows 701 to 707. 1,1 As the first PE 512 is activated, the data router 540 transfers the corresponding weight data element W 1,1 Routed to input port 516. Based on the same principle, each set of input data elements (A 1,2 ,A 1,3 ,A 2,2 and A 2,3 ),(A 1,3 ,A 1,4 ,A 2,3 and A 2,4 ),(A 2,1 ,A 2,2 ,A 3,1 and A 3,2 ),(A 2,2 ,A 2,3 ,A 3,2 and A 3,3 ),(A 2,3 ,A 2,4 ,A 3,3 and A 3,4 ),(A 3,1 ,A 3,2 ,A 4,1 and A 4,2 ),(A 3,2 ,A 3,3 ,A 4,2 and A 4,3 ),(A 3,3 ,A 3,4 ,A 4,3 and A 4,4 ), corresponding to the weight data elements W 1,2 , W 1,3 ,W 2,1 ,W 2,2 ,W 2,3 ,W 3,1 ,W 3,2 and W 3,3 , each stored in a memory cell of a corresponding PE in the second, third, fourth, fifth, sixth, seventh, eighth, and ninth PEs, and the corresponding activations (weight data elements) are fed into array 510. In some embodiments, the nine PEs may be arranged along a single column of array 510, while other configurations are contemplated.

[0059] Fig.10 An example circuit diagram of a column-based write circuit 550 according to various embodiments is shown. The column-based write circuit 550 is operably coupled between the data router 540 and the write port 514 and is configured to selectively perform a column-based write-back operation based on a control signal 563. It should be understood that Fig.10 The circuit diagram of FIG. 5 has been simplified, and thus, the column write circuit 550 may include any of a variety of other components while remaining within the scope of the present disclosure.

[0060] As shown in the figure, the column write circuit 550 writes to the input port 514 ( Fig.10 5 (not shown) is coupled to the array 510. The array 510 is shown as having "Y" columns and "X" rows of PE 512, where X and Y can each be an integer equal to or greater than 2. In some embodiments, the column-based write circuit 550 can include a plurality of MUXs 1012, 1014, 1016, etc., respectively coupled to the rows of the array 510, and a plurality of AND gates 1022, 1024, 1026, etc., respectively coupled to the columns of the array 510. The rows of the array 510 can be coupled to MUX 1018, which is configured to select one of the rows based on a "ROW_SEL" signal, and the columns of the array 510 can be coupled to MUX 1028, which is configured to select one of the columns based on "COL_SEL" information.

[0061] To selectively enable the above-mentioned column-based write-back operation, MUXs 1012 to 1016 are all controlled by control signal 563. Control signal 563 is sometimes referred to as a "COL_EN" signal, which indicates in some way whether the layer type of the corresponding neural network includes an attention layer that requires a transposition function, etc. For example, when processing an attention-related mechanism, control signal 563 can be converted between a first logic state and a second logic state to selectively enable the column-based write-back operation. In another example, when the layer type of the neural network is not associated with any attention-related mechanism or does not require a transposition function, control signal 563 can be maintained at a constant logic state to disable the column-based write-back operation.

[0062] Specifically, each of the MUXs 1012 to 1016 may have a first input, a second input, and an output. Each of the first and second inputs is configured to be connected to the output terminal through the write port 514 ( Figure 5 ) receives a plurality of data elements. That is, the data elements received through the first or second input are configured to be programmed into the array 510 (or PE 512). In various embodiments, the first input is configured to receive a plurality of data elements (e.g., Figure 4 The key weight matrix W shown in K The second input is configured to also receive the weight data element of the corresponding activation (e.g., Figure 4When the control signal 563 indicates that the column-based write-back operation is enabled, the MUXs 1012 to 1016 can each select the data element received from the second input and forward it to its output; when the control signal 563 indicates that the column-based write-back operation is disabled, the MUXs 1012 to 1016 can each select the data element received from the first input and forward it to its output.

[0063] Fig.11 An example block diagram of a portion of an attention mechanism is shown, according to various embodiments, where a key weight matrix “W K "Process the input tensor "X" to generate the matrix "K", and use the transposed matrix K T The query matrix "Q" is processed to generate a pre-normalized weight score matrix "A". It should be understood that for the purpose of illustration, Fig.11 The instance of is simplified, and the dimension of each matrix can be equal to any other value.

[0064] exist Fig.11 In the example, the input tensor X, key weight matrix W K The dimensions of the query matrix Q are both 2 × 2. The input tensor X has input data elements arranged in 2 columns and 2 rows. 1,1 , X 1,2 , X 2,1 and X 2,2 ; Key weight matrix W K With weight data elements W arranged in 2 columns and 2 rows K1,1 , W K1,2 , W K2,1 and W K2,2 ; The query matrix Q has weight data elements Q arranged in 2 columns and 2 rows 1,1 , Q 1,2 , Q 2,1 and Q 2,2 Based on the above attention mechanism (such as Figure 4 ), the matrix K consists of K arranged in 2 columns and 2 rows 1,1 , K 1,2 , K 2,1 and K 2,2 Composed by combining the input tensor X with the key weight matrix W K Multiply (K = X·W K ) is generated. Next, by combining the query matrix Q with the transposed matrix K T Multiply (A = Q·K T ), generates A arranged in 2 columns and 2 rows 1,1 , A 1,2 , A2,1 and A 2,2 The matrix A is composed of.

[0065] Fig.12 According to various embodiments, how the column-based write circuit 550 processes Fig.11 The input tensor X, key weight matrix W shown in K and query matrix Q to read the intermediate results from array 510 and perform the column-based write-back operation. Fig.12 , array 510 is shown as having four PEs, 512A, 512B, 512C, and 512D, arranged in a 2×2 array, and each of PEs 512A to 512D includes two storage cells.

[0066] First, the key weight matrix W K The data element W K1,1 ,W K1,2 ,W K2,1 ,W K2,2 are programmed into four PEs 512 respectively. Specifically, the key weight matrix W K The first line (W K1,1 and W K1,2 ) are written into the first storage unit of PE 512A and 512B of the first row respectively, and the key weight matrix W K The second line (W K2,1 and W K2,2 ) are written into the first storage unit of PE 512C and 512D of the second row respectively. In other words, the key weight matrix W K The data elements of are written row by row into the array 510. The first row (X) of the input tensor X that can be received through the input port 516 1,1 and X 1,2 ) and the first column (W) of the data element stored in the first storage unit K1,1 and W K2,1 ) and multiplied by the second column (W stored in the first memory cell K1,2 and W K2,2 ) to generate intermediate results respectively. For example, K 1,1 =X 1,1 ×W K1,1 +X 1,2 ×W K2,1 , and K 1,2 =X 1,1 ×W K1,2 +X 1,2 ×W K2,2 .

[0067] When generating or reading data element K 1,1 and K 1,2At the same time, the column-based write circuit 550 can write these data elements (intermediate results) back to the array 510 by column, as shown by arrow 1201. For example, data element K 1,1 The second storage unit of the first PE (e.g., 512A) in the first column of PEs is written, data element K 1,2 is written to the second memory cell of the second PE (e.g., 512C) in the same first column of PEs. Next, the second row (X) of the input tensor X received via input port 516 may be 2,1 and X 2,2 ) and the first column data element (W) stored in the first storage unit K1,1 and W K2,1 ) and multiplied by the second column data element (W) stored in the first storage unit K1,2 and W K2,2 ) to generate intermediate results, such as data element K 2,1 and K 2,2 Similarly, the column write circuit 550 can write these data elements K 2,1 and K 2,2 The data element K is written back to the array 510 in a column-wise manner, as indicated by arrow 1203. For example, the data element K 2,1 is written into the second storage unit of the first PE (e.g., 512B) in the second column of PEs, and the data element K 2,2 Written into the second storage unit of the second PE (eg, 512D) in the same second column of PEs.

[0068] With data element K 1,1 , K 1,2 , K 2,1 and K 2,2 In the column-wise write-back array 510, data element K 1,1 , K 1,2 , K 2,1 and K 2,2 is equivalently transposed in array 510. For example, data element K 1,2 Already from ( Fig.11 The first position at the intersection of the first row and the second column of the key matrix K in the key matrix K is changed to the second position at the intersection of the second row and the first column of the matrix (formed by PEs 512A to 512D), corresponding to the data element K. 2,1 Thus, the first row (Q 1,1 and Q 1,2 ) can be combined with the first column data element (K 1,1 and K 1,2 ) and multiplied by the second column data element (K) stored in the second storage unit 2,1 and K 2,2) are multiplied to generate data elements A respectively 1,1 and A 1,2 For example, A 1,1 =Q 1,1 ×K 1,1 +Q 1,2 ×K 1,2 and A 1,2 =Q 1,1 ×K 2,1 +Q 1,2 ×K 2,2 Similarly, data element A 2,1 and A 2,2 This can be done by replacing the second row of the query matrix Q ( Q2,1 and Q 2,2 ) are respectively related to the data element (K 1,1 and K 1,2 ) and data elements (K 2,1 and K 2,2 ) are multiplied together to generate.

[0069] Fig.13 1 is a flowchart of an example method 1300 for operating a CIM circuit according to various embodiments of the present disclosure. The operations of the method 1300 may be performed by the above components (e.g., Figure 5 and Figure 6 ) is performed, and therefore, some of the reference numbers used above may be reused in the following discussion of method 1300. For example, method 1300 is mainly directed to operations performed by controller 560 and data router 540. It should be understood that method 1300 has been simplified, and therefore, may be described in detail in detail. Fig.13 Additional operations are provided before, during, and after method 1300, and some other operations may only be briefly described herein.

[0070] The method 1300 begins with an operation 1310 of identifying a layer type of a neural network for processing a plurality of input data elements and a plurality of weight data elements. For example, the controller 560 may identify such a layer type and provide a control signal 561 to the data router 540. As a non-limiting example, the control signal 561 may be provided in a first logic state when the layer type (first type) is a regular convolutional layer or an attention layer, and may be provided in a second logic state when the layer type (second type) is a deep convolutional layer. It should be noted that any of the first or second logic states may indicate any of a variety of other layer types of a neural network while still within the scope of the present disclosure.

[0071] The method 1300 proceeds to operation 1320, in response to identifying the first type, storing a single element of the weight data element in a storage unit of the corresponding processing element (PE). Continuing with the same example, in response to the control signal 561 provided at the first logic state (e.g., the regular convolution layer or the attention layer), the data router 540 (or Figure 6 The first MUX 610 in the non-limiting implementation of can select the weight data element received from the weight buffer 530 to be forwarded to the write port 514 of the array 510. At the same time, the data router 540 (or its second MUX 620) can select the input data element received from the activation buffer 520 to be forwarded to the input port 516 of the array 510. The weight data elements received through the write port 514 can be programmed into different PEs respectively. Specifically, each PE having multiple storage units can store the corresponding weight data element in one of its multiple storage units. In some embodiments, the stored weight data elements can be multiplied with a subset of the input data elements (e.g., multiple input data elements), which can be fed into the array 510 row by row.

[0072] Method 1300 proceeds to operation 1330, in response to identifying the second type, storing the plurality of input data elements in a plurality of storage units of a corresponding processing element (PE). Continuing with the same example, in response to providing control signal 561 at a second logic state (e.g., a deep convolutional layer), data router 540 (or Figure 6 The first MUX 610 in the non-limiting implementation of the data router 540 (or its second MUX 620) can select the input data elements received from the activation buffer 520 to be forwarded to the write port 514 of the array 510. At the same time, the data router 540 (or its second MUX 620) can select the weight data elements received from the weight buffer 530 to be forwarded to the input port 516 of the array 510. The input data elements received through the write port 514 can be programmed into different PEs respectively. Specifically, each PE having multiple storage units can store corresponding input data elements in its multiple storage units respectively. In some embodiments, the stored input data elements can be multiplied with a subset of the weight data elements (e.g., a single element in the weight data elements), which can be fed into the array 510 row by row.

[0073] Fig.14 1400 is a flowchart of an example method 1400 for operating a CIM circuit according to various embodiments of the present disclosure. The operations of the method 1400 may be performed by the above components (e.g. Figure 5 and Fig.10), and therefore, some of the reference numbers used above may be reused in the following discussion of method 1400. For example, method 1400 is mainly directed to operations performed by controller 560 and column write circuit 550. It should be understood that method 1400 has been simplified, and therefore, may be described in detail in detail. Fig.14 Additional operations are provided before, during, and after method 1400, and some other operations may only be briefly described herein. For example, method 1400 can be selectively performed after method 1300.

[0074] Method 1400 begins with operation 1410, identifying a layer type including an attention layer of a neural network for processing a plurality of input data elements and a plurality of weight data elements. In some embodiments, operation 1410 may be equivalent to or a portion of operation 1310 of method 1300. For example, controller 560 may identify the attention layer and provide control signals 561 and 563 to data router 540 and column write circuit 550, respectively. When the layer type includes an attention layer, control signal 563 may be provided to switch between a first logic state and a second logic state, and when the layer type does not include an attention layer, control signal 563 may be provided as fixed in a second logic state.

[0075] Method 1400 continues with operation 1420 where the intermediate results are read out row by row from the memory array. Fig.12 As a representative example, these intermediate results correspond to the first row of the input tensor X (e.g., X 1,1 and X 1,2 ) and the key weight matrix W K (For example, W K1,1 , W K1,2 , W K2,1 and W K2,2 ) by multiplying the matrix K (for example, K 1,1 and K 1,2 ). For example, before reading out the data elements of the matrix K, the column write circuit 550 may respond to the control signal 563 provided in the second logic state to write the key weight matrix (W K1,1 , W K1,2 , W K2,1 and W K2,2 ) are written row by row into the memory array 510. The key weight matrix W K (W K1,1 , W K1,2 )'s first row (W K1,1 ,W K1,2 ) can be stored in the corresponding first storage unit of the first row PE (512A and 512B), the key weight matrix W K The second line (W K2,1 , WK2,2 ) can be stored in the corresponding first storage unit of the second row PE (512C and 512D). Next, the data element K 1,1 and K 1,2 PE 512A to 512D can be based on K 1,1 =X 1,1 ×W K1,1 +X 1,2 ×W K2,1 and K 1,2 =X 1,1 ×W K1,2 +X 1,2 ×W K2,2 is generated and received by the column write circuit 550 through the adder peripheral circuit 570.

[0076] Method 1400 continues to operation 1430 where the intermediate results are written back to the memory array column by column. Fig.12 In the same example of FIG. 1 , the column-based write circuit 550 receives the data element K 1,1 and K 1,2 At the same time, the column write circuit 5.5 can change to the first logic state in response to the control signal 563 to write the data element back to the array 510 by column. 1,1 and K 1,2 can be written back to the corresponding second storage unit of the first column PE (512A and 512C). In various embodiments, operations 1420 and 1430 can be performed one or more times so that the second row of the input tensor X (e.g., X 2,1 and X 2,2 ) and the key weight matrix W K Multiply, as the intermediate result (K 2,1 and K 2,2 ), these intermediate results are written back to the corresponding second storage units of the second column PE (512B and 512D).

[0077] In one aspect of the present disclosure, a memory circuit is disclosed. The memory circuit includes a first buffer configured to store a plurality of first data elements; a second buffer configured to store a plurality of second data elements; a controller configured to generate a control signal based on a layer type; an array including a plurality of processing elements (PEs), each PE including a plurality of storage cells; and a data router configured to receive the control signal and determine whether to store a corresponding first data element of the plurality of first data elements or a corresponding second data element of the plurality of second data elements in a storage cell of each PE based on the control signal.

[0078] In some embodiments, the first data element comprises a weight data element and the second data element comprises an input data element.

[0079] In some embodiments, the data router includes: a first multiplexer having a first input connected to the second buffer, a second input connected to the first buffer, and a first output connected to an input port of the array; and a second multiplexer having a third input connected to the second buffer, a fourth input connected to the first buffer, and a second output connected to a write port of the array.

[0080] In some embodiments, the first multiplexer is configured to select a data element received from one of the first input and the second input based on a logically inverted version of the control signal, and the second multiplexer is configured to select a data element received from one of the third input and the fourth input based on the control signal.

[0081] In some embodiments, when the control signal indicates that the layer type is a regular convolutional layer or an attention layer, the first multiplexer is configured to output the second data element to the input port, and the second multiplexer is set to output the first data element to the write port.

[0082] In some embodiments, in response to the indication of the control signal, each PE is configured to store a single corresponding data element of the plurality of the first data elements.

[0083] In some embodiments, when the control signal indicates that the layer type is a depthwise convolutional layer, the first multiplexer is configured to output the first data element to the input port, and the second multiplexer is configured to output the second data element to the write port.

[0084] In some embodiments, in response to the indication of the control signal, each PE is configured to store a corresponding plurality of second data elements of the plurality of second data elements.

[0085] In some embodiments, the number of the plurality of second data elements stored in each PE is determined based on at least one of: an arrangement of the first data elements corresponding to a window size; an arrangement of the second data elements; and a stride size.

[0086] In some embodiments, a plurality of said second data elements are stored along a single column of storage cells in each PE.

[0087] In some embodiments, the PEs are each configured to perform at least one multiplication operation on one or more of the plurality of first data elements and one or more of the plurality of second data elements. In another aspect of the present disclosure, a memory circuit is disclosed. The memory circuit includes an array including a plurality of processing elements (PEs). Each PE includes a plurality of storage cells. Each PE is configured to selectively store (i) a single first data element of the plurality of first data elements in one of the corresponding storage cells based on a control signal indicating a layer type; or (ii) store a plurality of second data elements of the plurality of second data elements in the corresponding storage cells, respectively.

[0088] In some embodiments, the first data element comprises a weight data element and the second data element comprises an input data element.

[0089] In some embodiments, the memory circuit further comprises: a data router configured to receive the control signal and determine whether to store a single first data element among the plurality of first data elements or to store a plurality of second data elements among the plurality of second data elements based on the control signal.

[0090] In some embodiments, the data router includes: a first multiplexer having a first input configured to receive at least one of the second data elements, a second input configured to receive at least one of the first data elements, and a first output connected to an input port of the array; and a second multiplexer having a third input configured to receive at least one of the second data elements, a fourth input configured to receive at least one of the first data elements, and a second output connected to a write port of the array.

[0091] In some embodiments, when the control signal indicates that the layer type is a regular convolutional layer or an attention layer, the first multiplexer is configured to output the at least one second data element received through the first input to the input port, and the second multiplexer is configured to output the at least one first data element received through the fourth input to the write port.

[0092] In some embodiments, when the control signal indicates that the layer type is a depthwise convolutional layer, the first multiplexer is configured to output the at least one first data element received through the second input to the input port, and the second multiplexer is configured to output the at least one second data element received through the third input to the write port.

[0093] In some embodiments, the PEs are each configured to perform at least a multiplication operation on one or more of the plurality of the first data elements and one or more of the plurality of the second data elements.

[0094] In another aspect of the present disclosure, a method for operating an in-memory computing circuit is disclosed. The method includes identifying a layer type of a neural network for processing a plurality of input data elements and a plurality of weight data elements. The method includes, in response to the layer type being a first type, storing a single weight data element of the plurality of weight data elements in one of a plurality of storage units of a corresponding processing element. The method includes, in response to the layer type being a second type, storing a plurality of input data elements of the plurality of input data elements in a plurality of storage units of a corresponding processing element, respectively.

[0095] As used herein, the terms "approximately" and "approximately" generally refer to a value of a given quantity that may vary depending on a particular technology node associated with the subject semiconductor device. Based on the particular technology node, the term "approximately" may refer to a value of a given quantity that varies, for example, within a range of 10% to 30% of a value (e.g., +10%, ±20%, or ±30% of a value).

[0096] The foregoing summarizes the features of several embodiments so that those skilled in the art can better understand aspects of the present disclosure. Those skilled in the art should understand that they can easily use the present disclosure as a basis for designing or modifying other processes and structures to achieve the same purposes and / or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they can be subjected to various changes, substitutions and modifications without departing from the spirit and scope of the present disclosure.

Claims

1. A memory circuit, comprising: a first buffer configured to store a plurality of first data elements; a second buffer configured to store a plurality of second data elements; a controller configured to generate a control signal based on the layer type; an array including a plurality of processing elements (PEs), each processing element including a plurality of storage units; as well as A data router is configured to receive the control signal and determine whether to store a corresponding first data element of the plurality of first data elements or a corresponding second data element of the plurality of second data elements in a storage unit of each processing element based on the control signal.

2. The memory circuit according to claim 1, wherein: The first data element comprises a weight data element and the second data element comprises an input data element.

3. The memory circuit according to claim 1, wherein: The data router comprises: a first multiplexer having a first input connected to the second buffer, a second input connected to the first buffer, and a first output connected to an input port of the array; and A second multiplexer has a third input connected to the second buffer, a fourth input connected to the first buffer, and a second output connected to a write port of the array.

4. The memory circuit according to claim 3, wherein: The first multiplexer is configured to select a data element received from one of the first input and the second input based on a logically inverted version of the control signal, and the second multiplexer is configured to select a data element received from one of the third input and the fourth input based on the control signal.

5. A memory circuit comprising: an array, comprising a plurality of processing elements (PEs); Wherein, each processing element includes a plurality of storage units; and Each processing element is configured to selectively store (i) a single first data element from among multiple first data elements in a storage unit of a corresponding storage unit; or (ii) a plurality of second data elements from among multiple second data elements in respective corresponding storage units based on a control signal indicating a layer type.

6. The memory circuit according to claim 5, wherein: The first data element comprises a weight data element and the second data element comprises an input data element.

7. The memory circuit according to claim 5, further comprising: A data router is configured to receive the control signal and determine whether to store a single first data element among the plurality of first data elements or to store a plurality of second data elements among the plurality of second data elements based on the control signal.

8. A method for operating a memory circuit, comprising: identifying a layer type of a neural network for processing a plurality of input data elements and a plurality of weight data elements; responsive to the layer type being a first type, storing a single weight data element of the plurality of weight data elements in one of a plurality of storage units of a corresponding processing element; as well as In response to the layer type being the second type, a plurality of the input data elements are stored in a plurality of the storage units of the corresponding processing element, respectively.

9. The method according to claim 8, wherein: The first type includes regular convolutional layers or attention layers, while the second type includes depthwise convolutional layers.