Storage and calculation function fusion macro cell supporting bidirectional ping-pong read-write
By designing a fusion macro unit that supports bidirectional ping-pong reading and writing, the problem that the fusion macro unit of the fusion macro unit of the comprehension function does not support ViT attention layer redundant memory access and two computing modes is solved, and the effect of reducing redundant memory access and improving computing efficiency is achieved.
Patent Information
- Application Number
- CN202510097598.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-20
AI Technical Summary
The memory function fusion macro unit does not support attention layer redundant memory access and two computing modes of the visual converter model (ViT), resulting in insufficient configurability of computing resources.
A fusion macro unit of memory and calculation function that supports bidirectional ping-pong reading and writing is designed, including a ping-pong transpose mode controller, input weight controller, transistor, memory and accumulator. By optimizing the memory path of intermediate results, it supports online transpose matrix multiplication and normal matrix multiplication.
Effectively reduce redundant memory access, optimize data flow scheduling, improve hardware energy efficiency, and support efficient matrix computing to improve computing efficiency.
Smart Images

Figure CN120179600A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of memory - computing function integration, and particularly to a memory - computing function integration macro - cell that supports bidirectional ping - pong read - write. Background Art
[0002] The Vision Transformer (ViT) is a tool applied in the field of computer vision for processing image recognition tasks based on deep - learning models. The Vision Transformer model can divide an image into multiple small patches, and then linearly embed these small patches into a dimensional space to form a 1D vector sequence, making the vector sequence compatible with the transformer architecture. Therefore, the Vision Transformer model can have excellent performance in computer vision tasks and has gradually become the mainstream model.
[0003] The self - attention mechanism in ViT can involve the calculation of query (Q), key (K), and value (V) matrices. The Vision Transformer can implement the calculation of query, key, and value matrices through matrix multiplication. For example, in the self - attention mechanism, the query matrix performs a dot - product operation with the transpose of the key matrix to calculate the attention scores, which are used to weight the value matrix to obtain a weighted output.
[0004] However, due to the redundant memory access between the generation stage of query, key, and value matrices and the attention calculation stage (QK T , AV), and the co - existence of online transposed matrix multiplication and normal matrix multiplication operators, it poses a severe challenge to the configurability of computing resources, resulting in the lack of effective support of the memory - computing function integration macro - cell for the redundant memory access of the ViT attention layer and the two operation modes. Summary of the Invention
[0005] In view of this, an embodiment of this application provides a memory - computing function integration macro - cell that supports bidirectional ping - pong read - write to solve the problem that the memory - computing function integration macro - cell does not support redundant memory access of the ViT attention layer and the two operation modes.
[0006] According to one aspect of this application, a memory - computing function integration macro - cell that supports bidirectional ping - pong read - write is provided and applied to the Vision Transformer model. The unit includes:
[0007] A ping - pong transpose mode controller configured to generate control signals;
[0008] A first transistor including a first connection end, a second connection end, and a third connection end; the first connection end is connected to the ping - pong transpose mode controller through a global word line;
[0009] An input weight controller configured to perform weight update and calculation;
[0010] A second transistor, including a fourth connection terminal, a fifth connection terminal, and a sixth connection terminal; the fourth connection terminal is connected to the input weight controller through a global bit line;
[0011] A memory, including a first read / write terminal, a second read / write terminal, a third read / write terminal, a fourth read / write terminal, and an output terminal; the first read / write terminal is connected to the second connection terminal through a vertical word line; the second read / write terminal is connected to the third connection terminal through a horizontal word line; the third read / write terminal is connected to the fifth connection terminal through a horizontal bit line; the fourth read / write terminal is connected to the sixth connection terminal through a vertical bit line;
[0012] An accumulator, connected to the output terminal;
[0013] The ping-pong transpose mode controller is further configured to:
[0014] When performing macro update in the normal mode, generate a first control signal, where the first control signal is used to connect the vertical word line to the global word line and connect the vertical bit line to the global bit line;
[0015] When performing macro update in the transpose mode, generate a second control signal, where the second control signal is used to connect the horizontal word line to the global word line and connect the horizontal bit line to the global bit line.
[0016] Optionally, it further includes a vertical bit line decoder. The fifth connection terminal is connected to the vertical bit line decoder through the horizontal bit line. The vertical bit line decoder is configured to disperse a first number of the horizontal bit lines into a second number of sub-horizontal bit lines; the first number is equal to the number of the second transistors; the second number is equal to the number of the first transistors.
[0017] Optionally, the input weight controller is further connected to the horizontal word line. The ping-pong transpose mode controller is further configured to:
[0018] When the input data is valid, select a horizontal word line from the horizontal bit lines to write the input data column by column;
[0019] By controlling the global word line and the global bit line, perform data writing in the normal mode and data writing in the transpose mode when updating data.
[0020] Optionally, the memory is a static random access memory; the memory includes a multiplexer, a storage sub-unit, a first selector, and a second selector;
[0021] The first read / write terminal and the second read / write terminal are arranged on the multiplexer; the multiplexer is further connected to the storage sub-unit;
[0022] The third read / write terminal and the fourth read / write terminal are disposed on the first selection controller; the first selection controller is further connected to the storage sub-unit;
[0023] The second selection controller is connected to the storage sub-unit and the output terminal.
[0024] Optionally, the memory includes at least two of the storage sub-units, and among the at least two storage sub-units, a part of the storage sub-units are configured as computing blocks; another part of the storage sub-units are configured as storage blocks; the ping-pong transpose mode controller is further configured to:
[0025] Generate a processing signal according to the operating mode, the processing signal including a first signal value and a second signal value;
[0026] During weight update, if the processing signal is the first signal value, the weight is stored in the computing block through the multiplexer;
[0027] If the processing signal is the second signal value, the weight is stored in the storage block through the multiplexer.
[0028] Optionally, the ping-pong transpose mode controller is further configured to:
[0029] Generate an activation signal according to the operating mode, the activation signal including a third signal value and a fourth signal value;
[0030] During execution of the calculation, if the activation signal is the third signal value, the activation is multiplied by the computing block through the multiplexer;
[0031] If the activation signal is the fourth signal value, the activation is multiplied by the storage block through the multiplexer.
[0032] Optionally, the ping-pong transpose mode controller is further configured to:
[0033] During execution of in-memory computing, control the activation to be written into the memory in a bit-serial manner along the vertical bit line direction.
[0034] Optionally, the ping-pong transpose mode controller is further configured to:
[0035] Control the activation and the pre-stored data to perform bit-by-bit multiplication through an AND gate to obtain an operation result;
[0036] Use an adder to accumulate the operation result along the vertical word line to obtain accumulated data;
[0037] Store the accumulated data in the accumulator.
[0038] Optionally, the ping-pong transpose mode controller is further configured to:
[0039] Obtain the activation accuracy, where the activation accuracy includes a preset number of bit accuracy values;
[0040] Set the accumulation time according to the activation accuracy, where the accumulation time is equal to the number of clock cycles of the preset number of bit accuracy values;
[0041] Set the accumulation time to the accumulator.
[0042] According to another aspect of the present application, there is provided a visual transformer model acceleration system, including the memory-computation function fusion macro cell that supports bidirectional ping-pong reading and writing.
[0043] By means of the above technical solutions, the embodiments of the present application provide a memory-computation function fusion macro cell that supports bidirectional ping-pong reading and writing. The cell includes a ping-pong transpose mode controller, an input weight controller, a first transistor, a second transistor, a memory, and an accumulator. When performing macro update in the normal mode, the ping-pong transpose mode controller can connect the vertical word line to the global word line and the vertical bit line to the global bit line through the first control signal. When performing macro update in the transpose mode, the ping-pong transpose mode controller connects the horizontal word line to the global word line and the horizontal bit line to the global bit line through the second control signal. By optimizing the macro cell storage path of the intermediate result and supporting efficient online transpose matrix multiplication and normal matrix multiplication, the data flow scheduling of the visual transformer model is optimized, the redundant memory access is effectively reduced, and the hardware energy efficiency is improved.
[0044] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Description of the Drawings
[0045] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0046] Figure 1 It is a schematic structural diagram of a memory-computation function fusion macro cell that supports bidirectional ping-pong reading and writing provided by the embodiment of the present application;
[0047] Figure 2 It is a schematic structural diagram of a transistor provided by the embodiment of the present application;
[0048] Figure 3 It is a schematic structural diagram of a memory provided by the embodiment of the present application;
[0049] Figure 4Schematic diagram of the multiplexer structure provided by the embodiment of the present application;
[0050] Figure 5 Schematic diagram of the HBL decoder structure provided by the embodiment of the present application;
[0051] Figure 6 Schematic diagram of the control flow of the normal mode and transpose mode provided by the embodiment of the present application;
[0052] Figure 7 Schematic diagram of the connection effect between word lines and bit lines in the normal mode provided by the embodiment of the present application;
[0053] Figure 8 Schematic diagram of the connection effect between word lines and bit lines in the transpose mode provided by the embodiment of the present application;
[0054] Figure 9 Schematic diagram of the accumulation structure provided by the embodiment of the present application;
[0055] Figure 10 Schematic diagram of the structure of the visual transformer model acceleration system provided by the embodiment of the present application. Detailed implementation manners
[0056] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0057] In the embodiment of the present application, the computing-in-memory (CIM) technology, also known as in-memory computing technology, is a computing architecture that integrates computing and storage functions. By integrating computing units and storage units together, a tightly coupled structure is formed. For example, a processing unit is embedded in a memory or a storage unit is embedded in a processing unit. The computing-in-memory technology can directly execute computing operations in the memory, reduce frequent data transmission, and improve computing efficiency.
[0058] The computing-in-memory macro cell is a processing unit designed by applying the computing-in-memory technology. The computing-in-memory macro cell can be a chip, an integrated circuit, a microprocessor, etc. For example, the computing-in-memory macro cell can integrate computing and storage functions on the same chip to improve data processing efficiency and reduce energy consumption. The computing-in-memory macro cell can embed computing capabilities inside the storage array to achieve a tight combination of data storage and computing, thereby enhancing the memory access bandwidth and reducing the memory access power consumption.
[0059] In some embodiments, the memory - computing integrated macro - cell may include a computing sub - unit and a storage medium. Among them, the computing sub - unit, also known as the controller, control sub - unit, etc., is responsible for executing computing tasks; the storage medium is used to store data. The storage media involved in the memory - computing integrated macro - cell may include: Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Resistive Random Access Memory (RRAM), Magnetic Random Access Memory (MRAM), etc. By tightly integrating the computing sub - unit and the storage medium, the integration of data storage and computing can be achieved, thereby improving data - processing efficiency and reducing energy consumption.
[0060] The memory - computing integrated macro - cell can be applied to Artificial Intelligence (AI) algorithms to provide operator acceleration for vector - matrix multiplication in AI algorithms, including being applied to AI algorithms such as Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN). For example, the memory - computing integrated macro - cell can provide a computing power of more than 1000 TOPS and a high energy efficiency of more than 10 - 100 TOPS / W in the field of AI algorithms.
[0061] According to the different specific application fields of AI algorithms, the memory - computing integrated macro - cell can also be designed specifically according to the design requirements of specific application fields to improve the data - processing efficiency in specific application fields. In some embodiments, the memory - computing integrated macro - cell can be applied to the Vision Transformer model. The Vision Transformer model is a collection of deep - learning model software applications and hardware devices used to perform tasks such as image recognition in the field of computer vision. For example, the Vision Transformer model can be based on the Transformer architecture of Natural Language Processing (NLP), convert image data into sequences, and use the self - attention mechanism to capture the relationships in the images.
[0062] When performing image processing, the Vision Transformer (ViT) model can divide an image into multiple small patches, and then linearly embed these patches into a dimensional space to form a sequence of 1D vectors. That is, each patch can be serialized into a vector, mapped to a smaller dimension through a single matrix multiplication, and the input image can be decomposed into a series of patches, enabling the vector sequence corresponding to the image segmentation result to be compatible with the Transformer architecture. Then, feature extraction is performed based on the vector sequence to identify feature targets in the image, that is, the self-attention mechanism is used to capture local and global relationships within the image, and the Transformer encoder is used to process these vector embeddings, enabling ViT to distinguish patterns present in the image by capturing context and hierarchical features based on datasets such as ImageNet in image classification tasks.
[0063] For example, the Vision Transformer model can be used to perform object detection, that is, by using the self-attention mechanism, ViT is used to enhance the ability of the object detection model to identify and locate objects in the image, improving the detection performance. The Vision Transformer model can also utilize its ability to divide the image into multiple parts to capture dependencies in application fields such as autonomous driving and medical imaging, thereby accurately outlining object boundaries. The Vision Transformer model can also be applied to video recognition tasks, achieving excellent performance in video understanding tasks by combining spatio-temporal attention mechanisms.
[0064] When the Vision Transformer model performs image processing, it uses matrices such as Query, Key, and Value. These matrices are derived from the input embedding representation and are used to calculate attention scores and weighted outputs. For example, the Query matrix can be generated by multiplying the input embedding by a learnable weight matrix. The Key matrix can be generated by multiplying the input embedding by another learnable weight matrix. The Value matrix is generated by multiplying the input embedding by a third learnable weight matrix.
[0065] Therefore, in vision tasks, the processing of the Query, Key, and Value matrices by the Vision Transformer model can include stages such as matrix generation and attention calculation. Among them, the attention calculation stage refers to the calculation of attention and value. Since in the self-attention mechanism, the attention scores can be used to weight the value matrix, the Vision Transformer model can first calculate the attention scores, that is, use the query matrix Q and the key matrix K to calculate the attention scores.
[0066] For example, the dot product of the query matrix Q and the transpose of the key matrix K (Q·K T) to calculate the attention scores. The calculated attention scores can also be divided by a scaling factor, such as the square root of the key matrix dimension, to avoid the problem of vanishing gradients caused by the overly large dot product results in a high-dimensional space. Then apply the Softmax function to convert the attention scores into a probability distribution through the Softmax function. The Softmax function ensures that the sum of all attention scores is 1. Then calculate the weighted values, that is, use the attention weights to perform a weighted sum on the value matrix V. During the weighted sum process, each value vector can be multiplied by its corresponding attention weight, and then all the weighted value vectors are summed up.
[0067] Due to the redundant memory access between the generation stage of the Query, Key, and Value matrices and the attention calculation stage (QKT, AV) in some vision transformer models, and during the matrix multiplication calculation process, there are certain requirements for the configurability of computing resources. It makes the memory-computation function fusion macrocell in some vision transformer models lack effective support for the redundant memory access and two operation modes of the ViT attention layer.
[0068] Based on this, in order to optimize the data flow scheduling of the vision transformer model and reduce redundant memory access, in this embodiment, a memory-computation function fusion macrocell supporting bidirectional ping-pong reading and writing is provided, such as Figure 1 shown, the memory-computation function fusion macrocell supporting bidirectional ping-pong reading and writing can be applied to the vision transformer model to perform vision processing tasks.
[0069] The memory-computation function fusion macrocell supporting bidirectional ping-pong reading and writing includes: a ping-pong transpose mode controller, an input weight controller, a first transistor, a second transistor, a memory, an accumulator, and word lines (WordLine, WL) and bit lines (BitLine, BL) for connecting each component. Among them, the word line is a control line used to select a specific row in the memory array. For example, in an SRAM array, a word line spans a row and is connected to all memories in that row through access transistors. The row address decoder is connected to all word lines and is used to activate the correct word line according to the row address. When the word line is activated, it allows data to be transferred between the bit line and the memory, thereby realizing data reading or writing.
[0070] A bit line is a line used to transfer data in a memory array. In an SRAM array, a pair of bit lines spans a column and connects to all the cells in that column. Bit lines can be used for data input and output of the memory. In a read operation, the bit lines can be precharged to a known voltage, and then by activating the target word line, the data in the memory will keep the bit lines at a high level or pull them low, thus generating a voltage difference in the bit line pair, which can be detected by a sense amplifier and converted into a digital signal. In a write operation, the bit lines are set to the data (0 or 1) to be written, and then by activating the target word line, the memory will change its state according to the voltage of the bit lines.
[0071] The ping-pong transpose mode controller is a control unit that enables efficient read and write operations in data stream processing. When performing data read and write operations, the memory is divided into two types, similar to the two buffers in a ping-pong operation. When one part of the memory is performing a read operation, the other part of the memory can perform a write operation simultaneously, and vice versa. By means of parallel processing, the data throughput is increased and the waiting time is reduced, thus improving the overall performance.
[0072] For example, the ping-pong transpose mode controller can be implemented by a state machine, which can control the data writing and reading between the two memories. When the data enable signal is high, the data is first written into TS1. When TS1 is full, the controller will switch to write TS2 and read data from TS1 simultaneously, and so on, realizing writing in one buffer while reading from the other buffer, facilitating seamless buffering and processing of data.
[0073] To perform data read and write operations, the ping-pong transpose mode controller is also configured to generate control signals according to the actual operating mode and calculation requirements, so as to control the connection state of the word lines and bit lines through the control signals.
[0074] The input weight controller is used to control the content related to weight update and weight calculation in the visual conversion process. That is, the input weight controller is configured to perform weight update and calculation. Among them, weight calculation and weight update involve the adjustment of weight parameters by the neural network during the training process. The purpose of weight calculation and weight update is to minimize the loss function, calculate the gradient of the loss with respect to the model parameters through the backpropagation algorithm, and use optimizers such as Adam and SGD to update the model parameters according to these gradients.
[0075] In some embodiments, the weights can be embodied as a weight matrix, which is used to calculate query vectors, key vectors, and value vectors. During the model construction phase, the weight matrix is randomly initialized with some values. At each step of training, the model performs forward propagation on the input data to calculate the predicted output, and then calculates the loss based on the predicted output and the true labels. The gradient of the loss with respect to the model parameters is calculated through the backpropagation algorithm, and finally the optimizer is used to update the model parameters according to the calculated gradient. The updated model parameters can include the parameters used to calculate the attention weights.
[0076] In some embodiments, the weight update method in the memory-computation fusion system may include obtaining a target weight matrix, receiving the calculation result obtained after the memory-computation fusion array mapped to the target weight matrix performs forward calculation, and performing an inverse solution process on the calculation result to obtain an inverse solution weight matrix. An updated weight matrix is obtained based on the target weight matrix and the inverse solution weight matrix, where the updated weight matrix is used to update the weights of the memory-computation fusion array mapped to the target weight matrix. By performing an inverse solution process on the calculation result of the forward calculation, an inverse solution weight matrix that captures the non-ideal factors in the fusion forward calculation is obtained, and an updated weight matrix for guiding the actual weight update is obtained based on the inverse solution weight matrix and the target weight matrix, thereby improving the accuracy of weight update and the performance of the system.
[0077] As Figure 2 shown, the first transistor includes a first connection terminal, a second connection terminal, and a third connection terminal. Among them, the first connection terminal, the second connection terminal, and the third connection terminal, as the three connection terminals of the first transistor, can be controlled by a control signal to achieve the connection state between two of the connection terminals. For example, the first connection terminal is the base of the first transistor, the second connection terminal is the emitter of the first transistor, and the third connection terminal is the collector of the first transistor.
[0078] As Figure 2 shown, the second transistor includes a fourth connection terminal, a fifth connection terminal, and a sixth connection terminal. Similar to the first transistor, the fourth connection terminal, the sixth connection terminal, and the fifth connection terminal as the three connection terminals of the second transistor can also be controlled by a control signal to achieve the connection state between two of the connection terminals. For example, the fourth connection terminal is the base of the second transistor, the fifth connection terminal is the emitter of the second transistor, and the sixth connection terminal is the collector of the second transistor.
[0079] The ping-pong transpose mode controller and the input weight controller can control the connection state of the bit line and the word line in the memory-computation integrated macro cell through the first transistor and the second transistor, so as to perform corresponding data processing in different operating modes (or calculation stages). Among them, the bit line can include a global bit line (GBL), a horizontal bit line (HBL), and a vertical bit line (VBL). Similarly, the word line can include a global word line (GWL), a horizontal word line (HWL), and a vertical word line (VWL).
[0080] As components of the memory array, the bit line and the word line can be used to connect the outputs of the memories in the same row or the same column. Thus, when the corresponding bit line or word line is activated, the corresponding memory is connected to the word line or the bit line to realize data reading or writing. For example, in the memory array, the horizontal bit line is perpendicular to the word line and is used to connect the outputs of all memories in the same column. When the word line is activated, the corresponding memory is connected to the bit line to realize data reading or writing. In the memory-computation integrated macro cell, the horizontal bit line can be used to realize parallel processing and transmission of data, thereby improving the calculation efficiency and speed, and allowing calculations such as logical operations to be performed inside the memory array on the bit line without moving the data to an external computing unit, realizing in-memory computing.
[0081] Therefore, the first connection end is connected to the ping-pong transpose mode controller through the global word line. The fourth connection end is connected to the input weight controller through the global bit line. That is, the first connection end of the first transistor can be connected to the global word line GWL. The fourth connection end of the second transistor is connected to the global bit line GBL.
[0082] As Figure 3 shown, the memory, as the basic data storage and calculation unit of the memory-computation integrated macro cell, can be used to store the original data in the calculation process. The memory can include a first read / write end, a second read / write end, a third read / write end, a fourth read / write end, and an output end. Among them, the first read / write end is connected to the second connection end through the vertical word line VWL; the second read / write end is connected to the third connection end through the horizontal word line HWL; the third read / write end is connected to the fifth connection end through the horizontal bit line HBL; the fourth read / write end is connected to the sixth connection end through the vertical bit line VBL.
[0083] In some embodiments, the memory is a static random access memory (SRAM); the memory includes a multiplexer, a storage sub-unit, a first selector, and a second selector. Among them, the multiplexer can serve as the memory access transistors for controlling the read and write operations of the memory and selecting specific storage sub-units in the read and write operations. Taking 6T SRAM as an example, the memory may include six transistors. As Figure 4 shown, M&D represents the multiplexer, which may include a multiplexer (MUX) and a demultiplexer (DEMUX). The multiplexer is a multi-input single-output device that can select one output from multiple input signals according to the select signal. The demultiplexer is a single-input multi-output device that can distribute one input signal to one of multiple output lines according to the select signal. 6T represents the transistors for storing data, that is, the storage sub-unit. Therefore, the first read / write terminal and the second read / write terminal are provided on the multiplexer; the multiplexer is also connected to the storage sub-unit.
[0084] Obviously, the memory may include at least two storage sub-units. Among at least two storage sub-units, a part of the storage sub-units are configured as computing blocks; another part of the storage sub-units are configured as storage blocks. For example, as Figure 3 shown, when the memory may include 2 storage sub-units, namely TS1 and TS2. Then when the storage sub-unit TS1 is used as a computing block, the storage sub-unit TS2 can be used as a storage block.
[0085] The first selector and the second selector are used to perform selection control. Therefore, the third read / write terminal and the fourth read / write terminal are provided on the first selector; the first selector is also connected to the storage sub-unit; the second selector is connected to the storage sub-unit and the output terminal.
[0086] The memory-computation function integration macro cell may include multiple memories, and the multiple memories may be arranged in a rectangular array form. That is, the multiple memories in the memory-computation function integration macro cell can be arranged in a set number of rows and a set number of columns. Then, in order to perform data read / write and computation on the memories in a specific row and a specific column. The memory-computation function integration macro cell may also include a set number of column first transistors and a set number of row second transistors.
[0087] For example, the multiple memories included in the memory - computing function integration macro - cell can form a memory array with 64 rows and 512 columns. Then the memory - computing function integration macro - cell can include 512 first transistors and 64 second transistors. Correspondingly, the global word lines connecting the multiple first transistors can be numbered as GWL0, GWL1, GWL2, ……, GWL511 respectively. Similarly, the global bit lines connecting the second transistors can be numbered as GBL0, GBL1, GBL2, ……, GBL63 respectively.
[0088] The read - write terminals of multiple memories in the same row or the same column can be connected to the same bit line and word line, so as to be controlled by the ping - pong transpose mode controller or the input weight controller to perform data transmission, read - write, and calculation in the visual transformation algorithm. For example, the first write terminals of 64 memories in the first column can be connected to the vertical word line VWL0; the first write terminals of 64 memories in the second column can be connected to the vertical word line VWL1, and so on. Similarly, the fourth read - write terminals of 512 memories in the first row are connected to the same vertical bit line VBL0; the fourth read - write terminals of 512 memories in the second row are connected to the same vertical bit line VBL1, and so on.
[0089] As Figure 5 shown, in some embodiments, the memory - computing function integration macro - cell further includes a vertical bit - line decoder (HBL Decoder), and the fifth connection terminal is connected to the vertical bit - line decoder through the horizontal bit line HBL. And the vertical bit - line decoder is configured to disperse the first number of the horizontal bit lines HBL into the second number of sub - horizontal bit lines HBLL. Wherein, the first number is equal to the number of the second transistors; the second number is equal to the number of the first transistors.
[0090] For example, the memory - computing function integration macro - cell includes 512 first transistors and 64 second transistors. Then the first number is equal to 64, and the second number is equal to 512. Correspondingly, the horizontal bit lines HBL connected to the fifth connection terminal are HBL0 to HBL63, that is, HBL[63:0]. Then the HBL Decoder can divide HBL into 512 sub - horizontal bit lines HBLL, that is, HBLL[511:0].
[0091] Therefore, in some embodiments, the third read - write terminals of the memories can be connected to the sub - horizontal lines HBLL. For example, the third read - write terminals of multiple memories in the first column are connected to the sub - horizontal line HBLL0, and the third read - write terminals of multiple memories in the second column are connected to the sub - horizontal line HBLL1.
[0092] Moreover, the horizontal word line HWL can also be connected to an input weight controller for transmitting weight-related data, and the number of horizontal word lines HWL is also set according to the data processing bit width of the memory-computation function integrated macro cell. For example, if the data processing bandwidth of the memory-computation function integrated macro cell is 8 bits, the second read / write terminals of multiple memories in the first row can be connected to HWL[7:0], and the second read / write terminals of multiple memories in the second row can be connected to HWL[15:8].
[0093] The output terminals of multiple memories can be connected to an accumulator. The accumulator is a register in the memory-computation function integrated macro cell for storing the result or intermediate result of an arithmetic or logical operation. In some embodiments, multiple memories can be connected to the accumulator at their output terminals in the form of a bit-serial adder tree, so that when in-memory computing is performed, activation is written in a bit-serial manner along the VBL direction. The activation and pre-stored data in the memory-computation function integrated macro cell are multiplied bit by bit through an AND gate, and the operation result is accumulated along the VWL by an adder tree and then stored in the accumulator.
[0094] It can be seen that in the above embodiments, the design of the memory-computation function integrated macro cell structure for ViT operations allows data storage and calculation operations to be performed in parallel. That is, when generating the Query, Key, and Value matrices, the generated intermediate results are not written back to memory, but are directly written back to the storage column of the ping-pong block in transposed or normal mode through the write-back path within the macro cell. When performing attention calculations (QK T 、AV), the input activation is directly fed to the storage column to complete matrix operations.
[0095] To support bidirectional ping-pong read / write, as Figure 6 shown, the ping-pong transpose mode controller is further configured as:
[0096] S100: Generate a first control signal when performing macro update in the normal mode.
[0097] Wherein, the first control signal is used to connect the vertical word line VWL to the global word line GWL and connect the vertical bit line HBL to the global bit line GBL.
[0098] By using the design of the global word line GWL and the global bit line GBL, transpose mode and normal mode input configurations are realized. That is, in the normal calculation mode and the transpose calculation mode, a bidirectional ping-pong memory-computation function integrated macro cell (TWPP CIM) is configured to adapt to tasks such as transpose matrix operations Q×K T and normal matrix operations (AV, FFN, Proj) of the vision transformer model, without the need for an additional transpose unit for data preprocessing.
[0099] Such as Figure 7As shown, the memory - computing function - integrated macro - cell can achieve two - way decoding of the two - way ping - pong memory - computing function - integrated macro - cell according to a preset algorithm logic. When performing macro - update in the normal mode, the ping - pong transpose mode controller can generate a control signal with T = False. After inputting the control signal to the first transistor and the second transistor, it can control the first transistor to connect the vertical word line VWL to the global word line GWL, and control the second transistor to connect the vertical bit line VBL to the global bit line GBL.
[0100] Among them, the preset algorithm logic can define the input parameters, output parameters, related constants, and algorithm programs of the memory - computing function - integrated macro - cell in various operating modes according to the computing requirements of visual conversion. For example, the two - way decoding algorithm can define the input parameter quantity: Addr_In is the address signal waiting to be decoded, which is an N - bit binary number, that is, [N:0]; Mode is the write - mode selection signal, which can be the normal mode (norm) or the transpose mode (transpose); Cycle is the number of clock cycles determined according to the data width (4 - bit, 8 - bit, or 16 - bit), that is, 4 / 8 / 16 clock cycles for 4 - bit / 8 - bit / 16 - bit; P represents the operation precision signal, which is used to determine the calculation method of the column address; T represents the transpose - mode signal, which is used to select ordinary writing or transpose writing. The output parameter quantity: Row_Addr_Out represents the row - address signal; Col_Addr_Out represents the column - address signal. Constants: Row_Width represents the row width of the CIM macro; Col_Width represents the column width of the CIM macro, etc.
[0101] Based on the defined input quantities, output quantities, and constants, the control program of the two - way decoding algorithm can be written. For example: if T is false; then Col_Addr_Out←Addr_In; Row_Addr_Out∈Row_Width. Else if T is true; Col_Addr_Out bit_group ←Col_Addr_group; return Row_Addr_Out,Col_Addr_Out.
[0102] That is, for the decoding process: if T is false (normal mode), then directly output Addr_In as the column address, and the row address is extracted from Addr_In, and the specific range is determined by Row_Width. If T is true (transpose mode), then a more complex decoding process is required: first calculate the column - address count Col_Count, which is obtained by shifting Col_Width to the right It is achieved by bits. Then, by looping through each bit group (bit_group) and each group (n_group), the column address Col_Addr_Out of each bit group is calculated. Among them, the column address Col_Addr_Out can be achieved by shifting the group number n_group by bits and adding the bit group number bit_group, and then extracting the corresponding bits from Addr_In. The row address Row_Addr_Out is extracted from the high-order part of Addr_In, and the specific range is from N to bits. Finally, the algorithm returns the decoded row address and column address as the return value.
[0103] It can be seen that based on the above two-way decoding algorithm, in the memory interface, the input address signal (Addr_In) can be decoded into a row address (Row_Addr_Out) and a column address (Col_Addr_Out). In order to generate the correct row address and column address according to the input address signal and operation accuracy in different write modes, so as to correctly store or read data in the memory.
[0104] S200: Generate a second control signal when performing macro update in the transpose mode.
[0105] Among them, the second control signal is used to connect the horizontal word line HWL to the global word line GWL, and connect the horizontal bit line VBL to the global bit line GBL. For example, as Figure 8 shown, when performing macro update in the transpose mode, based on the decoding result of the preset HBL table, the ping-pong transpose mode controller can generate a control signal with T = True. After inputting the control signal to the first transistor and the second transistor, it can control the first transistor to connect the vertical word line HWL to the global word line GWL, and control the second transistor to connect the horizontal bit line HBL to the global bit line GBL.
[0106] It can be seen that in the above embodiment, by designing a memory-computation function fusion macro cell that supports a two-way ping-pong read-write mechanism, it can adapt to tasks such as visual transformer transpose matrix operations (Q×K T ) and normal matrix operations (AV, FFN, Proj), etc. By optimizing the main layer operation path, it allows the generated intermediate calculation results to be directly written back to the storage column of the ping-pong block without writing back to the memory, reducing redundant memory access. And it supports two types of matrix calculation modes, Q×K T and A×V, without an additional transpose buffer.
[0107] In some embodiments, when the input data is valid, the ping-pong transpose mode controller may select a horizontal word line (HWL) from the horizontal bit lines (HBLs) to write the input data column by column; by controlling the global word line (GWL) and the global bit line (GBL), perform data writing in the normal mode and data writing in the transpose mode when updating the data.
[0108] When the input data is valid, by selecting the HWL from the HBLs and writing column by column, and by controlling the GWL and the GBL, the compute-in-memory function integrated macro cell can support bidirectional ping-pong reading and writing, so as to achieve normal writing and transpose writing when updating the data.
[0109] In some embodiments, the ping-pong transpose mode controller is further configured to: generate a processing signal PPU according to the operating mode. Wherein, the processing signal includes a first signal value and a second signal value. For example, the first signal value is 1 and the second signal value is 0. Then during weight update, the ping-pong transpose mode controller may generate a processing signal of PPU = 1 or PPU = 0 according to the operating mode.
[0110] During weight update, if the processing signal is the first signal value, the weight is stored in the computing block through the multiplexer. That is, when PPU is 1, the weight can be stored in the computing block through the multiplexer. If the processing signal is the second signal value, the weight is stored in the storage block through the multiplexer. That is, when PPU is 0, the weight is stored in the storage block through the multiplexer.
[0111] Similarly, the ping-pong transpose mode controller can also control the data processing process during weight calculation according to the operating mode, that is, the ping-pong transpose mode controller is further configured to: generate an activation signal (PPC) according to the operating mode. Wherein, the activation signal includes a third signal value and a fourth signal value. Similar to the processing signal PPU, the third signal value can be 1 and the fourth signal value can be 0.
[0112] During the execution of the calculation, if the activation signal is the third signal value, the activation is multiplied by the computing block through the multiplexer; if the activation signal is the fourth signal value, the activation is multiplied by the storage block through the multiplexer. That is, during the calculation operation, when PPC is 1, the activation is multiplied by the computing block through the multiplexer. On the contrary, when PPC is 0, the activation is multiplied by the storage block through the multiplexer.
[0113] In some embodiments, when performing in-memory computing, the control activation is performed in a bit-serial manner, and data is written to the memory along the vertical bit line (VBL) direction. In the bit-serial manner, starting from the least significant bit, a pair of bits in two operands are processed in each cycle. This manner allows a result bit to be written back on each bit line in each cycle, which is suitable for arithmetic operations in the SRAM array, that is, the computing mode is completed in a bit-serial format rather than a bit-parallel format. The bit-serial manner can reduce the number of simultaneously activated bit lines, thereby reducing power consumption.
[0114] As Figure 9 shown, in some embodiments, the ping-pong transpose mode controller is further configured to: control the activation and pre-stored data to perform bit-by-bit multiplication through an AND gate to obtain an operation result. Then, an adder is used to accumulate the operation result along the vertical word line (VWL) to obtain accumulated data. Then, the accumulated data is stored in the accumulator.
[0115] That is, the activation and pre-stored data in the memory-computation function fusion macro cell can perform bit-by-bit multiplication through an AND gate, and after obtaining the operation result, it is accumulated along the vertical word line VWL by an adder tree, and then stored in the accumulator.
[0116] When performing the accumulation operation, the ping-pong transpose mode controller can also obtain the activation precision, where the activation precision includes a preset number of bit precision values. For example, the activation precision can include 4 / 8 / 16-bit precision. Then, the accumulation time is set according to the activation precision. Wherein, the accumulation time is equal to the number of clock cycles of the preset number of bit precision values. For example, for 4 / 8 / 16-bit precision, it is 4 / 8 / 16 clock cycles. Then, the accumulation time is set for the accumulator.
[0117] By applying the technical solutions provided in the above embodiments, the memory-computation function fusion macro cell allows the simultaneous execution of weight update and calculation using the ping-pong mode and multiplexer design, realizing the support for bidirectional ping-pong read and write. And, through digital architecture design, front-end design, and back-end design of digital circuits, an integrated circuit architecture and chip including the memory-computation function fusion macro cell can be obtained. For example, the process technology of the integrated circuit architecture and chip can adopt the 28nm process, and then the packaged chip is tested for power consumption and performance.
[0118] For the memory-computation function fusion macro cell, it can be tested in multiple Vision Transformer (ViT) models based on the CIFAR10, CIFAR100, and ImageNet 1K datasets. According to the test results, after adopting the memory-computation function fusion macro cell, the computational acceleration ratio of the ViT model in the attention layer can be increased by 30.8 times, and the energy efficiency is increased by 1.13 times. By supporting bidirectional ping-pong read and write operations, the data access in the attention layer is reduced by 4.6 times, significantly reducing redundant memory access and improving computational efficiency.
[0119] Further, as a specific implementation of the memory - computing function fusion macro - cell that supports two - way screen office reading and writing in the above - mentioned embodiments, the embodiments of the present application provide a visual transformer model acceleration system, as Figure 10 shown. This visual transformer model acceleration system includes the memory - computing function fusion macro - cell that supports two - way ping - pong reading and writing provided in the above - mentioned embodiments.
[0120] It should be noted that for other corresponding descriptions of each functional unit involved in the visual transformer model acceleration system provided in the embodiments of the present application, reference can be made to the corresponding descriptions in the memory - computing function fusion macro - cell that supports two - way ping - pong reading and writing provided in the above - mentioned embodiments, and details are not described herein again.
[0121] The technical features of the above - mentioned embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above - mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0122] The above - mentioned embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it cannot be understood as a limitation on the patent scope of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A storage and computing function fusion macro unit supporting bidirectional ping-pong reading and writing, characterized in that: Applied to the visual converter model, the unit includes: a ping-pong transposition mode controller configured to generate a control signal; A first transistor comprises a first connection terminal, a second connection terminal and a third connection terminal; the first connection terminal is connected to the ping-pong transposition mode controller via a global word line; Input weight controller, which is configured to perform weight updates and calculations; A second transistor includes a fourth connection terminal, a fifth connection terminal and a sixth connection terminal; the fourth connection terminal is connected to the input weight controller through a global bit line; A memory, comprising a first read / write terminal, a second read / write terminal, a third read / write terminal, a fourth read / write terminal and an output terminal; the first read / write terminal is connected to the second connection terminal via a vertical word line; the second read / write terminal is connected to the third connection terminal via a horizontal word line; the third read / write terminal is connected to the fifth connection terminal via a horizontal bit line; the fourth read / write terminal is connected to the sixth connection terminal via a vertical bit line; an accumulator connected to the output end; The ping-pong transposition mode controller is further configured to: When performing a macro refresh in a normal mode, generating a first control signal for connecting the vertical word line to the global word line and connecting the vertical bit line to the global bit line; When a macro refresh is performed in the transposition mode, a second control signal is generated, the second control signal being used to connect the horizontal word line to the global word line and the horizontal bit line to the global bit line.
2. The unit according to claim 1, characterized in that Also comprising a vertical bit line decoder, the fifth connection end is connected to the vertical bit line decoder through the horizontal bit line, and the vertical bit line decoder is configured to disperse the first number of the horizontal bit lines into a second number of sub-horizontal bit lines; The first number is equal to the second number of transistors; the second number is equal to the first number of transistors.
3. The unit according to claim 2, characterized in that The input weight controller is also connected to the horizontal word line; the ping-pong transposition mode controller is further configured as: When input data is valid, selecting a horizontal word line from the horizontal bit lines to write the input data in columns; By controlling the global word lines and the global bit lines, data writing in a normal mode and data writing in a transposed mode are performed when updating data.
4. The unit according to claim 1, characterized in that The memory is a static random access memory; the memory includes a multiplexer, a storage subunit, a first selector and a second selector; The first read-write end and the second read-write end are arranged on the multiplexer; the multiplexer is also connected to the storage subunit; The third read / write terminal and the fourth read / write terminal are arranged on the first selector; the first selector is also connected to the storage subunit; The second selector is connected to the storage subunit and the output end.
5. The unit according to claim 4, characterized in that The memory comprises at least two storage subunits, among which a part of the storage subunits are configured as computing blocks; another part of the storage subunits are configured as storage blocks; and the ping-pong transposition mode controller is further configured as: generating a processed signal according to the operating mode, the processed signal comprising a first signal value and a second signal value; During weight update, if the processed signal is a first signal value, the weight is stored in the calculation block through the multiplexer; If the processed signal is a second signal value, the weight is stored in the storage block through the multiplexer.
6. The unit according to claim 4, characterized in that The ping-pong transposition mode controller is further configured to: generating an activation signal according to the operating mode, the activation signal comprising a third signal value and a fourth signal value; During the execution of the calculation, if the activation signal is a third signal value, activation is multiplied by the calculation block through the multiplexer; If the activation signal is a fourth signal value, activation is multiplied by the storage block through the multiplexer.
7. The unit according to claim 1, characterized in that The ping-pong transposition mode controller is further configured to: When performing in-memory computations, control activations are written to the memory in a bit-serial manner, along the vertical bit line direction.
8. The unit according to claim 1, characterized in that The ping-pong transposition mode controller is further configured to: Control activation and pre-stored data to perform bit-by-bit multiplication through an AND gate to obtain an operation result; Accumulating the operation results along the vertical word lines using an adder to obtain accumulated data; The accumulated data is stored in the accumulator.
9. The unit according to claim 8, characterized in that The ping-pong transposition mode controller is further configured to: Acquire activation precision, where the activation precision includes a preset number of digits of precision value; The accumulation time is set according to the activation precision, and the accumulation time is equal to the preset number of digit precision value clock cycles; The accumulation time is set to the accumulator.
10. A visual converter model acceleration system, characterized in that: It comprises a storage and computing function fusion macro unit supporting bidirectional ping-pong reading and writing as described in any one of claims 1-9.