Data storage and calculation function fusion method and device based on ping-pong assembly line, visual converter-oriented model accelerator and acceleration system

By adopting the data storage and computing function fusion method of the ping-pong pipeline in the visual converter model, data storage and computing are unified in the same component for parallel processing, which solves the problem of data transmission delay, improves data processing efficiency and real-time performance, and meets the deployment requirements of deep learning.

CN120743841APending Publication Date: 2025-10-03INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510621236.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing technologies, the data transmission process of the visual converter model is complex and has high latency, resulting in data processing bottlenecks and failing to meet the real-time and large-scale deployment requirements of deep learning.

Method used

A data storage and computing function fusion method based on ping-pong pipeline is adopted to unify data storage and computing in the same component for parallel processing. The data flow scheduling is optimized through ping-pong pipeline scheduling, reducing redundant memory access and data transmission links.

Benefits of technology

It improves the timeliness and efficiency of data processing, reduces the resource requirements for data storage and computing, and meets the real-time and large-scale deployment requirements of deep learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743841A_ABST
    Figure CN120743841A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data storage and calculation function fusion method and device based on a ping-pong assembly line, a visual converter-oriented model accelerator and an acceleration system, and aims to solve the technical problem of data processing bottleneck caused by delay in a related data transmission process. The method comprises the following steps of: respectively updating a key vector and a value vector which are obtained by operation into a table tennis block which is used for a first storage unit and is correspondingly arranged in a corresponding storage and calculation function fusion unit through a transposition mode and a conventional mode, and not writing the key vector and the value vector back to a buffer area or a memory like a related accelerator, so that memory redundancy is reduced, and the memory efficiency is improved. And thus, an attention score vector and an attention output result can be obtained by directly utilizing the key vector and the value vector which are transferred in the table tennis block of the corresponding storage and calculation function fusion unit in the subsequent calculation operation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method and device for fusing data storage and computing functions based on a ping-pong pipeline, a model accelerator for a visual converter, and an acceleration system. Background Art

[0002] The Vision Transformer (ViT) model is a tool for processing image recognition tasks based on deep learning models. To improve the computational efficiency of the ViT model when processing visual tasks such as images and videos, the ViT model has high computational complexity, involving a large number of matrix multiplications and data storage and data operations in the attention mechanism. Current processing methods, such as the fusion of data storage and computing functions, are based on an architecture that independently distinguishes between storage units and computing units. During the entire data processing process, data transmission must pass through multiple nodes, with complex and long paths, resulting in various delays. This delay in data transmission creates a bottleneck in data processing, reducing the timeliness and efficiency of data processing, and cannot meet the current real-time and large-scale deployment requirements of deep learning. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide a data storage and computing function fusion method and device based on the ping-pong pipeline, a visual converter model accelerator and an acceleration system to solve the data processing bottleneck caused by delays in the relevant data transmission process, reduce the timeliness and processing efficiency of data processing, and cannot meet the current real-time and large-scale deployment requirements of deep learning.

[0004] According to a first aspect of the present application, a data storage and computing function fusion method based on a ping-pong pipeline is provided, which is applied to a visual converter model accelerator, including a first storage and computing function fusion component and a second storage and computing function fusion component, the first storage and computing function fusion component including a first storage and computing function fusion unit and a second storage and computing function fusion unit that can process operations and storage in parallel, the second storage and computing function fusion component including a third storage and computing function fusion unit and a fourth storage and computing function fusion unit that can process operations and storage in parallel; the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively include a pong block for the first storage unit, a ping block for the second storage unit, and a multiplier-accumulator shared by the pong block for the first storage unit and the ping block for the second storage unit.

[0005] The method includes: inputting input data into the first storage and computing function fusion unit and the second storage and computing function fusion unit, using a multiply-accumulator to perform operation processing on the input data and the key weights in the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain a key vector; at the same time, updating the key vector in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to the transposition update mode; inputting a query vector into the first storage and computing function fusion component, and performing operation processing on the query vector and the key vector after transposition in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain an attention score vector ; Input the input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and use a multiplier-accumulator to perform operation processing on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain a value vector; at the same time, update the value vector in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to the conventional update mode; input the attention score vector into the second storage and computing function fusion component, and perform operation processing on the attention score vector and the value vector in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain an attention output result.

[0006] As an implementation method, the first storage and computing function fusion unit includes a first pong block and a first ping-pong block; the second storage and computing function fusion unit includes a second pong block and a second ping-pong block; the third storage and computing function fusion unit includes a third pong block and a third ping-pong block; and the fourth storage and computing function fusion unit includes a fourth pong block and a fourth ping-pong block.

[0007] As another implementation manner, the method further includes: updating the key weights to the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to a conventional update manner.

[0008] As another implementation manner, the method further includes: updating the value weights to the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to a conventional updating manner.

[0009] As another implementation method, the visual converter model accelerator also includes a third storage and computing function fusion component, which includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate a query vector based on the query weight stored in the fifth storage and computing function fusion unit.

[0010] As another implementation, the method further includes: inputting the attention score vector obtained in the first storage function fusion component into the arithmetic processing unit to perform a scaling operation on the attention score vector, and using the Softmax function to convert the attention score vector after the scaling operation into a probability distribution. According to the second aspect of the present application, a data storage function fusion device based on a ping-pong pipeline is provided, which is applied to a model accelerator for a visual converter, including a first storage function fusion component and a second storage function fusion component, the first storage function fusion component including a first storage function fusion unit and a second storage function fusion unit that can process operations and storage in parallel, and the second storage function fusion component including a third storage function fusion unit and a fourth storage function fusion unit that can process operations and storage in parallel; the first storage function fusion unit to the fourth storage function fusion unit respectively include a pong block for the first storage unit, a ping block for the second storage unit, and a multiplication and accumulation device shared by the pong block for the first storage unit and the ping block for the second storage unit; the device includes:

[0011] The vector storage and computing function fusion module is used to input input data into the first storage and computing function fusion unit and the second storage and computing function fusion unit, and use a multiplier-accumulator to perform calculations on the input data and the key weights in the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain a key vector; at the same time, the key vector is updated in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to the transpose update mode.

[0012] The attention score determination module is used to input the query vector into the first storage and computing function fusion component, and perform calculations on the query vector and the key vector transposed in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain an attention score vector.

[0013] The vector storage and computing function fusion module is also used to input input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and use a multiplier-accumulator to perform calculations on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain a value vector; at the same time, the value vector is updated in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to the conventional update mode.

[0014] The output module is used to input the attention score vector into the second storage and computing function fusion component, and perform calculations on the attention score vector and the value vectors in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain the attention output result.

[0015] According to a third aspect of the present application, a visual converter model accelerator is provided, comprising a first storage and computing function fusion component and a second storage and computing function fusion component, the first storage and computing function fusion component comprising a first storage and computing function fusion unit and a second storage and computing function fusion unit capable of processing operations and storage in parallel, the second storage and computing function fusion component comprising a third storage and computing function fusion unit and a fourth storage and computing function fusion unit capable of processing operations and storage in parallel; the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively comprise a pong block for the first storage unit, a ping block for the second storage unit, and a multiplier-accumulator shared by the pong block for the first storage unit and the ping block for the second storage unit, the visual converter model accelerator is configured to execute a data storage and computing function fusion method based on the ping-pong pipeline as any implementation of the first aspect described above.

[0016] In one implementation, the first storage and computing function fusion unit includes a first pong block and a first ping-pong block; the second storage and computing function fusion unit includes a second pong block and a second ping-pong block; the third storage and computing function fusion unit includes a third pong block and a third ping-pong block; the fourth storage and computing function fusion unit includes a fourth pong block and a fourth ping-pong block; the visual converter model accelerator also includes a third storage and computing function fusion component, and the third storage and computing function fusion component includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate the query vector based on the query weight stored in the fifth storage and computing function fusion unit; the visual converter model accelerator is configured to execute the data storage and computing function fusion method based on the ping-pong pipeline as any implementation of the first aspect.

[0017] According to the fourth aspect of the present application, a visual converter model acceleration system is provided, comprising a visual converter model accelerator as in any implementation of the third aspect, configured to execute a data storage and computing function fusion method based on a ping-pong pipeline as in any implementation of the first aspect.

[0018] With the help of the above technical solution, an embodiment of the present application provides a data storage and computing function fusion method based on the ping-pong pipeline, in which the key vector and value vector obtained by calculation are updated to the pong block corresponding to the corresponding storage and computing function fusion unit through the transposition mode and the normal mode respectively. Unlike the related accelerator, the key vector and value vector will no longer be written back to the buffer or memory, reducing memory redundancy, so that in subsequent calculation operations, the transposed key vector and value vector in the pong block of the corresponding storage and computing function fusion unit can be directly used to obtain the attention score vector and attention output result.

[0019] Through this implementation, a storage-computation fusion component is used to integrate data storage and data operations for the same type of data in the attention mechanism into a single component. This allows different types of data operations to be performed near the data storage location, shortening data transmission links and reducing redundant memory accesses. Furthermore, through ping-pong pipeline scheduling, data operations and data storage are processed simultaneously in parallel within the ping and pong blocks of the same storage-computation fusion unit. By calling different storage-computation fusion components and scheduling their results, the data flow scheduling of the visual transformer model is optimized, reducing the resources required for data storage and data operations, and improving the timeliness and efficiency of data processing.

[0020] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 A schematic diagram of the structure of a visual converter model accelerator provided in an embodiment of the present application;

[0023] Figure 2 A flowchart of a method for integrating data storage and computing functions based on a ping-pong pipeline provided in an embodiment of the present application;

[0024] Figure 3 A schematic diagram of a data storage and computing function fusion device based on a ping-pong pipeline provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0026] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0027] Before introducing in detail the data storage and computing function integration method based on the ping-pong pipeline provided in the embodiment of the present application, a brief introduction to the application scenarios and implementation architecture involved in the embodiment of the present application is first given.

[0028] First, the application scenarios of this application are introduced as follows.

[0029] With the widespread application of visual transformer models in computer vision, the computational complexity and memory access requirements of the ViT model have placed high demands on hardware architecture. Related hardware accelerators have not adequately addressed the redundant data access and computational complexity issues in ViT algorithms, particularly during matrix operations. Related acceleration methods often fail to effectively reduce redundant memory access, resulting in low hardware utilization. Therefore, a new architecture design and acceleration approach are urgently needed to optimize data flow scheduling and improve hardware resource utilization.

[0030] Based on this, the present application proposes a data storage and computing function fusion method based on ping-pong pipeline, which uses a storage and computing function fusion component to integrate the data storage and data operation of the same type of data in the attention mechanism into one component for execution, so that different types of data operations are performed near the data storage location, shortening the data transmission link and reducing redundant memory access. Moreover, through the ping-pong pipeline scheduling method, the ping-pong block of the same storage and computing function fusion unit performs parallel alternating processing of data processing and data loading. By calling different storage and computing function fusion components, the results generated by different storage and computing function fusion components are scheduled, the data flow scheduling of the visual converter model is optimized, the resources required for data storage and data operation are reduced, and the timeliness and processing efficiency of data processing are improved.

[0031] Next, the implementation framework of this application is introduced as follows.

[0032] The structural diagram of the visual converter model accelerator of this application is as follows Figure 1As shown, it includes a first storage and computing function fusion component 11, a second storage and computing function fusion component 12, a third storage and computing function fusion component 13 and a control bus. The first storage and computing function fusion component 11, the second storage and computing function fusion component 12, the third storage and computing function fusion component 13 and the control bus are connected through communication. Among them, the first storage and computing function fusion component 11 includes a first storage and computing function fusion unit 111 and a second storage and computing function fusion unit 112 that can process calculations and storage in parallel. The first storage and computing function fusion unit 111 includes a first pong block 1111 for the first storage unit and a first ping block 1112 for the second storage unit; the second storage and computing function fusion unit 112 includes a second pong block 1121 for the first storage unit and a second ping block 1122 for the second storage unit.

[0033] It can be understood that the pong block for the first storage unit and the ping block for the second storage unit share a multiplication and accumulation device.

[0034] The second storage and computing function fusion component 12 includes a third storage and computing function fusion unit 121 and a fourth storage and computing function fusion unit 122. The third storage and computing function fusion unit 121 includes a third pong block 1211 and a third ping block 1212; the fourth storage and computing function fusion unit 122 includes a fourth pong block 1221 and a fourth ping block 1222.

[0035] The third storage and computing function fusion component 13 includes a fifth storage and computing function fusion unit 131 and an arithmetic processing unit 132; the fifth storage and computing function fusion unit 131 is used to generate a query vector based on the query weight stored in the fifth storage and computing function fusion unit.

[0036] The control bus includes an input bus 141 , a weight bus 142 , and a result bus 143 .

[0037] In the embodiments of the present application, the above-mentioned computing and storage function fusion components adopt computing in memory (CIM) technology, also known as in-memory computing technology, which is a computing architecture that integrates computing and storage functions. By integrating the computing unit with the storage unit, a tightly coupled structure is formed. For example, the processing unit is embedded in the memory or the storage unit is embedded in the processing unit. The computing and storage function fusion technology can perform computing operations directly in the memory, reduce frequent data transmission, and improve computing efficiency.

[0038] The aforementioned storage-computing fusion component can be a processing unit designed using storage-computing fusion technology. It can be a chip, integrated circuit, microprocessor, or other device. For example, it can integrate computing and storage functions on the same chip to improve data processing efficiency and reduce energy consumption. It can embed computing power within the storage array, achieving a close integration of data storage and computing, thereby increasing memory access bandwidth and reducing memory access power consumption.

[0039] The storage-computing fusion technology within the aforementioned storage-computing fusion component is the core of the accelerator. It integrates data storage and computing functions within the same physical unit, reducing the frequent data transfer between the storage and computing units, thereby reducing data transmission latency and power consumption. In ViT accelerators, non-volatile memories (such as RRAM and PCRAM) are typically used to implement the storage-computing fusion function. These memories can not only store model parameters and input data, but also perform computational operations such as matrix multiplication directly within the storage unit.

[0040] In some embodiments, the storage and computing function fusion component may include a computing subunit and a storage medium. Among them, the computing subunit is also called a controller, a control subunit, etc., which is responsible for executing computing tasks; the storage medium is used to store data. The storage media involved in the storage and computing function fusion component may include: static random access memory (SRAM), dynamic random access memory (DRAM), resistive random access memory (RRAM), magnetic random access memory (MRAM), etc. By tightly integrating the computing subunit with the storage medium, the fusion of data storage and computing can be achieved, thereby improving data processing efficiency and reducing energy consumption.

[0041] The storage and computing function fusion component can be applied to artificial intelligence (AI) algorithms to provide operator acceleration of vector-matrix multiplication for AI algorithms, including applications in AI algorithms such as convolutional neural networks (CNN) and recurrent neural networks (RNN). For example, the storage and computing function fusion component can provide a computing power of more than 1000TOPS and a high energy efficiency of more than 10 to 100TOPS / W in the field of AI algorithms. For ease of understanding, the data storage and computing function fusion method based on the ping-pong pipeline provided in this application is specifically introduced below with reference to the accompanying drawings. The data storage and computing function fusion method based on the ping-pong pipeline can be applied to the above-mentioned vision converter model accelerator.

[0042] Figure 2 is a flow chart showing a method for integrating data storage and computing functions based on a ping-pong pipeline according to an exemplary embodiment. Figure 2 As shown, the method includes the following steps.

[0043] S21, input the input data into the first storage and computing function fusion unit and the second storage and computing function fusion unit, and use a multiplier-accumulator to perform operation processing on the input data and the key weights in the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain a key vector; at the same time, the key vector is updated in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to the transpose update mode.

[0044] In some embodiments, according to a conventional update mode, the key weights are loaded or updated into the first ping block and the second ping block based on the weight bus, so that the key weights are directly called during the calculation process.

[0045] S22, input the query vector into the first storage and computing function fusion component, and perform operation processing on the query vector and the key vector transposed in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain an attention score vector.

[0046] S23, input the input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and use a multiplier-accumulator to perform operation processing on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain a value vector; at the same time, the value vector is updated in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to the conventional update mode.

[0047] As an implementation mode, the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively include pong blocks for the first storage unit and ping blocks for the second storage unit. Specifically, it can be understood that the first storage and computing function fusion unit includes the first pong block and the first ping block; the second storage and computing function fusion unit includes the second pong block and the second ping block; the third storage and computing function fusion unit includes the third pong block and the third ping block; and the fourth storage and computing function fusion unit includes the fourth pong block and the fourth ping block.

[0048] As an embodiment, the method further includes: updating the key weights to the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to a conventional updating method.

[0049] As another embodiment, the method further includes: updating the value weights to the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to a conventional updating method.

[0050] The key vector and value vector obtained by the operation are updated to the pong block corresponding to the corresponding storage and computing function fusion unit through the transposition mode and the normal mode respectively. Unlike the related accelerator, the key vector and value vector are no longer written back to the buffer or memory, reducing memory redundancy, so that the transposed key vector and value vector in the pong block of the corresponding storage and computing function fusion unit can be directly used in subsequent operation processes to obtain the attention score vector and attention output result.

[0051] In some embodiments, the value weights are loaded or updated into the third ping block and the fourth ping block according to a conventional update mode, so that the value weights can be directly called during the calculation process.

[0052] In some embodiments, the value vector output by the ping block can be updated to the corresponding pong block in transpose mode and normal mode respectively, and while updating and storing, it can also participate in the calculation in the computing block to achieve storage and calculation function fusion processing.

[0053] As a specific implementation method, in order to ensure that the first storage and computing function fusion component and the second storage and computing function fusion component can run in parallel, the key weights are loaded into the first ping-pong block and the second ping-pong block, and the key weights in the first ping-pong block and the second ping-pong block are updated according to the conventional update mode; at the same time, according to the conventional update mode, the value weights are loaded into the third ping-pong block and the fourth ping-pong block, and the value weights in the third ping-pong block and the fourth ping-pong block are updated.

[0054] S24, input the attention score vector into the second storage and computing function fusion component, and perform calculations on the attention score vector and the value vectors in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain the attention output result.

[0055] In order to reduce the consumption of transportation resources for data operations between vectors, the query vector is directly input into the first storage and computing function fusion component, so that the query vector is directly processed with the key vector in the first storage and computing function fusion component, thereby improving the computational efficiency of the attention score vector.

[0056] In some embodiments, in order to reduce memory access, all input data are input through a unified input bus, and all weight adjustment data are adjusted through a unified weight bus and input into the corresponding storage and computing function fusion unit, and all output results are output accordingly through a unified output bus.

[0057] Specifically, all input data is input into the input bus, so that all inputs are input into each component unit (such as the first storage and computing function fusion unit and the second storage and computing function fusion unit) through the same input bus. At the same time, all weight adjustment information used for weight modification or dynamic weight adjustment and dynamically set weights (such as query weights used to adjust weights) are uniformly set through the weight bus.

[0058] Through the above implementation, a storage and computation function fusion component is used to integrate data storage and data operations for the same type of data in the attention mechanism into a single component for execution. This allows different types of data operations to be performed near the data storage location, shortening the data transmission link and reducing redundant memory access. Furthermore, through ping-pong pipeline scheduling, data processing and data loading are performed in parallel within the same storage and computation function fusion component. By calling different storage and computation function fusion components and scheduling the results generated by these components, the data flow scheduling of the visual converter model is optimized, reducing the resources required for data storage and data operations, and improving the timeliness and efficiency of data processing.

[0059] The Vision Transformer model uses query, key, and value matrices when performing image processing. These matrices are derived from the input embedding representation and are used to compute attention scores and weighted outputs. For example, the query matrix can be generated by multiplying the input embedding with a learnable weight matrix. The key matrix can be generated by multiplying the input embedding with another learnable weight matrix. The value matrix is ​​generated by multiplying the input embedding with a third learnable weight matrix.

[0060] Matrices such as query, key, and value can also be represented by vectors, which are called "query vector, key vector, and value vector" respectively.

[0061] As a key vector generation method, the above S21 is specifically implemented through the following steps: first, the input data is quantized into an embedding vector to reduce the number of bits representing the input data, thereby reducing the resources required for storage and calculation, while maintaining the calculation accuracy to a certain extent; then, the embedding vector is further input into the first ping-pong block and the second ping-pong block in parallel, so that the embedding vector and the key weight are multiplied and added to obtain the key vector; and, according to the transpose update mode, the key vector is stored in the first ping-pong block and the second ping-pong block to obtain the transposed key vector.

[0062] The computing blocks and pong blocks of the storage and computing function fusion units in the above-mentioned storage and computing function fusion components are connected to the multiplier and adder tree so that the pong blocks and ping blocks share the multiplier and adder tree, thereby making the pong blocks and ping blocks as storage space have computing functions.

[0063] In the above-mentioned key vector implementation, the basic scheduling unit of the ping-pong pipeline is the storage-and-calculation unit. Within this unit, the pong block and the ping block share a MAC (Multiply Accumulator) and adder tree. When the ping block performs a calculation, the pong block stores its results, implementing an update operation during the calculation. During the next round of QKT, the pong block's data is calculated with the input through MAC. The calculation is performed by MACing the input Q and the KT in the pong block. The above-mentioned ping-pong pipeline utilizes the storage-and-calculation unit to perform data storage and calculation operations in parallel, reducing data latency and improving data processing efficiency.

[0064] As a method for generating a value vector, the above S23 is specifically implemented through the following steps: first, the embedding vector is input in parallel into the third ping-pong block and the fourth ping-pong block, and a multiplication and accumulation operation is performed on the embedding vector and the value weight to obtain a value vector; according to the normal update mode, the value vector is stored in the third ping-pong block and the fourth ping-pong block.

[0065] The ping-pong pipeline scheduling described above allows embedding vectors to be fed into the third and fourth blocks in parallel, allowing the two blocks to operate in parallel. While the third block is performing a multiplication-accumulation operation on an embedding vector and a value weight, the fourth block can simultaneously perform the same operation on the next set of embedding vectors, and vice versa. This parallel computing approach significantly reduces the overall attention computation time and improves the accelerator's computational throughput.

[0066] For example, in related sequential computations, completing a complete attention calculation can take a long time. However, with ping-pong pipeline scheduling, the parallel operation of the two ping-pong blocks can roughly halve the computation time, significantly increasing the accelerator's image data processing speed. Simultaneously, the resulting value vector is stored in the third and fourth pong blocks according to a conventional update pattern, effectively utilizing storage resources. The two pong blocks can store value vectors in parallel, avoiding waste of storage resources. For example, while the third pong block is storing the current value vector, the fourth pong block can read data or perform other operations, improving the efficiency of storage resource utilization. Furthermore, this storage method facilitates subsequent data processing and access. When performing subsequent attention score vector calculations or other operations, the value vector can be read directly from the corresponding pong block, reducing data search and transmission time.

[0067] As an embodiment, the visual converter model accelerator also includes a third storage and computing function fusion component, which includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate a query vector based on the query weight stored in the fifth storage and computing function fusion unit.

[0068] As another embodiment, the method also includes: inputting the attention score vector obtained in the first storage and calculation function fusion component into the arithmetic processing unit to perform a scaling operation on the attention score vector, and using the Softmax function to convert the attention score vector after the scaling operation into a probability distribution.

[0069] In the above embodiment, the weights in other storage and computing function fusion units are updated through the fifth storage and computing function fusion unit, which enables the system to flexibly adjust the computing parameters according to different task requirements or data characteristics. Compared with the fixed weight computing architecture, this dynamic update capability enhances the adaptability and flexibility of the system and can better handle diverse computing tasks.

[0070] Based on this implementation, in the fifth storage and computing function fusion unit, a multiplication and accumulation operation is performed on the input data and the query weight to obtain a query vector; and the attention score matrix obtained in the first storage and computing function fusion component is input into the arithmetic processing unit.

[0071] In some embodiments, the fifth storage and computing function fusion unit is also called a dynamically schedulable storage and computing function fusion macro unit.

[0072] Furthermore, the specific process of determining the attention score vector includes the following steps. First, the query vector is input into the first storage and calculation function fusion unit and the second storage and calculation function fusion unit to multiply the query vector with the transposed key vector to obtain an attention score matrix representing the similarity between the query vector and the key vector; in the arithmetic processing unit, a scaling operation is performed on the attention score matrix to obtain a scaling operation result, and the scaling operation result is converted into a probability distribution using the Softmax function to obtain the attention score vector.

[0073] The aforementioned ping-pong pipeline-based integrated storage and computation Visual Translator (ViT) accelerator combines integrated storage and computation technology with a ping-pong pipeline scheduling strategy to efficiently handle the complex computational tasks in the ViT model. The following describes the workflow in detail.

[0074] First, initialization and preloading.

[0075] Initialize the CIM macro and preload the weight data into the corresponding ping blocks of the storage and computation fusion units (e.g., the first and second storage and computation fusion units). Using a "ping block" and "pong block" storage method, data is processed and stored synchronously, achieving efficient data management and fast access. During the calculation process, data can be flexibly switched and updated between different pong blocks, ensuring continuity and efficiency.

[0076] Specifically, two ping-pong storage and computing function fusion macro units (the first storage and computing function fusion unit and the second storage and computing function fusion unit) preload Wk into the ping block in a conventional manner, while the other two ping-pong storage and computing function fusion macro units (the third storage and computing function fusion unit and the fourth storage and computing function fusion unit) preload the value weight Wv into the ping block in a conventional manner. Among them, the dynamically schedulable storage and computing function fusion macro unit can preload the query weight Wq in advance.

[0077] Second, update the weights and perform multiplication and accumulation.

[0078] Update the weights in the storage and computation function fusion unit, and then use the activation value X as input data and the key weight Wk and value weight Wv in the computation block (ping block) to perform multiplication and accumulation operations to generate the key vector K and value vector V.

[0079] Specifically, after the first and second storage-computing function fusion units update the key weight Wk, and the third and fourth storage-computing function fusion units update the value weight Wv, the input data X[X0…Xdim-1] is activated and used as the input of the above four storage-computing function fusion unit macros, and multiplication and accumulation operations are performed with the key weight Wk and value weight Wv in the calculation block (ping block) to generate a key vector [K0…Kdim-1] and a value vector [V0…Vdim-1]. After the key vectors [K0…Kdim-1] are quantized, they no longer need to be written back to the buffer or memory as in the accelerator in the related art.

[0080] Third, quantify and store the results.

[0081] The key vector K and the value vector V are quantized and then stored in the storage block (pong block) of the storage and computation function fusion unit in transposed update mode and regular update mode respectively.

[0082] Specifically, the key vector K and the value vector V are stored in the pong blocks of the first and second storage and computing function fusion units in transposed update mode, serving as weights for the next stage of pipeline operations. Similarly, after the value vectors [V0…Vdim-1] are quantized, they are stored in the storage blocks of the third and fourth storage and computing function fusion units in normal update mode, serving as weights for subsequent pipeline operations.

[0083] Fourth, generate a query vector.

[0084] The activation value X is multiplied by the query weight in the dynamically schedulable storage and computing function fusion macro unit to generate the query vector Q.

[0085] Specifically, the activation value X[X0…Xdim-1] is multiplied by the query weight Wq in the dynamically schedulable storage and computing function fusion macro unit to generate a query vector Q[Q0…Qdim-1].

[0086] Fifth, calculate the scoring matrix.

[0087] The query vector Q is fed as input to the first storage and computing function fusion unit and the second storage and computing function fusion unit, and multiplied with the transpose of the key vector K stored in the storage block (pong block) (here it is simply assumed to be K) to generate the score matrix A.

[0088] Specifically, the query vector Q is directly fed as input into the first storage-computation function fusion unit and the second storage-computation function fusion unit, and multiplied with the transposed key vector KT stored in the storage block (pong-block) to generate the score matrix A[A0...An-1].

[0089] Sixth, perform Scale&Softmax operation.

[0090] The score matrix A performs Scale & Softmax operations in the Arithmetic Processing Unit (APU).

[0091] Seventh, calculation of the final result.

[0092] The score matrix A is used as the input of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and is multiplied with the value vector V in the storage block (pong block). The result is quantized and sent to the inter-core accumulator to generate the final result, such as SA[SA0...SAdim-1].

[0093] In some embodiments, ViT divides the image into multiple small blocks and treats these small blocks as sequences to be input into the Transformer architecture for processing, with the attention mechanism as a core component. The storage and computing function fusion technology aims to closely integrate storage and computing units to reduce the data transmission overhead between storage and computing modules. Ping-pong pipeline scheduling is a strategy to improve computing efficiency. By using two sets of storage units in parallel, data reading, computing, and writing operations can be performed in parallel, thereby hiding the delay of data transmission.

[0094] When performing image processing, the visual transformer model ViT can divide the image into multiple small patches (patches), and then linearly embed these small patches into a dimensional space to form a 1D vector sequence. That is, each block can be serialized into a vector, mapped to a smaller dimension through a single matrix multiplication, and the input image can be decomposed into a series of blocks, so that the vector sequence corresponding to the image segmentation result is compatible with the Transformer architecture. Feature extraction is then performed based on the vector sequence to identify feature targets in the image, that is, the self-attention mechanism is used to capture local and global relationships within the image, and the Transformer encoder is used to process these vector embeddings, so that ViT can identify patterns in images based on datasets such as ImageNet in image classification tasks by capturing contextual and hierarchical features.

[0095] Based on this, in ViT, the attention mechanism determines the importance of features at different positions by calculating the similarity between the query vector Q (hereinafter referred to as query or Q), the key vector K (hereinafter referred to as key or K), and the value vector V (hereinafter referred to as value or V). The specific calculation steps are as follows.

[0096] First, calculate the similarity: given the query vector Q and key vector K, calculate the similarity score between them through the dot product operation.

[0097] In the attention calculations of the ViT (Vision Transformer) accelerator, which integrates storage and computation functions based on ping-pong pipeline scheduling, scaling is used to alleviate the problem of excessively large dot product results. In the attention mechanism, the dot product operation between the query and the key may produce a large value, which will cause the input value of the Softmax function to differ too much, resulting in the gradient of the Softmax function becoming extremely small, affecting the training effect of the model. Therefore, it is necessary to scale the dot product result. The specific scaling operation is shown in the following formula.

[0098]

[0099] Here, dk is the dimension of the key vector.

[0100] Second, calculate the attention score vector. Apply the softmax function to the similarity score to convert the score into a probability distribution, namely the attention score vector.

[0101] Third, weighted summation: Multiply the attention score vector with the value vector and sum them to get the final attention output.

[0102] The specific process of the above attention calculation based on ping-pong pipeline scheduling is as follows.

[0103] First, the data loading stage. Ping-Pong storage structure: Use two groups of storage units (such as SRAM), which are respectively recorded as storage unit A and storage unit B of the first storage and computing function fusion unit. When one group of storage units is reading and calculating data, the other group of storage units can simultaneously perform data writing operations to realize parallel data processing. Data prefetching: Before starting the calculation, the query matrix Q, key matrix K, and value matrix V are prefetched from the external memory (such as DRAM) into the storage unit. For example, first load a portion of Q, K, and V data into storage unit A.

[0104] Second, similarity calculation stage. Pipeline operation: After the data in storage unit A is ready, similarity calculation (i.e. QK T At the same time, the next batch of Q, K, and V data is loaded into memory unit B. In the integrated storage and computing accelerator, in-memory computing technology is used to perform matrix multiplication directly within the memory unit, reducing data movement. Block computing: To reduce computational complexity and storage requirements, matrices are typically divided into blocks for calculation. For example, the Q and K matrices can be divided into multiple small blocks, and the similarity score of each block is calculated sequentially.

[0105] Third, the attention score vector calculation phase. Parallel computing: After the similarity calculation is complete, a softmax operation is performed on the obtained similarity scores to obtain the attention score vector. Due to the parallel computing capabilities of the integrated storage and computing accelerator, softmax calculations can be performed on multiple similarity scores simultaneously. Ping-pong switching: After the similarity calculation and attention score vector calculation are completed for the data in storage unit A, the data is switched to storage unit B and the same operation is performed, while the new data is loaded into storage unit A.

[0106] Fourth, the weighted summation stage. Data reuse: Using the calculated attention score vector, perform a weighted summation on the value matrix V. This process fully utilizes the integrated storage and computation features to reduce data read and write times. Output: The weighted summation result is used as the final attention output and stored in the designated storage unit.

[0107] Through the above implementation, on the one hand, the waiting time for data transmission is reduced through ping-pong pipeline scheduling, so that the computing unit can continue to perform operations, thereby improving the overall computing efficiency of the accelerator. On the other hand, the storage and computing function integration technology reduces the frequent movement of data between the storage and computing modules, reduces the energy consumption of data transmission, and thus improves the energy efficiency ratio. Therefore, the attention calculation process of the ViT accelerator with integrated storage and computing functions based on ping-pong pipeline scheduling improves computing efficiency and energy efficiency through reasonable scheduling and storage and computing function integration technology, providing strong support for the efficient implementation of visual transformers.

[0108] In order to realize the above functions, the data storage and computing function fusion device includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy for those skilled in the art to realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. The embodiments disclosed herein also provide a method such as Figure 3The data storage and computing function fusion device based on the ping-pong pipeline shown is applied to a visual converter model accelerator. The visual converter model accelerator includes a first storage and computing function fusion component and a second storage and computing function fusion component. The first storage and computing function fusion component includes a first storage and computing function fusion unit and a second storage and computing function fusion unit that can process operations and storage in parallel. The second storage and computing function fusion component includes a third storage and computing function fusion unit and a fourth storage and computing function fusion unit that can process operations and storage in parallel. The first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively include a pong block for the first storage unit, a ping block for the second storage unit, and a multiplication and accumulation device shared by the pong block for the first storage unit and the ping block for the second storage unit. The device includes: a vector storage and computing function fusion module 301, an attention score determination module 302, and an output module 303.

[0109] The vector storage and computing function fusion module 301 is used to input input data into the first storage and computing function fusion unit and the second storage and computing function fusion unit, and use a multiplier-accumulator to perform calculations on the input data and the key weights in the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain a key vector; at the same time, the key vector is updated in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to the transpose update mode.

[0110] The attention score determination module 302 is used to input the query vector into the first storage and computing function fusion component, and perform calculations on the query vector and the key vector transposed in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit to obtain an attention score vector.

[0111] The vector storage and computing function fusion module 301 is also used to input input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and use a multiplier-accumulator to perform calculations on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain a value vector; at the same time, the value vector is updated in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to the conventional update mode.

[0112] The output module 303 is used to input the attention score vector into the second storage and computing function fusion component, and perform calculations on the attention score vector and the value vectors in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain the attention output result.

[0113] As an implementation mode, the first storage and computing function fusion unit includes a first pong block and a first ping-pong block; the second storage and computing function fusion unit includes a second pong block and a second ping-pong block; the third storage and computing function fusion unit includes a third pong block and a third ping-pong block; and the fourth storage and computing function fusion unit includes a fourth pong block and a fourth ping-pong block.

[0114] As another embodiment, the device is further used to: update the key weights to the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to a conventional updating method.

[0115] As another embodiment, the device is further configured to update the value weights to the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to a conventional updating method.

[0116] As another embodiment, the visual converter model accelerator also includes a third storage and computing function fusion component, which includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate a query vector based on the query weight stored in the fifth storage and computing function fusion unit.

[0117] As another embodiment, the device is also used to: input the attention score vector obtained in the first storage and calculation function fusion component into the arithmetic processing unit to perform a scaling operation on the attention score vector, and use the Softmax function to convert the attention score vector after the scaling operation into a probability distribution.

[0118] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data storage and computing function integration method based on ping-pong pipeline, characterized in that: The invention is applied to a model accelerator for a visual converter, comprising a first storage and computing function fusion component and a second storage and computing function fusion component, wherein the first storage and computing function fusion component comprises a first storage and computing function fusion unit and a second storage and computing function fusion unit capable of processing operations and storage in parallel, and the second storage and computing function fusion component comprises a third storage and computing function fusion unit and a fourth storage and computing function fusion unit capable of processing operations and storage in parallel; the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively comprise a pong block for a first storage unit, a ping block for a second storage unit, and a multiplication and accumulation device shared by the pong block for the first storage unit and the ping block for the second storage unit; the method comprises: Inputting input data into the first storage and computing function fusion unit and the second storage and computing function fusion unit, performing arithmetic processing on the input data and the key weights in the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit using the multiplication and accumulation unit to obtain a key vector; and simultaneously updating the key vector in the pong blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit according to a transposition update mode; Inputting a query vector into the first storage and computation function fusion component, and performing computation on the query vector and the key vector after transposition in the pong blocks of the first storage and computation function fusion unit and the second storage and computation function fusion unit to obtain an attention score vector; Inputting the input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, performing arithmetic processing on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit using the multiplication and accumulation unit to obtain a value vector; and simultaneously updating the value vector in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to a conventional update mode; The attention score vector is input into the second storage and computing function fusion component, and the attention score vector is processed with the value vector in the pong block of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain an attention output result.

2. The method according to claim 1, characterized in that The first storage and computing function fusion unit includes a first pong block and a first ping-pong block; the second storage and computing function fusion unit includes a second pong block and a second ping-pong block; the third storage and computing function fusion unit includes a third pong block and a third ping-pong block; the fourth storage and computing function fusion unit includes a fourth pong block and a fourth ping-pong block.

3. The method according to claim 2, characterized in that The method further comprises: The key weight is updated to the ping blocks of the first storage and computing function fusion unit and the second storage and computing function fusion unit in a conventional updating manner.

4. The method according to claim 3, characterized in that The method further comprises: The value weight is updated to the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit in a conventional updating manner.

5. The method according to any one of claims 1 to 4, characterized in that The visual converter model accelerator also includes a third storage and computing function fusion component, which includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate the query vector based on the query weight stored in the fifth storage and computing function fusion unit.

6. The method according to claim 5, characterized in that The method further comprises: The attention score vector obtained in the first storage and calculation function fusion component is input into the arithmetic processing unit to perform a scaling operation on the attention score vector, and the Softmax function is used to convert the attention score vector after the scaling operation into a probability distribution.

7. A data storage and computing function fusion device based on ping-pong pipeline, characterized in that: The invention is applied to a model accelerator for a visual converter, comprising a first storage and computing function fusion component and a second storage and computing function fusion component, wherein the first storage and computing function fusion component comprises a first storage and computing function fusion unit and a second storage and computing function fusion unit capable of processing operations and storage in parallel, and the second storage and computing function fusion component comprises a third storage and computing function fusion unit and a fourth storage and computing function fusion unit capable of processing operations and storage in parallel; the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively comprise a pong block for a first storage unit, a ping block for a second storage unit, and a multiplication and accumulation device shared by the pong block for the first storage unit and the ping block for the second storage unit; the device comprises: a vector storage and computation function fusion module, configured to input input data into the first storage and computation function fusion unit and the second storage and computation function fusion unit, perform computational processing on the input data and the key weights in the ping blocks of the first storage and computation function fusion unit and the second storage and computation function fusion unit using the multiplication and accumulation unit to obtain a key vector; and simultaneously update the key vector in the pong blocks of the first storage and computation function fusion unit and the second storage and computation function fusion unit according to a transposition update mode; an attention score determination module, configured to input a query vector into the first storage and computation function fusion component, and perform computation on the query vector and the key vector transposed in the pong blocks of the first storage and computation function fusion unit and the second storage and computation function fusion unit to obtain an attention score vector; The vector storage and computing function fusion module is further configured to input the input data into the third storage and computing function fusion unit and the fourth storage and computing function fusion unit, and use the multiplication and accumulator to perform operation processing on the input data and the value weights in the ping blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain a value vector; and simultaneously, update the value vector in the pong blocks of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit according to a conventional update mode; An output module is used to input the attention score vector into the second storage and computing function fusion component, and to perform calculations on the attention score vector and the value vector in the pong block of the third storage and computing function fusion unit and the fourth storage and computing function fusion unit to obtain an attention output result.

8. A visual converter model accelerator, characterized in that: It includes a first storage and computing function fusion component and a second storage and computing function fusion component, the first storage and computing function fusion component includes a first storage and computing function fusion unit and a second storage and computing function fusion unit that can process operations and storage in parallel, the second storage and computing function fusion component includes a third storage and computing function fusion unit and a fourth storage and computing function fusion unit that can process operations and storage in parallel; the first storage and computing function fusion unit to the fourth storage and computing function fusion unit respectively include a pong block for the first storage unit, a ping block for the second storage unit and a multiplication and accumulation unit shared by the pong block for the first storage unit and the ping block for the second storage unit, the visual converter model accelerator is configured to execute the data storage and computing function fusion method based on the ping-pong pipeline as claimed in claim 1.

9. The vision-oriented converter model accelerator according to claim 8, characterized in that: The first storage and computing function fusion unit includes a first pong block and a first ping-pong block; the second storage and computing function fusion unit includes a second pong block and a second ping-pong block; the third storage and computing function fusion unit includes a third pong block and a third ping-pong block; the fourth storage and computing function fusion unit includes a fourth pong block and a fourth ping-pong block; the vision-oriented converter model accelerator also includes a third storage and computing function fusion component, and the third storage and computing function fusion component includes a fifth storage and computing function fusion unit and an arithmetic processing unit; the fifth storage and computing function fusion unit is used to generate the query vector based on the query weight stored in the fifth storage and computing function fusion unit; the vision-oriented converter model accelerator is configured to execute the data storage and computing function fusion method based on the ping-pong pipeline as described in any one of claims 2 to 6.

10. A visual converter model acceleration system, characterized in that: It comprises the visual converter model accelerator as described in claim 8 or 9, and is configured to execute the data storage and computing function fusion method based on the ping-pong pipeline as described in any one of claims 1 to 6.