Photoelectric Transformer accelerator collaboration method and device
By chunking the matrix of the Transformer model and working together between the general computing chip and the optical computing chip, the problem of high storage bottlenecks and data movement costs in the optical computing architecture is solved, and efficient photoelectric Transformer accelerator calculation is achieved.
Patent Information
- Application Number
- CN202510262267.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
AI Technical Summary
The existing Transformer model has problems such as storage bottlenecks and high data movement costs in self-attention computing. Especially in optical computing architectures, limited SRAM capacity is difficult to meet the needs, resulting in increased energy consumption and reduced computing efficiency.
By blocking the query matrix Q, key matrix K and value matrix V, and working together between the general computing chip and the optical computing chip, the high-speed parallel computing power of the optical computing chip is used, combining nonlinear processing and block transmission, resource utilization is optimized.
The number of memory access times is reduced, the computing efficiency is improved, the data transmission delay and bandwidth pressure is reduced, and the computing time is shortened, and the efficient calculation of the photoelectric Transformer accelerator is realized.
Smart Images

Figure CN120258041A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of deep learning hardware acceleration, and in particular to a method and device for collaborative operation of an optoelectronic Transformer accelerator. Background Art
[0002] In the fields of high-performance computing, optical computing, and deep learning hardware acceleration, the storage bottleneck and data movement cost are one of the key challenges. As the scale of the Transformer model continues to grow, its self-attention calculation requires storing a large amount of intermediate activation values and attention weights. In an optical computing architecture, the limited SRAM capacity is difficult to meet this demand, and frequent data transmission will lead to increased energy consumption and reduced computing efficiency. This direct storage and computing method is obviously not efficient enough, so a better optimization strategy is needed. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a method and device for collaborative operation of an optoelectronic Transformer accelerator, which reduces the number of memory accesses and meets the demand for accurate calculation of attention under limited computing resources of optical computing.
[0004] In a first aspect, a method for collaborative operation of an optoelectronic Transformer accelerator is provided, including: based on the size M of the static random access memory (SRAM) of the optical chip on the optical computing chip, and the feature dimension d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, dividing the query matrix Q into T r sub-blocks, and dividing the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1; initializing the output matrix O, the normalization factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; interacting with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: respectively transmitting the j-th sub-blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j-th sub-blocks K j and V j, communicate with the optical computing chip in a cyclic manner to execute the following steps until all sub-blocks in the query matrix are transmitted: Transmit the i-th sub-block Q of the query matrix Q to the optical computing chip; Receive the i-th sub-block Q of the query matrix Q calculated by the optical computing chip and the local attention score matrix S between the j-th sub-block K of the key matrix K; Perform non-linear processing on the local attention score matrix S to obtain the corresponding local attention weight matrix P; Calculate the maximum value m of each row in the local attention score matrix S and the sum l of each row in the local attention weight matrix P; Update the maximum value tracking vector m based on the following formula: m = max(m, m), where m is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; Update the normalization factor vector l based on the following formula: where l is the current latest normalization factor vector l, and each value in the normalization factor vector l corresponds to a row in the local attention score matrix; Update the output matrix O based on the following formula: where O is the current latest output matrix. i Transmit to the optical computing chip; Receive the local attention score matrix S between the i-th sub-block Q of the query matrix Q calculated by the optical computing chip and the j-th sub-block K of the key matrix K i and the j-th sub-block K of the key matrix K j ij ; Perform non-linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; Calculate the maximum value m of each row in the local attention score matrix S ij ij and the sum l of each row in the local attention weight matrix P ij ij ; Update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; Update the normalization factor vector l based on the following formula: where l i is the current latest normalization factor vector l, and each value in the normalization factor vector l corresponds to a row in the local attention score matrix; Update the output matrix O based on the following formula:
[0005] where O i is the current latest output matrix.
[0006] In a possible implementation, the size of each sub-block in the query Q matrix is Br, and the size of each sub-block in the key matrix K and the V matrix is B c ,
[0007] In a possible implementation, the method further includes: After updating the maximum value tracking vector m, the normalization factor vector l, and the output matrix O each time, store the updated maximum value tracking vector m, the normalization factor vector l, and the output matrix O.
[0008] In a second aspect, an optoelectronic Transformer accelerator cooperation device is provided, including: a partitioning module, configured to partition the query matrix Q into T based on the size M of the static random access memory SRAM on the optical chip of the optical computing chip and the feature dimension d of the query matrix Q, the key matrix K, and the value matrix V in the Transformer modelr sub - blocks, and dividing the key matrix K and the value matrix V into T c sub - blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1; an initialization module for initializing the output matrix O, the attribution factor vector l, and the maximum - value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; a communication module for interacting with the optical computing chip to loop - execute the following steps until all sub - blocks in the key matrix K and the value matrix V are transmitted: respectively transmitting the j - th sub - blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j - th sub - blocks K j and V j , communicating and interacting with the optical computing chip to loop - execute the following steps until all sub - blocks in the query matrix are transmitted: transmitting the i - th sub - block Q i in the query matrix Q to the optical computing chip; receiving the local attention score matrix S i between the i - th sub - block Q j of the query matrix Q calculated by the optical computing chip and the j - th sub - block K ij of the key matrix K; performing non - linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; calculating the maximum value m ij of each row in the local attention score matrix S ij and the sum l ij of each row in the local attention weight matrix P ij ; updating the maximum - value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum - value tracking vector m, and each value in the maximum - value tracking vector m corresponds to a row in the local attention score matrix; updating the attribution factor vector l based on the following formula: where l i is the current latest attribution factor vector l, and each value in the attribution factor vector l corresponds to a row in the local attention score matrix; updating the output matrix O based on the following formula: Among them, O i is the current latest output matrix.
[0009] In a possible implementation, the size of each sub-block in the query Q matrix is B r , and the size of each sub-block in the key matrix K and the V matrix is B c ,
[0010] In a possible implementation, the device further includes: a storage module, configured to store the updated maximum value tracking vector m, attribution factor vector l, and output matrix O after each update of the maximum value tracking vector m, attribution factor vector l, and output matrix O.
[0011] In a third aspect, an optoelectronic Transformer accelerator collaboration device is provided, including: a general computing chip, including a general-purpose processor and a global static random access memory SRAM, the global static random access memory SRAM stores program instructions executable by the general-purpose processor, and the general-purpose processor can execute the following steps by calling the program instructions: based on the size M of the optical chip static random access memory SRAM on the optical computing chip, and the feature dimensions d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, divide the query matrix Q into T r sub-blocks, and divide the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1; initialize the output matrix O, attribution factor vector l, and maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; communicate with the optical computing chip to loop and execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: respectively transmit the j-th sub-blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j-th sub-blocks K j and V j , communicate with the optical computing chip to loop and execute the following steps until all sub-blocks in the query matrix are transmitted: transmit the i-th sub-block Q i in the query matrix Q to the optical computing chip; receive the i-th sub-block Q of the query matrix Q calculated by the optical computing chipi and the j-th sub-block K j of the key matrix K to obtain a local attention score matrix S ij ; perform a non-linear process on the local attention score matrix S ij to obtain a corresponding local attention weight matrix P ij ; calculate the maximum value m ij of each row in the local attention score matrix S ij and the sum l ij of each row in the local attention weight matrix P ij ; update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; update the normalization factor vector l based on the following formula: where l i is the current latest normalization factor vector l, and each value in the normalization factor vector l corresponds to a row in the local attention score matrix; update the output matrix O based on the following formula: where O i is the current latest output matrix.
[0012] In a possible implementation, the device further includes an optical computing chip, including an optical processor and an optical SRAM, where the optical SRAM stores program instructions executable by the optical processor, and the optical processor can execute the following steps by invoking the program instructions: the computing chip calculates the i-th sub-block Q i of the query matrix Q and the j-th sub-block K j of the key matrix K to obtain a local attention score matrix S ij ; transmit the local attention score matrix S ij to the general computing chip.
[0013] In a possible implementation, the size of each sub-block in the query Q matrix is B r , and the size of each sub-block in the key matrix K and the V matrix is B c .
[0014] In a possible implementation, the global static random access memory SRAM is further configured to: after updating the maximum value tracking vector m, the normalization factor vector l, and the output matrix O each time, store the updated maximum value tracking vector m, the normalization factor vector l, and the output matrix O.
[0015] Fourthly, a computer-readable storage medium is provided, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the optoelectronic Transformer accelerator cooperation method as described in the first aspect and any possible implementation manner in the first aspect.
[0016] The above technical solutions have the following technical effects:
[0017] Firstly, the general-purpose computing chip performs block processing on the matrix according to the size M of the SRAM on the optical computing chip and the characteristic dimensions d of the query matrix Q, the key matrix K, and the value matrix V. This method fully considers the characteristic of the limited SRAM capacity of the optical computing chip, avoids the calculation being blocked due to excessive data volume exceeding the SRAM's bearing capacity, enables the optical computing chip to efficiently process the block-divided matrix data, and realizes the reasonable allocation and collaborative work of the general-purpose computing chip and the optical computing chip in terms of computing resources.
[0018] Secondly, after the matrix is divided into multiple sub-blocks, each sub-block can be transmitted to the optical computing chip for calculation when needed, reducing unnecessary memory occupation. During the calculation process, only the currently needed sub-block will be loaded into the SRAM of the optical computing chip, and the memory space can be released in a timely manner after the calculation is completed to make room for the calculation of subsequent sub-blocks, thereby improving the overall memory usage efficiency.
[0019] Thirdly, by transmitting data in blocks, only a relatively small sub-block of data is transmitted between the general-purpose computing chip and the optical computing chip each time, reducing the data transmission delay and bandwidth pressure. This method avoids the transmission bottleneck that may be caused by transmitting a large amount of data at one time, enables the data to flow more efficiently between the two chips, and further improves the overall computing performance.
[0020] Finally, the optical computing chip has the ability of high-speed parallel computing. When calculating the local attention score matrix S i between the sub-block Q j of the query matrix Q and the sub-block K ij of the key matrix K, it can give full play to its advantages and quickly complete operations such as matrix multiplication. Compared with the traditional general-purpose computing chip calculating alone, the calculation time is greatly shortened. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the drawings without creative efforts.
[0022] Figure 1 A schematic block diagram of the optoelectronic Transformer accelerator collaboration method according to an embodiment of the present application is shown.
[0023] Figure 2 A schematic block diagram of the optoelectronic Transformer accelerator collaboration device according to an embodiment of the present application is shown.
[0024] Figure 3 Another schematic block diagram of the optoelectronic Transformer accelerator collaboration device according to an embodiment of the present application is shown. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0026] In today's information age, with the rapid development of fields such as natural language processing and computer vision, the Transformer model has achieved excellent results in many tasks with its powerful sequence modeling ability and parallel computing advantages. However, the computational complexity of the Transformer model is relatively high. Especially in its self-attention mechanism, a large number of matrix operations are involved, including multiplications between the query matrix Q, the key matrix K, and the value matrix V, and attention score calculations, etc. This makes the model face huge challenges in terms of computational resources and time when processing large-scale data and long sequences.
[0027] When traditional general-purpose computing chips (such as CPUs and GPUs) process these matrix operations, due to their characteristics based on electronic signal transmission and processing, there are problems such as bandwidth bottlenecks, high power consumption, and limited processing speed. With the continuous growth of data volume and the continuous expansion of model scale, these problems become more prominent, seriously affecting the training and inference efficiency of the Transformer model.
[0028] As an emerging computing technology, optical computing chips have advantages such as high speed, high bandwidth, and low power consumption, and can provide a more efficient computing method for matrix operations. Optical signals have extremely fast speeds and parallel processing capabilities during transmission, which can significantly improve the efficiency of operations such as matrix multiplication. However, optical computing chips also have some limitations. For example, the capacity of the static random access memory (SRAM) inside them is relatively limited, and it is impossible to store and process large-scale matrix data at one time.
[0029] In order to give full play to the respective advantages of general-purpose computing chips and optical computing chips and achieve their collaborative work to efficiently process matrix operations in the Transformer model, it has become a hot issue in current research. How to reasonably block these matrices according to the size of the SRAM on the optical computing chip and the characteristic dimensions of the query matrix Q, key matrix K, and value matrix V, and efficiently perform data transmission and computing task allocation between the general-purpose computing chip and the optical computing chip is the key to solving this problem. At the same time, during the blocked calculation process, it is necessary to accurately update the output matrix, as well as the relevant attribution factor vectors and maximum value tracking vectors, to ensure the accuracy and consistency of the final calculation results.
[0030] Therefore, the embodiments of the present application provide a collaborative method for an optoelectronic Transformer accelerator, aiming to solve the problems existing in the prior art and improve the computing efficiency and performance of the Transformer model.
[0031] Figure 1 The schematic flowchart of a collaborative method for an optoelectronic Transformer accelerator provided by the embodiments of the present application is shown. Optionally, this method can be executed by a general-purpose computing chip. Specifically, this method can be implemented through the interactive communication between the general-purpose computing chip and the optical computing chip, and the embodiments of the present application do not limit this. As Figure 1 shown, the method 100 includes the following parts or all of the content.
[0032] S110, based on the size M of the static random access memory SRAM of the optical chip on the optical computing chip and the characteristic dimension d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, divide the query matrix Q into T r sub-blocks, and divide the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1.
[0033] S120, initialize the output matrix O, the attribution factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ;
[0034] S130, interact with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted.
[0035] Specifically, the step S130 may include: respectively transmitting the j-th sub-blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j-th sub-blocks K j and V j , communicating and interacting with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the query matrix are transmitted: transmitting the i-th sub-block Qi in the query matrix Q to the optical computing chip; receiving the local attention score matrix S i between the i-th sub-block Q j in the query matrix Q calculated by the optical computing chip and the j-th sub-block K ij in the key matrix K; performing non-linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; calculating the maximum value m ij of each row in the local attention score matrix S ij and the sum l ij of each row in the local attention weight matrix Pij; updating the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; updating the normalization factor vector l based on the following formula: l new = e mi-mnew l i + e mij-mnew l ij , where l i is the current latest normalization factor vector l, and each value in the normalization factor vector l corresponds to a row in the local attention score matrix; updating the output matrix O based on the following formula: O new = diag(l new )(diag(l -1 ))O i + e mi-mnew P i V mij-mnew ij j i j , where O i is the current latest output matrix.
[0036] In the embodiments of the present application, an optical computing chip uses photons as information carriers for data transmission and calculation. It is based on optical principles, such as optical interference, diffraction, polarization and other phenomena, to achieve logical operations and data processing. According to the technical principle, optical technology chips are mainly divided into all-optical computing chips and optoelectronic hybrid computing chips, and the embodiments of the present application are mainly implemented based on all-optical computing chips. General computing chips are mainly based on electronic principles, and use the movement of electrons in semiconductor materials (such as silicon) to achieve data storage, transmission and processing. General computing chips generally include a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU), etc., and the embodiments of the present application do not limit the type of general computing chips.
[0037] Due to the limited capacity of the static random access memory (SRAM) of the optical chip in the optical computing chip, it is impossible to store and process the entire query matrix, key matrix and value matrix at one time. Therefore, these matrices need to be divided into multiple sub-blocks so that the data volume of each sub-block can adapt to the size of the SRAM, thereby realizing the effective transmission and calculation of data between the general computing chip and the optical computing chip. Specifically, in S110, it is known that the query matrix Q ∈ R N*d 、the key matrix K ∈ R N*d and the value matrix V ∈ R N*d , where N is the number of rows of the matrix, representing the length of the sequence; d is the number of columns of the matrix, that is, the feature dimension, which means that each element is represented by a d-dimensional vector. The size of the SRAM on the optical computing chip is M, and the unit is usually byte (Byte). When performing matrix partitioning, it is necessary to consider the amount of matrix data that the SRAM can accommodate. The size of the partition should be determined according to the SRAM size M and the feature dimension d of the matrix, and at the same time, it is necessary to ensure that the number of sub-blocks Tr and Tc after partitioning are positive integers greater than 1 for block calculation. In some embodiments, the size of each sub-block in the query matrix Q is Br, and the size of each sub-block in the key matrix K and the V matrix is B c , It should be understood that the embodiments of the present application do not limit the partitioning method and the size of the partitions.
[0038] For example, assume that we have a query matrix Q with a size of (1024, 128), a key matrix K with a size of (1024, 128), and a value matrix V with a size of (1024, 128). To adapt to the SRAM on the optical computing, the sub-block size of K and V is set to B c = 128, and then the sub-block size B of Q is calculated r=min(128,128)=128. Divide Q into 1024 / 128=8 sub-blocks, each of which is (128,128). Similarly, divide K and V into 8 sub-blocks, each of which is (128,128). At this point, the general computing chip has completed the block operation on the query matrix Q, key matrix K, and value matrix V, so that subsequent block calculations can be performed on the optical computing chip.
[0039] In S120, the general computing chip needs to initialize the output matrix O, the attribution factor vector l, and the maximum tracking vector m. These initialized data will be continuously updated in the subsequent block calculation process to obtain the final calculation result. The dimension of the matrix O is N*d, and the dimensions of the vectors l and m are both N, where N is the number of rows of the query matrix Q, the key matrix K, and the value matrix V, representing the length of the sequence, and d is the feature dimension. Each value in the attribution factor vector l and the maximum tracking vector m corresponds to a row in the matrix Q, V, K, or Q.
[0040] For example, the output matrix O is initialized to an all-zero matrix with a size of (1024, 128). The normalization factor l is initialized to an all-zero vector with a size of (1024, 1). The maximum tracking variable m is initialized to negative infinity with a size of (1024, 1). The normalization factor l and the maximum value m are also divided into blocks in the same way, each block size is (128, 1).
[0041] Next, in step S130, the general computing chip and the optical computing chip will interact cyclically until all sub-blocks of the key matrix K and the value matrix V are processed. j and V j When , all sub-blocks of the query matrix Q will be processed in a loop. That is to say, in step S130, there are two loops, an outer loop and an inner loop. For the outer loop, for j from 1 to T c Iterate, and each iteration will j and V j Transmitted to the optical computing chip, for example, first transmitted to the optical SRAM, and then transmitted to the optical computing core (PTC). For the inner loop, i ranges from 1 to T r Iterate and set the current Q i Transmitted to the optical computing chip, for example, first transmitted to the optical chip SRAM, and then transmitted to the optical processor.
[0042] The inner loop process in step S130 will be described in detail below.
[0043] For each pair of key matrix sub-blocks K j Sum value matrix sub-block V j , the general computing chip will loop through the query matrix Qi all sub - blocks until all sub - blocks of the query matrix are transmitted. The optical computing chip receives the query matrix sub - block Q i and the key matrix sub - block K j After that, using its high - speed computing ability, it calculates the local attention score matrix between them. The specific calculation method is usually achieved through matrix multiplication, that is, S ij = Q i K j T , in practical applications, in order to make the model more stable, it may also be divided by The optical computing chip transmits the calculated local attention score matrix S ij back to the general - purpose computing chip. After receiving the local attention score matrix S ij , the general - purpose computing chip performs non - linear processing on it, usually using the Softmax function to obtain the corresponding local attention weight matrix P ij . The Softmax function can convert the score matrix into a probability distribution, making the sum of each row of elements equal to 1. The general - purpose computing chip calculates the maximum value m ij of each row in the local attention score matrix S ij and the sum l ij of each row in the local attention weight matrix P ij , which can be achieved by traversing each row of the matrix. The general - purpose computing chip updates the maximum - value tracking vector m based on the formula m new = max(m i , m ij ). Here, the value of the corresponding row in the current maximum - value tracking vector m i is compared with m ij , and the larger value is updated into the maximum - value tracking vector m, so as to ensure that the maximum - value tracking vector m always records the maximum value of each row of all local attention score matrices. The general - purpose computing chip updates the normalization factor vector l based on the formula , and the normalization factor vector l records the cumulative value of the sum of each row of all local attention weight matrices. The general - purpose computing chip updates the output matrix O based on the corresponding formula , so that the results calculated by each sub - block can be gradually accumulated, and finally the complete output matrix is obtained. So far, one inner loop is completed.
[0044] When all sub - blocks of a round of query matrix Q are transmitted, one outer loop is completed. When all sub - blocks of a round of key matrix K and value matrix V are transmitted, the entire loop process ends, and the final output matrix O is output.
[0045] Optionally, in the embodiments of the present application, the method 100 may further include: after each update of the maximum value tracking vector m, the attribution factor vector l, and the output matrix O, the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O are stored. For example, the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O after each update may be stored in the memory in the general computing chip, i.e., the global SRAM.
[0046] Figure 2 shows a schematic block diagram of an optoelectronic Transformer accelerator cooperation device. As Figure 2 shown, the optoelectronic Transformer accelerator cooperation device 200 may include the following parts or all of the content.
[0047] The block module 210 is configured to divide the query matrix Q into T r sub-blocks, and divide the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1.
[0048] The initialization module 220 is configured to initialize the output matrix O, the attribution factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 .
[0049] The communication module 230 is configured to interact with the optical computing chip to loop through the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: respectively transmit the j-th sub-block K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j-th sub-block K j and V j , communicate and interact with the optical computing chip to loop through the following steps until all sub-blocks in the query matrix are transmitted: transmit the i-th sub-block Q i in the query matrix Q to the optical computing chip; receive the i-th sub-block Q i of the query matrix Q calculated by the optical computing chip and the j-th sub-block K of the key matrix Kj The local attention score matrix S between ij ; For the local attention score matrix S ij Perform nonlinear processing to obtain the corresponding local attention weight matrix P ij ; Calculate the local attention score matrix S ij The maximum value m in each row ij And the local attention weight matrix P ij The sum of each row in ij ; Update the maximum tracking vector m based on the following formula: m new =max(m i ,m ij ), where m i is the current latest maximum tracking vector m, each value in the maximum tracking vector m corresponds to a row in the local attention score matrix; the attribution factor vector l is updated based on the following formula: Among them, l i is the latest attribution factor vector l, each value in the attribution factor vector l corresponds to a row in the local attention score matrix; update the output matrix O based on the following formula: Among them, O i is the latest output matrix.
[0050] Optionally, in the embodiment of the present application, the size of each sub-block in the query Q matrix is B r , the size of each sub-block in the key matrix K and the V matrix is B c ,
[0051] Optionally, in an embodiment of the present application, the device also includes: a storage module, used to store the updated maximum tracking vector m, attribution factor vector l and output matrix O after each update of the maximum tracking vector m, attribution factor vector l and output matrix O.
[0052] Based on the same idea, the present application embodiment also provides another optoelectronic transformer accelerator coordination device. Figure 3 As shown, the optoelectronic Transformer accelerator collaborative device 300 includes a general computing chip 310, which includes a general processor 311 and a global SRAM 312. The global SRAM 312 stores program instructions that can be executed by the general processor 311. The general processor 311 calls the program instructions to perform the following steps:
[0053] Based on the size M of the static random access memory (SRAM) on the optical computing chip and the feature dimension d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, divide the query matrix Q into T r sub - blocks, and divide the key matrix K and the value matrix V into T c sub - blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1;
[0054] Initialize the output matrix O, the normalization factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ;
[0055] Interact with the optical computing chip to loop and execute the following steps until all sub - blocks in the key matrix K and the value matrix V are transmitted:
[0056] Transmit the j - th sub - blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip respectively, and for the j - th sub - blocks K j and V j , interact with the optical computing chip to loop and execute the following steps until all sub - blocks in the query matrix are transmitted:
[0057] Transmit the i - th sub - block Q i in the query matrix Q to the optical computing chip;
[0058] Receive the local attention score matrix S i between the i - th sub - block Q j of the query matrix Q and the j - th sub - block K ij of the key matrix K calculated by the optical computing chip;
[0059] Perform non - linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ;
[0060] Calculate the maximum value m ij of each row in the local attention score matrix S ij and the sum l ij of each row in the local attention weight matrix P ij ;
[0061] Update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix;
[0062] Update the attribution factor vector l based on the following formula: where l i is the current latest attribution factor vector l, and each value in the attribution factor vector l corresponds to a row in the local attention score matrix;
[0063] Update the output matrix O based on the following formula: where O i is the current latest output matrix.
[0064] Optionally, in the embodiments of the present application, the device further includes an optical computing chip 320, including an optical processor 321 and an optical SRAM 322. The optical SRAM stores program instructions executable by the optical processor, and the optical processor can execute the following steps by invoking the program instructions:
[0065] The computing chip calculates the local attention score matrix S i between the i-th sub-block Q j of the query matrix Q and the j-th sub-block K ij of the key matrix K;
[0066] Transmit the local attention score matrix S ij to the general computing chip.
[0067] Optionally, in the embodiments of the present application, the size of each sub-block in the query Q matrix is B r , and the size of each sub-block in the key matrix K and the V matrix is B c .
[0068] Optionally, in the embodiments of the present application, the global static random access memory SRAM 312 is further used for:
[0069] After updating the maximum value tracking vector m, the attribution factor vector l, and the output matrix O once each, store the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O.
[0070] It should be noted that for the detailed content of the device-side embodiments, reference may be made to the method-side embodiments. For the sake of brevity, it will not be elaborated here.
[0071] Based on the same concept, an embodiment of the present application further provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the above various method embodiments.
[0072] Although the present application has been described with reference to the preferred embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various technical features mentioned in each embodiment can be combined in any manner. The present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. An optoelectronic Transformer accelerator cooperation method, comprising: Based on the size M of the optical chip static random access memory SRAM on the optical computing chip and the feature dimension d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, divide the query matrix Q into T r sub-blocks, and divide the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1; Initialize the output matrix O, the attribution factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; Interacting and communicating with the optical computing chip to cyclically execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: Transmit the j-th sub-blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip respectively, and for the j-th sub-blocks K j and V j , communicate and interact with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the query matrix have been transmitted: Transmit the i-th sub-block Q i in the query matrix Q to the optical computing chip; Receive the i-th sub-block Q of the query matrix Q calculated by the optical computing chip i and the j-th sub-block K of the key matrix K j The local attention score matrix S between them ij ; Perform a non-linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; Calculate the maximum value m of each row in the local attention score matrix S ij and the sum l of each row in the local attention weight matrix P ij ; ij ij ; Update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; Update the attribution factor vector l based on the following formula: where l i is the current latest attribution factor vector l, and each value in the attribution factor vector l corresponds to a row in the local attention score matrix; Update the output matrix O based on the following formula: where O i is the current latest output matrix.
2. The method according to claim 1, wherein Each sub-block in the query Q matrix has a size of B r , and each sub-block in the key matrix K and the V matrix has a size of B c , 3. The method according to claim 1 or 2, characterized in that, The method further includes: After each update of the maximum value tracking vector m, the attribution factor vector l, and the output matrix O, storing the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O.
4. An optoelectronic Transformer accelerator cooperation device, comprising: Chunking module, configured to divide the query matrix Q into T r sub-chunks, and divide the key matrix K and the value matrix V into T c sub-chunks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K and V, and T r and T c are both positive integers greater than 1; Initialization module, used to initialize the output matrix O, the attribution factor vector l, and the maximum value tracking vector m, where, O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; A communication module for interacting and communicating with the optical computing chip to cyclically execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: Transmit the j-th sub-blocks K j and V j in the key matrix K and the value matrix V respectively to the optical computing chip, and for the j-th sub-blocks K j and V j , communicate and interact with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the query matrix are transmitted: Transmit the i-th sub-block Q in the query matrix Q i to the optical computing chip; Receive the \(i\)-th sub-block \(Q_{(i)}\) of the query matrix \(Q\) calculated by the optical computing chip i and the \(j\)-th sub-block \(K_{(j)}\) of the key matrix \(K\) j The local attention score matrix \(S\) between them ij ; Perform a non-linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; Calculate the maximum value m of each row in the local attention score matrix S ij and the sum l of each row in the local attention weight matrix P ij ; ij ij ; Update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; Update the attribution factor vector l based on the following formula: where l i is the current latest attribution factor vector l, and each value in the attribution factor vector l corresponds to a row in the local attention score matrix; Update the output matrix O based on the following formula: where O i is the current latest output matrix.
5. The device according to claim 4, characterized in that Each sub - block in the query Q matrix has a size of B r , each sub - block in the key matrix K and the V matrix has a size of B c , 6. The device according to claim 5, characterized in that The device further includes: A storage module for storing the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O after each update of the maximum value tracking vector m, the attribution factor vector l, and the output matrix O.
7. An optoelectronic Transformer accelerator collaborative device, comprising: A general-purpose computing chip, including a general-purpose processor and a global static random access memory SRAM, the global static random access memory SRAM stores program instructions executable by the general-purpose processor, and the general-purpose processor can execute the following steps by calling the program instructions: Based on the size M of the optical chip static random access memory SRAM on the optical computing chip and the feature dimension d of the query matrix Q, key matrix K, and value matrix V in the Transformer model, divide the query matrix Q into T r sub-blocks, and divide the key matrix K and the value matrix V into T c sub-blocks respectively, where Q ∈ R N*d , K ∈ R N*d , V ∈ R N* d, N is the number of rows of Q, K, and V, and T r and T c are both positive integers greater than 1; Initialize the output matrix O, the attribution factor vector l, and the maximum value tracking vector m, where O ∈ R N*d , l ∈ R N*1 and m ∈ R N*1 ; Interacting and communicating with the optical computing chip to cyclically execute the following steps until all sub-blocks in the key matrix K and the value matrix V are transmitted: Respectively transmit the j-th sub-blocks K j and V j in the key matrix K and the value matrix V to the optical computing chip, and for the j-th sub-blocks K j and V j , communicate and interact with the optical computing chip to repeatedly execute the following steps until all sub-blocks in the query matrix have been transmitted: Transfer the i-th sub-block Q i in the query matrix Q to the optical computing chip; Receive the i-th sub-block Q of the query matrix Q calculated by the optical computing chip i and the j-th sub-block K of the key matrix K j between the local attention score matrix S ij ; Perform a non-linear processing on the local attention score matrix S ij to obtain the corresponding local attention weight matrix P ij ; Calculate the maximum value m of each row in the local attention score matrix S ij and the sum l of each row in the local attention weight matrix P ij ; ij ij ; Update the maximum value tracking vector m based on the following formula: m new = max(m i , m ij ), where m i is the current latest maximum value tracking vector m, and each value in the maximum value tracking vector m corresponds to a row in the local attention score matrix; Update the attribution factor vector l based on the following formula: where l i is the current latest attribution factor vector l, and each value in the attribution factor vector l corresponds to a row in the local attention score matrix; Update the output matrix O based on the following formula: where O i is the current latest output matrix.
8. The device according to claim 7, characterized in that, The device further includes an optical computing chip, including an optical processor and an optical SRAM, the optical SRAM stores program instructions executable by the optical processor, and the optical processor can execute the following steps by calling the program instructions: The computing chip computes the i-th sub-block Q of the query matrix Q i and the j-th sub-block K of the key matrix K j to obtain the local attention score matrix S ij ; Transfer the local attention score matrix S ij to the general computing chip.
9. The device according to claim 7, wherein Each sub - block in the query Q matrix has a size of B r and each sub - block in the key matrix K and the V matrix has a size of B c , 10. The device according to any one of claims 7 to 9, characterized in that The global static random access memory SRAM is further used for: After each update of the maximum value tracking vector m, the attribution factor vector l, and the output matrix O, storing the updated maximum value tracking vector m, the attribution factor vector l, and the output matrix O.
Citation Information
Cited By
Large model data processing method and device, equipment and medium
CN121029652A
Data processing method and device of large model, equipment and medium
CN121029652B