MEMORY ACCESS ADAPTIVE SELF-ATTENTION MECHANISM FOR A TRANSFORMER MODEL
The memory access adaptive self-attention mechanism optimizes transformer model inference speed and accuracy by dynamically choosing between sparse and canonical self-attention based on total execution time and accuracy, addressing the inefficiencies in existing sparse algorithms.
Patent Information
- Application Number
- DE112022007966
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-09-04
AI Technical Summary
Transformer models face limited inference speed due to matrix computations in self-attention operations, particularly when sparse self-attention algorithms do not adequately consider memory access and calculation times, leading to potential increased execution times compared to canonical methods.
A memory access adaptive self-attention mechanism that compares the total execution time of sparse and canonical self-attention operations, determining the optimal approach based on minimizing the sum of memory access and calculation times while ensuring preset accuracy, using a variable k for sparse self-attention input matrix generation.
This approach enhances transformer model inference speed and accuracy by selecting the most efficient self-attention mechanism, balancing execution time and accuracy through adaptive memory access and calculation considerations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments described herein relate generally to neural network technology and, more particularly, to a memory access adaptive self-attention mechanism for a transformer model. BACKGROUND
[0002] Time series prediction is a critical component across many domains, such as sensor network monitoring, energy and smart grid management, economics and finance, and disease spread analysis. In these scenarios, a significant amount of time series data on past behavior can be used to make a long-term forecast, namely long-sequence time prediction (LSTF). Transformer models demonstrate superior performance over recurrent neural network (RNN) models in capturing long-range dependencies. A self-attention mechanism for transformer models can reduce the maximum length of network signal migration paths and avoid recurrent structures, thus demonstrating great potential for LSTF problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The various advantages of the embodiments will become apparent to those skilled in the art by reading the following description and the appended claims and by referring to the following drawings. Fig. 1 illustrates an exemplary procedure for determining a self-attention operation for a transformer model in accordance with some embodiments of the present disclosure; Fig. 2 illustrates another exemplary procedure for determining a self-attention operation for a transformer model in accordance with some embodiments of the present disclosure; Fig. 3A shows pseudocodes of an exemplary matrix multiplication algorithm according to some embodiments of the present disclosure; Fig. 3B shows pseudocodes of an exemplary top-k data selection algorithm for generating a sparse self-attention input matrix according to some embodiments of the present disclosure; Fig. 3C illustrates an example procedure of an example self-attention operation based on a sparse self-attention input matrix obtained by an example top-k data selection algorithm, according to some embodiments of the present disclosure; Fig. 4 is a flowchart of an exemplary procedure for implementing a memory access adaptive self-attention operation for a transformer model according to some embodiments of the present disclosure; Fig. 5 is a block diagram of an example processor platform structured to execute and / or instantiate machine-readable instructions and / or operations to implement example procedures in accordance with some embodiments of the present disclosure; Fig. 6 is a block diagram of an exemplary implementation of the processor circuitry of Fig. 5. Fig. 7 is a block diagram of another exemplary implementation of the processor circuitry of Fig. 5. DETAILED DESCRIPTION
[0004] Various aspects of the exemplary embodiments are described using terms commonly employed by those skilled in the art to convey the substance of the disclosure to others skilled in the art. However, it will be apparent to those skilled in the art that many alternative embodiments may be practiced using portions of the described aspects. For purposes of illustration, specific numbers, materials, and configurations are set forth to provide a basic understanding of the exemplary embodiments. However, it will be apparent to those skilled in the art that alternative embodiments may be practiced without the specific details. In other instances, well-known features may be omitted or simplified to avoid obscuring the exemplary embodiments.
[0005] Furthermore, various operations are described sequentially as multiple discrete operations in a manner that is helpful for understanding the example embodiments; however, the order of description should not be construed to imply that these operations are necessarily order-dependent. In particular, it is not necessary that these operations be performed in the order of presentation.
[0006] Transformer models have demonstrated superior performance in capturing long-range dependence and are widely used to solve LSTF problems. Although canonical transformer models show greatly improved accuracy in LSTF, the inference speed of transformer models is still a challenge for high-performance applications such as network traffic forecasting.
[0007] A major reason for the limited inference speed lies in the matrix computations involved in the self-attention operation of the transformer model. To reduce the time of matrix computations, sparse self-attention algorithms based on top-k data selection have been applied to some developed transformer models. These algorithms can generate a sparse self-attention input matrix by selecting a portion of an initial self-attention input matrix and then computing a sparse approximation of a self-attention operation for the transformer model. The value of k can determine the matrix computation complexity of the self-attention operation.However, current methods for selecting the value of k do not consider the additional time, such as memory access time and computation time, that the selection of the value of k incurs, which can sometimes exceed the reduced matrix multiplication time. As a result, the total execution time of the self-attention operation based on a sparse self-attention input matrix may increase compared to that of the canonical self-attention operation based on the initial self-attention input matrix, and the corresponding execution time of the transformer model may increase.
[0008] In view of this problem, according to some embodiments in the disclosure, it is proposed to compare the total execution time of the top-k data selection for generating the sparse self-attention input matrix and the self-attention operation based on the sparse self-attention input matrix and the execution time of the canonical self-attention operation based on the initial self-attention input matrix for the transformer model and to determine whether to perform the self-attention operation for the transformer model based on the sparse self-attention input matrix or the initial self-attention input matrix.
[0009] Fig. 1 illustrates an exemplary procedure for determining a self-attention operation for a transformer model according to some embodiments of the present disclosure; As shown in Fig. 1. Given initial self-attention input matrices for the transformer model, for example, a query matrix Q, a key matrix K, and a value matrix V, the first execution time T1 of top-k data selection for generating the sparse self-attention input matrix can be estimated based on a top-k data selection algorithm applied to the transformer model in step S101; the second execution time T2 of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix can be estimated in step S102; the third execution time T3 of performing the self-attention operation based on the initial self-attention input matrix can be estimated in step S103.and the sum of the first execution time T1 and the second execution time T2 may be compared with the third execution time in step S104 to determine whether to perform the self-attention operation for the transformer model based on the sparse self-attention input matrix (i.e., using top-k data selection-based sparse self-attention in the transformer model) or to perform the self-attention operation for the transformer model based on the initial self-attention input matrix (i.e., using canonical self-attention in the transformer model).;
[0010] It is noted that the top-k data selection algorithm applied to the transformer model can be any existing or future algorithm for selecting a number k of dominant data items from the initial self-attention input matrix to generate the sparse self-attention input matrix.
[0011] For example, a transformer-based model for LSTF, called Informer, is proposed by Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. in "Informer: Beyond Efficient Transformers for Long Sequence Time Series Forecasting," arXiv:2012.07436, March 28, 2021. In Informer, a probabilistic sparse (ProbSparse) self-attention mechanism is proposed to efficiently replace the canonical self-attention mechanism, which computes the attention probability distribution of a query on specific data and then selects the number of dominant queries as the value of k for top-k data selection to obtain a sparse self-attention input matrix that approximates the initial self-attention input matrix.
[0012] The ProbSparse self-attention mechanism achieves O(LlogL) time complexity for matrix computation. Compared to O(L 2) time complexity of canonical self-attention in matrix computation, the ProbSparse self-attention mechanism can greatly improve matrix computation performance. However, the Informer does not consider the memory access time and computation time brought by the top-k data selection algorithm. The total execution time of the Informer may in some cases be longer than that of the Transformer model with the canonical self-attention mechanism. According to the exemplary method presented in Fig. 1, the total execution time of the ProbSparse self-attention operation, including the memory access time and the comparison time associated with the top-k data selection, can be estimated and compared with the execution time of the canonical self-attention operation using the initial self-attention input matrix to determine whether to use the ProbSparse self-attention operation or the canonical self-attention operation in the transformer model.
[0013] In another example, a query selector-transformer model is proposed by Jacek Klimek, Jakub Klimek, Witold Kraskiewicz, and Mateusz Topolewski in "Long-Term Series Forecasting with Query Selector - Efficient Model of Sparse Attention," arXiv: 2107.08687v1, July 19, 2021. The query selector selects a predefined number ℓ of queries that give the largest dot products with keys, replaces the usual self-attention input matrix K with a column-constant matrix K' of elements equal to the mean of ℓ largest elements in column K, and constructs Q' by selecting ℓ rows of the usual self-attention input matrix Q with indices equal to indices of ℓ columns of K' with the highest common value of the given column and setting the remaining rows to zero. In this way, the generated sparse self-attention input matrix can be used in the self-attention operation for the transformer model.
[0014] In the query selector-transformer model, although a predefined number ℓ is used in the top-k data selection algorithm to sparse the self-attention input matrix and then accelerate the matrix multiplication calculation, the top-k data selection algorithm incurs significant memory access time and additional computation time. As a result, the total execution time of the query selector-transformer model may, in some cases, be longer than that of the transformer model with the canonical self-attention mechanism. According to the exemplary method presented in Fig. 1, the total execution time of the self-attention operation based on the generated sparse self-attention input matrix, including the memory access time and the comparison time associated with the top-k data selection, can be estimated and compared with the execution time of the canonical self-attention operation without using the sparse self-attention input matrix to determine whether to use the self-attention operation based on the generated sparse self-attention input matrix or the canonical self-attention operation in the transformer model.
[0015] According to some embodiments of the present disclosure, the value of k may be a variable and may be selected to minimize the sum of the first execution time of the top-k data selection for generating the sparse self-attention input matrix and the second execution time of performing the self-attention operation for the transformer model based on the sparse self-attention input matrix while ensuring that a preset accuracy is met.
[0016] Fig. 2 shows another exemplary procedure for determining a self-attention operation for a transformer model according to some embodiments of the present disclosure; In the exemplary procedure of Fig. 2, the value of k is to be a variable x, and steps S201 to S205 can be performed to obtain a transformer model with a high speed and a high accuracy for inference.
[0017] In step S201, the first execution time function T1(x) of the top-k data selection for generating the sparse self-attention input matrix may be estimated based on a top-k data selection algorithm applied to the transformer model. In step S202, the second execution time T2(x) of performing a self-attention operation for the transformer model may be estimated based on the sparse self-attention input matrix. In step S203, the value of k may be selected to minimize a sum of T1(x) and T2(x) while satisfying a preset accuracy. For example, k may be greater than or equal to c × InL Q where c is a constant sampling factor and LQ is a row number of an input query matrix for the transformer model. It has been proven that k>= cxlnL Qcan ensure the accuracy of the transformer model when using sparse self-attention based on top-k data selection. For example, the constant sampling factor c can be set to 2 or a larger number. Assume that when the value of k is equal to v (i.e., x=v), the sum of T1(x) and T2(x) is the minimum. In step S204, the third execution time T3 of performing the self-attention operation may be estimated based on the initial self-attention input matrix; and in step S205, the sum of T1(v) and T2(v) may be compared with the third execution time T3 to determine whether the self-attention operation is appropriate for the transformer model based on the sparse self-attention input matrix (i.e.,x=v, using top-k data selection (k=V) based sparse self-attention in the transformer model) or the self-attention operation for the transformer model based on the initial self-attention input matrix (i.e., using the canonical self-attention input matrix in the transformer model, that is, x=L. Q ) should be carried out.
[0018] Next, with reference to the Fig. 3A to Fig. 3C, an exemplary embodiment is provided to illustrate how to estimate the execution time of the self-attention operation for the transformer model.
[0019] In the following description, T Speicher specify the time for each data transfer between memory and register, T vergleich can specify the time for comparing two data elements, T Multiplikationcan specify the time for multiplying two data elements, T Addition can specify the time required for two data elements to act. The total time can be obtained from tests on a hardware platform on which the transformer model operates.
[0020] Fig. Figure 3A shows pseudocodes of a general matrix multiplication algorithm. During the execution of the code C[i,j] += A[i, t]*B[t, j], the operations may include loading A[i, t] and B[t, j] into memory, multiplying A[i, t] and B[t, j], adding the multiplication result to C[i, j], and then storing C[i, j] in memory. That is, the execution of the code C[i,j] += A[i, t]*B[t, j] may involve three memory access operations, one multiplication operation, and one addition operation, so the execution time of s times the code C[i,j] += A[i, t]*B[t, j] can be calculated as follows. Tc[i,j]=s*(3*TStorage+TMultiplication+TAddition)
[0021] Thus, the total execution time of T Matrix_Multiplikation of matrix multiplication (C=A*B) as m*n* T c[i,j] calculated and represented by the following equation. TMatrix_Multiplication=m*n*s*(3*TStorage+TMultiplication+TAdding)
[0022] As described above, the execution time of top-k data selection for generating the sparse self-attention input matrix can be estimated based on a top-k data selection algorithm applied to the transformer model. Top-k data selection algorithms can be constructed from sorting algorithms to select k dominant data items. Classic sorting algorithms include Bubble Sort, Quick Sort, and Heap Sort. Each algorithm has a different time complexity of sorting. To estimate the execution time of top-k data selection, three main operations can be considered: loading data items from memory, storing data items to memory, and comparing data items. For each sorting operation, it may be necessary to load two data items from memory into registers and then compare them.
[0023] Using the HeapSort algorithm as an example, the estimation of the execution time of top-k data selection can be done with reference to Fig. 3B, which shows the pseudocode of the HeapSort algorithm. Assume that the length of an array to be sorted is L, and a number of k dominant data elements are to be selected from the array. If k >= L, all data elements of the array are selected, and no selection operation is required. If k < L, a heap with k data elements can be constructed first. The time complexity of constructing the heap can be O(k*logk). For each comparison of two data elements, four memory accesses may be required, which include loading the two data elements from memory and storing the two data elements into memory. Thus, the time complexity of the total memory access to construct the heap can be O(4*k*logk). For the left (Lk) data elements, the time complexity of the comparison can be O((Lk)*logk). Then, the complexity of the memory access can be O(4*(Lk)*logk).Thus, the total execution time of the top-k data selection algorithm for a column of the matrix can be estimated as follows. Ttopk_column=(k*logk+(L−k)*logk)*TVariety+(4*k*logk+4*(L−k)*logk)*TStorage Ttopk_column=L*logk*TVariety+4*L*logk*TStorage
[0024] Since the self-attention input matrix A (e.g., the query matrix) can include D columns, the first execution time of the top-k data selection algorithm for the matrix can be estimated by the following equation. TMatrix_Top−k=D*(L*logk*Tcomparison+4*L*logk*TStorage)
[0025] In addition to the first execution time of the top-k data selection for generating the sparse self-attention input matrix, the second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix must be estimated to obtain the total execution time of the top-k data selection-based sparse self-attention operation. The estimation of the second execution time and the total execution time of the top-k data selection-based sparse self-attention operation can be performed by referring to Fig. 3C, which illustrates an exemplary procedure of an exemplary self-attention operation based on a sparse self-attention input matrix obtained by an exemplary top-k data selection algorithm, according to some embodiments of the present disclosure.
[0026] As in Fig. As shown in Figure 3C, the model presented in “Long-Term Series Forecasting with Query Selector - Efficient Model of Sparse Attention” by Jacek Klimek, Jakub Klimek, Witold Kraskiewicz, and Mateusz Topolewski, arXiv:2107.08687v1, July 19, 202, can be taken as an exemplary model to illustrate the estimation of the execution time of performing the self-attention operation based on top-k data selection.
[0027] To estimate the total execution time of the sparse self-attention operation based on top-k data selection, the main time cost of the Fig. 3C, e.g., calculating the execution time of the top-k selection in line 3 and the matrix multiplication in line 4, line 10, and line 11.
[0028] The code in line 3 selects k (=l) dominant data elements and accumulates them for each column, and the execution time can be given by T Zeile3 = T Matrix_Top-k + TAddition * l * D. Based on Equation 2, the execution time of the code in line 3 can be calculated as follows. TLine3=D*(L*Logl*TComparison+4*L*Logl*TStore)+TAddition*l*D
[0029] The code in line 4 is the matrix multiplication of the sparse key matrix K̂ ∈ R lxD and the transpose of the query matrix Q ∈ R LxD Based on Equation 1, the execution time of the code in line 4 can be calculated as follows. Trow4=l*D*L*(3*Tstorage+Tmultiplication+TAddition)
[0030] The code in line 10 is the matrix multiplication of the sparse query matrix Q̂ ∈ R l×D and the key matrix K ∈ R LxD Based on Equation 1, the execution time of the code in line 10 can be calculated as follows. Trow10=l*D*L*(3*Tstorage+Tmultiplication+TAddition)
[0031] The main time cost of the code in line 11 is the execution time of the matrix multiplication of Q˙K∈RlXL and the value matrix V∈R LxE Based on Equation 1, the execution time of the code in line 11 can be calculated as follows. Trow11=l*E*L*(3*Tstorage+Tmultiplication+TAddition)
[0032] As a result, the total core time cost of the sparse self-attention operation based on top-k data selection can be expressed as T spärliche_Selbstaufmerksamkeit = T Zeile3 + T Zeile4 + T Zeile10 + T Zeile11 estimated and represented by the following equation. Tspare_self-attention(l)=D*(L*Logl*Tcompare+4*L*Logl*TStorage)+ TAddition*l*D +l*L*(2*D+E)*(3*TStorage+TMultiplication+TAddition)
[0033] In some embodiments, the value of k (= l) in Equation 3 may be a variable and may be selected to minimize the estimated total execution time of the top-k data selection-based sparse self-attention operation while ensuring that a preset accuracy is met. That is, the value of l may be selected to ensure the minimum value of T spärliche_Selbstaufmerkeit to obtain as represented by equation 3 under the condition l>= c*lnL Q , where c is a constant sampling factor and L Q is a row number of the input query matrix for the transformer model. In “Informer: Beyond Efficient Transformers for Long Sequence Time Series Forecasting” by Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W., arXiv:2012.07436, March 28, 2021, it was proved that if l>= c×lnL Qis, the accuracy of a transformer model with sparse self-attention based on top-k data selection will not be lower than that of the transformer model with canonical self-attention. For example, the constant sampling factor c can be set to 2. That is, l>=2lnL Q .
[0034] Based on equation 3, if l=2lnL Q is the minimum value of T spärliche_Selbstaufmerksamkeit obtained and presented as follows. min(Tspare_self-attention(l))=Tspare_self-attention(2lnLQ)=D*(L*log(2lnLQ)*Tcomparison+4*L*log(2lnLQ)*TStorage)+TAddition*2lnLQ*D+2lnLQ*L *(2*D+E)*(3*TStorage+TMultiplication+TAddition)
[0035] Next, the third execution time of performing the canonical self-attention operation (Q*KT / d)*V based on the initial self-attention input matrix by calculating the matrix multiplication time and the memory access time of the canonical self-attention operation (Q*KT / d)*V are calculated and summed. Based on Equation 1, the third execution time of performing the canonical self-attention operation can be represented as follows. Tcanonical self-attention=L*D*L*(3*TStorage+TMultiplication+TAddition)+L*E*L*(3*TStorage+TMultiplication+TAddition)=L*L*(D+E)*(3*TStorage+TMultiplication + TAddition)
[0036] Then min(T spärliche_Selbstaufmerksamkeit (l)) and T kanonischeSelf-attention can be compared to determine whether the sparse self-attention operation based on top-k data selection should be performed on the transformer model or the canonical self-attention operation should be performed on the transformer model. If min(T spärliche_Selbstaufmerksamkeit (l)) less than T kanonische Selbstaufmerksamkeit the value of l can be used to obtain min(T spärliche-Selbstaufmerksamkeit (l)) can be set as the value of k for the top-k data selection, and the sparse self-attention operation based on top-k data selection can be performed for the transformer model, or otherwise, the canonical self-attention operation (Q*KT / d)*V based on the initial self-attention input matrix for the transformer model.
[0037] After selecting a suitable self-attention mechanism for the transformer model, the transformer model can be trained with the selected self-attention mechanism to obtain weights for the model and then used for inference with high accuracy and high speed.
[0038] As illustrated above, the embodiments of the present disclosure can provide the transformer model with high accuracy and high speed for inference based on a comparison of the total execution time, including memory access time, of the sparse self-attention operation based on top-k data selection and the execution time of the canonical self-attention operation. In other words, a memory access-adaptive self-attention mechanism is proposed for the transformer model.
[0039] To illustrate an overall principle of the memory access adaptive self-attention mechanism for the transformer model, the following is presented below with reference to Fig. 4 describes an exemplary procedure for implementing a memory access adaptive self-attention operation for a transformer model according to some embodiments of the present disclosure. The procedure may be implemented by processor circuitry and may include operations 410 through 440.
[0040] In operation 410, the processor circuitry may estimate the first execution time of selecting a number k of dominant data elements from an initial self-attention input matrix for a transformer model to generate a sparse self-attention input matrix.
[0041] In operation 420, the processor circuitry may estimate the second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix.
[0042] In operation 430, the processor circuitry may estimate the third execution time of performing the self-attention operation based on the initial self-attention input matrix.
[0043] In operation 440, the processor circuitry may perform the self-attention operation based on the first execution time, the second execution time, and the third execution time.
[0044] According to some embodiments, prior to performing the self-attention operation, the processor circuitry may determine a value of the number k to minimize a sum of the first execution time and the second execution time under a condition that a preset accuracy of the self-attention operation is met. In this case, the processor circuitry may perform the self-attention operation: performing the self-attention operation based on the sparse self-attention input matrix under a condition that the sum of the first execution time and the second execution time is less than the third execution time, or otherwise performing the self-attention operation based on the initial self-attention input matrix.
[0045] According to some embodiments, the first execution time may include memory access time for data transfer between memory and registers and comparison time for data comparison.
[0046] According to some embodiments, the second execution time may include a memory access time for data transfer between memory and registers and a matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix.
[0047] According to some embodiments, the third execution time may include a memory access time for data transfer between memory and registers and a matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the initial self-attention input matrix.
[0048] According to some embodiments, the initial self-attention input matrix may include a query matrix Q, a key matrix K, and a value matrix V
[0049] According to some embodiments, the number k may be greater than or equal to c×lnL Q where c is a constant sampling factor and L Q is a row number of an input query matrix for the transformer model.
[0050] Fig. 5 is a block diagram of an example processor platform 500 structured to execute and / or instantiate machine-readable instructions and / or operations to implement example procedures, according to some embodiments of the present disclosure. The processor platform 500 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smartphone, a tablet such as an iPad™), an internet appliance, a DVD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset (e.g., an AR (augmented reality) headset, a VR (virtual reality) headset, etc.), or other wearable device, or any other type of computing device.
[0051] The processor platform 500 of the illustrated example includes processor circuitry 512. The processor circuitry 512 of the illustrated example is hardware. For example, the processor circuitry 512 may be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers of any desired family or manufacturer. The processor circuitry 512 may be implemented by one or more semiconductor-based (e.g., silicon-based) devices.
[0052] The processor circuitry 512 of the illustrated example includes a local memory 513 (e.g., a cache, registers, etc.). The processor circuitry 512 of the illustrated example is in communication via a bus 518 with a main memory, which includes a volatile memory 514 and a non-volatile memory 516. The volatile memory 514 can be implemented by SDRAM (Synchronous Dynamic Random Access Memory), DRAM (Dynamic Random Access Memory), RDRAM® (RAMBUS ® Dynamic Random Access Memory (DRAM) and / or any other type of RAM device. Non-volatile memory 516 may be implemented by flash memory and / or any other desired type of storage device. Access to main memory 514, 516 of the illustrated example is controlled by a memory controller 517.
[0053] The processor platform 500 of the illustrated example also includes interface circuitry 520. The interface circuitry 520 may be implemented in hardware according to any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB) interface, a Bluetooth® interface, a Near Field Communication (NFC) interface, a PCI interface, and / or a PCIe interface.
[0054] In the illustrated example, one or more input devices 522 are connected to the interface circuitry 520. The one or more input devices 522 enable a user to input data and / or commands into the processor circuitry 512. The one or more input devices 522 may be implemented, for example, by an audio sensor, a microphone, a camera (still image or video), a keyboard, a button, a mouse, a touchscreen, a trackpad, a trackball, an isopoint device, and / or a speech recognition system.
[0055] Also connected to the interface circuitry 520 of the illustrated example is one or more output devices 524. The output devices 524 may be implemented, for example, by display devices (e.g., a light-emitting diode (LED), an organic light-emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube (CRT) display, an in-place switching (IPS) display, a touchscreen, etc.), a haptic output device, a printer, and / or speakers. The interface circuitry 520 of the illustrated example thus typically includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit, such as a GPU.
[0056] The interface circuitry 520 of the illustrated example also includes a communication device, such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface, to enable data exchange with external machines (e.g., data processing devices of any type) via a network 526. Communication may occur, for example, via an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a wireless line-of-sight system, a cellular telephone system, an optical connection, etc.
[0057] The processor platform 500 of the illustrated example also includes one or more mass storage devices 528 for storing software and / or data. Examples of such mass storage devices 528 include, but are not limited to, magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disk drives, RAID (Redundant Array of Independent Disks) systems, solid-state storage devices such as flash memory devices, and DVD drives.
[0058] The machine-executable instructions 532 may be stored in the mass storage device 528, in the volatile memory 514, in the non-volatile memory 516, and / or on a removable non-volatile computer-readable storage medium, such as a CD or DVD.
[0059] Fig. 6 is a block diagram of another exemplary implementation of the processor circuitry 512 of Fig. 5. In this example, the processor circuitry 512 of Fig. 5 is implemented by a microprocessor 600. For example, the microprocessor 600 may implement multi-core hardware circuitry, such as a CPU, a DSP, a GPU, an XPU, etc. Although it may include any number of example cores 602 (e.g., 1 core), the microprocessor 600 of this example is a multi-core semiconductor device including N cores. The cores 602 of the microprocessor 600 may operate independently or may cooperate to execute machine-readable instructions. For example, machine code corresponding to a firmware program, an embedded software program, or a software program may be executed by one of the cores 602, or may be executed by multiple cores 602 at the same time or at different times.In some examples, the machine code corresponding to the firmware program, the embedded software program, or the software program is divided into threads and executed in parallel by two or more of the cores 602. The software program may correspond to some or all of the machine-readable instructions and / or operations discussed herein.
[0060] The cores 602 may communicate through example bus 604. In some examples, bus 604 may implement a communication bus to enable communication in association with one or more of the cores 602. For example, bus 604 may implement an I2C (Inter-Integrated Circuit) bus and / or an SPI (Serial Peripheral Interface) bus and / or a PCI bus and / or a PCIe bus. Additionally or alternatively, bus 604 may implement any other type of data processing or electrical bus. The cores 602 may receive data, instructions, and / or signals from one or more external devices through example interface circuitry 606. The cores 602 may output data, instructions, and / or signals to the one or more external devices through interface circuitry 606. Although the cores 602 of this example include example local memory 620 (e.g.,L1 cache (Level 1), which may be divided into an L1 data cache and an L1 instruction cache), the microprocessor 600 also includes an example shared memory 610 that may be shared between the cores (e.g., L2 cache (Level 2)) to enable fast access to data and / or instructions. Data and / or instructions may be transferred (e.g., shared) by writing to and / or reading from the shared memory 610. The local memory 620 of each of the cores 602 and the shared memory 610 may be part of a hierarchy of memory devices that may include multiple levels of cache memories and main memory (e.g., the main memory 614, 616 of . Fig. 6). Typically, higher memory levels in the hierarchy have lower access times and smaller storage capacities than lower memory levels. Changes in the different levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherence policy.
[0061] Each core 602 may be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuitry. Each core 602 includes a control unit circuitry 614, an AL (arithmetic and logic) circuitry (sometimes referred to as ALU) 616, a plurality of registers 618, the L1 cache 620, and exemplary bus 622. Other structures may be present. For example, each core 602 may include a vector unit circuitry, a SIMD (single instruction multiple data) unit circuitry, an LSU (load / store unit) circuitry, a branch / jump unit circuitry, an FPU (floating-point unit), etc. The control unit circuitry 614 includes semiconductor-based circuitry structured to control (e.g., coordinate) data movement within the corresponding core 602.The AL circuitry 616 includes semiconductor-based circuitry structured to perform one or more mathematical and / or logical operations on the data in the corresponding core 602. The AL circuitry 616 of some examples performs integer operations. In other example cases, the AL circuitry 616 also performs floating-point operations. In still other examples, the AL circuitry 616 may include first AL circuitry that performs integer-based operations and second AL circuitry that performs floating-point operations. In some examples, the AL circuitry 616 may be referred to as an ALU (Arithmetic Logic Unit). Registers 618 are semiconductor-based structures for storing data and / or instructions, such asthe results of one or more of the operations performed by the AL circuitry 616 of the corresponding core 602. The registers 618 may include, for example, vector registers, SIMD registers, general-purpose registers, flag registers, segment registers, machine-specific registers, instruction pointer registers, control registers, debug registers, memory management registers, machine check registers, etc. The registers 618 may be arranged in a bank, as shown in FIG. Fig. 6. Alternatively, registers 618 may be organized in any other arrangement, format, or structure, including distribution within core 602 to reduce access time. Bus 620 may implement an I2C bus, an SPI bus, a PCI bus, and / or a PCIe bus.
[0062] Each core 602 and / or more generally, the microprocessor 600 may include additional and / or alternative structures to those shown and described above. For example, one or more clock circuits, one or more power supplies, one or more power gates, one or more CHAs (Cache Home Agents), one or more CMSs (Converged / Common Mesh Stops), one or more shifters (e.g., barrel shifters), and / or other circuitry may be present. The microprocessor 600 is a semiconductor device fabricated to include many interconnected transistors to implement the structures described above in one or more ICs (integrated circuits) contained in one or more packages. The processor circuitry may include and / or cooperate with one or more accelerators.In some examples, accelerators are implemented by logic circuitry to perform certain tasks faster and / or more efficiently than can be performed by a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed here. A GPU or other programmable device can also be an accelerator. Accelerators can be located on the processor circuitry, in the same chip package as the processor circuitry, and / or in one or more packages separate from the processor circuitry.
[0063] Fig. 7 is a block diagram of another exemplary implementation of the processor circuitry 512 of Fig. 5. In this example, the processor circuitry 600 is implemented by the FPGA circuitry 700. The FPGA circuitry 700 may, for example, be used to perform operations that would otherwise be performed by the exemplary microprocessor 600 of Fig. 6, which executes corresponding machine-readable instructions. However, once configured, the FPGA circuitry 700 instantiates the machine-readable instructions in hardware and can therefore often perform the operations faster than would be possible with a general-purpose microprocessor executing the corresponding software.
[0064] In contrast to the microprocessor 600 described above from Fig. 6 (which is a general-purpose device that may be programmed to execute some or all of the instructions disclosed herein, but whose connections and logic are fixed after manufacture) includes the FPGA circuitry 700 of the example of Fig. 7, in particular, interconnections and logic circuitry that can be configured and / or interconnected in different ways after fabrication for instantiation. In particular, the FPGA 700 can be viewed as an array of logic gates, interconnections, and switches. The switches can be programmed to change the way the logic gates are interconnected by the interconnections, effectively forming one or more dedicated logic circuits (unless and until the FPGA circuit 700 is reprogrammed). The configured logic circuits enable the logic gates to cooperate in different ways to perform various operations on the data received from the input circuits. These operations may correspond to some or all of the software represented by the operations discussed herein.Accordingly, the FPGA circuitry 700 may be structured to effectively instantiate some or all of the machine-readable instructions representing the operations discussed herein as dedicated logic circuits to perform the operations corresponding to these software instructions in a dedicated manner analogous to an ASIC. Therefore, the FPGA circuitry 700 may perform the operations corresponding to some or all of the operations discussed herein faster than the general-purpose microprocessor can.
[0065] In the example of Fig. 7, the FPGA circuitry 700 is structured to be programmed (and / or reprogrammed one or more times) by an end user via a hardware description language (HDL) such as Verilog. The FPGA circuitry 700 of Fig. 7 includes exemplary input / output (I / O) circuitry 702 for receiving and / or outputting data from exemplary configuration circuitry 704 and / or external hardware (e.g., external hardware circuitry) 706. For example, configuration circuitry 704 may implement interface circuitry that may receive machine-readable instructions for configuring FPGA circuitry 700 or one or more portions thereof. In some such examples, configuration circuitry 704 may receive the machine-readable instructions from a user, a machine (e.g., hardware circuitry (e.g., programmed or dedicated circuitry) that may implement an AI / ML (artificial intelligence / machine learning) model to generate the instructions), etc. In some examples, external hardware 706 may interface microprocessor 600 with Fig. 6. The FPGA circuitry 700 also includes an array of example logic gate circuits 708, a plurality of example configurable interconnects 710, and an example storage circuit 712. The logic gate circuitry 708 and the interconnects 710 are configurable to instantiate one or more operations discussed herein and / or other desired operations. Fig. The logic gate circuit 708 shown in Figure 7 is fabricated in groups or blocks. Each block contains semiconductor-based electrical structures that can be configured into logic circuits. In some examples, the electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that represent basic building blocks for logic circuits. Electrically controllable switches (e.g., transistors) are present within each of the logic gate circuits 708 to enable configuration of the electrical structures and / or the logic gates to form circuits for performing desired operations. The logic gate circuit 708 may include other electrical structures, such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.
[0066] The interconnections 710 of the illustrated example are conductive paths, traces, vias, or the like that may include electrically controllable switches (e.g., transistors) whose state may be changed by programming (e.g., using an HDL instruction language) to enable or disable one or more connections between one or more of the logic gate circuitry 708 to program desired logic circuits.
[0067] The memory circuitry 712 of the illustrated example is structured to store one or more results of one or more of the operations performed by the corresponding logic gates. The memory circuitry 712 may be implemented by registers or the like. In the illustrated example, the memory circuitry 712 is distributed among the logic gate circuitry 708 to facilitate access and increase execution speed.
[0068] The exemplary FPGA circuit arrangement 700 of Fig. 7 also includes example circuitry 714 for dedicated operations. In this example, the circuitry 714 for dedicated operations includes special-purpose circuitry 716 that can be invoked to implement frequently used functions to avoid the need to program those functions in the field. Examples of such special-purpose circuitry 716 include memory control circuitry (e.g., DRAM), PCIe control circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of special-purpose circuitry may also be present. In some examples, the FPGA circuitry 700 may also include example programmable general-purpose circuitry 718, such as an example CPU 720 and / or an example DSP 722.Additionally or alternatively, there may also be further programmable general-purpose circuitry 718, such as a GPU, an XPU, etc., which may be programmed to perform further operations.
[0069] Although Fig. 6 and Fig. 7 two exemplary implementations of the processor circuitry 512 of Fig. 5, many other approaches are conceivable. For example, a modern FPGA circuitry, as mentioned above, may include an on-board CPU, such as one or more of the exemplary CPUs 720 of Fig. 7. Thus, the processor circuitry 512 of Fig. 5 additionally by combining the exemplary microprocessor 600 of Fig. 6 and the exemplary FPGA circuit arrangement 700 of Fig. 7. In some such hybrid examples, a first portion of the machine-readable instructions may be executed by one or more of the cores 602 of Fig. 6 and a second part of the machine-readable instructions can be executed by the FPGA circuitry 700 of Fig. 7 can be executed.
[0070] In some examples, the processor circuitry 512 may be Fig. 5 in one or more encapsulations. For example, the processor circuitry 600 of Fig. 6 and / or the FPGA circuit arrangement 700 of Fig. 7 in one or more packages. In some examples, an XPU may be implemented by the processor circuitry 512 of Fig. 5, which may be located in one or more packages. For example, the XPU may include a CPU in one package, a DSP in another package, a GPU in another package, and an FPGA in yet another package.
[0071] In various embodiments, the operations explained herein may be implemented as hardware (e.g., logic circuitry), software, firmware, or combinations thereof, which may be provided as a computer program product, e.g., including a tangible (e.g., non-transitory) machine-readable or computer-readable medium having stored thereon instructions (or software sequences) used to program a computer to perform a process explained herein. The machine-readable medium may include a storage device. Additionally, such computer-readable media may be downloaded as a computer program product, wherein the program is executed from a remote computer (e.g., a server) to a requesting computer (e.g.,a client) by means of data signals provided in a carrier wave or other propagation medium over a communications link (for example, a bus, modem, or network connection).
[0072] Although examples have been described in language specific to structural features and / or methodological acts, it is understood that the claimed subject matter is not intended to be limited to the specific features or acts described. Instead, the specific features and acts are disclosed as example forms of implementing the claimed subject matter. Additional notes and examples: Example 1 includes a device comprising: interface circuitry; and processor circuitry coupled to the interface circuitry and configured to obtain an initial self-attention input matrix for a transformer model received via the interface circuitry; estimating a first execution time of selecting a number k of dominant data elements from the initial self-attention input matrix to generate a sparse self-attention input matrix; estimating the second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix; estimating the third execution time of performing the self-attention operation based on the initial self-attention input matrix;and performing the self-attention operation based on the first execution time, the second execution time, and the third execution time; Example 2 includes the apparatus of Example 1, wherein the processor circuitry, prior to performing the self-attention operation, is further configured to determine a value of the number k for minimizing a sum of the first execution time and the second execution time under a condition that a preset accuracy of the self-attention operation is satisfied. Example 3 includes the apparatus of Example 2, wherein the processor circuitry is configured to perform the self-attention operation by performing the self-attention operation based on the sparse self-attention input matrix under a condition that the sum of the first execution time and the second execution time is less than the third execution time, or otherwise performing the self-attention operation based on the initial self-attention input matrix. Example 4 includes the device of any of examples 1 to 3, wherein the first execution time includes memory access time for data transfer between memory and registers and comparison time for a data comparison. Example 5 includes the apparatus of any of Examples 1 to 4, wherein the second execution time includes memory access time for a data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. Example 6 includes the apparatus of any of Examples 1 to 5, wherein the third execution time includes memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. Example 7 includes the setup of any of Examples 1 to 6, wherein the initial self-attention input matrix comprises a query matrix Q, a key matrix K, and a value matrix V Example 8 includes the device of any of Examples 1 to 7, wherein the number k is greater than or equal to c×lnL Q where c is a constant sampling factor and L Q is a row number of an input query matrix for the transformer model. Example 9 includes a method comprising: estimating a first execution time of selecting a number k of dominant data items from an initial self-attention input matrix for a transformer model to generate a sparse self-attention input matrix; estimating a second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix; estimating a third execution time of performing the self-attention operation based on the initial self-attention input matrix; and performing the self-attention operation based on the first execution time, the second execution time, and the third execution time. Example 10 includes the method of Example 9, wherein the method further comprises, prior to performing the self-attention operation: determining a value of the number k to minimize a sum of the first execution time and the second execution time under a condition that a preset accuracy of the self-attention operation is satisfied. Example 11 includes the method of Example 10, wherein performing the self-attention operation comprises: performing the self-attention operation based on the sparse self-attention input matrix under a condition that the sum of the first execution time and the second execution time is less than the third execution time, or otherwise performing the self-attention operation based on the initial self-attention input matrix. Example 12 includes the method of any of Examples 9 to 11, wherein the first execution time includes memory access time for data transfer between memory and registers and comparison time for data comparison. Example 13 includes the method of any of Examples 9 to 12, wherein the second execution time includes memory access time for a data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. Example 14 includes the method of any of Examples 9 to 13, wherein the third execution time includes memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. Example 15 includes the method of any of Examples 9 to 14, wherein the initial self-attention input matrix comprises a query matrix Q, a key matrix K, and a value matrix V Example 16 includes the method of any of Examples 9 to 15, wherein the number k is greater than or equal to c×lnL Q where c is a constant sampling factor and L Q is a row number of an input query matrix for the transformer model. Example 17 includes a computer-readable medium having instructions stored thereon, the instructions, when executed by processor circuitry, causing the processor circuitry to perform any of the methods of Examples 9 to 16. Example 18 includes a device comprising means for performing a method of Examples 9 to 16.
[0073] Various techniques, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy disks, CD-ROMs, hard disks, a non-transitory computer-readable storage medium, or any other machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes a means for performing the various techniques. The non-transitory computer-readable storage medium may be a computer-readable storage medium that does not include a signal. In the case of program code execution on programmable computers, the computing system may include a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.The volatile and non-volatile memory and / or storage elements may be a RAM, EPROM, flash drive, optical drive, magnetic hard disk, solid-state drive, or other medium for storing electronic data. One or more programs implementing or utilizing the various techniques described herein may employ an application programming interface (API), reusable controllers, and the like. Such programs may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, the one or more programs may, if desired, be implemented in assembly or machine language. In any event, the language may be a compiled or interpreted language and may be combined with hardware implementations. Example systems or devices may include, but are not limited to:Laptop computers, tablet computers, desktop computers, smartphones, computer terminals and servers, storage databases, and other electronics that utilize circuitry and programmable memory, such as household appliances, smart televisions, digital video disc (DVD) players, heating, ventilation, and air conditioning (HVAC) controls, light switches, and the like.
[0074] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, certain embodiments that may be practiced. These embodiments are also referred to herein as "examples." Such examples may include elements in addition to those shown or described. However, the inventors also contemplate examples in which only the elements shown or described are provided. In addition, the present inventors also contemplate examples using any combination or permutation of these elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof) or with respect to other examples (or one or more aspects thereof) shown or described herein.
[0075] All publications, patents, and patent specifications cited in this document are incorporated herein by reference in their entirety, as if individually incorporated by reference. In the event of any conflicting usage between this document and the documents incorporated by reference, the usage in the incorporated reference(s) shall be considered supplemental to that of this document; in the event of irreconcilable conflict, the usage in this document shall prevail.
[0076] Throughout this document, the terms "a," "an," or "an" are used as is customary in patent documents to include one or more than one, regardless of any other instances or uses of "at least one" or "one or more." Throughout this document, the term "or" is used to refer to a non-exclusive or, such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise specified. Throughout the appended claims, the terms "including" and "in which" are used as plain English equivalents of the respective terms "comprising" and "wherein."Also, in the following claims, the terms "including" and "comprising" are used openly, meaning that a system, apparatus, article, or process that includes elements in addition to those listed after such a term in a claim is still considered within the scope of that claim. Furthermore, in the following claims, the terms "first," "second," and "third," etc., are used only as labels and are not intended to impose numerical requirements on their objects.
[0077] The foregoing description is intended to be illustrative and not restrictive. Thus, the examples described above (or one or more aspects thereof) may also be used in combination with one another. Other embodiments may be used, for example, by one of ordinary skill in the art upon review of the foregoing description. The abstract is intended to enable the reader to quickly ascertain the essence of the technical disclosure and is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, in the above detailed description, various features may be grouped together to streamline the disclosure. This should not be construed to imply that an unclaimed disclosed feature is essential to a claim.Rather, the subject matter of the invention may lie in fewer than all features of a particular disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment. The scope of the embodiments should be determined by reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature
[0000] Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., und Zhang, W. in "Informer: Beyond Efficient Transformierer for Long Sequence Time Series Forecasting", arXiv:2012.07436, 28. März 2021
[0011] Jacek Klimek, Jakub Klimek, Witold Kraskiewicz und Mateusz Topolewski in „Long-Term Series Forecasting with Query Selector - Efficient Model of Sparse Attention“, arXiv: 2107.08687v1, 19. Juli 2021
[0013] Long-Term Series Forecasting with Query Selector - Efficient Model of Sparse Attention“ von Jacek Klimek, Jakub Klimek, Witold Kraskiewicz und Mateusz Topolewski, arXiv: 2107.08687v1, 19. Juli 202
[0026] Informer: Beyond Efficient Transformierer for Long Sequence Time Series Forecasting“ von Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., und Zhang, W., arXiv:2012.07436, 28. März 2021
[0033]
Claims
[1] A device comprising: an interface circuitry; and a processor circuitry coupled to the interface circuitry and configured to Obtaining an initial self-attention input matrix for a transformer model received via the interface circuitry; estimating a first execution time of selecting a number k of dominant data items from the initial self-attention input matrix to generate a sparse self-attention input matrix; Estimating a second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix; Estimating a third execution time of performing the self-attention operation based on the initial self-attention input matrix; and Performing the self-attention operation based on the first execution time, the second execution time, and the third execution time. [2] The device of claim 1, wherein the processor circuitry is further configured, prior to performing the self-attention operation, to determine a value of the number k for minimizing a sum of the first execution time and the second execution time under a condition that a preset accuracy of the self-attention operation is satisfied. [3] The device of claim 2, wherein the processor circuitry is configured to perform the self-attention operation by performing the self-attention operation based on the sparse self-attention input matrix under a condition that the sum of the first execution time and the second execution time is less than the third execution time, or otherwise performing the self-attention operation based on the initial self-attention input matrix. [4] Device according to one of claims 1 to 3, wherein the first execution time comprises memory access time for data transfer between memory and registers and comparison time for data comparison. [5] The device according to any one of claims 1 to 3, wherein the second execution time includes a memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. [6] The device according to any one of claims 1 to 3, wherein the second execution time includes a memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. [7] Device according to one of claims 1 to 3, wherein the initial self-attention input matrix comprises a query matrix Q, a key matrix K and a value matrix V [8] Device according to one of claims 1 to 3, wherein the number k is greater than or equal to c×lnL Q where c is a constant sampling factor and L Q is a row number of an input query matrix for the transformer model. [9] Method comprising: Estimating a first execution time of selecting a number k of dominant data items from an initial self-attention input matrix for a transformer model to generate a sparse self-attention input matrix; Estimating a second execution time of performing a self-attention operation for the transformer model based on the sparse self-attention input matrix; Estimating a third execution time of performing the self-attention operation based on the initial self-attention input matrix; and Performing the self-attention operation based on the first execution time, the second execution time, and the third execution time. [10] The method of claim 9, wherein the method further comprises, prior to performing the self-attention operation: Determining a value of the number k for minimizing a sum of the first execution time and the second execution time under a condition that a preset accuracy of the self-attention operation is satisfied. [11] The method of claim 10, wherein performing the self-attention operation comprises: Performing the self-attention operation based on the sparse self-attention input matrix under a condition that the sum of the first execution time and the second execution time is less than the third execution time, or otherwise performing the self-attention operation based on the initial self-attention input matrix. [12] Method according to one of claims 9 to 11, wherein the first execution time comprises memory access time for data transfer between memory and registers and comparison time for data comparison. [13] The method of any one of claims 9 to 11, wherein the second execution time includes a memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. [14] The method of any one of claims 9 to 11, wherein the second execution time comprises a memory access time for data transfer between memory and registers and matrix multiplication time of matrix multiplication operations involved in the self-attention operation based on the sparse self-attention input matrix. [15] A method according to any one of claims 9 to 11, wherein the initial self-attention input matrix comprises a query matrix Q, a key matrix K and a value matrix V [16] Method according to one of claims 9 to 11, wherein the number k is greater than or equal to c×lnL Q where c is a constant sampling factor and L Q is a row number of an input query matrix for the transformer model. [17] A computer-readable medium having instructions stored thereon, the instructions, when executed by processor circuitry, causing the processor circuitry to perform any method according to claims 9 to 16. [18] Apparatus comprising means for carrying out a method according to claims 9 to 16.