Memory access adaptive self-attention mechanism for transducer models
By dynamically selecting k values in the transformer model and optimizing the execution time of sparse self-attention operations, the problem of limited inference speed in LSTF is solved by transformer model, achieving more efficient inference speed and accuracy.
Patent Information
- Application Number
- CN202280100542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-05-13
AI Technical Summary
The transformer model has limited inferring speed in long-sequence time series prediction (LSTF), mainly due to the long matrices calculation time involved in self-attention operations. The current method used to select k values fails to effectively consider memory access time and calculation time.
A memory access adaptive self-attention mechanism is proposed. By comparing the execution time of self-attention operations based on sparse self-attention input matrix and the initial self-attention input matrix, a suitable k value is dynamically selected to minimize the total execution time.
By optimizing the selection of k values, the total execution time of self-attention operations of the transformer model is reduced, the inference speed is improved, and the preset accuracy is ensured.
Smart Images

Figure CN119998815A_ABST
Abstract
Description
Technical Field
[0001] Embodiments described herein relate generally to neural network techniques, and more particularly to a memory-accessed adaptive self-attention mechanism for transformer models. Background Art
[0002] Time series forecasting is a key element in many fields, such as sensor network monitoring, energy and smart grid management, economics and finance, and disease spread analysis. In these scenarios, a large amount of time series data about past behaviors can be used to make long-term predictions, namely long-range time series forecasting (LSTF). Transformer models have better performance than recurrent neural network (RNN) models in capturing long-range dependencies. The self-attention mechanism used in transformer models can reduce the maximum travel path length of network signals and avoid repeated structures, so transformer models show great potential for LSTF problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Various advantages of the embodiments will become apparent to those skilled in the art upon reading the following description and appended claims, and by referring to the following drawings, in which:
[0004] Figure 1 An example process for determining a self-attention operation of a transformer model according to some embodiments of the present disclosure is shown;
[0005] Figure 2 Another example process for determining a self-attention operation of a transformer model according to some embodiments of the present disclosure is shown;
[0006] Figure 3A shows pseudo code for an example matrix multiplication algorithm according to some embodiments of the present disclosure;
[0007] Figure 3B shows pseudo code for an example top-k data selection algorithm for generating a sparse self-attention input matrix according to some embodiments of the present disclosure;
[0008] Figure 3C An example process of an example self-attention operation based on a sparse self-attention input matrix obtained by an example top-k data selection algorithm according to some embodiments of the present disclosure is shown;
[0009] Figure 4 A flowchart illustrating an example process for implementing a memory access adaptive self-attention operation of a transformer model according to some embodiments of the present disclosure;
[0010] Figure 5is a block diagram of an example processor platform configured to execute and / or instantiate machine-readable instructions and / or operations to implement example processes according to some embodiments of the present disclosure;
[0011] Figure 6 yes Figure 5 A block diagram of an example implementation of a processor circuit.
[0012] Figure 7 yes Figure 5 A block diagram of another example implementation of a processor circuit. DETAILED DESCRIPTION
[0013] The various aspects of the illustrative embodiments will be described using terms commonly used by those skilled in the art to convey the essence of the present disclosure to other persons skilled in the art. However, it will be apparent to those skilled in the art that many alternative embodiments may be practiced using portions of the described aspects. For purposes of explanation, specific numbers, materials, and configurations are set forth to provide a thorough understanding of the illustrative embodiments. However, it will be apparent to those skilled in the art that alternative embodiments may be practiced without the specific details. In other cases, known features may be omitted or simplified to avoid obscuring the illustrative embodiments.
[0014] Furthermore, various operations are described as multiple discrete operations in sequence in a manner that is most helpful for understanding the illustrative embodiments; however, the order of description should not be interpreted as implying that these operations are necessarily order-dependent. Specifically, these operations are not necessarily performed in the order presented.
[0015] Transformer models show excellent performance in capturing long-range dependencies and are widely used to solve LSTF problems. Although the regularized Transformer model greatly improves the accuracy of LSTF, the inference speed of the Transformer model is still a problem for high-performance applications such as network traffic prediction.
[0016] One of the main reasons for the limited inference speed is the matrix calculation involved in the self-attention operation of the transformer model. In order to reduce the time of matrix calculation, sparse self-attention algorithms based on top-k data selection have been applied to some evolved transformer models. These algorithms can create a sparse self-attention input matrix by selecting a portion of the initial self-attention input matrix, and then calculate a sparse approximation of the self-attention operation of the transformer model. The value of k can determine the matrix calculation complexity of the self-attention operation. However, the current method for selecting the value of k does not take into account the additional time brought by selecting the k value, such as memory access time and calculation time, which may sometimes be more than the reduced matrix multiplication time. Therefore, compared with the normalized self-attention operation based on the initial self-attention input matrix, the total execution time of the self-attention operation based on the sparse self-attention input matrix may increase, and then the corresponding execution time of the transformer model may increase.
[0017] In view of this problem, according to some embodiments of the present disclosure, it is proposed to compare the total execution time of the top-k data selection for generating a sparse self-attention input matrix and the self-attention operation based on the sparse self-attention input matrix with the execution time of the normalized self-attention operation based on the initial self-attention input matrix of the transformer model, and determine whether to perform the self-attention operation of the transformer model based on the sparse self-attention input matrix or the initial self-attention input matrix.
[0018] Figure 1 An example process of a self-attention operation for determining a transformer model according to some embodiments of the present disclosure is shown. Figure 1 As shown, given an initial self-attention input matrix of a transformer model, such as a query matrix Q, a key matrix K, and a value matrix V, at step S101, a first execution time T1 of top-k data selection for generating a sparse self-attention input matrix can be estimated based on a top-k data selection algorithm applied to the transformer model; at step S102, a second execution time T2 of performing a self-attention operation of the transformer model based on the sparse self-attention input matrix can be estimated; at step S103, a third execution time T3 of performing a self-attention operation based on the initial self-attention input matrix can be estimated; and at step S104, a sum of the first execution time T1 and the second execution time T2 can be compared with the third execution time to determine whether to perform the self-attention operation of the transformer model based on the sparse self-attention input matrix (i.e., using sparse self-attention based on top-k data selection in the transformer model) or to perform the self-attention operation of the transformer model based on the initial self-attention input matrix (i.e., using normalized self-attention in the transformer model).
[0019] It should be noted that the top-k data selection algorithm applied to the transformer model can be any existing or future algorithm for selecting a number k of dominant data elements from an initial self-attention input matrix to generate a sparse self-attention input matrix.
[0020] For example, Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. proposed a transformer-based model (named Informer) for LSTF in "Informer: Beyond efficient transformer for long sequence timeseries forecasting" (arXiv:2012.07436) on March 28, 2021. In Informer, a probabilistic sparse (ProbSparse) self-attention mechanism is proposed to effectively replace the regularized self-attention mechanism, which calculates the attention probability distribution of the query on the specific data, and then selects the number of dominant queries as the value of k for top-k data selection to obtain a sparse self-attention input matrix, which is approximate to the initial self-attention input matrix.
[0021] ProbSparse self-attention mechanism achieves O(LlogL) time complexity in matrix calculation. Compared with the O(LlogL) time complexity of normalized self-attention in matrix calculation, 2 ), the ProbSparse self-attention mechanism can greatly improve the matrix calculation performance. However, Informer does not consider the memory access time and calculation time brought by the top-k data selection algorithm. The total execution time of Informer may be longer than the total execution time of the Transformer model using the normalized self-attention mechanism in some cases. Figure 1 The example process shown can estimate the total execution time of the ProbSparse self-attention operation (which includes the memory access time and comparison time associated with the top-k data selection) and compare it with the execution time of the normalized self-attention operation using the initial self-attention input matrix to determine whether to use the ProbSparse self-attention operation or the normalized self-attention operation in the Transformer model.
[0022] In another example, Jacek Klimek, Jakub Klimek, Witold Kraskiewicz, and Mateusz Topolewski proposed a query selector transformer model in "Long-term series forecasting with query selector-efficient model of sparse attention" (arXiv:2107.08687v1) on July 19, 2021. The query selector performs the following operations: selects a predefined number of queries that give the maximum scalar product with the key; replace the regular self-attention input matrix K with a column constant matrix K', where the elements of K' are equal to the columns of K The mean of the largest elements; and the matrix Q' is constructed by selecting from the regular self-attention input matrix Q The row whose index equals K' that has the highest common value for a given column The indices of the columns are added and the remaining rows are set to zero. In this way, the generated sparse self-attention input matrix can be used in the self-attention operation of the Transformer model.
[0023] In the query selector transformer model, although a predefined number of To make the self-attention input matrix sparse and then speed up the matrix multiplication calculation, the top-k data selection algorithm brings a lot of memory access time and additional computation time. Therefore, the total execution time of the query selector transformer model may be longer than that of the transformer model using the normalized self-attention mechanism in some cases. Figure 1 The example process shown can estimate the total execution time of the self-attention operation based on the generated sparse self-attention input matrix (the total execution time includes the memory access time and comparison time associated with the top-k data selection), and compare it with the execution time of the normalized self-attention operation without using the sparse self-attention input matrix to determine whether to use the self-attention operation based on the generated sparse self-attention input matrix or the normalized self-attention operation in the transformer model.
[0024] According to some embodiments of the present disclosure, the value of k can be a variable and can be selected to minimize the sum of a first execution time for top-k data selection for generating a sparse self-attention input matrix and a second execution time for performing the self-attention operation of the transformer model based on the sparse self-attention input matrix, while ensuring that a preset accuracy is met.
[0025] Figure 2Another example process for determining a self-attention operation of a transformer model according to some embodiments of the present disclosure is shown. Figure 2 In the example process of , it is assumed that the value of k is a variable x, and steps S201 to S205 may be performed to implement a converter model with high speed and high accuracy for inference.
[0026] At step S201, a first execution time function T1(x) of top-k data selection for generating a sparse self-attention input matrix can be estimated based on a top-k data selection algorithm applied to a transformer model. At step S202, a second execution time T2(x) of a self-attention operation of a transformer model based on the sparse self-attention input matrix can be estimated. At step S203, a value of k can be selected to minimize the sum of T1(x) and T2(x) while satisfying a preset accuracy. For example, k can be greater than or equal to c×lnL Q , where c is a constant sampling factor and L Q is the number of rows in the input query matrix of the Transformer model. It has been shown that when using sparse self-attention based on top-k data selection, k>=c×lnL Q The accuracy of the transformer model can be ensured. For example, the constant sampling factor c can be set to 2 or a larger number. Assume that when the value of k is equal to v (ie, x=v), the sum of T1(x) and T2(x) is minimized. At step S204, a third execution time T3 for performing a self-attention operation based on the initial self-attention input matrix can be estimated; and at step S205, the sum of T1(v) and T2(v) can be compared with the third execution time T3 to determine whether to perform the self-attention operation of the transformer model based on a sparse self-attention input matrix (ie, x=v, using sparse self-attention based on top-k (k=v) data selection in the transformer model), or to perform the self-attention operation of the transformer model based on the initial self-attention input matrix (ie, using normalized self-attention in the transformer model, i.e., x=L Q ).
[0027] Next, an example embodiment is provided to illustrate how to refer to FIG. 3A to FIG. 3C The example code shown here estimates the execution time of the self-attention operation of the Transformer model.
[0028] In the following description, T memory It can represent the time of each data transfer between the memory and the register, T compare Can represent the comparison time of two data elements, T multiply It can represent the time of multiplying two data elements, T addIt can represent the time to add two data elements. All the times can be obtained from testing the hardware platform on which the transformer model operates.
[0029] Figure 3A Pseudocode of a general matrix multiplication algorithm is shown. During the execution of code C[i, j]+=A[i, t]*B[t, j], operations may include: loading A[i, t] and B[t, j] into memory, multiplying A[i, t] and B[t, j], adding the multiplication result to C[i, j], and then storing C[i, j] into memory. That is, the execution of code C[i, j]+=A[i, t]*B[t, j] may include three memory access operations, one multiplication operation, and one addition operation, so the execution time of code C[i, j]+=A[i, t]*B[t, j] for s times may be calculated as follows.
[0030] T c[i,j] =s*(3*T memory +T multiply +T add )
[0031] Therefore, the T of matrix multiplication (C = A*B) is matrix_multiply The total execution time can be calculated as m*n*T c[i,j] And is expressed by the following equation.
[0032] T matrix_multiply =m*n*s*(3*T memory +T multiply +T add )(Equation 1)
[0033] As described above, the execution time of the top-k data selection for generating the sparse self-attention input matrix can be estimated based on the top-k data selection algorithm applied to the transformer model. The top-k data selection algorithm can be constructed by a sorting algorithm to select k dominant data elements. Classical sorting algorithms include BubbleSort, QuickSort, and HeapSort, etc. Each algorithm has a different sorting time complexity. In order to estimate the execution time of top-k data selection, three main operations can be considered, such as loading data elements from memory, storing data elements to memory, and comparing data elements. For each sorting operation, it may be necessary to load two data elements from memory into registers respectively and then compare them.
[0034] Taking the HeapSort algorithm as an example, you can refer to Figure 3B To describe the estimated execution time for top-k data selection, Figure 3BThe pseudo-code of the HeapSort algorithm is shown. Assume that the length of the array to be sorted is L, and the number of dominant data elements to be selected from the array is k. If k >= L, all data elements of the array will be selected and no selection operation is required. If k < L, a heap with k data elements can be constructed first. The time complexity of constructing the heap can be O(k * log k). For each comparison of two data elements, four memory accesses may be required, which includes loading two data elements from the memory and storing two data elements into the memory. Therefore, the time complexity of the total memory access for constructing the heap can be O(4 * k * log k). For the remaining (L - k) data elements, the time complexity of comparison may be O((L - k) * log k). Then the complexity of memory access can be O(4 * (L - k) * log k). Therefore, the total execution time of the top-k data selection algorithm for a column of the matrix can be estimated as follows.
[0035] T topk_column =(k * log k+(L - k) * log k) * T compare +(4 * k * log k+4 * (L - k) * log k) * T memory
[0036] T topk_column =L * log k * T compare +4 * L * log k * T memory
[0037] Since the self-attention input matrix A (e.g., the query matrix) can include D columns, the first execution time of the top-k data selection algorithm for this matrix can be estimated by the following equation.
[0038] T matrix_top-k =D * (L * log k * T compare +4 * L * log k * T memory (Equation 2)
[0039] In addition to the first execution time for generating the top-k data selection of the sparse self-attention input matrix, the second execution time of performing the self-attention operation of the transformer model based on the sparse self-attention input matrix also needs to be estimated to obtain the total execution time of the sparse self-attention operation based on top-k data selection. Reference can be made to Figure 3C to describe the estimation of the second execution time and the total execution time of the sparse self-attention operation based on top-k data selection. Figure 3C shows an example process of an example self-attention operation based on a sparse self-attention input matrix obtained by an example top-k data selection algorithm according to some embodiments of the present disclosure.
[0040] like Figure 3C As shown, the query selector described in "Long-term series forecasting with Query Selector-efficient model of sparse attention" by Jacek Klimek, Jakub Klimek, Witold Kraskiewicz, and Mateusz Topolewski (arXiv:2107.08687v1, July 19, 2021) can be used as an example to illustrate the estimation of the execution time of the self-attention operation based on top-k data selection.
[0041] To estimate the total execution time of the sparse self-attention operation based on top-k data selection, we can calculate Figure 3C The main time costs of the algorithm shown are, for example, the execution time of computing the top-k selection on line 3 and the matrix multiplications on lines 4, 10, and 11.
[0042] The code on line 3 selects k (= l) leading data elements for each column and accumulates them. The execution time can be given by T line3 =T matrix_top-k +T add *l*D represents. Based on Equation 2, the execution time of the code on line 3 can be calculated as follows.
[0043] T line3 =D*(L*logl*T compare +4*L*logl*T memory )+T add *l*D
[0044] The code on line 4 is the sparse key matrix and the query matrix Q∈R LxD Based on Equation 1, the execution time of the code on Line 4 can be calculated as follows.
[0045] T line4 =l*D*L*(3*T memory +T multiply +T add )
[0046] The code on line 10 is the sparse query matrix and bond matrix K∈R LxD Based on Equation 1, the execution time of the code on Line 10 can be calculated as follows.
[0047] T line10 =l*D*L*(3*Tmemory +T multiply +T add )
[0048] The main time cost of the code on line 11 is The sum matrix V∈R LxE Based on Equation 1, the execution time of the code on line 11 can be calculated as follows.
[0049] T line11 =l*E*L*(3*T memory +T multiply +T add )
[0050] Therefore, the total main time cost of the sparse self-attention operation based on top-k data selection can be estimated as T sparse_self-attention =T line3 +T line4 +T line10 +T line11 , and is expressed by the following equation.
[0051] T sparse_self-attention (l) = D*(L*logl*T compare +4*L*logl*T memory )+T add *l*D
[0052] +l*L*(2*D+E)*(3*T memory +T multiply +T add )(Formula 3)
[0053] In some embodiments, the value of k(=l) in Equation 3 may be variable and may be selected to minimize the estimated total execution time of the sparse self-attention operation based on top-k data selection while ensuring that a preset accuracy is met. That is, when l>=c*lnL Q In this case, the value of l can be selected to obtain T expressed by Equation 3 sparse_self-attention The minimum value of , where c is a constant sampling factor, and L QThe number of rows of the input query matrix of the transformer model. It has been proved in Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W., "Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting" (arXiv:2012.07436, March 28, 2021) that when l>=c×lnL Q When , the accuracy of the Transformer model using sparse self-attention based on top-k data selection will not be lower than the accuracy of the Transformer model using normalized self-attention. For example, the constant sampling factor c can be set to 2. That is, l>=2lnL Q .
[0054] Based on equation 3, when l = 2lnL Q When T sparse_self-attention and is expressed as follows.
[0055] min(T sparse_self-attention (l)) = T sparse_self-attention (2lnL Q )=D*(L*log(2lnL Q )*T compare +4*L*log(2lnL Q )*T memory )+T add *2lnL Q *D+2lnL Q *L*(2*D+E)*(3*T memory +T multiply +T add )
[0056] Next, by normalizing the self-attention operation The matrix multiplication time and memory access time are calculated and summed to estimate the normalized self-attention operation based on the initial self-attention input matrix. The third execution time of . Based on Equation 1, the third execution time of performing the normalized self-attention operation can be expressed as follows.
[0057] T canonical self-attention =L*D*L*(3*T memory +T multiply +T add )+L*E*L*(3*T memory +T multiply +T add )=L*L*(D+E)*(3*Tmemory +T multiply +T add )
[0058] Then, min(T sparse_self-attention (l)) and T canonical self-attention A comparison is made to determine whether to perform the sparse self-attention operation based on top-k data selection of the transformer model or the normalized self-attention operation of the transformer model. sparse_self-attention (l)) is less than T canonical self-attention , then used to obtain min(T sparse_self-attention The value of l in (l)) can be set to the value of k for top-k data selection, and a sparse self-attention operation based on top-k data selection can be performed for the Transformer model. Otherwise, a normalized self-attention operation based on the initial self-attention input matrix can be performed for the Transformer model.
[0059] After selecting an appropriate self-attention mechanism for the Transformer model, the Transformer model utilizing the selected self-attention mechanism can be trained to obtain the weights of the model and then used to perform inference with high accuracy and high speed.
[0060] As described above, embodiments of the present disclosure can provide a transformer model with high accuracy and high speed for inference based on a comparison of the total execution time including the memory access time of the sparse self-attention operation based on top-k data selection and the execution time of the normalized self-attention operation. In other words, a memory access adaptive self-attention mechanism is proposed for the transformer model.
[0061] To illustrate the overall idea of the memory access adaptive self-attention mechanism of the transformer model, we will refer to Figure 4 4. An example process for implementing a memory access adaptive self-attention operation of a transformer model according to some embodiments of the present disclosure is described. The process may be implemented by a processor circuit and may include operations 410 to 440.
[0062] At operation 410 , the processor circuit may estimate a first execution time for selecting a number k of dominant data elements from an initial self-attention input matrix of a transformer model to generate a sparse self-attention input matrix.
[0063] At operation 420 , the processor circuit may estimate a second execution time for performing a self-attention operation of the transformer model based on the sparse self-attention input matrix.
[0064] At operation 430 , the processor circuit may estimate a third execution time to perform the self-attention operation based on the initial self-attention input matrix.
[0065] At operation 440 , the processor circuit may perform a self-attention operation based on the first execution time, the second execution time, and the third execution time.
[0066] According to some embodiments, before performing the self-attention operation, the processor circuit may determine a value of a number k for minimizing the sum of the first execution time and the second execution time while satisfying a preset accuracy of the self-attention operation. In this example, the processor circuit may perform the self-attention operation: if the sum of the first execution time and the second execution time is less than the third execution time, perform the self-attention operation based on the sparse self-attention input matrix, otherwise, perform the self-attention operation based on the initial self-attention input matrix.
[0067] According to some embodiments, the first execution time may include a memory access time for data transfer between a memory and a register and a comparison time for data comparison.
[0068] According to some embodiments, the second execution time may include a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on a sparse self-attention input matrix.
[0069] According to some embodiments, the third execution time may include a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on an initial self-attention input matrix.
[0070] According to some embodiments, the initial self-attention input matrix may include a query matrix Q, a key matrix K, and a value matrix V.
[0071] According to some embodiments, the number k may be greater than or equal to c×lnL Q , where c is a constant sampling factor and L Q The number of rows in the input query matrix for the transformer model.
[0072] Figure 5 1 is a block diagram of an example processor platform 500 configured to execute and / or instantiate machine-readable instructions and / or operations to implement example processes according to some embodiments of the present disclosure. The processor platform 500 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smart phone, an iPad, etc.), or a processor. TM, such as a tablet computer), an Internet device, a DVD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset (e.g., an augmented reality (AR) headset, a virtual reality (VR) headset, etc.) or other wearable device, or any other type of computing device.
[0073] The processor platform 500 of the illustrated example includes a processor circuit 512. The processor circuit 512 of the illustrated example is hardware. For example, the processor circuit 512 can be implemented by one or more integrated circuits, logic circuits, FPGA microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. The processor circuit 512 can be implemented by one or more semiconductor-based (e.g., silicon-based) devices.
[0074] The processor circuit 512 of the illustrated example includes a local memory 513 (e.g., cache, registers, etc.). The processor circuit 512 of the illustrated example communicates with a main memory including a volatile memory 514 and a non-volatile memory 516 via a bus 518. The volatile memory 514 may be a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), Dynamic Random Access Memory The non-volatile memory 516 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 514, 516 of the illustrated example is controlled by a memory controller 517.
[0075] The processor platform 500 of the illustrated example also includes an interface circuit 520. The interface circuit 520 may be implemented by hardware according to any type of interface standard (e.g., an Ethernet interface, a universal serial bus (USB) interface, interface, near field communication (NFC) interface, PCI interface and / or PCIe interface).
[0076] In the example shown, one or more input devices 522 are connected to the interface circuit 520. The input device(s) 522 allow a user to input data and / or commands into the processor circuit 512. The input device 522 may be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touch screen, a trackpad, a trackball, an isopoint device, and / or a voice recognition system.
[0077] One or more output devices 524 are also connected to the interface circuit 520 of the illustrated example. The output device 524 can be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube (CRT) display, an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Therefore, the interface circuit 520 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit such as a GPU.
[0078] The interface circuitry 520 of the illustrated example also includes communication devices, such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces, that facilitate data exchange with external machines (e.g., any type of computing device) over a network 526. Communications may occur over, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, an optical connection, and the like.
[0079] The processor platform 500 of the illustrated example also includes one or more mass storage devices 528 for storing software and / or data. Examples of such mass storage devices 528 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disk drives, redundant array of independent disks (RAID) systems, solid-state storage devices such as flash memory devices, and DVD drives.
[0080] The machine-executable instructions 532 may be stored in the mass storage device 528, in the volatile memory 514, in the non-volatile memory 516, and / or on a removable, non-transitory computer-readable storage medium such as a CD or DVD.
[0081] Figure 6 yes Figure 5 5. In this example, Figure 5The processor circuit 512 is implemented by the microprocessor 600. For example, the microprocessor 600 can implement a multi-core hardware circuit, such as a CPU, a DSP, a GPU, an XPU, etc. Although it can include any number of example cores 602 (for example, 1 core), the microprocessor 600 of this example is a multi-core semiconductor device including N cores. The cores 602 of the microprocessor 600 can operate independently or can cooperate to execute machine-readable instructions. For example, the machine code corresponding to a firmware program, an embedded software program, or a software program can be executed by one core 602, or can be executed by multiple cores 602 simultaneously or at different times. In some examples, the machine code corresponding to the firmware program, the embedded software program, or the software program is decomposed into threads and executed in parallel by two or more cores 602. The software program may correspond to part or all of the machine-readable instructions and / or operations discussed herein.
[0082] The core 602 can communicate via an example bus 604. In some examples, the bus 604 can implement a communication bus to implement communications associated with (one or more) cores 602. For example, the bus 604 can implement at least one of the following bus protocols: an inter-integrated circuit (I2C) bus, a serial peripheral interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the bus 604 can implement any other type of computing or electrical bus. The core 602 can obtain data, instructions, and / or signals from one or more external devices via an example interface circuit 606. The core 602 can output data, instructions, and / or signals to one or more external devices via the interface circuit 606. Although the core 602 of this example includes an example local memory 620 (e.g., a level 1 (L1) cache that can be divided into an L1 data cache and an L1 instruction cache), the microprocessor 600 also includes an example shared memory 610 (e.g., a level 2 (L2 cache)) that can be shared by the cores for high-speed access to data and / or instructions. Data and / or instructions may be transferred (eg, shared) by writing to and / or reading from the shared memory 610. The local memory 620 of each core 602 and the shared memory 610 may be a plurality of levels of cache memory and main memory (eg, Figure 6 Cache memory is part of a hierarchy of storage devices (main memory 614, 616). Generally, memory at higher levels in the hierarchy exhibits less access time and has less storage capacity than memory at lower levels. Changes in the various levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherence policy.
[0083] Each core 602 may be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuit. Each core 602 includes a control unit circuit 614, an arithmetic and logic (AL) circuit (sometimes referred to as an ALU) 616, a plurality of registers 618, an L1 cache 620, and an example bus 622. Other structures may exist. For example, each core 602 may include a vector unit circuit, a single instruction multiple data (SIMD) unit circuit, a load / store unit (LSU) circuit, a branch / jump unit circuit, a floating point unit (FPU) circuit, etc. The control unit circuit 614 includes a semiconductor-based circuit configured to control (e.g., coordinate) data movement within the corresponding core 602. The AL circuit 616 includes a semiconductor-based circuit configured to perform one or more mathematical and / or logical operations on data within the corresponding core 602. The AL circuit 616 of some examples performs integer-based operations. In other examples, the AL circuit 616 also performs floating-point operations. In other other examples, the AL circuit 616 may include a first AL circuit that performs integer-based operations and a second AL circuit that performs floating-point operations. In some examples, AL circuit 616 may be referred to as an arithmetic logic unit (ALU). Register 618 is a semiconductor-based structure to store data and / or instructions, such as the results of one or more operations performed by AL circuit 616 of corresponding core 602. For example, register 618 may include vector register(s), SIMD register(s), general register(s), flag register(s), segment register(s), machine-specific register(s), instruction pointer register(s), control register(s), debug register(s), memory management register(s), machine check register(s), etc. Register 618 may be arranged in a manner such as Figure 6 Alternatively, registers 618 may be organized in any other arrangement, format, or structure, including being distributed throughout core 602 to reduce access time. Bus 620 may implement at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.
[0084] Each core 602 and / or (more generally) microprocessor 600 may include additional and / or alternative structures to the structures shown and described above. For example, there may be one or more clock circuits, one or more power supplies, one or more power gates, one or more cache master agents (CHA), one or more converged / commonmesh stops (CMS), one or more shifters (e.g., (one or more) barrel shifters) and / or other circuits. Microprocessor 600 is a semiconductor device manufactured to include many interconnected transistors to implement the above structure in one or more integrated circuits (ICs) contained in one or more packages. The processor circuit may include one or more accelerators and / or collaborate with one or more accelerators. In some examples, the accelerator is implemented by a logic circuit that can be more efficient and / or perform certain tasks faster than a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. GPUs or other programmable devices may also be accelerators. Accelerators may be on a processor circuit, in the same chip package as a processor circuit, and / or in one or more packages separated from a processor circuit.
[0085] Figure 7 yes Figure 5 5. In this example, the processor circuit 600 is implemented by an FPGA circuit 700. For example, the FPGA circuit 700 may be used to perform the following operations: These operations may be performed by a processor circuit 512 executing corresponding machine-readable instructions. Figure 6 However, once configured, FPGA circuit 700 instantiates machine-readable instructions in hardware and, therefore, can generally perform operations faster than they could be performed by a general-purpose microprocessor executing corresponding software.
[0086] More specifically, Figure 6 In contrast to the microprocessor 600 of FIG. 1 (which is a general-purpose device that can be programmed to perform some or all of the operations disclosed herein but whose interconnections and logic circuits are fixed once manufactured), Figure 7The example FPGA circuit 700 includes interconnections and logic circuits that can be configured and / or interconnected in different ways to instantiate after manufacturing. In particular, FPGA 700 can be considered as an array of logic gates, interconnections, and switches. The switches can be programmed to change how the logic gates are interconnected by interconnections, thereby effectively forming one or more special logic circuits (unless and until the FPGA circuit 700 is reprogrammed). The configured logic circuit enables the logic gates to cooperate in different ways to perform different operations on the data received by the input circuit. These operations can correspond to some or all of the software represented by the operations discussed herein. Therefore, FPGA circuit 700 can be constructed to effectively instantiate some or all of the machine-readable instructions representing the operations discussed herein as special logic circuits, performing operations corresponding to those software instructions in a special manner similar to ASIC. Therefore, FPGA circuit 700 can perform operations corresponding to some or all of the operations discussed herein faster than a general-purpose microprocessor.
[0087] exist Figure 7 In the example of FPGA circuit 700, FPGA circuit 700 is configured to be programmed (and / or reprogrammed one or more times) by an end user via a hardware description language (HDL) such as Verilog. Figure 7 FPGA circuit 700 includes example input / output (I / O) circuit 702 to obtain data from and / or output data to example configuration circuit 704 and / or external hardware (e.g., external hardware circuit) 706. For example, configuration circuit 704 can implement interface circuitry that can obtain machine-readable instructions to configure FPGA circuit 700 or portions thereof. In some such examples, configuration circuit 704 can obtain machine-readable instructions from a user, a machine (e.g., a hardware circuit (e.g., programmed or dedicated circuitry) that can implement an artificial intelligence / machine learning (AI / ML) model to generate instructions), etc. In some examples, external hardware 706 can implement Figure 6 The microprocessor 600 of the embodiment of the present invention. The FPGA circuit 700 also includes an array of example logic gate circuits 708, a plurality of example configurable interconnects 710, and example storage circuits 712. The logic gate circuits 708 and the interconnects 710 may be configured to instantiate one or more operations discussed herein and / or other desired operations. Figure 7The logic gate circuit 708 shown in is manufactured in groups or blocks. Each block includes a semiconductor-based electrical structure that can be configured as a logic circuit. In some examples, these electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that provide basic building blocks for logic circuits. Electrically controllable switches (e.g., transistors) are present in each logic gate circuit 708 so that the electrical structure and / or logic gate can be configured to form a circuit that performs the desired operation. The logic gate circuit 708 may include other electrical structures, such as a lookup table (LUT), a register (e.g., a flip-flop or latch), a multiplexer, etc.
[0088] The interconnect 710 of the illustrated example is a conductive path, trace, via, etc., which may include an electrically controllable switch (e.g., a transistor), the state of which may be changed by programming (e.g., using an HDL instruction language) to activate or disable one or more connections between one or more of the logic gate circuits 708, thereby programming the desired logic circuit.
[0089] The storage circuit 712 of the illustrated example is configured to store (one or more) results of one or more operations performed by the corresponding logic gates. The storage circuit 712 can be implemented by registers, etc. In the illustrated example, the storage circuit 712 is distributed between the logic gate circuits 708 to facilitate access and improve execution speed.
[0090] Figure 7 The example FPGA circuit 700 also includes an example dedicated operation circuit 714. In this example, the dedicated operation circuit 714 includes a dedicated circuit 716, which can be called to implement common functions to avoid the need to program those functions in the field. Examples of such dedicated circuits 716 include memory (e.g., DRAM) controller circuits, PCIe controller circuits, clock circuits, transceiver circuits, memories, and multiplier-accumulator circuits. There may be other types of dedicated circuits. In some examples, the FPGA circuit 700 may also include an example general-purpose programmable circuit 718, such as an example CPU 720 and / or an example DSP 722. Additionally or alternatively, there may be other general-purpose programmable circuits 718 that can be programmed to perform other operations, such as GPUs, XPUs, etc.
[0091] although Figure 6 and Figure 7 Shows Figure 5 These are two example implementations of the processor circuit 512, but many other approaches are contemplated. For example, as described above, modern FPGA circuits may include an onboard CPU, such as one or more Figure 7 The example CPU is 720. Therefore, Figure 5 The processor circuit 512 may additionally be combined by Figure 6An example microprocessor 600 and Figure 7 In some such hybrid examples, the first portion of the machine-readable instructions may be implemented by one or more Figure 6 The core 602 is executed, and the second portion of the machine-readable instructions may be executed by Figure 7 The FPGA circuit 700 executes.
[0092] In some examples, Figure 5 The processor circuit 512 may be in one or more packages. For example, Figure 6 The processor circuit 600 and / or Figure 7 The FPGA circuit 700 can be in one or more packages. In some examples, the XPU can be composed of Figure 5 The processor circuit 512 implements, Figure 5 The processor circuit 512 of the XPU can be in one or more packages. For example, the XPU can include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in yet another package.
[0093] In various embodiments, the operations discussed herein may be implemented as hardware (e.g., logic circuits), software, firmware, or a combination thereof, which may be provided as a computer program product, for example, including one or more tangible (e.g., non-volatile) machine-readable or computer-readable media having instructions (or software processes) stored thereon for programming a computer to perform the processes discussed herein. The machine-readable medium may include a storage device.
[0094] In addition, such computer-readable media may be downloaded as a computer program product, where the program may be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a communication link (e.g., a bus, a modem, or a network connection) as a data signal provided in a carrier wave or other propagation medium.
[0095] Although multiple embodiments have been described using language specific to structural features and / or methodological acts, it should be understood that the claimed subject matter may not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as sample forms of implementing the claimed subject matter.
[0096] Additional notes and examples:
[0097] Example 1 includes an apparatus comprising: an interface circuit; and a processor circuit coupled to the interface circuit and configured to: obtain an initial self-attention input matrix of a transformer model received via the interface circuit; estimate a first execution time for selecting a number k of dominant data elements from the initial self-attention input matrix to generate a sparse self-attention input matrix; estimate a second execution time for performing a self-attention operation of the transformer model based on the sparse self-attention input matrix; estimate a third execution time for performing a self-attention operation based on the initial self-attention input matrix; and perform the self-attention operation based on the first execution time, the second execution time, and the third execution time.
[0098] Example 2 includes the apparatus of Example 1, wherein, before performing the self-attention operation, the processor circuit is further configured to determine a value of a number k for minimizing the sum of the first execution time and the second execution time while satisfying a preset accuracy of the self-attention operation.
[0099] Example 3 includes the apparatus of Example 2, wherein the processor circuit is configured to perform a self-attention operation in the following manner: when the sum of a first execution time and a second execution time is less than a third execution time, the self-attention operation is performed based on a sparse self-attention input matrix; otherwise, the self-attention operation is performed based on an initial self-attention input matrix.
[0100] Example 4 includes the apparatus of any one of Examples 1 to 3, wherein the first execution time comprises: a memory access time for data transfer between a memory and a register and a comparison time for data comparison.
[0101] Example 5 includes an apparatus of any one of Examples 1 to 4, wherein the second execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on a sparse self-attention input matrix.
[0102] Example 6 includes an apparatus of any one of Examples 1 to 5, wherein the third execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on an initial self-attention input matrix.
[0103] Example 7 includes the apparatus of any one of Examples 1 to 6, wherein the initial self-attention input matrix includes a query matrix Q, a key matrix K, and a value matrix V.
[0104] Example 8 includes the apparatus of any one of Examples 1 to 7, wherein the number k is greater than or equal to c×lnL Q , where c is a constant sampling factor and L QThe number of rows in the input query matrix for the transformer model.
[0105] Example 9 includes a method, comprising: estimating a first execution time for selecting a number k of dominant data elements from an initial self-attention input matrix of a transformer model to generate a sparse self-attention input matrix; estimating a second execution time for performing a self-attention operation of the transformer model based on the sparse self-attention input matrix; estimating a third execution time for performing the self-attention operation based on the initial self-attention input matrix; and performing the self-attention operation based on the first execution time, the second execution time, and the third execution time.
[0106] Example 10 includes the method of Example 9, wherein, before performing the self-attention operation, the method further includes: determining a value of a number k for minimizing the sum of the first execution time and the second execution time while satisfying a preset accuracy of the self-attention operation.
[0107] Example 11 includes the method of Example 10, wherein performing the self-attention operation includes: when the sum of the first execution time and the second execution time is less than the third execution time, performing the self-attention operation based on a sparse self-attention input matrix; otherwise, performing the self-attention operation based on an initial self-attention input matrix.
[0108] Example 12 includes the method of any one of Examples 9 to 11, wherein the first execution time includes: a memory access time for data transfer between a memory and a register and a comparison time for data comparison.
[0109] Example 13 includes the method of any one of Examples 9 to 12, wherein the second execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on a sparse self-attention input matrix.
[0110] Example 14 includes the method of any one of Examples 9 to 13, wherein the third execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in a self-attention operation based on an initial self-attention input matrix.
[0111] Example 15 includes the method of any one of Examples 9 to 14, wherein the initial self-attention input matrix includes a query matrix Q, a key matrix K, and a value matrix V.
[0112] Example 16 includes the method of any one of Examples 9 to 15, wherein the number k is greater than or equal to c×lnL Q , where c is a constant sampling factor and L Q The number of rows in the input query matrix for the transformer model.
[0113] Example 17 includes a computer readable medium having instructions stored thereon, wherein the instructions, when executed by a processor circuit, cause the processor circuit to perform any of the methods of Examples 9-16.
[0114] Example 18 includes an apparatus comprising means for performing any of the methods of Examples 9-16.
[0115] Various technologies or some aspects or parts thereof can take the form of program codes (i.e., instructions) embodied in tangible media, such as floppy disks, CD-ROMs, hard disk drives, non-transient computer-readable storage media, or any other machine-readable storage media, wherein, when the program code is loaded into a machine (e.g., a computer) and executed by the machine, the machine becomes a device for implementing each technology. Non-transient computer-readable storage media can be computer-readable storage media that do not include signals. In the case of executing program codes on a programmable computer, a computing system can include a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. Volatile and non-volatile memory and / or storage elements can be RAM, EPROM, flash drives, optical drives, magnetic hard disk drives, solid-state drives, or other media for storing electronic data. One or more programs that can implement or utilize various technologies described herein can use application programming interfaces (APIs), reusable controls, etc. Such programs can be implemented with high-level procedures or object-oriented programming languages to communicate with a computer system. However, if necessary, (one or more) programs can be implemented with assembly language or machine language. In any case, the language may be a compiled or interpreted language and may be combined with a hardware implementation. Exemplary systems or devices may include, but are not limited to, laptop computers, tablet computers, desktop computers, smart phones, computer terminals and servers, storage databases, and other electronic devices using circuit systems and programmable memory, such as household appliances, smart televisions, digital video disk (DVD) players, heating, ventilation and air conditioning (HVAC) controllers, and light switches, etc.
[0116] The above detailed description includes reference to the accompanying drawings that form a part of the detailed description. The accompanying drawings illustrate specific embodiments that can be put into practice by way of illustration. These embodiments are also referred to as "examples" in this article. Such examples may include elements other than the elements shown or described. However, the inventors also contemplate examples in which only the elements shown or described are provided. In addition, the inventors also contemplate examples using any combination or permutation of the elements shown or described relative to the specific examples (or one or more aspects thereof) or other examples (or one or more aspects thereof) shown or described herein.
[0117] All publications, patents, and patent documents mentioned in this document are incorporated herein by reference in their entirety, as if individually incorporated by reference. In the event of an inconsistency between the usages of this document and a document incorporated by reference, the usage in the incorporated reference(s) should be considered supplementary to that of this document; for irreconcilable inconsistencies, the usage in this document controls.
[0118] In this document, as is common in patent documents, the terms "a" or "an" are used to include one or more than one, independent of any other instance or usage of "at least one" or "one or more". In this document, unless otherwise stated, the term "or" is used to refer to a non-exclusive or, such that "A or B" includes "A but not B", "B but not A", and "A and B". In the appended claims, the terms "including" and "wherein" are used as the plain English equivalents of the respective terms "comprising" and "wherein". Moreover, in the appended claims, the words "including" and "comprising" are open-ended, i.e., systems, devices, articles, or processes that include elements other than those listed after the word in the claim are still deemed to fall within the scope of the claim. In addition, in the appended claims, the terms "first", "second", and "third", etc., are used merely as labels and are not intended to impose numerical requirements on their objects.
[0119] The above description is intended to be illustrative, not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. In reviewing the above description, other embodiments may be used, for example, by a person of ordinary skill in the art. The abstract allows the reader to quickly determine the essence of the technical disclosure, and it is understood that the abstract will not be used to interpret or limit the scope or meaning of the claims. In addition, in the above specific embodiments, various features may be combined together to simplify the present disclosure. This should not be interpreted as indicating that the disclosed features that are not claimed for protection are essential for any claim. On the contrary, the subject matter of the invention may include features less than all the features of a particular disclosed embodiment. Therefore, the claims below are hereby incorporated into the specific embodiments, and each claim is based on itself as a separate embodiment. The scope of the embodiment should be determined with reference to the attached claims and the entire scope given by such claims as being equivalent.
Claims
1. A device comprising: Interface circuit; and a processor circuit coupled to the interface circuit and configured to: Obtaining an initial self-attention input matrix of a transformer model received via the interface circuit; estimating a first execution time of: selecting k number of dominant data elements from the initial self-attention input matrix to generate a sparse self-attention input matrix; estimating a second execution time of performing a self-attention operation of the transformer model based on the sparse self-attention input matrix; estimating a third execution time of performing the self-attention operation based on the initial self-attention input matrix; as well as The self-attention operation is performed based on the first execution time, the second execution time, and the third execution time.
2. The device according to claim 1, wherein: Prior to performing the self-attention operation, the processor circuit is further configured to: Under the condition that a preset accuracy of the self-attention operation is satisfied, a value of the number k for minimizing the sum of the first execution time and the second execution time is determined.
3. The device according to claim 2, wherein: The processor circuit is configured to perform the self-attention operation by: When the sum of the first execution time and the second execution time is less than the third execution time, the self-attention operation is performed based on the sparse self-attention input matrix; otherwise, the self-attention operation is performed based on the initial self-attention input matrix.
4. The device according to any one of claims 1 to 3, wherein: The first execution time includes: a memory access time for data transmission between a memory and a register and a comparison time for data comparison.
5. The device according to any one of claims 1 to 3, wherein: The second execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in the self-attention operation based on the sparse self-attention input matrix.
6. The device according to any one of claims 1 to 3, wherein: The third execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in the self-attention operation based on the initial self-attention input matrix.
7. The device according to any one of claims 1 to 3, wherein: The initial self-attention input matrix includes a query matrix Q, a key matrix K and a value matrix V.
8. The device according to any one of claims 1 to 3, wherein: The number k is greater than or equal to c×lnL Q , where c is a constant sampling factor and L Q The number of rows in the input lookup matrix for the transformer model.
9. A method comprising: estimating a first execution time for performing the following operations: selecting a number k of leading data elements from an initial self-attention input matrix of a transformer model to generate a sparse self-attention input matrix; estimating a second execution time of performing a self-attention operation of the transformer model based on the sparse self-attention input matrix; estimating a third execution time of performing the self-attention operation based on the initial self-attention input matrix; as well as The self-attention operation is performed based on the first execution time, the second execution time, and the third execution time.
10. The method according to claim 9, wherein: Before performing the self-attention operation, the method further includes: Under the condition that a preset accuracy of the self-attention operation is satisfied, a value of the number k for minimizing the sum of the first execution time and the second execution time is determined.
11. The method according to claim 10, wherein: Performing the self-attention operation includes: When the sum of the first execution time and the second execution time is less than the third execution time, the self-attention operation is performed based on the sparse self-attention input matrix; otherwise, the self-attention operation is performed based on the initial self-attention input matrix.
12. The method according to any one of claims 9 to 11, wherein: The first execution time includes: a memory access time for data transmission between a memory and a register and a comparison time for data comparison.
13. The method according to any one of claims 9 to 11, wherein: The second execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in the self-attention operation based on the sparse self-attention input matrix.
14. The method according to any one of claims 9 to 11, wherein: The third execution time includes: a memory access time for data transfer between a memory and a register and a matrix multiplication time for a matrix multiplication operation involved in the self-attention operation based on the initial self-attention input matrix.
15. The method according to any one of claims 9 to 11, wherein: The initial self-attention input matrix includes a query matrix Q, a key matrix K and a value matrix V.
16. The method according to any one of claims 9 to 11, wherein: The number k is greater than or equal to c×lnL Q , where c is a constant sampling factor and L Q The number of rows in the input lookup matrix for the transformer model.
17. A computer readable medium having instructions stored thereon, wherein: The instructions, when executed by a processor circuit, cause the processor circuit to perform a method according to any one of claims 9 to 16.
18. An apparatus comprising means for performing the method according to any one of claims 9 to 16.
Citation Information
Cited By
Coagulant addition prediction method based on fusion of sparse coding and graph space-time attention
CN120911711A
A method for coagulant dosing prediction by fusing sparse coding and graph spatio-temporal attention
CN120911711B