Neural network accelerator and data processing method for neural network accelerator

WO2026200907A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/085587
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2026085587_01102026_PF_FP_ABST
    Figure CN2026085587_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and provides a neural network accelerator and a data processing method for the neural network accelerator. The accelerator comprises a matrix computation unit and at least one vector computation unit. The matrix computation unit comprises a matrix computation core and a matrix post-processing module, and a direct connection path is established between the matrix post-processing module and the at least one vector computation unit. The matrix computation core is used for performing matrix computation in a neural network model, a result of the matrix computation is a first matrix, and the first matrix is divided according to the number of vector computation units to obtain first sub-matrices corresponding to the respective vector computation units. The matrix post-processing module is used for sending, via the direct connection path with the vector computation unit, a first sub-matrix corresponding to the vector computation unit. The vector computation unit is used for performing vector computation on the basis of the received first sub-matrix. The present application can improve the training and inference efficiency of large language models. The present application is used for data processing for neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

Neural network accelerators and their data processing methods

[0001] This application claims priority to Chinese Patent Application No. 202510390670.7, filed on March 28, 2025, entitled “Neural Network Accelerator and Data Processing Method for Neural Network Accelerator”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of data processing technology, and in particular to a neural network accelerator and a data processing method for the neural network accelerator. Background Technology

[0003] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Neural networks (NNs), as an important branch of AI, are network structures that mimic the behavioral characteristics of animal neural networks for information processing. They are widely used in various fields such as image classification, image segmentation, object recognition, image enhancement, and audio recognition.

[0004] Large language models (LLMs) based on the Transformer neural network architecture are increasingly being used in fields such as chatbots, translation, search, text-to-image processing, and video parsing. Accelerated training and inference of LLMs in computing clusters are becoming a crucial technology enabling intelligent transformation of production ecosystems across various industries.

[0005] Among them, fusion operator (FO) technology is a deep learning model optimization technique that combines multiple independent computational operations (operators) into a composite operation, reducing intermediate data storage and memory access overhead during the computation process, thereby significantly improving computational efficiency and reducing GPU memory usage. To further accelerate the training and inference of large language models, improving the performance and energy efficiency of fusion operators has become a pressing issue. Summary of the Invention

[0006] This application provides a neural network accelerator and a data processing method for the neural network accelerator, which can improve the performance and energy efficiency of the fusion operator, thereby further improving the training and inference efficiency of large language models.

[0007] In a first aspect, this application provides a neural network accelerator, which includes a matrix computation unit and at least one vector computation unit. The matrix computation unit includes a matrix computation kernel and a matrix post-processing module. The matrix post-processing module is connected to the at least one vector computation unit via a direct connection. The matrix computation kernel is used to perform matrix computations in the neural network model. The result of the matrix computation is a first matrix. The first matrix is ​​divided according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit. The matrix post-processing module is used to send the first sub-matrix corresponding to the vector computation unit through the direct connection with the vector computation unit. The vector computation unit is used to perform vector computations based on the received first sub-matrix.

[0008] In this embodiment, since a direct connection is established between the matrix post-processing module and at least one vector computation unit, data interaction between the matrix computation unit and at least one vector computation unit can be achieved with a single instruction. The matrix computation unit and the vector computation unit do not need to exchange data and signals through a bus, and there is no bandwidth contention between them. This improves the data interaction efficiency between the matrix computation unit and the vector computation unit, thereby improving the performance and energy efficiency of the fusion operator, and further enhancing the training and inference efficiency of large language models. Furthermore, this embodiment completes data and signal interaction within the subsystem using on-chip static random-access memory (SRAM), reducing the number of bus accesses by the fusion operator compared to related technologies, further improving the performance and energy efficiency of the fusion operator.

[0009] In one possible implementation, the matrix computation kernel includes a cache module, in which a first matrix is ​​stored. The first matrix is ​​divided along a target direction according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit. The target direction includes either the matrix row direction or the matrix column direction. A matrix post-processing module is specifically used to alternately read the first sub-matrix corresponding to each vector computation unit from the cache module.

[0010] For example, assuming the bandwidth of the caching module is a first bandwidth, the matrix post-processing module can read data of the first bandwidth size from the caching module each time.

[0011] For example, assuming there are two vector computation units, the first matrix can be divided into two first sub-matrices along the central axis of the target direction.

[0012] In this embodiment of the application, by dividing the first matrix, a single instruction can simultaneously complete the data transmission from one matrix calculation unit to at least one vector calculation unit, thereby simplifying instruction programming and improving the performance of the fusion operator.

[0013] In one possible implementation, the matrix post-processing module is further configured to perform a first data rearrangement on each read first submatrix; the vector calculation unit is further configured to perform a second data rearrangement on the first submatrix after the first data rearrangement; and the vector calculation unit is specifically configured to perform vector calculation based on the first submatrix after the second data rearrangement.

[0014] For example, the matrix post-processing module reads multiple matrix blocks of a first bandwidth size each time, and these multiple matrix blocks of the first bandwidth size read each time are called a matrix block group. The result of performing the first data rearrangement on the first submatrix is ​​that the elements in a single matrix block within each matrix block group are consecutive, and in every two adjacent matrix blocks, the last element of the previous matrix block is consecutive with the first element of the next matrix block. Furthermore, in two adjacent matrix block groups, the last element of the previous matrix block group is consecutive with the first element of the next matrix block group.

[0015] The vector computation unit rearranges the first submatrix after rearranging the first received data, and the result of the second data rearrangement is that the elements in each column are consecutive, and the last element of the first column and the first element of the second column are consecutive in two adjacent columns.

[0016] In this embodiment of the application, by rearranging the first data and the second data, the data interaction and signal interaction performance between the matrix calculation unit and the vector calculation unit can be effectively improved, thereby effectively improving the performance of the fusion operator technology.

[0017] In one possible implementation, the matrix post-processing module is also used to send a synchronization signal to the vector computation unit via a direct connection to the vector computation unit.

[0018] In this embodiment, the matrix post-processing unit and the vector calculation unit can achieve high-speed synchronous signal interaction through a direct connection path. Compared with related technologies, this improves the signal interaction efficiency, thereby improving the performance and energy efficiency of the fusion operator, and further improving the training and inference efficiency of the large language model.

[0019] Secondly, this application provides a data processing method for a neural network accelerator. The accelerator includes a matrix computation unit and at least one vector computation unit. The matrix computation unit includes a matrix computation kernel and a matrix post-processing module. A direct connection is established between the matrix post-processing module and the at least one vector computation unit. The method includes: the matrix computation kernel performing matrix computation in a neural network model, the result of which is a first matrix; the first matrix being divided according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit; the matrix post-processing module sending the first sub-matrix corresponding to the vector computation unit through the direct connection with the vector computation unit; and the vector computation unit performing vector computation based on the received first sub-matrix.

[0020] In one possible implementation, the matrix computation kernel includes a cache module, a first matrix is ​​stored in the cache module, and the first matrix is ​​divided equally along the target direction according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit. The target direction includes the matrix row direction or the matrix column direction. The process of the matrix post-processing module sending the first sub-matrix corresponding to the vector computation unit through a direct connection with the vector computation unit includes: the matrix post-processing module alternately reading the first sub-matrix corresponding to each vector computation unit from the cache module.

[0021] In one possible implementation, the method further includes: the matrix post-processing module performing a first data rearrangement on each read first submatrix; the vector calculation unit performing a second data rearrangement on the received first data rearranged first submatrix; the process of the vector calculation unit performing vector calculation based on the received first submatrix includes: the vector calculation unit performing vector calculation based on the first submatrix after the second data rearrangement.

[0022] In one possible implementation, the method further includes: the matrix post-processing module sending a synchronization signal to the vector computation unit via a direct connection with the vector computation unit.

[0023] Thirdly, this application provides a processing apparatus for a neural network accelerator, comprising: a matrix computation module, a transceiver module, and a vector computation module. The accelerator includes a matrix computation unit and at least one vector computation unit. The matrix computation unit includes a matrix computation kernel and a matrix post-processing module. A direct connection is established between the matrix post-processing module and the at least one vector computation unit. The matrix computation module is used to perform matrix computations in the neural network model. The result of the matrix computation is a first matrix, which is divided according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit. The transceiver module is used to send the first sub-matrix corresponding to the vector computation unit through the direct connection between the matrix post-processing module and the vector computation unit. The vector computation module is used to perform vector computation based on the received first sub-matrix.

[0024] In one possible implementation, the matrix computation kernel includes a cache module, in which a first matrix is ​​stored. The first matrix is ​​divided along a target direction according to the number of vector computation units to obtain a first sub-matrix corresponding to each vector computation unit. The target direction includes either the matrix row direction or the matrix column direction. A transceiver module is used to alternately read the first sub-matrix corresponding to each vector computation unit from the cache module.

[0025] In one possible implementation, the matrix calculation module is further configured to perform a first data rearrangement on each read first submatrix; the vector calculation module is further configured to perform a second data rearrangement on the first submatrix after the first data rearrangement; and the vector calculation module is specifically configured to perform vector calculation based on the first submatrix after the second data rearrangement.

[0026] In one possible implementation, the transceiver module is also used to send a synchronization signal to the vector computation unit via a direct connection between the matrix post-processing module and the vector computation unit.

[0027] Fourthly, this application provides an electronic device comprising: one or more processors; a memory for storing one or more computer programs or instructions; and, when the one or more computer programs or instructions are executed by the one or more processors, causing the one or more processors to implement the method as described in any of the second aspects.

[0028] Fifthly, this application provides an electronic device including a processor for performing the method as described in any one of the second aspects.

[0029] In a sixth aspect, this application provides an electronic device comprising: a processing circuit and an interface circuit; wherein the interface circuit is used to couple with a memory external to the electronic device and to provide a communication interface for the processing circuit to access the memory; the processing circuit is used to execute program instructions in the memory to implement the method as described in any of the second aspects.

[0030] In practical implementation, the electronic device can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, gate circuit, flip-flop, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be output to, for example, but not limited to, a transmitter and transmitted by the transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.

[0031] In a seventh aspect, this application provides a computer-readable storage medium storing program code, which, when executed by a processor, implements the method as described in any one of the second aspects.

[0032] Eighthly, this application provides a chip comprising: at least one processor. The at least one processor is configured to perform the method as described in any one of the second aspects.

[0033] Optionally, the chip also includes memory. At least one processor is used to execute code in the memory, and when the at least one processor executes the code, it causes the chip to implement the method as described in any one of the second aspects.

[0034] Alternatively, the chip described above can also be an integrated circuit.

[0035] Ninthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method as described in any one of the second aspects. Attached Figure Description

[0036] Figure 1 is a schematic diagram of the basic architecture of a large language model provided in an embodiment of this application;

[0037] Figure 2 is a schematic diagram of the computational flow of a fast attention mechanism provided in an embodiment of this application;

[0038] Figure 3 is a schematic diagram of the structure of a neural network accelerator provided in an embodiment of this application;

[0039] Figure 4 is a schematic diagram of a matrix post-processing module provided in an embodiment of this application;

[0040] Figure 5 is a schematic diagram of a data rearrangement provided in an embodiment of this application;

[0041] Figure 6 is a schematic diagram of the division of a first matrix provided in an embodiment of this application;

[0042] Figure 7 is a schematic diagram of another first matrix partitioning provided in an embodiment of this application;

[0043] Figure 8 is a schematic diagram of a first data rearrangement and a second data rearrangement provided in an embodiment of this application;

[0044] Figure 9 is a schematic diagram of another first data rearrangement and second data rearrangement provided in the embodiments of this application;

[0045] Figure 10 is a flowchart illustrating a data processing method for a neural network accelerator provided in an embodiment of this application;

[0046] Figure 11 is a schematic diagram of the computational data flow of a fusion operator provided in an embodiment of this application;

[0047] Figure 12 is a block diagram of a processing device for a neural network accelerator provided in an embodiment of the application;

[0048] Figure 13 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0051] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0052] This application embodiment can be applied to neural network models, such as large language models. Please refer to Figure 1, which is a schematic diagram of the basic architecture of a large language model provided by this application embodiment. Figure 1 shows the flash attention module, which is used to perform a series of calculations based on the input vector and output the results. The input vector includes: query (Q), key (K), and value (V). As shown in Figure 1, Q, K, and V are first transposed. Then, the flash attention module performs matrix calculations on the transposed Q and K, and performs vector calculations (SoftMax) on the matrix calculation results to obtain attention weights. Then, the model parameters are locally updated based on the attention weights, and regularization (dropout) is performed. Then, the attention weights are weighted and summed with V to obtain the attention context (AttnContext). The model parameters are globally updated based on the attention context, and the attention context is transposed and output (Attn Output).

[0053] The output of the fast attention module is regularized, then subjected to residual connections (Add) and layer normalization (Layer normalization). The results of residual connections and layer normalization are then fed into a feedforward neural network (FNN) or a mixture of experts (MOE) model. The input to the FNN or MOE model is processed by an intermediate layer and an activation function (Gelu), and the output is then added together (Add).

[0054] The architecture of the large language model shown in Figure 1 is for illustrative purposes only, and the embodiments of this application do not limit the architecture of the large language model used.

[0055] Referring to the description in Figure 1, the fast attention module needs to use the results of matrix calculations during vector computation. Typically, matrix and vector calculations are performed on different cores, thus requiring data and signal interaction between the matrix and vector calculation units. This application provides a neural network accelerator that improves the efficiency of data and signal interaction between the matrix and vector calculation units, thereby accelerating the hardware computation of fusion operators in large language model training or inference. The fusion operator includes the fast attention technique shown in Figure 1, which will be described below. Please refer to Figure 2, which is a schematic diagram of the computation flow of a fast attention mechanism provided in this application embodiment. The fast attention technique divides the matrix and vector calculation blocks into smaller granularities and sends them to different cores (e.g., different AI acceleration cores) for computation. As shown in Figure 2, Q is divided into 8 Q blocks and sent to 8 cores. Each core performs the computation of Q blocks with K and V in parallel, obtaining the final output O.

[0056] For example, please refer to Figure 3, which is a schematic diagram of the structure of a neural network accelerator provided in an embodiment of this application. The neural network accelerator can execute part or all of the computation process of a neural network model. The neural network accelerator of this embodiment can be deployed on cloud service devices or terminal devices, such as computers, servers, vehicles, drones, or mobile phones. For example, the neural network accelerator can be deployed on a system-on-a-chip (SoC) in a terminal device such as a mobile phone, such as the AI ​​processor module in a cloud processor SoC. The AI ​​processor can be, for example, a neural network processing unit (NPU).

[0057] The neural network accelerator can be set up in the fast attention module shown in Figure 1 to complete the data and signal interaction between the matrix calculation unit and the vector calculation unit.

[0058] The neural network accelerator includes a matrix multiplication unit (MM Unit) and at least one vector computation unit (Vec Unit). Figure 3 shows two vector computation units, Vec Unit0 and Vec Unit1. The matrix multiplication unit may include a matrix multiplication core (MM Core) and a matrix post-processing module, with a direct connection established between the matrix post-processing module and at least one vector computation unit.

[0059] A matrix computation kernel is used to perform matrix computations in a neural network model. Exemplarily, matrix computations include matrix multiplication or convolution operations. For example, a matrix computation kernel can be used to perform multiplication operations in a fully connected layer or convolution operations in a convolutional layer of a neural network model.

[0060] For example, the matrix computation core can be implemented using a matrix computation core found in an existing neural network processor. For instance, the matrix computation core can perform matrix computations using the matrix computation section of the fast attention module in Figure 1. This application does not limit the specific structure of the matrix computation core, as long as the core is capable of performing matrix computations in the neural network model.

[0061] The matrix post-processing module is used to implement data and signal interaction with the vector computation unit. The vector computation unit is used to perform vector operations in the neural network model based on the results of matrix computation.

[0062] The matrix post-processing module can obtain the matrix calculation result from the source address (source, src), process the matrix calculation result, and send the processed data to the destination address (destination, dst).

[0063] In this context, `src` can be one or more, and `dst` can be one or more. The configuration of `src` and `dst` depends on the hardware architecture of the neural network accelerator. Furthermore, the matrix post-processing module can also obtain the parameters for vector operations from the addresses of the parameters used in vector computation. Any of the following can be indicated by parameters in the instructions or registers in the post-processing unit: the address of `dst`, `src`, or the parameters for the vector operations.

[0064] For example, as shown in Figure 1, the matrix computation kernel may include a matrix computation module and a first memory. The matrix computation module is used to perform matrix computations in a neural network model.

[0065] The first memory can be used to store data related to matrix calculations. For example, the first memory can be used to store at least one of the following: input data for matrix calculations, parameters of a neural network model, or results of matrix calculations. The parameters of the neural network model include weight parameters for matrix calculations.

[0066] It should be noted that the first memory may also include multiple memories, which are used to store the input data for matrix calculation, the parameters of the neural network model, and the results of matrix calculation, respectively.

[0067] For example, the first memory includes an input memory and a result memory. The input memory stores the input data for matrix computation and the parameters of the neural network model, providing the input matrix to the matrix computation unit. For example, the input memory can be a data buffer. The input matrix can include a left matrix and a right matrix, corresponding to the input data and weight parameters for the matrix computation, respectively. The result memory stores the results of the matrix computation. For example, the result memory can include a result buffer.

[0068] For example, as shown in Figure 3, the first memory may include: a level 1 buffer (L1 Buffer), a level 0 AB buffer (L0 AB Buffer), and a level 0 C buffer (L0 C Buffer). The level 1 buffer is used to store the input data for matrix calculation, the level 0 AB buffer is used to store the parameters of the neural network model, and the level 0 C buffer is used to store the results of matrix calculation.

[0069] The matrix post-processing module is used to retrieve the results of matrix calculations from a first memory, such as a level-zero C buffer. In this case, src may include the level-zero C buffer in the first memory.

[0070] Please refer to Figure 4, which is a schematic diagram of a matrix post-processing module provided in an embodiment of this application. The matrix post-processing module may include a data rearrangement submodule, a second memory, a format conversion submodule (transpose buffer, TP BUF), a data selection submodule, and a first scalar unit (SU). The data selection submodule may be a data multiplexer (DMUX), and the second memory may be a column buffer (Col BUF). The format conversion submodule may include a matrix transpose circuit implemented by a register file. The data rearrangement submodule may be a cross bar.

[0071] For example, the vector computation unit may include a third memory, a second scalar unit, and a vector computation module. The third memory is used to store data related to vector operations (e.g., the results of matrix computations). For example, after the matrix post-processing module obtains the results of matrix computations, it can write the results into the third memory. The third memory can be a unified buffer (UB).

[0072] In this system, a direct path is established between the data selection submodule and each vector calculation unit, and a direct path is established between the first scalar unit and each second scalar unit.

[0073] It should be noted that the neural network accelerator shown in Figure 3 or Figure 4 is merely an example. In specific implementations, those skilled in the art should understand that the accelerator in Figure 3 or Figure 4 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the accelerator in Figure 3 or Figure 4 may also include hardware devices for implementing other additional functions. Moreover, those skilled in the art should understand that the accelerator in Figure 3 or Figure 4 may only include the devices necessary for implementing the embodiments of this application, and does not necessarily include all the devices shown in Figure 3 or Figure 4.

[0074] For example, a matrix computation kernel performs matrix computations in a neural network model. The result of the matrix computation is a first matrix, which is then divided according to the number of vector computation units to obtain a first submatrix corresponding to each vector computation unit. A matrix post-processing module sends the first submatrix corresponding to each vector computation unit through a direct connection to the vector computation unit. The vector computation unit performs vector computations based on the received first submatrix.

[0075] In one possible implementation, the matrix computation kernel includes a cache module (which may be the aforementioned first memory, such as a level-zero C buffer in the first memory), and the first matrix is ​​stored in the cache module. The first matrix is ​​divided equally along the target direction according to the number of vector computation units, resulting in a first sub-matrix corresponding to each vector computation unit. The target direction may include the matrix row direction or the matrix column direction. A matrix post-processing module is specifically used to alternately read the first sub-matrix corresponding to each vector computation unit from the cache module.

[0076] The matrix post-processing module can read data from the cache module according to the bandwidth of the cache module, that is, it reads data of the cache module's bandwidth size each time. The bandwidth of the cache module is a first bandwidth. When alternately reading the first sub-matrix corresponding to each vector calculation unit, it can sequentially read the elements of the first bandwidth size in the first sub-matrix corresponding to each vector calculation unit. Then, it continues to sequentially read the elements of the first bandwidth size in the first sub-matrix corresponding to each vector calculation unit until the first sub-matrix corresponding to each vector calculation unit has been read. The bandwidth of the cache module is configurable, and this embodiment does not limit it.

[0077] For example, assuming there are two vector computation units, the first matrix can be divided into two first sub-matrices along the central axis of the target direction.

[0078] By dividing the first matrix, a single instruction can simultaneously transmit data from one matrix calculation unit to at least one vector calculation unit, thereby simplifying instruction programming and improving the performance of the fusion operator.

[0079] For example, as shown in Figure 4, the data rearrangement submodule can rearrange the data read from the cache module and then cache it in the second memory. For instance, the data rearrangement submodule can rearrange the data according to the size of the matrix block. Please refer to Figure 5, which is a schematic diagram of data rearrangement provided in an embodiment of this application. Each time the data rearrangement submodule reads from the cache module, it rearranges the read data from four sets of 4 rows and 4 columns of elements into one set of 4 rows and 16 columns of elements. In this embodiment of the application, one row of 16 columns of elements can be referred to as a matrix block.

[0080] In one possible implementation, please refer to Figure 6, which is a schematic diagram of the partitioning of a first matrix provided in an embodiment of this application. Figure 6 illustrates the first matrix with the target direction as the matrix row direction as an example. Figure 6 shows the first matrix, where one matrix block represents 1 row and 16 elements. It is assumed that the data size of four matrix blocks is the same as the bandwidth of the cache module. For example, assuming that the size of one element is 32 bits and the bandwidth of the cache module is 256 bytes, then the size of one matrix block is 64 bytes, and the size of four matrix blocks is 256 bytes.

[0081] As shown in Figure 6, the first matrix is ​​divided into two first sub-matrices along the row direction (i.e., the central axis of side M). Each first sub-matrix includes 8 rows and 64 columns, or 32 matrix blocks. Rows 1 to 8 are the first sub-matrix corresponding to one vector calculation unit, and rows 9 to 16 are the first sub-matrix corresponding to another vector calculation unit. The matrix post-processing module can read the first 16 columns of elements from rows 1 to 4 (the first column matrix block of rows 1 to 4) from the cache module in the first instance. It can read the first 16 columns of elements from rows 9 to 12 (the first column matrix block of rows 9 to 12) from the cache module in the second instance. It can read the first 16 columns of elements from rows 5 to 8 (the first column matrix block of rows 5 to 8) from the cache module in the fourth instance. It can read the first 16 columns of elements from rows 13 to 16 (the first column matrix block of rows 13 to 16) from the cache module in the fourth instance. The reading process of the matrix blocks in columns 2 to 4 can refer to the reading process of the matrix block in column 1, and will not be described in detail here in the embodiments of this application.

[0082] Please refer to Figure 7, which is a schematic diagram of another first matrix partitioning provided in an embodiment of this application. Figure 7 illustrates the first matrix with the target direction as the matrix column direction as an example. Figure 7 shows the first matrix, where one matrix block represents 1 row and 16 elements. It is assumed that the data size of the four matrix blocks is the same as the bandwidth of the cache module.

[0083] As shown in Figure 7, the first matrix is ​​divided into two first sub-matrices along the column direction (i.e., the central axis of the N sides). Each first sub-matrix includes 12 rows and 32 columns, or 24 matrix blocks. Columns 1 and 2 correspond to the first sub-matrix of one vector calculation unit, and columns 3 and 4 correspond to the first sub-matrix of another vector calculation unit. The matrix post-processing module can first read the first 16 columns of elements from rows 1 to 4 from the cache module, i.e., the first column matrix block of rows 1 to 4. Secondly, it can read the 33rd to 48th columns of elements from rows 1 to 4 from the cache module, i.e., the third column matrix block of rows 1 to 4. Thirdly, it can read the first 16 columns of elements from rows 5 to 8 from the cache module, i.e., the first column matrix block of rows 5 to 8. Fourthly, it can read the 33rd to 48th columns of elements from rows 5 to 8 from the cache module, i.e., the third column matrix block of rows 5 to 8. The process of reading the remaining elements of the two first submatrices can refer to the aforementioned reading process, and will not be repeated here in the embodiments of this application.

[0084] In one possible implementation, the matrix post-processing module is further configured to perform a first data rearrangement on each read first submatrix. The vector calculation unit is further configured to perform a second data rearrangement on the received first submatrix after the first data rearrangement. Specifically, the vector calculation unit is configured to perform vector calculations based on the first submatrix after the second data rearrangement.

[0085] As shown in Figure 4, the matrix post-processing module can perform the first data rearrangement on each first submatrix read through the format conversion submodule.

[0086] For example, the matrix post-processing module reads multiple matrix blocks of a first bandwidth size each time, and these multiple matrix blocks of the first bandwidth size read each time are called a matrix block group. The result of performing the first data rearrangement on the first submatrix is ​​that the elements in a single matrix block within each matrix block group are consecutive, and in every two adjacent matrix blocks, the last element of the previous matrix block is consecutive with the first element of the next matrix block. Furthermore, in two adjacent matrix block groups, the last element of the previous matrix block group is consecutive with the first element of the next matrix block group.

[0087] The vector computation unit rearranges the first submatrix after rearranging the first received data, and the result of the second data rearrangement is that the elements in each column are consecutive, and the last element of the first column and the first element of the second column are consecutive in two adjacent columns.

[0088] For example, please refer to Figure 8, which is a schematic diagram of a first data rearrangement and a second data rearrangement provided in an embodiment of this application. Figure 8 is illustrated using the partitioning method shown in Figure 6 as an example. Figure 8(a) shows the matrix shown in Figure 6 transposed, where a matrix block is transposed from the original 16 columns to 16 rows. Columns 1 to 8 are the first submatrix corresponding to one vector calculation unit, and columns 9 to 16 are the first submatrix corresponding to another vector calculation unit.

[0089] Specifically, the four matrix block groups in the first row represent the 1st, 3rd, 2nd, and 4th reads from the cache module, respectively. The four matrix block groups in the second row represent the 5th, 7th, 6th, and 8th reads from the cache module, respectively. The four matrix block groups in the third row represent the 9th, 11th, 10th, and 12th reads from the cache module, respectively. The four matrix block groups in the fourth row represent the 9th, 11th, 10th, and 12th reads from the cache module, respectively.

[0090] Figure 8(a) shows the result after the matrix post-processing module performs the first data rearrangement: the elements in each first submatrix are continuous in the direction indicated by the arrows. Figure 8(b) shows the result after a vector computation unit performs the second data rearrangement on the received first submatrix: the first submatrix corresponding to this vector computation unit is continuous in the direction indicated by the arrows. Figure 8(c) shows the result after another vector computation unit performs the second data rearrangement on the received first submatrix: the first submatrix corresponding to this vector computation unit is continuous in the direction indicated by the arrows.

[0091] It should be noted that Figure 8 exemplarily shows the results of the first and second data rearrangements of the first four rows in the first matrix, and the rearrangement of the matrix block groups read from the cache module for the first to fourth times is represented by different filling methods. Other parts can be referred to accordingly, and the embodiments of this application will not be described in detail here.

[0092] For example, please refer to Figure 9, which is a schematic diagram of another first data rearrangement and second data rearrangement provided in the embodiments of this application. Figure 9 is illustrated using the partitioning method shown in Figure 7 as an example. Figure 9(a) shows the matrix shown in Figure 7 transposed, where a matrix block is transposed from the original 16 columns to 16 rows. Rows 1 to 4 are the first sub-matrix corresponding to one vector calculation unit, and rows 5 to 8 are the first sub-matrix corresponding to another vector calculation unit.

[0093] Among them, the matrix block groups in rows 1 to 8 are read from the cache module for the 1st, 3rd, 5th, 7th, 2nd, 4th, 6th and 8th times, respectively.

[0094] Figure 9(a) shows the result after the matrix post-processing module performs the first data rearrangement: the elements in each first sub-matrix are continuous in the direction indicated by the arrows. Figure 9(b) shows the result after a vector computation unit performs the second data rearrangement on the received first sub-matrix: the first sub-matrix corresponding to this vector computation unit is continuous in the direction indicated by the arrows. Figure 9(c) shows the result after another vector computation unit performs the second data rearrangement on the received first sub-matrix: the first sub-matrix corresponding to this vector computation unit is continuous in the direction indicated by the arrows.

[0095] It should be noted that Figure 9 illustrates the rearrangement of the matrix block groups read from the cache module from the 1st to the 4th time using different filling methods. Other parts can be referred to accordingly, and the embodiments of this application will not be described in detail here.

[0096] In this embodiment, for the format conversion submodule, the arrangement efficiency of each first submatrix is ​​higher after the first data rearrangement. For the vector calculation unit, the arrangement efficiency of each first submatrix is ​​higher after the second data rearrangement. Therefore, by rearranging the first and second data, the data interaction and signal interaction performance between the matrix calculation unit and the vector calculation unit can be effectively improved, thereby effectively improving the performance of the fusion operator technology.

[0097] After completing the first data rearrangement, the format conversion submodule sends the corresponding first submatrix to each vector computation unit. In one possible implementation, the size of the read port of the cache module of the matrix post-processing module can be equal to the size of the write ports of the multiple vector computation units of the matrix post-processing module. That is, the matrix post-processing unit reads data of the first bandwidth size from the cache module each time and sends the first bandwidth size of data to each vector computation unit each time. As shown in Figure 8 or Figure 9, the data is written to the format conversion submodule in the order rearranged by the data rearrangement submodule. When writing 4 matrix blocks along the N direction, the format conversion submodule outputs 4 rows of 256 bytes of data (i.e., 4 rows of matrix block groups) and writes them to two vector computation units simultaneously in 4 steps. It should be noted that the format conversion submodule outputs one row of matrix block group at a time and writes it to the corresponding vector computation unit.

[0098] In one possible implementation, the matrix post-processing module is further configured to send a synchronization signal to the vector computation unit via a direct connection to the vector computation unit. The vector computation unit can also send a synchronization signal to the matrix post-processing module via a direct connection to the matrix post-processing module. As shown in Figure 4, a synchronization signal can be transmitted between the first scalar unit and the second scalar unit. Exemplarily, the synchronization signal can be used to instruct the matrix computation unit to start matrix computation, or it can be used to instruct the vector computation unit to start vector computation after waiting for the first matrix to be completely written to the vector computation unit (enabling the vector computation unit to perform lockstep execution).

[0099] For example, the synchronization signal may include at least one of the following: a pulse signal, a flag identity (flag_id), and a counter. The pulse signal is used to indicate that the signal receiver has started receiving the synchronization signal, the flag_id is used to indicate the computation flow, and each flag_id corresponds to at least one counter. When the counter value becomes 0, it indicates that the message receiver can start the corresponding computation. For example, the synchronization signal may include the SET_INTRA_BLOCK and WAIT_INTRA_BLOCK instructions.

[0100] In this embodiment, the matrix post-processing unit and the vector calculation unit can achieve high-speed synchronous signal interaction through a direct connection path. Compared with related technologies, this improves the signal interaction efficiency, thereby improving the performance and energy efficiency of the fusion operator, and further improving the training and inference efficiency of the large language model.

[0101] This application provides a data processing method for a neural network accelerator. The description of the accelerator can be found in the foregoing embodiments, and will not be repeated here. Please refer to Figure 10, which is a flowchart illustrating a data processing method for a neural network accelerator provided in this application. This method may include the following processes:

[0102] 1001. The matrix calculation kernel performs matrix calculations in the neural network model. The result of the matrix calculation is the first matrix. The first matrix is ​​divided according to the number of vector calculation units to obtain the first sub-matrix corresponding to each vector calculation unit.

[0103] In one possible implementation, the matrix computation kernel includes a cache module, in which a first matrix is ​​stored. The first matrix is ​​divided equally along the target direction according to the number of vector computation units, resulting in a first submatrix corresponding to each vector computation unit. The target direction includes either the matrix row direction or the matrix column direction.

[0104] 1002. The matrix post-processing module sends the first sub-matrix corresponding to the vector calculation unit through a direct connection with the vector calculation unit.

[0105] In one possible implementation, the matrix post-processing module alternately reads the first sub-matrix corresponding to each vector computation unit from the cache module. Then, it sends the first sub-matrix corresponding to the vector computation unit through a direct connection to the vector computation unit.

[0106] 1003. The vector calculation unit performs vector calculations based on the received first submatrix.

[0107] In one possible implementation, the matrix post-processing module can further perform a first data rearrangement on each read first submatrix. The vector calculation unit can further perform a second data rearrangement on the received first submatrix after the first data rearrangement. The vector calculation unit then performs vector calculations based on the first submatrix after the second data rearrangement.

[0108] In one possible implementation, the matrix post-processing module can also send a synchronization signal to the vector computation unit via a direct connection to the vector computation unit.

[0109] The implementation of this method and the effects it can achieve can be found in the description of the aforementioned neural network accelerator, and will not be repeated here in the embodiments of this application.

[0110] The following describes the computational data flow of the fusion operator when applied in a neural network model according to the embodiments of this application. Please refer to Figure 11, which is a schematic diagram of the computational data flow of a fusion operator provided in an embodiment of this application. As shown in Figure 11, the memory transfer engine (MTE) sequentially reads the input vectors Q1, K1, K2, K3, V1, K4, V2, V3, and V4 from the memory into the kernel (e.g., the AI ​​computation kernel). Here, Qn represents a part of Q after being split, Kn represents a part of K after being split, and Vn represents a part of V after being split.

[0111] After the MTE moves Q1 and K1 into the kernel, the matrix computation kernel performs matrix computation Q1K1 on Q1 and K1, resulting in S1. After the MTE moves K2 into the kernel, the matrix computation kernel performs matrix computation Q1K2 on Q1 and K2, resulting in S2. After the MTE moves K3 into the kernel, the matrix computation kernel performs matrix computation Q1K3 on Q1 and K3, resulting in S3. After the MTE moves K4 into the kernel, the matrix computation kernel performs matrix computation Q1K4 on Q1 and K4, resulting in S4.

[0112] After obtaining S1 from the matrix calculation kernel, the matrix post-processing module of the neural network accelerator provided in this application embodiment divides S1 into matrix calculation results corresponding to vector calculation unit 0 and vector calculation unit 1 respectively, and sends their respective matrix calculation results to vector calculation units 0 and 1 through direct connections. After the transmission is completed, the matrix post-processing module sends a synchronization signal to vector calculation units 0 and 1 through direct connections to instruct them to start vector calculation. Vector calculation units 0 and 1 perform vector calculation SoftMax1 based on the received portion of S1. Similarly, the same process is performed on S2, S3, and S4, with vector calculation units 0 and 1 performing corresponding vector calculations SoftMax2, SoftMax3, and SoftMax4 based on the received portions of S2, S3, and S4.

[0113] After SoftMax1 completes its calculation, the matrix computation core performs a matrix calculation (P1V1) on the outputs P1 and V1 of SoftMax1. After SoftMax2 completes its calculation, the matrix computation core performs a matrix calculation (P2V2) on the outputs P2 and V2 of SoftMax2. After SoftMax3 completes its calculation, the matrix computation core performs a matrix calculation (P3V3) on the outputs P3 and V3 of SoftMax3. After SoftMax4 completes its calculation, the matrix computation core performs a matrix calculation (P4V4) on the outputs P4 and V4 of SoftMax4.

[0114] After the calculations of P1V1 and P2V2 are completed, vector calculation unit 0 and vector calculation unit 1 perform vector calculations G1U1 for the result G1 of P1V1 and the result U1 of P2V2, respectively. Then, they perform vector calculations G2U2 for the result G1U1 and the result U2 of P3V3, respectively. Finally, they perform vector calculations G3U3 for the result G2U2 and the result U3 of P4V4, respectively.

[0115] In summary, the neural network accelerator and its processing method provided in this application include a matrix computation unit and at least one vector computation unit. The matrix computation unit comprises a matrix computation kernel and a matrix post-processing module, with a direct connection established between the matrix post-processing module and the at least one vector computation unit. The matrix computation kernel performs matrix computations in the neural network model, resulting in a first matrix. This first matrix is ​​divided according to the number of vector computation units, yielding a first sub-matrix corresponding to each vector computation unit. The matrix post-processing module sends the first sub-matrix corresponding to each vector computation unit through the direct connection. The vector computation unit performs vector computation based on the received first sub-matrix. Because the matrix post-processing module has a direct connection with the at least one vector computation unit, data interaction between the matrix computation unit and the at least one vector computation unit can be achieved with a single instruction, improving the data interaction efficiency between the matrix computation unit and the vector computation unit. This, in turn, improves the performance and energy efficiency of the fusion operator, further enhancing the training and inference efficiency of large language models. Furthermore, in this embodiment, data and signal interaction is completed within the subsystem using on-chip static random-access memory (SRAM), which reduces the number of bus accesses by the fusion operator compared to related technologies, thereby further improving the performance and energy efficiency of the fusion operator.

[0116] The order of the methods provided in the embodiments of this application can be adjusted appropriately, and the process can also be added or removed as appropriate. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and the embodiments of this application do not limit this.

[0117] The foregoing primarily describes the processing method of the neural network accelerator provided in this application from the perspective of the device. It is understood that, in order to achieve the above functions, the device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] This application embodiment can divide the device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one device. The integrated modules can be implemented in hardware or software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0119] Figure 12 is a block diagram of a processing device for a neural network accelerator according to an embodiment of the application. When each functional module is divided according to its corresponding function, the processing device 1200 of the neural network accelerator may include: a matrix calculation module 1201, a transceiver module 1202, and a vector calculation module 1203. The accelerator includes a matrix calculation unit and at least one vector calculation unit. The matrix calculation unit includes a matrix calculation kernel and a matrix post-processing module. A direct connection is established between the matrix post-processing module and the at least one vector calculation unit. Exemplarily, the processing device of this neural network accelerator may be a neural network accelerator, or a chip within a neural network accelerator, or other combined devices or components having the aforementioned processing device functions of a neural network accelerator. Exemplarily, the processing device 1200 of the neural network accelerator may be a processor (or processing circuit), such as a baseband processor, which may include one or more central processing units (CPUs). The processing device 1200 of the neural network accelerator may also be a processor (or processing circuit) of a chip system, which may include one or more central processing units.

[0120] For example, the matrix calculation module 1201 is used to perform matrix calculations in the neural network model. The result of the matrix calculation is a first matrix. The first matrix is ​​divided according to the number of vector calculation units to obtain a first sub-matrix corresponding to each vector calculation unit. The transceiver module 1202 is used to send the first sub-matrix corresponding to the vector calculation unit through a direct connection between the matrix post-processing module and the vector calculation unit. The vector calculation module 1203 is used to perform vector calculations based on the received first sub-matrix.

[0121] In conjunction with the above scheme, the matrix calculation kernel includes a cache module, a first matrix is ​​stored in the cache module, the first matrix is ​​divided equally along the target direction according to the number of vector calculation units, and a first sub-matrix corresponding to each vector calculation unit is obtained. The target direction includes the matrix row direction or the matrix column direction. The transceiver module 1202 is specifically used to alternately read the first sub-matrix corresponding to each vector calculation unit from the cache module.

[0122] In conjunction with the above scheme, the matrix calculation module 1201 is also used to perform a first data rearrangement on each read first submatrix; the vector calculation module 1203 is also used to perform a second data rearrangement on the first submatrix after the first data rearrangement; the vector calculation module 1203 is specifically used to perform vector calculation based on the first submatrix after the second data rearrangement.

[0123] In conjunction with the above scheme, the transceiver module 1202 is also used to send a synchronization signal to the vector calculation unit through the direct connection path between the matrix post-processing module and the vector calculation unit.

[0124] Optionally, the communication device provided in FIG12 may further include a storage module 1204, mainly used for storing software programs. The matrix calculation module 1201 or the vector calculation module 1203 can call the software program stored in the storage module 1204 to implement the method described in any of the embodiments of this application. The storage module 1204 includes media with storage functions such as hard disk, RAM, and ROM.

[0125] Figure 13 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device 1300 can be a neural network accelerator or a chip or functional module in a neural network accelerator. As shown in Figure 13, the electronic device 1300 includes a processor 1301, a transceiver 1302, and a communication line 1303.

[0126] The processor 1301 is used to execute any step in the aforementioned method embodiments, and when performing processes such as sending the corresponding first sub-matrix to the vector calculation unit, it can selectively call the transceiver 1302 and the communication line 1303 to complete the corresponding operation.

[0127] Furthermore, the electronic device 1300 may also include a memory 1304. The processor 1301, the memory 1304, and the transceiver 1302 can communicate with each other via a communication line 1303.

[0128] Transceiver 1302 is used to communicate with other devices or other communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Transceiver 1302 can be a module, circuit, transceiver, or any device capable of enabling communication.

[0129] The transceiver 1302 is mainly used for sending and receiving data, and may include a transmitter and a receiver to send and receive data, respectively. For example, a neural network model can be obtained through the transceiver 1302. Operations other than sending and receiving data are implemented by the processor, such as performing matrix calculations or vector calculations.

[0130] The communication line 1303 may include a path for transmitting information between various components of the device 1300 (e.g., processor 1301, transceiver 1302, and memory 1304).

[0131] In one design, the processor can be viewed as a logic circuit, and the transceiver as an interface circuit.

[0132] Memory 1304 is used to store instructions. These instructions may be computer programs. When the instructions stored in memory 1304 are executed by processor 1301, processor 1301 performs various steps of the data processing method of this embodiment. Specifically, processor 1301 may execute methods 1001 to 1003 described above.

[0133] The memory 1304 can be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).

[0134] It should be noted that the memory 1304 can exist independently of the processor 1301, or it can be integrated with the processor 1301. The memory 1304 can be used to store instructions, program code, or some data, etc. The memory 1304 can be located inside or outside the electronic device 1300, without restriction.

[0135] The processor 1301 includes the matrix calculation unit and the vector calculation unit shown in Figure 4.

[0136] The processor 1301 may be a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute relevant programs to implement the data processing method of the method embodiment of this application.

[0137] The processor 1301 can also be an integrated circuit chip with signal processing capabilities.

[0138] The processor 1301 described above can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 1304. The processor 1301 reads the information in memory 1304 and, in conjunction with its hardware, completes the functions required by the units included in the device shown in FIG4, or executes the data processing method of the method embodiments of this application.

[0139] As an optional implementation, the electronic device 1300 also includes an output device 1305 and an input device 1306. For example, the input device 1306 is a device such as a keyboard, mouse, microphone, or joystick, and the output device 1305 is a device such as a display screen or speaker.

[0140] It should be noted that the electronic device 1300 can be a chip system or a device with a similar structure to that shown in FIG13. The chip system may include chips or other discrete devices. Actions, terms, etc., involved in the various embodiments of this application can be referred to interchangeably without limitation. The message names or parameter names in the messages used for interaction between devices in the embodiments of this application are merely examples; other names may be used in specific implementations without limitation. Furthermore, the composition structure shown in FIG13 does not constitute a limitation on the electronic device 400. In addition to the components shown in FIG13, the electronic device 1300 may include more or fewer components than shown in FIG13, or combine certain components, or have different component arrangements.

[0141] The processor and transceiver described in this application can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits, mixed-signal ICs, application-specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductors (CMOS), n-metal-oxide-semiconductor (NMOS), p-type metal oxide semiconductors (PMOS), bipolar junction transistors (BJTs), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0142] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute any of the methods described in the embodiments of this application.

[0143] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by a computer or a communication-enabled device using computer programs or instructions to control related hardware. The computer program or set of instructions can be stored in the aforementioned computer-readable storage medium. When executed, the computer program or set of instructions can include the processes described in the above method embodiments. The computer-readable storage medium can be an internal storage unit of the terminal device in any of the foregoing embodiments, such as the hard disk or memory of the terminal device. The aforementioned computer-readable storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Further, the aforementioned computer-readable storage medium can include both internal storage units and external storage devices of the terminal device. The aforementioned computer-readable storage medium is used to store the aforementioned computer program or instructions, as well as other programs and data required by the terminal device. The aforementioned computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0144] This application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in the memory through the data interface and executes the data processing method of this application.

[0145] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the data processing method in the embodiments of this application.

[0146] This application also provides a system-on-a-chip (SoC), which includes the neural network accelerator described in this application.

[0147] This application also provides an electronic device, which includes the neural network accelerator described in this application.

[0148] It should be understood that the processor in the embodiments of this application can be a CPU, but it can also be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0149] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache.

[0150] By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0151] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0155] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0156] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0157] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A neural network accelerator, characterized in that, The accelerator includes a matrix computing unit and at least one vector computing unit. The matrix computing unit includes a matrix computing kernel and a matrix post-processing module. A direct path is established between the matrix post-processing module and the at least one vector computing unit. The matrix calculation kernel is used to perform matrix calculations in the neural network model. The result of the matrix calculation is a first matrix. The first matrix is ​​divided according to the number of vector calculation units to obtain a first sub-matrix corresponding to each vector calculation unit. The matrix post-processing module is used to send the first sub-matrix corresponding to the vector calculation unit through a direct connection with the vector calculation unit; The vector calculation unit is used to perform vector calculations based on the received first submatrix.

2. The accelerator according to claim 1, characterized in that, The matrix calculation kernel includes a cache module, the first matrix is ​​stored in the cache module, and the first matrix is ​​divided equally along the target direction according to the number of vector calculation units to obtain the first sub-matrix corresponding to each vector calculation unit. The target direction includes the matrix row direction or the matrix column direction. The matrix post-processing module is specifically used to alternately read the first sub-matrix corresponding to each vector calculation unit from the cache module.

3. The accelerator according to claim 1 or 2, characterized in that, The matrix post-processing module is also used to perform a first data rearrangement on each of the first sub-matrices read; The vector calculation unit is further configured to perform a second data rearrangement on the first submatrix after rearranging the received first data; The vector calculation unit is specifically used to perform the vector calculation based on the first submatrix after the second data is rearranged.

4. The accelerator according to any one of claims 1 to 3, characterized in that, The matrix post-processing module is also used to send a synchronization signal to the vector calculation unit through a direct connection with the vector calculation unit.

5. A data processing method for a neural network accelerator, characterized in that, The accelerator includes a matrix computation unit and at least one vector computation unit. The matrix computation unit includes a matrix computation kernel and a matrix post-processing module. A direct path is established between the matrix post-processing module and the at least one vector computation unit. The method includes: The matrix calculation kernel performs matrix calculations in the neural network model, and the result of the matrix calculation is a first matrix. The first matrix is ​​divided according to the number of vector calculation units to obtain a first sub-matrix corresponding to each vector calculation unit. The matrix post-processing module sends the first sub-matrix corresponding to the vector calculation unit through a direct connection with the vector calculation unit; The vector calculation unit performs vector calculations based on the received first submatrix.

6. The method according to claim 5, characterized in that, The matrix calculation kernel includes a cache module, the first matrix is ​​stored in the cache module, and the first matrix is ​​divided equally along the target direction according to the number of vector calculation units to obtain the first sub-matrix corresponding to each vector calculation unit. The target direction includes the matrix row direction or the matrix column direction. The matrix post-processing module sends the first sub-matrix corresponding to the vector calculation unit through a direct connection with the vector calculation unit, including: The matrix post-processing module alternately reads the first sub-matrix corresponding to each vector calculation unit from the cache module.

7. The method according to claim 5 or 6, characterized in that, The method further includes: The matrix post-processing module performs a first data rearrangement on each of the first sub-matrices read; The vector calculation unit performs a second data rearrangement on the first sub-matrix after rearranging the received first data; The vector calculation unit performs vector calculations based on the received first sub-matrix, including: The vector calculation unit performs the vector calculation based on the first submatrix after rearranging the second data.

8. The method according to any one of claims 5 to 7, characterized in that, The method further includes: The matrix post-processing module sends a synchronization signal to the vector calculation unit through a direct connection with the vector calculation unit.

9. An electronic device, characterized in that, The device includes: One or more processors; Memory, used to store one or more computer programs or instructions; When the one or more computer programs or instructions are executed by the one or more processors, the one or more processors perform the method as described in any one of claims 5 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code that, when executed by a computer or processor, causes the computer or processor to perform the method as described in any one of claims 5 to 8.