Matrix memory reading method applied to sublimation NPU
By preprocessing the memory arrangement of the matrix, the problem that matrix multiplication operations in heterogeneous computing platforms cannot make full use of hardware bandwidth is solved, efficient matrix memory reading is achieved, and system computing performance is improved.
Patent Information
- Application Number
- CN202510391568.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
In heterogeneous computing platforms, especially in heterogeneous computing architectures composed of CPU and GPU, traditional matrix multiplication operations cannot fully utilize the hardware bandwidth due to the complex layout of input matrix data and the diverse data types, and cannot meet the high throughput and high efficiency computing needs.
By preprocessing the matrix with memory arrangement, the matrix is divided into regular and non-regular matrices, the matrix calculation unit and vector calculation unit are used to preprocess the regular matrix with memory data arrangement, and the non-regular matrix is judged based on the equivalent bandwidth quantization preprocessing benefits, and the matrix multiplication operation is performed using an efficient memory read interface.
The memory bandwidth utilization in memory access bottleneck scenarios has been optimized, the efficiency of matrix multiplication operations has been improved, and the overall computing performance of the system has been improved.
Smart Images

Figure CN120335994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heterogeneous parallel computing, and particularly to a matrix memory reading method applied to Ascend NPU. Background Art
[0002] With the progress of science and technology, the demand for computing resources in fields such as artificial intelligence, high-performance computing, and the Internet of Things is continuously increasing. Hardware resources provide the necessary computing power. The traditional single computing architecture based on the CPU as the main computing core has low efficiency in processing parallel computing tasks, such as the core computing operation of neural networks, matrix multiplication, and cannot meet the computing requirements of high throughput and high efficiency. At the same time, with the slowdown of Moore's Law, the improvement of the computing performance of a single CPU core has gradually slowed down.
[0003] Due to its high performance in processing parallel tasks, the GPU is gradually applied to the above high-computing-power demand scenarios. The CPU and the GPU form a typical heterogeneous computing platform, where the CPU is responsible for processing complex control logic, task scheduling, and serialization operations, and the GPU is responsible for executing large-scale parallel computing tasks. The heterogeneous computing architecture divides the computing tasks into multiple subtasks and assigns them to different processing units, providing better performance in specific tasks, being able to support both high-throughput computing and low-latency computing, and significantly reducing energy consumption by selecting the appropriate processor to execute specific tasks.
[0004] The design concept of Ascend NPU (Neural Processing Unit) is significantly different from that of traditional CPUs and GPUs. It is specially optimized for neural networks, adopts the self-developed Da Vinci architecture, uses AI Core as the computing core, and hides latency through strategies such as data reuse, multi-level caching, and pipeline parallelism. The new heterogeneous parallel computing system composed of the CPU and the NPU provides more energy-efficient support for the above scenarios. Building an Ascend operator library is the core task of improving the Ascend NPU ecosystem, and it is necessary to ensure the high efficiency of its operators, especially the core step of the operator library, matrix multiplication. However, the performance optimization of matrix multiplication faces problems such as complex input matrix data layout and diverse data types, and cannot utilize all the bandwidth of the hardware, so it cannot fully exert the computing power performance of the hardware. Therefore, an efficient matrix memory reading method is needed to improve the overall computing performance of the system. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a matrix memory reading method applied to Ascend NPU, which preprocesses the memory layout of the matrix, rearranges the memory of the input regular and irregular matrices, and can achieve high-bandwidth memory reading during matrix multiplication operations, thereby improving the overall computing efficiency of the system. The object of the present invention can be achieved by adopting the following technical solutions:
[0006] A matrix memory reading method applied to Ascend NPU, comprising the following steps:
[0007] S1. Divide the matrix into a regular matrix and an irregular matrix according to the characteristic information of the matrix. The regular matrix includes a regular left matrix and a regular right matrix;
[0008] S2. Use the matrix operation unit to perform preprocessing on the memory arrangement of the regular left matrix, and arrange the memory data of the regular left matrix into the format required for matrix multiplication in the L1 cache;
[0009] S3. Use the vector operation unit to perform preprocessing on the memory arrangement of the regular right matrix, and arrange its memory data into the format required for matrix multiplication in the L1 cache;
[0010] S4. Judge whether to use the vector operation unit to perform preprocessing on the memory data arrangement of the irregular matrix according to the equivalent bandwidth quantization preprocessing benefit; when the equivalent bandwidth after preprocessing is greater than that before preprocessing, perform preprocessing on the memory data arrangement of the irregular matrix, otherwise do not perform preprocessing on the memory data arrangement of the irregular matrix;
[0011] Performing preprocessing on the memory data arrangement of the irregular matrix includes: arranging the interval between the main dimensions into a multiple of the device memory aligned reading requirement, and at the same time making the memory inside the block continuously stored;
[0012] S5. Use the high-efficiency memory reading interface to read the preprocessed matrix into the cache for matrix multiplication operation.
[0013] Specifically, the characteristic information of the matrix includes the data type of the matrix, the data arrangement of the matrix, and the dimension information of the matrix. The data type of the matrix includes integer type and various precision floating-point types; the data arrangement of the matrix includes row-major and column-major; the dimension information of the matrix is the number of elements on two axes.
[0014] Specifically, using the matrix operation unit to perform preprocessing on the memory arrangement of the regular left matrix and arranging the memory data of the regular left matrix into the format required for matrix multiplication in the L1 cache specifically includes:
[0015] Allocate a space in the global memory with the same size as the original matrix to store the preprocessed result matrix;
[0016] Traverse each block for matrix operation in turn, and use the cache-coherent transfer instruction from global memory to L1 cache to arrange the current block from the arrangement format on the global memory into the format required for matrix multiplication in the L1 cache: row-major inside the fractal matrix, and small-Z large-N format with row-major between fractal matrices;
[0017] Copy the fractal matrix to the corresponding position in the global memory result matrix space using a continuous memory copy instruction, and then perform the memory layout preprocessing for the next block until the memory layout preprocessing of all blocks is completed.
[0018] Specifically, the preprocessing of the memory layout of the regular right matrix using the vector operation unit to arrange its memory data into the format required for matrix multiplication in the L1 cache includes:
[0019] Allocate a space in the global memory with the same size as the original matrix to store the preprocessed result matrix;
[0020] Traverse each block for matrix operation in turn, obtain the starting address of the block in the source matrix and the starting address in the destination matrix according to the block number, use the interval to copy the block of the copy instruction to the UB cache for continuous row-first storage, and then store it continuously in column-first order after transposing the block;
[0021] Copy a column of data to the global memory at intervals of 32 bytes, so that the memory data layout in each block is arranged into a small N large Z format with column-first inside the fractal matrix and row-first between the fractal matrices, and then perform the memory layout preprocessing for the next block until the memory layout preprocessing of all blocks for matrix operation is completed.
[0022] Specifically, step S2 and step S3 are performed simultaneously, and the matrix operation unit and the vector operation unit are used to perform the preprocessing of the memory layout of the regular left matrix and the regular right matrix respectively.
[0023] Specifically, arranging the interval between the main dimensions into a multiple of the device memory alignment read requirement and enabling continuous memory storage inside the block includes:
[0024] Recalculate the size of the global memory space for storing the result matrix, arrange the interval between the main dimensions into the memory space size required by the matrix after being a multiple of the device memory alignment read requirement, and allocate a corresponding size of space in the global memory space to store the preprocessed result matrix;
[0025] Traverse each block for matrix operation in turn, use the interval memory copy instruction to copy the current block to the UB cache for continuous storage, copy the regular block to the global memory for continuous storage, copy the irregular tail block to the global memory for interval storage, set the row interval to the size of the main dimension of the block, and use the continuous memory copy instruction to copy the fractal to the corresponding position in the global memory result matrix space; then perform the memory layout preprocessing for the next block until the memory layout preprocessing of all blocks for matrix operation is completed.
[0026] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0027] The present invention provides a matrix memory reading method applied to Ascend NPU. According to the characteristic information of the matrix, the matrix is divided into a regular matrix and an irregular matrix. The memory data arrangements of the regular left matrix and the regular right matrix are sorted out into the format required for matrix multiplication in the L1 cache. According to the equivalent bandwidth quantization preprocessing benefit, it is judged whether to use the vector operation unit to preprocess the memory data arrangement of the irregular matrix. When the equivalent bandwidth after preprocessing is greater than the original bandwidth before preprocessing, the memory data arrangement of the irregular matrix is preprocessed, and the preprocessed matrix is read into the cache using an efficient memory reading interface for matrix multiplication operation. By rearranging the memory of the input regular and irregular matrices, the present invention optimizes the problems that the memory bandwidth cannot be efficiently utilized in the memory access bottleneck scenario and the occupancy ratios of the MTE2 pipeline and the MTE1 pipeline are too high. By using multi-block parallel processing and the parallel processing of the matrix operation unit and the vector operation unit, the time required for preprocessing the input matrix is shortened. At the same time, the benefits and costs brought by the preprocessing are quantitatively weighed to achieve efficient reading of the input matrix memory, thereby improving the overall computing performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0029] Figure 1 It is a flowchart of a matrix memory reading method applied to Ascend NPU in Embodiment 1;
[0030] Figure 2 It is a schematic diagram of the memory arrangement requirements of the matrix for matrix multiplication operation in Embodiment 1;
[0031] Figure 3 It is a schematic diagram of preprocessing the memory arrangement of the regular left matrix in Embodiment 1;
[0032] Figure 4 It is a schematic diagram of preprocessing the memory arrangement of the regular right matrix in Embodiment 1;
[0033] Figure 5 It is a schematic diagram of preprocessing the memory arrangement of the irregular matrix in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, the technical solution of the present invention will be further described in detail in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. The implementation manners of the present invention are not limited thereto. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0035] Embodiment 1:
[0036] As Figure 1 shown, a flowchart of a matrix memory reading method applied to Ascend NPU. The matrix memory reading method applied to Ascend NPU described in the present invention specifically includes:
[0037] S1. Divide the matrix into a regular matrix and an irregular matrix according to the characteristic information of the matrix.
[0038] When Ascend NPU processes large-scale data, it requires an efficient memory access mechanism. Memory-aligned access can make full use of the characteristics of the hardware, reduce the overhead of memory access, and improve the data reading and storage speed. Ascend NPU has alignment requirements for memory access, in units of bytes.
[0039] The matrix characteristic information mainly includes the data type of the matrix, the data arrangement of the matrix, and the dimension information of the matrix. The data type of the matrix includes integer type and various precision floating-point types. Elements of different data types occupy different spaces in memory, which is the basis for calculating the total number of bytes occupied by matrix elements. For example, an 8-bit integer type element occupies 1 byte, while a 32-bit single-precision floating-point type element occupies 4 bytes. The data arrangement of the matrix includes row-major and column-major. The data arrangement of the matrix describes the storage order of matrix elements in memory. In row-major arrangement, the elements of the matrix are stored in memory in the order of rows, that is, all elements of the first row are stored first, then all elements of the second row are stored, and so on; while in column-major arrangement, the elements are stored in the order of columns. The data arrangement method affects the layout of the matrix in memory; the dimension information of the matrix is the number of elements on two axes. In a matrix structure, there are usually two dimensions, rows and columns. The number of elements on the row axis describes the vertical scale of the matrix, while the number of elements on the column axis defines the horizontal scale range of the matrix. The number of elements on these two axes together determines the specific shape and size characteristics of the matrix.
[0040] First, determine the main dimension of the matrix based on the eigen-information of the root matrix. The main dimension of the matrix can be understood as the dimension with the fastest change in the matrix. For example, for a two-dimensional matrix, if it is arranged in row-major order, then the row dimension is the main dimension because when traversing the matrix elements, the change in rows is continuous; if it is arranged in column-major order, the column dimension is the main dimension. In practical applications, determine the main dimension according to the specific data arrangement and calculation requirements to better analyze and process the memory access situation of the matrix.
[0041] Calculate the total number of bytes occupied by the matrix elements. The total number of bytes occupied by the matrix elements can be calculated by multiplying the number of bytes of each matrix element by the number of elements. The number of bytes of each matrix element is determined by the data type of the matrix, and the number of elements is determined by the dimension information of the matrix. For example, for a two-dimensional matrix of size m×n, if it is arranged in row-major order and the element data type is 32-bit single-precision floating-point number (each element occupies 4 bytes), then the total number of bytes occupied by the matrix elements is m×n×4 bytes.
[0042] Then, divide the matrix into a regular matrix and an irregular matrix based on whether the product of the number of bytes of each matrix element and the number of elements is a multiple of the memory alignment access requirement. If the total number of bytes occupied by the matrix elements is a multiple of the number of bytes required for memory alignment access, then the matrix is called a regular matrix; otherwise, it is called an irregular matrix. For example, assume that the memory alignment access requirement of the Ascend NPU is 16 bytes. For the above two-dimensional matrix of size m×n, if m×n×4 is a multiple of 16, then the matrix is a regular matrix; otherwise, it is an irregular matrix. Since the total number of bytes occupied by the elements of an irregular matrix does not meet the memory alignment access requirement, unaligned situations may occur during memory access, resulting in additional overhead. For an irregular matrix, some preprocessing is usually required to make it meet the memory alignment requirement.
[0043] In matrix operations, the source matrix and the destination matrix are usually very large. Processing the entire matrix at once will not only consume a large amount of memory bandwidth but also may lead to a decrease in cache hit rate. Therefore, divide the matrix into multiple small blocks and process each small block separately. Considering the number of computing cores of the hardware and the sizes of each level of cache, partition the matrices participating in matrix multiplication operations so that when performing memory arrangement preprocessing on the matrices, all computing cores can be utilized as much as possible and all cache spaces can be utilized as much as possible, and formulate a suitable computing task allocation strategy so that each computing core is assigned an approximate number of tasks to achieve the best system parallelism.
[0044] S2. Use the matrix operation unit to perform preprocessing on the memory arrangement of the regular left matrix, and organize the memory data arrangement of the regular left matrix into the format required for matrix multiplication in the L1 cache.
[0045] Such as Figure 2As shown in the figure, it is a schematic diagram of the memory layout requirements for the matrix in the matrix multiplication operation in Embodiment 1. The Ascend NPU has fixed requirements for the memory layout of the matrix in the matrix multiplication operation. The data layout requirement of the left matrix on the L0A cache is row-major inside the fractal matrix and small-Z large-Z format with row-major between fractal matrices. However, the Ascend NPU does not directly provide a follow-on conversion instruction from row-major to small-Z large-Z format. Therefore, it is necessary to first organize the data layout of the left matrix on the L1 cache into a small-Z large-N format with row-major inside the fractal matrix and column-major between fractal matrices as an intermediate format to meet the matrix memory layout requirements for the matrix multiplication operation. However, due to the limitations of bandwidth and the efficiency of follow-on conversion instructions, the proportion of the MTE2 pipeline from global memory to L1 cache is too high, showing an obvious memory access bottleneck. At the same time, since there is only a memory transfer path from L1 cache to global memory and no memory transfer path from L0 cache to global memory, the memory layout of the left matrix is pre-organized into the small-Z large-N format required by the L1 cache in advance with the help of the follow-on conversion instruction of the matrix operation unit to improve the efficiency of memory reading and thus enhance the overall computing efficiency.
[0046] As Figure 3 shown, it is a schematic diagram of preprocessing the memory layout for the regular left matrix in Embodiment 1. The matrix operation unit is used to preprocess the memory layout of the regular left matrix, which specifically includes:
[0047] First, a space of the same size as the original matrix is allocated in the global memory to store the result matrix after preprocessing. Then, each block for matrix operation is traversed in turn, and the follow-on conversion instruction from global memory to L1 cache is used to organize the layout format of the current block from the global memory into the format required for matrix multiplication in the L1 cache: row-major inside the fractal matrix and small-Z large-N format with row-major between fractal matrices.
[0048] Then, the fractal matrix is copied to the corresponding position in the global memory result matrix space using the continuous memory copy instruction, and then the memory layout preprocessing of the next block is performed until the memory layout preprocessing of all blocks is completed. In this process, the double buffering of the L1 cache can be used to parallelize the copying out of the previous block and the copying in of the next block to improve the overall processing efficiency.
[0049] S3. Use the vector operation unit to preprocess the memory layout of the regular right matrix and organize its memory data layout into the format required for matrix multiplication in the L1 cache.
[0050] Ascend NPU also has fixed requirements for the memory layout of the right matrix in matrix multiplication operations. The data layout requirement of the right matrix on the L0B cache is the small-N large-Z format with column-major inside the fractal matrix and row-major between fractal matrices. However, Ascend NPU does not directly provide the in-path conversion instruction from row-major to the small-N large-Z format. Therefore, it is necessary to first arrange the data layout of the right matrix on the L1 cache into the small-Z large-Z format with row-major inside the fractal matrix and row-major between fractal matrices, and then use the transpose instruction to arrange it into the target format on the L0 cache. However, both the MTE2 pipeline from global memory to L1 cache and the MTE1 pipeline from L1 cache to L0 cache have too high a proportion, and there is also an obvious memory access bottleneck. Therefore, the vector operation unit is used to arrange the memory layout of the right matrix into the small-N large-Z format required by the L0 cache, improve the memory reading efficiency, and thus improve the overall computing efficiency.
[0051] As Figure 4 shown, it is a schematic diagram of preprocessing the memory layout of the regular right matrix in Embodiment 1. The vector operation unit is used to preprocess the memory layout of the regular right matrix and arrange its memory data layout into the format required for matrix multiplication in the L1 cache, specifically including:
[0052] First, a space with the same size as the original matrix is opened in the global memory to store the result matrix after preprocessing;
[0053] Then, each block for matrix operation is traversed in turn. According to the block number, the starting address of the block in the source matrix and the starting address in the destination matrix are obtained. The copy instruction is used to copy the block into the UB (Unified Buffer) cache in a continuous row-major storage, and then the block is transposed and stored in a continuous column-major manner;
[0054] Finally, a column of data is copied to the global memory at intervals of 32 bytes (i.e., the size of a cache line), so that the memory data layout in each block is arranged as follows: column-major within the fractal matrix and row-major between the fractal matrices in the small N large Z format. Then, the memory layout preprocessing of the next block is performed until the memory layout preprocessing of all blocks for matrix operations is completed. During this process, the double buffering of the UB cache can parallelize the transpose and copy-out of the previous block and the copy-in of the next block to improve the overall processing efficiency. This strategy can efficiently complete the processing and transmission of matrix data, improve data locality and cache hit rate, and make full use of the hardware data transmission bandwidth, thereby improving the performance of the entire matrix operation. Since the L1 cache, L0 cache, and UB cache all have capacity limitations, it is impossible to put the entire matrix into the matrix operation unit and vector operation unit for processing at one time. The matrix needs to be divided into blocks according to the number of computing cores of the Ascend NPU, the sizes of each level of cache, and the block shape to obtain multiple sub-blocks that can make the best use of the cache space. Each computing core of the Ascend NPU reads and processes one block each time, so that the number of blocks to be processed by each computing core is as consistent as possible, making full use of the computing resources of the system and thus improving the overall computing efficiency.
[0055] Specifically, step S2 and step S3 are carried out simultaneously, and the matrix operation unit and vector operation unit are used to perform the preprocessing of the memory layout of the regular left matrix and regular right matrix respectively. Since the adjustment of the matrix memory data layout before matrix multiplication needs to operate on the entire matrix, and matrix multiplication cannot start until all matrix processing is completed, the time-consuming of these preprocessing steps cannot be masked by the pipeline parallel method of the main matrix multiplication operation, and its time-consuming significantly affects the overall computing performance. Therefore, the matrix operation unit and vector operation unit are used to process the left matrix and right matrix simultaneously, and the benefit of computing parallelism can mask the impact brought by the shared bandwidth, thereby improving the overall computing efficiency.
[0056] S4. Judge whether to use the vector operation unit to perform the preprocessing of the memory data layout of the irregular matrix according to the quantization preprocessing benefit of the equivalent bandwidth; when the equivalent bandwidth after preprocessing is greater than the original bandwidth, perform the preprocessing of the memory data layout of the irregular matrix, otherwise do not perform the preprocessing of the memory data layout of the irregular matrix.
[0057] Specifically, the preprocessing of the memory data layout of the irregular matrix includes: arranging the interval between the main dimensions into a multiple of the device memory aligned reading requirement, and at the same time making the memory inside the block continuously stored.
[0058] All of the preprocessing operations for the memory data layout of an irregular matrix need to be performed using a vector arithmetic unit. Compared with the preprocessing of a regular left matrix and a regular right matrix, using a matrix arithmetic unit and a vector arithmetic unit in parallel reduces the waiting time for the main matrix multiplication operation. The preprocessing of an irregular matrix has a greater impact on the overall calculation time. Therefore, it is necessary to judge whether to use a vector arithmetic unit to preprocess the memory data layout of an irregular matrix according to the equivalent bandwidth quantization of the preprocessing benefit. By comparing the bandwidth improvement brought by the preprocessing and the equivalent bandwidth reduction caused by the preprocessing time, when there is a theoretical benefit, that is, when the equivalent bandwidth after preprocessing is greater than the original bandwidth, preprocessing can be carried out to achieve better results.
[0059] Calculate the equivalent bandwidth and the original bandwidth. When the equivalent bandwidth is greater than the original bandwidth, it proves that the preprocessing has a positive benefit. When the equivalent bandwidth after preprocessing is greater than the original bandwidth, preprocess the memory data layout of the irregular matrix. The calculation methods of the equivalent bandwidth and the original bandwidth are as follows:
[0060] Original bandwidth = data access volume / data access time;
[0061] Equivalent bandwidth = data access volume / (data access time + preprocessing time);
[0062] That is, when the equivalent bandwidth is greater than the original bandwidth, it proves that the preprocessing has a positive benefit, and preprocessing is adopted in the corresponding scenario.
[0063] For an irregular matrix, unaligned memory access will significantly reduce the bandwidth, and its impact on the overall computing efficiency is more obvious in the scenario of a memory access bottleneck. Therefore, the inside of the block of the irregular matrix is organized into continuous memory storage to avoid the problem of non-aligned intervals between multiple rows and columns of the original matrix, and continuous memory access instructions are used to read the matrix with high bandwidth. However, different from the preprocessing of a regular left matrix and a regular right matrix, the temporarily allocated global memory space for storing the result matrix needs to be larger than the space for storing the original matrix, which is equivalent to organizing the interval between the main dimensions into multiples of the device memory alignment reading requirements to store the extra space required for the tail block.
[0064] To accurately evaluate the benefits brought by this preprocessing, it is also necessary to quantify according to the equivalent bandwidth. The equivalent bandwidth refers to the actual data transmission rate during the data transmission process, taking into account factors such as data read and write time and transmission delay. By measuring and comparing the equivalent bandwidth before and after preprocessing, it is possible to clearly understand the degree of improvement in system performance brought by the preprocessing work. If the equivalent bandwidth after preprocessing is significantly greater than the original bandwidth before preprocessing, it means that this preprocessing method is effective and can optimize the system performance; on the contrary, if the improvement in the equivalent bandwidth is not obvious or even decreases.
[0065] As Figure 5 shown, it is a schematic diagram of preprocessing the memory layout for an irregular matrix in Embodiment 1. For an irregular matrix, a vector operation unit is used to preprocess the memory layout:
[0066] First, recalculate the global memory space size of the stored result matrix, which is the memory space size required for the matrix after arranging the intervals between the main dimensions into multiples of the device memory alignment reading requirements, and allocate a corresponding space in the global memory space to store the preprocessed result matrix;
[0067] Then, traverse each block for matrix operations in sequence. Use the interval memory copy instruction to copy the current block continuously to the UB cache, copy the regular block continuously to the global memory, copy the irregular tail block to the global memory at intervals, set the row interval to the size of the main dimension of the block, and then use the continuous memory copy instruction to copy this fractal to the corresponding position in the global memory result matrix space. Then, perform the memory layout preprocessing for the next block until the memory layout preprocessing of all blocks for matrix operations is completed. In this process, double buffering can also be used to improve efficiency.
[0068] Step Five: Use an efficient memory reading interface to read the preprocessed matrix into the cache and perform matrix multiplication operations, which can improve the overall computing efficiency.
[0069] Specifically, the preprocessed matrix includes the preprocessed regular matrix and the preprocessed irregular matrix. For the preprocessed regular left matrix, the continuous memory copy instruction is used to replace the on-path conversion instruction for the copy from the global memory to the L1 cache; for the preprocessed regular right matrix, the continuous memory copy instruction is used to replace the on-path conversion instruction and the transpose copy instruction for the copy from the global memory to the L1 cache and from the L1 cache to the L0 cache; for the preprocessed irregular matrix, the continuous memory copy instruction is used to replace the unaligned copy instruction for the copy from the global memory to the L1 cache, and then matrix operations are performed. The above process is looped to calculate the value of each result block until the computing task is completed.
[0070] Compared with the preprocessed regular left matrix, if the unpreprocessed regular left matrix needs to use the on-the-fly conversion instruction to copy data from the global memory to the L1 cache, while the preprocessed regular left matrix uses a continuous memory copy instruction with higher execution efficiency; compared with the preprocessed regular right matrix, if the unpreprocessed regular right matrix needs to use the on-the-fly conversion instruction to copy data from the global memory to the L1 cache and then use the transpose copy instruction to copy data from the L1 cache to the L0 cache, while the preprocessed regular right matrix uses a continuous memory copy instruction with higher execution efficiency; compared with the preprocessed irregular matrix, if the unpreprocessed irregular matrix needs to use the unaligned copy instruction to copy data from the global memory to the L1 cache, while the preprocessed irregular matrix uses an aligned memory copy instruction with higher execution efficiency. More efficient memory copying can improve the overall computing efficiency.
[0071] This embodiment provides a matrix memory reading method applied to Ascend NPU. The matrix is divided into a regular matrix and an irregular matrix according to the characteristic information of the matrix; for the regular left matrix, the matrix operation unit is used to arrange its memory data into the format required for matrix multiplication in the L1 cache, and for the regular right matrix, the vector operation unit is used to arrange its memory data into the format required for matrix multiplication in the L2 cache; for the irregular matrix, the vector operation unit is used to arrange the interval between the main dimensions into a multiple of the device memory aligned reading requirement, and at the same time, the memory inside the block is continuously stored; the equivalent bandwidth is used as the basis for whether to preprocess the irregular matrix; the processed matrix is read into the cache through an efficient memory reading interface for computing tasks. It can perform memory rearrangement on the input regular and irregular matrices, optimize the problem that the memory bandwidth cannot be efficiently utilized in the memory access bottleneck scenario and the over-high occupancy ratios of the MTE2 pipeline and the MTE1 pipeline, and use multi-block parallel processing and the parallel processing of the matrix operation unit and the vector operation unit to further shorten the time required for preprocessing the input matrix. At the same time, it quantitatively weighs the benefits and costs brought by the preprocessing, so as to achieve efficient reading of the input matrix memory, and then improve the overall computing performance of the system. The present invention can be applied to a heterogeneous parallel computing system composed of a CPU and an NPU, and can improve the overall computing efficiency of the system.
[0072] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A matrix memory reading method applied to Ascend NPU, characterized in that It includes the following steps: S1. Divide the matrix into a regular matrix and an irregular matrix according to the characteristic information of the matrix. The regular matrix includes a regular left matrix and a regular right matrix; S2. Use the matrix operation unit to perform preprocessing on the memory layout of the regular left matrix, and arrange the memory data layout of the regular left matrix into the format required for matrix multiplication in the L1 cache; S3. Use the vector operation unit to perform preprocessing on the memory layout of the regular right matrix, and arrange its memory data layout into the format required for matrix multiplication in the L1 cache; S4. Judge whether to use the vector operation unit to perform preprocessing on the memory data layout of the irregular matrix according to the equivalent bandwidth quantization preprocessing benefit; when the equivalent bandwidth after preprocessing is greater than the original bandwidth, perform preprocessing on the memory data layout of the irregular matrix, otherwise do not perform preprocessing on the memory data layout of the irregular matrix; Performing preprocessing on the memory data layout of the irregular matrix includes: arranging the interval between the main dimensions into a multiple of the device memory aligned read requirement, and at the same time making the memory inside the block continuously stored; S5. Use the high-efficiency memory read interface to read the preprocessed matrix into the cache and perform matrix multiplication operations.
2. The matrix memory reading method applied to Ascend NPU according to claim 1, wherein: The characteristic information of the matrix includes the data type of the matrix, the data layout of the matrix, and the dimension information of the matrix. The data type of the matrix includes integer type and various precision floating-point types; the data layout of the matrix includes row-major and column-major; the dimension information of the matrix is the number of elements on two axes.
3. The matrix memory reading method applied to Ascend NPU according to claim 1, wherein: The use of the matrix operation unit to perform preprocessing on the memory layout of the regular left matrix and arrange the memory data layout of the regular left matrix into the format required for matrix multiplication in the L1 cache specifically includes: Allocate a space in the global memory with the same size as the original matrix to store the preprocessed result matrix; Traverse each block for matrix operations in turn, and use the cache-coherent transfer instruction from the global memory to the L1 cache to arrange the current block from the layout format on the global memory into the format required for matrix multiplication in the L1 cache: row-major inside the fractal matrix, and small-Z large-N format with row-major between fractal matrices; Use the continuous memory copy instruction to copy the fractal matrix to the corresponding position in the global memory result matrix space, and then perform the memory layout preprocessing of the next block until the memory layout preprocessing of all blocks is completed.
4. The matrix memory reading method applied to Ascend NPU according to claim 3, wherein: The use of the vector operation unit to perform preprocessing on the memory layout of the regular right matrix and arrange its memory data layout into the format required for matrix multiplication in the L1 cache includes: Allocate a space in the global memory with the same size as the original matrix to store the preprocessed result matrix; Traverse each block for matrix operations in turn, obtain the starting address of the block in the source matrix and the starting address in the destination matrix according to the block number, use the interval to copy the instruction in blocks and store them continuously in row-major order in the UB cache, and then store them continuously in column-major order after transposing the block; Copy a column of data to the global memory at intervals of 32 bytes, so that the memory data arrangement in each block is sorted into a small N large Z format with column-major inside the fractal matrix and row-major between the fractal matrices, and then perform the preprocessing of the memory arrangement for the next block until the memory arrangement preprocessing of all blocks for matrix operations is completed.
5. The matrix memory reading method applied to Ascend NPU according to claim 4, wherein: Steps S2 and S3 are performed simultaneously, and the matrix operation unit and the vector operation unit are used to perform the preprocessing of the memory arrangement for the regular left matrix and the regular right matrix respectively.
6. The matrix memory reading method applied to Ascend NPU according to claim 1, wherein: The interval between the main dimensions is sorted into a multiple of the device memory alignment read requirement, and at the same time, the memory inside the block is continuously stored, including: Recalculate the global memory space size of the stored result matrix, sort the interval between the main dimensions into the memory space size required by the matrix after being a multiple of the device memory alignment read requirement, and allocate a corresponding size of space in the global memory space to store the preprocessed result matrix; Traverse each block for matrix operations in turn, use the interval memory copy instruction to copy the current block to the UB cache for continuous storage, copy the regular block to the global memory for continuous storage, copy the irregular tail block to the global memory at intervals, set the row interval to the size of the main dimension of the block, and use the continuous memory copy instruction to copy the fractal to the corresponding position in the global memory result matrix space; then perform the preprocessing of the memory arrangement for the next block until the memory arrangement preprocessing of all blocks for matrix operations is completed.
7. A matrix memory reading method applied to Ascend NPU according to claim 1, characterized in that: Use the high-efficiency memory read interface to read the preprocessed matrix into the cache and perform matrix multiplication operations, including: For the preprocessed regular left matrix, use the continuous memory copy instruction instead of the on-the-fly conversion instruction for the copy from the global memory to the L1 cache; for the preprocessed regular right matrix, use the continuous memory copy instruction instead of the on-the-fly conversion instruction and the transpose copy instruction for the copy from the global memory to the L1 cache and from the L1 cache to the L0 cache; for the preprocessed irregular matrix, use the continuous memory copy instruction instead of the unaligned copy instruction for the copy from the global memory to the L1 cache, and then perform matrix operations. Loop the above process to calculate the value of each result block until the calculation task is completed.
Citation Information
Cited By
Data processing method and device, equipment and storage medium
CN121350398A