A data processing method, device and electronic device applied to a graphics processor

By performing two preset operations on the to-process matrix in the GPU, the problem of low applicability of the GPU in matrix multiplication operations is solved, and more matrix arrangement methods are supported, which improves the flexibility and efficiency of calculations.

CN114626969BActive Publication Date: 2025-05-13ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011465448.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-14
Publication Date
2025-05-13
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

In the prior art, when performing matrix multiplication operations, GPUs can only perform transformation processing for specific integer types of matrices, resulting in low applicability.

Method used

By reading the to-processed matrix from the memory of the GPU into the registers of the streaming multiprocessor, and performing two preset operations in the shared memory, the first preset operation and the second preset operation, the matrix is ​​converted to meet the requirements of the matrix multiplication dedicated computing unit.

Benefits of technology

It improves the applicability of GPU in matrix multiplication operations, allowing it to handle more types of matrix arrangement methods, thereby improving the flexibility and efficiency of calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626969B_ABST
    Figure CN114626969B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method applied to a GPU, comprising: reading a matrix to be processed from a memory corresponding to a target GPU into a register in a target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor for performing matrix calculations on the matrix to be processed; in the process of reading the matrix to be processed from the register into a shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed; in the process of reading the initial matrix from the shared memory into the register, performing a second preset operation on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of a matrix-specific computing unit. The data processing method applied to a graphics processing unit (GPU) can transform the matrix to be processed into a target matrix that meets the matrix multiplication operation requirements of the GPU, thereby improving the applicability of the GPU for matrix multiplication operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly to a data processing method for a graphics processor. The present application also relates to a graphics processor device, an electronic device, and a storage medium. The present application also relates to another data processing method for a GPU. Background Art

[0002] GPU (Graphics Processing Unit), also known as display core, visual processor, display chip, is a microprocessor used for image processing and image-related computing work on personal computers, tablet computers, smart phones and other computing devices, such as: performing general-purpose calculations such as matrix multiplication, etc. When using GPU for matrix multiplication calculations, the matrix multiplication dedicated computing unit used for matrix operations in the GPU has certain restrictions on the arrangement of the input matrix. The matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the streaming multiprocessor. For example, when the integer type of the element is INT8 (integer) of the input matrix, it is required that the arrangement form of the first matrix in the two multiplied matrices is a row-major matrix, and the arrangement form of the second matrix is ​​a column-major matrix.

[0003] In practical applications, the arrangement of input matrices is varied. At this time, only by transforming the input matrix into a pre-set arrangement can the GPU perform further matrix multiplication operations on the input matrix. In the prior art, when using the GPU for matrix multiplication operations, matrix transformation processing can only be performed on matrices whose element integer type is a specific bit size, resulting in low applicability of the GPU in the prior art for matrix multiplication operations. Summary of the invention

[0004] The present application provides a data processing method, device, electronic device and storage medium applied to a graphics processor to improve the applicability of a GPU for performing matrix multiplication operations.

[0005] The present application provides a data processing method applied to a graphics processing unit (GPU), comprising:

[0006] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0007] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0008] In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0009] Optionally, reading the matrix to be processed from the memory corresponding to the target GPU into a register in the target streaming multiprocessor includes:

[0010] Obtaining a first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed;

[0011] According to the first multiple, starting from the first column, the elements of the same row in the matrix to be processed, which are each a multiple of the first multiple, are read into the register in sequence.

[0012] Optionally, obtaining the first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed includes:

[0013] Obtain the number of bits of the register;

[0014] Obtain the number of digits of the integer type corresponding to the elements in the matrix to be processed;

[0015] The first multiple is determined according to the number of bits of the register and the number of bits of the integer type corresponding to the elements in the matrix to be processed.

[0016] Optionally, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed includes:

[0017] Obtaining a first read instruction for the matrix to be processed, wherein the first read instruction is a read instruction for a first integer type;

[0018] Obtain a second multiple of the number of bits corresponding to the first integer type relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed;

[0019] Determine a first preset operation for the matrix to be processed according to the second multiple, wherein the first preset operation is to sequentially distribute and merge elements of each row of the matrix that is a multiple of the second multiple into one row starting from the first row;

[0020] Starting from the first row, the elements of each second multiple row in the matrix to be processed are sequentially distributed and merged into one row, and read into the shared memory to obtain the initial matrix.

[0021] Optionally, the second multiple is twice;

[0022] The step of merging the elements of every second multiple row in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix comprises: merging the elements of every two rows in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix.

[0023] Optionally, the second multiple is two, including: the number of bits corresponding to the first integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits;

[0024] Alternatively, the number of bits corresponding to the first integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 16 bits.

[0025] Optionally, the second multiple is one;

[0026] The step of distributing the elements of each second multiple row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix includes: distributing the elements of each row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix.

[0027] Optionally, performing a second preset operation on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the target GPU includes:

[0028] Obtaining a second read-in instruction for the initial matrix, wherein the second read-in instruction carries a second integer type for which the second read-in instruction is intended;

[0029] Obtain a third multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix;

[0030] Determine a second preset operation for the initial matrix according to the third multiple, wherein the second preset operation is to perform matrix transposition on each third multiple column in the matrix as a column starting from the first column, obtain a target matrix, and read the target matrix into a register;

[0031] Starting from the first column, each column of the third multiple in the initial matrix is ​​taken as a column to perform matrix transposition to obtain the target matrix, and the target matrix is ​​read into the register.

[0032] Optionally, the third multiple is two;

[0033] The method of starting from the first column, taking each third multiple column in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register includes: starting from the first column, taking every two columns in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register.

[0034] Optionally, the second multiple is two, including: the third multiple is two, including: the number of bits corresponding to the third integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​8 bits;

[0035] Alternatively, the number of bits corresponding to the third integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​16 bits.

[0036] Optionally, the number of rows and the number of columns of the matrix to be processed are multiples of a natural number power of two.

[0037] Optionally, the matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order, and the first matrix and the second matrix are the first matrix and the second matrix of two multiplication matrices to be subjected to matrix multiplication operation.

[0038] In another aspect, the present application further provides a data processing device, applied to a GPU, the device comprising:

[0039] A matrix reading unit to be processed is used to read the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, and the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0040] A first preset operation execution unit is used to execute a first preset operation on the matrix to be processed in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, so as to obtain an initial matrix corresponding to the matrix to be processed;

[0041] A second preset operation execution unit is used to perform a second preset operation on the initial matrix during the process of reading the initial matrix from the shared memory into the register, so as to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit, and the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operations in the target streaming multiprocessor.

[0042] In another aspect, the present application further provides an electronic device, applied to a GPU, comprising:

[0043] Processor; and

[0044] The memory is used to store a program of the text recognition method. After the device is powered on and the program of the data processing method applied to the GPU is passed, the following steps are performed:

[0045] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0046] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0047] In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0048] On the other hand, the present application further provides a storage medium applied to a GPU, wherein the storage medium applied to the GPU stores a program of a data processing method applied to the GPU, and the program is executed by a processor to perform the following steps:

[0049] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0050] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0051] In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0052] On the other hand, the present application also provides a data processing method applied to a GPU, comprising:

[0053] Obtaining a matrix to be processed that needs to be subjected to a matrix multiplication operation, wherein the matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order, and the first matrix and the second matrix are a first matrix and a second matrix of two multiplication matrices to be subjected to a matrix multiplication operation;

[0054] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0055] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0056] In the process of reading the initial matrix from the shared memory into the register, performing a second preset operation on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of a matrix multiplication dedicated computing unit, wherein the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor;

[0057] The matrix multiplication dedicated computing unit performs a matrix multiplication operation on the target matrix.

[0058] Compared with the prior art, this application has the following advantages:

[0059] The present application provides a data processing method applied to a graphics processing unit (GPU), first, a matrix to be processed is read from a memory corresponding to a target GPU into a register in a target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor for performing matrix calculations on the matrix to be processed; then, in the process of reading the matrix to be processed from the register into a shared memory corresponding to the register, a first preset operation is performed on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed; finally, in the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of a matrix multiplication dedicated calculation unit, wherein the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operations in the target streaming multiprocessor. The data processing method applied to a graphics processing unit (GPU) provided by the present application can transform the matrix to be processed into a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit by sequentially performing two different preset operations on the matrix to be processed that does not meet the GPU matrix multiplication operation requirements, thereby improving the applicability of the GPU for matrix multiplication operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 A schematic diagram of a scenario of a data processing method applied to a GPU provided in the first embodiment of the present application.

[0061] Figure 2A flowchart of a data processing method applied to a GPU provided in the first embodiment of the present application.

[0062] Figure 3 A schematic diagram of a matrix processing process to be processed provided in the first embodiment of the present application.

[0063] Figure 4 A schematic diagram of an initial matrix processing process provided in the first embodiment of the present application.

[0064] Figure 5 A schematic diagram of a target matrix processing process provided in the first embodiment of the present application.

[0065] Figure 6 This is a schematic diagram of a data processing device provided in the second embodiment of the present application.

[0066] Figure 7 This is a schematic diagram of an electronic device provided in the third embodiment of the present application.

[0067] Figure 8 A flowchart of a data processing method applied to a GPU provided in the fifth embodiment of the present application. DETAILED DESCRIPTION

[0068] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0069] First embodiment

[0070] In order to more clearly demonstrate the data processing method applied to GPU provided by the embodiment of the present application, the application scenario of the data processing method applied to GPU provided by the first embodiment of the present application is first introduced. Figure 1 , which is a scenario diagram of a data processing method applied to a GPU provided in the first embodiment of the present application.

[0071] The data processing method applied to GPU provided in the first embodiment of the present application, in practical application, when performing matrix multiplication operation on a matrix, first, execute step S101: input the matrix to be processed into the memory of the target GPU; second, execute step S102: read the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor; third, execute step S103: in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, execute the first preset operation on the matrix to be processed, and obtain the initial matrix corresponding to the matrix to be processed; fourth, execute step S104: in the process of reading the initial matrix from the shared memory into the register, execute the second preset operation on the initial matrix, and obtain the target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit; fifth, execute step S105: perform matrix multiplication operation on the target matrix. Among them, the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operation in the target streaming multiprocessor.

[0072] The so-called matrix to be processed is a first matrix arranged in column-major order, a second matrix arranged in row-major order, and a first matrix arranged in column-major order and a second matrix arranged in row-major order. The so-called first matrix and second matrix are the first matrix and the second matrix of the two multiplication matrices to be multiplied by the matrix multiplication operation. That is to say, the matrix to be processed is the first matrix arranged in column-major order and / or the second matrix arranged in row-major order.

[0073] The so-called target matrix is ​​a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the so-called matrix multiplication operation requirements of the matrix multiplication dedicated computing unit are matrix multiplication operation requirements predetermined according to the hardware conditions of the target GPU. For example, when the integer type of the elements is an input matrix of INT8, it is required that the arrangement form of the first matrix in the two multiplied matrices is a row-major matrix, and the arrangement form of the second matrix is ​​a column-major matrix. At this time, the arrangement form of the first matrix in the two multiplied matrices is a row-major matrix, and the arrangement form of the second matrix is ​​a column-major matrix, which is the matrix multiplication operation requirement of the matrix multiplication dedicated computing unit.

[0074] In order to better understand the data processing method applied to the GPU provided by the first embodiment of the present application, before introducing the specific execution method of the data processing method applied to the GPU provided by the first embodiment of the present application, the related concepts related to the GPU and integer types involved in the first embodiment of the present application are first introduced.

[0075] The GPU-related concepts involved in the first embodiment of the present application are: the target streaming multiprocessor (SM) is a computing unit used in the GPU. When the GPU executes a task, it will poll the SM through the task allocation unit to see if there are enough resources to execute the new task. If there are, a new task will be allocated to the SM. If not, the next SM will be checked. The factors that determine whether a new task can be allocated to the SM include: the memory of the shared memory used by each task, the number of registers used by each task, and some other restrictions.

[0076] When executing the assigned tasks, SM schedules threads through thread clusters (warps) to execute the same instructions in parallel. Generally, each warp contains 32 threads. During specific execution, according to the design of the GPU, each warp can schedule preset threads at a time to execute the same instructions in parallel. In the current common GPU, each warp can first schedule 16 threads at a time to execute the same instructions in parallel. After the execution of the first 16 threads is completed, another 16 threads are scheduled to execute the same instructions in parallel. In the first embodiment of the present application, there is no specific limitation on the number of threads that can be scheduled at a time for each warp. In addition, generally, each warp can also contain 64 threads. In the first embodiment of the present application, there is no limitation on the threads contained in the warp. When introducing the specific execution steps of the data processing method applied to the GPU provided in the first embodiment of the present application, the first embodiment of the present application is described in detail by taking each warp containing 32 threads as an example. In general, each thread corresponds to a register, and multiple registers can also be used. In addition, multiple warps share a shared memory. The register is the GPU on-chip cache. The number of bits of the register is 32 bits (binary digit, binary unit), that is, the register can store a 32-bit file. Specifically, registers and shared memory are generally SM resources. Assuming that an SM has 65536 registers and 1024 threads run on an SM, then a thread can use 64 registers. In addition, different types of GPUs have a predetermined limit on the number of registers available to a single thread.

[0077] The number of bits of an integer type refers to the number of bits in a number represented in binary form. For example, int8 means an integer with 8 bits in binary, and int16 means an integer with 16 bits in binary, and so on.

[0078] Regarding the data processing method applied to a GPU provided in the first embodiment of the present application, the following is combined with Figure 2 Provide explanation.

[0079] Figure 2A flowchart of a data processing method applied to a GPU provided in the first embodiment of the present application. Figure 2 The data processing method applied to a GPU shown includes: steps S201 to S303.

[0080] In step S201, a matrix to be processed is read from a memory corresponding to a target GPU into a register in a target streaming multiprocessor, where the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculations on the matrix to be processed.

[0081] The so-called target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculations on a matrix to be processed.

[0082] In the first embodiment of the present application, the matrix to be processed is a first matrix arranged in column-major order and / or a second matrix arranged in row-major order. The number of rows and columns of the so-called matrix to be processed is a natural number power of two. For two multiplication matrices that can perform matrix multiplication operations, the number of columns of the first matrix needs to be equal to the number of rows of the second matrix. For example, the first matrix is ​​a matrix of 8 rows and 16 columns, and the second matrix is ​​a matrix of 16 rows and 8 columns. The so-called matrix arranged in row-major order means that each row of the matrix is ​​placed continuously in the memory, and the so-called matrix arranged in column-major order means that each column of the matrix is ​​placed continuously in the memory.

[0083] When using the target GPU to perform matrix multiplication operations, the matrix to be operated needs to be first input into the memory corresponding to the target GPU. In the first embodiment of the present application, if the two multiplication matrices in the matrix multiplication operation include the matrix to be processed, then in the process of executing the matrix multiplication operation, the matrix to be processed needs to be first read from the memory corresponding to the target GPU into the register in the target streaming multiprocessor. For the specific execution process, please refer to Figure 3 , which is a schematic diagram of a matrix processing process to be processed provided in the first embodiment of the present application.

[0084] The specific implementation method of reading the matrix to be processed from the memory corresponding to the target GPU into the register of the target streaming multiprocessor is as follows: first, obtain the first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed. Then, according to the first multiple, read the elements of the same row in the matrix to be processed into the register in sequence starting from the first column.

[0085] In the first embodiment of the present application, when obtaining the first multiple, it is necessary to first obtain the number of bits of the register and the number of bits of the integer type corresponding to the elements in the matrix to be processed, and then determine the first multiple according to the number of bits of the register and the number of bits of the integer type corresponding to the elements in the matrix to be processed. Specifically, taking the second matrix in which the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits and the matrix arrangement is in row-major order and has 16 rows and 8 columns as an example, the specific implementation method of reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor is described, and the number of bits of other registers, the number of bits of the integer type corresponding to the elements in the matrix to be processed, or the number of rows and columns matrix will not be repeated here.

[0086] Since the number of bits in the register is 32 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits, at this time, the first multiple of the number of bits in the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed is 32 bits / 8 bits=4. In other words, 4 elements can be read into the register. In the process of reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, it is necessary to read every 4 elements of the same row in the matrix to be processed into the register starting from the first column. Figure 3 As shown, the first 4 columns of elements in the 0th row of the matrix to be processed with 16 rows and 8 columns are read into a register r0, the last 4 columns of elements in the 0th row are read into a register r1, the first 4 columns of elements in the first row of the matrix to be processed with 16 rows and 8 columns are read into a register r2, the last 4 columns of elements in the first row are read into a register r3… until all elements in the matrix to be processed with 16 rows and 8 columns are read into the registers respectively.

[0087] Please refer to Figure 2 In step S202, in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, a first preset operation is performed on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed.

[0088] The first preset operation is to merge the elements of each second multiple row in the matrix into one row in turn starting from the first row. The second multiple is the multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed.

[0089] In the first embodiment of the present application, the specific process of performing the first preset operation on the matrix to be processed and obtaining the initial matrix corresponding to the matrix to be processed is as follows: first, obtain a first read-in instruction for the matrix to be processed, and the first read-in instruction is a read-in instruction for the first integer type. Secondly, obtain the second multiple of the number of bits corresponding to the first integer type relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed. Thirdly, determine the first preset operation for the matrix to be processed based on the second multiple. Finally, starting from the first row, the elements of each second multiple row in the matrix to be processed are sequentially distributed and merged into one row, and read into the shared memory to obtain the initial matrix.

[0090] The first read instruction is specifically the prmt instruction. The prmt instruction has two input registers, one output register, and an index. It can extract 4 bytes from the 8 bytes of data in the two input registers according to the index and put them in the output register. The form of this instruction can be: prmt.b32d,a,b,index. For prmt.b32d,a,b,index, d is a 32-bit output register, a and b are 32-bit input registers, and index is a 32-bit register composed of 4 int8 data. The 4 bytes in register a are numbered 0, 1, 2, and 3 from low to high, and the 4 bytes in register b are numbered 4, 5, 6, and 7 from low to high. Assuming that index is 0x5140 (that is, the 4 int8s contained in index are 5, 1, 4, and 0 respectively), the data contained in the output d are byte 0 of register a, byte 4 of register b, byte 1 of register a, and byte 5 of register b from low to high. In the first embodiment of the present application, the prmt instruction may specifically be: prmt.b32 regd, r0, r2.

[0091] The second multiple can be two times. Generally, the number of bits corresponding to the first integer type is 16 bits, and the number of bits corresponding to the integer type of the elements in the matrix to be processed is 8 bits; or, the number of bits corresponding to the first integer type is 32 bits, and the number of bits corresponding to the integer type of the elements in the matrix to be processed is 16 bits. The following specifically takes the case where the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits, and the matrix is ​​arranged in row-major order with 16 rows and 8 columns, the number of bits corresponding to the first integer type is 16 bits, and the number of bits of the register is 32 bits as an example to illustrate the specific process of obtaining the initial matrix corresponding to the matrix to be processed. Please refer to Figure 4, which is a schematic diagram of an initial matrix processing process provided in the first embodiment of the present application. At this time, the first column element (0) in r0 is used as the first column element of the merged row 0, the first column element (8) in r2 is used as the second column element of the merged row 0, the second column element (1) in r0 is used as the third column element of the merged row 0, and the second column element (9) in r2 is used as the fourth column element of the merged row 0… The corresponding steps are performed in sequence until the elements in the matrix to be processed are read from the register into the shared memory corresponding to the register. Correspondingly, the elements of the merged matrix row 0 are, starting from the first column: 0, 8, 1, 9, 2, 10, 3, 11…; the elements of the first row of the merged matrix are, starting from the first column: 16, 24, 17, 25, 18, 26, 19, 27… The elements of the second row of the merged matrix are, starting from the first column… The elements of the merged matrix row 7 are, starting from the first column: 112, 120, 113, 121, 114, 122, 115, 123…

[0092] In addition, the second multiple can also be one. In this case, starting from the first row, the elements of each second multiple row in the matrix to be processed are sequentially distributed and merged into one row, and read into the shared memory. The specific implementation process of obtaining the initial matrix is: starting from the first row, the elements of each row in the matrix to be processed are sequentially distributed and merged into one row, and read into the shared memory to obtain the initial matrix.

[0093] Please refer to Figure 2 In step S203, in the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0094] The second preset operation is to transpose the matrix starting from the first column and taking every third multiple column in the matrix as a column to obtain a target matrix, and read the target matrix into the register. The third multiple is the third multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the initial matrix.

[0095] In the first embodiment of the present application, the specific process of performing the second preset operation on the initial matrix to obtain the target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit is as follows: first, obtain the second read-in instruction for the matrix to be processed, and the second read-in instruction carries the second integer type targeted by the second read-in instruction. Secondly, obtain the third multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix. Again, based on the third multiple, determine the second preset operation for the matrix to be processed. Finally, starting from the first column, take each third multiple column in the initial matrix as a column to perform matrix transposition, obtain the target matrix, and read the target matrix into the register.

[0096] The second read instruction is specifically the ldmatrix instruction, which is used to read matrix blocks from shared memory into registers. For the matrix calculation dedicated unit, the matrix operation is performed by a warp together. Before doing the matrix calculation, all threads of a warp need to read the data of the matrix into the register. An ldmatrix can read 1, 2, or 4 matrices into the register at a time. If 1 matrix is ​​read, threads 0 to 7 in the warp pass the starting address of 8 rows of matrix data to the ldmatrix instruction, and the line0 register of threads 8 to 31 is ignored. If 4 matrices are read, threads 0 to 7 in the warp input the actual address of the 8 rows of data of the first matrix to ldmatrix, threads 8 to 15 input the address of the 8 rows of the second matrix, and threads 16 to 23 and 24 to 31 respectively input the actual address of the 8 rows of the third and fourth matrices to ldmatrix. The form of this instruction can be as follows: ldmatrix.sync.aligned.m8n8*1.trans.shared.b16redm(line0). Specifically, the initial matrix is ​​8 rows and 16 columns, so thread 0 of a warp is responsible for reading the 4 numbers of row 0, column 0 to 3 into the register, thread 1 is responsible for reading the 4 numbers of row 0, column 4 to 7 into the register, and thread 4 is responsible for reading the 4 numbers of row 1, column 0 to 3 into the register. After reading, the data of 32 threads of a warp are pieced together to form a target matrix of 8 rows and 16 columns.

[0097] The third multiple is two. In general, the number of bits corresponding to the second integer type is 16 bits, and the number of bits corresponding to the integer type of the elements in the initial matrix is ​​8 bits; or, the number of bits corresponding to the second integer type is 32 bits, and the number of bits corresponding to the integer type of the elements in the initial matrix is ​​16 bits. At this time, starting from the first column, every third multiple column in the matrix to be processed is taken as a column to perform matrix transposition, obtain the target matrix, and read the target matrix into the register. The specific implementation method is: starting from the first column, every two columns in the matrix to be processed are taken as a column to perform matrix transposition, obtain the target matrix, and read the target matrix into the register.

[0098] The following specifically takes the case where the number of bits of the integer type corresponding to the elements in the initial matrix is ​​8 bits, the matrix is ​​arranged in row-major order, the second matrix has 16 rows and 8 columns, the number of bits corresponding to the first integer type is 16 bits, and the number of bits of the register is 32 bits as an example to illustrate the specific process of obtaining the target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit. Please refer to Figure 5 , which is a schematic diagram of a target matrix processing process provided in the first embodiment of the present application. The first and second columns of row 0 are used as the first column of row 0, the third and fourth columns of row 0 are used as the second column of row 0, the fifth and third columns of row 0 are used as the third column of row 0, the seventh and eighth columns of row 0 are used as the fourth column of row 0... and transpose. At this time, the target matrix is: the first column of row 0 is (0, 8), (16, 24)..., the first column of row 1 is (1, 9), (17, 25)..., the first column of row 2 is (2, 10), (18, 26)..., the first column of row 3 is (3, 11), (19, 27)...

[0099] After completing the transposition of the initial matrix, the transposed initial matrix is ​​read into the register to obtain the target matrix. For the specific implementation process of reading the transposed initial matrix into the register, please refer to step S201, the process of reading the target matrix from the memory corresponding to the target GPU into the register in the target stream multiprocessor, which will not be repeated here. In the first embodiment of the present application, the distribution of the target matrix after reading the target matrix from the memory corresponding to the target GPU into the register in the target stream multiprocessor is as follows: the first 4 columns of elements (0, 8, 16, 24) in the 0th row are read into a register r0, the 4-8 columns of elements (32, 40, 48, 56) in the 0th row are read into a register r1, the 4-8 columns of elements (64, 72, 80, 88) in the 0th row are read into a register r2, the 4-8 columns of elements (96, 104, 112, 120) in the 0th row are read into a register r3..., and all elements in the target matrix are read into the registers respectively.

[0100] In the first embodiment of the present application, after the target matrix is ​​read into the register, the SM can perform a multiplication operation on the target matrix by calling a warp.

[0101] In the first embodiment of the present application, a data processing method applied to a graphics processing unit (GPU) is provided. First, a matrix to be processed is read from a memory corresponding to a target GPU into a register in a target streaming multiprocessor, and the target streaming multiprocessor is a streaming multiprocessor for performing matrix calculations on the matrix to be processed; then, in the process of reading the matrix to be processed from the register into a shared memory corresponding to the register, a first preset operation is performed on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed; finally, in the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of a matrix multiplication dedicated calculation unit, and the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operations in the target streaming multiprocessor. The data processing method applied to a graphics processing unit (GPU) provided in the first embodiment of the present application can transform the matrix to be processed into a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit by sequentially performing two different preset operations on the matrix to be processed that does not meet the matrix multiplication operation requirements of the GPU, thereby improving the applicability of the GPU for matrix multiplication operations.

[0102] Second embodiment

[0103] Corresponding to the data processing method applied to a GPU provided in the first embodiment of the present application, the second embodiment of the present application further provides a data processing device, which is applied to a GPU. Since the device embodiment is basically similar to the first embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the first embodiment. The device embodiment described below is only illustrative.

[0104] Please refer to Figure 6 , which is a schematic diagram of a data processing device provided in the second embodiment of the present application.

[0105] The data processing device provided in the second embodiment of the present application includes:

[0106] A matrix reading unit 601 to be processed is used to read the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, where the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0107] A first preset operation execution unit 602 is used to execute a first preset operation on the matrix to be processed in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, so as to obtain an initial matrix corresponding to the matrix to be processed;

[0108] The second preset operation execution unit 603 is used to perform a second preset operation on the initial matrix during the process of reading the initial matrix from the shared memory into the register, so as to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit, and the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operations in the target streaming multiprocessor.

[0109] Optionally, the matrix to be processed reading unit 601 is specifically used to obtain a first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed; according to the first multiple, starting from the first column, read each first multiple of the elements in the same row in the matrix to be processed into the register.

[0110] Optionally, obtaining the first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed includes:

[0111] Obtain the number of bits of the register;

[0112] Obtain the number of digits of the integer type corresponding to the elements in the matrix to be processed;

[0113] The first multiple is determined according to the number of bits of the register and the number of bits of the integer type corresponding to the elements in the matrix to be processed.

[0114] Optionally, the first preset operation execution unit 602 is specifically used to obtain a first read instruction for the matrix to be processed, the first read instruction being a read instruction for a first integer type; obtaining a second multiple of the number of bits corresponding to the first integer type relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed; determining a first preset operation for the matrix to be processed based on the second multiple, the first preset operation being to merge the elements of each second multiple row in the matrix into one row in sequence starting from the first row; and to merge the elements of each second multiple row in the matrix to be processed into one row in sequence starting from the first row, and reading them into the shared memory to obtain the initial matrix.

[0115] Optionally, the second multiple is twice;

[0116] The step of merging the elements of every second multiple row in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix comprises: merging the elements of every two rows in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix.

[0117] Optionally, the second multiple is two, including: the number of bits corresponding to the first integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits;

[0118] Alternatively, the number of bits corresponding to the first integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 16 bits.

[0119] Optionally, the second multiple is one;

[0120] The step of distributing the elements of each second multiple row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix includes: distributing the elements of each row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix.

[0121] Optionally, the second preset operation execution unit 603 is specifically used to obtain a second read instruction for the initial matrix, the second read instruction carries a second integer type targeted by the second read instruction; obtain a third multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix; determine a second preset operation for the initial matrix based on the third multiple, the second preset operation being to transpose the matrix starting from the first column, taking each third multiple column in the matrix as a column, to obtain a target matrix, and to read the target matrix into a register; to transpose the matrix starting from the first column, taking each third multiple column in the initial matrix as a column, to obtain the target matrix, and to read the target matrix into the register.

[0122] Optionally, the third multiple is two;

[0123] The method of starting from the first column, taking each third multiple column in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register includes: starting from the first column, taking every two columns in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register.

[0124] Optionally, the third multiple is two, including: the number of bits corresponding to the third integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​8 bits;

[0125] Alternatively, the number of bits corresponding to the third integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​16 bits.

[0126] Optionally, the number of rows and the number of columns of the matrix to be processed are multiples of a natural number power of two.

[0127] Optionally, the matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order, and the first matrix and the second matrix are the first matrix and the second matrix of two multiplication matrices to be subjected to matrix multiplication operation.

[0128] Third embodiment

[0129] Corresponding to the data processing method applied to a GPU provided in the first embodiment of the present application, the third embodiment of the present application further provides an electronic device applied to a GPU. Since the third embodiment is basically similar to the first embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the first embodiment. The third embodiment described below is only illustrative.

[0130] Please refer to Figure 7 , which is a schematic diagram of an electronic device provided in the third embodiment of the present application.

[0131] The electronic device provided in the third embodiment of the present application includes:

[0132] Processor 710;

[0133] and a memory 702 for storing a program of the text recognition method. After the device is powered on and passes the program of the data processing method applied to the GPU, the following steps are performed:

[0134] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0135] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0136] In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0137] It should be noted that the detailed description of the electronic device provided in the third embodiment of the present application can refer to the relevant description of the data processing method applied to the GPU provided in the first embodiment of the present application, and will not be repeated here.

[0138] Fourth embodiment

[0139] Corresponding to the data processing method applied to the GPU provided in the first embodiment of the present application, the third embodiment of the present application further provides a storage medium applied to the GPU. Since the fourth embodiment is basically similar to the first embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the first embodiment. The fourth embodiment described below is only illustrative.

[0140] The storage medium provided in the fourth embodiment of the present application stores a program of a data processing method applied to a GPU, and the program is executed by a processor to perform the following steps:

[0141] Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed;

[0142] In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed;

[0143] In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0144] It should be noted that the detailed description of the storage medium provided in the fourth embodiment of the present application can refer to the relevant description of the data processing method applied to the GPU provided in the first embodiment of the present application, and will not be repeated here.

[0145] Fifth embodiment

[0146] Corresponding to the data processing method applied to the GPU provided in the first embodiment of the present application, the fifth embodiment of the present application also provides another data processing method applied to the GPU. Since the fifth embodiment is basically similar to the first embodiment, the description is relatively simple, and the relevant parts refer to the partial description of the first embodiment. The fifth embodiment described below is only illustrative.

[0147] Figure 8 A flowchart of a data processing method applied to a GPU provided in the fifth embodiment of the present application. Figure 8 The data processing method applied to a GPU shown includes: steps S801 to S805.

[0148] Step S801: Obtain a matrix to be processed that requires matrix multiplication operation.

[0149] The matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order. The first matrix and the second matrix are the first matrix and the second matrix of the two matrix multiplication operations to be performed. That is, the matrix to be processed is the first matrix arranged in column-major order and / or the second matrix arranged in row-major order.

[0150] In the fifth embodiment of the present application, the so-called row-major matrix arrangement means that each row of the matrix is ​​placed continuously in the memory, and the so-called column-major matrix arrangement means that each column of the matrix is ​​placed continuously in the memory.

[0151] Step S802: read the matrix to be processed from the memory corresponding to the target GPU into the register of the target streaming multiprocessor, where the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed.

[0152] In the fifth embodiment of the present application, the register is a GPU on-chip cache, and the number of bits of the register is 32 bits (binary digit, binary unit), that is, the register can store a file of a size of 32 bits.

[0153] Step S803: in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, a first preset operation is performed on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed.

[0154] In the fifth embodiment of the present application, the so-called first preset operation is to sequentially distribute and merge the elements of every second multiple row in the matrix into one row starting from the first row.

[0155] Step S804: In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor.

[0156] In the fifth embodiment of the present application, the so-called target matrix is ​​a target matrix that meets the matrix multiplication operation requirements of the target GPU, and the so-called target GPU matrix multiplication operation requirements are matrix multiplication operation requirements predetermined according to the hardware conditions of the target GPU. For example, when the integer type of the elements is an input matrix of INT8, it is required that the arrangement form of the first matrix in the two multiplied matrices is a row-major matrix, and the arrangement form of the second matrix is ​​a column-major matrix. At this time, the arrangement form of the first matrix in the two multiplied matrices is a row-major matrix, and the arrangement form of the second matrix is ​​a column-major matrix, which is the target GPU matrix multiplication operation requirement.

[0157] The so-called second preset operation is to transpose the matrix starting from the first column and taking every third multiple column in the matrix as a column to obtain a target matrix, and read the target matrix into a register.

[0158] Step S805: Execute a matrix multiplication operation on the target matrix through the dedicated matrix multiplication computing unit.

[0159] Although the present application is disclosed as above in the preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

[0160] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0161] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0162] 1. Computer readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include non-transitory media such as modulated data signals and carrier waves.

[0163] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A data processing method applied to a graphics processing unit (GPU), characterized in that: include: Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed; In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed, wherein the first preset operation is to sequentially distribute and merge the elements of each second multiple row in the matrix to be processed into one row starting from the first row, and the second multiple is a multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed; In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor, wherein the second preset operation is to transpose the matrix starting from the first column and taking every third multiple column in the initial matrix as a column to obtain a target matrix, and read the target matrix into the register, and the third multiple is a multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix.

2. The data processing method applied to GPU according to claim 1, characterized in that: The step of reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor includes: Obtaining a first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed; According to the first multiple, starting from the first column, the elements of the same row in the matrix to be processed, which are each a multiple of the first multiple, are read into the register in sequence.

3. The data processing method applied to GPU according to claim 2, characterized in that: The obtaining of the first multiple of the number of bits of the register relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed comprises: Obtain the number of bits of the register; Obtain the number of digits of the integer type corresponding to the elements in the matrix to be processed; The first multiple is determined according to the number of bits of the register and the number of bits of the integer type corresponding to the elements in the matrix to be processed.

4. The data processing method applied to GPU according to claim 1, characterized in that: The performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed includes: Obtaining a first read instruction for the matrix to be processed, wherein the first read instruction is a read instruction for a first integer type; Obtain a second multiple of the number of bits corresponding to the first integer type relative to the number of bits of the integer type corresponding to the elements in the matrix to be processed; Determining a first preset operation for the matrix to be processed according to the second multiple; Starting from the first row, the elements of each second multiple row in the matrix to be processed are sequentially distributed and merged into one row, and read into the shared memory to obtain the initial matrix.

5. The data processing method applied to GPU according to claim 4, characterized in that: The second multiple is twice; The step of merging the elements of every second multiple row in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix comprises: merging the elements of every two rows in the matrix to be processed into one row in sequence and at intervals starting from the first row, and reading the elements into the shared memory to obtain the initial matrix.

6. The data processing method applied to GPU according to claim 5, characterized in that: The second multiple is two, including: the number of bits corresponding to the first integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 8 bits; Alternatively, the number of bits corresponding to the first integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the matrix to be processed is 16 bits.

7. The data processing method applied to GPU according to claim 4, characterized in that: The second multiple is one; The step of distributing the elements of each second multiple row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix includes: distributing the elements of each row in the matrix to be processed starting from the first row in sequence into a row at intervals, reading the elements into the shared memory, and obtaining the initial matrix.

8. The data processing method applied to GPU according to claim 1, characterized in that: The performing a second preset operation on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the target GPU includes: Obtaining a second read-in instruction for the initial matrix, wherein the second read-in instruction carries a second integer type for which the second read-in instruction is intended; Obtain a third multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix; Determining a second preset operation for the initial matrix according to the third multiple; Starting from the first column, each column of the third multiple in the initial matrix is ​​taken as a column to perform matrix transposition to obtain the target matrix, and the target matrix is ​​read into the register.

9. The data processing method applied to GPU according to claim 8, characterized in that: The third multiple is two times; The method of starting from the first column, taking each third multiple column in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register includes: starting from the first column, taking every two columns in the initial matrix as a column to perform matrix transposition to obtain the target matrix, and reading the target matrix into the register.

10. The data processing method applied to GPU according to claim 9, characterized in that: The third multiple is two, including: the number of bits corresponding to the second integer type is 16 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​8 bits; Alternatively, the number of bits corresponding to the second integer type is 32 bits, and the number of bits of the integer type corresponding to the elements in the initial matrix is ​​16 bits.

11. The data processing method applied to GPU according to claim 1, characterized in that: The number of rows and the number of columns of the matrix to be processed are multiples of a natural number power of two.

12. The data processing method applied to GPU according to claim 11, characterized in that: The matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order, and the first matrix and the second matrix are the first matrix and the second matrix of two multiplication matrices to be subjected to matrix multiplication operation.

13. A data processing device, characterized in that: Applied to a GPU, the device comprises: a to-be-processed matrix reading unit, used to read the to-be-processed matrix from a memory corresponding to a target GPU into a register in a target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the to-be-processed matrix; A first preset operation execution unit is used to execute a first preset operation on the matrix to be processed in the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, so as to obtain an initial matrix corresponding to the matrix to be processed, wherein the first preset operation is to sequentially distribute and merge the elements of each second multiple row in the matrix to be processed into one row starting from the first row, and the second multiple is a multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed; A second preset operation execution unit is used to perform a second preset operation on the initial matrix during the process of reading the initial matrix from the shared memory into the register to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated calculation unit, wherein the matrix multiplication dedicated calculation unit is a calculation unit dedicated to matrix operations in the target streaming multiprocessor, wherein the second preset operation is to transpose the matrix starting from the first column and treat every third multiple column in the initial matrix as a column to obtain a target matrix, and read the target matrix into the register, wherein the third multiple is a multiple of the number of bits corresponding to the second integer type relative to the number of bits corresponding to the integer type of the elements in the initial matrix.

14. An electronic device, characterized in that: Applied to GPU, including: processor; and a memory for storing a program of a text recognition method. After the device is powered on and the program of the data processing method applied to the GPU is executed, the following steps are performed: Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed; In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed, wherein the first preset operation is to sequentially distribute and merge the elements of each second multiple row in the matrix to be processed into one row starting from the first row, and the second multiple is a multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed; In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor, wherein the second preset operation is to transpose the matrix starting from the first column and taking every third multiple column in the initial matrix as a column to obtain a target matrix, and read the target matrix into the register, and the third multiple is a multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix.

15. A storage medium, characterized in that: A program for a data processing method applied to a GPU is stored, and the program is executed by a processor to perform the following steps: Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed; In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed, wherein the first preset operation is to sequentially distribute and merge the elements of each second multiple row in the matrix to be processed into one row starting from the first row, and the second multiple is a multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed; In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, and the matrix multiplication dedicated computing unit is a computing unit dedicated to matrix operations in the target streaming multiprocessor, wherein the second preset operation is to transpose the matrix starting from the first column and taking every third multiple column in the initial matrix as a column to obtain a target matrix, and read the target matrix into the register, and the third multiple is a multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix.

16. A data processing method applied to a GPU, characterized in that: include: Obtaining a matrix to be processed that needs to be subjected to a matrix multiplication operation, wherein the matrix to be processed includes at least one of a first matrix arranged in column-major order and a second matrix arranged in row-major order, and the first matrix and the second matrix are a first matrix and a second matrix of two multiplication matrices to be subjected to a matrix multiplication operation; Reading the matrix to be processed from the memory corresponding to the target GPU into the register in the target streaming multiprocessor, wherein the target streaming multiprocessor is a streaming multiprocessor used to perform matrix calculation on the matrix to be processed; In the process of reading the matrix to be processed from the register into the shared memory corresponding to the register, performing a first preset operation on the matrix to be processed to obtain an initial matrix corresponding to the matrix to be processed, wherein the first preset operation is to sequentially distribute and merge the elements of each second multiple row in the matrix to be processed into one row starting from the first row, and the second multiple is a multiple of the number of bits corresponding to the first integer type relative to the number of bits corresponding to the integer type of the elements in the matrix to be processed; In the process of reading the initial matrix from the shared memory into the register, a second preset operation is performed on the initial matrix to obtain a target matrix that meets the matrix multiplication operation requirements of the matrix multiplication dedicated computing unit, the matrix multiplication dedicated computing unit being a computing unit dedicated to matrix operations in the target streaming multiprocessor, wherein the second preset operation is to perform matrix transposition on every third multiple column in the initial matrix starting from the first column as a column to obtain a target matrix, and read the target matrix into the register, the third multiple being a multiple of the number of bits corresponding to the second integer type relative to the number of bits of the integer type corresponding to the elements in the initial matrix; The matrix multiplication dedicated computing unit performs a matrix multiplication operation on the target matrix.

Citation Information

Patent Citations

  • Method and device of improving matrix multiplication calculation performance of graphics processing unit (GPU)

    CN107622037A

  • Data processing method and device and electronic equipment

    CN112069460A