Method for accessing data in shared memory, electronic equipment, storage medium and computer program product
By determining the number of memory blocks occupied by the target matrix in shared memory and selecting the appropriate address hashing algorithm, the memory block conflict problem that data memory access in shared memory is easily caused, and the performance of shared memory is significantly improved.
Patent Information
- Application Number
- CN202510414829.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Existing solutions for data access in shared memory are prone to memory block conflicts, resulting in memory access delay and data read and write efficiency, which seriously affects the performance of shared memory.
By determining the number of memory blocks occupied by each row of data in the shared memory of the target matrix, and determining the data access layout based on the predetermined address hashing algorithm, memory block conflicts are avoided.
It effectively avoids the occurrence of memory block conflicts in shared memory and significantly improves the performance of shared memory.
Smart Images

Figure CN119938362A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of data processing, and more specifically to a method, electronic device, storage medium, and computer program product for accessing data in a shared memory. Background Art
[0002] Artificial intelligence (AI) processors are increasingly used to process large-scale data and intensive computing tasks because they can execute multiple threads simultaneously and process vector data and tensor data, making them efficient in parallel computing.
[0003] However, when using AI processors (such as graphics processing units (GPUs)) to perform computing tasks such as matrix multiplication and accumulation (MMA), existing solutions for data access in shared memory usually only access data in shared memory based on data layout such as in the initial matrix used for matrix multiplication and accumulation. In this case, when data related to the transposed matrix corresponding to the initial matrix is accessed in the shared memory, a bank conflict will inevitably occur, which increases the latency of memory access and reduces the efficiency of data reading and writing, thereby seriously affecting the performance of the shared memory.
[0004] In summary, the existing solutions for accessing data in shared memory have the following shortcomings: memory block conflicts are prone to occur, and the performance of shared memory cannot be fully utilized. Summary of the invention
[0005] In response to the above problems, the present invention provides a method, electronic device, storage medium and computer program product for accessing data in a shared memory, so that when data is accessed in the shared memory based on a matrix with any data type or its corresponding transposed matrix, the occurrence of memory block conflicts in the shared memory can be avoided, thereby significantly improving the shared memory performance.
[0006] According to a first aspect of the present invention, a method for accessing data in a shared memory is provided, comprising: determining the number of memory blocks occupied by each row of data of a target matrix in the shared memory; determining a data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data of the target matrix in the shared memory according to a predetermined address hash algorithm, so as to store the data in the target matrix in the shared memory based on the data access layout.
[0007] In some embodiments, the shared memory includes multiple memory blocks with the same bandwidth. In these embodiments, determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory includes: determining the total number of bytes of each row of data of the target matrix; determining the bandwidth size of the memory block; and determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory based on the total number of bytes of each row of data of the target matrix and the bandwidth size of the memory block.
[0008] In some embodiments, determining the data access layout for the target matrix includes: performing binary conversion on the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory; selecting an address hash algorithm to be adopted from predetermined address hash algorithms based on the position of the first "1" from low to high in the binary conversion result; and determining the data access layout for the target matrix based at least on the address hash algorithm to be adopted.
[0009] In some embodiments, selecting an address hash algorithm to be adopted from predetermined address hash algorithms includes: when the position of the first "1" in the binary conversion result satisfies a first predetermined condition, the address hash algorithm to be adopted includes performing a hash operation on no more than a first number of data addresses; when the position of the first "1" in the binary conversion result satisfies a second predetermined condition, the address hash algorithm to be adopted includes performing a hash operation on a second number of data addresses.
[0010] In some embodiments, the first predetermined condition is that the position of the first "1" in the binary conversion result is any one of the third position, the fourth position or the fifth position from the low position to the high position in the binary conversion result.
[0011] In some embodiments, the second predetermined condition is that the position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low position to the high position in the binary conversion result.
[0012] In some embodiments, the first number is related to the position of the first "1" in the binary conversion result, and the second number is five.
[0013] In some embodiments, the first number is equal to the number of bits of the first "1" in the binary conversion result from the lower bit to the higher bit in the binary conversion result minus one.
[0014] In some embodiments, at least based on the number of memory blocks occupied by each row of data of the target matrix in the shared memory, according to a predetermined address hash algorithm, determining the data access layout for the target matrix includes: determining a value of a stride parameter for data access; selecting an address hash algorithm to be adopted from predetermined address hash algorithms based on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory; and determining the data access layout for the target matrix based on the address hash algorithm to be adopted.
[0015] According to a second aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention.
[0016] According to a third aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of the first aspect of the present invention.
[0017] According to a fourth aspect of the present invention, there is provided a computer program product comprising a computer program, wherein the computer program is used to perform the method of the first aspect of the present invention when a machine is used.
[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements.
[0020] Figure 1 An exemplary architectural diagram of a processor according to an embodiment of the present invention is shown.
[0021] Figure 2 An exemplary architecture diagram of a stream processor cluster according to an embodiment of the present invention is shown.
[0022] Figure 3 An exemplary architecture diagram of a shared memory according to an embodiment of the present invention is shown.
[0023] Figure 4A An exemplary schematic diagram is shown of the memory access layout of a matrix whose data type is 32-bit data in a shared memory.
[0024] Figure 4BAn exemplary schematic diagram is shown of the memory access layout of a matrix whose data type is 32-bit data in a shared memory.
[0025] Figure 4C An exemplary schematic diagram is shown of the memory access layout of a matrix whose data type is 32-bit data in a shared memory.
[0026] Figure 5 A schematic framework diagram of a controller for controlling data access to a shared memory according to an embodiment of the present invention is shown.
[0027] Figure 6 A flow chart of a method for accessing data from a shared memory according to an embodiment of the present invention is shown.
[0028] Figure 7 A schematic diagram showing a data access layout of a target matrix according to an embodiment of the present invention is shown.
[0029] Figure 8 A schematic diagram showing a data access layout of a target matrix according to yet another embodiment of the present invention is shown.
[0030] Fig. 9 A schematic diagram showing part of the data access layout of a target matrix according to another embodiment of the present invention.
[0031] Fig.10 A schematic diagram showing a data access layout of a target matrix according to an embodiment of the present invention is shown.
[0032] Fig.11 A schematic diagram showing a data access layout of a target matrix according to yet another embodiment of the present invention is shown. DETAILED DESCRIPTION
[0033] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0034] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0035] Artificial Intelligence (AI) processors, such as Graphics Processing Unit (GPU), General-purpose Graphics Processing Unit (GPGPU), Tensor Processing Unit (TPU), Neural Network Processing Unit (NPU), Deep Learning Processing Unit (DPU), Accelerated Processing Unit (APU), etc., are increasingly widely used to process large-scale data and intensive computing tasks because they can execute multiple threads and large amounts of data simultaneously, allowing efficient parallel computing.
[0036] Figure 1 An exemplary architecture diagram of a processor 100 according to an embodiment of the present invention is shown. It should be understood that the processor 100 may further include additional units and / or modules not shown, and the scope of the present invention is not limited in this respect.
[0037] Regarding the processor 100 , it may be any appropriate type of artificial intelligence processor such as a graphics processing unit (GPU), a general purpose graphics processing unit (GPGPU), etc.
[0038] For example, a graphics processing unit (GPU) Figure 1 The processor 100 shown may include multiple streaming processor clusters (SPCs) 110 and global memory 120, wherein the global memory 120 may be configured to be accessed by the multiple streaming processor clusters 110 to implement data reading and writing. In some other cases, the processor 100 includes multiple global memories 120, so that each of the multiple streaming processor clusters 110 corresponds to one of the multiple global memories 120 to implement data reading and writing.
[0039] The global memory 120 may be a high bandwidth memory (HBM) or a double data rate (DDR) synchronous dynamic random access memory, or a high-speed cache that works in cooperation with the high bandwidth memory or the synchronous dynamic random access memory.
[0040] With respect to the plurality of stream processor clusters 110, in some embodiments, the processor 100 may include, but is not limited to, such as 4, 8, 12, or 16 stream processor clusters.
[0041] In an embodiment of the present invention, a plurality of stream processor clusters 110 configured in the processor 100 may share a secondary cache (L2 Cache) ( Figure 1 ) and global memory 120.
[0042] Furthermore, for each stream processor cluster, multiple computing units (Computation Unit, CU) may be configured therein, and each computing unit may include a vector core (Vector Core, VC) unit and a tensor core (Tensor Core, TC) unit for computing implementation. Specifically, Figure 2 Shows Figure 1 An exemplary architectural diagram of a stream processor cluster 110 in FIG.
[0043] like Figure 2 As shown, the stream processor cluster 110 may include multiple computing units (CUs) 210, wherein each computing unit 210 includes multiple vector core (VC) units 220 and one tensor core (TC) unit 230. Further, these multiple vector core (VC) units 220 and one tensor core (TC) unit 230 may share data through a vector register (Vector Register) 240 and a shared memory (Shared Memory) 250. Specifically, in combination Figure 2 In the stream processor cluster 110 shown, the shared memory 250 can be configured between the vector register 240 and the global memory 120. When the data in the global memory 120 is loaded into the vector register 240, or when the data in the vector register 240 is stored in the global memory 120, the data will first be sent to the shared memory 250, and operations such as transposition will be performed on the data when reading and writing the shared memory 250.
[0044] For example, if data in the global memory 120 is to be loaded into the vector register 240, the data in the global memory 120 (e.g., high bandwidth memory HBM) must first be moved to the shared memory 250, and then moved to the vector register 240 via the shared memory 250, so that the computing core (e.g., vector core unit 220 or tensor core unit 230) in the computing unit 210 can perform calculations based on the data in the vector register 240. For another example, when the computing unit 210 completes the calculation and obtains the corresponding calculation result, if the calculation result is to be stored from the vector register 240 to the global memory 120, the calculation result must first be moved from the vector register 240 to the shared memory 250, and then moved to the global memory 120 via the shared memory 250, so as to realize the storage of the calculation result in the global memory 120. It can be seen that the data access (i.e., storage and reading) efficiency of the shared memory 250 determines the efficiency of the processor as a system in processing data.
[0045] In order to improve the data access efficiency of the shared memory 250, the shared memory 250 can be divided into multiple memory blocks (Banks) of the same size, and these memory blocks can be accessed simultaneously, so that the memory access bandwidth of the shared memory 250 is significantly increased. Figure 3 Describe the architecture of shared memory 250.
[0046] Figure 3 An exemplary architectural diagram of the shared memory 250 according to an embodiment of the present invention is shown.
[0047] like Figure 3 As shown, the shared memory 250 is divided into 32 memory blocks (i.e., Bank-0 to Bank-31), and these 32 memory blocks can be accessed simultaneously. For example, for a processor with a single instruction multiple thread (SIMT)-32 architecture, each thread warp (wrap) of the processor includes 32 threads, and these 32 threads can access 32 memory blocks of the shared memory 250 respectively, that is, each thread can access a separate memory block and process the data in the memory block it accesses accordingly.
[0048] like Figure 3 As shown, the bandwidth of each memory block of the shared memory 250 can be 32 bits. Taking data of a double word (Dword) data type as an example, since a data of a double word data type occupies 32 bits (i.e., 32 bits), each address in each memory block can be used to store a data of a double word data type. If each memory block has N (where N is an integer greater than or equal to 1) different addresses for storing data, then correspondingly, each memory block can store at most N data of a double word data type. Therefore, for Figure 3The shared memory 250 shown may have 32*N data of double word data type.
[0049] However, when accessing data in the shared memory 250, for the same memory block, only one data of a double word data type can be read from it or one data of a double word data type can be written into it in the same clock cycle. Therefore, when multiple different requests expect to access data in different addresses in the same memory block at the same time in the same clock cycle, such as performing a read operation or a write operation at the same time, a bank conflict will occur. In this case, the bank conflict will increase the latency of memory access and reduce the efficiency of data reading and writing, thereby seriously affecting the performance of the shared memory 250.
[0050] In particular, when using an AI processor to perform computing tasks, for example, when performing operations such as matrix multiplication accumulation (MMA) in a GPU, it is usually necessary to perform read and write operations on data in a shared memory based on both an initial matrix for matrix multiplication accumulation and its corresponding transposed matrix. In this case, if data is written to the shared memory based only on the data layout in the initial matrix for matrix multiplication accumulation, when the data of the transposed matrix corresponding to the initial matrix is read from the shared memory, a memory block conflict will inevitably occur.
[0051] Take the calculation task on 32-bit data as an example. FIG. 4A to FIG. 4C The diagram shows an exemplary memory access layout of a matrix whose data type is 32-bit data in the shared memory 250 .
[0052] As described above, since the bandwidth of each memory block in the shared memory 250 is 32 bits, for 32-bit data, each data can be stored in one entry of the memory block of the shared memory 250. If a basic matrix with a dimension of 8×4 (i.e., eight rows and four columns) is used as the initial tiled matrix, for a matrix 400 with a dimension of 16×8 (i.e., sixteen rows and eight columns) and a data type of 32-bit data, it can be divided into four basic matrices with a dimension of 8×4, such as Figure 4A shown.
[0053] Further, since each basic matrix includes 32 data, and the shared memory 250 includes 32 memory blocks, the 32 data in each basic matrix can be stored in different memory blocks of the shared memory 250, and each data occupies one entry of the memory block, such as Figure 4B shown.
[0054] However, when the data read from the shared memory needs to be transposed to obtain a matrix with a dimension of 8×4, the original data is stored in the shared memory as a basic matrix with a dimension of 4×8, such as Figure 4C shown. Figure 4C In , four basic matrices of dimension 4×8 form a matrix 400' of dimension 16×8. Figure 4C The matrix 400 ′ before transposition includes four basic matrices with dimensions of 4×8 (ie, four rows and eight columns), namely, a first basic matrix 401 , a second basic matrix 402 , a third basic matrix 403 , and a fourth basic matrix 404 .
[0055] Further by Figure 4C It can be seen that if the data in the shared memory 250 is to be accessed based on the first basic matrix 401, since the data corresponding to the first basic matrix 401 are data 0 to data 15 and data 32 to data 47, however, referring to Figure 4B It can be seen that, for example, data 0 and data 32 are stored in different entries of the same memory block of the shared memory 250, that is, data 0 and data 32 are both stored in Bank-0. As described above, only data in one address in the same memory block can be accessed in the same clock cycle, so if data 0 and data 32 stored in Bank-0 are to be accessed, at least two cycles are required to read data 0 and data 32. Similarly, data 1 and data 33 in the first basic matrix 401 are stored in the same memory block of the shared memory 250, data 2 and data 34 are stored in the same memory block of the shared memory 250, and so on. It can be seen that if the data related to the first basic matrix 401 is to be accessed in the shared memory, at least two cycles are required to complete.
[0056] Similarly, for Figure 4C The second basic matrix 402, the third basic matrix 403, and the fourth basic matrix 404 in the same manner require at least two cycles to complete their respective data accesses. In other words, when performing data accesses to the shared memory for the transposed matrix, due to the memory block conflict, the processor only uses part of the memory read and write bandwidth in one cycle, thereby reducing the data access efficiency and increasing the latency of data access in the shared memory.
[0057] From the above, it can be seen that the existing data access scheme for shared memory will inevitably cause memory block conflicts when accessing data for matrices that need to be transposed, which increases the latency of memory access and reduces data reading and writing efficiency, thereby seriously affecting the performance of shared memory and further reducing the performance of AI processors.
[0058] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, an example embodiment of the present invention proposes a scheme for data access in a shared memory, by determining the number of memory blocks occupied by each row of data of a target matrix in the shared memory; at least based on the number of memory blocks occupied by each row of data of the target matrix in the shared memory, according to a predetermined address hash algorithm, a data access layout for the target matrix is determined, so that the data in the target matrix is stored in the shared memory based on the data access layout, so that when an initial matrix or a corresponding transposed matrix for matrix multiplication and accumulation having any data type is read from the shared memory, the occurrence of memory block conflicts in the shared memory can be avoided, thereby significantly improving the shared memory performance.
[0059] The following will be combined Figures 5 to 8 A scheme for accessing data in a shared memory according to the present invention is described in detail.
[0060] Figure 5 A schematic framework diagram of a controller 500 for controlling data access to a shared memory according to an embodiment of the present invention is shown. It should be understood that the controller 500 may further include additional units and / or modules not shown, and the scope of the present invention is not limited in this respect.
[0061] Regarding the controller 500, it may be configured to receive requests for a shared memory such as Figure 3 According to an embodiment of the present invention, the memory access request may be a write request, and the write request may include information such as data to be written into the shared memory, an address of the write data, etc. According to an embodiment of the present invention, the memory access request may also be a read request.
[0062] like Figure 5 As shown, the controller 500 may include a hash operation unit 510 to perform a hash operation based on a memory access request received by the controller 500. According to an embodiment of the present invention, for example, the hash operation unit 510 may determine the number of memory blocks occupied by each row of data of the target matrix related to the data to be written into the shared memory in the shared memory based on the memory access request received by the controller 500; and based on at least the number of memory blocks occupied by each row of data of the target matrix in the shared memory, perform a hash operation on the address of the write data in the memory access request according to a predetermined address hash algorithm to determine the data access layout for the current target matrix, so that the controller 500 can then store the data in the target matrix into the shared memory 250 based on the determined data access layout.
[0063] The following will be combined Figure 6 A scheme for accessing data in a shared memory according to an embodiment of the present invention is described in detail.
[0064] Figure 6 The flowchart of the method 600 for accessing data in a shared memory according to an embodiment of the present invention is shown. It should be understood that the method 600 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this respect.
[0065] In step 601 , the controller 500 determines the number of memory blocks occupied by each row of data of the target matrix in the shared memory.
[0066] Regarding the target matrix, data therein may be stored into a shared memory so that data corresponding to the matrix for the matrix multiplication-accumulation operation can be subsequently read from the shared memory.
[0067] According to an embodiment of the present invention, the target matrix may have different matrix sizes. Specifically, the size of each row of data in the target matrix may be determined according to the first size of the target matrix in the column dimension. For example, the first size of the target matrix in the column dimension may be denoted as K. In some embodiments, K may be 16 bytes, that is, the size of each row of data in the target matrix is 16 bytes. In some other embodiments, K may be 32 bytes, that is, the size of each row of data in the target matrix is 32 bytes. In other embodiments, K may also be 48 bytes, 64 bytes, or other multiples of 4 bytes, etc., which are not limited by the present invention.
[0068] Further, according to the inventive concept of the present invention, the number of memory blocks occupied by each row of data of the target matrix in the shared memory can be determined at least according to the size of each row of data in the target matrix. Specifically, according to an embodiment of the present invention, determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory can include: determining the total number of bytes of each row of data of the target matrix; determining the bandwidth size of the memory block; and determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory based on the total number of bytes of each row of data of the target matrix and the bandwidth size of the memory block.
[0069] Regarding the bandwidth size of the memory block, Figure 3 Taking the shared memory 250 shown as an example, the bandwidth size of each memory block in the shared memory 250 packet is 32 bits, that is, each row address in a memory block can be used to access 32 bits of data, that is, it can be used to access data of a double word (Dword) data type.
[0070] According to an embodiment of the present invention, then the size of each row of data in the target matrix can be divided by the bandwidth size of the memory block to calculate the number of memory blocks that each row of data of the target matrix occupies in the shared memory. In certain embodiments, the total number of bytes of each row of data in the target matrix can be divided by the bandwidth size of the memory block to determine the number of memory blocks that each row of data of the target matrix occupies in the shared memory. In other embodiments, it is also possible to determine how many Dwords each row of data in the target matrix can be equivalent to based on the total number of bytes of each row of data in the target matrix, and then determine the number of memory blocks that each row of data of the target matrix occupies in the shared memory based on the number of Dwords that each row of address can access in the memory block.
[0071] For example, taking the example that the size of each row of data in the above target matrix is 16 bytes, since the total number of bytes of each row of data in the target matrix is 16 bytes (i.e., 128 bits, equivalent to 4 Dwords), for a memory block with a bandwidth size of 32 bits in the shared memory 250, each row of data in the target matrix can occupy 4 memory blocks in the shared memory 250. That is, if the size of each row of data in the target matrix is 16 bytes, the number of memory blocks occupied by each row of data in the shared memory (such as the shared memory 250) is 4. Similarly, when the total number of bytes of each row of data of the target matrix is 32 bytes (i.e., 256 bits, equivalent to 8 Dwords), each row of data will occupy 8 memory blocks in the shared memory; when the total number of bytes of each row of data of the target matrix is 48 bytes (i.e., 384 bits, equivalent to 12 Dwords), each row of data will occupy 12 memory blocks in the shared memory; when the total number of bytes of each row of data of the target matrix is 64 bytes (i.e., 512 bits, equivalent to 16 Dwords), each row of data will occupy 16 memory blocks in the shared memory; and so on.
[0072] In step 602, the controller 500 determines the data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data of the target matrix in the shared memory according to a predetermined address hash algorithm, so as to store the data in the target matrix into the shared memory based on the data access layout.
[0073] Regarding the predetermined address hash algorithm, it may be a predetermined hash operation performed on the storage address corresponding to the data read from the target matrix, so that the data can be stored in the corresponding memory block in the shared memory based on the storage address obtained by the hash operation. For example, in an embodiment of the present invention, the hash operation may include an XOR operation on the storage address, which will be described in detail below and will not be repeated here.
[0074] According to the inventive concept of the present invention, different address hash algorithms can be used for different numbers of memory blocks occupied by each row of data of the target matrix in the shared memory. For example, in some embodiments of the present invention, binary conversion can be performed on the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory to determine the address hash algorithm to be used based on the binary conversion result.
[0075] Regarding the binary conversion of the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory, according to an embodiment of the present invention, it can be corresponding to determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory, and converting the determined quantity value into a binary representation.
[0076] Continuing with the above example, for example, if the number of memory blocks occupied by each row of data in the target matrix in the shared memory is 4, it can be converted into binary form, and the result after binary conversion can be represented as b0100. Similarly, if the number of memory blocks occupied by each row of data in the target matrix in the shared memory is 8, the result after binary conversion can be represented as b1000; if the number of memory blocks occupied by each row of data in the target matrix in the shared memory is 12, the result after binary conversion can be represented as b1100; if the number of memory blocks occupied by each row of data in the target matrix in the shared memory is 16, the result after binary conversion can be represented as b10000.
[0077] On this basis, the position of the first "1" can be determined according to the binary conversion result. For example, for the binary conversion result b0100, the first "1" is in the third position (also called the "third digit") from the low bit to the high bit (that is, the digits from right to left) in the binary conversion result.
[0078] Regarding the "low bit" and the "high bit", usually, taking the binary conversion result as an example, the further to the left the digit is, the higher the digit is; the further to the right the digit is, the lower the digit is.
[0079] Similarly, for the binary conversion result b1000, the first "1" is in the fourth position from the low bit to the high bit in the binary conversion result; for the binary conversion result b1100, the first "1" is in the third position from the low bit to the high bit in the binary conversion result; for the binary conversion result b10000, the first "1" is in the fifth position from the right to the left in the binary conversion result.
[0080] According to the inventive concept of the present invention, the address hash algorithm to be adopted can then be determined based on the position of the first "1" in the binary conversion result. Specifically, the address hash algorithm to be adopted can be selected from the predetermined address hash algorithms based on the position of the first "1" in the binary conversion result, so as to determine the data access layout for the target matrix based at least on the address hash algorithm to be adopted.
[0081] Regarding the data access layout, it may refer to the access layout of the data in the target matrix in the shared memory.
[0082] Specifically, in an embodiment of the present invention, when the position of the first "1" in the binary conversion result is the same, the same address hash algorithm can be used. For example, for the binary conversion result b0100 and the binary conversion result b1100, the first "1" in the binary conversion result is located at the third position from the low position to the high position (from right to left) in the binary conversion result. Therefore, for the binary conversion result b0100 and the binary conversion result b1100, the same address hash algorithm can be selected from the predetermined address hash algorithm to obtain the data access layout of the target matrix.
[0083] As described above, in an embodiment of the present invention, the address hash algorithm may refer to performing a hash operation on the storage address corresponding to the data read from the target matrix. For different address hash algorithms, it may refer to performing a hash operation on the storage address corresponding to the data read from the target matrix based on different rules.
[0084] Specifically, according to an embodiment of the present invention, the address hash algorithm to be adopted can be determined by judging whether the position of the first "1" in the binary conversion result meets a predetermined condition.
[0085] In some embodiments, when the position of the first "1" in the binary conversion result satisfies a first predetermined condition, for example, the position of the first "1" in the binary conversion result is any one of the third position, the fourth position or the fifth position from the low position to the high position in the binary conversion result, the address hash algorithm to be adopted may include: performing hash operations on data addresses not exceeding a first number.
[0086] Regarding the first quantity, it may be related to the position of the first "1" in the binary conversion result. According to an embodiment of the present invention, the first quantity may be equal to the number of bits of the first "1" in the binary conversion result from the low bit to the high bit in the binary conversion result minus one. For example, in one example, if the position of the first "1" in the binary conversion result is the third position from the low bit to the high bit in the binary conversion result, the first quantity is equal to 2, that is, when the position of the first "1" in the binary conversion result is the third position from the low bit to the high bit in the binary conversion result, the address hash algorithm to be adopted may include two hash operations performed respectively for two groups of data addresses. In another example, if the position of the first "1" in the binary conversion result is the fifth position from the low bit to the high bit in the binary conversion result, the first quantity is equal to 4, that is, when the position of the first "1" in the binary conversion result is the fifth position from the low bit to the high bit in the binary conversion result, the address hash algorithm to be adopted may include four hash operations performed respectively for four groups of data addresses.
[0087] Specifically, according to an embodiment of the present invention, taking the first "1" in the third position from the low bit to the high bit in the binary conversion result as an example, for embodiments that meet this condition, in these embodiments, the address hash algorithm adopted may be to perform hash operations on the first digit and the second digit in the binary representation of the storage address corresponding to the data read from the target matrix, respectively, to obtain the storage address after hash conversion, so that the data in the target matrix can then be stored in the shared memory based on the storage address after hash conversion. For example, the first digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[0]) and the seventh digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[6]) can be hashed, such as an exclusive OR operation; and the second digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[1]) and the sixth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[5]) can be hashed, such as an exclusive OR operation. From this, we can then get the following Figure 7 The data access layout is shown.
[0088] Specifically, according to another embodiment of the present invention, taking the example that the first "1" mentioned above is in the fifth position from the low bit to the high bit in the binary conversion result, for the embodiments that meet this condition, in these embodiments, the address hash algorithm adopted can be to perform hash operations on the first digit, the second digit, the third digit and the fourth digit in the binary representation of the storage address corresponding to the data read from the target matrix, respectively, to obtain the storage address after hash conversion, so that the data in the target matrix can then be stored in the shared memory based on the storage address after hash conversion. For example, a hash operation such as an exclusive-OR operation may be performed on the first digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[0]) and the ninth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[8]); a hash operation such as an exclusive-OR operation may be performed on the second digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[1]) and the eighth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[7]). Calculation; perform a hash operation such as an exclusive OR operation on the third digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[2]) and the sixth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[5]); and perform a hash operation such as an exclusive OR operation on the fourth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[3]) and the seventh digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[6]). Thus, the following can be obtained. Figure 8 The data access layout is shown.
[0089] In some embodiments, when the position of the first "1" in the binary conversion result satisfies a second predetermined condition, for example, the position of the first "1" in the binary conversion result is the sixth position or any subsequent position from the low position to the high position in the binary conversion result, the address hash algorithm to be adopted may include: performing a hash operation on a second number of data addresses.
[0090] Regarding the second number, according to the inventive concept of the present invention, it can be five. That is, in an embodiment of the present invention, when the position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low position to the high position in the binary conversion result (such as, the seventh position from the low position to the high position, the tenth position from the low position to the high position, etc.), the address hash algorithm to be adopted includes five hash operations performed on five groups of data addresses respectively. Specifically, the address hash algorithm to be adopted can be: hash operations are performed on the first digit to the fifth digit in the binary representation of the storage address corresponding to the data read from the target matrix, respectively, to obtain the storage address after hash conversion, so that the data in the target matrix can then be stored in the shared memory based on the storage address after hash conversion. However, it should be understood that in this case, when the position of the first "1" in the binary conversion result is different, the five groups of data addresses to be hashed are different five groups of data addresses.
[0091] For example, taking the first "1" being at the sixth position from the right to the left in the binary conversion result as an example, for an embodiment that satisfies this condition, the address hash algorithm to be adopted may be: the first digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[0]) and the tenth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[9]) may be subjected to a hash operation such as an exclusive OR operation; the second digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[1]) and the ninth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[8]) may be subjected to a hash operation such as an exclusive OR operation; the binary representation of the storage address corresponding to the data read from the target matrix may be subjected to a hash operation such as an exclusive OR operation. The third digit in the binary representation of the data read from the target matrix (for example, recorded as addr_bin[2]) is hashed with the sixth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[5]) by a hash operation such as an exclusive OR operation; the fourth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[3]) is hashed with the seventh digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[6]) by a hash operation such as an exclusive OR operation; and the fifth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[4]) is hashed with the eighth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[7]) by a hash operation such as an exclusive OR operation.
[0092] For another example, taking the first "1" being at the eighth position from the right to the left in the binary conversion result, for an embodiment that satisfies this condition, the address hash algorithm to be adopted may be: the first digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[0]) and the twelfth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin
[11] ) may be subjected to a hash operation such as an exclusive-OR operation; the second digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[1]) and the eleventh digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin
[10] ) may be subjected to a hash operation such as an exclusive-OR operation; the binary representation of the storage address corresponding to the data read from the target matrix may be subjected to a hash operation such as an exclusive-OR operation. A hash operation such as an exclusive OR operation is performed on the third digit in the binary representation (e.g., recorded as addr_bin[2]) and the eighth digit in the binary representation of the storage address corresponding to the data read from the target matrix (e.g., recorded as addr_bin[7]); a hash operation such as an exclusive OR operation is performed on the fourth digit in the binary representation of the storage address corresponding to the data read from the target matrix (e.g., recorded as addr_bin[3]) and the ninth digit in the binary representation of the storage address corresponding to the data read from the target matrix (e.g., recorded as addr_bin[8]); and a hash operation such as an exclusive OR operation is performed on the fifth digit in the binary representation of the storage address corresponding to the data read from the target matrix (e.g., recorded as addr_bin[4]) and the tenth digit in the binary representation of the storage address corresponding to the data read from the target matrix (e.g., recorded as addr_bin[9]).
[0093] From the above, it can be seen that when the position of the first "1" in the binary conversion result is the sixth position or any of the subsequent positions from right to left in the binary conversion result, that is, the size of each row of data in the target matrix is greater than or equal to 128 bytes, the address hash algorithm to be adopted may include: performing hash operations such as XOR operations on the first to fifth digits in the binary representation of the storage address corresponding to the data read from the target matrix and the five digits starting from the digit corresponding to the position of the aforementioned first "1" in the binary conversion result in the binary representation of the storage address corresponding to the data read from the target matrix, so as to determine the data access layout for the target matrix.
[0094] The following will be combined Figure 7 and Figure 8An example of a data access layout for a target matrix obtained based on the address hash algorithm according to an embodiment of the present invention is described.
[0095] Figure 7 A schematic diagram of a data access layout 700 of a target matrix according to an embodiment of the present invention is shown, wherein each row of data of the target matrix occupies four memory blocks in a shared memory.
[0096] like Figure 7 As shown, each small block corresponds to the size of a Dword data, and the number in each small block represents the index of the memory block used to store the data in the small block, indicating which memory block in the shared memory the data in the small block is stored in. Taking the first row as an example, the data in the first row and first column of the small block can be stored in a file such as Figure 3 In the memory block Bank-0 of the shared memory 250, the data in the first row and second column of the small block can be stored in a file such as Figure 3 In the memory block Bank-1 of the shared memory 250, the data in the first row and third column of the small block can be stored in a file such as Figure 3 In the memory block Bank-2 of the shared memory 250, the data in the first row and fourth column of the small block can be stored in a file such as Figure 3 In the memory block Bank-3 of the shared memory 250. For another example, for a small block with a number of 16 in the small block, the data in the small block can be stored accordingly such as Figure 3 The shared memory 250 is in the memory block Bank-16.
[0097] Based on Figure 7 As shown in the data access layout 700, if the data in the target matrix is stored in the shared memory 250 based on the data access layout 700, for matrices with different data types or their corresponding transposed matrices used for matrix multiplication and accumulation operations, the data in the matrix can be accessed within one cycle.
[0098] For example, if data access to a shared memory is to be implemented with a matrix operation of eight rows and four columns (e.g., a non-transpose operation for matrices of all data types, or a transpose data operation for a 16-bit data matrix), it is obvious that data in any sub-matrix of dimension 8×4 (i.e., eight rows and four columns) in the data access layout 700 can be accessed in 32 different memory blocks respectively, such as Figure 7 That is, no matter how the submatrix with a dimension of 8×4 is selected, the numbers in the small blocks of the submatrix will not be repeated, that is, the 32 groups of data in the submatrix will be necessarily indicated and stored in different 32 memory blocks in the shared memory 250, respectively, without any memory block conflict.
[0099] For example, if a matrix operation of sixteen rows and two columns is used to implement data access to a shared memory (e.g., a transposed data operation for an 8-bit data matrix), similarly, data in any sub-matrix of 16×2 dimensions in the data access layout 700 can also be accessed in correspondence with 32 different memory blocks, such as Figure 7 As shown in sub-matrix 702A or sub-matrix 702B.
[0100] For another example, if the data access of the shared memory is to be implemented by a matrix operation of thirty-two rows and one column (for example, a transposed data operation for a 4-bit data matrix), similarly, the data in any sub-matrix of dimension 32×1 in the data access layout 700 can also be accessed corresponding to 32 different memory blocks respectively, such as Figure 7 As shown in sub-matrix 703 in .
[0101] It can be seen from the above that the data access layout 700 for the target matrix obtained according to the embodiment of the present invention can avoid the occurrence of memory block conflicts in the shared memory when reading the initial matrix or the corresponding transposed matrix for matrix multiplication and accumulation having any data type from the shared memory, thereby significantly improving the shared memory performance.
[0102] Figure 8 A schematic diagram of a data access layout 800 of a target matrix according to yet another embodiment of the present invention is shown, wherein each row of data of the target matrix occupies 16 memory blocks in a shared memory.
[0103] With the above Figure 7 similar, Figure 8 Each small block in the table corresponds to the size of a Dword data, and the number in each small block represents the storage index for the data in the small block, indicating which memory block in the shared memory the data in the small block is stored in. For details, please refer to the above Figure 7 The description is not repeated here.
[0104] Similarly, based on Figure 8 As shown in the data access layout 800, if the data in the target matrix is stored in the shared memory 250 based on the data access layout 800, for matrices with different data types or their corresponding transposed matrices used for matrix multiplication and accumulation operations, the data in the matrix can also be accessed within one cycle.
[0105] like Figure 8As shown, the data in any sub-matrix with a dimension of 8×4 (i.e., eight rows and four columns) in the data access memory layout 800 can be accessed corresponding to 32 different memory blocks, as shown in sub-matrix 801A, sub-matrix 801B, or sub-matrix 801C; similarly, the data in any sub-matrix with a dimension of 16×2 in the data access memory layout 800 can also be accessed corresponding to 32 different memory blocks, as shown in sub-matrix 802A or sub-matrix 802B; and the data in any sub-matrix with a dimension of 32×1 in the data access memory layout 800 can also be accessed corresponding to 32 different memory blocks, as shown in sub-matrix 803A or sub-matrix 803B.
[0106] From the above, it can be seen that based on the data access layout 800, if the data access of the shared memory is implemented by a matrix operation of eight rows and four columns, or the data access of the shared memory is implemented by a matrix operation of sixteen rows and two columns, or the data access of the shared memory is implemented by a matrix operation of thirty-two rows and one column, the occurrence of memory block conflicts in the shared memory can be avoided, thereby significantly improving the shared memory performance.
[0107] In summary, the present invention proposes a solution for data access in a shared memory, by determining a data access layout according to a predetermined address hash algorithm, and storing the data in a shared memory based on the determined data access layout, so that when an initial matrix or a corresponding transposed matrix for matrix multiplication and accumulation having any data type is read from the shared memory, memory block conflicts in the shared memory can be avoided, thereby significantly improving the shared memory performance.
[0108] According to another concept of the present invention, based on the above-mentioned solution for accessing data in a shared memory, a stride parameter may be further introduced so as to determine the data access layout of a target matrix in combination with the stride parameter.
[0109] Regarding the stride parameter, it can be recorded as "stride", which indicates the stride of reading the input activation when performing a convolution operation in a convolutional neural network. Corresponding to reading the matrix data in the shared memory, the stride parameter can indicate the stride in the number of rows when reading the matrix data. For example, taking the stride parameter as 1 (i.e., stride=1) as an example, it means that when performing a convolution operation in a convolutional neural network, the stride in the input channel dimension when reading the input activation is 1, that is, the stride in the number of rows when accessing the matrix in the shared memory is 1. In other words, "stride=1" can mean that when reading the initial matrix or the corresponding transposed matrix for matrix multiplication and accumulation from the shared memory based on the determined data access layout, the difference in the number of row indices of two consecutive rows of data read is 1. For example, the above Figure 7 and Figure 8 The data access layouts of the target matrix shown are all suitable for data access with a stride parameter of 1.
[0110] However, when the stride parameter is greater than 1, since the data will be read across rows, if the data in the target matrix is stored in the shared memory based only on the data access layout of the target matrix determined in the above manner, a memory block conflict will occur when the data is read from the shared memory.
[0111] Fig. 9 A partial schematic diagram of the data access layout of a target matrix according to another embodiment of the present invention is shown, wherein each row of data of the target matrix occupies 32 memory blocks in the shared memory, and the stride parameter is 1.
[0112] With the above Figure 7 and / or Figure 8 similar, Fig. 9 Each small block in the table corresponds to the size of a Dword data, and the number in each small block represents the index of the memory block used to store the data in the small block, indicating which memory block in the shared memory the data in the small block is stored in. For details, please refer to the above about Figure 7 or Figure 8 The description is not repeated here.
[0113] Based on Fig. 9 In the data access layout 900 shown in FIG. 1 , if the data access of the shared memory is implemented by a matrix operation of eight rows and four columns, when the stride parameter is 1, the data in any sub-matrix of dimension 8×4 (i.e., eight rows and four columns) in the data access layout 900 can be accessed respectively corresponding to 32 different memory blocks, such as Fig. 9 As shown in the sub-matrix 901 in FIG. Therefore, the 32 groups of data in the sub-matrix 901 will be necessarily indicated and stored in different 32 memory blocks in the shared memory 250 respectively, and no memory block conflict will occur.
[0114] However, when the stride parameter is greater than 1 (eg, the stride parameter is 2), if a sub-matrix with the same dimension of 8×4 is read from the shared memory based on the current data access layout 900 (eg, Fig. 9 Obviously, the 32 groups of data in the sub-matrix 901 include a plurality of data repeatedly indicated to the same memory block, so a memory block conflict will occur during the data reading process.
[0115] Therefore, in order to at least partially resolve the memory block conflict caused by reading data when the stride parameter is greater than 1, as described above, the method for data access in shared memory provided by the present invention may further include: determining the data layout in the shared memory based on the stride parameter and the data size of each row of the determined target matrix, and then storing the data in the target matrix in the shared memory based on the determined data layout. Specifically, according to an embodiment of the present invention, it is possible to: determine the value of the stride parameter for data access (denoted as stride); select the address hash algorithm to be adopted from the predetermined address hash algorithm based on the product (denoted as stride*row_size) of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory; and determine the data access layout for the target matrix based on the address hash algorithm to be adopted.
[0116] As mentioned above about Figure 7 , Figure 8 , and / or Fig. 9 In the example embodiment, since the value of the stride parameter stride=1, stride*row_size=1*row_size=row_size, the address hash algorithm to be adopted can be directly selected based on the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory, that is, by binary-converting the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory as described above, and selecting the address hash algorithm to be adopted according to the position of the first "1" in the result after binary conversion. For details, please refer to the description above and will not be repeated here.
[0117] However, when the value of the stride parameter is greater than 1, according to an embodiment of the present invention, similarly, a binary conversion can be performed on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory, and the address hash algorithm to be adopted is selected based on the position of the first "1" in the result after binary conversion.
[0118] The address hash algorithm to be used is selected based on the position of the first "1" in the result after binary conversion, which is similar to the above Figure 6 The contents in the description are similar and will not be repeated here.
[0119] The following will be combined Figure 10 to Figure 11 Detailed description.
[0120] Fig.10A schematic diagram of a data access layout 1000 of a target matrix according to an embodiment of the present invention is shown, wherein each row of data of the target matrix occupies 32 memory blocks in the shared memory (row_size=32) and the stride parameter is 2 (stride=2).
[0121] like Fig.10 In the illustrated embodiment, since row_size=32 and stride=2, the value of the stride parameter and the product of the number of memory blocks occupied by each row of data of the target matrix in the shared memory are equal to row_size*stride=32*2=64. On this basis, binary conversion is performed on the value of the stride parameter and the product of the number of memory blocks occupied by each row of data of the target matrix in the shared memory (row_size*stride=64), and the result after binary conversion can be b1000000, wherein the position of the first "1" is the seventh position from the low position to the high position in the binary conversion result.
[0122] In this embodiment, since the position of the first "1" is the seventh position from the low position to the high position in the binary conversion result, it can be known that the position of the first "1" in the binary conversion result satisfies the second predetermined condition, that is, the position of the first "1" in the binary conversion result is the sixth position or any subsequent position from the low position to the high position in the binary conversion result. Therefore, the address hash algorithm to be adopted includes: performing hash operations on the second number of data addresses.
[0123] For example, a hash operation such as an exclusive-OR operation may be performed on the first digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[0]) and the eleventh digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin
[10] ); a hash operation such as an exclusive-OR operation may be performed on the second digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[1]) and the tenth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[9]); a hash operation such as an exclusive-OR operation may be performed on the third digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[2]) and the tenth digit in the binary representation of the storage address corresponding to the data read from the target matrix (for example, recorded as addr_bin[3]). A hash operation such as an exclusive OR operation is performed on the seventh digit (for example, recorded as addr_bin[6]) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation is performed on the fourth digit (for example, recorded as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eighth digit (for example, recorded as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix; and a hash operation such as an exclusive OR operation is performed on the fifth digit (for example, recorded as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix and the ninth digit (for example, recorded as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix.
[0124] and Fig. 9 Compared with the target matrix data access layout 900, due to Fig.10 The value of the stride parameter corresponding to the embodiment shown is 2, that is, the stride in the number of rows when accessing data of the matrix in the shared memory is 2. Therefore, when reading the initial matrix or the corresponding transposed matrix used for matrix multiplication and accumulation from the shared memory, the difference in the number of row indices of two consecutive rows of data read is 2. Therefore, according to the inventive concept of the present invention, in order to avoid memory block conflicts caused by reading data across rows, the selected hash algorithm to be adopted can enable every two consecutive rows of data to use the same address hash result when storing the data in the target matrix in the shared memory, that is, the address hash results of every two rows are the same. For example, the data of rows 0 and 1 use the same address hash result, the data of rows 2 and 3 use the same address hash result, the data of rows 4 and 5 use the same address hash result, and so on. Thus, it can be obtained as follows Fig.10 The data access layout 1000 of the target matrix is shown in FIG.
[0125] From the above, we can see that based on Fig.10 The data access layout 1000 shown in FIG. 1 stores the data in the target matrix into the shared memory and then reads the matrix with different data types or its corresponding transposed matrix for the matrix multiplication and accumulation operation from the shared memory, such as Fig.10 The sub-matrix 1001 with a dimension of 15×4 and a stride parameter of 2, or the sub-matrix 1002 with a dimension of 31×2 and a stride parameter of 2 shown in the figure can access the data in the matrix within one cycle, thereby avoiding memory block conflicts.
[0126] Fig.11 A schematic diagram of a data access layout 1100 of a target matrix according to yet another embodiment of the present invention is shown, wherein each row of data of the target matrix occupies 32 memory blocks in the shared memory (row_size=32) and the stride parameter is 4 (stride=4).
[0127] like Fig.11 In the embodiment shown, since row_size=32 and stride=4, the value of the stride parameter and the product of the number of memory blocks occupied by each row of data of the target matrix in the shared memory are equal to row_size*stride=32*4=128. On this basis, binary conversion is performed on the value of the stride parameter and the product of the number of memory blocks occupied by each row of data of the target matrix in the shared memory (row_size*stride=64), and the result after binary conversion is b10000000, where the position of the first "1" is the eighth position from the low position to the high position in the binary conversion result. Fig.10 Similar to the embodiment shown, the position of the first "1" in the binary conversion result of this embodiment also satisfies the second predetermined condition, so the address hash algorithm to be adopted includes: performing a hash operation on a second number of data addresses.
[0128] For example, a hash operation such as an exclusive-OR operation may be performed on the first digit (for example, recorded as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix and the twelfth digit (for example, recorded as addr_bin
[11] ) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive-OR operation may be performed on the second digit (for example, recorded as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eleventh digit (for example, recorded as addr_bin
[10] ) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive-OR operation may be performed on the third digit (for example, recorded as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eleventh digit (for example, recorded as addr_bin
[11] ) in the binary representation of the storage address corresponding to the data read from the target matrix; and a hash operation such as an exclusive-OR operation may be performed on the third digit (for example, recorded as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eleventh digit (for example, recorded as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix. The eighth digit (for example, recorded as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix is subjected to a hash operation such as an exclusive OR operation; the fourth digit (for example, recorded as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix is subjected to a hash operation such as an exclusive OR operation with the ninth digit (for example, recorded as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix; and the fifth digit (for example, recorded as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix is subjected to a hash operation such as an exclusive OR operation with the tenth digit (for example, recorded as addr_bin[9]) in the binary representation of the storage address corresponding to the data read from the target matrix.
[0129] Furthermore, due to Fig.11 The value of the stride parameter corresponding to the embodiment shown is 4, so when reading the initial matrix or the corresponding transposed matrix for matrix multiplication and accumulation from the shared memory, the difference in the number of row indices of two consecutive rows of data read is 4. Similarly, in order to avoid memory block conflicts caused by reading data across rows, in this case, the selected hash algorithm to be adopted enables every four consecutive rows of data to use the same address hash result when storing the data in the target matrix in the shared memory, that is, the address hash results of every four rows are the same. For example, the data in rows 0, 1, 2, and 3 use the same address hash result, the data in rows 4, 5, 6, and 7 use the same address hash result, the data in rows 8, 9, 10, and 11 use the same address hash result, and so on. Thus, the following can be obtained: Fig.11 The data access layout 1100 of the target matrix is shown in FIG.
[0130] Therefore, based on Fig.11The data access layout 1100 shown in FIG. 1 stores the data in the target matrix into the shared memory and then reads the matrix with different data types or its corresponding transposed matrix for the matrix multiplication and accumulation operation from the shared memory, such as Fig.11 The sub-matrix 1101A or the sub-matrix 1101B shown in FIG. 1 , with a dimension of 29×4 and a stride parameter of 4, can implement memory access of the data in the matrix within one cycle, thereby avoiding memory block conflicts.
[0131] In summary, the solution for data access in shared memory provided by the present invention can also take into account the stride parameter on the basis of determining the data access layout of the target matrix, so as to store the data in the target matrix in the shared memory based on the expanded data access layout, so that memory block conflicts will not occur when cross-row data is read in the shared memory, thereby improving the shared memory performance.
[0132] In addition, the various processes and processing described herein, such as method 600, may be performed at a computing device. The computing device includes, for example: at least one processor (at least one graphics processor and at least one central processing unit); and a memory connected to at least one processor in communication; wherein the memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor. In some embodiments, method 600 may be implemented as a computer software program or program product, which is tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed on a computing device via a read-only memory (ROM) and / or a communication unit. When the computer program is loaded into a random access memory (RAM) and executed by an AI processor such as a GPU, GPGPU, or NPU, one or more actions of method 600 described above may be performed.
[0133] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for executing various aspects of the present invention. The computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof.
[0134] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. Various aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each box of the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.
[0135] These computer-readable program instructions can be provided to a central processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the central processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0136] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the function involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0137] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0138] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0139] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for accessing data in a shared memory, characterized in that: include: Determine the number of memory blocks occupied by each row of data of the target matrix in the shared memory; At least based on the number of memory blocks occupied by each row of data of the target matrix in the shared memory, a data access layout for the target matrix is determined according to a predetermined address hash algorithm, so that the data in the target matrix is stored in the shared memory based on the data access layout.
2. The method according to claim 1, characterized in that The shared memory includes a plurality of memory blocks having the same bandwidth; as well as Determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory includes: Determine the total number of bytes of each row of data of the target matrix; Determining the bandwidth size of the memory block; and The number of memory blocks occupied by each row of data of the target matrix in the shared memory is determined based on the total number of bytes of each row of data of the target matrix and the bandwidth size of the memory block.
3. The method according to claim 1, characterized in that Determining a data access layout for the target matrix includes: Performing binary conversion on the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory; Selecting an address hash algorithm to be adopted from the predetermined address hash algorithms based on the position of the first "1" in the binary conversion result; and A data access layout for the target matrix is determined based at least on the address hash algorithm to be adopted.
4. The method according to claim 3, characterized in that Selecting an address hash algorithm to be adopted from the predetermined address hash algorithms includes: When the position of the first "1" in the binary conversion result satisfies a first predetermined condition, the address hash algorithm to be adopted includes performing a hash operation on data addresses not exceeding a first number; When the position of the first "1" in the binary conversion result meets the second predetermined condition, the address hash algorithm to be adopted includes performing a hash operation on a second number of data addresses.
5. The method according to claim 4, characterized in that The first predetermined condition is: The position of the first "1" in the binary conversion result is any one of the third position, the fourth position or the fifth position from the low position to the high position in the binary conversion result.
6. The method according to claim 4, characterized in that The second predetermined condition is: The position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low position to the high position in the binary conversion result.
7. The method according to claim 4, characterized in that The first number is related to the position of the first "1" in the binary conversion result, and the second number is five.
8. The method according to claim 7, characterized in that The first number is equal to the number of bits of the first "1" in the binary conversion result from the low bit to the high bit in the binary conversion result minus one.
9. The method according to claim 1, characterized in that: At least based on the number of memory blocks occupied by each row of data of the target matrix in the shared memory, and according to a predetermined address hash algorithm, determining a data access layout for the target matrix includes: Determine the value of the stride parameter used for data fetching; Selecting an address hash algorithm to be adopted from the predetermined address hash algorithms based on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory; and Based on the address hash algorithm to be adopted, a data access layout for the target matrix is determined.
10. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that The invention comprises a computer program which, when executed by a machine, performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Memory management method and device
CN103077126A
Data processing method and device and electronic equipment
CN112069460A
Data storage method, data reading method, electronic equipment and storage medium
CN117785759A
Matrix computation method, computation device, and processor
WO2021036729A1
Multi-core system-based task scheduling method and apparatus, and related product
WO2024198863A1
Cited By
Memory access method and device, electronic equipment and storage medium
CN120561033A
Artificial intelligence chip, data multicast method and device, equipment and storage medium
CN121166611A