A method, an electronic device, a storage medium, and a computer program product for accessing data in shared memory

By determining the data access layout of the target matrix in shared memory, the address hashing algorithm is used to solve the memory block conflict problem, improving the data access efficiency and performance of shared memory.

CN119938362BActive Publication Date: 2025-07-29SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510414829.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-29
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

Existing shared memory data access schemes are prone to memory block conflicts, resulting in memory access delay and low data read and write efficiency, especially when processing matrix multiplication and accumulation tasks, affecting shared memory performance.

Method used

By determining the number of memory blocks occupied by each row of data in the shared memory of the target matrix, a predetermined address hashing algorithm is used to determine the data memory access layout to avoid memory block conflicts and improve data storage efficiency.

Benefits of technology

It significantly improves the performance of shared memory, avoids memory block conflicts, and improves data reading and writing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938362B_ABST
    Figure CN119938362B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to a method, an electronic device, a storage medium, and a computer program product for accessing data in a shared memory. The method includes: determining the number of memory blocks occupied by each row of data of a target matrix in the shared memory; and determining a data access layout for the target matrix based at least on the number of memory blocks occupied by each row of data of the target matrix in the shared memory according to a predetermined address hashing algorithm, so as to store the data in the target matrix into the shared memory based on the data access layout. The method for accessing data in the shared memory provided by the present invention can avoid the occurrence of memory block conflicts in the shared memory when accessing data in the shared memory based on a matrix with any data type or its corresponding transposed matrix, and significantly improve the performance of the shared memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of data processing, and more particularly to a method, an electronic device, a storage medium, and a computer program product for accessing data in shared memory. Background Art

[0002] Artificial intelligence (AI) processors, due to their ability to execute multiple threads simultaneously and process vector and tensor data, enable efficient parallel computing and are thus increasingly widely used to process large-scale data and intensive computing tasks.

[0003] However, when using an AI processor (such as a graphics processing unit (GPU)) to execute a computing task such as matrix multiply-accumulate (MMA), existing schemes for accessing data in shared memory usually only perform data access in the shared memory based on the data layout in the initial matrix for matrix multiply-accumulate. In this case, when accessing data related to the transposed matrix corresponding to the initial matrix in the shared memory, memory bank conflicts will inevitably occur, increasing the memory access latency, reducing the data read / write efficiency, and thus seriously affecting the performance of the shared memory.

[0004] In summary, the deficiencies of existing schemes for accessing data in shared memory are: memory bank conflicts are likely to occur, and the performance of the shared memory cannot be fully utilized. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method, an electronic device, a storage medium, and a computer program product for accessing data in shared memory, which can avoid the occurrence of memory bank conflicts in the shared memory and significantly improve the performance of the shared memory when accessing data in the shared memory based on a matrix with any data type or its corresponding transposed matrix.

[0006] According to a first aspect of the present invention, there is provided a method for accessing data in shared memory, including: determining the number of memory blocks occupied by each row of data of a target matrix in the shared memory; and determining a data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data of the target matrix in the shared memory according to a predetermined address hashing algorithm, so as to store the data in the target matrix into the shared memory based on the data access layout.

[0007] In some embodiments, the shared memory includes multiple memory blocks with the same bandwidth. In these embodiments, determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory includes: determining the total number of bytes of each row of data in the target matrix; determining the bandwidth size of the memory block; and determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory based on the total number of bytes of each row of data in the target matrix and the bandwidth size of the memory block.

[0008] In some embodiments, determining the data access layout for the target matrix includes: performing binary conversion on the value of the number of memory blocks occupied by each row of data in the target matrix in the shared memory; selecting an address hashing algorithm to be used from a predetermined address hashing algorithm based on the position of the first "1" from the low bit to the high bit in the binary conversion result; and determining the data access layout for the target matrix based at least on the address hashing algorithm to be used.

[0009] In some embodiments, selecting an address hashing algorithm to be used from a predetermined address hashing algorithm includes: when the position of the first "1" in the binary conversion result satisfies a first predetermined condition, the address hashing algorithm to be used includes performing a hashing operation on no more than a first number of data addresses; when the position of the first "1" in the binary conversion result satisfies a second predetermined condition, the address hashing algorithm to be used includes performing a hashing operation on a second number of data addresses.

[0010] In some embodiments, the first predetermined condition is that the position of the first "1" in the binary conversion result is any one of the third, fourth, or fifth positions from the low bit to the high bit in the binary conversion result.

[0011] In some embodiments, the second predetermined condition is that the position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low bit to the high bit in the binary conversion result.

[0012] In some embodiments, the first number is related to the position of the first "1" in the binary conversion result, and the second number is five.

[0013] In some embodiments, the first number is equal to the number of bits of the first "1" in the binary conversion result from the low bit to the high bit minus one.

[0014] In some embodiments, based at least on the number of memory blocks occupied by each row of data of the target matrix in the shared memory, according to a predetermined address hashing algorithm, determining the data access layout for the target matrix includes: determining the value of the stride parameter for data access; selecting, based on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory, an address hashing algorithm to be adopted from the predetermined address hashing algorithm; and determining the data access layout for the target matrix based on the address hashing algorithm to be adopted.

[0015] According to a second aspect of the present invention, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to the first aspect of the present invention.

[0016] According to a third aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to the first aspect of the present invention.

[0017] According to a fourth aspect of the present invention, there is provided a computer program product, including a computer program, the computer program being for the method according to the first aspect of the present invention when executed by a machine.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more obvious. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0020] Figure 1 An exemplary architecture diagram of a processor according to an embodiment of the present invention is shown.

[0021] Figure 2 An exemplary architecture diagram of a stream processor cluster according to an embodiment of the present invention is shown.

[0022] Figure 3 An exemplary architecture diagram of a shared memory according to an embodiment of the present invention is shown.

[0023] Figure 4A An exemplary schematic diagram of the access layout of a matrix with a data type of 32-bit data in the shared memory is shown.

[0024] Figure 4BAn exemplary schematic diagram showing the memory access layout of a matrix with a data type of 32-bit data in shared memory.

[0025] Figure 4C An exemplary schematic diagram showing the memory access layout of a matrix with a data type of 32-bit data in shared memory.

[0026] Figure 5 A schematic framework diagram of a controller for controlling data memory access of shared memory according to an embodiment of the present invention is shown.

[0027] Figure 6 A flowchart of a method for data memory access of shared memory according to an embodiment of the present invention is shown.

[0028] Figure 7 A schematic diagram showing the memory access layout of a target matrix according to an embodiment of the present invention is shown.

[0029] Figure 8 A schematic diagram showing the memory access layout of a target matrix according to another embodiment of the present invention is shown.

[0030] Figure 9 A schematic diagram showing a part of the memory access layout of a target matrix according to another embodiment of the present invention is shown.

[0031] Figure 10 A schematic diagram showing the memory access layout of a target matrix according to an embodiment of the present invention is shown.

[0032] Figure 11 A schematic diagram showing the memory access layout of a target matrix according to another embodiment of the present invention is shown. Detailed implementation manners

[0033] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0034] As used herein, the term "including" and its variants mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.

[0035] Artificial Intelligence (AI) processors, such as Graphics Processing Unit (GPU), General-purpose Graphics Processing Unit (GPGPU), Tensor Processing Unit (TPU), Neural Network Processing Unit (NPU), Deep Learning Processing Unit (DPU), Accelerated Processing Unit (APU), etc., due to their characteristics of being able to execute multiple threads and a large amount of data simultaneously, enable efficient parallel computing, and thus are increasingly widely used to process tasks such as large-scale data and intensive computing tasks.

[0036] Figure 1 An exemplary architecture diagram of the processor 100 according to an embodiment of the present invention is shown. It should be understood that the processor 100 may further include additional units and / or modules not shown, and the scope of the present invention is not limited in this regard.

[0037] Regarding the processor 100, it may be any suitable type of artificial intelligence processor such as a Graphics Processor (GPU), General Graphics Processor (GPGPU), etc.

[0038] Taking a Graphics Processing Unit (GPU) as an example, as Figure 1 The illustrated processor 100 may include multiple Streaming Processing Clusters (SPC) 110 and a global memory 120, where the global memory 120 may be configured to be accessed by the multiple streaming processing clusters 110 to achieve data reading and writing. In some other cases, the processor 100 includes multiple global memories 120, such that each of the multiple streaming processing clusters 110 corresponds to one of the multiple global memories 120 to achieve data reading and writing.

[0039] Regarding the global memory 120, it may be a High Bandwidth Memory (HBM) or a Double Data Rate (DDR) synchronous dynamic random access memory, or a cache Cache that works in cooperation with the high bandwidth memory and the synchronous dynamic random access memory.

[0040] Regarding multiple stream processor clusters 110, in some embodiments, the processor 100 may include, but is not limited to, such as 4, 8, 12, or 16 stream processor clusters.

[0041] In an embodiment of the present invention, multiple stream processor clusters 110 configured in the processor 100 may share a secondary cache (L2 Cache) (not shown in Figure 1 here) and the global memory 120 through a crossbar of the processor 100.

[0042] Furthermore, for each stream processor cluster, multiple computation units (CUs) may be configured therein, and each computation unit may include a vector core (VC) unit and a tensor core (TC) unit for computing implementation. Specifically, Figure 2 is shown Figure 1 an exemplary architecture diagram of the stream processor cluster 110 in

[0043] As Figure 2 shown, the stream processor cluster 110 may include multiple computation units (CUs) 210, where each computation unit 210 includes multiple vector core (VC) units 220 and one tensor core (TC) unit 230. Furthermore, these multiple vector core (VC) units 220 and one tensor core (TC) unit 230 may achieve data sharing through a vector register 240 and a shared memory 250. Specifically, in combination with Figure 2 the stream processor cluster 110 shown, the shared memory 250 may be configured between the vector register 240 and the global memory 120. When data in the global memory 120 is loaded into the vector register 240, or when data in the vector register 240 is stored into the global memory 120, these data will first be sent to the shared memory 250, and operations such as transpose will be performed on these data when reading and writing the shared memory 250.

[0044] For example, to load data from the global memory 120 into the vector register 240, the data in the global memory 120 (e.g., high-bandwidth memory HBM) needs to be first transferred to the shared memory 250, and then transferred from the shared memory 250 to the vector register 240, so that the computing cores (e.g., vector core unit 220 or tensor core unit 230) in the computing unit 210 can perform calculations based on the data in the vector register 240. Another example is that when the computing unit 210 finishes the calculation and obtains the corresponding calculation result, if the calculation result is to be stored from the vector register 240 to the global memory 120, the calculation result needs to be first transferred from the vector register 240 to the shared memory 250, and then transferred from the shared memory 250 to the global memory 120 to achieve the storage of the calculation result in the global memory 120. Thus, it can be seen that the data access (i.e., storage and reading) efficiency of the shared memory 250 determines the data processing efficiency of the processor as a system.

[0045] To improve the data access efficiency of the shared memory 250, the shared memory 250 can be divided into multiple memory blocks (Banks) of the same size, and these memory blocks can be accessed simultaneously, significantly increasing the memory access bandwidth of the shared memory 250. The following will combine Figure 3 to describe the architecture of the shared memory 250.

[0046] Figure 3 shows an exemplary architecture diagram of the shared memory 250 according to an embodiment of the present invention.

[0047] As Figure 3 shown, the shared memory 250 is divided into 32 memory blocks (i.e., Bank-0 to Bank-31), and these 32 memory blocks can be accessed simultaneously. For example, for a processor with a single instruction multiple thread (SIMT)-32 architecture, each warp in the processor includes 32 threads, and these 32 threads can respectively access the 32 memory blocks of the shared memory 250, that is, each thread can access a separate memory block and correspondingly process the data in the memory block it accesses.

[0048] As Figure 3 shown, the bandwidth of each memory block of the shared memory 250 can be 32 bits. Taking data of the data type double word (Dword) as an example, since a data of the data type double word occupies 32 bits (i.e., 32 bits), each address in each memory block can be used to store a data of the data type double word. If each memory block has, for example, N (where N is an integer greater than or equal to 1) different addresses for storing data, then correspondingly, each memory block can store at most N data of the data type double word. Therefore, for Figure 3The shared memory 250 shown can have 32*N data items of double-word data type.

[0049] However, when accessing the data in the shared memory 250, for the same memory bank, only one data item of double-word data type can be read from it or one data item of double-word data type can be written into it within the same clock cycle. Therefore, when multiple different requests expect to access the data at different addresses in the same memory bank simultaneously within the same clock cycle, such as performing read operations simultaneously or write operations simultaneously, it will lead to a bank conflict. In this case, the bank conflict will increase the latency of memory access and reduce the data read / write efficiency, thus seriously affecting the performance of the shared memory 250.

[0050] In particular, when using an AI processor to execute a computing task, for example, when performing operations such as matrix multiplication accumulation (MMA) in a GPU, it is usually necessary to read and write data in the shared memory based on both a certain initial matrix for matrix multiplication accumulation and its corresponding transposed matrix. In this case, if the data is written into the shared memory only based on the data layout in the initial matrix for matrix multiplication accumulation, a bank conflict will necessarily occur when reading the data of the transposed matrix corresponding to the initial matrix from the shared memory.

[0051] Taking a computing task for 32-bit data as an example. Figures 4A to 4C A schematic diagram showing an exemplary memory access layout of a matrix of 32-bit data type in the shared memory 250 is shown.

[0052] As mentioned above, since the bandwidth of each memory bank in the shared memory 250 is 32 bits, for 32-bit data, each data item can be stored in one entry of the memory bank in the shared memory 250. If a basic matrix with a dimension of 8×4 (i.e., eight rows and four columns) is used as the initial tiling matrix, for a matrix 400 with a dimension of 16×8 (i.e., sixteen rows and eight columns) and a data type of 32-bit data, it can be divided into four basic matrices with a dimension of 8×4, as Figure 4A shown.

[0053] Furthermore, since each basic matrix includes 32 data items and the shared memory 250 includes 32 memory banks, the 32 data items in each basic matrix can be separately stored in different memory banks in the shared memory 250, and each data item occupies one entry of the memory bank, as Figure 4B shown.

[0054] However, when it is necessary to perform a transpose operation on the data read from the shared memory to obtain a matrix with dimensions of 8×4, the original data is stored in the shared memory as a basic matrix with dimensions of 4×8, as Figure 4C shown. Figure 4C In, four basic matrices with dimensions of 4×8 form a matrix 400' with dimensions of 16×8. Referring to Figure 4C , the matrix 400' before transposition includes four basic matrices with dimensions of 4×8 (i.e., four rows and eight columns), namely the first basic matrix 401, the second basic matrix 402, the third basic matrix 403, and the fourth basic matrix 404.

[0055] Furthermore, from Figure 4C , it can be known that if it is necessary to access the data in the shared memory 250 based on the first basic matrix 401, since the data corresponding to the first basic matrix 401 is data 0 to data 15 and data 32 to data 47, however, referring to Figure 4B , it can be known that, for example, data 0 and data 32 are stored in different entries of the same memory block in the shared memory 250, that is, both data 0 and data 32 are stored in Bank-0. As described above, only the data in one address in the same memory block can be accessed within the same clock cycle. Therefore, if it is necessary to access data 0 and data 32 stored in Bank-0, at least two cycles are required to achieve the reading of data 0 and data 32. Similarly, data 1 and data 33 in the first basic matrix 401 are stored in the same memory block in the shared memory 250, data 2 and data 34 are stored in the same memory block in the shared memory 250, and so on. It can be seen from this that if it is necessary to implement the access of data related to the first basic matrix 401 in the shared memory, at least two cycles are required to complete.

[0056] Similarly, for Figure 4C the second basic matrix 402, the third basic matrix 403, and the fourth basic matrix 404 in, it also takes at least two cycles to complete the respective corresponding data access. In other words, when performing data access to the shared memory for the transposed matrix, due to memory block conflicts, the processor only uses part of the memory read / write bandwidth within one cycle, thereby reducing the data access efficiency and increasing the latency of data access in the shared memory.

[0057] As can be seen from the above, the existing data access scheme for the shared memory will inevitably cause memory block conflicts when performing data access to the matrix that needs to be transposed, resulting in an increase in the latency of memory access and a decrease in the data read / write efficiency, thereby seriously affecting the performance of the shared memory and further reducing the performance of the AI processor.

[0058] To at least partially address one or more of the above and other potential problems, example embodiments of the present invention propose a scheme for accessing data in shared memory. By determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory; at least based on the number of memory blocks occupied by each row of data in the target matrix in the shared memory, according to a predetermined address hashing algorithm, determine the data access layout for the target matrix, so as to store the data in the target matrix into the shared memory based on the data access layout, such that when reading from the shared memory an initial matrix or a corresponding transposed matrix for matrix multiply-accumulate with any data type, the occurrence of memory block conflicts in the shared memory can be avoided, significantly improving the performance of the shared memory.

[0059] The following will be combined with Figures 5 to 8 Describe in detail the scheme for accessing data in shared memory according to the present invention.

[0060] Figure 5 FIG. shows a schematic framework diagram of a controller 500 for controlling data access to shared memory according to an embodiment of the present invention. It should be understood that the controller 500 may further include additional units and / or modules not shown, and the scope of the present invention is not limited in this regard.

[0061] Regarding the controller 500, it can be configured to receive an access request for a shared memory (such as Figure 3 shared memory 250). According to an embodiment of the present invention, the access request may be a write request, and the write request may include information such as the data to be written to the shared memory, the address of the written data, etc. According to an embodiment of the present invention, the access request may also be a read request.

[0062] As Figure 5 shown, the controller 500 may include a hashing operation unit 510 to perform a hashing operation based on the access request received by the controller 500. According to an embodiment of the present invention, for example, the hashing operation unit 510 may determine, based on the access request received by the controller 500, the number of memory blocks occupied by each row of data in the target matrix related to the data to be written to the shared memory; and at least based on the number of memory blocks occupied by each row of data in the target matrix in the shared memory, perform a hashing operation on the address of the written data in the access request according to a predetermined address hashing algorithm, so as to determine the data access layout for the current target matrix, such that the controller 500 can then store the data in the target matrix into the shared memory 250 based on the determined data access layout.

[0063] The following will be combined with Figure 6 Describe in detail the scheme for accessing data in shared memory according to an embodiment of the present invention.

[0064] Figure 6 FIG. 600 is a flowchart of a method for accessing data in a shared memory according to an embodiment of the present invention. It should be understood that method 600 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.

[0065] In step 601, the controller 500 determines the number of memory blocks occupied by each row of data of the target matrix in the shared memory.

[0066] Regarding the target matrix, the data therein can be stored in the shared memory so that subsequently, data corresponding to the matrix for matrix multiply-accumulate operations can be read from the shared memory.

[0067] According to an embodiment of the present invention, the target matrix may have different matrix sizes. Specifically, the size of each row of data in the target matrix can be determined according to the first size of the target matrix in the column dimension. For example, the first size of the target matrix in the column dimension can be denoted as K. In some embodiments, K can be 16 bytes, that is, the size of each row of data in the target matrix is 16 bytes. In still other embodiments, K can be 32 bytes, that is, the size of each row of data in the target matrix is 32 bytes. In other embodiments, K can also be 48 bytes, 64 bytes, or other multiples of 4 bytes, etc., and the present invention is not limited thereto.

[0068] Further, according to the inventive concept of the present invention, the number of memory blocks occupied by each row of data of the target matrix in the shared memory can be determined at least according to the size of each row of data in the target matrix. Specifically, according to an embodiment of the present invention, determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory may include: determining the total number of bytes of each row of data of the target matrix; determining the bandwidth size of the memory block; and based on the total number of bytes of each row of data of the target matrix and the bandwidth size of the memory block, determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory.

[0069] Regarding the bandwidth size of the memory block, taking the shared memory 250 shown above Figure 3 as an example, the bandwidth size of each memory block in the shared memory 250 packet is 32 bits, that is, each row address in a memory block can be used for accessing 32-bit data, that is, for accessing a data of the data type double word (Dword).

[0070] According to an embodiment of the present invention, the number of memory blocks occupied by each row of data in the target matrix in the shared memory can then be calculated by dividing the size of each row of data in the target matrix by the bandwidth size of the memory block. In some embodiments, the number of memory blocks occupied by each row of data in the target matrix in the shared memory can be determined by dividing the total number of bytes of each row of data in the target matrix by the bandwidth size of the memory block. In still other embodiments, based on the total number of bytes of each row of data in the target matrix, the number of Dwords equivalent to each row of data in the target matrix can be determined first, and then based on the number of Dwords that can be accessed by each row address in the memory block, the number of memory blocks occupied by each row of data in the target matrix in the shared memory can be determined.

[0071] For example, taking the example where the size of each row of data in the above target matrix is 16 bytes, since the total number of bytes of each row of data in the target matrix is 16 bytes (i.e., 128 bits, equivalent to 4 Dwords), for a memory block with a bandwidth size of 32 bits in the shared memory 250, each row of data in the target matrix can occupy 4 memory blocks in the shared memory 250. That is to say, if the size of each row of data in the target matrix is 16 bytes, the number of memory blocks occupied by each row of data in the shared memory (such as the shared memory 250) is 4. Similarly, for the case where the total number of bytes of each row of data in the target matrix is 32 bytes (i.e., 256 bits, equivalent to 8 Dwords), the number of memory blocks occupied by each row of data in the shared memory will be 8; for the case where the total number of bytes of each row of data in the target matrix is 48 bytes (i.e., 384 bits, equivalent to 12 Dwords), the number of memory blocks occupied by each row of data in the shared memory will be 12; for the case where the total number of bytes of each row of data in the target matrix is 64 bytes (i.e., 512 bits, equivalent to 16 Dwords), the number of memory blocks occupied by each row of data in the shared memory will be 16; and so on.

[0072] In step 602, the controller 500 determines the data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data in the target matrix in the shared memory according to a predetermined address hashing algorithm, so as to store the data in the target matrix into the shared memory based on the data access layout.

[0073] Regarding the predetermined address hashing algorithm, it can be a hashing operation performed in advance on the storage address corresponding to the data read from the target matrix, so that the data can be stored into the corresponding memory block in the shared memory based on the hashed storage address. For example, in an embodiment of the present invention, the hashing operation may include an exclusive OR operation on the storage address, which will be described in detail below and will not be elaborated here for the time being.

[0074] According to the inventive concept of the present invention, different address hashing algorithms can be adopted according to the different numbers of memory blocks occupied by each row of data of the target matrix in the shared memory. For example, in some embodiments of the present invention, binary conversion can be performed on the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory, so as to determine the address hashing algorithm to be adopted based on the binary conversion result.

[0075] Regarding the binary conversion of the value of the number of memory blocks occupied by each row of data of the target matrix in the shared memory, according to an embodiment of the present invention, it may be corresponding to determining the number of memory blocks occupied by each row of data of the target matrix in the shared memory, and converting the determined numerical value into a binary representation form.

[0076] Continuing with the above example, for example, if the number of memory blocks occupied by each row of data of the target matrix in the shared memory is 4, binary conversion can be performed on it, and the result after binary conversion can be represented as b0100. Similarly, if the number of memory blocks occupied by each row of data of the target matrix in the shared memory is 8, the result after its binary conversion can be represented as b1000; if the number of memory blocks occupied by each row of data of the target matrix in the shared memory is 12, the result after its binary conversion can be represented as b1100; if the number of memory blocks occupied by each row of data of the target matrix in the shared memory is 16, the result after its binary conversion can be represented as b10000.

[0077] On this basis, the position of the first "1" can be determined according to the result after binary conversion. For example, for the binary conversion result b0100, the first "1" is in the third position (also called the "third digit") from the low order to the high order (i.e., the digits from right to left) in this binary conversion result.

[0078] Regarding "low order" and "high order", generally, taking the binary conversion result as an example, the more to the left the digit is, the higher the digit; the more to the right the digit is, the lower the digit.

[0079] Similarly, for the binary conversion result b1000, the first "1" is in the fourth position from the low order to the high order in this binary conversion result; for the binary conversion result b1100, the first "1" is in the third position from the low order to the high order in this binary conversion result; for the binary conversion result b10000, the first "1" is in the fifth position from right to left in this binary conversion result.

[0080] According to the inventive concept of the present invention, the address hashing algorithm to be adopted can then be determined based on the position of the first "1" in the binary conversion result. Specifically, based on the position of the first "1" in the binary conversion result, the address hashing algorithm to be adopted can be selected from a predetermined set of address hashing algorithms, so as to determine the data access layout for the target matrix based at least on the address hashing algorithm to be adopted.

[0081] Regarding the data access layout, it can refer to the access layout of the data in the target matrix in the shared memory.

[0082] Specifically, in an embodiment of the present invention, when the positions of the first "1" in the binary conversion results are the same, the same address hashing algorithm can be adopted. For example, for the binary conversion result b0100 and the binary conversion result b1100, the first "1" in their binary conversion results is both at the third position from the low order to the high order (from right to left) in the binary conversion result. Therefore, for the cases of the binary conversion result b0100 and the binary conversion result b1100, the same address hashing algorithm can be selected from the predetermined set of address hashing algorithms to obtain the data access layout of the target matrix.

[0083] As described above, in an embodiment of the present invention, the address hashing algorithm can refer to performing a hashing operation on the storage address corresponding to the data read from the target matrix. For different address hashing algorithms, it can refer to performing a hashing operation on the storage address corresponding to the data read from the target matrix based on different rules.

[0084] Specifically, according to an embodiment of the present invention, the address hashing algorithm to be adopted can be determined by judging whether the position of the first "1" in the binary conversion result meets a predetermined condition.

[0085] In some embodiments, when the position of the first "1" in the binary conversion result meets a first predetermined condition, for example, the position of the first "1" in the binary conversion result is any one of the third, fourth, or fifth position from the low order to the high order in the binary conversion result, the address hashing algorithm to be adopted can include: performing a hashing operation on no more than a first number of data addresses.

[0086] Regarding the first quantity, it may be related to the position of the first "1" in the binary conversion result. According to an embodiment of the present invention, the first quantity may be equal to the number of bits of the first "1" in the binary conversion result minus one, starting from the lower bit to the higher bit in the binary conversion result. For example, in one example, if the position of the first "1" in the binary conversion result is the third position starting from the lower bit to the higher bit in the binary conversion result, the first quantity is equal to 2. That is to say, when the position of the first "1" in the binary conversion result is the third position starting from the lower bit to the higher bit in the binary conversion result, the address hashing algorithm to be adopted may include two hashing operations respectively performed on two groups of data addresses. In another example, if the position of the first "1" in the binary conversion result is the fifth position starting from the lower bit to the higher bit in the binary conversion result, the first quantity is equal to 4. That is to say, when the position of the first "1" in the binary conversion result is the fifth position from right to left in the binary conversion result, the address hashing algorithm to be adopted may include four hashing operations respectively performed on four groups of data addresses.

[0087] Specifically, according to an embodiment of the present invention, taking the above-mentioned first "1" being in the third position starting from the lower bit to the higher bit in the binary conversion result as an example, for the embodiments satisfying this condition, in these embodiments, the address hashing algorithm adopted may be to perform hashing operations respectively on the first digit and the second digit in the binary representation of the storage address corresponding to the data read from the target matrix, so as to obtain the storage address after hash conversion, so that then the data in the target matrix can be stored in the shared memory based on the storage address after hash conversion. For example, a hashing operation such as an exclusive OR operation may be performed on the first digit (for example, denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix and the seventh digit (for example, denoted as addr_bin[6]) in the binary representation of the storage address corresponding to the data read from the target matrix; and a hashing operation such as an exclusive OR operation may be performed on the second digit (for example, denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix and the sixth digit (for example, denoted as addr_bin[5]) in the binary representation of the storage address corresponding to the data read from the target matrix. Thus, then the data access layout as shown in the appendix Figure 7 can be obtained.

[0088] Specifically, according to another embodiment of the present invention, taking the above first "1" being in the fifth position from the low bit to the high bit in the binary conversion result as an example, for embodiments satisfying this condition, in these embodiments, the address hashing algorithm adopted may be to perform hashing operations on the first digit, the second digit, the third digit, and the fourth digit in the binary representation of the storage address corresponding to the data read from the target matrix respectively, so as to obtain the storage address after hash conversion, such that then the data in the target matrix can be stored in the shared memory based on the storage address after hash conversion. For example, the first digit (e.g., denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the ninth digit (e.g., denoted as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix; the second digit (e.g., denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the eighth digit (e.g., denoted as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix; the third digit (e.g., denoted as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the sixth digit (e.g., denoted as addr_bin[5]) in the binary representation of the storage address corresponding to the data read from the target matrix; and the fourth digit (e.g., denoted as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the seventh digit (e.g., denoted as addr_bin[6]) in the binary representation of the storage address corresponding to the data read from the target matrix. Thus, then the data access layout as shown in the appendix Figure 8 can be obtained.

[0089] In some embodiments, when the position of the first "1" in the binary conversion result satisfies a second predetermined condition, for example, the position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low bit to the high bit in the binary conversion result, the address hashing algorithm to be adopted may include: performing a hashing operation on a second quantity of data addresses.

[0090] Regarding the second quantity, according to the inventive concept of the present invention, it may be five. That is, in an embodiment of the present invention, when the position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions (such as the seventh position from the low order to the high order, the tenth position from the low order to the high order, etc.) in the binary conversion result from the low order to the high order, the address hashing algorithm to be adopted includes five hashing operations respectively performed on five groups of data addresses. Specifically, the address hashing algorithm to be adopted may be: performing hashing operations respectively on the first digit to the fifth digit in the binary representation of the storage address corresponding to the data read from the target matrix to obtain the hashed storage address, so that then the data in the target matrix can be stored in the shared memory based on the hashed storage address. It should be understood, however, that in this case, when the position of the first "1" in the binary conversion result is different, the five groups of data addresses for which the hashing operations are to be performed are different five groups of data addresses.

[0091] For example, taking the case where the first "1" is in the sixth position from the right to the left in the binary conversion result, for embodiments satisfying this condition, the address hashing algorithm to be adopted can be as follows: The first digit (for example, denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the tenth digit (for example, denoted as addr_bin[9]) in the binary representation of the storage address corresponding to the data read from the target matrix; The second digit (for example, denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the ninth digit (for example, denoted as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix; The third digit (for example, denoted as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the sixth digit (for example, denoted as addr_bin[5]) in the binary representation of the storage address corresponding to the data read from the target matrix; The fourth digit (for example, denoted as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the seventh digit (for example, denoted as addr_bin[6]) in the binary representation of the storage address corresponding to the data read from the target matrix; and The fifth digit (for example, denoted as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the eighth digit (for example, denoted as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix.

[0092] For another example, taking the case where the first "1" is in the eighth position from the right to the left in the binary conversion result as an example, for embodiments satisfying this condition, the address hashing algorithm to be adopted can be as follows: The first digit (e.g., denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the twelfth digit (e.g., denoted as addr_bin

[11] ) in the binary representation of the storage address corresponding to the data read from the target matrix; the second digit (e.g., denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the eleventh digit (e.g., denoted as addr_bin

[10] ) in the binary representation of the storage address corresponding to the data read from the target matrix; the third digit (e.g., denoted as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the eighth digit (e.g., denoted as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix; the fourth digit (e.g., denoted as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the ninth digit (e.g., denoted as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix; and the fifth digit (e.g., denoted as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix can be subjected to a hashing operation such as an exclusive OR operation with the tenth digit (e.g., denoted as addr_bin[9]) in the binary representation of the storage address corresponding to the data read from the target matrix.

[0093] As can be seen from the above, when the position of the first "1" in the binary conversion result is any one of the sixth position or later positions from the right to the left in the binary conversion result, that is, when the size of each row of data in the target matrix is greater than or equal to 128 bytes, the address hashing algorithm to be adopted can include: The first to fifth digits in the binary representation of the storage address corresponding to the data read from the target matrix are respectively subjected to a hashing operation such as an exclusive OR operation with the five digits starting from the digit corresponding to the position of the aforementioned first "1" in the binary conversion result in the binary representation of the storage address corresponding to the data read from the target matrix, to determine the data access layout for the target matrix.

[0094] Next, it will be combined with Figure 7 and Figure 8Describe an example of the data access layout for the target matrix obtained based on the above address hashing algorithm according to an embodiment of the present invention.

[0095] Figure 7 FIG. shows a schematic diagram of the data access layout 700 of the target matrix according to an embodiment of the present invention, wherein each row of data of the target matrix occupies 4 memory blocks in the shared memory.

[0096] As Figure 7 shown, each small block corresponds to the size of a Dword data, and the number in each small block represents the index of the memory block used for storing the data in the small block, to indicate which memory block in the shared memory the data in the small block is stored in. Taking the first row as an example, the data in the small block of the first row and the first column can be stored in the memory block Bank-0 of the shared memory 250 such as Figure 3 , the data in the small block of the first row and the second column can be stored in the memory block Bank-1 of the shared memory 250 such as Figure 3 , the data in the small block of the first row and the third column can be stored in the memory block Bank-2 of the shared memory 250 such as Figure 3 , and the data in the small block of the first row and the fourth column can be stored in the memory block Bank-3 of the shared memory 250 such as Figure 3 . Also, for example, for the small block with the number 16 in the small block, the data in the small block can be correspondingly stored in the memory block Bank-16 of the shared memory 250 such as Figure 3 .

[0097] Based on the data access layout 700 as Figure 7 shown, if the data in the target matrix is stored in the shared memory 250 based on this data access layout 700, for matrices with different data types or their corresponding transposed matrices used for matrix multiply-accumulate operations, the data access in the matrix can be achieved within one cycle.

[0098] For example, if it is required to perform data access to the shared memory with an 8-row and 4-column matrix operation (for example, for non-transposed operations of all data type matrices, or for transposed data operations of 16-bit data matrices), obviously, the data in any sub-matrix with a dimension of 8×4 (i.e., 8 rows and 4 columns) in the data access layout 700 can be accessed corresponding to 32 different memory blocks respectively, as shown by the sub-matrix 701A or sub-matrix 701B in Figure 7 . That is to say, no matter how the sub-matrix with a dimension of 8×4 is selected, the numbers in the small blocks of the sub-matrix will not repeat. That is to say, the 32 groups of data in the sub-matrix will be necessarily indicated and stored in 32 different memory blocks in the shared memory 250 respectively, without memory block conflicts.

[0099] For example, to implement data access to shared memory through matrix operations with 16 rows and 2 columns (e.g., transpose data operations for an 8-bit data matrix), similarly, data in any sub-matrix with dimensions of 16×2 in the data access layout 700 can also be accessed corresponding to 32 different memory blocks respectively, as shown in Figure 7 sub-matrix 702A or sub-matrix 702B in

[0100] For another example, to implement data access to shared memory through matrix operations with 32 rows and 1 column (e.g., transpose data operations for a 4-bit data matrix), similarly, data in any sub-matrix with dimensions of 32×1 in the data access layout 700 can also be accessed corresponding to 32 different memory blocks respectively, as shown in Figure 7 sub-matrix 703 in

[0101] As can be seen from the above, for the data access layout 700 of the target matrix obtained according to the embodiments of the present invention, when reading an initial matrix for matrix multiply-accumulate or its corresponding transposed matrix with any data type from the shared memory, the occurrence of memory block conflicts in the shared memory can be avoided, thereby significantly improving the performance of the shared memory.

[0102] Figure 8 FIG. shows a schematic diagram of the data access layout 800 of the target matrix according to another embodiment of the present invention, where each row of data of the target matrix occupies 16 memory blocks in the shared memory.

[0103] Similar to the above Figure 7 case, Figure 8 each small block corresponds to the size of a Dword data, and the number in each small block represents the storage index for the data in that small block, to indicate which memory block in the shared memory the data in that small block is stored in. For specific reference, refer to the above description about Figure 7 here, which will not be elaborated further.

[0104] Similarly, based on the data access layout 800 as shown in Figure 8 if the data in the target matrix is stored in the shared memory 250 based on this data access layout 800, for matrices with different data types for matrix multiply-accumulate operations or their corresponding transposed matrices, the data access in the matrix can also be achieved within one cycle.

[0105] As shown in Figure 8As shown, the data in any sub-matrix with a dimension of 8×4 (i.e., eight rows and four columns) in the data memory access layout 800 can be accessed corresponding to 32 different memory blocks respectively, as shown in sub-matrix 801A, sub-matrix 801B or sub-matrix 801C; similarly, the data in any sub-matrix with a dimension of 16×2 in the data memory access layout 800 can also be accessed corresponding to 32 different memory blocks respectively, as shown in sub-matrix 802A or sub-matrix 802B; and the data in any sub-matrix with a dimension of 32×1 in the data memory access layout 800 can also be accessed corresponding to 32 different memory blocks respectively, as shown in sub-matrix 803A or sub-matrix 803B.

[0106] As can be seen from the above, based on the data memory access layout 800, if the data access of the shared memory is to be implemented by a matrix operation with eight rows and four columns, or by a matrix operation with sixteen rows and two columns, or by a matrix operation with thirty-two rows and one column, the occurrence of memory block conflicts in the shared memory can be avoided, thus significantly improving the performance of the shared memory.

[0107] In summary, the present invention proposes a scheme for data access in a shared memory. By determining the data memory access layout according to a predetermined address hashing algorithm and storing the data in the shared memory based on the determined data memory access layout, when reading an initial matrix or a corresponding transposed matrix for matrix multiplication and accumulation with any data type from the shared memory, the occurrence of memory block conflicts in the shared memory can be avoided, significantly improving the performance of the shared memory.

[0108] According to another concept of the present invention, based on the above scheme for data access in a shared memory, a stride parameter can be further introduced to jointly determine the data memory access layout of the target matrix in combination with the stride parameter.

[0109] Regarding the stride parameter, it can be denoted as "stride", which represents the stride for reading the input activation during the convolution operation in a convolutional neural network (Convolution). Corresponding to reading the matrix data in the shared memory, the stride parameter can represent the stride in the number of rows when reading the matrix data. For example, taking the stride parameter as 1 (i.e., stride = 1) as an example, it means that during the convolution operation in the convolutional neural network, the stride in the input channel dimension when reading the input activation is 1, that is, the stride in the number of rows when accessing the matrix data in the shared memory is 1. In other words, "stride = 1" can mean that when reading the initial matrix or the corresponding transposed matrix for matrix multiplication and accumulation from the shared memory based on the determined data memory access layout, the difference in the row indices of two consecutive rows of data read is 1. For example, the above Figure 7 andFigure 8 The data access layout of the target matrix shown is applicable to data access with a stride parameter of 1.

[0110] However, when the stride parameter is greater than 1, since data reading will be performed across rows, if the data access layout of the target matrix determined only based on the foregoing method is used to store the data in the target matrix into the shared memory, memory bank conflicts will occur when reading data from the shared memory.

[0111] Figure 9 FIG. shows a partial schematic diagram of the data access layout of the target matrix according to another embodiment of the present invention. Among them, the number of memory banks occupied by each row of data in the target matrix in the shared memory is 32, and the stride parameter is 1.

[0112] Similar to the above Figure 7 and / or Figure 8 Similar, Figure 9 In, each small block corresponds to the size of a Dword data, and the number in each small block represents the index of the memory bank used for storing the data in the small block, to indicate which memory bank in the shared memory the data in the small block is stored in. For details, reference can be made to the above description of Figure 7 or Figure 8 which will not be elaborated here.

[0113] Based on the data access layout 900 as shown in Figure 9 , if data access to the shared memory is implemented with an 8-row and 4-column matrix operation, when the stride parameter is 1, the data in any sub-matrix with a dimension of 8×4 (i.e., 8 rows and 4 columns) in the data access layout 900 can be accessed corresponding to 32 different memory banks, as shown in the sub-matrix 901 in Figure 9 . Therefore, the 32 groups of data in the sub-matrix 901 will be necessarily indicated and stored in 32 different memory banks in the shared memory 250 respectively, without memory bank conflicts occurring.

[0114] However, when the stride parameter is greater than 1 (such as, the stride parameter is 2), if data in a sub-matrix with the same dimension of 8×4 (for example, the sub-matrix 901 in Figure 9 ) is read from the shared memory based on the current data access layout 900, obviously, the 32 groups of data in the sub-matrix 901 include multiple data that are repeatedly indicated to the same memory bank, so memory bank conflicts will occur during the data reading process.

[0115] Therefore, to at least partially solve the memory block conflicts caused by reading data when the stride parameter is greater than 1, as described above, the method for accessing data in shared memory provided by the present invention may further include: determining the data layout in the shared memory based on the stride parameter and the data size of each row of the determined target matrix, and then storing the data in the target matrix into the shared memory based on the determined data layout. Specifically, according to an embodiment of the present invention, it is possible to: determine the value of the stride parameter for data access (denoted as stride); select the address hashing algorithm to be adopted from a predetermined address hashing algorithm based on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data in the target matrix in the shared memory (denoted as row_size) (denoted as stride * row_size); and determine the data access layout for the target matrix based on the address hashing algorithm to be adopted.

[0116] As described above regarding Figure 7 , Figure 8 , and / or Figure 9 In the exemplary embodiments of, since the value of the stride parameter stride = 1, stride * row_size = 1 * row_size = row_size, thus the address hashing algorithm to be adopted can be directly selected based on the value of the number of memory blocks occupied by each row of data in the target matrix in the shared memory, that is, by performing a binary conversion on the value of the number of memory blocks occupied by each row of data in the target matrix in the shared memory as described above, and selecting the address hashing algorithm to be adopted according to the position of the first "1" in the result after the binary conversion. For details, refer to the description above and will not be elaborated here.

[0117] However, in the case where the value of the stride parameter is greater than 1, according to an embodiment of the present invention, similarly, a binary conversion can be performed on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data in the target matrix in the shared memory, and the address hashing algorithm to be adopted can be selected according to the position of the first "1" in the result after the binary conversion.

[0118] Regarding selecting the address hashing algorithm to be adopted according to the position of the first "1" in the result after the binary conversion, it is similar to the content in the description above regarding Figure 6 and will not be elaborated here.

[0119] The following will be described in detail in conjunction with Figures 10 to 11 Detailed description.

[0120] Figure 10A schematic diagram of the data access layout 1000 of the target matrix according to an embodiment of the present invention is shown. Among them, the number of memory blocks occupied by each row of data of the target matrix in the shared memory is 32 (row_size = 32), and the stride parameter is 2 (stride = 2).

[0121] In the embodiment as Figure 10 shown, since row_size = 32 and stride = 2, the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory, that is, row_size * stride = 32 * 2 = 64. On this basis, performing binary conversion on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data of the target matrix in the shared memory (row_size * stride = 64), the result after binary conversion can be obtained as b1000000. Among them, the position of the first "1" is the seventh position from the low bit to the high bit in this binary conversion result.

[0122] In this embodiment, since the position of the first "1" is the seventh position from the low bit to the high bit in this binary conversion result, it can be known that the position of the first "1" in the binary conversion result meets the second predetermined condition, that is, the position of the first "1" in the binary conversion result is any one of the sixth position or later positions from the low bit to the high bit in the binary conversion result. Therefore, the address hashing algorithm to be adopted includes: performing hashing operation on the second quantity of data addresses.

[0123] For example, a hash operation such as an exclusive OR operation can be performed on the first digit (e.g., denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eleventh digit (e.g., denoted as addr_bin

[10] ) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the second digit (e.g., denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix and the tenth digit (e.g., denoted as addr_bin[9]) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the third digit (e.g., denoted as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix and the seventh digit (e.g., denoted as addr_bin[6]) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the fourth digit (e.g., denoted as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eighth digit (e.g., denoted as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix; and a hash operation such as an exclusive OR operation can be performed on the fifth digit (e.g., denoted as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix and the ninth digit (e.g., denoted as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix.

[0124] Compared with Figure 9 the data access layout 900 of the target matrix shown, since Figure 10 the value of the stride parameter corresponding to the embodiment shown is 2, that is, the stride in the number of rows is 2 when accessing data of the matrix in the shared memory, so when reading the initial matrix or the corresponding transposed matrix for matrix multiply-accumulate from the shared memory, the difference in the row index numbers of two consecutive rows of data read is 2. Thus, according to the inventive concept of the present invention, in order to avoid memory block conflicts caused by reading data across rows, the selected hash algorithm can make it such that when storing the data in the target matrix into the shared memory, the same address hash result is used for every two consecutive rows of data, that is, the address hash results of every two rows are the same. For example, the data in row 0 and row 1 use the same address hash result, the data in row 2 and row 3 use the same address hash result, the data in row 4 and row 5 use the same address hash result, and so on. Thus, the data access layout 1000 of the target matrix as shown in Figure 10 can be obtained.

[0125] As described above, based on the data access layout 1000 as shown in Figure 10 the data in the target matrix is stored in the shared memory and then the matrix with different data types or its corresponding transposed matrix for matrix multiply-accumulate operations is read from the shared memory, such as Figure 10 the sub-matrix 1001 with a dimension of 15×4 and a stride parameter of 2, or the sub-matrix 1002 with a dimension of 31×2 and a stride parameter of 2 as shown in

[0126] Figure 11 can achieve the data access in the matrix within one cycle, thus avoiding memory bank conflicts.

[0127] As in Figure 11 the embodiment shown, since row_size = 32 and stride = 4, the product of the value of the stride parameter and the number of memory blocks occupied by each row of data in the target matrix in the shared memory, that is, row_size * stride = 32 * 4 = 128. On this basis, binary conversion is performed on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data in the target matrix in the shared memory (row_size * stride = 64), and the result after binary conversion can be obtained as b10000000, where the position of the first "1" is the eighth position from the low bit to the high bit in the binary conversion result. Similar to Figure 10 the embodiment shown, the position of the first "1" in the binary conversion result of this embodiment also satisfies the second predetermined condition, so the address hashing algorithm to be adopted includes: performing hashing operations on the second quantity of data addresses.

[0128] For example, a hash operation such as an exclusive OR operation can be performed on the first digit (e.g., denoted as addr_bin[0]) in the binary representation of the storage address corresponding to the data read from the target matrix and the twelfth digit (e.g., denoted as addr_bin

[11] ) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the second digit (e.g., denoted as addr_bin[1]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eleventh digit (e.g., denoted as addr_bin

[10] ) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the third digit (e.g., denoted as addr_bin[2]) in the binary representation of the storage address corresponding to the data read from the target matrix and the eighth digit (e.g., denoted as addr_bin[7]) in the binary representation of the storage address corresponding to the data read from the target matrix; a hash operation such as an exclusive OR operation can be performed on the fourth digit (e.g., denoted as addr_bin[3]) in the binary representation of the storage address corresponding to the data read from the target matrix and the ninth digit (e.g., denoted as addr_bin[8]) in the binary representation of the storage address corresponding to the data read from the target matrix; and a hash operation such as an exclusive OR operation can be performed on the fifth digit (e.g., denoted as addr_bin[4]) in the binary representation of the storage address corresponding to the data read from the target matrix and the tenth digit (e.g., denoted as addr_bin[9]) in the binary representation of the storage address corresponding to the data read from the target matrix.

[0129] Further, since Figure 11 the value of the stride parameter corresponding to the embodiment shown is 4, when reading the initial matrix or the corresponding transposed matrix for matrix multiplication and accumulation from the shared memory, the difference in the row index numbers of two consecutive rows of data read is 4. Similarly, in order to avoid memory block conflicts caused by reading data across rows, in this case, the selected hash algorithm can ensure that when storing the data in the target matrix into the shared memory, the same address hash result is used for every four consecutive rows of data, that is, the address hash results of every four rows are the same. For example, the data in rows 0, 1, 2, and 3 use the same address hash result, the data in rows 4, 5, 6, and 7 use the same address hash result, the data in rows 8, 9, 10, and 11 use the same address hash result, and so on. Thus, the data access layout 1100 of the target matrix as shown in Figure 11 can be obtained.

[0130] Thus, based on as Figure 11The data access layout 1100 shown stores the data in the target matrix into the shared memory and then reads from the shared memory a matrix with different data types or its corresponding transposed matrix for matrix multiply-accumulate operations, such as Figure 11 the sub-matrix 1101A or sub-matrix 1101B shown in Figure 11 with dimensions of 29×4 and a stride parameter of 4, which can achieve data access in the matrix within one cycle, thus avoiding memory bank conflicts.

[0131] In summary, the solution for data access in the shared memory provided by the present invention, based on the determined data access layout of the target matrix, can also take the stride parameter as a consideration factor, so as to store the data in the target matrix into the shared memory based on the extended data access layout, such that no memory bank conflicts occur when reading data across rows in the shared memory, thereby improving the performance of the shared memory.

[0132] In addition, each of the processes and treatments described herein, such as method 600, can be executed at a computing device. The computing device includes, for example: at least one processor (at least one graphics processor and at least one central processor); and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor. In some embodiments, method 600 can be implemented as a computer software program or program product, which is tangibly embodied in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computing device via a read-only memory (ROM) and / or a communication unit. When the computer program is loaded into a random-access memory (RAM) and executed by an AI processor such as a GPU, GPGPU, NPU, etc., one or more actions of method 600 described above can be executed.

[0133] The present invention can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention. The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.

[0134] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to an embodiment of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.

[0135] These computer-readable program instructions can be provided to a central processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the central processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0136] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart, and the combinations of blocks in the block diagrams and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.

[0137] It should be understood that the various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. No limitations are set forth herein.

[0138] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

[0139] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for accessing data in shared memory, characterized in that Including: Determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory; Determining the data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data in the target matrix in the shared memory according to a predetermined address hashing algorithm, so as to store the data in the target matrix into the shared memory based on the data access layout, where it includes Performing binary conversion on the value of the number of memory blocks occupied by each row of data in the target matrix in the shared memory, and based on the position of the first "1" in the binary conversion result, selecting the address hashing algorithm to be adopted from the predetermined address hashing algorithm.

2. The method according to claim 1, wherein The shared memory includes multiple memory blocks with the same bandwidth; And Determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory includes: Determining the total number of bytes of each row of data in the target matrix; Determining the bandwidth size of the memory block; and Based on the total number of bytes of each row of data in the target matrix and the bandwidth size of the memory block, determining the number of memory blocks occupied by each row of data in the target matrix in the shared memory.

3. The method according to claim 1, wherein Determining the data access layout for the target matrix includes: Determining the data access layout for the target matrix based on at least the address hashing algorithm to be adopted.

4. The method according to claim 3, characterized in that, Selecting the address hashing algorithm to be adopted from the predetermined address hashing algorithm includes: When the position of the first "1" in the binary conversion result satisfies the first predetermined condition, the address hashing algorithm to be adopted includes performing hashing operations on no more than the first number of data addresses; When the position of the first "1" in the binary conversion result satisfies the second predetermined condition, the address hashing algorithm to be adopted includes performing hashing operations on the second number of data addresses.

5. The method according to claim 4, characterized in that, The first predetermined condition is: The position of the first "1" in the binary conversion result is any one of the third, fourth, or fifth positions from the low bit to the high bit in the binary conversion result.

6. The method according to claim 4, characterized in that The second predetermined condition is: The position of the first "1" in the binary conversion result is any one of the sixth position or subsequent positions from the low bit to the high bit in the binary conversion result.

7. The method according to claim 4, characterized in that, The first number is related to the position of the first "1" in the binary conversion result, and the second number is five.

8. The method according to claim 7, characterized in that, The first number is equal to the number of bits of the first "1" in the binary conversion result counted from the low bit to the high bit in the binary conversion result minus one.

9. The method according to claim 1, wherein Determining the data access layout for the target matrix based on at least the number of memory blocks occupied by each row of data in the target matrix in the shared memory according to a predetermined address hashing algorithm includes: Determining the value of the stride parameter for data access; Based on the product of the value of the stride parameter and the number of memory blocks occupied by each row of data in the target matrix in the shared memory, selecting the address hashing algorithm to be adopted from the predetermined address hashing algorithm; and Based on the address hashing algorithm to be adopted, determining the data access layout for the target matrix.

10. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.

11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-9.

12. A computer program product, characterized in that, Comprising a computer program which, when executed by a machine, executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Data processing method and device and electronic equipment

    CN112069460A

  • Data storage method, data reading method, electronic equipment and storage medium

    CN117785759A