Data read-write method, processor, electronic equipment and storage medium
By interleaving the target tensor data in a general-purpose graphics processor and adopting a loop operation with balanced access, the memory channel conflict problem caused by the mismatch between the computing unit task division direction and data arrangement is solved, thereby improving bandwidth utilization and performance.
Patent Information
- Application Number
- CN202511254201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-04
AI Technical Summary
In general-purpose graphics processors, memory channel conflicts caused by the mismatch between the task division direction of the computing unit and the data arrangement direction lead to low bandwidth utilization. Existing technologies usually alleviate this problem by modifying the task division method of the computing unit or increasing the reserved space of the video memory, but the effect is limited.
By interleaving the target tensor data in the memory at a first granularity along a first direction, dividing the tasks of the M computing units along a second direction, and adopting a loop operation at a second granularity, the access of the computing units is balanced to N memory channels, thereby optimizing bandwidth utilization.
When the task division of computing units is limited, it can effectively alleviate memory channel conflicts, improve bandwidth utilization, and enhance the overall performance of memory-intensive operators.
Smart Images

Figure CN120762918A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data reading and writing method, a processor, an electronic device, and a storage medium. Background Art
[0002] As computing devices become increasingly powerful and the amount of data they need to process grows, central processing units (CPUs) and coprocessors continue to develop rapidly. The most common coprocessor is the graphics processing unit (GPU), which primarily handles computations related to graphics display. Later, GPUs evolved to handle general-purpose computing tasks, giving rise to general-purpose graphics processing units (GPGPUs). GPGPU chips not only provide powerful graphics processing capabilities but can also be used to perform highly parallel scientific computing, data analysis, machine learning, and other compute-intensive tasks.
[0003] General-purpose graphics processors are typically equipped with high-speed graphics memory and multiple parallel computing units. High-speed graphics memory has higher memory bandwidth than the CPU, which helps accelerate data-intensive applications. Multiple parallel computing units transmit data to and from the high-speed graphics memory through multiple memory channels. During data transmission, due to algorithm design or task partitioning, data access by computing units may be concentrated in only a few memory channels during program execution, while other memory channels are idle. In this case, bandwidth utilization is lower than when the memory channels are fully loaded. This situation is called memory channel conflict. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a data reading and writing method for M computing units to read and write a target tensor from a memory, wherein the data of the target tensor is interleaved in the memory at a first granularity according to a first direction of the target tensor, the tasks of the M computing units are divided along a second direction of the target tensor, and the M computing units read and write the data from the memory through N memory channels to perform computing tasks. The method includes: coordinately representing multiple sub-blocks obtained by dividing the target tensor according to the first granularity; performing read and write operations on the multiple sub-blocks at a second granularity, wherein the following loop operation is performed N times for each second granularity until all data of the second granularity is read and written: based on the coordinate representation and the index of each computing unit, determining the target sub-block of the multiple sub-blocks of each computing unit in this loop operation, so that the access of the M computing units to the memory is balanced to the N memory channels; each computing unit performs read and write operations on the target sub-block through the corresponding memory channel, and N and M are both positive integers.
[0005] For example, in at least one data reading and writing method provided in the present disclosure, based on the coordinate representation and the index of the computing unit, the target sub-block among the multiple sub-blocks of the current loop operation of each computing unit is calculated, including: taking the remainder of the first value and N as the coordinate of the target sub-block in the second direction, and the index of the computing unit as the coordinate of the target sub-block in the first direction, wherein the first value is obtained based on the index of the computing unit; and determining the target sub-block based on the coordinate representation of the multiple sub-blocks.
[0006] For example, in at least one data reading and writing method provided in the present disclosure, for the first loop operation in N loop operations, the first value is the index of the computing unit; for the loop operation after the first loop operation, the first value is the sum of the previous first value and 1.
[0007] For example, in at least one data reading and writing method provided in the present disclosure, each computing unit performs read and write operations on the target sub-block through the corresponding memory channel, including: based on the first granularity, converting the coordinates of the target sub-block of the current loop operation into the index position of the target tensor; judging whether the index position exceeds the range of the target tensor; in response to the index position not exceeding the range of the target tensor, each computing unit performs read and write operations on the target sub-block through the corresponding memory channel.
[0008] For example, in at least one data reading and writing method provided in the present disclosure, the first direction is the row direction, the second direction is the column direction, and the second granularity is [M, N]; or the first direction is the column direction, the second direction is the row direction, and the second granularity is [N, M].
[0009] For example, in at least one data read-write method provided by the present disclosure, performing read-write operations on the plurality of sub-blocks at the second granularity includes: determining a first maximum value of the plurality of sub-block coordinate values in the first direction and a second maximum value of the plurality of sub-block coordinate values in the second direction based on the size of the target tensor; starting from the sub-block with the minimum coordinate value in the first direction and the second direction, performing the loop operation; after the loop operation is completed, performing the loop operation after the coordinate value in the second direction increases by M each time until the coordinate value in the second direction reaches the maximum value; and then performing the loop operation after the coordinate value in the first direction increases by N each time until the coordinate value in the first direction reaches the maximum value, so that the data in the tensor is completely read and written.
[0010] For example, in at least one data read-write method provided by the present disclosure, in the case that the size of the first granularity in the second direction is greater than the number of threads processed by each computing unit at a time, each computing unit performs read-write operations on the target sub-block through a corresponding memory channel, including: dividing the first granularity in the second direction into a plurality of sub-target sub-blocks according to the number of threads, and sequentially performing read-write operations on the plurality of sub-target sub-blocks.
[0011] For example, in at least one data read-write method provided by the present disclosure, the second granularity is determined according to the number of computing units, the number of memory channels, the task division direction of the computing unit, and the data arrangement.
[0012] At least one embodiment of the present disclosure further provides a processor, including an instruction parsing unit and an execution unit, wherein the instruction parsing unit is configured to receive and parse a data read-write instruction, the data read-write instruction being used for N computing units to read and write a target tensor from a memory; and the execution unit executes the data read-write method provided by any one of the embodiments of the present disclosure after the instruction parsing unit parses the data read-write instruction.
[0013] At least one embodiment of the present disclosure further provides an electronic device, including: a memory, which non-transiently stores computer executable instructions; and a processor, which is configured to run the computer executable instructions, wherein the computer executable instructions, when run by the processor, implement the data read-write method provided by any one of the embodiments of the present disclosure.
[0014] At least one embodiment of the present disclosure further provides a computer readable storage medium, which non-transiently stores computer executable instructions, and the computer executable instructions, when executed by a processor, implement the data read-write method provided by any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0016] Figure 1 A schematic structural diagram of a general-purpose graphics processing unit (GPGPU); Figure 2 A schematic diagram of the structure of HBM storage data is shown; Figure 3 A schematic diagram showing the relationship between multiple data storage units and multiple memory channels is shown; Figure 4A A schematic diagram showing a row-major data layout; Figure 4B A schematic diagram showing a column-major data layout; Figure 5 A schematic diagram showing a memory channel conflict; Figure 6A Shows a schematic flow chart of a data reading and writing method provided by at least one embodiment of the present disclosure; Figure 6B A flow chart of a method for cyclic operation provided by at least one embodiment of the present disclosure is shown; Figure 7A A schematic diagram showing a data reading and writing method provided by at least one embodiment of the present disclosure; Figure 7B A schematic diagram illustrating another data reading and writing method provided by at least one embodiment of the present disclosure is shown; Figure 8 At least one embodiment of the present disclosure provides a Figure 6B Method flow chart of step S22; Figure 9 At least one embodiment of the present disclosure provides a Figure 6A Flowchart of the method of step S20; Figure 10 A schematic diagram of another data reading and writing method provided by at least one embodiment of the present disclosure is shown; Figure 11 A schematic block diagram of a data reading and writing device provided in at least one embodiment of the present disclosure; Figure 12 A schematic structural diagram of a processor provided for at least one embodiment of the present disclosure; Figure 13 A schematic block diagram of an electronic device provided in accordance with an embodiment of the present disclosure; and Figure 14 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0018] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0019] The present disclosure is described below through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of known functions and known components. When any component of an embodiment of the present disclosure appears in more than one figure, the component is represented by the same or similar reference numeral in each figure.
[0020] Figure 1 A schematic structural diagram of a general-purpose graphics processing unit (GPGPU).
[0021] like Figure 1 As shown, the general purpose graphics processor is actually an array of programmable multiprocessors. For example, the programmable multiprocessor can be a streaming processor cluster (SPC), including Figure 1 Streaming processor clusters 1, ..., and M are shown, where M is a positive integer greater than 1. In a general-purpose graphics processor, one streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Multiple streaming processor clusters share data through a global cache or global memory.
[0022] like Figure 1 As shown, taking stream processor cluster 1 as an example, a stream processor cluster includes multiple computing units (ComputeUnit, referred to as CU), for example Figure 1 In the calculation unit 1, calculation unit 2, ..., calculation unit N, N is a positive integer. Each calculation unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A calculation unit includes multiple cores (also called calculation cores or calculation cores), each of which includes an arithmetic logic unit (ALU), a floating-point calculation unit, etc. The calculation core is used to perform specific calculation tasks. In addition, the calculation unit also includes registers (such as Figure 3 The register file in the computing unit and the shared memory are used to hierarchically store source data and destination data related to computing tasks. The shared memory in a computing unit is used to share data between the cores of the computing unit.
[0023] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processor (or parallel computing processor), and then distributed through the thread block distribution module ( Figure 1 (not shown) distributes multiple thread blocks to various compute units. All threads in a thread block must be assigned to the same compute unit for execution. When performing operations on large matrices or tensors, they are typically divided into sub-blocks that match the hardware characteristics. Each sub-block is processed by a thread block, and each thread within a thread block is responsible for a specific element in the block. Thread blocks are then scheduled to compute units for execution, and a single compute unit can accommodate multiple thread blocks simultaneously.
[0024] At the same time, thread blocks are split into minimum execution warps (or simply warps), each of which contains a fixed number of threads (or less than this fixed number), for example, 32 threads. Multiple thread blocks can be executed in the same compute unit or in different compute units.
[0025] In each computational unit, the warp scheduling / dispatching module ( Figure 1 (not shown) schedules and allocates thread warps so that the multiple compute cores of the compute unit can execute the warps. Depending on the number of compute cores in the compute unit, multiple warps in a thread block can execute simultaneously or in a time-sharing manner. Multiple threads in each warp execute the same instruction. Memory execution instructions are sent to the shared memory within the compute unit or further to the mid-level cache, global cache, or global memory for read and write operations.
[0026] Global cache / memory, also known as video memory, is a core component supporting the large-scale parallel computing capability of GPGPU, and its design directly affects the throughput, latency and energy efficiency of GPGPU. GPGPU video memory technology evolves with computing needs, and the current mainstream types can be divided into Graphics Double Data Rate (GDDR) and High Bandwidth Memory (HBM). GDDR is a video memory technology optimized for graphics processing and parallel computing, widely used in mid-end GPUs (such as game graphics cards and entry-level computing cards); HBM adopts a 3D stacked architecture and through-silicon via (TSV) technology, focusing on ultra-high bandwidth and energy efficiency ratio, and is the core video memory solution for high-end GPGPU (such as AI training cards and supercomputer accelerator cards).
[0027] Video memory (e.g., HBM) is usually divided into multiple data storage units (also referred to as "sections"). Continuous memory data is interleaved in different data storage units at a certain granularity (e.g., 256B, 512B, etc.), which is referred to as a sub-block in embodiments of the present disclosure.
[0028] Figure 2 A structure diagram of HBM storage data is shown.
[0029] As shown in Figure 2 , the HBM is divided into 16 data storage units, namely data storage unit S0~data storage unit S15. For example, the HBM is 32G, and each data storage unit can store 2G of data. For example, the data is interleaved in the 16 data storage units at a granularity of 512B, i.e., each sub-block includes 512B of data. As shown in Figure 2 , sub-block B0 stores 0~511B of data in data storage unit S0, sub-block B1 stores 512~1023B of data in data storage unit S1, and so on, until sub-block B15 is stored in data storage unit S15, and then sub-block B16 is stored in data storage unit S0, and so on.
[0030] The channel that reads data from the video memory is called a memory channel. Typically, a data storage unit can be accessed by multiple memory channels, and a memory channel can also access multiple data storage units. However, each sub-block has only one memory channel associated with it. The memory channel is the hardware data transmission path that connects the computing unit to the video memory. Its core function is to enable high-speed data exchange between the computing unit and the video memory. The hardware components of the memory channel include, for example, a video memory controller, a data bus, a physical interface, an arbiter, and so on. Different GPU or GPGPU architectures (such as GDDR and HBM) use different technologies in channel design. This disclosure does not specifically limit the hardware structure of the memory channel.
[0031] Figure 3 A schematic diagram showing the relationship between multiple data storage units and multiple memory channels is shown.
[0032] like Figure 3 As shown, data storage units S0 through S3 correspond to memory channels C0 through C3; data storage units S4 through S7 correspond to memory channels C4 through C7. Memory channel C0 can access each of data storage units S0 through S3; memory channel C1 can access each of data storage units S0 through S3; memory channel C2 can access each of data storage units S0 through S3; and memory channel C3 can access each of data storage units S0 through S3. Data storage unit S0 is accessible by each of memory channels C0 through C3; data storage unit S1 is accessible by each of memory channels C0 through C3; data storage unit S2 is accessible by each of memory channels C0 through C3; and data storage unit S3 is accessible by each of memory channels C0 through C3. Similarly, the data storage units S4 to S7 and the memory channels C4 to C7 are similar to the data storage units S0 to S3 and the memory channels C0 to C3 described above.
[0033] From a software perspective, the data structure processed is a tensor. The data layout of a two-dimensional tensor includes row-major, column-major, and plain. Different data layouts result in the same logical location but different physical locations in memory.
[0034] Figure 4A A schematic diagram showing a row-major data layout; Figure 4B A schematic diagram showing a column-major data layout.
[0035] like Figure 4AAs shown, a [H, W] tensor, each data in the tensor is, for example, 4 bytes. The [H, W] tensor is divided into 16 sub-blocks, each of which is, for example, a small 4×32 matrix, that is, each sub-block is 512 bytes. For example, if the data in the [H, W] tensor is arranged in row-first order, then sub-blocks B0, B1, B2, and B3 are first arranged in data storage unit S0, data storage unit S1, data storage unit S2, and data storage unit S3 in the row direction, and then the data in the second and third rows are arranged in the data storage units in sequence.
[0036] like Figure 4B As shown, if the data in the [H, W] tensor is laid out in column priority, sub-blocks B0, B1, B2, and B3 are first arranged in the data storage unit S0, S1, S2, and S3 in sequence along the column direction, and then the data in the second and third columns are arranged in sequence in the data storage units.
[0037] For example, Figure 4A The sub-block B1 and Figure 4B The sub-blocks B4 in the [H, W] tensor are all the data in the 0th to 3rd rows and 32nd to 63rd columns, but in Figure 4A In the layout of , the data of rows 0 to 3 and columns 32 to 63 are stored in the data storage unit S1. Figure 4B In the layout, the data in rows 0 to 3 and columns 32 to 63 are stored in the data storage unit S4.
[0038] In GPGPUs, multi-threaded parallel read / write and computation are performed on a compute unit basis. Typically, a compute unit accesses one or more consecutive sub-blocks. The number of compute units is typically multiples greater than the number of memory channels. When data access by a compute unit is concentrated on only a few memory channels for a period of time, leaving other memory channels idle, bandwidth utilization is lower than when the memory channels are fully loaded. This situation is called a memory channel conflict.
[0039] In some large-model or high-performance task scenarios, the algorithm principle requires that all computing unit processing tasks be divided in the row or column direction. For example, during the unpermutation calculation of the Mixture of Experts (MOE), the parallel computing task needs to process a two-dimensional tensor [H, W], where the H direction represents the number of tokens and the W direction represents the vector representation of each token. The unpermutation algorithm will accumulate the tokens of different experts to the corresponding place based on the index. In order to ensure the correctness of the accumulation, each expert must have synchronization operations. Therefore, the computing unit tasks need to be divided in the W direction, so that only internal synchronization of the CU is required. Otherwise, global computing unit synchronization will be introduced, causing greater performance degradation. For more information about the mixed expert model, please refer to relevant materials, and this disclosure will not go into details.
[0040] When the task division direction of the computing unit is restricted, for example, all tasks are arranged in the row direction and the tensor data layout is column-first, the data accessed by all computing units is concentrated in certain memory channels, causing some memory channels to be congested and some memory channels to be idle, resulting in low overall bandwidth utilization.
[0041] Figure 5 A schematic diagram showing a memory channel conflict.
[0042] like Figure 5 As shown in the figure, the task division of the computing unit is based on the row direction of the tensor, that is, the data in the tensor is allocated to the computing unit for computing tasks according to the row direction of the tensor. The data layout of the tensor is column-first. The computing units CU0~CU7 execute computing tasks in parallel, so the computing units CU0~CU7 all read data from memory channel C0, while memory channel C1~memory channel C3 are idle. Therefore, when the task division direction of the computing unit is different from the data arrangement direction, it may cause memory channel conflicts, resulting in low overall bandwidth utilization.
[0043] In related technologies, memory channel conflicts are typically mitigated by modifying the task partitioning method of the computing unit or adding video memory padding. The embodiments provided by this disclosure can optimize bandwidth utilization in high-performance scenarios where the task partitioning method of the computing unit is limited and no video memory is added.
[0044] At least one embodiment of the present disclosure provides a data read-write method for M computing units to read and write a target tensor from a memory, data of the target tensor is interleaved in the memory according to a first direction of the target tensor with a first granularity, task division of the M computing units is along a second direction of the target tensor, the M computing units read and write data from the memory through N memory channels to perform a computing task, the method comprises: performing coordinate representation on a plurality of sub-blocks obtained by dividing the target tensor according to the first granularity; performing read-write operation on the plurality of sub-blocks with a second granularity, for each second granularity, performing N times of the following loop operation until data of the second granularity is read and written completely: based on the coordinate representation and an index of each computing unit, calculating a target sub-block in the plurality of sub-blocks of each computing unit in this loop operation, so that access of the M computing units to the memory is balanced to the N memory channels; each computing unit performs read-write operation on the target sub-block through a corresponding memory channel, and N and M are positive integers. The data read-write method can alleviate memory channel conflict under the condition that task division of the computing unit is limited, thereby improving bandwidth utilization.
[0045] Figure 6A A schematic flowchart of a data read-write method provided by at least one embodiment of the present disclosure is shown; Figure 6B A method flowchart of a loop operation provided by at least one embodiment of the present disclosure is shown. The data read-write method is for M computing units to read and write a target tensor from a memory, data of the target tensor is interleaved in the memory according to a first direction of the target tensor with a first granularity, task division of the M computing units is along a second direction of the target tensor, and the M computing units read and write data from the memory through N memory channels to perform a computing task.
[0046] For example, the memory is the above-described video memory, and the target tensor is, for example, a tensor of [H, W] as shown in Figure 4A and Figure 4B The tensor can be arranged according to column priority as shown in Figure 4A , or arranged according to row priority as shown in Figure 4B For example, if the tensor is arranged according to row priority as shown in Figure 4A , the first direction is the W direction (i.e., the row direction), and the second direction is the H direction (i.e., the column direction). For example, if the tensor is arranged according to column priority as shown in Figure 4B , the first direction is the H direction, and the second direction is the W direction.
[0047] For example, the target tensor is interleaved in the video memory according to a granularity of 512B, as shown in the example of Figure 2 , and 512B is an example of the first granularity. For interleaved arrangement, refer to the description of Figure 2 above. In the following, N=4 and M=8 are taken as examples to illustrate some embodiments of the present disclosure.
[0048] As Figure 6A shown, the data read-write method includes steps S10-S20.
[0049] Step S10, the target tensor is divided into a plurality of sub-blocks according to a first granularity, and the plurality of sub-blocks are represented by coordinates.
[0050] Step S20, read-write operation is performed on the plurality of sub-blocks with a second granularity.
[0051] For each second granularity, N times of Figure 6B As shown in the loop operation until the data of the second granularity is read and written.
[0052] As Figure 6B shown, the loop operation includes steps S21 and S22.
[0053] Step S21, based on the coordinate representation and the index of the calculation unit, the target sub-block in the plurality of sub-blocks of each calculation unit in this loop operation is calculated, so that the access of M calculation units to the display is balanced to N memory channels.
[0054] Step S22, each calculation unit performs read-write operation on the target sub-block through the corresponding memory channel.
[0055] Based on the memory access characteristics of HBM and calculation units, for the problem of low bandwidth utilization caused by the fact that the calculation unit access is "stacked" in some memory channels while other memory channels are idle in the scenario where the calculation unit task partition is limited, the embodiment converts the data reading position algorithm (i.e., the position of the target sub-block) by software, so that the access of different calculation units to HBM is evenly distributed to different memory channels under the premise of logical invariance, thereby improving the bandwidth utilization and improving the overall performance of the memory-intensive operator.
[0056] Figure 7A A schematic diagram of a data read-write method provided by at least one embodiment of the present disclosure is shown. The foregoing steps will be described below. Figure 7A
[0057] As Figure 7A shown, the target tensor [H, W] is divided into a plurality of sub-blocks according to a granularity of 512B, and 512B is an example of the first granularity. The first granularity can be determined according to the number of threads processed in parallel by the calculation unit and the shared memory capacity. For example, the calculation unit performs 32 threads in parallel, and the first granularity can be P x 32, i.e., P rows and 32 columns as a sub-block, P being a positive integer, and the size of P can be determined according to the shared memory capacity, for example, P = 4.
[0058] For step S10, for example Figure 7A The multiple sub-blocks in are represented by coordinates according to their arrangement positions in the row and column directions. For example, Figure 7A The coordinates of the sub-block BLOCK in the first row and first column are (0,0); the coordinates of the sub-block BLOCK in the second row and first column are (1,0); the coordinates of the sub-block BLOCK in the first row and second column are (0,1); the coordinates of the sub-block BLOCK in the second row and second column are (1,1); and so on.
[0059] Regarding step S20, in some embodiments of the present disclosure, the second granularity is determined according to the number of computing units, the number of memory channels, the task division direction of the computing units, and the data arrangement.
[0060] For example, the first direction is the row direction, the second direction is the column direction, and the second granularity is [M, N]; or the first direction is the column direction, the second direction is the row direction, and the second granularity is [N, M].
[0061] For example, Figure 7A As shown in the figure, the number of computing units is 8, the number of memory channels is 4, the computing unit tasks are divided along the rows of the tensor, and the data is arranged in column-first order, so the second granularity can be [4, 8]. That is, read and write operations are performed in units of 4 rows and 8 columns. For example, when multiple computing units read data from the video memory, the eight computing units (CU0~CU7) first read and write BLOCK(0, 0)~BLOCK(0, 7), BLOCK(1, 0)~BLOCK(1, 7), BLOCK(2, 0)~BLOCK(2, 7), BLOCK(3, 0)~BLOCK(3, 7), and BLOCK(4, 0)~BLOCK(4, 7); and then read and write BLOCK(0, 8)~BLOCK(0, 15), BLOCK(1, 8)~BLOCK(1, 15), BLOCK(2, 8)~BLOCK(2, 15), BLOCK(3, 8)~BLOCK(3, 15), and BLOCK(4, 8)~BLOCK(4, 15) (not shown). After the traversal in the row direction is completed, the traversal in the column direction is performed at a granularity of 4 rows and 8 columns until all sub-blocks are read and written. Figure 7A Only part of the data in the tensor (6 rows and 8 columns) is shown. In practice, the size of the tensor can be larger or smaller. The embodiments of the present disclosure do not limit the size of the tensor.
[0062] For another example, if the number of computing units is 8, the number of memory channels is 4, the task division direction of the computing units is the column direction of the tensor, and the data is arranged in row priority, then the second granularity can be [8, 4].
[0063] For each second granularity, execute Figure 6B The loop operation shown is as follows. In step S21, for example, the coordinates of the target sub-block of this loop operation are calculated based on the index of each calculation unit, and then the target sub-block is located based on the coordinate representation. In step S22, a read / write operation is performed on the target sub-block. For example, the target sub-block is read from the data storage unit of the video memory through the corresponding memory channel.
[0064] In some embodiments of the present disclosure, step S21 includes: taking the remainder of the first value and N as the coordinate of the target sub-block in the second direction, the index of the calculation unit as the coordinate of the target sub-block in the first direction, the first value being obtained based on the index of the calculation unit; and determining the target sub-block based on the coordinate representation of multiple sub-blocks.
[0065] For example, when the tensor data layout is column-major and the computational unit tasks are split in the row direction, that is, the first direction is the column direction and the second direction is the row direction, if the index of the computational unit is CU_idx, then the coordinates of the target sub-block are (R%N, CU_idx), where R is obtained based on the index of the computational unit. For example, R is the index of the computational unit itself, or it is calculated from the index of the computational unit. In the embodiments of the present disclosure, "%" represents the remainder, that is, the remainder is calculated.
[0066] In some embodiments of the present disclosure, for the first loop operation among M loop operations, the first value is the index of the computing unit; for loop operations after the first loop operation, the first value is the sum of the previous first value and 1.
[0067] like Figure 7AAs shown, since there are four memory channels, four loops are performed. In each loop, memory access by multiple compute units is balanced across the four memory channels. That is, through four loops, each compute unit sequentially accesses four memory channels. For example, in the first loop, the first value is the compute unit index, and the coordinates of the target sub-block are (CU_idx%N, CU_idx). That is, compute unit CU0 accesses BLOCK(0, 0) through memory channel C0; compute unit CU1 accesses BLOCK(1, 1) through memory channel C1; compute unit CU2 accesses BLOCK(2, 2) through memory channel C2; compute unit CU3 accesses BLOCK(3, 3) through memory channel C3; compute unit CU4 accesses BLOCK(0, 4) through memory channel C0, and so on. That is, in the first loop, the sub-blocks marked with dashed boxes in the tensor are read and written through the corresponding memory channels. For the second loop operation and subsequent loop operations, the coordinates of the target sub-block in the second direction of each loop operation are the remainder of the sum of the coordinates in the second direction of the previous loop operation and 1 and N, that is, the target sub-block coordinates are ((CU_idx+1)%N,CU_idx). For example, for the second loop operation, the target sub-block coordinates are ((CU_idx+1)%N,CU_idx). For example, computing unit CU0 accesses BLOCK ((0+1)%4,0) = BLOCK (1,0) through memory channel C0; computing unit CU1 accesses BLOCK ((1+1)%4,1) = BLOCK (2,1) through memory channel C1; and so on. That is, in the second loop operation, the sub-blocks marked with solid boxes in the tensor are read and written through the corresponding memory channels. In the third loop operation, compute unit CU0 accesses BLOCK ((1+1)%4,0) = BLOCK (2,0) through memory channel C0; compute unit CU1 accesses BLOCK ((2+1)%4,1) = BLOCK (3,1) through memory channel C1; compute unit CU2 accesses BLOCK ((3+1)%4,2) = BLOCK (0,2) through memory channel C0; and so on. That is, in the third loop operation, the sub-blocks indicated by the oval marks in the tensor are read and written through the corresponding memory channels.
[0068] Figure 7B A schematic diagram of another data reading and writing method provided by at least one embodiment of the present disclosure is shown.
[0069] For example, in Figure 7B In the example, the data layout of the tensor is row-major order and the computing unit tasks are split in the column direction, that is, the first direction is the row direction and the second direction is the column direction. If the index of the computing unit is CU_idx, the coordinates of the target sub-block are (CU_idx, R%N), where R is obtained based on the index of the computing unit.
[0070] like Figure 7B As shown, since there are four memory channels, four loops are performed. In each loop, memory access by multiple compute units is balanced across the four memory channels. That is, through four loops, each compute unit sequentially accesses four memory channels. For example, in the first loop operation, the first value is the compute unit index, and the coordinates of the target sub-block are (CU_idx, CU_idx%N). That is, compute unit CU0 accesses BLOCK(0, 0) through memory channel C0; compute unit CU1 accesses BLOCK(1, 1) through memory channel C1; compute unit CU2 accesses BLOCK(2, 2) through memory channel C2; compute unit CU3 accesses BLOCK(3, 3) through memory channel C3; compute unit CU4 accesses BLOCK(4, 0) through memory channel C0, and so on. That is, in the first loop operation, the sub-blocks marked with dashed boxes in the tensor are read and written through the corresponding memory channels. For the second loop operation and subsequent loop operations, the coordinates of the target sub-block in the second direction of each loop operation are the remainder of the sum of the coordinates in the second direction of the previous loop operation and 1, plus N. That is, the target sub-block coordinates are (CU_idx, (CU_idx+1)%N). For example, for the second loop operation, the target sub-block coordinates are (CU_idx, (CU_idx+1)%N). For example, compute unit CU0 accesses BLOCK (0, (0+1)%4) = BLOCK (0, 1) through memory channel C0; compute unit CU1 accesses BLOCK (1, (1+1)%4) = BLOCK (1, 2) through memory channel C1; and so on. That is, in the second loop operation, the sub-blocks marked with solid lines in the tensor are read and written through the corresponding memory channels. For the third loop operation, computing unit CU0 accesses BLOCK (0, (1+1)%4,) = BLOCK (0, 2) through memory channel C0; computing unit CU1 accesses BLOCK (1, (2+1)%4,) = BLOCK (1, 3) through memory channel C1; and so on.
[0071] It's important to note that data reads originate from multiple compute units. For example, compute units CU0 and CU4 read the two blocks marked by dashed boxes, respectively. These two blocks are read through memory channel C0, while compute units CU1 and CU5 use memory channel C1. To clarify, the criterion for determining whether a memory channel conflict exists is whether the memory channels are fully utilized. Although there is contention between compute units CU0 and CU4, all memory channels are fully utilized, so there is no memory channel conflict.
[0072] Figure 8At least one embodiment of the present disclosure provides a Figure 6B Flowchart of the method for step S22 in FIG.
[0073] like Figure 8 As shown, the method includes steps S801 to S803.
[0074] Step S801: Based on the first granularity, the coordinates of the target sub-block of this loop operation are converted into the index position of the target tensor.
[0075] Step S802: Determine whether the index position exceeds the range of the target tensor.
[0076] Step S803: In response to the index position not exceeding the range of the target tensor, each computing unit performs a read and write operation on the target sub-block through the corresponding memory channel.
[0077] In GPU or GPGPU, the physical memory opened is larger than the logical memory opened by the software. For example, the logical memory opened is [22,256]. Due to the alignment requirements of the physical memory, the physical memory actually opened is [24, 256]. Figure 7A Therefore, after each computing unit determines the index position of the target tensor corresponding to the target sub-block, it is necessary to determine whether the index position exceeds the range of the target tensor. This can improve the accuracy of reading data and avoid additional read and write operations and errors.
[0078] For step S801, for example, during software programming, each thread can calculate the corresponding coordinates it reads through programming, and the data [4, 32] corresponding to a sub-block can just correspond to the number of threads executed in parallel by a computing unit. Figure 7A In the example, a total of [4×6, 32×8] physical memory is opened up, that is, 6 rows and 8 columns of sub-blocks, each row of sub-blocks contains 4 rows and 32 columns of data, then the data logical coordinates corresponding to BLOCK(0, 3) (that is, the index position of the target tensor) are the data in the rectangle formed by [96, 0], [128, 0], [96, 4] and [128, 4].
[0079] In step S802, when a thread calculates that the index position of the target tensor it will read exceeds the range of logical memory (that is, the size of the target tensor), the target tensor is out of range. For example, if logical memory has data of size [4, 20], when a thread calculates that the index position of the target tensor it will read is [4, 23], the target tensor is out of range.
[0080] Regarding step S803 , when the index position does not exceed the range of the target tensor, each computing unit performs a read and write operation on the target sub-block through the corresponding memory channel.
[0081] Figure 9 At least one embodiment of the present disclosure provides a Figure 6A Flowchart of the method for step S20 in FIG.
[0082] like Figure 9 As shown, step S20 includes steps S901 to S904.
[0083] Step S901: Based on the size of the target tensor, determine a first maximum value of a plurality of sub-block coordinate values in a first direction and a second maximum value in a second direction.
[0084] Step S102: starting from the sub-block with the smallest coordinate value in the first direction and the second direction, perform a loop operation.
[0085] Step S903: After the loop operation is completed, the coordinate value in the second direction is increased by M each time, and then the loop operation is performed until the coordinate value in the second direction reaches a maximum value.
[0086] Step S904: After the coordinate value in the first direction increases by N each time, a loop operation is performed until the coordinate value in the first direction reaches a maximum value, and then all the data in the tensor are read and written.
[0087] For step S901, according to the size of the target tensor, the number of sub-blocks in the first direction and the second direction is determined, thereby determining the maximum value of the sub-block coordinate value in the first direction (i.e., the first maximum value) and the maximum value of the sub-block coordinate value in the second direction (i.e., the second maximum value).
[0088] For example, if the size of the target tensor is [H, W], the size of each sub-block is [4, 32], and the coordinate values of the sub-blocks start from [0, 0] and increase with a step size of 1, then the number of sub-blocks in the first direction is H / 4 rounded up, and the first maximum value is (H / 4 rounded up - 1), and the number of sub-blocks in the second direction is W / 32 rounded up, and the second maximum value is (W / 32 rounded up - 1).
[0089] For step S902, the sub-block with the smallest coordinate value is BLOCK(0, 0), and the second granularity is performed starting from the sub-block BLOCK(0, 0). Figure 6B The cycle operation shown.
[0090] For steps S903 and S904, after the loop operation is completed, the next second granularity is executed. For example, the second granularity is first traversed in the row direction and then in the column direction.
[0091] For example, the first maximum value is represented as block_n, the second maximum value is represented as block_m, and the coordinates of the sub-block are represented as [block_n_id, block_m_id]. Figure 10 The method shown can be implemented using a For loop. For example, the code can be schematically represented as follows:
[0092] For (block_n_id = 0; block_n_id < block_n; block_n_id += N)
[0093] { for (block_m_id = 0; block_m_id < block_m; block_m_id += M)
[0094] { implement Figure 6B The loop operation shown is
[0095] In some embodiments of the present disclosure, when the size of the first granularity in the second direction is larger than the number of threads processed by each computing unit at one time, each computing unit performs read and write operations on the target sub-block through the corresponding memory channel, including: dividing the first granularity into multiple sub-target sub-blocks in the second direction according to the number of threads, and performing read and write operations on the multiple sub-target sub-blocks in sequence.
[0096] For example, if the tensor data layout is linear, that is, the tensor data is arranged sequentially along a single dimension. For example, data with sequence numbers [0, 127] is arranged in data storage unit S0, data with sequence numbers [128, 255] is arranged in data storage unit S1, and so on. If the data layout is linear, then the row-wise data volume corresponding to a data storage unit is larger, typically corresponding to data of size [1, block_w], where block_w is the number of data included in the first granularity, for example, 128. If the number of threads executed in parallel by a computation unit is 32, then a data storage unit, or a sub-block, needs to be executed multiple times (for example, four times). Therefore, the computation unit needs to align the starting position of data read and write in the W direction to block_w, that is, the computation unit needs to loop in the row direction until block_w is processed. For example, if the number of threads is 32 and the first granularity is [1, 128], then in the row direction, the target sub-block is divided into four sub-target sub-blocks, and read and write operations are performed on each of the four sub-target sub-blocks in sequence. This embodiment is similar to the above Figure 7B The embodiment described in Figure 7B Based on the embodiment of the present invention, it is necessary to perform 4 cycles of read and write operations in each sub-block until the read and write of the sub-block is completed.
[0097] The data reading and writing methods provided by the embodiments of this disclosure adapt to the hardware characteristics of GPGPU products by interleaving the memory access locations of computing units and balancing the use of memory channels. This method fully utilizes the bandwidth of on-chip high-speed storage and improves overall operator performance. The above embodiments of this disclosure are applied to the unpermutation process. For example, the input tensor is a matrix tensor of [B, H, W], where B represents the batch size, H represents the number of tokens × topk, where topk represents the number of experts processing a token, and W represents the length of the vectorized representation of a token. Assuming the number of tokens is 4096 and topk=2, the shape of the input tensor is [1, 8192, 8192], the shape of the output tensor is [1, 4096, 8192], and the data precision is brain floating point 16 (bf16). Experimental verification shows that if traditional read and write methods are used to store memory channel conflicts, it takes 2.4ms to complete unpermutation. If the read and write methods provided by the embodiments of the present disclosure are used, there is no memory channel conflict and it takes 1.0ms to complete unpermutation, thus improving performance by 2.4X.
[0098] It should be noted that any embodiment of the present disclosure can be adaptively optimized for specific scenarios such as different HBM architectures and different numbers of computing units, and is generalizable; it is universally applicable to some products of GPGPU video memory including HBM, GDDR, and DDR.
[0099] Figure 10 A schematic diagram of a data reading and writing method provided by at least one embodiment of the present disclosure is shown.
[0100] like Figure 10 As shown in the figure, the data of the target tensor is arranged in the memory in an interleaved manner of [1, 128] along the row direction. The tasks of the 8 computing units are divided along the column direction of the target tensor. The 8 computing units read and write data from the memory through 4 memory channels to perform computing tasks. The number of threads executed in parallel by the computing units is 32.
[0101] A block with M rows and N columns is an inner loop unit, such as Figure 11The dashed box represents an inner loop unit. The following loop operations are performed in the inner loop unit: 1) the initial block processed by each calculation unit cu_idx is block(cu_idx, cu_idx % N); 2) each calculation unit performs the block (cu_idx, (cu_idx + 1) % N) each time; 3) convert the target sub-block into the index position of the target tensor; 4) in block_w, loop through the data, judge whether the index position exceeds the range of the target tensor, and if not, perform read-write operations. 5) perform 2), 3), and 4) N times until the MxN sub-blocks are processed. For a complete [H, W] tensor, the above For loop is implemented.
[0102] In some embodiments of the present disclosure, the data of the target tensor is arranged in the memory in the [1, 128] staggered direction in the row direction, and the task division of the 8 calculation units is in the row direction of the target tensor. In this case, the data arrangement direction and the calculation unit division direction are the same, and the starting coordinates read by each calculation unit will not conflict as long as they are aligned to block_w. In this case, for a [H, W] tensor, the read-write operation can be performed according to the following steps: 1) calculate the total number of sub-blocks blocks_num; 2) perform the following For loop. For (block_id = cu_idx; block_id < blocks_num; block_id += N){ convert block_id into an index position; perform a block_w loop in a target sub-block}. For (block_id = cu_idx; block_id < blocks_num; block_id += N), which means that the sub-block coordinate block_id is declared and initialized to cu_idx, and if block_id is less than blocks_num, then block_id is assigned to block_id + N.
[0103] At least one embodiment of the present disclosure also provides a data read-write device. Figure 11 A schematic block diagram of a data read-write device provided by at least one embodiment of the present disclosure is shown.
[0104] As Figure 11As shown, the data reading and writing device 100 includes a coordinate determination module 110 and a loop module 120. The device 100 is used for M computing units to read and write a target tensor from a memory, wherein the data of the target tensor is interleaved in the memory at a first granularity along a first direction of the target tensor, the tasks of the M computing units are divided along a second direction of the target tensor, and the M computing units read and write the data from the memory through N memory channels to perform computing tasks.
[0105] The coordinate determination module 110 is configured to perform coordinate representation on the multiple sub-blocks obtained by dividing the target tensor according to the first granularity.
[0106] The loop module 120 performs read and write operations on the plurality of sub-blocks at a second granularity, and performs N loop operations for each second granularity until all data at the second granularity is read and written.
[0107] The loop module 120 includes a target sub-block determining unit 121 and a reading and writing unit 122 .
[0108] The target sub-block determination unit 121 is configured to calculate the target sub-block among the multiple sub-blocks of each computing unit in this loop operation based on the coordinate representation and the index of each computing unit, so that the access of the M computing units to the memory is balanced to the N memory channels.
[0109] The read / write unit 122 is configured such that each computing unit performs read / write operations on the target sub-block through a corresponding memory channel, where N and M are both positive integers.
[0110] For example, the coordinate determination module 110 and the loop module 120 include codes and programs stored in a memory. The coordinate determination module 110 and the loop module 120 are implemented as, for example, a central processing unit (CPU) or other forms of processing units with data reading and writing capabilities and / or instruction execution capabilities. The processing unit can be a general-purpose processor, and can also be a single-chip microcomputer, a microprocessor, a digital signal processor, a dedicated image processing chip, or a field programmable logic array, etc. The coordinate determination module 110 and the loop module 120 execute the codes and programs to implement some or all of the functions of the coordinate determination module 110 and the loop module 120 as described above. For example, the coordinate determination module 110 and the loop module 120 can be a circuit board or a combination of multiple circuit boards for implementing the functions as described above. In an embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processors; and (3) firmware stored in the memories that can be executed by the processors.
[0111] It should be noted that the coordinate determination module 110 and the circulation module 120 can be used to implement Figure 6A The coordinate determination module 110 and the loop module 120 are shown in steps S10 and S20; therefore, for a detailed description of the functions that can be implemented by the coordinate determination module 110 and the loop module 120, reference can be made to the description of steps S10 and S20 in the embodiment of the above-mentioned data reading and writing method. In addition, the data reading and writing device 100 can achieve similar technical effects as the above-mentioned data reading and writing method, which will not be described in detail here.
[0112] It should be noted that in at least one embodiment of the present disclosure, the data reading and writing device 100 may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or can be constructed in other applicable ways.
[0113] For example, the data reading and writing device 100 may be implemented in hardware, software, or a combination of hardware and software, and this disclosure does not impose any specific limitations on this.
[0114] The data reading and writing method provided by at least one embodiment of the present disclosure can achieve technical effects similar to the data reading and writing method described above, and will not be described in detail here.
[0115] The data reading and writing method and data reading and writing device provided by at least one embodiment of the present disclosure can be applied to different systems or devices, such as Figure 13 The electronic device 300 shown. The electronic device 300 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR device, a VR device, a vehicle-mounted terminal, etc., and can also be a server, etc. The data reading and writing method provided in at least one embodiment of the present disclosure can be applied to scenarios involving convolution operations such as CPU, high-performance computing (HPC) and artificial intelligence (AI) in the electronic device 300, such as tensor computing units, etc. Of course, the present disclosure is not limited to this, and any scenario, device, apparatus, etc. involving backpropagation calculations can adopt the data reading and writing method or data reading and writing apparatus provided in at least one embodiment of the present disclosure.
[0116] In some embodiments, the data reading and writing device provided in at least one embodiment of the present disclosure may be a chip, for example, a system-on-a-chip (SoC). The SoC includes a processor, which may be a single-core processor or a multi-core processor, a memory, and an I / O interface. The processor may load data and applications from the memory and then process the data, for example, performing backpropagation calculations.
[0117] It should be noted that the target tensor of the data reading and writing method or data reading and writing device provided in at least one embodiment of the present disclosure may have different specific physical meanings depending on the application scenario. For example, the data reading and writing method provided in at least one embodiment of the present disclosure may be applied in fields such as speech processing, image processing, text processing, and video processing.
[0118] For example, in the field of speech processing, the target tensor can be the input and output parameters of convolution operations involved in tasks such as feature extraction, speech enhancement, and speech recognition.
[0119] For example, in the field of image processing, the target tensor can be the input and output parameters of convolution operations involved in tasks such as image recognition, feature extraction, image segmentation, target detection, image classification, and scene reconstruction.
[0120] For example, in the field of text processing, the target tensor can be the input and output parameters of convolution operations involved in tasks such as text classification, sentiment analysis, and text generation.
[0121] For example, in the field of video processing, the target tensor can be the relevant parameters in the image processing field as mentioned above, or the input and output parameters of convolution operations unique to the video processing field, such as the optical flow operator (used to estimate motion between video frames) and the target tracking operator (used to track specific targets in the video).
[0122] Of course, the present disclosure is not limited to this. For other application scenarios or fields, as long as convolution operations are required, the data reading and writing method described in at least one embodiment of the present disclosure can be applied, and they will not be described one by one here.
[0123] Figure 12 A schematic structural diagram of a processor provided in at least one embodiment of the present disclosure. Figure 12 As shown, the processor 200 includes an instruction parsing unit 201 and an execution unit 202 .
[0124] For example, the instruction parsing unit 201 is used to receive and parse data read and write instructions, which are used by the M computing unit to read and write target tensors from the memory. After the instruction parsing unit parses the data read and write instructions, the execution unit 202 executes the data read and write method provided by any embodiment of the present disclosure.
[0125] For example, the processor 200 may use the aforementioned Figure 1 The graphics processor or general graphics processor architecture shown.
[0126] For example, after receiving a data read or write instruction, the processor parses the data read or write instruction, such as decoding the data read or write instruction, generates a microinstruction, and sends the microinstruction to the instruction dispatch unit; the instruction dispatch unit sends the microinstruction to the corresponding scheduling queue according to the microinstruction category; the scheduler in the processor selects the appropriate microinstruction from the scheduling queue for execution based on the priority, dependency, and other factors of the microinstruction; in response to the microinstruction, when all or the required input data (such as activation tensors and convolution kernel tensors) are prepared, the execution unit reads the data and executes the relevant operations of the data read or write instruction.
[0127] For example, the execution unit 202 may include hardware involved in executing data read and write instructions, such as a multiplier, a register, a queue, an ALU (Arithmetic Logic Unit), and the like.
[0128] For example, when the execution unit 202 executes a data read and write instruction, it includes performing coordinate representation on the multiple sub-blocks obtained by dividing the target tensor according to the first granularity; performing read and write operations on the multiple sub-blocks at the second granularity, and performing the following loop operation N times for each second granularity until all the data of the second granularity is read and written: based on the coordinate representation and the index of each computing unit, calculating the target sub-block among the multiple sub-blocks of this loop operation for each computing unit, so that the access of the M computing units to the memory is balanced to the N memory channels; each computing unit performs read and write operations on the target sub-block through the corresponding memory channel, and N and M are both positive integers.
[0129] Regarding the specific process of using the execution unit 202 to execute the data read and write instructions, reference may be made to steps S10-S20 and other related contents in the aforementioned data read and write method, and repeated details will be omitted.
[0130] The processor provided by at least one embodiment of the present disclosure can achieve technical effects similar to those of the aforementioned data reading and writing method, and the repeated parts will not be repeated.
[0131] Figure 13 This is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 13 As shown, the electronic device 300 is suitable for implementing the data reading and writing method provided by the embodiment of the present disclosure. It should be noted that Figure 13 The components of the electronic device 300 shown are merely exemplary and non-limiting. The electronic device 300 may also have other components according to actual application requirements.
[0132] like Figure 13As shown, the electronic device 300 may include a processing device 301 (eg, a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in a memory to implement various functions.
[0133] For example, when the computer readable instructions are executed by the processing device 301, one or more steps of the data reading and writing method according to any of the above embodiments can be executed. It should be noted that for a detailed description of the processing process of the data reading and writing method, reference can be made to the relevant description in the embodiments of the above data reading and writing method.
[0134] For example, the memory may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory. For example, computer-readable instructions may be loaded from storage device 308 into RAM 303 to execute the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, and the like. Various applications and various data, such as style images and various data used and / or generated by the applications, may also be stored in the computer-readable storage media.
[0135] For example, the processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output I / O interface 305 is also connected to the bus 304.
[0136] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other electronic devices wirelessly or by wire to exchange data. Figure 13While an electronic device 300 is shown with various devices, it should be understood that implementation or presence of all illustrated devices is not required, and the electronic device 300 may alternatively implement or possess more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions. The processing device 301 may be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit (GPU), etc., which has data reading and writing capabilities and / or program execution capabilities. The central processing unit (CPU) may be of an X86, ARM, or RISC-V architecture. The GPU may be directly integrated into the SOC, directly integrated into the motherboard, or built into the motherboard's northbridge chip.
[0137] Figure 14 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. Figure 14 As shown, the storage medium 400 may be a non-transitory computer-readable storage medium, and one or more computer-executable instructions 401 may be non-transitory stored on the storage medium 400. For example, when the computer-executable instructions 401 are executed by a processor, one or more steps in the data reading and writing method described above may be performed.
[0138] For example, the storage medium 400 may be applied to the electronic device 300 . For example, the storage medium 400 may include the storage device 308 in the electronic device 300 .
[0139] For example, the storage device may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer-executable instructions may be stored on the computer-readable storage medium, and the processor may execute the computer-executable instructions to implement various functions of the processor. The storage medium may also store various application programs and various data.
[0140] For example, the storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other applicable storage media.
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0142] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0143] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0144] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0145] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0146] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
[0147] Regarding this disclosure, the following points need to be explained:
[0148] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0149] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0150] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A data reading and writing method for M computing units to read and write target tensors from memory, wherein: The data of the target tensor is interleaved in the memory at a first granularity along a first direction of the target tensor, the tasks of the M computing units are divided along a second direction of the target tensor, and the M computing units read and write the data from the memory through N memory channels to perform computing tasks, The method comprises: Performing coordinate representation on a plurality of sub-blocks obtained by dividing the target tensor according to the first granularity; Performing read and write operations on the multiple sub-blocks at a second granularity, wherein the following loop operation is performed N times for each second granularity until all data at the second granularity is read and written: Determine, based on the coordinate representation and the index of each computing unit, a target sub-block among the multiple sub-blocks to be looped by each computing unit, so that access to the memory by the M computing units is balanced across the N memory channels; Each computing unit uses the corresponding memory channel to perform read and write operations on the target sub-block. Wherein, N and M are both positive integers.
2. The method according to claim 1, wherein Calculating a target sub-block among the multiple sub-blocks of the current loop operation of each computing unit based on the coordinate representation and the index of the computing unit includes: using a remainder of the first value and N as the coordinate of the target sub-block in the second direction, and an index of the calculation unit as the coordinate of the target sub-block in the first direction, wherein the first value is obtained based on the index of the calculation unit; and The target sub-block is determined based on the coordinate representations of the multiple sub-blocks.
3. The method according to claim 2, wherein: For a first loop operation in N loop operations, the first value is an index of the computing unit; For the loop operation after the first loop operation, the first value is the sum of the previous first value and 1.
4. The method according to claim 1, wherein Each computing unit uses a corresponding memory channel to perform read and write operations on the target sub-block, including: Based on the first granularity, convert the coordinates of the target sub-block in the current loop operation into an index position of the target tensor; Determine whether the index position exceeds the range of the target tensor; In response to the index position not exceeding the range of the target tensor, each computing unit performs a read and write operation on the target sub-block through a corresponding memory channel.
5. The method according to claim 1, wherein The first direction is a row direction, the second direction is a column direction, and the second granularity is [M, N]; or The first direction is a column direction, the second direction is a row direction, and the second granularity is [N, M].
6. The method according to claim 1, wherein Performing read and write operations on the plurality of sub-blocks at the second granularity includes: Determining, based on a size of the target tensor, a first maximum value of the plurality of sub-block coordinate values in the first direction and a second maximum value in the second direction; Starting from the sub-block with the smallest coordinate value in the first direction and the second direction, performing the loop operation; After the loop operation is completed, the coordinate value in the second direction is increased by M each time, and the loop operation is performed until the coordinate value in the second direction reaches the maximum value; After the coordinate value in the first direction increases by N each time, the loop operation is performed until the coordinate value in the first direction reaches the maximum value, and then all the data in the tensor is read and written.
7. The method according to claim 1, wherein When the size of the first granularity in the second direction is greater than the number of threads processed by each computing unit at one time, each computing unit performs a read and write operation on the target sub-block through a corresponding memory channel, including: Divide the first granularity into a plurality of sub-target sub-blocks in the second direction according to the number of threads, Read and write operations are performed on the multiple sub-target sub-blocks in sequence.
8. The method according to claim 1, wherein The second granularity is determined according to the number of computing units, the number of memory channels, the task division direction of the computing units, and the data arrangement.
9. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is used to receive and parse data read and write instructions, and the data read and write instructions are used by the M computing unit to read and write target tensors from the memory; as well as The execution unit executes the data reading and writing method according to any one of claims 1 to 8 after the instruction parsing unit parses the data reading and writing instruction.
10. An electronic device comprising: a memory that non-transitorily stores computer-executable instructions; a processor configured to execute the computer-executable instructions, Wherein, when the computer executable instructions are executed by the processor, the data reading and writing method according to any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data reading and writing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Storage scheduling method and device for data among multiple kinds of storages
CN104375895A
Method and device for optimizing tensor calculation performance
CN116775277A
Hardware resource allocation method, electronic equipment and storage medium
CN118312327A
Parallel computing method and system of tensor protocol and computer equipment
CN119718687A
Processor, chip product, computer equipment and tensor calculation method
CN119883375A
Cited By
Data processing method and device, computer equipment, readable storage medium and program product
CN121210159A