Tensor load store method, device and storage medium
By breaking down tensor loading and storage procedures into sub-tensor operations and combining them with the memory data layout of basic processing units and thread-level parallel architecture, the problem of low efficiency in tensor loading and storage is solved, thereby improving hardware resource utilization and overall computing performance.
Patent Information
- Application Number
- CN202511247922.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-02
AI Technical Summary
In high-performance computing and artificial intelligence accelerators, the loading and storage operations of tensors are inefficient, resulting in low utilization of hardware resources and affecting overall computing performance.
The tensor loading and storage process is broken down into sub-loading and storage processes of multiple sub-tensors. Based on the memory data layout and hardware attributes of pre-encapsulated basic processing units, the instruction set is optimized, and a thread-level parallel architecture is used for data loading and storage.
It improves the utilization of hardware resources and the performance of tensor loading and storage, especially significantly improving the efficiency of data loading and storage when dealing with large-scale tensors.
Smart Images

Figure CN120762600B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of artificial intelligence chip, and in particular to a tensor loading and storing method, device and storage medium. BACKGROUND
[0002] In high-performance computing and artificial intelligence accelerators, data movement usually consumes more resources and time than the computation itself. The Load and Store operations of tensors, as the bridge between the computation core and the memory, directly affect the overall computing performance.
[0003] In the related art, simple memory access instructions are used to directly read and write tensor data, which causes some hardware resources in artificial intelligence to be underutilized, resulting in low utilization of chip hardware resources and affecting the performance of tensor loading and storing. SUMMARY
[0004] Embodiments of the present application provide a tensor loading and storing method, device and storage medium, which are used to improve the utilization of hardware resources and the performance of tensor loading and storing.
[0005] In one aspect, the present application provides a tensor loading and storing method applied to an artificial intelligence chip, which comprises:
[0006] Based on a pre-packaged basic processing unit, the process of loading and storing a target tensor is split into a plurality of sub-tensor corresponding sub-loading and storing processes; each sub-loading and storing process comprises the following steps:
[0007] Based on a first memory data layout of the basic processing unit, the destination physical address of each data item in the sub-tensor is obtained; the first memory data layout is adapted to the hardware properties of the artificial intelligence chip;
[0008] For each data item, based on the destination physical address and the logical layout of the source tensor containing the data item, the source physical address of the data item is obtained;
[0009] The data item is loaded from the source physical address and stored in the destination physical address.
[0010] In one aspect, the present application provides a tensor loading and storing device applied to an artificial intelligence chip, which comprises:
[0011] A splitting module is configured to split, based on a pre-packaged basic processing unit, the process of loading and storing a target tensor into a plurality of sub-tensor corresponding sub-loading and storing processes;
[0012] An execution module is configured to execute each sub-loading and storing process, specifically performing the following steps:
[0013] obtain a destination physical address of each data item in the sub-tensor based on a first memory data layout of the basic processing unit; the first memory data layout is adapted to hardware attributes of the artificial intelligence chip;
[0014] for each data item, obtain a source physical address of the data item based on the destination physical address and a logical layout of a source tensor containing the data item;
[0015] load the data item from the source physical address and store the data item in the destination physical address.
[0016] Optionally, the method further comprises an instruction optimization module;
[0017] The instruction optimization module is specifically configured to:
[0018] pre-package a set of original instructions of the basic processing unit;
[0019] optimize the set of original instructions through an expression system to obtain a set of target instructions, the set of target instructions being used to execute each of the sub-load-store processes.
[0020] Optionally, the execution module is specifically configured to:
[0021] for each data item, obtain a target physical index corresponding to the data item based on a first memory data layout of the basic processing unit;
[0022] map the destination physical address of the data item based on the target physical index.
[0023] Optionally, the target physical index comprises a register memory index and a thread memory index; and the destination physical address comprises a register memory address and a thread memory address.
[0024] The execution module is specifically configured to:
[0025] map the register memory address of the data item based on the register memory index;
[0026] map the thread memory address of the data item based on the thread memory index.
[0027] Optionally, each of the plurality of sub-tensors corresponds to a basic processing unit.
[0028] The execution module is further configured to:
[0029] The first memory data layout of the basic processing unit is combined with a target arrangement mode of the basic processing unit corresponding to each of the plurality of sub-tensors to obtain a second memory data layout of the target tensor.
[0030] Optionally, the execution module is specifically configured to:
[0031] Based on the target physical address, the second memory data layout, and a logical layout of the source tensor, a target logical address of the data item is obtained.
[0032] Based on the target logical address of the data item, an original logical address of the data item in the source tensor is mapped and obtained.
[0033] Based on the original logical address of the data item, a source physical address of the data item in the original memory is mapped and obtained.
[0034] Optionally, the target physical address includes a register memory address and a thread memory address.
[0035] The execution module is specifically configured to:
[0036] Based on the register memory address, the second memory data layout, and the logical layout of the source tensor, a register logical address of the data item is obtained.
[0037] Based on the thread memory address, the second memory data layout, and the logical layout of the source tensor, a thread logical address of the data item is obtained.
[0038] The register logical address and the thread logical address are accumulated to obtain the target logical address of the data item.
[0039] In an aspect, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored on the memory and running on the artificial intelligence chip, and the artificial intelligence chip executes the computer program to implement the steps of the tensor loading and storing method.
[0040] In an aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program executable by a computer device, and when the computer program runs on the computer device, the computer device executes the steps of the tensor loading and storing method.
[0041] In an aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer device, the computer device executes the steps of the tensor loading and storing method.
[0042] In the embodiment of the present application, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-tensor corresponding sub-loading and storing processes; the first memory data layout of the basic processing unit is adapted to the hardware properties of the artificial intelligence chip, that is, the first memory data layout of the basic processing unit meets the hardware alignment granularity of the artificial intelligence chip, and at the same time, when the data is loaded and stored according to the first memory data layout of the basic processing unit for processing, the resource (such as bandwidth) usage of the artificial intelligence chip reaches the upper limit value, therefore, when the basic processing unit is used as the minimum data unit for loading and storing tensor data, it is more consistent with the hardware memory structure of the artificial intelligence chip, which not only makes full use of the hardware resources of the chip and improves the utilization rate of the hardware resources, but also improves the performance of tensor loading and storing; secondly, a plurality of threads can perform loading and storing operations in parallel, and this thread-level parallel architecture based on the basic processing unit can significantly improve the efficiency of data loading and storing, and the effect is more significant when processing large-scale tensors. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0044] Figure 1 A structural schematic diagram of an artificial intelligence chip is provided for the embodiment of the present application;
[0045] Figure 2 A flowchart of a tensor loading and storing method is provided for the embodiment of the present application;
[0046] Figure 3A A schematic diagram of a data layout is provided for the embodiment of the present application;
[0047] Figure 3B A schematic diagram of a data layout is provided for the embodiment of the present application;
[0048] Figure 4A A schematic diagram of a physical memory data layout is provided for the embodiment of the present application;
[0049] Figure 4B A schematic diagram of a physical memory data layout is provided for the embodiment of the present application;
[0050] Figure 5 A schematic diagram of a physical memory data layout is provided for the embodiment of the present application;
[0051] Figure 6A schematic diagram of a physical memory data layout provided for an embodiment of the present application;
[0052] Figure 7 A flowchart of an address mapping method provided for an embodiment of the present application;
[0053] Figure 8 A flowchart of an address mapping method provided for an embodiment of the present application;
[0054] Figure 9 A structural schematic diagram of a tensor load storage device provided for an embodiment of the present application;
[0055] Figure 10 A structural schematic diagram of a computer device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the objectives, technical solutions and beneficial effects of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0057] The terms "first", "second", "third", etc. in the text are only for the purpose of description, and cannot be understood as explicitly or implicitly indicating relative importance or implicitly indicating the number of the indicated technical features.
[0058] Reference Figure 1 It is a structural diagram of an artificial intelligence chip applicable to an embodiment of the present application, which comprises at least a display memory 101 and a plurality of computing unit memories 102. Each computing unit memory 102 comprises a plurality of execution unit memories 103; each execution unit memory 103 comprises an on-chip cache 104 and a plurality of register memories 105; and each register memory 105 comprises a plurality of thread memories.
[0059] The display memory 101 can be a high bandwidth memory (HBM) or other types of memory. The on-chip cache 104 is a temporary memory, which has a smaller capacity than the display memory 101 but a faster data exchange speed than the display memory 101.
[0060] In an embodiment of the present application, a basic processing unit (Trait) is pre-packaged, which is the smallest data unit of tensor segmentation. The first memory data layout (i.e. the physical memory data layout) of the basic processing unit refers to a memory data arrangement mode arranged according to certain rules to meet hardware requirements.
[0061] The first memory data layout of the basic processing unit is adapted to the hardware attribute of the artificial intelligence chip, that is, the first memory data layout of the basic processing unit meets the basic requirements of the hardware alignment granularity of the artificial intelligence chip, and meanwhile, when the artificial intelligence chip processes (such as loads and stores) the basic processing unit according to the first memory data layout, the resource (such as bandwidth) usage of the artificial intelligence chip reaches the upper limit value, such as "full bandwidth".
[0062] In the embodiment of the present application, the source tensor is stored in the display memory 101, and the data in the source tensor is loaded from the display memory 101 to the register memory 105 to obtain the target tensor; the target tensor is all or part of the data of the source tensor.
[0063] In a specific implementation, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-loading and storing processes corresponding to the respective sub-tensors; each sub-loading and storing process includes the following steps:
[0064] Based on the first memory data layout of the basic processing unit, the destination physical address of each data item in the target storage space in the sub-tensor is obtained. In actual application, each data item corresponds to a physical index in the first memory data layout, and each physical index can be converted into a destination physical address. Based on the destination physical address, the source physical address of each data item in the display memory 101 is calculated, and the data item is loaded from the source physical address to the destination physical address.
[0065] In the embodiment of the present application, the artificial intelligence chip 100 includes a plurality of physical memory levels; such as thread memory, register memory, execution unit memory, and calculation unit memory; the data layouts of the plurality of physical memory levels are combined to obtain the first memory data layout of the basic processing unit, wherein the data layout of each physical memory level is adapted to the hardware attribute of the artificial intelligence chip.
[0066] In some embodiments, the first memory data layout of the basic processing unit is obtained by combining the register memory data layout and the thread memory data layout. The register memory data layout refers to the arrangement manner of the data stored in the plurality of register memories 105; and the thread memory data layout refers to the arrangement manner of the data saved in the plurality of thread memories.
[0067] In some embodiments, the first memory data layout of the basic processing unit is obtained by combining the execution unit memory data layout, the register memory data layout, and the thread memory data layout. The execution unit memory data layout refers to the arrangement manner of the data stored in the plurality of execution unit memories 103; the register memory data layout refers to the arrangement manner of the data stored in the plurality of register memories 105; and the thread memory data layout refers to the arrangement manner of the data saved in the plurality of thread memories.
[0068] Of course, the basic processing unit can also be in other forms, and the present application does not make specific limitations.
[0069] Since the first memory data layout of the basic processing unit is adapted to the hardware properties of the artificial intelligence chip 100, that is, the first memory data layout of the basic processing unit meets the hardware alignment granularity of the artificial intelligence chip 100, and at the same time, when loading and storing data according to the first memory data layout of the basic processing unit, the resource (such as bandwidth) usage of the artificial intelligence chip 100 reaches the upper limit value, such as "full bandwidth", therefore, when loading and storing tensor data with the basic processing unit as the minimum data unit, the underlying hardware resources can be fully utilized; and multiple threads can perform loading and storing operations in parallel, and such a thread-level parallel architecture based on the basic processing unit can significantly improve the efficiency of data loading and storing, and the effect is more significant when processing large-scale tensors.
[0070] The artificial intelligence chip 100 in the present application can include other structures in addition to the above-mentioned structure, and the present application does not make specific limitations.
[0071] The artificial intelligence chip 100 can be a graphics processing unit (GPU), a general-purpose computing on graphics processing units (GPGPU), a domain specific architecture (DSA), etc.
[0072] The present application is based on Figure 1 The architecture diagram of the artificial intelligence chip is shown, and a tensor loading and storing method is provided, which can be applied to various scenes. For example, image processing scene, speech processing scene, text processing scene, etc. In different application scenarios, the physical meaning of the source tensor and the target tensor can be different.
[0073] For example, in the text processing scene, the source tensor and the target tensor can be text data used in text generation, text recognition, etc.
[0074] For example, in the speech processing scene, the source tensor and the target tensor can be speech data used in speech enhancement, speech recognition, speech synthesis, etc.
[0075] For example, in the image processing scene, the source tensor and the target tensor can be image data used in image preprocessing, image segmentation, target detection, etc.
[0076] The flow of the tensor loading and storing method will be introduced in detail below, please refer to Figure 2The method is executed by an artificial intelligence chip, and the method comprises the following steps:
[0077] In step 201, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-tensor corresponding sub-loading and storing processes.
[0078] Specifically, the basic processing unit is the smallest data unit of tensor splitting. The first memory data layout of the basic processing unit refers to a memory data arrangement mode arranged in a certain rule meeting the hardware requirements. The first memory data layout of the basic processing unit is the physical memory data layout of the basic processing unit.
[0079] The first memory data layout of the basic processing unit is adapted to the hardware attribute of the artificial intelligence chip, that is, the first memory data layout of the basic processing unit meets the basic requirements of the hardware alignment granularity of the artificial intelligence chip, and at the same time, when the artificial intelligence chip loads and stores data according to the first memory data layout, the resource (such as bandwidth) usage of the artificial intelligence chip reaches the upper limit value, such as "full bandwidth".
[0080] For different types of artificial intelligence chips, the hardware attribute of the artificial intelligence chip is different; accordingly, the first memory data layout of the basic processing unit packaged according to the hardware attribute of the artificial intelligence chip is also different.
[0081] The first memory data layout of the basic processing unit is obtained by combining a plurality of physical memory level data layouts, and each physical memory level data layout comprises a plurality of dimensions; for example, the physical memory level data layout can be a matrix arrangement mode, in which the row (W) dimension priority or the column (H) dimension priority can be further set.
[0082] The data layout is explained below.
[0083] The data layout (layout) can be characterized by shape and stride; wherein, the shape refers to the length in each dimension; the stride refers to the distance that needs to be moved in the storage space when accessing the next element along a dimension.
[0084] For example, referring to Figure 3A In the W dimension priority case, the elements of the first row are accessed first, then the elements of the second row are accessed, and so on.
[0085] At this time, the data layout is represented as (sh(h, w), st(w, 1)); where sh(h, w) represents a shape, i.e., a length of h in the H dimension and a length of w in the W dimension; and st(w, 1) represents a stride, i.e., a distance of w in the storage space when accessing a next element in the H dimension (when accessing a next element in the H dimension, the current element and the next element are located in different rows, and in the W dimension priority case, the row where the current element is located needs to be accessed first before the row where the next element is located, and therefore, the distance in the storage space is w instead of 1). When accessing a next element in the W dimension, the distance in the storage space is 1.
[0086] For example, referring to Figure 3B In the H dimension priority case, the elements in the first column are accessed first, and then the elements in the second column are accessed, and so on.
[0087] At this time, the data layout is represented as (sh(h, w), st(w, 1)); where sh(h, w) represents a shape, i.e., a length of h in the H dimension and a length of w in the W dimension; and st(w, 1) represents a stride, i.e., a distance of w in the storage space when accessing a next element in the H dimension (when accessing a next element in the H dimension, the current element and the next element are located in different rows, and in the W dimension priority case, the row where the current element is located needs to be accessed first before the row where the next element is located, and therefore, the distance in the storage space is w instead of 1). When accessing a next element in the W dimension, the distance in the storage space is 1.
[0088] In some embodiments, data layouts of a plurality of physical memory levels are obtained; the data layout of each physical memory level is adapted to a hardware attribute of the artificial intelligence chip; and the data layouts of the plurality of physical memory levels are combined to obtain a first memory data layout of the basic processing unit.
[0089] Specifically, a shape and a stride of each data layout of the plurality of physical memory levels are obtained; and the shapes of the data layouts of the plurality of physical memory levels are combined in a descending order of priority to obtain a target shape of the first memory data layout of the basic processing unit.
[0090] From the data layouts of the plurality of physical memory levels, a data layout with the highest priority is obtained; and then the strides under each of the other data layouts are converted into strides under the data layout with the highest priority. The stride of the data layout with the highest priority and the converted other strides are combined in a descending order of priority to obtain a target stride of the first memory data layout of the basic processing unit.
[0091] The target shape and the target stride are combined to obtain the first memory data layout of the basic processing unit.
[0092] In some embodiments, the data layout of the plurality of physical memory levels comprises: a thread memory data layout and a register memory data layout; the thread memory data layout and the register memory data layout are adapted to the hardware properties of the artificial intelligence chip. The thread memory data layout and the register memory data layout are combined to obtain the first memory data layout of the basic processing unit.
[0093] Specifically, a first shape and a first step size of the thread memory data layout are obtained; and a second shape and a second step size of the register memory data layout are obtained.
[0094] Since the priority of the thread memory data layout is higher than the priority of the register memory data layout, the first shape and the second shape are combined in order of priority from high to low to obtain a target shape of the first memory data layout.
[0095] The data layout with the highest priority (i.e., the thread memory data layout) is obtained, and the second step size is converted into a third step size under the thread memory data layout based on the first number k1 of thread memories contained in the register memory; in actual application, the values of the second step size in each dimension are multiplied by the first number k1 to obtain the third step size.
[0096] The first step size and the third step size are combined to obtain a target step size of the first memory data layout.
[0097] For example, referring to Figure 4A , the register memory data layout is a matrix arrangement, i.e., the arrangement of data stored by the register memory r0, the register memory r1, the register memory r2, and the register memory r3 is a matrix arrangement, and the matrix ordering is W dimension first; the register memory data layout can be represented as (sh(2, 2), st(2, 1)), where sh(2, 2) represents a matrix arrangement.
[0098] The thread memory data layout is a matrix arrangement, i.e., the arrangement of data stored by the thread memory thr0, the thread memory thr1, …, and the thread memory thr31 is a matrix arrangement, and the matrix ordering is W dimension first; the thread memory data layout can be represented as (sh(8, 4), st(4, 1)).
[0099] sh(8, 4) and sh(2, 2) are combined to obtain a target shape of the first memory data layout of the basic processing unit, which is sh((8, 4), (2, 2)).
[0100] The second step length of the register memory data layout is st(2, 1), and the register memory includes 32 (the first number k1) thread memories, so the value of the second step length in each dimension is multiplied by the first number k1, to obtain a third step length st(64, 32).
[0101] The first step length st(4, 1) of the thread memory data layout is combined with the third step length st(64, 32) to obtain a target step length st((4, 1), (64, 32)) of the first memory data layout.
[0102] In summary, the first memory data layout of the basic processing unit finally obtained is (sh((8, 4), (2, 2)), st((4, 1), (64, 32))).
[0103] In some embodiments, the data layout of the plurality of physical memory levels includes: a thread memory data layout, a register memory data layout, and an execution unit memory data layout; the thread memory data layout, the register memory data layout, and the execution unit memory data layout are adapted to the hardware properties of the artificial intelligence chip. The thread memory data layout, the register memory data layout, and the execution unit memory data layout are combined to obtain a first memory data layout of a basic processing unit.
[0104] Specifically, a fourth shape and a fourth step length of the execution unit memory data layout are obtained; the respective priorities of the execution unit memory data layout, the register memory data layout, and the thread memory data layout are sequentially increased.
[0105] The first shape, the second shape, and the fourth shape are combined in order of priority from high to low to obtain a target shape of the first memory data layout.
[0106] The data layout with the highest priority (i.e., the thread memory data layout) is obtained, and based on the first number k1 of thread memories included in the register memory and the second number k2 of register memories included in the execution unit memory, the fourth step length is converted into a fifth step length under the thread memory data layout.
[0107] In actual applications, the value of the fourth step length in each dimension is sequentially multiplied by the first number k1 and the second number k2 to obtain the fifth step length. When the fourth step length is st(w, 1), the fifth step length is When the fourth step length is st(1, h), the fifth step length is .
[0108] The first step length, the third step length, and the fifth step length are combined to obtain a target step length of the first memory data layout.
[0109] For example, referring to Figure 4B, the register memory data layout is represented as: (sh(2, 2), st(2, 1)); the thread memory data layout is represented as: (sh(8, 4), st(4, 1)). The execution unit memory data layout is a matrix arrangement, that is, the arrangement of data stored in the execution unit memory eu0, the execution unit memory eu1, the execution unit memory eu2 and the execution unit memory eu3 is a matrix arrangement, and the matrix arrangement is dimension W first; the execution unit memory data layout can be represented as: (sh(4, 1), st(1, 1)). The execution unit memory data layout is a matrix arrangement, that is, the arrangement of data stored in the execution unit memory eu0, the execution unit memory eu1, the execution unit memory eu2 and the execution unit memory eu3 is a matrix arrangement, and the matrix arrangement is dimension W first; the execution unit memory data layout can be represented as: (sh(4, 1), st(1, 1)).
[0110] The execution unit memory data layout is a matrix arrangement, that is, the arrangement of data stored in the execution unit memory eu0, the execution unit memory eu1, the execution unit memory eu2 and the execution unit memory eu3 is a matrix arrangement, and the matrix arrangement is dimension W first; the execution unit memory data layout can be represented as: (sh(4, 1), st(1, 1)).
[0111] The fourth step of the execution unit memory data layout is st(1, 1), the register memory includes 32 (i.e., the first number k1) thread memories, and the execution unit memory includes 4 (i.e., the second number k2) register memories, so the value of the fourth step in each dimension is multiplied by the first number k1 and the second number k2 in turn to obtain the fifth step st(128, 128).
[0112] The first step st(4, 1), the third step st(64, 32) and the fifth step st(128, 128) are combined to obtain the target step st((4, 1), (64, 32), (128, 128)) of the first memory data layout.
[0113] In summary, the first memory data layout of the basic processing unit finally obtained is: (sh((8, 4), (2, 2), (4, 1)), st((4, 1), (64, 32), (128, 128))).
[0114] In some embodiments, the first memory data layout of the basic processing unit is also different when the data types are different, where the data types include: FP32 (32-bit floating point number), FP16 (16-bit floating point number), etc. For example, when the data type is FP32, each thread memory saves one data point in the data item; when the data type is FP16, each thread memory saves two data points in the data item.
[0115] At this point, the data layout across multiple physical memory levels includes: grid (Dem) memory data layout, thread memory data layout, register memory data layout, and execution unit memory data layout; the Dem memory data layout is adapted to the hardware attributes of the AI chip. Combining the Dem memory data layout, thread memory data layout, register memory data layout, and execution unit memory data layout yields the first memory data layout of the basic processing unit.
[0116] Specifically, the sixth shape and sixth step of the Dem memory data layout are obtained; the priorities of the execution unit memory data layout, register memory data layout, thread memory data layout, and Dem memory data layout increase sequentially.
[0117] The sixth shape, the first shape, the second shape, and the fourth shape are combined in descending order of priority to obtain the target shape of the first memory data layout.
[0118] Obtain the highest priority data layout (i.e., the Dem memory data layout). Based on the third quantity k3 of Dem memory in the thread memory, convert the first step length under the thread memory data layout to the seventh step length under the Dem memory data layout.
[0119] Based on the third quantity k3 of Dem memory in thread memory and the first quantity k1 of thread memory contained in register memory, the second step size under the register memory data layout is converted into the eighth step size under the Dem memory data layout.
[0120] Based on the third quantity k3 of Dem memory in thread memory, the first quantity k1 of thread memory contained in register memory, and the second quantity k2 of register memory contained in execution unit memory, the fourth step size under the execution unit memory data layout is converted into the ninth step size under the Dem memory data layout.
[0121] The sixth, seventh, eighth, and ninth step lengths are combined to obtain the target step length for the first memory data layout.
[0122] For example, see Figure 5 ,set up Figure 4B Each thread in the process stores two data points in memory, designated dem0 and dem1. The two data points are arranged according to... Given a matrix arrangement, and the matrix is sorted in W-dimensional priority, the memory data layout of Dem can be represented as: (sh(1,2), st(2,1)).
[0123] The sh (8, 4), sh (2, 2), sh (4, 1) and sh (1, 2) are combined to obtain a target shape of the first memory data layout of the basic processing unit: sh ((8, 4), (2, 2), (4, 1), (1, 2)).
[0124] The thread memory saves two data points, that is, the third number k3=2; and the first number k1=32 and the second number k2=4 can be known from the example shown in the formula (1). Figure 4B
[0125] Therefore, based on the third number k3, the first step length st (4, 1) under the thread memory data layout is converted into the seventh step length st (8, 2) under the Dem memory data layout.
[0126] Based on the third number k3 and the first number k1, the second step length st (2, 1) under the register memory data layout is converted into the eighth step length st (128, 64) under the Dem memory data layout.
[0127] Based on the third number k3, the first number k1 and the second number k2, the st (1, 1) under the execution unit memory data layout is converted into the ninth step length st (256, 256) under the Dem memory data layout.
[0128] The sixth step length, the seventh step length, the eighth step length and the ninth step length are combined to obtain the target step length st ((2, 1), (8, 2), (128, 64), (256, 256)) of the first memory data layout.
[0129] In summary, the first memory data layout of the basic processing unit is finally obtained as: (sh ((8, 4), (2, 2), (4, 1), (1, 2)), st ((2, 1), (8, 2), (128, 64), (256, 256))).
[0130] It should be noted that the first memory data layout of the basic processing unit is not limited to the above several forms, but can also be in other forms, and the present application does not limit this.
[0131] Based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-tensor corresponding sub-loading and storing processes; wherein when the target tensor is an integer multiple of the basic processing unit, each sub-tensor corresponds to a basic processing unit; when the target tensor is not an integer multiple of the basic processing unit, in order to ensure the hardware performance, it is still split into a plurality of basic processing units for loading and storing, at this time, the last processed basic processing unit will exceed the boundary of the target tensor, and the part exceeding the boundary will not perform the tensor loading and storing operation.
[0132] Before performing the load store, the first memory data layout of the basic processing unit is combined with the target arrangement mode of the basic processing unit corresponding to each of the plurality of sub-tensors to obtain a second memory data layout of the target tensor. The priority of the first memory data layout is higher than the priority of the target arrangement mode. The specific combination mode is the same as the combination mode of obtaining the first memory data layout, which will not be described here.
[0133] For example, referring to Figure 6 , it is set to split the target tensor into two basic processing units for load storage, i.e., basic processing unit T0 and basic processing unit T1, which are arranged in a matrix according to dimension priority, and the target arrangement mode of the two basic processing units can be represented as: (sh(1, 2), st(2, 1)); and it is set that the first memory data layout of each basic processing unit is: (sh((1, 2), (8, 4), (2, 2), (4, 1)), st((2, 1), (8, 2), (128, 64), (256, 256))).
[0134] The first memory data layout is combined with the target arrangement mode to obtain the second memory data layout of the target tensor: ((sh(1, 2), (8, 4), (2, 2), (4, 1)) (1, 2), st((2, 1), (8, 2), (128, 64), (256, 256) (2048, 1024))).
[0135] The following describes each sub-load storage process (i.e., step 202 shown in Figure 2 ), which includes the following steps:
[0136] Step 2021, based on the first memory data layout of the basic processing unit, obtaining the destination physical address of each data item in the sub-tensor.
[0137] Specifically, when loading tensor data from the video memory to the register memory, the target storage space is the register memory. The destination physical address refers to the physical storage address of the data item in the register memory after performing the load store.
[0138] In some embodiments, based on the first memory data layout of the basic processing unit, the target physical index corresponding to the data item is obtained; and then based on the target physical index, the destination physical address of the data item is mapped and obtained.
[0139] The target physical index includes a register memory index and a thread memory index. The register memory index is obtained from a register memory data layout. For example, the register memory data layout is an arrangement of data stored in a register memory r0, a register memory r1, a register memory r2, and a register memory r3, where r0, r1, r2, and r4 are respective register memory indexes of the four register memories.
[0140] The thread memory index is obtained from a thread memory data layout. For example, the thread memory data layout is an arrangement of data stored in a thread memory thr0, a thread memory thr1, …, and a thread memory thr31, where thr0, thr1, …, and thr31 are respective thread memory indexes of the 32 thread memories.
[0141] The target physical address includes a register memory address and a thread memory address. The register memory address is a physical address of the register memory. The thread memory address is a physical address of the thread memory. Based on the register memory index, the register memory address of the data item is mapped. Based on the thread memory index, the thread memory address of the data item is mapped.
[0142] At step 2022, for each data item, a source physical address of the data item is obtained based on the target physical address and a logical layout of a source tensor containing the data item.
[0143] Specifically, the source tensor is stored in an original memory, which can be a video memory or an on-chip cache. The source physical address of the data item refers to a physical storage address of the data item in the original memory before the load store is performed.
[0144] In some embodiments, based on the target physical address, a second memory data layout, and a logical layout of the source tensor, a target logical address of the data item is obtained. Based on the target logical address of the data item, an original logical address of the data item in the source tensor is mapped. Based on the original logical address of the data item, a source physical address of the data item in the original memory is mapped.
[0145] Specifically, the logical layout of the source tensor is input by a user. The target physical address includes a register memory address and a thread memory address. Based on the register memory address, the second memory data layout, and the logical layout of the source tensor, a register logical address of the data item is obtained. Based on the thread memory address, the second memory data layout, and the logical layout of the source tensor, a thread logical address of the data item is obtained. The register logical address and the thread logical address are added to obtain the target logical address of the data item. The target logical address refers to a logical address of the data item in the target tensor.
[0146] Then, a starting logical address of the source tensor and a starting logical address of the target tensor are obtained, the source tensor including a plurality of dimensions, the source tensor and the target tensor including the same dimensions.
[0147] When the starting logical address of the target tensor is the same as the starting logical address of the source tensor, a target logical address of the data item is taken as an original logical address of the data item.
[0148] When the starting logical address of the target tensor is different from the starting logical address of the source tensor, it is determined that the starting logical address of the target tensor and the starting logical address of the source tensor are address offsets in the plurality of dimensions respectively; and then the original logical address of the data item is obtained by accumulating the address offsets in the plurality of dimensions respectively on the basis of the target logical address of the data item.
[0149] For example, referring to Figure 7 , the target tensor 701 and the source tensor 702 are both two-dimensional tensors, and the starting logical address of the target tensor 701 and the starting logical address of the source tensor 702 are both (0, 0).
[0150] For the data item 703, the target logical address of the data item 703 is (1, 1), and the original logical address of the data item 703 is also (1, 1).
[0151] For example, referring to Figure 8 , the target tensor 701 and the source tensor 702 are both two-dimensional tensors, the starting logical address of the source tensor 702 is (0, 0), and the starting logical address of the target tensor 701 is (0, 5). In the W dimension, the address offset of the starting logical address of the target tensor 701 compared with the starting logical address of the source tensor 702 is 5, that is, the address offset in the W dimension is 5. In the H dimension, the address offset of the starting logical address of the target tensor 701 compared with the starting logical address of the source tensor 702 is 0, that is, the address offset in the H dimension is 0.
[0152] Therefore, for the data item 703, the target logical address of the data item 703 is (1, 1), and the original logical address of the data item 703 is obtained by accumulating the address offset in the W dimension on the starting logical address of the data item 703 in the W dimension and accumulating the address offset in the H dimension on the starting logical address of the data item 703 in the H dimension, that is, (1, 6).
[0153] Then, the original logical address of the data item is mapped to the source physical address of the data item in the original memory on the basis of the mapping relationship between the logical address and the physical address. In actual application, the source physical address of each data item included in the sub-tensor is calculated in parallel by a plurality of threads.
[0154] In step 2023, the data item is loaded from the source physical address and stored in the destination physical address.
[0155] Specifically, a plurality of threads parallel loads a plurality of data items from a source physical address to a destination physical address. This thread-level parallel architecture based on the basic processing unit can fully utilize the parallel computing capability of the underlying hardware, thereby improving the overall performance of the tensor load and store.
[0156] In the embodiments of the present application, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-tensor corresponding sub-loading and storing processes; the first memory data layout of the basic processing unit is adapted to the hardware properties of the artificial intelligence chip, that is, the first memory data layout of the basic processing unit meets the hardware alignment granularity of the artificial intelligence chip, and at the same time, when the data is loaded and stored according to the first memory data layout of the basic processing unit, the resource (such as bandwidth) usage of the artificial intelligence chip reaches the upper limit value. Therefore, when the basic processing unit is used as the minimum data unit for loading and storing tensor data, it is more consistent with the hardware memory structure of the artificial intelligence chip. This not only makes full use of the hardware resources of the chip and improves the utilization rate of hardware resources, but also improves the performance of tensor loading and storage. Secondly, a plurality of threads can perform loading and storing operations in parallel. This thread-level parallel architecture based on the basic processing unit can significantly improve the efficiency of data loading and storing, and the effect is more significant when processing large-scale tensors.
[0157] In some embodiments, the original instruction set of the pre-packaged basic processing unit is obtained; and then the original instruction set is optimized by an expression system to obtain a target instruction set, which is used to execute each sub-loading and storing process.
[0158] Specifically, the original instruction set includes instructions that need to be executed in the process of calculating the source physical address of each data item from the destination physical address of each data item. The number of these instructions is relatively large and there are some redundant instructions. Therefore, the original instruction set is optimized at the instruction level by the expression system. The instruction set optimization specifically includes instruction merging and instruction rearrangement, thereby reducing the programming complexity, making it easier to optimize the computing performance, and improving the computing performance.
[0159] Based on the same technical concept, the embodiments of the present application provide a structural diagram of a tensor load and store device, as shown in Figure 9 The tensor load and store device 900 includes:
[0160] The splitting module 901 is configured to split, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor into a plurality of sub-tensor corresponding sub-loading and storing processes.
[0161] The execution module 902 is configured to execute each sub-loading and storing process, specifically performing the following steps:
[0162] obtain a destination physical address of each data item in the sub-tensor based on a first memory data layout of the basic processing unit; the first memory data layout is adapted to hardware attributes of the artificial intelligence chip;
[0163] For each data item, obtain a source physical address of the data item based on the destination physical address and a logical layout of a source tensor containing the data item;
[0164] Load the data item from the source physical address and store the data item in the destination physical address.
[0165] Optionally, further comprising an instruction optimization module 903;
[0166] The instruction optimization module 903 is specifically configured to:
[0167] Pre-package a set of original instructions of the basic processing unit;
[0168] Optimize the set of original instructions through an expression system to obtain a set of target instructions, the set of target instructions being used to execute each of the sub-load and store processes.
[0169] Optionally, the execution module 902 is specifically configured to:
[0170] For each data item, obtain a target physical index corresponding to the data item based on a first memory data layout of the basic processing unit;
[0171] Map to obtain a destination physical address of the data item based on the target physical index.
[0172] Optionally, the target physical index comprises a register memory index and a thread memory index; and the destination physical address comprises a register memory address and a thread memory address.
[0173] The execution module 902 is specifically configured to:
[0174] Map to obtain a register memory address of the data item based on the register memory index;
[0175] Map to obtain a thread memory address of the data item based on the thread memory index.
[0176] Optionally, each of the plurality of sub-tensors corresponds to a basic processing unit.
[0177] The execution module 902 is further configured to:
[0178] The first memory data layout of the basic processing unit is combined with a target arrangement mode of the basic processing unit corresponding to each of the plurality of sub-tensors to obtain a second memory data layout of the target tensor.
[0179] Optionally, the execution module 902 is specifically used for:
[0180] Based on the target physical address, the second memory data layout and a logical layout of the source tensor, a target logical address of the data item is obtained.
[0181] Based on the target logical address of the data item, an original logical address of the data item in the source tensor is mapped and obtained.
[0182] Based on the original logical address of the data item, a source physical address of the data item in the original memory is mapped and obtained.
[0183] Optionally, the target physical address includes a register memory address and a thread memory address.
[0184] The execution module 902 is specifically used for:
[0185] Based on the register memory address, the second memory data layout and the logical layout of the source tensor, a register logical address of the data item is obtained.
[0186] Based on the thread memory address, the second memory data layout and the logical layout of the source tensor, a thread logical address of the data item is obtained.
[0187] The register logical address and the thread logical address are accumulated to obtain the target logical address of the data item.
[0188] In the embodiment of the application, based on the pre-packaged basic processing unit, the process of loading and storing the target tensor is split into a plurality of sub-loading and storing processes corresponding to each of the sub-tensors; the first memory data layout of the basic processing unit is adapted to the hardware properties of the artificial intelligence chip, that is, the first memory data layout of the basic processing unit meets the hardware alignment granularity of the artificial intelligence chip, and at the same time, when loading and storing data according to the first memory data layout of the basic processing unit for processing, the resource (such as bandwidth) usage of the artificial intelligence chip reaches the upper limit value, so that when loading and storing tensor data with the basic processing unit as the minimum data unit, the hardware memory structure of the artificial intelligence chip is more in line with the hardware memory structure of the artificial intelligence chip, which not only makes full use of the hardware resources of the chip and improves the utilization rate of the hardware resources, but also improves the performance of tensor loading and storing; secondly, a plurality of threads can perform loading and storing operations in parallel, and this thread-level parallel architecture based on the basic processing unit can significantly improve the efficiency of data loading and storing, and the effect is more significant when processing large-scale tensors.
[0189] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0190] Based on the same technical concept, the embodiments of the present application provide a computer device, as shown in Figure 10 The computer device includes at least one artificial intelligence chip 100, and a memory 1001 connected with the at least one artificial intelligence chip 100. In the embodiments of the present application, the specific connection medium between the artificial intelligence chip 100 and the memory 1001 is not limited, Figure 10 For example, the artificial intelligence chip 100 and the memory 1001 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0191] In the embodiments of the present application, the memory 1001 stores instructions executable by the at least one artificial intelligence chip 100. The at least one artificial intelligence chip 100 can execute the steps of the tensor loading and storing method by executing the instructions stored in the memory 1001.
[0192] The artificial intelligence chip 100 is the control center of the computer device, and can connect various parts of the computer device through various interfaces and lines, and realize tensor loading and storing by running or executing instructions stored in the memory 1001 and calling data stored in the memory 1001. Optionally, the artificial intelligence chip 100 can include one or more processing units. The artificial intelligence chip 100 can integrate an application processor and a modem processor. The application processor mainly processes the operating system, user interface and application program, etc. The modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the artificial intelligence chip 100. In some embodiments, the artificial intelligence chip 100 and the memory 1001 can be implemented on the same chip, and in some embodiments, they can also be implemented on separate chips respectively.
[0193] The artificial intelligence chip 100 can be a general processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as execution completed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0194] The memory 1001 is a non-volatile computer readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 1001 can include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card type memory, random access memory (RAM), static random access memory (SRAM), programmable read only memory (PROM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. The memory 1001 is any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer device, but is not limited thereto. The memory 1001 in the embodiments of the present application can also be a circuit or any other device capable of realizing a storage function, used for storing program instructions and / or data.
[0195] Based on the same inventive concept, the embodiments of the present application provide a computer readable storage medium storing a computer program executable by a computer device, which, when executed on the computer device, causes the computer device to perform the steps of the tensor loading and storing method described above.
[0196] Based on the same inventive concept, the embodiments of the present application provide a computer program product, which includes a computer program stored on a computer readable storage medium, and the computer program includes program instructions, which, when executed by a computer device, cause the computer device to perform the steps of the tensor loading and storing method described above.
[0197] Those skilled in the art will appreciate that embodiments of the present application can be readily used as software, hardware, or a combination of software and hardware. In a software embodiment, the methods can be tangibly embodied in a machine-readable storage medium having stored thereon instructions that can be used to program a processing machine (or multiple machines) to perform the methods. The term "storage medium" as used herein shall
[0198] The present application is described in reference to the drawings using a flow diagram and / or a block diagram of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flow diagram and / or block diagram, and combinations of blocks in the flow diagram and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing machine, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagram and / or block diagram block or blocks. Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks. Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks.
[0199] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow diagram and / or block diagram block or blocks. Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks. Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks.
[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow diagram and / or block diagram block or blocks. Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks. Figure 1 Figure 1 one or more functions specified in the flow diagram and / or block diagram block or blocks.
[0201] While preferred embodiments of the application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the foregoing description. Accordingly, it is intended that the appended claims be construed to include all such variations and modifications as fall within the scope of the present application.
[0202] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A tensor loading and storage method, characterized in that, Applications in artificial intelligence chips, including: Based on pre-encapsulated basic processing units, the process of loading and storing the target tensor is broken down into multiple sub-loading and storing procedures corresponding to each sub-tensor; each sub-loading and storing procedure includes the following steps: Based on the first memory data layout of the basic processing unit, the destination physical address of each data item in the sub-tensor is obtained; the first memory data layout is adapted to the hardware attributes of the artificial intelligence chip; the first memory data layout satisfies the hardware alignment granularity of the artificial intelligence chip, and when the stored data is loaded according to the first memory data layout, the resource usage of the artificial intelligence chip reaches the upper limit. For each data item, the source physical address of the data item is obtained based on the destination physical address and the logical layout of the source tensor containing the data item; The data item is loaded from the source physical address and stored at the destination physical address.
2. The method as described in claim 1, characterized in that, Also includes: The original instruction set of the basic processing unit is pre-encapsulated; The original instruction set is optimized using an expression system to obtain a target instruction set, which is used to execute each of the sub-loaded stored procedures.
3. The method as described in claim 1, characterized in that, The first memory data layout based on the basic processing unit, obtaining the destination physical address of each data item in the sub-tensor, includes: For each data item, the target physical index corresponding to the data item is obtained based on the first memory data layout of the basic processing unit; Based on the target physical index, the destination physical address of the data item is obtained through mapping.
4. The method as described in claim 3, characterized in that, The target physical index includes: a register memory index and a thread memory index; the destination physical address includes: a register memory address and a thread memory address. The process of mapping and obtaining the destination physical address of the data item based on the target physical index includes: Based on the register memory index, the register memory address of the data item is obtained by mapping; Based on the thread memory index, the thread memory address of the data item is obtained through mapping.
5. The method as described in claim 1, characterized in that, Each of the plurality of sub-tensors corresponds to a basic processing unit; the method further includes: The first memory data layout of the basic processing unit is combined with the target arrangement of the basic processing units corresponding to the plurality of sub-tensors to obtain the second memory data layout of the target tensor.
6. The method as described in claim 5, characterized in that, Obtaining the source physical address of the data item based on the destination physical address and the logical layout of the source tensor containing the data item includes: Based on the destination physical address, the second memory data layout, and the logical layout of the source tensor, the target logical address of the data item is obtained; Based on the target logical address of the data item, the original logical address of the data item in the source tensor is obtained by mapping. Based on the original logical address of the data item, the source physical address of the data item in the original memory is obtained by mapping.
7. The method as described in claim 6, characterized in that, The destination physical address includes: register memory address and thread memory address; Obtaining the target logical address of the data item based on the destination physical address, the second memory data layout, and the logical layout of the source tensor includes: Based on the register memory address, the second memory data layout, and the logical layout of the source tensor, the register logical address of the data item is obtained; Based on the thread memory address, the second memory data layout, and the logical layout of the source tensor, the thread logical address of the data item is obtained; The register logical address and the thread logical address are added together to obtain the target logical address of the data item.
8. A computer device comprising a memory, an artificial intelligence chip, and a computer program stored in the memory and running on the artificial intelligence chip, characterized in that, When the artificial intelligence chip executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, It stores a computer program that is executed by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium, the computer program including program instructions that, when executed by a computer device, cause the computer device to perform the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Execution method and device for memory handling operator and storage medium
CN118193410A
Tensor memory carrying method and device, storage medium and program product
CN118567580A