Data loading method, data storage method, processor, electronic equipment and medium
By determining multiple requests in a parallel processor to load tensors into the cache, the problem of discontinuous loading of tensors in certain dimensions is solved, efficient data loading and storage is achieved, and computing performance is improved.
Patent Information
- Application Number
- CN202510607351.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When the prior art loads and stores tensor data in parallel processors, it is impossible to effectively handle the problem of discontinuous loading of tensors in certain dimensions, resulting in insufficiency of data bandwidth and computation.
By obtaining the shape size, data storage format and starting coordinates of the tensor to be processed, combining the information of the original tensor, multiple requests are determined to load the tensor into the cache area, and the request is divided using state machine logic to ensure that request division is performed on dimensions that cannot be continuously loaded, improving data loading efficiency.
It realizes efficient loading and storing tensor data in parallel processors, improves data access bandwidth and hardware utilization of computing units, and improves computing performance.
Smart Images

Figure CN120144491A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] A tensor is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. For example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor can have any number of dimensions. Tensor operations are widely used in processors such as parallel processors.
[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processor devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of parallel processors require a large amount of data, and this data is generally stored in the storage components of parallel processors. For example, the storage components can be memories. Through data loading instructions, this data can be extracted from the storage components to the buffer for calculation, and through data storage instructions, the data in the buffer can be stored in the memory. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a data loading method for loading a tensor to be processed into a buffer. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, and a5, where a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The data loading method includes: obtaining the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, where the data storage format of the original tensor is the same as that of the tensor to be processed, and the shape and size of the original tensor are represented by b1, b2, b3, b4, and b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component; combining the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for loading the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the tensor to be processed into the buffer; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, requests are divided in the second dimension, the sub-data loaded by each request are all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request are from the original tensor and are continuously stored in the memory, the data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0005] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that continuous loading cannot be performed on the tensor to be processed in the first dimension.
[0006] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the data storage format of the to-be-processed tensor being NDHWC, the coordinates of each sub-data loaded by each request satisfy the following conditions: the coordinates of each dimension in the first dimension and dimensions lower than the first dimension in the to-be-processed tensor are different, but the coordinates of each dimension in the second dimension and dimensions higher than the second dimension are the same; where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension.
[0007] For example, in the data loading method provided by at least one embodiment of the present disclosure, in combination with the shape and size of the original tensor, the to-be-processed tensor, the data storage format of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor, determining a plurality of requests for loading the to-be-processed tensor includes: based on the starting coordinates of the to-be-processed tensor, the shape and size of the to-be-processed tensor, and the data storage format of the to-be-processed tensor, determining the first request sent among the plurality of requests and the initial state of the first request entering the state machine; based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the to-be-processed tensor, using the state machine to determine each request among the plurality of requests other than the first request.
[0008] For example, in the data loading method provided by at least one embodiment of the present disclosure, each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the length of the data loaded by the request. Based on the starting coordinates of the to-be-processed tensor, the shape and size of the to-be-processed tensor, and the data storage format of the to-be-processed tensor, determining the first request sent among the plurality of requests and the initial state of the first request entering the state machine includes: based on the starting coordinates of the to-be-processed tensor, determining the initial state of the first request entering the state machine; taking the starting coordinates of the to-be-processed tensor as the request initial coordinates corresponding to the first request; according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the to-be-processed tensor, determining the data read address of the first request; determining the starting address of writing the to-be-processed tensor in the buffer as the data write address of the first request; according to the starting coordinates of the to-be-processed tensor and the shape and size of the to-be-processed tensor, determining the length of the data loaded by the first request.
[0009] For example, in the data loading method provided by at least one embodiment of the present disclosure, determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.
[0010] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the initial state, combining the shape and size of the original tensor and the data storage format of the tensor to be processed, using the state machine to determine each of the multiple requests other than the first request includes: based on the initial state, using the state machine to determine the request initial coordinates corresponding to each of the requests and the data length loaded by each of the requests; according to the shape and size of the original tensor, the request initial coordinates corresponding to each of the requests, and the data storage format of the tensor to be processed, determining the data read addresses of each of the requests; according to the data length loaded by each of the requests, determining the data write addresses of each of the requests.
[0011] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: in response to the state of the current request being the first state and satisfying a first condition, determining that the next request of the current request enters the first state, where in response to the current request being the first request, the state of the current request is the initial state; in response to the next request entering the first state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is a first coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length loaded by the current request is the size of the to-be-processed tensor in the first dimension; where the first condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first condition includes that the sum of the fourth coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the second coordinate value and the fourth coordinate value is less than the second coordinate value, the remaining size is the number of remaining tensor data in the first dimension of the to-be-processed tensor that has not been requested to be loaded, and the first coordinate value is the coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension.
[0012] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine further includes a second state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests further includes: in response to the state of the current request being the first state and satisfying a second condition, determining that the next request of the current request enters the second state, in response to the next request entering the second state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the second coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged; determining the data length loaded by the current request based on the first coordinate value and the second coordinate value; where the second condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, or the second condition includes that the sum of the fourth coordinate value and the remaining size is greater than or equal to the second coordinate value.
[0013] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state and a second state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: in response to the state of the current request being the second state and satisfying a third condition, determining that the next request of the current request enters the first state, where, in response to the current request being the first request, the state of the current request is the initial state; in response to the next request entering the first state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is a first coordinate value, determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions other than the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining the data length loaded by the current request based on the first coordinate value and the size of the to-be-processed tensor in the first dimension; where the third condition includes that the first coordinate value is less than a second coordinate value and the sum of the first coordinate value and the size of the to-be-processed tensor in the first dimension is less than a third coordinate value, or, the third condition includes that the sum of the fourth coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the third coordinate value and the fourth coordinate value is less than the second coordinate value, the remaining size is the number of remaining tensor data in the first dimension of the to-be-processed tensor that has not been requested to be loaded, the first coordinate value is the coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the first dimension, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
[0014] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the initial state, the state machine is used to determine the request initial coordinates corresponding to the respective requests and the data lengths to be loaded by the respective requests, and further includes: in response to the state of the current request being the second state and satisfying the fourth condition, determining that the next request enters the second state; in response to the next request entering the second state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length to be loaded by the current request is the size of the to-be-processed tensor in the first dimension; wherein, the fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the to-be-processed tensor in the first dimension is less than the third coordinate value, or the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
[0015] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine further includes a third state. Based on the initial state, the state machine is used to determine the request initial coordinates corresponding to the respective requests and the data lengths to be loaded by the respective requests, and further includes: in response to the state of the current request being the second state and satisfying the fifth condition, determining that the next request enters the third state; in response to the next request entering the third state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the third coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged; determining the data length to be loaded by the current request based on the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the third coordinate value; wherein, the fifth condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is greater than or equal to the third coordinate value, or the fifth condition includes that the sum of the fourth coordinate value and the remaining size is greater than the third coordinate value.
[0016] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state, a second state, and a third state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: in response to the state of the current request being the third state, determining the state entered by the next request corresponding to the current request according to the first coordinate value, where in response to the current request being the first request, the state of the current request is the initial state; determining that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is the first coordinate value, and determining that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length loaded by the current request is determined based on the fourth coordinate value and the third coordinate value of the request initial coordinates corresponding to the current request in the first dimension, where the first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension, and the third coordinate value is the sum of the second coordinate value of the starting coordinate of the original tensor in the first dimension and the size of the original tensor in the first dimension.
[0017] For example, in the data loading method provided by at least one embodiment of the present disclosure, the multiple requests are sequentially sent, and the sub-data returned by each request is sequentially written into the buffer area to load the tensor to be processed into the buffer area, including: for any one request, in response to all the sub-data loaded by the any one request being located in the memory, sending the any one request to the memory; in response to all the sub-data loaded by the any one request not being located in the memory, converting the any one request into writing a plurality of predetermined values into the buffer area, where the number of the plurality of predetermined values is determined by the data length specified by the any one request.
[0018] For example, in the data loading method provided by at least one embodiment of the present disclosure, the data storage format of the tensor to be processed is NDHWC, or N(C / x)DHW(xC), where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, and x is a positive integer greater than 1. In response to the data storage format of the tensor to be processed being NDHWC, the coordinate value of the channel number dimension is incremented by 1 when updated; in response to the data storage format of the tensor to be processed being N(C / x)DHW(xC), the coordinate value of the channel number dimension is incremented by x when updated.
[0019] At least one embodiment of the present disclosure provides a data loading method, including: receiving a data loading instruction indicating to execute loading a tensor to be processed into a buffer, where the data loading instruction includes the shape dimensions of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in the 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component; after parsing the data loading instruction, use an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: combining the original tensor, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for loading the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the tensor to be processed into the buffer; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, divide the requests in the second dimension, the sub-data loaded by each request is all located in the memory or none of them is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory, the data storage format of the tensor to be processed indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0020] At least one embodiment of the present disclosure provides a data storage method for writing a tensor to be processed in a buffer into memory. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, and a5, where a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The data storage method includes: obtaining the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, where the data storage format of the original tensor is the same as that of the tensor to be processed, and the shape and size of the original tensor are represented by b1, b2, b3, b4, and b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component; combining the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for storing the tensor to be processed; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when storing the tensor to be processed, continuous storage cannot be performed in the first dimension but continuous storage can be performed in each dimension lower than the first dimension, requests are divided in the second dimension, and the sub-data to be written by each request either all belong to the data range of the original tensor or all do not belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory. The data storage format of the tensor to be processed indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0021] For example, in the data storage method provided by at least one embodiment of the present disclosure, combining the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for storing the tensor to be processed includes: determining the first request sent among the plurality of requests and the initial state of the first request entering a state machine based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed; and based on the initial state, combining the shape and size of the original tensor and the data storage format of the tensor to be processed, and using the state machine to determine each request other than the first request among the plurality of requests.
[0022] For example, in the data storage method provided by at least one embodiment of the present disclosure, the request initial coordinates corresponding to each request are used to determine the data storage address of the sub-data to be stored by the request in the memory, and at least some of the multiple requests are sequentially sent, and the data to be written indicated by each request is sequentially written into the memory to store the tensor to be processed, including: for any one of the requests, in response to the first request coordinate value of the request initial coordinates corresponding to the any one of the requests in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determining not to send the any one of the requests; in response to the first request coordinate value being greater than or equal to the second coordinate value and the first request coordinate value being less than the third coordinate value, determining to send the any one of the requests; wherein, the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0023] At least one embodiment of the present disclosure provides a processor, including an instruction parsing unit and an execution unit. Wherein, the instruction parsing unit is configured to receive and parse a data loading instruction, and the data loading instruction includes the shape and size of the to-be-processed tensor, the data storage format of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor in the coordinate system determined by the original tensor as input parameters. The shape and size of the to-be-processed tensor are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the to-be-processed tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The data storage format of the original tensor is the same as that of the to-be-processed tensor. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in the 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. After the instruction parsing unit parses the data loading instruction, the execution unit executes the data loading instruction. When the execution unit executes the data loading instruction, the following operations are included: combining the original tensor, the shape and size of the to-be-processed tensor, the data storage format of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor to determine a plurality of requests for loading the to-be-processed tensor; sequentially sending the plurality of requests, and writing the sub-data returned by each request into the buffer in order to load the to-be-processed tensor into the buffer. Wherein, in response to the size relationship between the to-be-processed tensor and the original tensor in the first dimension such that when loading the to-be-processed tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, requests are divided in the second dimension. The sub-data loaded by each request are all located in the memory or none of them are located in the memory. And in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request are from the original tensor and are continuously stored in the memory. The data storage format of the to-be-processed tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0024] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data loading method according to at least one embodiment of the present disclosure, or the data storage method according to at least one embodiment of the present disclosure.
[0025] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the data loading method according to at least one embodiment of the present disclosure, or the data storage method according to at least one embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0027] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU); Figure 2A It is a schematic structure of a tensor; Figure 2B It is a schematic diagram of the storage format of NDHWC; Figure 2C It is a schematic diagram of the storage format of N(C / 32)DHW(32C); Figure 3 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure; Figure 4A It is a schematic diagram of a tensor to be processed provided by an embodiment of the present disclosure; Figure 4B It is a schematic diagram of a tensor to be processed provided by another embodiment of the present disclosure; Figure 4C It is a schematic diagram of a tensor to be processed provided by another embodiment of the present disclosure; Figure 5 It is a schematic diagram of a state machine provided by an embodiment of the present disclosure; Figure 6 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure; Figure 7 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure; Figure 8 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure; Figure 9 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure; Figure 10 It is a schematic structural diagram of a processor provided by at least one embodiment of the present disclosure; Figure 11 It is a schematic structural diagram of a processor provided by at least one embodiment of the present disclosure; Figure 12 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0028] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0029] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, some details of known functions and known components are omitted in the present disclosure.
[0030] Figure 1 Schematic structural diagram of a general-purpose graphics processing unit (GPGPU).
[0031] As Figure 1 shown, a general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processors may be streaming processor clusters (SPCs), for example, including Figure 1 the streaming processor cluster 1 shown in the figure, …, the streaming processor cluster M, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing is performed between multiple streaming processor clusters through a global cache or global memory.
[0032] As Figure 1 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster includes multiple computing units, for example Figure 1The computing units 1, 2, …, N in it, where N is a positive integer. Each computing unit (abbreviated as CU) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or computing kernels), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores in the computing unit.
[0033] As Figure 1 shown, each streaming processor cluster also provides a buffer for caching data of the N computing units in the streaming processor cluster.
[0034] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processing unit (or called a parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called a thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.
[0035] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates the thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the intermediate-level cache or global cache or global memory (such as Figure 1 the high-bandwidth memory in it, High Bandwidth Memory, abbreviated as HBM) for read and write operations, etc.
[0036] As Figure 1As shown, general computing operations, such as computing operations on matrices, usually require a large amount of data, which is usually stored in a memory, such as can be stored in HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored in the memory (Store operation). The storage method of data in the memory affects the memory access bandwidth, and thus affects the hardware utilization rate of the computing unit.
[0037] For example, general computing operations include General Matrix Multiplication (GEMM for short). The data required for general matrix multiplication includes a first tensor and a second tensor.
[0038] For example, for the first tensor or the second tensor, its shape dimensions can be represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor data to be processed in 5 dimensions, and a1, a2, a3, a4, a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, the number of data samples captured in one training. The D dimension represents the depth, the H dimension represents the height of the input data, the W dimension represents the width of the input data, and the C dimension represents the number of channels. For example, taking the first tensor as an example, a1 can be the N dimension size, a2 can be the D dimension size, a3 can be the H dimension size, a4 can be the W dimension size, and a5 can be the C dimension size. Of course, the present disclosure does not make specific limitations on this.
[0039] The elements of the tensor are arranged in various formats in the memory (such as Figure 1 the memory), which is called the data storage format (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0040] For example, Figure 2A is a schematic structure of a tensor. In the Figure 2A shown tensor, a1 is the N dimension size and is equal to 1, a2 is the D dimension size and is equal to 1, a3 is the H dimension size and is equal to 5, a4 is the W dimension size and is equal to 4, and a5 is the C dimension size and is equal to 64.
[0041] For example, Figure 2A the pixel elements of the tensor in are represented as 0, 1, 2, 3,... and so on. The following describes different data storage formats with the tensor shown in Figure 2A as an example.
[0042] For example, the data storage format can include NDHWC, also known as the Linear mode. Figure 2BSchematic diagram of the storage format for NDHWC.
[0043] For example, for NDHWC, as Figure 2B shown, starting from the first element of the first channel (c0 in Figure 2B where a5 = 0), Figure 2B element 0 in Figure 2B ), then store the first element of the second channel (c1 in Figure 2B where a5 = 1), Figure 2B element 20 in Figure 2B ), and so on until the first elements of all channels are laid out. For example, after the first element of the 64th channel (c63 in Figure 2B where a5 = 63), Figure 2B element 1260 in Figure 2B ), select the second element of the first channel (c0 in Figure 2B where a5 = 0), Figure 2B element 1 in Figure 2B ), then store the second element of the second channel (c1 in Figure 2B where a5 = 1), Figure 2B element 21 in Figure 2B ), and so on until the second elements of all channels are laid out, and so on. Figure 2B element 1260 in Figure 2B ), select the second element of the first channel (c0 in Figure 2B where a5 = 0), Figure 2B element 1 in Figure 2B ), then store the second element of the second channel (c1 in Figure 2B where a5 = 1), Figure 2B element 21 in Figure 2B ), and so on until the second elements of all channels are laid out, and so on. Figure 2B element 21 in Figure 2B ), and so on until the second elements of all channels are laid out, and so on. Figure 2B element 21 in Figure 2B ), and so on until the second elements of all channels are laid out, and so on.
[0044] For example, the data storage format can also include N(C / x)DHW(xC), also known as the interleave mode, where x can be 8, 16, 32, etc. according to needs.
[0045] N(C / x)DHW(xC) is similar to NDHWC, but the key difference is that in the memory layout of N(C / x)DHW(xC), the a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.
[0046] Figure 2C Schematic diagram of the storage format for N(C / 32)DHW(32C).
[0047] As Figure 2C shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 2C c0 in Figure 2C ) to a5 = 31 ( Figure 2C c31 in Figure 2C ), and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.
[0048] In memory, the original tensor is continuously stored in memory in the data storage format as described above. The shape of the original tensor can be expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, for example, the shape of the partial tensor can be expressed as a1×a2×a3×a4×a5. Since it is a partial tensor inside the original tensor, it may not be possible to continuously extract in each dimension during extraction, and it is impossible to accurately obtain the amount of data and the data location requested when loading or storing data to optimally implement data loading or storing, which greatly reduces the data bandwidth when loading data and reduces the hardware computing efficiency.
[0049] At least one embodiment of the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium.
[0050] In at least one embodiment, the data loading method is used to load a tensor to be processed into a buffer area. The shape size of the tensor to be processed is represented by a1, a2, a3, a4, and a5, where a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The data loading method includes: obtaining the shape size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape size of the original tensor is represented by b1, b2, b3, b4, and b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. Combining the original tensor, the shape size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed, determining a plurality of requests for loading the tensor to be processed. Sequentially sending the plurality of requests to the memory, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area. In response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it is not possible to continuously load in the first dimension but it is possible to continuously load in each dimension lower than the first dimension, dividing the requests in the second dimension, the sub-data loaded by each request is either all located in the memory or none of it is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory. The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0051] In the data loading method provided by at least one embodiment of the present disclosure, requests are divided according to whether the tensor can be continuously loaded in the first dimension. Each request is used to load sub-data that is either all located in memory or all not located in memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.
[0052] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0053] Figure 3 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure.
[0054] As Figure 3 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S10 - S30.
[0055] For example, the data loading method provided by at least one embodiment of the present disclosure is used to load the tensor to be processed into a buffer area, and the buffer area is, for example, a buffer in a streaming processor cluster.
[0056] For example, the shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. In subsequent embodiments of the present disclosure, an example is described where a1 is the N - dimension size, a2 is the D - dimension size, a3 is the H - dimension size, a4 is the W - dimension size, and a5 is the C - dimension size.
[0057] For example, the original tensor is stored in memory. The tensor to be processed can be a partial tensor of the original tensor, or the tensor to be processed can be the original tensor itself.
[0058] Figure 4A It is a schematic diagram of the tensor to be processed provided by an embodiment of the present disclosure.
[0059] Figure 4A In, each solid - line small cube represents an element, and the tensor composed of multiple solid - line small cubes is the original tensor. Figure 4A What is shown is one batch and a tensor at a certain depth in one batch, but the original tensor may have multiple batches, and each batch may have multiple tensors in the depth dimension. Its structure is the same as Figure 4AThe same as that shown, and will not be described again here.
[0060] Assume Figure 4A The element pointed by the arrow is the origin of the coordinate system where the original tensor is located. C, W, and H respectively represent the three coordinate axes. row is the coordinate value in the H - dimension direction, col is the coordinate value in the W - dimension direction, c is the coordinate value in the C - dimension direction. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. The channel, width, and height values in the coordinates of other elements increase along the arrow direction. Of course, the original tensor can also include D - dimension and N - dimension coordinates, which are not shown here.
[0061] For example, the shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5. Assume b1 is the N - dimension size, b2 is the D - dimension size, b3 is the H - dimension size, b4 is the W - dimension size, and b5 is the C - dimension size as an example for description. In Figure 4A In the example shown, assume b1 = 1, b2 = 1, b3 = 4, b4 = 8, b5 = 8.
[0062] For example, in Figure 4A In the example of, the dashed - box is the tensor to be processed, which includes some elements in the original tensor. The dimensions of the tensor to be processed in each dimension are represented by a1, a2, a3, a4, a5. For example, assume that the N - dimension size a5 and the D - dimension size a4 are both equal to 1. Figure 4A In the example of, the coordinates of the upper - left - corner element of the tensor to be processed are c = 0, row = 0, col = 3, and a1 = 5, a2 = 3, a3 = 3.
[0063] For example, in some other embodiments, the tensor to be processed includes at least some elements of the original tensor and tensor elements not stored in memory. For example, in the padding mode, the tensor to be processed, in addition to including at least some content of the original tensor, also includes a plurality of elements with a predetermined value (such as 0) added at the edge of the original tensor.
[0064] Figure 4B It is a schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure.
[0065] Figure 4B In, the tensor composed of multiple solid - line cubes is the original tensor. Figure 4B In, the element at the upper - left - corner of the original tensor is the origin of the coordinate system determined based on the original tensor. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. C, W, and H respectively represent the three coordinate axes, and the specific meanings are the same as those in Figure 4A and will not be elaborated here.
[0066] For example, in Figure 4BIn the example, the dashed box is the tensor to be processed, which includes some elements (solid small cubes) in the original tensor. In addition, the tensor to be processed also includes dashed small cubes outside the edge of the original tensor. The dashed small cubes are obtained, for example, through a padding operation, and their values are, for example, 0. For example, assume that both the N-dimensional size a5 and the D-dimensional size a4 are equal to 1. Figure 4B In the example, the coordinates of the upper-left corner element of the tensor to be processed are c = 0, row = 0, col = -1, and a1 = 5, a2 = 3, a3 = 3.
[0067] Of course, in some other embodiments, the tensor to be processed may not even include any elements in the original tensor, and it can perform a data loading operation using the coordinate system determined by the original tensor.
[0068] Figure 4C This is a schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure.
[0069] Figure 4C In, the tensor composed of multiple solid cubes is the original tensor. Figure 4C In, the element at the upper-left corner of the original tensor is the origin of the coordinate system determined based on the original tensor. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. C, W, and H respectively represent the three coordinate axes, and their specific meanings are the same as those in Figure 4A which will not be elaborated here.
[0070] For example, in Figure 4C In the example, the dashed box is the tensor to be processed, which does not include any elements (solid small cubes) in the original tensor. The tensor to be processed is composed of dashed small cubes outside the edge of the original tensor. The dashed small cubes are obtained, for example, through a padding operation, and their values are, for example, 0. For example, assume that both the N-dimensional size a5 and the D-dimensional size a4 are equal to 1. Figure 4C In the example, the coordinates of the upper-left corner element of the tensor to be processed are c = 0, row = 0, col = -3, and a1 = 5, a2 = 3, a3 = 3.
[0071] As Figure 3 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S10 - S30.
[0072] First, in step S10, obtain the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.
[0073] The shape and size of the original tensor are represented by b1, b2, b3, b4, and b5. b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component. The relationship between the original tensor and the tensor to be processed is as described above. For example, refer to Figure 4A - Figure 4C the relevant description, which will not be elaborated here.
[0074] For example, the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format of the tensor to be processed or the original tensor is NDHWC, or both are N(C / x)DHW(xC), where x is a positive integer greater than 1. Taking NDHWC as an example, the C dimension is the lowest dimension during storage, followed by the H dimension, then the W dimension, then the D dimension, and finally the N dimension. Taking N(C / x)DHW(xC) as an example, the W dimension is the lowest dimension during storage, followed by the H dimension, then the D dimension, then the C dimension, and finally the N dimension. And for the W dimension, continuous xC is essentially considered.
[0075] For example, taking a certain element in the original tensor as the origin of the coordinate system to determine a coordinate system. For example, refer to Figure 4A - Figure 4C the embodiment of, taking the upper left corner vertex of the original tensor as the origin of the coordinate system. Of course, the present disclosure is not limited thereto.
[0076] The starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor are, for example, Figure 4A - Figure 4C the coordinates of the upper left corner vertex of the tensor to be processed in the embodiment of. The coordinate values of the starting coordinates of the tensor to be processed are the minimum coordinate values among the coordinate values of all elements in the tensor to be processed.
[0077] In step S20, in combination with the shape and size of the original tensor, the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed, determine multiple requests for loading the tensor to be processed.
[0078] In step S30, sequentially send multiple requests, and write the sub-data returned by each request into the buffer area in order to load the tensor to be processed into the buffer area.
[0079] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, the first dimension cannot be loaded continuously when loading the tensor to be processed, but each dimension lower than the first dimension can be loaded continuously, and the request is divided on the second dimension, each sub-data requested to be loaded is located in the memory or is not located in the memory at all, and the sub-data loaded in response to the request are all located in the memory, the sub-data requested to be loaded are from the original tensor and are stored continuously in the memory, and the data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
[0080] The object to be loaded, i.e., the tensor to be processed, is split into multiple different requests according to a certain rule and sent in sequence, and the returned data is sequentially written into the cache area, thereby efficiently loading the tensor to be processed into the memory. A principle when splitting is that the sub-data loaded by each request is either all in the memory (the sub-data requested to be loaded all belong to the data range of the original tensor) or not in the memory (the sub-data requested to be loaded do not belong to the data range of the original tensor). This is because when the requested sub-data is in the memory, a request is sent to the memory. When the requested sub-data is not in the memory, the request can be converted into an operation such as writing a predetermined value to the cache area by the hardware. This splitting method can more reasonably send corresponding loading requests to different hardware.
[0081] In addition, when splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension makes it impossible to load continuously in the first dimension when loading the tensor to be processed but can load continuously in dimensions lower than the first dimension, then further division of the request is performed on the second dimension, so that the request can be split more reasonably, the data stored continuously in the memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be loaded efficiently, the bandwidth when loading or storing data from the memory can be greatly improved, the performance can be improved, and the efficiency of the hardware computing unit can be improved.
[0082] For example, if the data storage format of the tensor to be processed is NDHWC, the coordinates of each sub-data requested to be loaded meet the following conditions: the coordinates of the first dimension and each dimension below the first dimension in the tensor to be processed are different, but the coordinates of the second dimension and each dimension above the second dimension are the same, wherein the data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
[0083] For example, taking NDHWC as an example, when the first dimension is C dimension, the second dimension is W dimension, or when the first dimension is W dimension, the second dimension is H dimension, or when the first dimension is H dimension, the second dimension is D dimension, and so on.
[0084] For example, in response to the size copy_t of the tensor to be processed in the first dimension not being equal to the size tensor_t of the original tensor in the first dimension, and / or the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.
[0085] Taking the data storage format as NDHWC and the second coordinate value as 0 as an example, assuming the first dimension is the C dimension and the second dimension is the W dimension, when any of the following conditions is satisfied, it is determined that it is not continuous in the C dimension: (1) The starting coordinate c_coord_b of the tensor to be processed in the C dimension is not equal to 0 (2) The size copy_c of the tensor to be processed in the C dimension is greater than the size tensor_c of the original tensor in the C dimension (3) The size copy_c of the tensor to be processed in the C dimension is less than the size tensor_c of the original tensor in the C dimension For example, the sub-data loaded by each request has different coordinates in the C dimension, but the same coordinates in the W dimension, H dimension, D dimension, and N dimension. At this time, it can be understood that without considering the depth direction, the original tensor is unfolded into a one-dimensional vector composed of multiple pixels (pixel) in the order of WHDN, and the sub-data loaded by each request is the data within one pixel, that is, the requests are divided in the W dimension, and the next request can be to load the data in the adjacent next pixel, that is, the requests are divided in a pixel-skipping manner.
[0086] Still taking the data storage format as NDHWC and the second coordinate value as 0 as an example, assuming the first dimension is the W dimension and the second dimension is the H dimension, when any of the following conditions is satisfied, it is determined that it is not continuous in the W dimension: (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0 (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension At this time, when the C dimension is continuous, the coordinates of the sub-data to be loaded by each request are different in the C dimension and the W dimension, but the same in the H dimension, D dimension, and N dimension. At this time, it can be understood that the sub-data loaded by each request belongs to the same row, that is, the requests are divided in the H dimension, and the sub-data loaded by different requests are located in different rows.
[0087] Assuming the first dimension is the H dimension and the second dimension is the D dimension, when any of the following conditions is satisfied, it is determined that it is not continuous in the H dimension: (1)The starting coordinate of the tensor to be processed in the H dimension is not equal to 0 (2)The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension (3)The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension At this time, when both the C dimension and the W dimension are continuous, the coordinates of the sub-data to be loaded for each request are different in the C dimension, the W dimension, and the H dimension, but the coordinates in the D dimension and the N dimension are the same.
[0088] Assume that the first dimension is the D dimension and the second dimension is the N dimension. When any of the following conditions holds, it is determined that the D dimension is discontinuous: (1)The starting coordinate of the tensor to be processed in the D dimension is not equal to 0 (2)The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension (3)The size copy_d of the tensor to be processed in the D dimension is less than the size tensor_d of the original tensor in the D dimension At this time, when the C dimension, the W dimension, the H dimension, and the D dimension are all continuous, the coordinates of the sub-data to be loaded for each request are different in the C dimension, the W dimension, the H dimension, and the D dimension, but the coordinates in the N dimension are the same, that is, the requests are partitioned in the D dimension.
[0089] Assume that the first dimension is the N dimension. When any of the following conditions holds, it is determined that the N dimension is discontinuous: (1)The starting coordinate of the tensor to be processed in the N dimension is not equal to 0 (2)The size copy_n of the tensor to be processed in the N dimension is greater than the size tensor_n of the original tensor in the N dimension (3)The size copy_n of the tensor to be processed in the N dimension is less than the size tensor_n of the original tensor in the N dimension At this time, when the C dimension, the W dimension, the H dimension, and the D dimension are all continuous, the coordinates of the sub-data to be loaded for each request are the same in the CWHDN dimensions. For the tensor data located in the memory, in essence, at this time, a request can be sent to the memory to load the data in the memory, and other requests can be used to load the data not located in the memory.
[0090] If the low dimensions, that is, the C dimension, the W dimension, the H dimension, and the D dimension are discontinuous, the N dimension must also be discontinuous.
[0091] The request partitioning logic for the tensor to be processed with the data storage format of N(C / x)DHW(xC) is the same as that of NDHWC. The difference is that the dimension arrangement of N(C / x)DHW(xC) is different from that of NDHWC. For N(C / x)DHW(xC), due to its special interleaved structure, it is considered that (xC) is necessarily continuous by default. Therefore, the lowest dimension is considered to be the W dimension, followed by the H dimension, then the D dimension, then the C dimension, and the highest is the N dimension.
[0092] For example, taking the data storage format as N(C / x)DHW(xC) and the second coordinate value as 0 as an example, assuming the first dimension is the D dimension and the second dimension is the C dimension, it is determined that it is discontinuous in the D dimension when any of the following conditions is met: (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0 (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension (3) The size copy_d of the tensor to be processed in the D dimension is less than the size tensor_d of the original tensor in the D dimension At this time, when both the W dimension and the H dimension are continuous, the coordinates of the sub-data to be loaded for each request are different in the W dimension, H dimension, and D dimension, but the coordinates in the N dimension are the same. That is, the requests are partitioned in the C dimension.
[0093] In addition, for N(C / x)DHW(xC), due to the particularity of the interleaving pattern (xC), the pixels are continuous. Therefore, the first dimension from low to high is W, H, D, C, N in turn. Regarding the conditions for non - continuous loading and the request partitioning principle when the data storage format is N(C / x)DHW(xC), reference can be made to the relevant content of NDHWC, which will not be specifically listed here one by one.
[0094] The following specifically describes the determination method for multiple requests for loading the tensor to be processed.
[0095] For example, in some embodiments, step S20 may include: determining the first request sent among multiple requests and the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed; based on the initial state, combining the shape and size of the original tensor and the data storage format of the tensor to be processed, using the state machine to determine each request other than the first request among multiple requests.
[0096] Each request includes three parameters: the data read address for indicating the starting position to read data from the memory, the data write address for indicating the starting position to write data to the buffer area, and the length of the data to be loaded.
[0097] The data read address of each request is obtained by summing up the sizes of all previous requests.
[0098] For example, taking the data storage format as NDHWC, the calculation formulas for the data read address Addr_1 and the data write address Addr_2 of the nth request are as follows: Addr_1 = u_addr_base + ((((n_coord × tensor_d + d_coord) × tensor_h + h_coord) × tensor_w + (Formula 1) w_coord) × tensor_c + c_coord) Addr_2 = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 (Formula 2) Where, u_addr_base represents the storage address in memory of the element at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in 5 dimensions, b_addr_base represents the starting address of the buffer for writing the tensor to be processed, and req_size_1, req_size_2, …, req_size_n-1 represent the data lengths loaded by each of the previous n - 1 sent requests.
[0099] For example, taking the data storage format as N(C / x)DHW(xC), the calculation formulas for the data read address Addr_3 and the data write address Addr_4 of the nth request are as follows: Addr_3 = u_addr_base + n_coord × tensor_c × tensor_d × tensor_h × tensor_w + c_coord × tensor_d × tensor_h × tensor_w + d_coord × tensor_h × tensor_w × xC + (Formula 3) h_coord × tensor_w × xC + w_coord × xC Addr_4 = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 (Formula 4) Wherein, xC represents the size occupied by x channels. The definitions of other parameters are the same as those in Formula 1 and Formula 2, and the repeated parts will not be elaborated here.
[0100] For example, in some embodiments, based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, determining the first request sent among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address where the tensor to be processed is written into the buffer as the data write address of the first request; and determining the data length loaded by the first request according to the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed.
[0101] For the first request sent, its corresponding request initial coordinates are the starting coordinates of the tensor to be processed. Thus, referring to the above formula, the data read address of the first request can be calculated. The starting address where the tensor to be processed is written into the buffer is used as the data write address of the first request.
[0102] The data length loaded by the first request is determined according to the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the absolute value of the first coordinate value t_coord_b.
[0103] Considering that the state machine has the advantages of a clear logical structure, being easy to maintain and expand, and being particularly suitable for processing logical scenarios with multiple conditions and multiple branches, in the present disclosure, the state machine is adopted to automatically update the request initial coordinates corresponding to each request and the data length loaded, avoiding complex conditional nesting, with clear logic, being easy to maintain, and having strong scalability.
[0104] Figure 5 Schematic diagram of a state machine provided by an embodiment of the present disclosure.
[0105] For example, the state machine includes a first state s0, a second state s1, and a third state s2. The initial state entered into the state machine is determined by the starting coordinate coord_b of the tensor to be processed.
[0106] For example, in response to the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, the initial state is determined to be the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, the initial state is determined to be the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.
[0107] Refer to Figure 5 , the tensor_t direction represents the data in the first dimension. The data with the coordinate in the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory ( Figure 5 the black part in Figure 5 ), and the data with the coordinate in the t dimension between the second coordinate value and the third coordinate value (
[0108] the white part in Figure 5 ) is stored in memory, which is the actual size of the original tensor data in the first dimension.
[0109] Figure 5 The 6 cases (① to ⑥) in
[0110] For example, Figure 5 in case ① in
[0111] For example, Figure 5 in case ② in
[0112] For example, Figure 5 in case ③ in Figure 5 , the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1.
[0113] For example, Figure 5 in case ④ in Figure 5 , the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1, and the third state s2.
[0114] For example, Figure 5 in case ⑤ in Figure 5 , the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.
[0115] For example, Figure 5 in case ⑥ in Figure 5 , the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.
[0116] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and combine with the trigger condition to determine the state that the second sent request enters. Then, based on the state that the second sent request enters, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and combine with the trigger condition to determine the state that the third sent request enters, and so on.
[0117] For example, in some embodiments, based on the initial state, combining the shape and size of the original tensor and the data storage format of the tensor to be processed, using the state machine to determine each request except the first request among multiple requests may include: based on the initial state, using the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request; according to the shape and size of the original tensor, the request initial coordinates corresponding to each request, and the data storage format of the tensor to be processed, determine the data read address of each request; according to the data length loaded by each request, determine the data write address of each request.
[0118] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state.
[0119] After determining the request initial coordinates corresponding to the current request, it can be judged whether the sub-data loaded by the current request is in the memory according to the request initial coordinates.
[0120] For example, in response to the coordinate value of the request initial coordinate corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the sub-data to be loaded by the current request is located in the memory. At this time, the data read address and data write address of the current request can be determined with reference to the above formula 1-4, and then the current request is sent to the memory to load the corresponding sub-data into the buffer.
[0121] For example, in response to the coordinate value of the request initial coordinate corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the sub-data to be loaded by the current request is not located in the memory. At this time, in response to all the sub-data to be loaded by the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values to the buffer, where the number of the multiple predetermined values is determined by the length of the data to be loaded specified by the current request. For example, the predetermined value is 0.
[0122] The process of using the state machine to determine the request initial coordinate and the length of the data to be loaded corresponding to each request is specifically described below.
[0123] When the first coordinate value of the tensor to be processed in the first dimension is less than the second coordinate value, it enters the first state s0.
[0124] For example, when the current request is the first request, the state of the current request is the initial state, the state machine outputs the request initial coordinate and the length of the data to be loaded corresponding to the current request, and prepares the request initial coordinate for the next state.
[0125] For example, in response to the state of the current request being the first state and satisfying the first condition, it is determined that the next request of the current request enters the first state. For example, the first condition includes that the sum of the size copy_t of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value. Taking the second coordinate value as 0 as an example, the first condition includes that the size copy_t of the tensor to be processed in the first dimension is less than the absolute value of the first coordinate value.
[0126] For example, if the current request is in the first state s0 and the size copy_t of the tensor to be processed in the first dimension is relatively small, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value, the state of the next request is still the first state s0. For example, in case ① in Figure 5 s0->s0.
[0127] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values in the lower other dimensions.
[0128] Specifically, for other dimensions, the update methods include incrementing by 1 or x, returning to the initial coordinates, and remaining unchanged. For example, for the second dimension, the coordinates of the second dimension for each request are incremented, and the third dimension and higher dimensions remain unchanged until the boundary of the tensor to be processed in the second dimension is reached. After that, the coordinates of the second dimension are updated back to the starting coordinates and the coordinates of the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension is incremented in this way. The coordinates of other dimensions lower than the first dimension remain unchanged.
[0129] For example, in this case, the length of the data loaded by the current request, req_size, is the size copy_t of the tensor to be processed in the first dimension.
[0130] For example, in response to the status of the current request being the first status and satisfying the second condition, it is determined that the next request of the current request enters the second status s1. For example, the second condition includes that the size copy_t of the tensor to be processed in the first dimension is greater than or equal to the absolute value of the first coordinate value t_coord_b. For example, in cases ② and ④ in Figure 5 s0->s1.
[0131] For example, if the current request is in the first status s0 and the size copy_t of the tensor to be processed in the first dimension is relatively large, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, the status of the next request jumps to the second status s1. For example, assuming the second coordinate value is 0, the second condition includes that the size copy_t of the tensor to be processed in the first dimension is greater than or equal to the absolute value of the first coordinate value t_coord_b.
[0132] In this case, it is determined that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is the second coordinate value, for example, 0, and it is determined that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension remain unchanged. Since it is a jump from the first status to the second status and the data in the first dimension has not been fully loaded, it is determined that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension remain unchanged.
[0133] For example, in this case, the length of the data loaded by the current request, req_size, is determined based on the first coordinate value and the second coordinate value, for example, it is the difference between the first coordinate value and the second coordinate value.
[0134] For example, in response to the state of the current request being the second state s1 and satisfying the third condition, it is determined that the next request of the current request enters the first state s0. The third condition includes that the first coordinate value t_coord_b is less than the second coordinate value and the sum of the first coordinate value t_coord_b and the size copy_t of the tensor to be processed in the first dimension is less than the third coordinate value (for example, the second coordinate value is 0, and the third coordinate value is the size tensor_t of the original tensor in the first dimension), for example Figure 5 In case ②, s1->s0. The sum of the first coordinate value t_coord_b and the size copy_t of the first dimension of the tensor to be processed is less than the third coordinate value, so the third state s2 will not be entered.
[0135] In this case, the coordinate value of the initial coordinate of the request corresponding to the next request in the first dimension is determined to be the first coordinate value t_coord_b, and the coordinate value of the initial coordinate of the request corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions. Specifically, for other dimensions, the update methods include self-increment (1 or x), return to the initial coordinate, and unchanged. For example, for the second dimension, the coordinate of the second dimension of each request is self-incremented, and the third dimension and higher dimensions remain unchanged until the tensor to be processed reaches the boundary of the second dimension, after which the coordinate of the second dimension is updated back to the starting coordinate and the coordinate of the third dimension is self-incremented. The third dimension is adjacent to the second dimension and higher than the second dimension, and each dimension increases in this way. The coordinates of other dimensions lower than the first dimension remain unchanged.
[0136] For example, in this case, the data length req_size currently requested to be loaded is determined based on the first coordinate value and the size copy_t of the tensor to be processed in the first dimension. For example, if the second coordinate value is 0, the data length req_size currently requested to be loaded is the sum of the first coordinate value and the size copy_t of the tensor to be processed in the first dimension.
[0137] For example, the state of the current request is the second state s1 and satisfies the fourth condition, and the next request is determined to enter the second state s1. The fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension. For example, if the current request is in the second state, it means that the requested sub-data are all in the memory, that is, they are all valid data. When the size copy_t of the tensor to be processed in the first dimension is relatively small, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the third coordinate value, the state of the next request is still the second state s1 and will not jump to the third state s2. For example, Figure 5 Situation ③.
[0138] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those in other dimensions. Specifically, for other dimensions, the update methods include incrementing (by 1 or x), returning to the initial coordinate, and remaining unchanged. For example, for the second dimension, the coordinates of each request in the second dimension are incremented, and the third dimension and higher dimensions remain unchanged until reaching the boundary of the tensor to be processed in the second dimension. After that, the coordinates in the second dimension are updated back to the starting coordinate and the coordinates in the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension is incremented in this way. The coordinate values of other dimensions lower than the first dimension remain unchanged.
[0139] For example, in this case, the data length req_size loaded by the current request is the size copy_t of the tensor to be processed in the first dimension.
[0140] For example, in response to the status of the current request being the second status s1 and satisfying the fifth condition, it is determined that the next request enters the third status s2. The fifth condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the third coordinate value. At this time, the size copy_t of the tensor to be processed in the first dimension is relatively large, and the status of the next request jumps to the third status s2. For example, in Figure 5 cases ④ and ⑤ in
[0141] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the third coordinate value, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged.
[0142] For example, in this case, the data length req_size loaded by the current request is determined based on the coordinate value t_coord and the third coordinate value of the request initial coordinate corresponding to the current request in the first dimension. For example, the data length req_size loaded by the current request is the difference between the third coordinate value and the coordinate value t_coord.
[0143] For example, in response to the status of the current request being the third status, the status that the next request of the current request enters is determined according to the first coordinate value.
[0144] For example, when the first coordinate value is less than the second coordinate value, it is determined that the next request enters the first status s0. For example, in Figure 5 case ④ in Figure 5In case ⑤ in [description], s2 -> s1. When the first coordinate value is greater than or equal to the third coordinate value, it is determined that the next request enters the third state s2. For example Figure 5 In case ⑥ in [description], s2 -> s2.
[0145] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those in other dimensions. Specifically, for other dimensions, the update methods include incrementing (by 1 or x), returning to the initial coordinate, or remaining unchanged. For example, for the second dimension, the coordinates of each request in the second dimension are incremented until reaching the boundary of the tensor to be processed in the second dimension, and then the coordinates are updated back to the starting coordinate and the coordinates of the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension is incremented in this way. The coordinate values of other dimensions lower than the first dimension remain unchanged.
[0146] For example, in this case, the data length req_size loaded by the current request is determined based on the coordinate value and the third coordinate value of the request initial coordinate corresponding to the current request in the first dimension. For example, the data length req_size loaded by the current request is the difference between the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the third coordinate value.
[0147] In some other embodiments, the state machine can also update the remaining size remain_copy_t of 5 dimensions. The remaining size represents the amount of data that has not been loaded in each dimension. When judging the state transition condition of the state machine, it is judged in combination with the relationship between the fourth coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the remaining size.
[0148] For example, taking the transition from the first state s0 to the first state s0 as an example, the first condition may include that the sum of the fourth coordinate value t_coord of the request initial coordinate corresponding to the current request in the first dimension and the remaining size remain_copy_t of the tensor to be processed in the first dimension is less than the second coordinate value, and the fourth coordinate value t_coord is less than the second coordinate value. The remaining size remain_copy_t is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded.
[0149] For example, taking the transition from the first state s0 to the second state s1 as an example, the second condition may include that the sum of the fourth coordinate value t_coord of the request initial coordinate corresponding to the current request in the first dimension and the remaining size remain_copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value.
[0150] For example, taking the jump from the second state s1 to the first state s0 as an example, the third condition may include that the sum of the fourth coordinate value t_coord of the request initial coordinates corresponding to the current request in the first dimension and the remaining size remain_copy_t of the tensor to be processed in the first dimension is less than or equal to the third coordinate value, and the first coordinate value is less than the second coordinate value.
[0151] For example, taking the jump from the second state s1 to the second state s1 as an example, the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value.
[0152] For example, taking the jump from the second state s1 to the third state s2 as an example, the fifth condition may include that the sum of the request initial coordinates corresponding to the current request and the remaining size remain_copy_t of the tensor to be processed in the first dimension is greater than the third coordinate value.
[0153] For example, for the third state s2, the state into which the next request of the current request enters is determined according to the first coordinate value, and specific reference can be made to the foregoing content.
[0154] Of course, as the coordinates continue to increase, when the remaining sizes of all dimensions are equal to 1, it is determined that the state machine can end the coordinate determination of each request, and the state machine can enter the Idle state.
[0155] The state transition conditions of the specific state machine and the updated coordinate content can be changed and set according to actual needs. The logic is similar to the state machine logic described above, and no further examples are given here.
[0156] It should be noted that the process of using the state machine to determine the request initial coordinates and the loaded data length corresponding to the request is basically the same for NDHWC and N(C / x)DHW(xC). The difference is that in response to the data storage format of the tensor to be processed being NDHWC, the coordinate value in the channel number dimension is incremented by 1 when updated, and in response to the data storage format of the tensor to be processed being N(C / x)DHW(xC), the coordinate value in the channel number dimension is incremented by x when updated.
[0157] For example, taking the data storage format as NDHWC below as an example, the data loading process when the first dimension is the C dimension is described. In this example, the first coordinate value is the coordinate value c_coord_b of the starting coordinate of the tensor to be processed in the C dimension, the second coordinate value is 0, the third coordinate value is the size tensor_c of the original tensor in the C dimension, and the predetermined value is 0.
[0158] For example, when the size copy_c of the tensor to be processed in the C dimension is less than or greater than the size tensor_c of the original tensor in the C dimension, and / or the first coordinate value c_coord_b is not equal to 0, it is determined that the tensor to be processed cannot be continuously loaded in the C dimension.
[0159] First, in step S10, obtain the shape dimensions of the tensor to be processed, i.e., copy_c / h / w / d / n, etc., the data storage format of the tensor to be processed, i.e., NDHWC, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, i.e., c / h / w / d / n_coord_b.
[0160] After that, in step S20, determine multiple requests for loading the tensor to be processed in combination with the above information.
[0161] Referring to the process described above, use a state machine to determine the request initial coordinates and the data length to be loaded corresponding to each request, and accordingly determine the data read address and data write address of each request. For example, in this example, the coordinates of the sub-data loaded by each request are different in the C dimension, but the same in the H dimension, W dimension, D dimension, and N dimension.
[0162] Table 1 shows the update process of the request initial coordinates and the data length to be loaded provided in an embodiment of the present disclosure.
[0163] Table 1
[0164] As shown in Table 1, if the first coordinate value c_coord_b is less than 0, it is determined that the first request enters the first state s0. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0165] If the first condition is satisfied, it is determined that the next request enters the first state s0. For the first condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the first coordinate value c_coord_b; the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is incremented by 1, and when it reaches the boundary of the to-be-processed tensor in the W dimension, it returns w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1 when w_coord returns w_coord_b. If it does not reach the boundary of the to-be-processed tensor in the H dimension, it remains unchanged, and when it reaches the boundary of the to-be-processed tensor in the H dimension, it returns h_coord_b; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the to-be-processed tensor in the D dimension, it remains unchanged, and when it reaches the boundary of the to-be-processed tensor in the D dimension, it returns d_coord_b; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, and remains unchanged in other cases.
[0166] Moreover, at this time, the data length loaded by the first request is the size copy_c of the to-be-processed tensor in the C dimension.
[0167] If the second condition is satisfied, it is determined that the next request enters the second state s1. For the second condition, refer to the foregoing description. At this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is 0; the coordinate values w_coord, h_coord, d_coord, and n_coord of the request initial coordinates corresponding to the next request in the W, H, D, and N dimensions respectively remain unchanged.
[0168] Moreover, at this time, the data length loaded by the first request is the absolute value of the first coordinate value.
[0169] Since the coordinate value c_coord of the request initial coordinate corresponding to the first request in the C dimension is less than 0, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing copy_c or |c_coord_b| zeros to the buffer area, where |c_coord_b| represents the absolute value of c_coord_b.
[0170] In addition, it should be noted that if it is determined that the coordinate value c_coord of the request initial coordinate corresponding to the request in the C dimension is less than 0, there is no need to calculate the data read address and data write address of the request, thus reducing the amount of calculation.
[0171] As shown in Table 1, when the first coordinate value c_coord_b is greater than or equal to 0 and less than tensor_c, it is determined that the first request enters the second state s1. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0172] If the third condition is satisfied, it is determined that the next request enters the first state s0. For the third condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the first coordinate value c_coord_b; the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is incremented by 1 and returns w_coord_b when reaching the boundary of the tensor to be processed in the W dimension; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1 when w_coord returns w_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the H dimension, and returns h_coord_b when reaching the boundary of the tensor to be processed in the H dimension; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the D dimension, and returns d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, and remains unchanged in other cases.
[0173] Moreover, at this time, the data length loaded by the first request is the sum of the size copy_c of the tensor to be processed in the C dimension and the first coordinate value c_coord_b.
[0174] If the fourth condition is satisfied, it is determined that the next request enters the second state s1. For the fourth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the first coordinate value c_coord_b; the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is incremented by 1, and when it reaches the boundary of the tensor to be processed in the W dimension, it returns w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1 when w_coord returns w_coord_b. If it does not reach the boundary of the tensor to be processed in the H dimension, it remains unchanged, and when it reaches the boundary of the tensor to be processed in the H dimension, it returns h_coord_b; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the tensor to be processed in the D dimension, it remains unchanged, and when it reaches the boundary of the tensor to be processed in the D dimension, it returns d_coord_b; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, and remains unchanged in other cases.
[0175] Moreover, at this time, the data length loaded by the first request is the size copy_c of the tensor to be processed in the C dimension.
[0176] If the fifth condition is satisfied, it is determined that the next request enters the third state s2. For the fifth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the third coordinate value tensor_c; the coordinate values w_coord, h_coord, d_coord of the request initial coordinates corresponding to the next request in the W, H, and D dimensions, and the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension all remain unchanged.
[0177] Moreover, at this time, the data length loaded by the first request is the difference between the size tensor_c of the original tensor in the C dimension and the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension.
[0178] Since the state of the first request is the second state s1 and the sub-data it loads is all in the memory, in step S30, the first request is sent to the memory to load the sub-data in the original tensor stored in the memory into the memory.
[0179] As shown in Table 1, when the first coordinate value c_coord_b is greater than or equal to tensor_c, it is determined that the first request enters the third state s2. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0180] Determine the state that the next request enters according to the first coordinate value c_coord_b. For example, when the first coordinate value c_coord_b is less than 0, it is determined that the next request enters the first state s0; when the first coordinate value c_coord_b is greater than or equal to 0 but less than tensor_c, it is determined that the next request enters the second state s1; when the first coordinate value c_coord_b is greater than or equal to tensor_c, it is determined that the next request enters the third state s2.
[0181] As shown in Table 1, regardless of which state the next request enters, the update logic of the request initial coordinates and the data length loaded by the first request are exactly the same.
[0182] For example, taking the next request entering the first state s0 as an example, at this time, the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is the first coordinate value c_coord_b; the coordinate value w_coord of the request initial coordinates corresponding to the next request in the W dimension is incremented by 1 and returns w_coord_b when reaching the boundary of the tensor to be processed in the W dimension; the coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is incremented by 1 when w_coord returns w_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the H dimension, and returns h_coord_b when reaching the boundary of the tensor to be processed in the H dimension; the coordinate value d_coord of the request initial coordinates corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the D dimension, and returns d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinates corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, and remains unchanged in other cases. The data length loaded by the first request is the difference between the size tensor_c of the original tensor in the C dimension and the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension.
[0183] Since the coordinate value c_coord of the request initial coordinates corresponding to the first request in the C dimension is greater than tensor_c, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing c_coord - tensor_c zeros to the buffer area.
[0184] After that, continue the above process to determine the request initial coordinates and the lengths of the data to be loaded corresponding to subsequent requests, and execute step S30 to send multiple requests. The specific process will not be elaborated here.
[0185] For example, taking the data storage format of N(C / x)DHW(xC) as an example, the data loading process when the first dimension is the W dimension will be described. In this example, the first coordinate value is the coordinate value w_coord_b of the starting coordinate of the tensor to be processed in the W dimension, the second coordinate value is 0, the third coordinate value is the size tensor_w of the original tensor in the W dimension, and the predetermined value is 0.
[0186] For example, if the size copy_w of the tensor to be processed in the W dimension is less than or greater than the size tensor_w of the original tensor in the W dimension, and / or the first coordinate value w_coord_b is not equal to 0, it is determined at this time that the tensor to be processed cannot be continuously loaded in the W dimension.
[0187] For example, first, in step S10, obtain the shape and size of the tensor to be processed, i.e., copy_c / h / w / d / n, etc., the data storage format of the tensor to be processed, i.e., N(C / x)DHW(xC), and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor, i.e., c / h / w / d / n_coord_b.
[0188] After that, in step S20, determine multiple requests for loading the tensor to be processed in combination with the above information.
[0189] Referring to the process described above, use a state machine to determine the request initial coordinates and the lengths of the data to be loaded corresponding to each request, and accordingly determine the data read address and data write address for each request. For example, in this example, the coordinates of the sub-data loaded by each request in the W dimension are different, but the coordinates in the H dimension, D dimension, and N dimension are the same. Also, the coordinates of the sub-data loaded by each request in the C dimension are different, which is caused by xC in the interleaving mode.
[0190] Table 2 shows the update process of the request initial coordinates and the lengths of the data to be loaded provided by another embodiment of the present disclosure.
[0191] Table 2
[0192] As shown in Table 2, if the first coordinate value w_coord_b is less than 0, it is determined that the first request enters the first state s0. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0193] If the first condition is satisfied, it is determined that the next request enters the first state s0. Refer to the foregoing description for the first condition. As shown in Table 2, at this time, the coordinate value w_coord of the initial request coordinate corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the initial request coordinate corresponding to the next request in the H dimension is incremented by 1, and when it reaches the boundary of the tensor to be processed in the H dimension, it returns h_coord_b; the coordinate value d_coord of the initial request coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the tensor to be processed in the D dimension, it remains unchanged, and when it reaches the boundary of the tensor to be processed in the D dimension, it returns d_coord_b; the coordinate value c_coord of the initial request coordinate corresponding to the next request in the C dimension is incremented by x when d_coord returns d_coord_b. If it does not reach the boundary of the tensor to be processed in the C dimension, it remains unchanged, and when it reaches the boundary of the tensor to be processed in the C dimension, it returns c_coord_b; the coordinate value n_coord of the initial request coordinate corresponding to the next request in the N dimension is incremented by 1 when c_coord returns c_coord_b, and remains unchanged in other cases.
[0194] Moreover, at this time, the data length loaded by the first request is the size copy_w of the tensor to be processed in the W dimension.
[0195] If the second condition is satisfied, it is determined that the next request enters the second state s1. Refer to the foregoing description for the second condition. At this time, the coordinate value w_coord of the initial request coordinate corresponding to the next request in the W dimension is 0; the coordinate values h_coord of the initial request coordinate corresponding to the next request in the H dimension, d_coord of the initial request coordinate corresponding to the next request in the D dimension, c_coord of the initial request coordinate corresponding to the next request in the C dimension, and n_coord of the initial request coordinate corresponding to the next request in the N dimension all remain unchanged.
[0196] Moreover, at this time, the data length loaded by the first request is the absolute value of the first coordinate value.
[0197] Since the coordinate value w_coord of the initial request coordinate corresponding to the first request in the W dimension is less than 0, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing copy_w or |w_coord_b| zeros to the buffer area.
[0198] As shown in Table 2, when the first coordinate value w_coord_b is greater than or equal to 0 and less than tensor_w, it is determined that the first request enters the second state s1. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0199] If the third condition is satisfied, it is determined that the next request enters the first state s0. Refer to the foregoing description for the third condition. As shown in Table 2, at this time, the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1, and when reaching the boundary of the W dimension of the tensor to be processed, h_coord_b is returned; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If the boundary of the D dimension of the tensor to be processed is not reached, it remains unchanged, and when reaching the boundary of the D dimension of the tensor to be processed, d_coord_b is returned; the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is incremented by x when d_coord returns d_coord_b. If the boundary of the C dimension of the tensor to be processed is not reached, it remains unchanged, and when reaching the boundary of the C dimension of the tensor to be processed, c_coord_b is returned; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when c_coord returns c_coord_b, and remains unchanged in other cases.
[0200] Moreover, at this time, the data length loaded by the first request is the sum of the size copy_w of the tensor to be processed in the W dimension and the first coordinate value w_coord_b.
[0201] If the fourth condition is satisfied, it is determined that the next request enters the second state s1. Refer to the foregoing description for the fourth condition. As shown in Table 2, at this time, the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1, and when reaching the boundary of the W dimension of the tensor to be processed, h_coord_b is returned; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If the boundary of the D dimension of the tensor to be processed is not reached, it remains unchanged, and when reaching the boundary of the D dimension of the tensor to be processed, d_coord_b is returned; the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is incremented by x when d_coord returns d_coord_b. If the boundary of the C dimension of the tensor to be processed is not reached, it remains unchanged, and when reaching the boundary of the C dimension of the tensor to be processed, c_coord_b is returned; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when c_coord returns c_coord_b, and remains unchanged in other cases.
[0202] Moreover, at this time, the length of the data loaded by the first request is the size copy_w of the to-be-processed tensor in the W dimension.
[0203] If the fifth condition is met, it is determined that the next request enters the third state s2. For the fifth condition, refer to the foregoing description. As shown in Table 2, at this time, the coordinate value w_coord of the initial request coordinate corresponding to the next request in the W dimension is the third coordinate value tensor_w; the coordinate value h_coord of the initial request coordinate corresponding to the next request in the H dimension, the coordinate value d_coord of the initial request coordinate corresponding to the next request in the D dimension, the coordinate value c_coord of the initial request coordinate corresponding to the next request in the C dimension, and the coordinate value n_coord of the initial request coordinate corresponding to the next request in the N dimension remain unchanged.
[0204] Moreover, at this time, the length of the data loaded by the first request is the difference between the size tensor_w of the original tensor in the W dimension and the coordinate value w_coord of the initial request coordinate corresponding to the next request in the W dimension.
[0205] Since the state of the first request is the second state s1 and the sub-data it loads are all in the memory, in step S30, the first request is sent to the memory to load the sub-data in the original tensor stored in the memory into the memory.
[0206] As shown in Table 2, when the first coordinate value w_coord_b is greater than or equal to tensor_w, it is determined that the first request enters the third state s2. The data reading address and data writing address of the first request refer to the foregoing description and will not be elaborated here.
[0207] Determine the state that the next request enters according to the first coordinate value w_coord_b. For example, when the first coordinate value w_coord_b is less than 0, it is determined that the next request enters the first state s0; when the first coordinate value w_coord_b is greater than or equal to 0 but less than tensor_w, it is determined that the next request enters the second state s1; when the first coordinate value w_coord_b is greater than or equal to tensor_w, it is determined that the next request enters the third state s2.
[0208] As shown in Table 2, regardless of which state the next request enters, the update logic of the initial request coordinates and the length of the data loaded by the first request are exactly the same.
[0209] For example, taking the case where the following request enters the first state s0, at this time, the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1, and when it reaches the boundary of the to-be-processed tensor in the W dimension, it returns h_coord_b; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b, and remains unchanged if it does not reach the boundary of the to-be-processed tensor in the D dimension, and returns d_coord_b when it reaches the boundary of the to-be-processed tensor in the D dimension; the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is incremented by x when d_coord returns d_coord_b, and remains unchanged if it does not reach the boundary of the to-be-processed tensor in the C dimension, and returns c_coord_b when it reaches the boundary of the to-be-processed tensor in the C dimension; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when c_coord returns c_coord_b, and remains unchanged in other cases. The data length loaded by the first request is the difference between the size tensor_w of the original tensor in the W dimension and the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension.
[0210] Since the coordinate value w_coord of the request initial coordinate corresponding to the first request in the W dimension is greater than tensor_w, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing w_coord - tensor_w zeros to the buffer area.
[0211] After that, continue the above process, determine the request initial coordinates and the data lengths loaded by subsequent requests, and execute step S30 to send multiple requests. The specific process will not be elaborated here.
[0212] In the above embodiment, requests are divided according to whether the tensor can be continuously loaded in the first dimension. Each request is used to load sub-data that are all in the memory or all not in the memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improves the bandwidth and efficiency during data access, and further improves the hardware utilization rate of the computing unit and the hardware performance.
[0213] At least one embodiment of the present disclosure further provides a data storage method for writing the to-be-processed tensor in the buffer area into the memory.
[0214] For example, the shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, and a5. a1, a2, a3, a4, and a5 respectively indicate the dimensions of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. For the relevant content of the tensor to be processed, reference can be made to the description of the foregoing data loading method, which will not be elaborated here.
[0215] Figure 6 Schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure.
[0216] As Figure 6 shown, the data storage method provided by at least one embodiment of the present disclosure at least includes steps S40 - S60.
[0217] First, in step S40, obtain the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.
[0218] The shape dimensions of the original tensor are represented by b1, b2, b3, b4, and b5. b1, b2, b3, b4, and b5 respectively indicate the dimensions of the original tensor in five dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0219] For the relevant descriptions of the original tensor and the relationship between the original tensor and the tensor to be processed, reference can be made to the relevant content of the foregoing data loading method, which will not be elaborated here.
[0220] For example, the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format of the tensor to be processed or the original tensor is both NDHWC, or both N(C / x)DHW(xC), where x is a positive integer greater than 1.
[0221] In step S50, combine the original tensor, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine multiple requests for storing the tensor to be processed.
[0222] In step S60, sequentially send at least some of the multiple requests, and write the data to be written indicated by each request into the memory in sequence to store the tensor to be processed into the memory.
[0223] For example, in some embodiments, step S50 may include: determining the first request sent among multiple requests and the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed; and determining each request other than the first request among the multiple requests by using the state machine based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed.
[0224] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension, requests are divided in the second dimension, and the sub-data to be written by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory.
[0225] For example, referring to Figure 5 , the tensor_t direction represents the data in the first dimension. The data with the coordinate of the t dimension less than the second coordinate value and greater than the third coordinate value is not in the memory, or rather, the data range of this part of the t dimension is not within the data range of the original tensor ( Figure 5 the part shown in black background in Figure 5 This part of the data can be called invalid data). The data with the coordinate of the t dimension between the second coordinate value and the third coordinate value (
[0226] the white part in Figure 5 ) is stored in the memory, or rather, the data range of this part of the t dimension is within the data range of the original tensor.
[0227] For example, the sub-data to be written all belonging to the data range of the original tensor means the coordinate range of the sub-data in any dimension. For example, referring to Figure 5 the t dimension in
[0228] is between the corresponding second coordinate value and the third coordinate value, that is, it belongs to the original tensor. At this time, the sub-data to be written, as a part of the original tensor, needs to be stored in the memory.
[0229] Such as Figure 4AIn the example shown, the tensor to be processed consists of partial tensors of the original tensor. At this time, all tensors to be processed in the buffer need to be stored at the corresponding positions in the original tensor to update the original tensor. Therefore, multiple requests need to be sent to the memory at this time to store the sub-data for each request at the corresponding positions in the memory.
[0230] As Figure 4B shown in the example, the tensor to be processed not only includes partial tensors of the original tensor, but also includes data outside the edge of the original tensor caused by padding operations and the like that are not stored in the memory. Therefore, requests need to be distinguished before being sent. If the sub-data for a request to be stored does not belong to the content of the original tensor, such as Figure 4B the small dotted cube in, there is no need to send a request, and these sub-data do not need to be stored in the memory; if the sub-data for a request to be stored belongs to the original tensor, such as Figure 4B the solid small cube in the tensor to be processed that belongs to the original tensor, a request is sent to the memory to store the sub-data at the corresponding position in the original tensor.
[0231] As Figure 4C shown in the example, the tensor to be processed does not include any elements of the original tensor. Then, there is no need to send a request at this time.
[0232] For example, in some embodiments, step S60 may include: for any request, in response to the first request coordinate value of the request initial coordinate corresponding to any request in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determining not to send any request; in response to the first request coordinate value being greater than or equal to the second coordinate value and the first request coordinate value being less than the third coordinate value, determining to send any request; where the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0233] For example, if the first request coordinate value of the request initial coordinate corresponding to the request is less than the second coordinate value in the first dimension, or the first request coordinate value is greater than or equal to the third coordinate value, this indicates that the sub-data for the request to be stored does not belong to the original tensor and does not need to be stored in the memory, and this request is not sent.
[0234] For example, if the first request coordinate value is greater than or equal to the second coordinate value and the first request coordinate value is less than the third coordinate value, this indicates that the sub-data for the request to be stored belongs to the original tensor and needs to be stored in the memory, and this request is sent to store the sub-data at the corresponding position in the original tensor.
[0235] The data storage method provided by at least one embodiment of the present disclosure divides requests according to whether tensors can be continuously stored in the first dimension. Each sub-data used for storage in each request is either entirely located in memory or not located in memory at all. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby enhancing the hardware utilization rate of the computing unit and improving the hardware performance.
[0236] At least one embodiment of the present disclosure also provides a data loading method. Figure 7 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure.
[0237] As Figure 7 shown, the data loading method provided by at least one embodiment of the present disclosure includes steps S201 - S202.
[0238] For example, in step S201, a data loading instruction for instructing to load the tensor to be processed into the buffer is received.
[0239] For example, the data loading instruction includes the shape and size of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0240] For the relevant descriptions of the tensor to be processed and the original tensor, reference can be made to the relevant descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0241] For example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction can be a Load instruction.
[0242] In step S202, after parsing the data loading instruction, the execution unit is used to execute the data loading instruction.
[0243] For example, after receiving a data loading instruction, the processor parses the data loading instruction, such as decoding the data loading instruction, generating microinstructions, and sending the microinstructions to the instruction distribution unit; the instruction distribution unit sends them to the corresponding scheduling queue according to the microinstruction category; in response to the microinstructions, when the input parameters are ready, the execution unit performs the relevant operations of the data loading instruction.
[0244] For example, step S202 may include: determining multiple requests for loading the tensor to be processed in combination with the shape and size of the original tensor, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending the multiple requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area.
[0245] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension, the requests are divided in the second dimension, the sub-data loaded by each request are all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request are from the original tensor and are continuously stored in the memory, the data storage format of the tensor to be processed indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.
[0246] Regarding "determining multiple requests for loading the tensor to be processed in combination with the shape and size of the original tensor, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed", reference can be made to the relevant content of the foregoing step S20, which will not be elaborated here.
[0247] Regarding "sequentially sending the multiple requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area", reference can be made to the relevant content of the foregoing step S30, which will not be elaborated here.
[0248] In the present disclosure, the data loading instruction is divided into multiple requests. According to whether the tensor can be continuously loaded in the first dimension, the data loading instruction is converted into multiple requests, and the sub-data for loading by each request are all located in the memory or none of them are located in the memory. The division of the data requests is more reasonable, more suitable for the storage of tensor data, greatly improves the bandwidth and efficiency during data access, and further improves the hardware utilization rate of the computing unit and the hardware performance.
[0249] At least one embodiment of the present disclosure further provides a data storage method. Figure 8 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure.
[0250] As Figure 8As shown, the data storage method provided by at least one embodiment of the present disclosure includes steps S203 - S204.
[0251] For example, in step S203, a data storage instruction indicating to execute writing a tensor to be processed in a buffer into memory is received.
[0252] For example, the data storage instruction includes the shape dimensions of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0253] For the relevant descriptions of the tensor to be processed and the original tensor, reference can be made to the relevant descriptions of the foregoing data storage method, and the repeated parts will not be elaborated.
[0254] For example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction can be a Store instruction.
[0255] In step S204, after parsing the data storage instruction, the execution unit is used to execute the data storage instruction.
[0256] For example, after receiving the data storage instruction, the processor parses the data storage instruction, for example, decodes the data storage instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameters are ready, the execution unit executes the relevant operations of the data storage instruction.
[0257] For example, step S204 may include: combining the original tensor, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine multiple requests for storing the tensor to be processed; sequentially sending at least some of the multiple requests, and writing the data to be written indicated by each request into memory in order to store the tensor to be processed.
[0258] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension. The requests are partitioned in the second dimension, and the sub-data to be written by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. Moreover, in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory.
[0259] Regarding "determining multiple requests for storing the tensor to be processed in combination with the shape and size of the original tensor and the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed", reference can be made to the relevant content of the foregoing step S50, which will not be elaborated here.
[0260] Regarding "sequentially sending at least some of the multiple requests and sequentially writing the data to be written indicated by each request into the memory to store the tensor to be processed", reference can be made to the relevant content of the foregoing step S60, which will not be elaborated here.
[0261] In the present disclosure, the data storage instruction is partitioned into multiple requests. According to whether the tensor can be continuously stored in the first dimension, the data storage instruction is converted into multiple requests, and the sub-data to be stored by each request is either all located in the memory or all not located in the memory. The partitioning of the data requests is more reasonable, more suitable for storing tensor data, greatly improving the bandwidth and efficiency during data memory access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.
[0262] Figure 9 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 9 shown, the electronic device 300 is, for example, suitable for implementing the data loading method or data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 9 the components of the electronic device 300 shown are only exemplary and not restrictive. According to actual application requirements, the electronic device 300 may also have other components.
[0263] As Figure 9 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.
[0264] For example, when the computer-readable instructions are run by the processing device 301, one or more steps in the data loading method according to any of the above embodiments, or one or more steps in the data storage method according to any of the above embodiments, can be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference can be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference can be made to the relevant descriptions in the embodiments of the data storage method above.
[0265] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory, etc. For example, computer-readable instructions can be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data can also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.
[0266] For example, the processing device 301, read-only memory (ROM) 302, and random access memory (RAM) 303 are connected to each other through the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0267] Generally, the following devices can be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 9An electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 may alternatively implement or have more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions. The processing device 301 may be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU, etc., which has data processing capabilities and / or program execution capabilities. The central processing unit (CPU) may be of the X86, ARM, RISC-V architecture, etc. The GPU may be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.
[0268] Figure 10 Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 10 shown, the processor 400 includes an instruction parsing unit 401 and a first execution unit 402.
[0269] For example, the instruction parsing unit 401 is used to receive and parse data loading instructions.
[0270] For example, the data loading instruction includes the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0271] For the relevant descriptions of the tensor to be processed and the original tensor, reference may be made to the relevant descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0272] For example, after the instruction parsing unit parses the data loading instruction, the first execution unit 402 executes the data loading instruction.
[0273] For example, when the first execution unit 402 executes a data loading instruction, it includes performing the following operations: determining a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area.
[0274] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, the requests are divided in the second dimension, and the sub-data loaded by each request is either all located in the memory or none of it is located in the memory. And in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory.
[0275] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send a data loading instruction for computing and processing to the processor (such as a CPU or GPU) through a unified encapsulated function library, the data loading instruction can carry the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters; when the processor receives the data loading instruction, the instruction parsing unit 401 parses the data loading instruction to obtain the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data loading instruction, the processor can store the input parameters in the data loading instruction in a register or memory, and when the first execution unit 402 performs computing and processing, it can obtain the input parameters from the register or memory.
[0276] Regarding the specific process of using the first execution unit 402 to execute the data loading instruction, reference can be made to steps S20 - S30 in the data loading method described above, and the repeated parts will not be elaborated.
[0277] The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method, and the repeated parts will not be elaborated.
[0278] Figure 11 It is a schematic structural diagram of the processor provided in at least one embodiment of the present disclosure. As Figure 11 shown, the processor 500 includes an instruction parsing unit 501 and a second execution unit 502.
[0279] For example, the instruction parsing unit 501 is used to receive and parse data storage instructions.
[0280] For example, the data storage instruction includes the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, as input parameters. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The data storage format of the original tensor is the same as that of the tensor to be processed. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0281] For the related descriptions of the tensor to be processed and the original tensor, reference can be made to the related descriptions of the aforementioned data loading method, and the repeated parts will not be elaborated.
[0282] For example, after the instruction parsing unit parses the data storage instruction, the second execution unit 502 executes the data storage instruction.
[0283] For example, when the second execution unit 502 executes the data storage instruction, it includes performing the following operations: determining a plurality of requests for storing the tensor to be processed by combining the original tensor, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed.
[0284] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, such that when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension, requests are divided in the second dimension. The sub-data to be written by each request all belong to the data range of the original tensor or all do not belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory.
[0285] Specifically, when upper-layer software based on a processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data storage instructions for computing and processing to the processor (such as a CPU or GPU) through a unified encapsulated function library, the data storage instructions can carry the shape and size of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; when the processor receives the data storage instruction, the instruction parsing unit 401 parses the data storage instruction to obtain the shape and size of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, and the processor schedules the operation unit to execute the data loading task for the input parameter. For example, after parsing the data storage instruction, the processor can store the input parameter in the data storage instruction into a register or memory, and when the second execution unit 502 executes the computing and processing, it can obtain the input parameter from the register or memory.
[0286] Regarding the specific process of using the second execution unit 502 to execute the data storage instruction, reference can be made to steps S50 - S60 in the data loading method described above, and the repeated parts will not be elaborated.
[0287] The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects to the aforementioned data storage method, and the repeated parts will not be elaborated.
[0288] Figure 12 It is a schematic diagram of a non-transitory computer-readable storage medium provided in at least one embodiment of the present disclosure. For example, as Figure 12 shown, the storage medium 600 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 601 can be non-temporarily stored on the storage medium 600. For example, when the computer-readable instructions 601 are executed by a processor, one or more steps in the data loading method described above can be executed. For example, when the computer-readable instructions 601 are executed by a processor, one or more steps in the data storage method described above can be executed.
[0289] For example, the storage medium 600 can be applied to the electronic device 300. For example, the storage medium 600 can include the storage device 308 in the electronic device 300.
[0290] For example, a storage device may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor may run the computer-readable instructions to implement various functions of the processor. Various application programs and various data may also be stored in the storage medium.
[0291] For example, the storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and may also be other applicable storage media.
[0292] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0293] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0294] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0295] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, it should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0296] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0297] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.
[0298] Regarding the present disclosure, the following points also need to be noted: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures can refer to the general design.
[0299] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0300] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A data loading method for loading a tensor to be processed into a buffer, wherein: The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, and a5, where a1, a2, a3, a4, and a5 respectively indicate the size of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include batch dimension, depth dimension, height dimension, width dimension, and channel number dimension. The data loading method comprises: Obtain the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, wherein the data storage format of the original tensor is the same as the data storage format of the tensor to be processed, and the shape and size of the original tensor are represented by b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component; Determine a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; Sending the multiple requests in sequence, and writing the sub-data returned by each request into the buffer area in sequence, so as to load the tensor to be processed into the buffer area; In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but each dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each sub-data requested to be loaded is located in the memory or is not located in the memory at all, and in response to the sub-data loaded by the request is located in the memory, the sub-data requested to be loaded comes from the original tensor and is stored continuously in the memory, The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
2. The data loading method according to claim 1, wherein: In response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be loaded continuously in the first dimension.
3. The data loading method according to claim 1, wherein: In response to the data storage format of the tensor to be processed being NDHWC, the coordinates of each sub-data requested to be loaded satisfy the following conditions: the coordinates of the tensor to be processed in the first dimension and each dimension lower than the first dimension are different, but the coordinates of the tensor in the second dimension and each dimension higher than the second dimension are the same; Among them, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension.
4. The data loading method according to claim 1, wherein: Determining multiple requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed includes: Determine, based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, a first request sent among the multiple requests and an initial state in which the first request enters a state machine; Based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed, the state machine is used to determine each request among the multiple requests except the first request.
5. The data loading method according to claim 4, wherein: Each request includes a data read address for indicating a starting position for reading data from the memory, a data write address for indicating a starting position for writing data to the cache, and a length of the data requested to be loaded. Determining a first request sent from the multiple requests and an initial state of the first request entering into a state machine based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, includes: Determining, based on the starting coordinates of the tensor to be processed, that the first request enters the initial state in the state machine; Using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; Determine a data reading address for the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; Determine a starting address in the buffer area for writing the tensor to be processed as a data writing address for the first request; The data length of the first request to load is determined according to the starting coordinates of the tensor to be processed and the shape size of the tensor to be processed.
6. The data loading method according to claim 5, wherein: Determining, based on the starting coordinates of the tensor to be processed, that the first request enters the initial state in the state machine, comprises: In response to a first coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension being less than a second coordinate value of the starting coordinate of the original tensor in the first dimension, determining that the initial state is the first state, In response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining that the initial state is a second state, wherein a difference between the second coordinate value and the third coordinate value is equal to a size of the original tensor in the first dimension, In response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.
7. The data loading method according to claim 4, wherein: Based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed, the state machine is used to determine each request except the first request in the multiple requests, including: Based on the initial state, using the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request; Determine a data reading address for each request according to the shape and size of the original tensor, the request initial coordinates corresponding to each request, and the data storage format of the tensor to be processed; The data write address of each request is determined according to the data length of each request to be loaded.
8. The data loading method according to claim 7, wherein: The state machine comprises a first state, Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request by using the state machine includes: In response to the state of the current request being the first state and satisfying the first condition, determining that the next request of the current request enters the first state, wherein in response to the current request being the first request, the state of the current request is the initial state; In response to the next request, entering the first state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is a first coordinate value, and determining that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions; Determine that the length of the data currently requested to be loaded is the size of the tensor to be processed in the first dimension; The first condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first condition includes that the sum of the fourth coordinate value of the requested initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the second coordinate value and the fourth coordinate value is less than the second coordinate value, and the remaining size is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded, The first coordinate value is a coordinate value of the starting coordinate of the tensor to be processed in the first dimension.
9. The data loading method according to claim 8, wherein: The state machine also includes a second state, Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request using the state machine, further comprising: In response to the state of the current request being the first state and satisfying a second condition, determining that the next request of the current request enters the second state, In response to the next request, entering the second state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the second coordinate value, and determining that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remains unchanged; Determine the data length of the current request to load based on the first coordinate value and the second coordinate value; The second condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, or the second condition includes that the sum of the fourth coordinate value and the remaining size is greater than or equal to the second coordinate value.
10. The data loading method according to claim 7, wherein: The state machine includes a first state and a second state, Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request by using the state machine includes: In response to the state of the current request being the second state and satisfying a third condition, determining that a next request of the current request enters the first state, wherein in response to the current request being the first request, the state of the current request is the initial state; In response to the next request, entering the first state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is a first coordinate value, and determining that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions; Determine the data length of the current request to load based on the first coordinate value and the size of the tensor to be processed in the first dimension; The third condition includes that the first coordinate value is less than the second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value, or the third condition includes that the sum of the fourth coordinate value of the requested initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the third coordinate value and the fourth coordinate value is less than the second coordinate value, and the remaining size is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded, The first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the first dimension, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
11. The data loading method according to claim 10, wherein: Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request using the state machine, further comprising: In response to the state of the current request being the second state and satisfying a fourth condition, determining that the next request enters the second state, In response to the next request, entering the second state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determining that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions; Determine that the length of the data currently requested to be loaded is the size of the tensor to be processed in the first dimension; Among them, the fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value, or, the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value.
12. The data loading method according to claim 11, wherein: The state machine also includes a third state, Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request using the state machine, further comprising: In response to the state of the current request being the second state and satisfying the fifth condition, determining that the next request enters the third state, In response to the next request, entering the third state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the third coordinate value, and determining that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remains unchanged; Determine the data length loaded by the current request based on the coordinate value of the initial coordinate of the request corresponding to the current request in the first dimension and the third coordinate value; Among them, the fifth condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the third coordinate value, or the fifth condition includes that the sum of the fourth coordinate value and the remaining size is greater than the third coordinate value.
13. The data loading method according to claim 7, wherein: The state machine includes a first state, a second state and a third state, Based on the initial state, determining the request initial coordinates corresponding to each request and the data length loaded by each request by using the state machine includes: In response to the state of the current request being the third state, determining the state to be entered by the next request of the current request according to the first coordinate value, wherein in response to the current request being the first request, the state of the current request is the initial state; Determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determine that the coordinate value of the request initial coordinate corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions; Determining the length of the data loaded by the current request is based on the fourth coordinate value and the third coordinate value of the initial coordinate of the request corresponding to the current request in the first dimension, The first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension, and the third coordinate value is the sum of the second coordinate value of the starting coordinate of the original tensor in the first dimension and the size of the original tensor in the first dimension.
14. The data loading method according to claim 1, wherein: Sending the multiple requests in sequence, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area, including: For any request, all sub-data loaded in response to any request are located in the memory, and any request is sent to the memory; In response to the fact that none of the sub-data loaded by any of the requests are located in the memory, the any of the requests is converted into writing a plurality of predetermined values into the cache area, wherein the number of the plurality of predetermined values is determined by the length of the loaded data specified by the any of the requests.
15. The method according to any one of claims 1 to 14, wherein: The data storage format of the tensor to be processed is NDHWC, or N(C / x)DHW(xC), where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, and x is a positive integer greater than 1. In response to the data storage format of the tensor to be processed being NDHWC, the coordinate value of the channel number dimension is incremented by 1 when updated; In response to the data storage format of the tensor to be processed being N(C / x)DHW(xC), the coordinate value of the channel number dimension is incremented by x when being updated.
16. A data loading method, comprising: Receive a data loading instruction instructing execution of loading a to-be-processed tensor into a cache area, wherein the data loading instruction includes the shape and size of the to-be-processed tensor as input parameters, the data storage format of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor in a coordinate system determined by the original tensor, the shape and size of the to-be-processed tensor is represented by a1, a2, a3, a4, and a5, a1, a2, a3, a4, and a5 respectively indicate the sizes of the to-be-processed tensor in five dimensions and are all positive integers, the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension, the data storage format of the original tensor is the same as the data storage format of the to-be-processed tensor, the shape and size of the original tensor is represented by b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component; After parsing the data loading instruction, using the execution unit to execute the data loading instruction, Wherein, using the execution unit to execute the data loading instruction includes: Determine a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; Sending the multiple requests in sequence, and writing the sub-data returned by each request into the buffer area in sequence, so as to load the tensor to be processed into the buffer area; In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but each dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each sub-data requested to be loaded is located in the memory or is not located in the memory at all, and in response to the sub-data loaded by the request is located in the memory, the sub-data requested to be loaded comes from the original tensor and is stored continuously in the memory, The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
17. A data storage method for writing a tensor to be processed in a buffer into a memory, wherein: The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, and a5, where a1, a2, a3, a4, and a5 respectively indicate the size of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include batch dimension, depth dimension, height dimension, width dimension, and channel number dimension. The data storage method comprises: Obtain the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, wherein the data storage format of the original tensor is the same as the data storage format of the tensor to be processed, and the shape and size of the original tensor are represented by b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component; Determine a plurality of requests for storing the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; Send at least part of the multiple requests in sequence, and write the to-be-written data indicated by each request into the memory in sequence, so as to store the to-be-processed tensor in the memory; In which, in response to a size relationship between the tensor to be processed and the original tensor in the first dimension, when storing the tensor to be processed, it cannot be stored continuously in the first dimension but can be stored continuously in dimensions lower than the first dimension, the request is divided in the second dimension, each sub-data requested for writing belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the sub-data requested for writing belonging to the data range of the original tensor, the sub-data requested for writing belongs to the original tensor and is stored continuously in the memory, The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
18. The data storage method according to claim 17, wherein: Determining multiple requests for storing the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed includes: Determine, based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, a first request sent among the multiple requests and an initial state in which the first request enters a state machine; Based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed, the state machine is used to determine each request among the multiple requests except the first request.
19. The data storage method according to claim 18, wherein: The request initial coordinates corresponding to each request are used to determine the data storage address of the sub-data for storage in the memory. Sending at least some of the multiple requests in sequence, and sequentially writing the to-be-written data indicated by each request into the memory to store the to-be-processed tensor, including: For any request, in response to a first request coordinate value of a request initial coordinate corresponding to any request in the first dimension being less than a second coordinate value of a starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to a third coordinate value, determining not to send any request; In response to the first request coordinate value being greater than or equal to the second coordinate value and the first request coordinate value being less than the third coordinate value, determining to send the any one request; The difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
20. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is used to receive and parse the data loading instruction, wherein the data loading instruction includes the shape size of the tensor to be processed as an input parameter, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, the shape size of the tensor to be processed is represented by a1, a2, a3, a4, and a5, a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers, the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension, the data storage format of the original tensor is the same as the data storage format of the tensor to be processed, the shape size of the original tensor is represented by b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the 5 dimensions and are all positive integers, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component; The execution unit executes the data loading instruction after the instruction parsing unit parses the data loading instruction. Wherein, when the execution unit executes the data loading instruction, it includes performing the following operations: Determine a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed; Send the multiple requests in sequence, and write the sub-data returned by each request into the buffer area in sequence, so as to load the tensor to be processed into the buffer area; In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but each dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each sub-data requested to be loaded is located in the memory or is not located in the memory at all, and in response to the sub-data loaded by the request is located in the memory, the sub-data requested to be loaded comes from the original tensor and is stored continuously in the memory, The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.
21. An electronic device, comprising: A memory non-transitorily stores computer executable instructions; a processor configured to execute the computer executable instructions, Wherein, when the computer executable instructions are executed by the processor, the data loading method according to any one of claims 1-16 or the data storage method according to any one of claims 17-19 is implemented.
22. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by the processor, the data loading method according to any one of claims 1 to 16 or the data storage method according to any one of claims 17 to 19 is implemented.
Citation Information
Patent Citations
Neural network hardware acceleration structure capable of parameterized generation and dynamic configuration
CN116933850A
Data processing method and device, processor, electronic equipment and storage medium
CN119089948A
Method and system for determining wind power in different weathers, storage medium and equipment
CN119149655A
Image transformation for machine learning
US20190236755A1
Method and apparatus for vectorized resource scheduling in distributed computing systems using tensors
US20210089363A1
Cited By
Data processing method, electronic equipment and storage medium
CN120653883A
Data processing method, electronic device, and storage medium
CN120653883B