Data Loading Method, Data Storage Method, Processor, Electronic Device and Medium

By dividing data requests into request groups in parallel processors and filling them, the calculation errors and memory access problems caused by data misalignment are solved, and data transmission efficiency and processor performance are improved.

CN120196566BActive Publication Date: 2025-07-22SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510660556.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

In parallel processors, when data is loaded and stored, the data is not aligned at a fixed multiple, resulting in calculation errors and memory access errors, which increases instruction complexity and resource consumption, and the software needs additional calculations to fill data to meet hardware alignment requirements.

Method used

By dividing multiple requests of the pending tensor into X load or storage request groups, and filling data that is not aligned by multiples, it meets hardware requirements and writes directly to the cache area, reducing software intervention and instruction complexity.

Benefits of technology

It improves the processor's data transmission efficiency, avoids calculation errors, simplifies the instruction process, and improves overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196566B_ABST
    Figure CN120196566B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a medium. The data loading method includes: determining a plurality of first requests for loading a tensor to be processed based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor; sequentially sending the first requests in the k-th loading request group among X loading request groups, sequentially obtaining data from a storage space, and obtaining first data returned by the k-th loading request group; in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, wherein the size of the padded second data is aligned in multiples of Y; and writing the second data into a buffer. The data loading method reduces the complexity of instructions and improves the overall performance of the processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are elastic and not exact. Therefore, an elastic data form is needed to describe various types of data, and thus the concept of a tensor is generated. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions (for example, four dimensions, five dimensions, etc.).

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processor devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of a parallel processor require a large amount of data, and this data is generally stored in the storage component of the parallel processor. For example, the storage component can be a high-speed memory. Through a data loading instruction, this data can be loaded from the storage component to the buffer inside the processor for calculation, and through a data storage instruction, the data in the buffer can be stored in the memory. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a data loading method for loading a tensor to be processed from a storage space where an original tensor is located into a buffer. Wherein, the tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading method includes: obtaining the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determining a plurality of first requests for loading the tensor to be processed. Wherein, the plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed, and the size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. X and Y are positive integers; sequentially sending the first requests in the k-th loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining the first data returned by the k-th loading request group, where k = 1, 2,..., X; in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y; writing the second data into the buffer.

[0005] For example, in the data loading method provided by at least one embodiment of the present disclosure, each of the plurality of first requests includes a data reading address for indicating the starting position of reading data from the storage space, a data writing address for indicating the starting position of writing data into the buffer, and the size of the data loaded by the first request. The request initial coordinates corresponding to the next first request are determined according to the request initial coordinates corresponding to the previous first request. The request initial coordinates corresponding to each first request are used to determine the data reading address, the data writing address, and the size of the data loaded by the first request.

[0006] For example, in the data loading method provided by at least one embodiment of the present disclosure, the loading method of the tensor to be processed includes a first loading method. In the first loading method, multiple dimensions of the tensor to be processed include a first dimension and a second dimension, and the data determined by the first dimension and the second dimension represents a pixel. In response to the loading method of the tensor to be processed being the first loading method, based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, multiple first requests for loading the tensor to be processed are determined, including: determining the number of pixels included in the tensor to be processed; based on the number of pixels included in the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, multiple first requests for loading the tensor to be processed are determined. Among them, the multiple first requests are divided into 1 target loading request group, and the number of pixels included in the first data loaded by the target loading request group is required to be aligned in multiples of Y.

[0007] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the loading method of the tensor to be processed being the first loading method, in response to the size of the first data returned by the kth loading request group not being aligned in multiples of Y, padding the first data to obtain second data, including: in response to the number of pixels included in the first data returned by the target loading request group not being aligned in multiples of Y, padding the first data to obtain the second data, where the number of pixels included in the padded second data is aligned in multiples of Y.

[0008] For example, in the data loading method provided by at least one embodiment of the present disclosure, multiple dimensions of the tensor to be processed include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The first dimension includes the width dimension, the second dimension includes the height dimension, and the channel number dimension does not calculate the number of pixels.

[0009] For example, in the data loading method provided by at least one embodiment of the present disclosure, the loading manner of the tensor to be processed includes a second loading manner. In the second loading manner, multiple dimensions of the tensor to be processed include a first dimension, and the data size of the tensor to be processed in the first dimension is required to be aligned in multiples of Y. In response to the loading manner of the tensor to be processed being the second loading manner, and the size of the first data returned in response to the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data includes: in response to the size of the first data returned in response to the k-th loading request group not being aligned in multiples of Y in the first dimension, padding the first data in the first dimension to obtain second data, where the size of the obtained second data in the first dimension is aligned in multiples of Y.

[0010] For example, in the data loading method provided by at least one embodiment of the present disclosure, the multiple dimensions of the tensor to be processed further include a second dimension and a third dimension. The data storage format of the tensor to be processed indicates that the first dimension has priority over the second dimension during storage or loading, and indicates that the third dimension has priority over the first dimension during storage or loading. The first dimension is adjacent to the second dimension, and the first dimension is adjacent to the third dimension. The data determined by the first dimension and the second dimension represents a pixel. In response to the tensor to be processed not being continuously loadable in the third dimension, each first request in the k-th loading request group loads sub-data corresponding to one pixel, or, in response to the tensor to be processed being continuously loadable in the third dimension, the first request in the k-th loading request group loads the first data.

[0011] For example, in the data loading method provided by at least one embodiment of the present disclosure, the multiple dimensions of the tensor to be processed include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The first dimension includes the width dimension, the second dimension includes the height dimension, and the third dimension includes the channel number dimension.

[0012] For example, in the data loading method provided by at least one embodiment of the present disclosure, the first dimension includes the batch dimension, the depth dimension, the height dimension, the width dimension, or the channel number dimension of the tensor to be processed.

[0013] For example, in the data loading method provided by at least one embodiment of the present disclosure, when the size of the first data returned in response to the k-th loading request group is not aligned with a multiple of Y, padding the first data to obtain second data includes: when the size of the first data returned in response to the k-th loading request group is not aligned with a multiple of Y, determining the size of the data to be padded for the first data based on the part of the size of the first data that is less than a multiple of Y; determining at least one second request for padding the first data based on the size of the data to be padded for the first data; sending the at least one second request to pad the first data to obtain the second data, where the size of the padding data requested by the at least one second request is equal to the size of the data to be padded for the first data.

[0014] For example, in the data loading method provided by at least one embodiment of the present disclosure, each second request in the at least one second request includes a data write address for indicating the starting position of filling data into the buffer area and the size of the padding data requested by the second request. The request initial coordinate of the first second request in the at least one second request is determined according to the request initial coordinate of the last first request in the k-th loading request group, and the request initial coordinate of the next second request is determined according to the request initial coordinate corresponding to the previous second request. When k is less than X, the request initial coordinate of the first first request in the (k + 1)-th loading request group is determined according to the request initial coordinate corresponding to the last second request in the at least one second request. The request initial coordinate corresponding to each first request in the multiple first requests is used to determine the data read address, data write address, and the size of the data loaded by the first request, and the request initial coordinate corresponding to each second request is used to determine the data write address of the second request and the size of the padding data requested by the second request.

[0015] For example, in the data loading method provided by at least one embodiment of the present disclosure, the sub-data requested by each second request in the at least one second request is 0 or other specified values.

[0016] For example, the data loading method provided by at least one embodiment of the present disclosure further includes: when the second data corresponding to the X-th loading request group in the X loading request groups is written into the buffer area, determining the data write address of the starting position of the next tensor to be processed in the buffer area based on the data write address of the starting position of the tensor to be processed in the buffer area and the size of the data after the tensor to be processed is padded.

[0017] At least one embodiment of the present disclosure provides another data loading method, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from a storage space where an original tensor is located into a buffer, where the tensor to be processed and the original tensor are multi-dimensional tensors, and a data storage format of the original tensor is the same as that of the tensor to be processed, and the data storage format is used to indicate a storage order and a dimension arrangement of the tensor in the storage space, and the data loading instruction includes a loading method of the tensor to be processed, a shape size of the tensor to be processed, the data storage format of the tensor to be processed, and a starting coordinate of the tensor to be processed in a coordinate system determined by the original tensor as input parameters; after parsing the data loading instruction, executing the data loading instruction using a first execution unit, where executing the data loading instruction using the first execution unit includes: determining a plurality of first requests for loading the tensor to be processed based on the loading method of the tensor to be processed, the shape size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor, where the plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed, and a size of data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y, and X and Y are positive integers; sequentially sending first requests in the k-th loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining first data returned by the k-th loading request group, where k = 1, 2,..., X; in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where a size of the padded second data is aligned in multiples of Y; and writing the second data into the buffer.

[0018] At least one embodiment of the present disclosure provides a data storage method for writing a tensor to be stored in a buffer into the storage space where the original tensor is located. The tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage method includes: obtaining the storage mode of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer; determining a plurality of third requests for storing the tensor to be stored based on the storage mode of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer. The plurality of third requests are divided into X storage request groups based on the storage mode of the tensor to be stored, and the size of the data stored in each storage request group of the X storage request groups is required to be aligned in multiples of Y, where X and Y are positive integers; sequentially sending the third requests in the k-th storage request group among the X storage request groups, sequentially obtaining data from the buffer, and obtaining the third data returned by the k-th storage request group, where k = 1, 2,..., X; in response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data, where the size of the padded fourth data is aligned in multiples of Y; in response to k being equal to 1, writing the third data returned by the first storage request group into the storage space; in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, writing the third data returned by the (k + 1)-th storage request group into the storage space.

[0019] For example, in the data storage method provided by at least one embodiment of the present disclosure, the step of in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, writing the third data returned by the (k + 1)-th storage request group into the storage space includes: in response to k being less than X, determining the data read address of the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data write address of the starting position of the third data in the storage space based on the fourth data corresponding to the k-th storage request group; based on the data read address of the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data write address of the starting position of the third data in the storage space, writing the third data returned by the (k + 1)-th storage request group into the storage space.

[0020] For example, in the data storage method provided by at least one embodiment of the present disclosure, each of the plurality of third requests includes a data read address for indicating a starting position for reading data from the buffer area, a data write address for indicating a starting position for writing data to the storage space, and a size of the data stored in the third request. The request initial coordinates corresponding to the next third request are determined according to the request initial coordinates corresponding to the previous third request. The request initial coordinates corresponding to each third request are used to determine the data read address, the data write address, and the size of the data stored in the third request.

[0021] For example, in the data storage method provided by at least one embodiment of the present disclosure, the storage method of the tensor to be stored includes a first storage method. In the first storage method, multiple dimensions of the tensor to be stored include a first dimension and a second dimension, and the data determined by the first dimension and the second dimension is represented as a pixel. In response to the storage method of the tensor to be stored being the first storage method, determining a plurality of third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer area includes: determining the number of pixels included in the tensor to be stored; determining a plurality of third requests for storing the tensor to be stored based on the number of pixels included in the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer area. Among them, the plurality of third requests are divided into 1 target storage request group, and the number of pixels included in the third data stored in the target storage request group is required to be aligned in multiples of Y.

[0022] For example, in the data storage method provided by at least one embodiment of the present disclosure, the storage method of the tensor to be stored includes a second storage method. In the second storage method, multiple dimensions of the tensor to be stored include a first dimension, and the data size of the tensor to be stored in the first dimension is required to be aligned in multiples of Y. In response to the storage method of the tensor to be stored being the second storage method, in response to the size of the third data returned by the kth storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data includes: in response to the size of the third data returned by the kth storage request group not being aligned in multiples of Y in the first dimension, padding the third data in the first dimension to obtain fourth data, where the size of the obtained fourth data in the first dimension is aligned in multiples of Y.

[0023] At least one embodiment of the present disclosure provides another data storage method, including: receiving a data storage instruction indicating to execute writing a tensor to be stored in a buffer into a storage space where an original tensor is located, where the tensor to be stored and the original tensor are multi-dimensional tensors, the data storage format of the original tensor is the same as that of the tensor to be stored, the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space, and the data storage instruction includes the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer as input parameters; after parsing the data storage instruction, executing the data storage instruction using a second execution unit, where executing the data storage instruction using the second execution unit includes: determining a plurality of third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, where the plurality of third requests are divided into X storage request groups based on the storage method of the tensor to be stored, and the size of the data stored in each of the X storage request groups is required to be aligned in multiples of Y, and X and Y are positive integers; sequentially sending the third requests in the k-th storage request group among the X storage request groups, sequentially obtaining data from the buffer, and obtaining third data returned by the k-th storage request group, where k = 1, 2,..., X; in response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data, where the size of the padded fourth data is aligned in multiples of Y; in response to k being equal to 1, writing the third data returned by the first storage request group into the storage space; in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, writing the third data returned by the (k + 1)-th storage request group into the storage space.

[0024] At least one embodiment of the present disclosure provides a processor, including a first instruction parsing unit and a first execution unit. Wherein, the first instruction parsing unit is configured to receive and parse a data loading instruction, where the data loading instruction is used to execute loading a tensor to be processed from a storage space where an original tensor is located to a buffer area. The tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading instruction includes the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor as input parameters. After the first instruction parsing unit parses the data loading instruction, the first execution unit executes the data loading instruction. When the first execution unit executes the data loading instruction, the following operations are included: determining a plurality of first requests for loading the tensor to be processed based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor. The plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed. The size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. X and Y are positive integers. Sequentially sending the first requests in the k-th loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining the first data returned by the k-th loading request group, where k = 1, 2,..., X. In response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y. Writing the second data into the buffer area.

[0025] At least one embodiment of the present disclosure provides another processor, including a second instruction parsing unit and a second execution unit. The second instruction parsing unit is configured to receive and parse a data storage instruction, where the data storage instruction is used to execute writing a tensor to be stored in a buffer into a storage space where an original tensor is located. The tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage instruction includes the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer as input parameters. After the second instruction parsing unit parses the data storage instruction, the second execution unit executes the data storage instruction. When the second execution unit executes the data storage instruction, the following operations are included: determining a plurality of third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer. The plurality of third requests are divided into X storage request groups based on the storage method of the tensor to be stored. The size of the data stored in each of the X storage request groups is required to be aligned in multiples of Y, where X and Y are positive integers. Sequentially sending the third requests in the k-th storage request group among the X storage request groups, sequentially obtaining data from the buffer, and obtaining third data returned by the k-th storage request group, where k = 1, 2,..., X. In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data, where the size of the padded fourth data is aligned in multiples of Y. In response to k being equal to 1, writing the third data returned by the first storage request group into the storage space. In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, writing the third data returned by the (k + 1)-th storage request group into the storage space.

[0026] At least one embodiment of the present disclosure provides an electronic device, including: one or more processors; a memory including one or more computer program modules; where the one or more computer program modules are stored in the memory and configured to be executed by the one or more processors, and the one or more computer program modules are used to implement the data loading method according to any embodiment of the present disclosure, or the data storage method according to any embodiment of the present disclosure.

[0027] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data loading method according to any embodiment of the present disclosure is implemented, or the data storage method according to any embodiment of the present disclosure is implemented. Description of the Drawings

[0028] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0029] Figure 1 Shows a schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU);

[0030] Figure 2 Shows a schematic structure of a tensor;

[0031] Figure 3A Shows a schematic diagram of the data storage format of NDHWC;

[0032] Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32;

[0033] Figure 4 Shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;

[0034] Figure 5A Shows a schematic diagram of an example of the first loading method of a tensor to be processed provided by at least one embodiment of the present disclosure;

[0035] Figure 5B Shows a schematic diagram of an example of the data loading method provided by at least one embodiment of the present disclosure;

[0036] Figure 6A Shows a schematic diagram of an example of the second loading method of a tensor to be processed provided by at least one embodiment of the present disclosure;

[0037] Figure 6B Shows a schematic diagram of another example of the data loading method provided by at least one embodiment of the present disclosure;

[0038] Figure 6C Shows a schematic diagram of yet another example of the data loading method provided by at least one embodiment of the present disclosure;

[0039] Figure 7 Another schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown;

[0040] Figure 8 A schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown;

[0041] Figure 9 Another schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown;

[0042] Figure 10 A schematic block diagram of a processor provided by at least one embodiment of the present disclosure is shown;

[0043] Figure 11 A schematic block diagram of another processor provided by at least one embodiment of the present disclosure is shown;

[0044] Figure 12 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown;

[0045] Figure 13 A schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure is shown;

[0046] Figure 14 A schematic block diagram of a storage medium provided by at least one embodiment of the present disclosure is shown. Detailed implementation manners

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0048] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a", "an" or "the" do not denote a limitation of quantity, but rather denote the presence of at least one. Terms such as "comprising" or "including" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0049] The following illustrates this disclosure through several specific embodiments. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of this disclosure appears in more than one drawing, the component is denoted by the same or similar reference numeral in each drawing.

[0050] Figure 1 A schematic structural diagram of a GPGPU is shown.

[0051] For example, as Figure 1 shown, a GPGPU is actually an array of programmable multi-processors. For example, the programmable multi-processors can be Streaming Processor Clusters (SPCs), such as including Figure 1 the Streaming Processor Cluster 1 shown, ……, the Streaming Processor Cluster M, where M is a positive integer greater than 1. In a general-purpose graphics processor, 1 Streaming Processor Cluster processes one computing task, or multiple Streaming Processor Clusters process one computing task. Data sharing among multiple Streaming Processor Clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0052] For example, as Figure 1 shown, taking the Streaming Processor Cluster 1 as an example, 1 Streaming Processor Cluster can include multiple Compute Units (CUs), such as Figure 1The computing units 1, 2, …, N in it, where N is a positive integer. Each computing unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit may include multiple cores (also known as computing cores or computing nuclei), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register bank in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit.

[0053] For example, in parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processing unit (or known as a parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in it). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0054] For example, in each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in it) schedules and allocates the thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0055] For example, as Figure 1 shown, general computing operations, such as computing operations on matrices in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as can be stored in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored to the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.

[0056] For example, general computing operations include General Matrix Multiplication (GEMM). As an example, the data for general matrix multiplication is represented as two 5-dimensional arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, there are other computing operations, which will not be listed one by one here. The data involved in these calculations is usually embodied in the form of tensors.

[0057] For example, for a certain tensor A in the cache area of the processor, its shape and size can be represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor data in 5 dimensions. Here, a1, a2, a3, a4, a5 are all positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples captured in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking tensor A as an example, a1 can be the size of the N dimension, a2 can be the size of the D dimension, a3 can be the size of the H dimension, a4 can be the size of the W dimension, and a5 can be the size of the C dimension. Of course, the present disclosure does not make specific limitations on the number of dimensions of the tensor and the type of each dimension.

[0058] As an example, Figure 2 shows a schematic structure of a tensor.

[0059] For example, in Figure 2 the shown tensor, a1 is the size of the N dimension and is equal to 1, a2 is the size of the D dimension and is equal to 1, a3 is the size of the H dimension and is equal to 5, a4 is the size of the W dimension and is equal to 4, and a5 is the size of the C dimension and is equal to 64. For example, Figure 2 the pixel elements of the tensor in

[0060] For example, the placement of a tensor in a storage component (such as memory or cache) can have multiple formats, called data storage formats (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The following uses Figure 2 the shown tensor to describe different data storage formats.

[0061] In the related art, the data storage format can include NDHWC, also known as the Linear mode. Figure 3A shows a schematic diagram of the data storage format of NDHWC.

[0062] For example, for the NDHWC linear mode, as Figure 3A shown, starting from the first element of the first channel (c0 in Figure 3A where a5 = 0), which is the element 0 in Figure 3A , then storing the first element of the second channel (c1 in Figure 3A where a5 = 1), which is the element 20 in Figure 3A , and so on until the first elements of all channels are laid out. For example, after the first element of the 64th channel (c63 in Figure 3A where a5 = 63), which is the element 1260 in Figure 3A , select the second element of the first channel (c0 in Figure 3A where a5 = 0), which is the element 1 in Figure 3A , then store the second element of the second channel (c1 in Figure 3A where a5 = 1), which is the element 21 in Figure 3A , and so on until the second elements of all channels are laid out, and so on. For example, in the NDHWC storage format, the C dimension is the lowest dimension during storage, followed by the H dimension, then the W dimension, then the D dimension, and finally the N dimension.

[0063] In the related art, the data storage format can also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. according to needs. The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but there is a key difference. In the memory layout of N(C / x)DHW(xC), a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0064] Figure 3B Shows a schematic diagram of the N(C / x)DHW(xC) data storage format, where x = 32.

[0065] For example, as Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0 in Figure 3BIt consists of c31) in it. The second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format. For example, in the storage format of N(C / x)DHW(xC), the W dimension is the lowest dimension during storage, followed by the H dimension, then the D dimension, then the C dimension, and finally the N dimension. And for the W dimension, continuous xC is substantially considered.

[0066] For example, for a certain tensor B in the memory, similar to a certain tensor A in the buffer, this tensor B can be stored in the memory in one of the two data storage formats described above. The shape dimensions of the tensor can be similarly expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the dimensions of the tensor B in these 5 dimensions and are all positive integers.

[0067] It should be noted that in the related art and possible future developments, the data storage format of tensors is not limited to the above-described two data storage formats of N(C / x)DHW(xC) and NDHWC. The method described in this disclosure does not limit this.

[0068] In the related art, during the calculation process of the processor, a large amount of calculation data (for example, in the form of tensors) may be generated. Multiple tensors generated during the calculation process can be stored in the memory or buffer. For example, the buffer here can refer to Figure 1 the buffer inside the streaming processor cluster shown in, and further, this data can also be transferred from the buffer or directly stored in the memory. For example, this memory can be Figure 1 the high - bandwidth memory HBM shown in. For example, the storage form of tensors in both the memory and the buffer can be any one of the above - described N(C / x)DHW(xC) data storage format and NDHWC data storage format, or can also be selected as other data storage formats according to the actual situation.

[0069] For example, the storage space of the memory is usually much larger than that of the buffer, but it is farther from the calculation unit, and the data transfer efficiency is lower than that of the buffer. If part of the data is temporarily stored in the buffer of the processor, the distance between this part of the data and the calculation unit of the processor can be shortened, and the data transfer efficiency can be improved. For example, correspondingly, this data can also be transferred from the buffer or directly stored in the memory. Thus, during the calculation process, according to factors such as technical requirements and the respective storage characteristics of the buffer and the memory, a large amount of data transfer needs to be carried out between the two.

[0070] In addition, when performing computational operations such as GEMM, hardware such as processors and memories has strict requirements for data alignment. For example, it is required that the data transferred between the buffer and the memory must be aligned in multiples of Y, where Y can be an integer that meets the hardware alignment requirements (such as 16, etc.). This is because hardware designs are typically based on fixed-size processing units (such as 16×16 matrix units) to achieve efficient parallel computing. However, in practical applications, the data used for loading or storing is not always aligned in multiples of Y; if not properly processed, the hardware cannot directly handle this misaligned data, resulting in computational errors.

[0071] To avoid computational errors caused by data misalignment, on the one hand, software needs to additionally calculate and pad (fill) virtual data (such as filling zeros or other specified values) before the data is transferred to the hardware to meet the hardware alignment requirements. This padding process in software requires additional calculations and memory operations in software. For example, when performing calculations, additional information is needed to indicate which data is virtual and does not need to participate in the actual calculation, which increases the instruction complexity, increases the time and resource consumption for executing instructions, not only makes the software and hardware designs more complex, but also may cause errors due to missing tag information.

[0072] On the other hand, in some cases, if the software fails to correctly align the data to a multiple of Y, the misaligned data will cause discontinuous memory access. For example, in memory (such as HBM, etc.), the addresses in the memory are regular, and each address strictly stores Y data units (such as Y pixel units); if the actual data is not aligned to Y, the hardware will still read in multiples of Y, resulting in invalid data (such as residual values or random values) being included in the calculation, thus producing incorrect results. In addition, for misaligned data, the memory access instruction for the next data to be accessed needs to re-determine the starting point coordinates of this data in the buffer or memory, that is, the software needs to recalculate the starting address of this data, further increasing the instruction complexity.

[0073] At least one embodiment of the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium.

[0074] In at least one embodiment, the data loading method is used to load a tensor to be processed from the storage space where the original tensor is located into a buffer. Here, the tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading method includes: obtaining the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determining a plurality of first requests for loading the tensor to be processed. Among them, the plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed. The size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. X and Y are positive integers; sequentially sending the first requests in the k-th loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining the first data returned by the k-th loading request group, where k = 1, 2,..., X; in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y; writing the second data into the buffer.

[0075] The method, processor, electronic device, and storage medium provided by at least one embodiment of the present disclosure can divide a plurality of first requests for loading a tensor to be processed into X loading request groups, and can pad the data requested by each loading request group so that the padded data is aligned in multiples of Y. Thus, the hardware can directly process the padded data without additional intervention from software, and the starting address of the next tensor to be processed can be automatically obtained after the padding operation is completed, without the software sending a separate instruction to calculate the next starting address. This significantly reduces the burden on the software, reduces the complexity of the instructions, ensures the efficient transmission and processing of data at the hardware level, and at the same time avoids risks such as calculation errors and memory access errors caused by unaligned data, improving the overall performance of the processor.

[0076] Next, at least one embodiment of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals in different drawings will be used to refer to the same elements that have been described.

[0077] The data loading method provided by at least one embodiment of the present disclosure is used to load a tensor to be processed from the storage space where the original tensor is located into a buffer. For example, the storage space where the original tensor is located can be memory or other memories (such as HBM, etc.), and the buffer can be a buffer in a processor such as a GPGPU. Specifically, it can be selected according to actual needs, and the embodiments of the present disclosure do not limit this.

[0078] For example, the tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. As an example, the data storage format of the tensor to be processed can be any one of NDHWC or N(C / x)DHW(xC). For the characteristics of the above two data storage formats, reference can be made to the description above Figure 3A - Figure 3B and will not be repeated here; the data storage format of the tensor to be processed can also be selected as other types of data storage formats according to actual needs, and the embodiments of the present disclosure do not limit this.

[0079] Figure 4 The schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown.

[0080] For example, as Figure 4 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S110 to S150.

[0081] Step S110: Obtain the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.

[0082] For example, in step S110, the loading method of the tensor to be processed may include multiple types. For example, in a continuous first loading method, through a data loading instruction, starting from a starting point, pixels can be continuously loaded (e.g., in a continuous manner such as pixel by pixel, skipping pixels, pixel row by pixel, etc.) from the storage space where the original tensor is located until the number of acquired pixels reaches the target value; this first loading method can be applicable to convolution operations in the computing units within the processor, and the embodiments of the present disclosure do not limit this. For example, in another discontinuous second loading method, through a data loading instruction, all or part of the data in the original tensor can be loaded into the buffer area, that is, the tensor to be processed is loaded as a whole from the storage space where the original tensor is located into the buffer area; this second loading method can be applicable to GEMM calculations such as matrix multiplication of two tensors as a whole in the computing units within the processor, and the embodiments of the present disclosure do not limit this. For example, the loading method of the tensor to be processed can also be selected as other methods according to actual needs, and the embodiments of the present disclosure do not limit this.

[0083] For example, in step S110, the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor can be the coordinates of the starting point for loading the tensor to be processed from the storage space where the original tensor is located, and can be specifically selected according to actual needs, and the embodiments of the present disclosure do not limit this.

[0084] For example, in step S110, the shape dimensions of the tensor to be processed may include the dimensions on each dimension of the tensor to be processed. As an example, the five dimensions of the tensor to be processed may include [N, D, H, W, C], and a1, a2, a3, a4, and a5 respectively indicate the dimensions of the tensor data on the five dimensions, and can be specifically referred to, for example, the description above Figure 2 and will not be repeated here.

[0085] Step S120: Based on the loading method of the tensor to be processed, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determine multiple first requests for loading the tensor to be processed.

[0086] For example, in step S120, based on the loading method of the tensor to be processed, the shape dimensions of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, relevant information of the tensor to be processed to be loaded in the storage space can be determined, so that multiple first requests can be determined to load the tensor to be processed.

[0087] For example, in some examples, each first request may include a data read address for indicating a starting position to read data from a storage space, a data write address for indicating a starting position to write data to a buffer, and a size of data loaded by the first request. For example, according to the request initial coordinates corresponding to the previous first request, the request initial coordinates corresponding to the next first request can be determined, and the request initial coordinates corresponding to each first request can be used to determine the data read address, data write address, and size of data loaded by the first request.

[0088] For example, in step S120, multiple first requests are divided into X loading request groups based on the loading manner of the tensor to be processed, and the size of data loaded by each loading request group is required to be aligned in multiples of Y, where X and Y are positive integers. For example, according to different loading manners of the tensor to be processed, the size of sub-data loaded by each first request is different, and the size of data required to be aligned in multiples of Y in the tensor to be processed is different. For example, by dividing multiple first requests into X loading request groups, it is convenient to perform padding processing on the data loaded by each loading request group in subsequent steps to meet the hardware alignment requirements.

[0089] Step S130: Sequentially send the first requests in the k-th loading request group among the X loading request groups, sequentially obtain data from the storage space, and obtain the first data returned by the k-th loading request group, where k = 1, 2,..., X.

[0090] For example, in step S130, the k-th loading request group includes one or more first requests, the first data returned by the k-th loading request group includes the data returned by each first request, and the size of the first data is required to be aligned in multiples of Y. For example, the k-th loading request group can be any one of the X loading request groups, that is, the first data returned by each loading request group is required to be aligned in multiples of Y.

[0091] Step S140: In response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, perform padding on the first data to obtain the second data.

[0092] For example, in step S140, if the size of the first data returned by the k-th loading request group is not aligned in multiples of Y (that is, the hardware alignment requirements are not met), the first data can be padded, and the size of the second data obtained by padding is aligned in multiples of Y. For example, for the second data obtained by padding, the padded part in the second data can be 0 or other specified values, that is, each padding value can be 0 or other specified values, and the embodiments of the present disclosure do not limit this.

[0093] Step S150: Write the second data into the buffer.

[0094] For example, in step S150, since the second data meets the alignment requirements of the hardware, the second data can be directly written into the buffer without additional calculations and memory operations by the software, improving the data transfer efficiency and reducing the complexity of the instructions.

[0095] It should be noted that when loading the tensor to be processed from the storage space where the original tensor is located, data located outside the edge of the original tensor due to padding operations, etc. in the storage space may be loaded as data in the tensor to be processed. At this time, the data loading operation can still be performed using the coordinate system determined by the original tensor; that is, the tensor to be processed may include all or part of the data in the original tensor, or the tensor to be processed may not include the data in the original tensor, which can be specifically selected according to actual needs, and the embodiments of the present disclosure do not limit this.

[0096] In some examples, the loading method of the tensor to be processed includes a first loading method. For example, in the first loading method, multiple dimensions of the tensor to be processed include a first dimension and a second dimension, and the data determined by the first dimension and the second dimension represents a pixel. For example, in response to the loading method of the tensor to be processed being the first loading method, Figure 4 step S120 of can further include the following steps S121 to S122.

[0097] Step S121: Determine the number of pixels included in the tensor to be processed;

[0098] Step S122: Based on the number of pixels included in the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determine multiple first requests for loading the tensor to be processed.

[0099] For example, the first loading method can be a continuous loading method, and the range of the tensor to be processed is determined by the number of pixels to be obtained. For example, in step S121, the number of pixels of the tensor to be processed that need to be loaded can be determined; in step S122, starting from the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, pixels are loaded from the storage space where the original tensor is located until the number of obtained pixels reaches the number of pixels included in the tensor to be processed. For example, in the first loading method, multiple first requests are divided into 1 target loading request group, and the number of pixels included in the first data loaded by the target loading request group is required to be aligned in multiples of Y.

[0100] For example, in response to the loading method of the tensor to be processed being the first loading method, Figure 4Step S140 may further include step S1401: in response to the number of pixels included in the first data returned by the target loading request group not being aligned in multiples of Y, padding the first data to obtain second data. For example, the number of pixels included in the padded second data is aligned in multiples of Y.

[0101] For example, in some examples, the multiple dimensions of the tensor to be processed may include a batch dimension N, a depth dimension D, a height dimension H, a width dimension W, and a number of channels dimension C. The first dimension may include the width dimension, the second dimension may include the height dimension, and the number of channels dimension does not count the number of pixels.

[0102] Figure 5A A schematic diagram showing an example of a first loading method of a tensor to be processed provided by at least one embodiment of the present disclosure.

[0103] For example, as Figure 5A shown, the data storage format of the original tensor is the same as that of the tensor to be processed. For example, in the first loading method, the multiple dimensions of the original tensor or the tensor to be processed include a first dimension W and a second dimension H. The values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel in the original tensor, W = 1 and H = 0 correspond to the second pixel in the original tensor. In the tensor, the C dimension does not affect the number of pixels.

[0104] For example, in the first loading method, all or part of the data in the original tensor can be continuously loaded into the buffer area as the tensor to be processed through a data loading instruction. Therefore, during the loading process, it is necessary to indicate the number of pixels included in the tensor to be processed, and it is required that the number of pixels be aligned in multiples of Y.

[0105] For example, as Figure 5A shown, the starting point of the original tensor is represented as (C = 0, W = 0, H = 0), for example, to indicate the position of the original tensor in the storage space. For example, the starting point of the tensor to be processed represented by the starting coordinates in the coordinate system determined by the original tensor is (C = 0, W = 3, H = 0), that is, starting from the 3rd pixel in the first row of the original tensor to obtain the tensor to be processed. For example, in step S121, determine the number of pixels copy_pixel_num included in the tensor to be processed; in step S122, as shown by the dotted line, starting from the starting point (C = 0, W = 3, H = 0), load pixels from the storage space where the original tensor is located until the number of obtained pixels reaches the number of pixels copy_pixel_num included in the tensor to be processed.

[0106] It should be noted that in Figure 5AIn the example, only the three dimensions W, H, and C are shown. It can be understood that if the data in these three dimensions still does not reach copy_pixel_num, data can be further obtained from higher dimensions (such as the batch dimension N, the depth dimension D, etc.). The embodiments of the present disclosure do not limit this. In addition, as Figure 5A shown, the C dimension itself does not affect the number of pixels. For example, in Figure 5A the example, the data storage format of the original tensor in the storage space is the NDWHC format, and the number of pixels of the tensor to be processed is copy_pixel_num = 28.

[0107] Figure 5B FIG. shows a schematic diagram of an example of the data loading method provided by at least one embodiment of the present disclosure. For example, Figure 5B is Figure 4 an example of the data loading method shown.

[0108] For example, Figure 5B the example of Figure 5A obtains the tensor to be processed in a similar first loading manner as shown. For example, in Figure 5B the example, both the tensor to be processed and the original tensor are in the NDHWC data storage format. The original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (C_dim = 8, the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Thus, the total number of pixels included in the original tensor can be obtained as 64.

[0109] For example, as Figure 5B shown, the tensor to be processed corresponds to the data in the squares where the solid arrows are located in the figure. For example, in step S110, the loading method of the tensor to be processed (i.e., the first loading method), the data storage format NDHWC of the tensor to be processed, the shape and size of the tensor to be processed (i.e., the sizes of each dimension), and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor (i.e., the pixel coordinates where the starting position of the solid arrow is located) are obtained. For example, the starting point coordinates of the tensor to be processed in the coordinate system determined by the original tensor to be obtained are represented as W_coord_b = 2, H_coord_b = 2, D_coord_b = 0, N_coord_b = 0, C_coord_b = 0. For example, in step S121, it is determined that the number of pixels included in the tensor to be processed is copy_pixel_num = 40.

[0110] For example, in step S122, based on the above information, multiple first requests for loading the tensor to be processed can be determined. For example, the multiple first requests are divided into 1 target loading request group based on the first loading method. For example, in step S130, the first requests in the target loading request group are sent to sequentially obtain data from the original tensors in the storage space, and the first data returned by the target loading request group (i.e., the pixel positions where the solid arrows are located in the figure) is obtained. Specifically, for example, taking the pixel coordinates of the starting position of the solid arrow as the starting point coordinates, according to the data arrangement method, in the order from low dimension to high dimension (i.e., in the order of C, W, H, D, N), the first data is obtained pixel by pixel in the arrow order until the number of pixels obtained is equal to the number of pixels copy_pixel_num included in the tensor to be processed. For example, in Figure 5B the example, the number of pixels copy_pixel_num included in the first data loaded by the target loading request group is required to be aligned in multiples of 16 (i.e., Y = 16).

[0111] For example, as Figure 5B shown, the number of pixels included in the first data returned by the target loading request group is the number of pixels copy_pixel_num = 40 included in the tensor to be processed. Since 40 is not a multiple of 16, that is, 16 - copy_pixel_num[3:0] is not equal to 0 (where copy_pixel_num[3:0] is the binary [3:0] bits of the number of pixels), the number of pixels included in the first data returned by the target loading request group is not aligned in multiples of 16. Therefore, in step S1401, the first data needs to be padded to obtain the second data, and the number of pixels to be padded is 16 - copy_pixel_num [3:0] = 8.

[0112] Specifically, for example, in Figure 5B the example, in step S1401, the first data is padded in the order of the dashed arrows, and the data in the padded part is Figure 5B represented as the pixels where the dashed arrows are located in the figure (the number of pixels in the padded part is 8). For example, the second data obtained by padding includes the first data and the data in the above - mentioned padded part, and the number of pixels included in the second data is 48, which can be aligned in multiples of 16. Each pixel in the above - mentioned padded part can be padded with 0 or other specified values, and the embodiments of the present disclosure do not limit this.

[0113] Further, for example, in step S150, since the second data meets the alignment requirements of the hardware (i.e., is aligned in multiples of 16), the second data can be directly written into the buffer for computing operations such as convolution in the processor.

[0114] It should be noted that Figure 5A and Figure 5B The first loading method, original tensor, tensor to be processed, data acquisition method, filling method, etc. shown are only exemplary. The first loading method, original tensor, tensor to be processed, data acquisition method, filling method, etc. can also be selected as other forms according to actual needs, and the embodiments of the present disclosure do not limit this; the tensor to be processed and the original tensor can also be in other data storage formats such as N(C / x)DHW(xC), and their implementation principles are basically similar to the example implementation methods provided by at least one embodiment of the present disclosure, and will not be elaborated here; in addition, in addition to Figure 5B loading pixels in a pixel-by-pixel manner in the example, pixels can also be loaded in a pixel-skipping manner by setting a stride (Stride), and its implementation principle is basically similar to the example implementation methods provided by at least one embodiment of the present disclosure, and will not be elaborated here.

[0115] In some examples, the loading method of the tensor to be processed may further include a second loading method. For example, in the second loading method, multiple dimensions of the tensor to be processed include a first dimension, and the data size of the tensor to be processed in the first dimension is required to be aligned in multiples of Y. For example, in response to the loading method of the tensor to be processed being the second loading method, Figure 4 step S140 of

[0116] may further include step S1402: In response to the size of the first data returned by the kth loading request group not being aligned in multiples of Y in the first dimension, padding the first data in the first dimension to obtain second data. For example, the size of the second data obtained by padding is aligned in multiples of Y in the first dimension.

[0117] For example, multiple dimensions of the tensor to be processed further include a second dimension and a third dimension. The data storage format of the tensor to be processed indicates that the first dimension has priority over the second dimension during storage or loading, and indicates that the third dimension has priority over the first dimension during storage or loading. The first dimension is adjacent to the second dimension, and the first dimension is adjacent to the third dimension. For example, in response to the tensor to be processed not being continuously loadable in the third dimension, the sub-data loaded by each first request in the kth loading request group is the data corresponding to one pixel; or, in response to the tensor to be processed being continuously loadable in the third dimension, the data loaded by the first request in the kth loading request group is the first data.

[0117] For example, in some examples, multiple dimensions of the tensor to be processed include a batch dimension N, a depth dimension D, a height dimension H, a width dimension W, and a channel number dimension C. The first dimension includes the width dimension W, the second dimension includes the height dimension H, and the third dimension includes the channel number dimension C.

[0118] For example, in other examples, in addition to the width dimension W, the first dimension may further include any one of the batch dimension N, the depth dimension D, the height dimension H, or the number of channels dimension C of the tensor to be processed. That is, in other different cases, it may be required that the tensor to be processed be aligned in multiples of Y on any one of the batch dimension N, the depth dimension D, the height dimension H, or the number of channels dimension C. Specifically, it can be selected according to the actual situation, and the embodiments of the present disclosure do not limit this. For example, at least one embodiment of the present disclosure mainly describes the case where the first dimension includes the width dimension W; in other cases, the implementation process and principle of the second loading method of the tensor to be processed are similar to those when the first dimension is the width dimension W, and will not be elaborated here.

[0119] Figure 6A FIG. shows a schematic diagram of an example of the second loading method of the tensor to be processed provided by at least one embodiment of the present disclosure.

[0120] For example, as Figure 6A shown, in the storage space, there is an original tensor, which may be stored in the memory in a data storage format such as Figure 3A and Figure 3B shown, such as N(C / x)DHW(xC) or NDHWC, that is, the original tensor is a 5D array, and only the three dimensions W, H, and C are schematically shown in Figure 6A . For example, the dashed box is the tensor to be processed, which includes some elements in the original tensor, and the data storage format of the tensor to be processed is the same as that of the original tensor.

[0121] For example, the shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and the sizes of the tensor to be processed in each dimension are represented by a1, a2, a3, a4, a5. Assuming that a1, b1 are the N dimension sizes, a2, b2 are the D dimension sizes, a3, b3 are the H dimension sizes, a4, b4 are the W dimension sizes, and a5, b5 are the C dimension sizes as an example for description. Assuming that the N dimension size and the D dimension size are both equal to 1, in Figure 6A the shown example, b1 = 1, b2 = 1, b3 = 8, b4 = 16, b5 = 16, a1 = 1, a2 = 1, a3 = 8, a4 = 12, a5 = 10.

[0122] For example, as Figure 6A shown, in the second loading method, multiple dimensions of the tensor to be processed include the first dimension W, the second dimension H, and the third dimension C. The data storage format of the tensor to be processed indicates that the first dimension W takes precedence over the second dimension H during storage or loading, and indicates that the third dimension C takes precedence over the first dimension W during storage or loading. The first dimension W is adjacent to the second dimension H, and the first dimension W is adjacent to the third dimension C. For example, the data determined by the first dimension W and the second dimension H represents a pixel, inFigure 6A In the example, W = 0 and H = 0 correspond to the first pixel in the original tensor, W = 1 and H = 0 correspond to the second pixel in the original tensor. In the tensor, the C dimension does not affect the number of pixels, and each small solid cube represents an element in a pixel.

[0123] For example, in the second loading method, all or part of the data in the original tensor can be loaded into the buffer as a tensor to be processed through a data loading instruction. Therefore, during the loading process, it is necessary to indicate the size of the tensor to be processed in each dimension, and it is required that the data size of the tensor to be processed in the first dimension W is aligned according to the multiple of Y. This is different from the first loading method described above (the first loading method needs to indicate the number of pixels included in the tensor to be processed and requires that the number of pixels is aligned according to the multiple of Y).

[0124] For example, as Figure 6A shown, the tensor to be processed can be loaded into the buffer as a whole from the original tensor. Specifically, the starting point of the original tensor to be loaded (C = 0, W = 0, H = 0), the starting point of the tensor to be processed in the coordinate system of the original tensor (C = 0, W = 3, H = 0), and the size of the tensor to be processed in each dimension can be indicated in the data loading instruction. Schematically, in the Figure 6A shown three-dimensional schematic diagram, the tensor to be processed is a cuboid determined by the above starting point and the size in each dimension. Through the above information, the tensor to be processed can be loaded from the storage space where the original tensor is located into the cache area of the processor to perform corresponding calculations on the tensor to be processed in the processor.

[0125] Figure 6B The schematic diagram shows another example of the data loading method provided by at least one embodiment of the present disclosure. For example, Figure 6B is Figure 4 another example of the data loading method shown.

[0126] For example, Figure 6B the example of Figure 6A obtains the tensor to be processed in a similar second loading method as shown in Figure 6B and loads the tensor to be processed into the buffer as a whole from the original tensor. For example, in the Figure 6B example, both the tensor to be processed and the original tensor are in the NDHWC data storage format, that is, including the first dimension W, the second dimension H, the third dimension C, the fourth dimension D, and the fifth dimension N. In

[0127] For example, inFigure 6B In the example, the size of the tensor to be processed in the C dimension is 10 (C_dim = 10), the size in the W dimension is 12 (W_dim = 12), the size in the H dimension is 8 (H_dim = 8), the size in the D dimension is 1 (D_dim = 1), and the size in the N dimension is 1 (N_dim = 1). For example, assuming the size of the original tensor in the third dimension C is tensor_c = 16, then the size of the tensor to be processed in the third dimension C, C_dim (C_dim = 10), is smaller than the size of the original tensor in the third dimension C, tensor_c (tensor_c = 16), indicating that the tensor to be processed in this example cannot be continuously loaded in the third dimension C. For example, in response to the inability to continuously load the tensor to be processed in the third dimension C, multiple first requests can be divided in the first dimension W. Each first request in each loading request group loads the sub-data corresponding to one pixel, and the next first request in this loading request group can be used to load the data in the adjacent next pixel, that is, the requests are divided in a pixel-skipping manner. For example, the sub-data loaded by the first requests in the same loading request group belongs to the same row in the W dimension, and the sub-data loaded by the first requests in different loading request groups is located in different rows in the W dimension.

[0128] For example, as Figure 6B shown, in step S110, the loading method of the tensor to be processed (i.e., the second loading method), the data storage format NDHWC of the tensor to be processed, the shape size of the tensor to be processed (i.e., the sizes of each dimension), and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor (i.e., the coordinates of the pixel where the starting position of the dashed arrow is located) are obtained. For example, the starting point coordinates of the tensor to be processed in the coordinate system determined by the original tensor to be obtained are represented as W_coord_b = 3, H_coord_b = 0, D_coord_b = 0, N_coord_b = 0, C_coord_b = 0.

[0129] For example, in step S120, based on the above information, multiple first requests for loading the tensor to be processed can be determined. For example, the multiple first requests are divided into 8 loading request groups based on the second loading method (i.e., X = H_dim = 8), and each loading request group is used to load each row of pixels in the W dimension. Each first request in this loading request group is used to correspondingly load one pixel in that row of pixels. For example, the data size of each loading request group in the first dimension W (i.e., the size of each row of pixels in the W dimension) is required to be aligned in multiples of 16 (i.e., Y = 16).

[0130] For example, in step S130, the first requests in each of the 8 loading request groups are sequentially sent, and data is sequentially fetched from the storage space to obtain the first data returned by each loading request group. For example, taking the first loading request group as an example, 12 first requests in the first loading request group are sent. Using the pixel coordinates at the starting position of the leftmost dashed arrow as the starting point coordinates, 12 pixels are fetched pixel by pixel from the original tensor in the W dimension in the order of the first dashed arrow (each first request corresponds to one pixel), and the first data returned by the first loading request group is obtained (i.e., the first row of pixels of the tensor to be processed in the W dimension). For example, in Figure 6B In the example of, the size Copy_w of the first data loaded by the first loading request group is required to be aligned in multiples of 16.

[0131] For example, as Figure 6B shown, the size of the first data returned by the first loading request group in the W dimension is 12. Since 12 is not a multiple of 16, that is, 16 - Copy_w[3:0] is not equal to 0 (where Copy_w[3:0] is the binary [3:0] bits of the size of the first data), the size of the first data returned by the first loading request group is not aligned in multiples of 16. Therefore, in step S1402, the first data needs to be padded to obtain the second data, and the size of the data to be padded is 16 - Copy_w[3:0] = 4.

[0132] For example, specifically, in Figure 6B In the example of, in step S1402, the first data is padded pixel by pixel in the order of the second dashed arrow, and the padded data is represented as the pixels in the shaded part where the second dashed arrow is located in Figure 6B (the size of the padded data is 4). For example, the second data obtained by padding includes the first data and the above-mentioned padded data (i.e., the data in the part where the two dashed arrows are located), and the size of the second data is 16, which can be aligned in multiples of 16. For example, each pixel in the padded part can be padded with 0 or other specified values, and the embodiments of the present disclosure do not limit this.

[0133] For example, further, in step S150, since the second data meets the alignment requirements of the hardware (i.e., is aligned in multiples of 16), the second data can be directly written into the buffer for computing operations such as GEMM in the processor.

[0134] It should be noted that in addition to Figure 6B loading pixels pixel by pixel in the example of, pixels can also be loaded in a way of skipping pixels by setting the stride, and its implementation principle is basically similar to the example implementation method provided by at least one embodiment of the present disclosure, and will not be elaborated here.

[0135] Figure 6C A schematic diagram showing another example of the data loading method provided by at least one embodiment of the present disclosure. For example, Figure 6C is Figure 4 another example of the data loading method shown.

[0136] For example, similar to the example of Figure 6B , the example of Figure 6C obtains the tensor to be processed in a similar second loading manner as shown in Figure 6A . Both the tensor to be processed and the original tensor are in the NDHWC data storage format. For example, the difference from the example of Figure 6B is that in the example of Figure 6C , the size of the tensor to be processed in the C dimension is 16 (C_dim = 16), and the sizes of other dimensions and the starting coordinates, etc. are the same as those in the example of Figure 6B , which will not be elaborated here.

[0137] For example, in the example of Figure 6C , assuming that the size of the original tensor in the third dimension C is tensor_c = 8, then the size C_dim (C_dim = 16) of the tensor to be processed in the third dimension C is equal to the size tensor_c (tensor_c = 16) of the original tensor in the third dimension C, indicating that the tensor to be processed in this example can be continuously loaded in the third dimension C. For example, the starting point coordinate W_coord_b = 3 of the tensor to be processed in the coordinate system determined by the original tensor (that is, the starting coordinate in the W dimension is not equal to 0), and it can be seen from Figure 6C that the size (W_dim = 12) of the tensor to be processed in the W dimension is smaller than the size tensor_w of the original tensor in the W dimension. For example, in this case, although the tensor to be processed can be continuously loaded in the third dimension C, it cannot be continuously loaded in the first dimension W. Thus, multiple first requests can be divided in the second dimension H, and each loading request group can include one first request, and the data loaded by this first request is the first data. For example, as shown in Figure 6C , the first data loaded by this first request can be a whole row of pixel data of the tensor to be processed in the W dimension.

[0138] For example, for the data loading process, the difference from the example of Figure 6B is that in Figure 6CIn the example of Figure 6B , in step S130, the first requests in each of the 8 loading request groups are sent in sequence. Taking the first loading request group as an example, the first request in the first loading request group is sent. Using the pixel coordinates at the starting position of the leftmost dotted arrow as the starting point coordinates, the first row of pixels in the W dimension is obtained from the original tensor, and the first data returned by the first loading request group is obtained (i.e., the first row of pixels in the W dimension of the tensor to be processed). For example, similar to

[0139] For example, as Figure 6C shown, the size Copy_w of the first data returned by the first loading request group is 12. Since 12 is not a multiple of 16, that is, 16 - Copy_w[3:0] is not equal to 0, the size of the first data returned by the first loading request group is not aligned according to a multiple of 16. For example, in step S1402, the first data is filled according to the second dotted arrow, and a whole row of data in the shaded part is filled each time. The size of the data filled each time is 16 - Copy_w[3:0] = 4. For example, the second data obtained by filling includes the first data and the filled data described above. The size of this second data is 16 and can be aligned according to a multiple of 16.

[0140] For example, except for the differences described above, Figure 6C in the data loading process of the example of Figure 6B is basically the same as that of the example of

[0141] and will not be elaborated here. Figure 6A and Figure 6B The second loading method, the original tensor, the tensor to be processed, the data acquisition method, the filling method, etc. shown are only exemplary. The second loading method, the original tensor, the tensor to be processed, the data acquisition method, the filling method, etc. can also be selected in other forms according to actual needs, and the embodiments of the present disclosure do not limit this; in addition, the tensor to be processed and the original tensor can also be in other data storage formats such as N(C / x)DHW(xC), and their implementation principles are basically similar to the example implementation methods provided by at least one embodiment of the present disclosure, and will not be elaborated here.

[0142] It can be understood that the above processor can be reasonably used according to the type of operation to be performed or the data processing characteristics, whether it is the continuous first data loading method or the block-based second data loading method, that is, it can support the adaptive switching between these two loading methods, and will not be further expanded here.

[0143] In some examples, for the specific filling process of the data to be processed, Figure 4Step S140 can further include the following steps S141 to S143.

[0144] Step S141: In response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, determine the size of the data to be filled for the first data based on the part of the size of the first data that is less than a multiple of Y;

[0145] Step S142: Determine at least one second request for filling the first data based on the size of the data to be filled for the first data;

[0146] Step S143: Send at least one second request to fill the first data to obtain second data.

[0147] For example, in step S141, if the size of the first data returned by the k-th loading request group does not meet the alignment requirement, the size of the data to be filled for the first data can be determined according to the part of the size of the first data that does not meet the alignment requirement; in step S142, one or more second requests are determined according to the size of the data to be filled for the first data; in step S143, the first data is filled with the filling data requested by the one or more second requests, and the size of the filling data requested by the one or more second requests is equal to the size of the data to be filled for the first data determined in step S141; thus, the size of the second data obtained in step S143 can be aligned in multiples of Y, so as to meet the hardware alignment requirement and can be used in the processor to execute corresponding computing tasks.

[0148] For example, each second request includes a data write address for indicating the starting position of filling data into the buffer area and the size of the filling data requested by the second request, so that the filling data requested by each second request can be loaded into the corresponding position in the buffer area.

[0149] For example, the request initial coordinate of the first second request in the at least one second request can be determined according to the request initial coordinate of the last first request in the k-th loading request group, and the request initial coordinate corresponding to the next second request can also be determined according to the request initial coordinate corresponding to the previous second request. For example, the request initial coordinate corresponding to each first request is used to determine the data read address, data write address, and the size of the data loaded by the first request, and the request initial coordinate corresponding to each second request is used to determine the data write address of the second request and the size of the filling data requested by the second request.

[0150] For example, in response to k being less than X (that is, the k-th loading request group is not the last loading request group of the data to be processed), the request initial coordinate corresponding to the first first request of the (k + 1)-th loading request group can be determined according to the request initial coordinate corresponding to the last second request in the at least one second request.

[0151] For example, in some examples, the sub-data requested by each second request is 0 or other specified values, and the embodiments of the present disclosure do not limit this.

[0152] For example, taking Figure 5B as an example, in step S141, since the number of pixels (copy_pixel_num = 40) included in the first data returned by the target load request group is not a multiple of 16, the size of the first data that does not meet the alignment requirement can be used to determine the size of the data to be filled in the first data, 16 - copy_pixel_num [3:0], that is, the number of pixels to be filled in the first data is copy_pixel_num [3:0] = 8. For example, in step S142, based on the size of the data to be filled in the first data, 16 - copy_pixel_num [3:0] = 8, one or more second requests are determined, and the number of pixels of the padding data requested by the one or more second requests is equal to the number of pixels of the data to be filled in the first data, 16 - copy_pixel_num [3:0] = 8. For example, in step S143, the padding data requested by the one or more second requests is used to fill the first data pixel by pixel in the order of the dashed arrows, and the sub-data requested by each second request corresponds to one padding pixel. For example, the padding data requested by the one or more second requests is Figure 5B represented as the pixels where the dashed arrows are located in

[0153] For example, taking Figure 6BTaking the first load request group in the example as an example, the size of the first data returned by the first load request group in the W dimension is 12. In step S141, since 12 is not a multiple of 16, the size of the part where the size of the first data does not meet the alignment requirement can be used to determine the size of the data to be filled for the first data as 16 - Copy_w[3:0] = 4, that is, the size of the data to be filled for the first data in the W dimension is 4. For example, in step S142, according to the size of the data to be filled for the first data as 16 - Copy_w[3:0] = 4, one or more second requests are determined, and the size of the filled data requested by the one or more second requests in the W dimension is equal to the size of the data to be filled for the first data as 16 - Copy_w[3:0] = 4. For example, in step S143, the filled data requested by the one or more second requests is used to fill the first data pixel by pixel in the order of the second dotted arrow, and each sub - data requested by each second request corresponds to one filled pixel. For example, the filled data requested by the one or more second requests is Figure 6B represented as the pixels in the shaded part where the second dotted arrow is located (the size of the filled data is 4). For example, the filled second data includes the first data and the above - mentioned filled data (that is, the data in the part where the two dotted arrows are located), and the size of the second data is 16, which can be aligned in multiples of 16, thus meeting the hardware alignment requirement and being able to execute computing tasks such as GEMM in the processor.

[0154] It should be noted that, in addition to Figure 5B and Figure 6B filling pixels in a pixel - by - pixel manner in the example, pixels can also be filled in a pixel - skipping manner by setting a step size. The implementation principle is basically similar to the example implementation method provided by at least one embodiment of the present disclosure, and will not be elaborated here.

[0155] For example, Figure 6C the filling process of the first load request group in the example of Figure 6B is similar to the example of Figure 6C The difference is that in step S142, according to the size of the data to be filled for the first data as 16 - Copy_w[3:0] = 4, one second request is determined; in step S143, the filled data requested by the second request is used to fill the first data according to the second dotted arrow, and the filled data requested by the second request corresponds to a whole row of data in the shaded part, and the size of the data filled each time is 16 - Copy_w[3:0] = 4. For example, Figure 6B the other steps of the filling process of the example of

[0156] It should be noted that Figure 5B , Figure 6A and Figure 6BThe filling method shown is only exemplary. The data filling principle and process of tensors to be processed in other implementation forms are basically similar to the example implementation provided by at least one embodiment of the present disclosure, and will not be elaborated here.

[0157] The data loading method provided by at least one embodiment of the present disclosure divides multiple first requests for loading tensors to be processed into X loading request groups, and can fill the data requested by each loading request group so that the filled data is aligned in multiples of Y. Thus, the hardware can directly process the filled data without additional intervention from software, thereby greatly reducing the burden on software, reducing the complexity of instructions, ensuring efficient transmission and processing of data at the hardware level, and at the same time avoiding risks such as calculation errors and memory access errors caused by misaligned data, and improving the overall performance of the processor.

[0158] In some examples, the data loading method provided by at least one embodiment of the present disclosure may further include step S160: in response to the second data corresponding to the Xth loading request group among the X loading request groups being written into the buffer area, based on the data write address of the starting position of the tensor to be processed in the buffer area and the size of the data after the tensor to be processed is filled, determine the data write address of the starting position of the next tensor to be processed in the buffer area.

[0159] For example, in step S160, after the second data corresponding to the Xth loading request group (i.e., the last loading request group of the tensor to be processed) is written into the buffer area, the current tensor to be processed is fully loaded into the processor's buffer area. For example, if the next tensor to be processed needs to be loaded from the storage space into the buffer area, based on the data write address of the starting position of the current tensor to be processed in the buffer area and the size of the data after it is filled, the data write address of the starting position of the next tensor to be processed in the buffer area can be calculated, so that starting from this data write address, in the same way as the current tensor to be processed, the next tensor to be processed can be loaded into the processor's buffer area to execute the next calculation task.

[0160] The data loading method provided by at least one embodiment of the present disclosure can automatically calculate the data write address of the starting position of the next tensor to be processed in the buffer area after completing the filling operation of the tensor to be processed, without the software sending a separate instruction to calculate the next starting address, thereby greatly reducing the burden on software, reducing the complexity of instructions, and improving the overall performance of the processor.

[0161] At least one embodiment of the present disclosure also provides another data loading method. Figure 7 Another schematic flowchart showing the data loading method provided by at least one embodiment of the present disclosure is shown.

[0162] For example, asFigure 7 As shown, the data loading method according to an embodiment of the present disclosure includes the following steps S201 to S202.

[0163] Step S201: Receive a data loading instruction instructing to execute loading a tensor to be processed from the storage space where the original tensor is located into the buffer.

[0164] For example, in step S201, the tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading instruction includes the loading method of the tensor to be processed as an input parameter, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.

[0165] For the relevant descriptions of the tensor to be processed and the original tensor, reference can be made to the relevant descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.

[0166] For example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction can be a Load instruction.

[0167] Step S202: After parsing the data loading instruction, use the first execution unit to execute the data loading instruction.

[0168] For example, after receiving the data loading instruction, the processor parses the data loading instruction, for example, decodes the data loading instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameter is ready, the execution unit executes the relevant operations of the data loading instruction.

[0169] For example, step S202 may include: determining a plurality of first requests for loading the tensor to be processed based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, where the plurality of first requests are divided into X groups of loading requests based on the loading method of the tensor to be processed, and the size of the data loaded by each group of loading requests in the X groups of loading requests is required to be aligned in multiples of Y, and here X and Y are positive integers; sequentially sending the first requests in the k-th group of loading requests in the X groups of loading requests, sequentially obtaining data from the storage space, and obtaining the first data returned by the k-th group of loading requests, where k = 1, 2,..., X; in response to the size of the first data returned by the k-th group of loading requests not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the second data obtained by padding is aligned in multiples of Y; and writing the second data into the buffer.

[0170] For example, the relevant content of the specific steps included in step S202 may refer to the relevant descriptions of steps S110 to S150 in the foregoing Figure 4 and will not be elaborated herein for the repeated parts.

[0171] In some examples, in addition to loading the tensor in the storage space where the original tensor is located into the buffer of the processor through a data loading instruction, the tensor in the buffer can also be stored into the storage space where the original tensor is located through a data storage instruction.

[0172] At least one embodiment of the present disclosure further provides a data storage method for writing the tensor to be stored in the buffer into the storage space where the original tensor is located. For example, the tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving the data from the buffer to the storage space of the memory (such as HBM, etc.), and the implementation principle thereof according to the embodiments of the present disclosure is similar to the above data loading method and will not be described repeatedly, and only the different parts will be described in detail.

[0173] Figure 8 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure. For example, as Figure 8 shown, the data storage method provided by at least one embodiment of the present disclosure includes the following steps S310 to S350.

[0174] Step S310: Obtain the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer.

[0175] For example, in step S310, the methods for obtaining the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, the starting coordinates of the tensor to be stored in the buffer, etc. are basically similar to those of the tensor to be processed in step S110 in Figure 4 and can specifically refer to the descriptions in the above text, which will not be elaborated here.

[0176] Step S320: Based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determine multiple third requests for storing the tensor to be stored.

[0177] For example, in step S320, the multiple third requests are divided into X storage request groups based on the storage method of the tensor to be stored, and the size of the data stored in each storage request group among the X storage request groups is required to be aligned in multiples of Y, where X and Y are positive integers. For example, the method for determining the multiple third requests and the division method of the storage request groups are basically similar to those of the first request in step S120 in Figure 4 and can specifically refer to the descriptions in the above text, which will not be elaborated here.

[0178] Step S330: Sequentially send the third requests in the k-th storage request group among the X storage request groups, sequentially obtain data from the buffer, and obtain the third data returned by the k-th storage request group, where k = 1, 2,..., X.

[0179] For example, in step S330, the specific method for obtaining data from the buffer based on the third request is basically similar to the specific method for obtaining data from the storage space based on the first request in step S130 in Figure 4 and can specifically refer to the descriptions in the above text, which will not be elaborated here.

[0180] Step S340: In response to the fact that the size of the third data returned by the k-th storage request group is not aligned in multiples of Y, pad the third data to obtain the fourth data.

[0181] For example, in step S340, the size of the obtained fourth data is aligned in multiples of Y. For example, the specific method for padding the third data to obtain the fourth data is basically similar to the specific method for padding the first data to obtain the second data in step S140 in Figure 4 and can specifically refer to the descriptions in the above text, which will not be elaborated here.

[0182] Step S350: In response to k being equal to 1, write the third data returned by the first storage request group into the storage space;

[0183] Step S360: In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, write the third data returned by the (k + 1)-th storage request group into the storage space.

[0184] For example, the difference from the data loading process is that during the data storage process, if the data requested for storage does not belong to the content of the original tensor, there is no need to send a request, and this data does not need to be stored in the storage space where the original tensor is located; if the data requested for storage belongs to the original tensor, a request is sent to the storage space to store the data at the corresponding position in the original tensor. For example, for each storage request group, since the data in the padding part does not belong to the content of the original tensor, only the non-padding part of the fourth data (i.e., the third data) can be written into the storage space; and the purpose of padding the third data is to calculate the starting address of the next storage request group, so that the third data requested by the next storage request group can be written into the correct position in the storage space.

[0185] For example, specifically, in step S350, when k is equal to 1, the starting address of the first storage request group is the starting address determined by the starting coordinates of the tensor to be stored in the buffer. Since this starting address is known, the third data returned by the first storage request group can be directly written into the buffer; in step S360, since the starting addresses of the second to the X-th storage request groups need to be calculated, when k is less than X, the starting address of the next storage request group can be calculated based on the fourth data corresponding to the k-th storage request group, so that the third data returned by the (k + 1)-th storage request group can be written into the correct position in the storage space, avoiding calculation errors of the storage address.

[0186] For example, specifically, Figure 8 Step S360 of can further include the following steps S361 to S362.

[0187] Step S361: In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, determine the data read address at the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data write address at the starting position in the storage space;

[0188] Step S362: Based on the data read address at the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data write address at the starting position in the storage space, write the third data returned by the (k + 1)-th storage request group into the storage space.

[0189] For example, to obtain the starting addresses of the second to the Xth storage request groups, in step S361, when k is less than X, the starting address of the (k + 1)th storage request group is calculated based on the fourth data corresponding to the kth storage request group. The starting address includes the data read address of the starting position of the third data requested by the (k + 1)th storage request group in the buffer and the data write address of the starting position in the storage space. Further, in step S362, based on the data read address of the starting position of the third data requested by the (k + 1)th storage request group in the buffer and the data write address of the starting position in the storage space calculated above, the third data returned by the (k + 1)th storage request group can be written to the correct position in the storage space.

[0190] For example, in some examples, each third request includes a data read address for indicating the starting position of reading data from the buffer, a data write address for indicating the starting position of writing data to the storage space, and the size of the data stored in the third request. For example, the request initial coordinates corresponding to the next third request are determined according to the request initial coordinates corresponding to the previous third request. The request initial coordinates corresponding to each third request are used to determine the data read address, the data write address, and the size of the data stored in the third request.

[0191] In some examples, the storage mode of the tensor to be stored includes a first storage mode. For example, in the first storage mode, multiple dimensions of the tensor to be stored include a first dimension and a second dimension, and the data determined by the first dimension and the second dimension is represented as a pixel. For example, Figure 8 Step S320 of

[0192] Step S321: Determine the number of pixels included in the tensor to be stored;

[0193] Step S322: Based on the number of pixels included in the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determine multiple third requests for storing the tensor to be stored.

[0194] For example, multiple third requests are divided into 1 target storage request group, and the number of pixels included in the third data stored in the target storage request group is required to be aligned in multiples of Y.

[0195] For example, in response to the storage mode of the tensor to be stored being the first storage mode, Figure 8Step S340 may further include step S3401: in response to the number of pixels included in the third data returned by the target storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data. For example, the number of pixels included in the obtained fourth data is aligned in multiples of Y.

[0196] For example, in some examples, the multiple dimensions of the tensor to be stored include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, the first dimension includes the width dimension, the second dimension includes the height dimension, and the number of channels dimension does not count the number of pixels.

[0197] In some examples, the storage method of the tensor to be stored includes a second storage method. For example, in the second storage method, the multiple dimensions of the tensor to be stored include a first dimension, and the data size of the tensor to be stored in the first dimension is required to be aligned in multiples of Y. For example, Figure 8 Step S340 may further include step S3402: in response to the size of the third data returned by the k-th storage request group in the first dimension not being aligned in multiples of Y, padding the third data in the first dimension to obtain fourth data. For example, the size of the obtained fourth data in the first dimension is aligned in multiples of Y.

[0198] For example, the multiple dimensions of the tensor to be stored further include a second dimension and a third dimension, the data storage format of the tensor to be stored indicates that the first dimension has priority over the second dimension during storage or loading, and indicates that the third dimension has priority over the first dimension during storage or loading, the first dimension is adjacent to the second dimension, and the first dimension is adjacent to the third dimension.

[0199] For example, the data represented by the first dimension and the second dimension is represented as one pixel. For example, in response to the tensor to be stored not being continuously storable in the third dimension, the sub-data stored by each third request in the k-th storage request group is the data corresponding to one pixel; or, in response to the tensor to be stored being continuously storable in the third dimension, the data stored by the third request in the k-th storage request group is the third data.

[0200] For example, in some examples, the multiple dimensions of the tensor to be processed include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, the first dimension includes the width dimension, the second dimension includes the height dimension, and the third dimension includes the number of channels dimension. For example, in some other examples, the first dimension may include any one of the batch dimension, the depth dimension, the height dimension, the width dimension, or the number of channels dimension of the tensor to be processed, and the embodiments of the present disclosure do not limit this.

[0201] In some examples, for the specific padding process of the data to be stored, Figure 8Step S340 may further include the following steps S341 to S343.

[0202] Step S341: In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, determine the size of the data to be filled for the third data based on the part of the size of the third data that is less than a multiple of Y.

[0203] Step S342: Based on the size of the data to be filled for the third data, determine at least one second request for filling the third data.

[0204] Step S343: Send at least one second request to fill the third data to obtain the fourth data. For example, the size of the filling data requested by at least one second request is equal to the size of the data to be filled for the third data.

[0205] It can be understood that for the first storage method, the second storage method, and the filling method of the data storage method provided by at least one embodiment of the present disclosure, they can be implemented in a similar manner to the first loading method, the second loading method, and the filling method of the data loading method provided by at least one embodiment of the present disclosure. Those skilled in the art can similarly apply the data loading process described above to the data storage process. Specifically, reference can be made to the description above and will not be elaborated here.

[0206] The data storage method provided by at least one embodiment of the present disclosure divides a plurality of third requests for storing a tensor to be stored into X storage request groups, and can fill the data requested by each storage request group so that the filled data is aligned in multiples of Y, thereby avoiding calculation errors of storage addresses. The hardware can correctly store the tensor to be stored without additional intervention by software, and the starting address of the next data to be stored can be automatically obtained after the filling operation is completed without the software separately sending an instruction to calculate the next starting address. This greatly reduces the burden on the software, reduces the complexity of instructions, ensures efficient storage of data at the hardware level, and at the same time avoids risks such as memory access errors caused by unaligned data, improving the overall performance of the processor.

[0207] At least one embodiment of the present disclosure also provides another data storage method. Figure 9 Another schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown.

[0208] For example, as Figure 9 shown, the data storage method according to an embodiment of the present disclosure includes the following steps S401 to S402.

[0209] Step S401: Receive a data storage instruction indicating to execute writing the tensor to be stored in the buffer into the storage space where the original tensor is located.

[0210] For example, in step S401, the tensor to be stored and the original tensor are multi-dimensional tensors. The data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage instruction includes the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer.

[0211] For the relevant descriptions of the tensor to be stored and the original tensor, reference can be made to the relevant descriptions of the foregoing data storage method, and the repeated parts will not be elaborated.

[0212] For example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction can be a Store instruction.

[0213] Step S402: After parsing the data storage instruction, use the second execution unit to execute the data storage instruction.

[0214] For example, after receiving the data storage instruction, the processor parses the data storage instruction, for example, decodes the data storage instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameters are ready, the execution unit performs the relevant operations of the data storage instruction.

[0215] For example, step S402 may include: determining a plurality of third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer. Among them, the plurality of third requests are divided into X storage request groups based on the storage method of the tensor to be stored. The size of the data stored in each storage request group among the X storage request groups is required to be aligned in multiples of Y. Here, X and Y are positive integers; sequentially send the third requests in the k-th storage request group among the X storage request groups, sequentially obtain data from the buffer, and obtain the third data returned by the k-th storage request group. Here, k = 1, 2,..., X; in response to the fact that the size of the third data returned by the k-th storage request group is not aligned in multiples of Y, fill the third data to obtain the fourth data, where the size of the filled fourth data is aligned in multiples of Y; in response to k being equal to 1, write the third data returned by the first storage request group into the storage space; in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, write the third data returned by the (k + 1)-th storage request group into the storage space.

[0216] For example, for the relevant content of the specific steps included in step S402, reference can be made to the foregoingFigure 8 For the relevant descriptions of steps S310 to S360, the repeated parts will not be elaborated again.

[0217] Figure 10 FIG. shows a schematic block diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 10 shown, the processor 200 includes a first instruction parsing unit 201 and a first execution unit 202.

[0218] For example, the first instruction parsing unit 201 is configured to receive and parse data loading instructions.

[0219] For example, the data loading instruction is used to load a tensor to be processed from the storage space where the original tensor is located into the buffer area. The tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading instruction includes the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters.

[0220] For the relevant descriptions of the tensor to be processed and the original tensor, reference can be made to the relevant descriptions of the foregoing data loading method, and the repeated parts will not be elaborated again.

[0221] For example, after the first instruction parsing unit 201 parses the data loading instruction, the first execution unit 202 executes the data loading instruction.

[0222] For example, when the first execution unit 202 executes the data loading instruction, it includes performing the following operations: determining a plurality of first requests for loading the tensor to be processed based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Among them, the plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed. The size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. Here, X and Y are positive integers; sequentially sending the first requests in the kth loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining the first data returned by the kth loading request group. Here, k = 1, 2,..., X; in response to the size of the first data returned by the kth loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y; writing the second data into the buffer area.

[0223] Specifically, when upper-layer software based on a processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data loading instructions for computing and processing to the processor (such as a CPU or GPU) through a unified encapsulated function library, the data loading instructions can carry the loading method of the tensor to be processed as an input parameter, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; when the processor receives the data loading instructions, the first instruction parsing unit 201 parses the data loading instructions to obtain the loading method of the tensor to be processed as an input parameter, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data loading instructions, the processor can store the input parameters in the data loading instructions in a register or memory, and the first execution unit 202 can obtain the input parameters from the register or memory when performing computing and processing.

[0224] Regarding the specific process of using the first execution unit 202 to execute the data loading instructions, reference can be made to steps S110~S160 in the data loading method described above, and repeated parts will not be elaborated.

[0225] The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method, and repeated parts will not be elaborated.

[0226] Figure 11 The schematic block diagram of another processor provided by at least one embodiment of the present disclosure is shown. As Figure 11 shown, the processor 300 includes a second instruction parsing unit 301 and a second execution unit 302.

[0227] For example, the second instruction parsing unit 301 is used to receive and parse data storage instructions.

[0228] For example, the data storage instructions are used to execute writing the tensor to be stored in the buffer into the storage space where the original tensor is located. The tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage instructions include the storage method of the tensor to be stored as an input parameter, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer.

[0229] Regarding the relevant descriptions of the tensor to be stored and the original tensor, reference can be made to the relevant descriptions of the foregoing data storage method, and repeated parts will not be elaborated.

[0230] For example, after the second instruction parsing unit 301 parses the data storage instruction, the second execution unit 302 executes the data storage instruction.

[0231] For example, when the second execution unit 302 executes the data storage instruction, it includes performing the following operations: determining a plurality of third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer area, where the plurality of third requests are divided into X storage request groups based on the storage method of the tensor to be stored, and the size of the data stored in each of the X storage request groups is required to be aligned in multiples of Y, and here X and Y are positive integers; sequentially sending the third requests in the k-th storage request group among the X storage request groups, sequentially obtaining data from the buffer area, and obtaining the third data returned by the k-th storage request group, where k = 1, 2,..., X; in response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain the fourth data, where the size of the padded fourth data is aligned in multiples of Y; in response to k being equal to 1, writing the third data returned by the first storage request group into the storage space; in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, writing the third data returned by the (k + 1)-th storage request group into the storage space.

[0232] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data storage instructions for computing and processing to the processor (such as CPU or GPU) through a unified encapsulated function library, the data storage instruction can carry the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer area as input parameters; when the processor receives the data storage instruction, the second instruction parsing unit 301 parses the data storage instruction to obtain the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer area as input parameters, and the processor schedules the operation unit to execute the data storage task for the input parameters. For example, after parsing the data storage instruction, the processor can store the input parameters in the data storage instruction in a register or memory, and when the second execution unit 302 performs computing and processing, it can obtain the input parameters from the register or memory.

[0233] Regarding the specific process of using the second execution unit 302 to execute the data storage instruction, reference can be made to steps S310 - S360 and the like in the data storage method described above, and the repeated parts will not be elaborated.

[0234] The processor provided by at least one embodiment of the present disclosure can achieve a technical effect similar to the foregoing data storage method, and the repeated parts will not be described again.

[0235] At least one embodiment of the present disclosure further provides an electronic device, which includes one or more processors and a memory; the memory includes one or more computer program modules; one or more computer program modules are stored in the memory and configured to be executed by one or more processors, and one or more computer program modules include those for implementing the data loading method or the data storage method provided by at least one embodiment of the present disclosure described above.

[0236] Figure 12 A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0237] For example, as Figure 12 shown, the electronic device 400 may include a processor 410 and a memory 420 connected to the processor 410. The number of processors 410 may be one or more; in addition, the processor 410 may further include a buffer. According to an embodiment of the present disclosure, the memory 420 may be implemented in the form of a high-bandwidth memory HBM, or may be implemented in other forms according to actual needs, and the embodiments of the present disclosure do not limit this. Specifically, according to at least one embodiment of the present disclosure, the processor 410 is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor 410, the data loading method according to the embodiment of the present disclosure is implemented to load a tensor to be processed from the storage space of the memory where the original tensor is located into the buffer, or the data storage method according to the embodiment of the present disclosure is implemented to write the tensor to be stored in the buffer into the storage space of the memory where the original tensor is located.

[0238] The processor 410 can perform various actions and processes according to a program stored in a non-transitory memory such as a non-volatile memory. Specifically, the processor 410 may refer to a processor chip capable of performing parallel computing. For example, it may be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 410 may also be implemented as other conventional types of processors, which are not limited herein.

[0239] Regarding the specific implementation processes of the data loading method and the data storage method, reference may be made to the above description and will not be repeated herein. The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects as the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0240] Figure 13 Schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.

[0241] For example, as Figure 13 shown, the electronic device 500 is, for example, suitable for implementing the data loading method or the data storage method provided in the embodiments of the present disclosure. It should be noted that Figure 13 the shown electronic device 500 is only an example and will not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0242] For example, as Figure 13As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 51, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 52 or a program loaded from a storage device 58 into a random access memory (RAM) 53. In the RAM 53, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 51, the ROM 52, and the RAM 53 are connected to each other through a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54. Generally, the following devices may be connected to the I / O interface 55: an input device 56 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 57 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 58 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 59. The communication device 59 may allow the electronic device 500 to communicate with other electronic devices wirelessly or wiredly to exchange data.

[0243] Although Figure 13 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 500 may alternatively implement or have more or fewer devices.

[0244] It should be noted that, for the sake of clarity and conciseness, the embodiments of the present disclosure do not show all the constituent units of the electronic device 400 / 500. To implement the necessary functions of the electronic device, those skilled in the art may provide and set other unshown constituent units according to specific needs, and the embodiments of the present disclosure do not limit this.

[0245] For the detailed description and technical effects of the electronic device 400 / 500, reference may be made to the relevant descriptions of the data loading method or the data storage method above, which will not be elaborated here.

[0246] At least one embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data loading method or the data storage method provided by at least one embodiment of the present disclosure is implemented.

[0247] Figure 14 It is a schematic diagram of a storage medium provided by at least one embodiment of the present disclosure.

[0248] For example, as Figure 14 shown, the storage medium 600 stores non-transitory computer-readable instructions 610. For example, when the non-transitory computer-readable instructions 610 are executed by a computer, one or more steps of the data loading method or the data storage method described above are executed.

[0249] In some examples, the storage medium 600 can be applied to Figure 12 the electronic device 400 shown, and can be implemented as, for example, the memory 420 in the electronic device 400. In other examples, the storage medium 600 can be applied to Figure 13 the electronic device 500 shown, and can be implemented as, for example, the storage device 58 in the electronic device 500.

[0250] For example, the storage device can include any combination of one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored on the computer-readable storage medium, and the processor can run the computer-readable instructions to implement various functions of the processor. Various application programs and various data, etc., can also be stored in the storage medium.

[0251] For example, the storage medium can include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and can also be other applicable storage media.

[0252] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the figures. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0253] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a unit does not constitute a limitation on the unit itself.

[0254] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SoCs), complex programmable logic devices (CPLDs), and the like.

[0255] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0256] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0257] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.

[0258] Regarding the present disclosure, the following points need to be noted:

[0259] (1) In the drawings of the embodiments of the present disclosure, only the structures related to the embodiments of the present disclosure are involved, and other structures can refer to the general design.

[0260] (2) Without conflict, the features in the same embodiment and different embodiments of the present disclosure can be combined with each other.

[0261] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A data loading method for loading a tensor to be processed from the storage space where the original tensor is located into a buffer, where The tensor to be processed and the original tensor are multi-dimensional tensors. The data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading method includes: Obtaining the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determining a plurality of first requests for loading the tensor to be processed, wherein the plurality of first requests are divided into X loading request groups based on the loading method of the tensor to be processed, and the size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. X and Y are positive integers; Sequentially sending the first requests in the k-th loading request group among the X loading request groups, sequentially obtaining data from the storage space, and obtaining the first data returned by the k-th loading request group, where k = 1, 2,..., X; In response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y; Writing the second data into the buffer.

2. The data loading method according to claim 1, wherein Each of the plurality of first requests includes a data read address for indicating the starting position of reading data from the storage space, a data write address for indicating the starting position of writing data into the buffer, and the size of the data loaded by the first request. Determining the request initial coordinates corresponding to the next first request according to the request initial coordinates corresponding to the previous first request. The request initial coordinates corresponding to each first request are used to determine the data read address, the data write address, and the size of the data loaded by the first request.

3. The data loading method according to claim 1, wherein, The loading method of the tensor to be processed includes a first loading method. In the first loading method, multiple dimensions of the tensor to be processed include a first dimension and a second dimension, and the data represented by the first dimension and the second dimension is represented as a pixel. In response to the loading method of the tensor to be processed being the first loading method, the determining, based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, a plurality of first requests for loading the tensor to be processed includes: Determining the number of pixels included in the tensor to be processed; Determine a plurality of first requests for loading the tensor to be processed based on the number of pixels included in the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Among them, the plurality of first requests are divided into 1 target loading request group, and the number of pixels included in the first data loaded by the target loading request group is required to be aligned in multiples of Y.

4. The data loading method according to claim 3, wherein In response to the loading method of the tensor to be processed being the first loading method, and in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, including: In response to the number of pixels included in the first data returned by the target loading request group not being aligned in multiples of Y, padding the first data to obtain the second data, where the number of pixels included in the padded second data is aligned in multiples of Y.

5. The data loading method according to claim 3, wherein, The multiple dimensions of the tensor to be processed include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The first dimension includes the width dimension, the second dimension includes the height dimension, and the channel number dimension does not count the number of pixels.

6. The data loading method according to claim 1, wherein, The loading method of the tensor to be processed includes a second loading method. In the second loading method, the multiple dimensions of the tensor to be processed include a first dimension, and the data size of the tensor to be processed in the first dimension is required to be aligned in multiples of Y. In response to the loading method of the tensor to be processed being the second loading method, and in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, padding the first data to obtain second data, including: In response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y in the first dimension, padding the first data in the first dimension to obtain second data, where the size of the padded second data in the first dimension is aligned in multiples of Y.

7. The data loading method according to claim 6, wherein, The multiple dimensions of the tensor to be processed further include a second dimension and a third dimension. The data storage format of the tensor to be processed indicates that the first dimension takes precedence over the second dimension during storage or loading, and indicates that the third dimension takes precedence over the first dimension during storage or loading. The first dimension is adjacent to the second dimension, and the first dimension is adjacent to the third dimension. The data determined by the first dimension and the second dimension represents one pixel. In response to the tensor to be processed not being continuously loadable in the third dimension, each first request in the k-th loading request group loads sub-data corresponding to one pixel, or In response to the tensor to be processed being continuously loadable in the third dimension, the first request in the k-th loading request group loads the first data.

8. The data loading method according to claim 7, wherein, The multiple dimensions of the tensor to be processed include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The first dimension includes the width dimension, the second dimension includes the height dimension, and the third dimension includes the number of channels dimension.

9. The data loading method according to claim 6, wherein, The first dimension includes the batch dimension, the depth dimension, the height dimension, the width dimension, or the number of channels dimension of the tensor to be processed.

10. The data loading method according to claim 1, wherein, When the size of the first data returned in response to the k-th loading request group is not aligned in multiples of Y, padding the first data to obtain second data includes: When the size of the first data returned in response to the k-th loading request group is not aligned in multiples of Y, determining the size of the data to be padded for the first data based on the portion of the size of the first data that is less than a multiple of Y; Based on the size of the data to be padded for the first data, determining at least one second request for padding the first data; Sending the at least one second request to pad the first data to obtain the second data, where the size of the padding data requested by the at least one second request is equal to the size of the data to be padded for the first data.

11. The data loading method according to claim 10, wherein, Each second request in the at least one second request includes a data write address for indicating the starting position for filling data into the buffer area, and the size of the padding data requested by the second request. Determining the request initial coordinate of the first second request in the at least one second request according to the request initial coordinate of the last first request in the k-th loading request group, and determining the request initial coordinate of the next second request according to the request initial coordinate corresponding to the previous second request. When k is less than X, determining the request initial coordinate of the first first request corresponding to the (k + 1)-th loading request group according to the request initial coordinate corresponding to the last second request in the at least one second request. The request initial coordinate corresponding to each first request in the multiple first requests is used to determine the data read address, the data write address, and the size of the data loaded by the first request. The request initial coordinate corresponding to each second request is used to determine the data write address of the second request and the size of the padding data requested by the second request.

12. The data loading method according to claim 10, wherein, The sub-data requested by each second request in the at least one second request is 0 or other specified values.

13. The data loading method according to claim 1, further comprising: When the second data corresponding to the X-th loading request group in the X loading request groups is written into the buffer area, determining the data write address of the starting position of the next tensor to be processed in the buffer area based on the data write address of the starting position of the tensor to be processed in the buffer area and the size of the data after the tensor to be processed is padded.

14. A data loading method, comprising: Receive a data loading instruction that instructs to load a tensor to be processed from the storage space where the original tensor is located into a buffer. Here, the tensor to be processed and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be processed. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data loading instruction includes the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters; After parsing the data loading instruction, use the first execution unit to execute the data loading instruction. Among them, using the first execution unit to execute the data loading instruction includes: Based on the loading method of the tensor to be processed, the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determine multiple first requests for loading the tensor to be processed. Among them, the multiple first requests are divided into X loading request groups based on the loading method of the tensor to be processed. The size of the data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y. X and Y are positive integers; Sequentially send the first requests in the k-th loading request group among the X loading request groups, and sequentially obtain data from the storage space to get the first data returned by the k-th loading request group, where k = 1, 2,..., X; In response to the fact that the size of the first data returned by the k-th loading request group is not aligned in multiples of Y, pad the first data to obtain second data, where the size of the padded second data is aligned in multiples of Y; Write the second data into the buffer.

15. A data storage method for writing a tensor to be stored in a buffer into the storage space where the original tensor is located, where The tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage method includes: Obtain the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer; Based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determine multiple third requests for storing the tensor to be stored. Among them, the multiple third requests are divided into X storage request groups based on the storage method of the tensor to be stored. The size of the data stored by each storage request group in the X storage request groups is required to be aligned in multiples of Y. X and Y are positive integers; Send the third request in the k-th storage request group among the X storage request groups in sequence, obtain data sequentially from the buffer, and obtain the third data returned by the k-th storage request group, where k = 1, 2, …, X; In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, pad the third data to obtain the fourth data, where the size of the obtained fourth data is aligned in multiples of Y; In response to k being equal to 1, write the third data returned by the 1st storage request group into the storage space; In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, write the third data returned by the (k + 1)-th storage request group into the storage space.

16. The data storage method according to claim 15, wherein, The step of "in response to k being less than X, based on the fourth data corresponding to the k-th storage request group, write the third data returned by the (k + 1)-th storage request group into the storage space" includes: In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, determine the data reading address of the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data writing address of the starting position in the storage space; Based on the data reading address of the starting position of the third data requested by the (k + 1)-th storage request group in the buffer and the data writing address of the starting position in the storage space, write the third data returned by the (k + 1)-th storage request group into the storage space.

17. The data storage method according to claim 15, wherein, Each of the multiple third requests includes a data reading address for indicating the starting position of reading data from the buffer, a data writing address for indicating the starting position of writing data into the storage space, and the size of the data stored in the third request, Determine the request initial coordinates corresponding to the next third request according to the request initial coordinates corresponding to the previous third request, and the request initial coordinates corresponding to each third request are used to determine the data reading address, data writing address and the size of the data stored in the third request.

18. The data storage method according to claim 15, wherein, The storage method of the tensor to be stored includes the first storage method, In the first storage method, multiple dimensions of the tensor to be stored include a first dimension and a second dimension, and the data represented by the first dimension and the second dimension is represented as a pixel, In response to the storage method of the tensor to be stored being the first storage method, the step of determining multiple third requests for storing the tensor to be stored based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer includes: Determine the number of pixels included in the tensor to be stored; Based on the number of pixels included in the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determine multiple third requests for storing the tensor to be stored. Among them, the multiple third requests are divided into 1 target storage request group, and the number of pixels included in the third data stored in the target storage request group is required to be aligned in multiples of Y.

19. The data storage method according to claim 15, wherein, The storage method of the tensor to be stored includes a second storage method. In the second storage method, multiple dimensions of the tensor to be stored include a first dimension, and the data size of the tensor to be stored in the first dimension is required to be aligned in multiples of Y. In response to the storage method of the tensor to be stored being the second storage method, and in response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data includes: In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y in the first dimension, padding the third data in the first dimension to obtain fourth data, where the size of the obtained fourth data in the first dimension is aligned in multiples of Y.

20. A data storage method, including: Receiving a data storage instruction indicating to execute writing the tensor to be stored in a buffer into the storage space where the original tensor is located, where the tensor to be stored and the original tensor are multi-dimensional tensors, the data storage format of the original tensor is the same as the data storage format of the tensor to be stored, the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space, and the data storage instruction includes the storage method of the tensor to be stored, the shape size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer as input parameters; After parsing the data storage instruction, using a second execution unit to execute the data storage instruction. Among them, using the second execution unit to execute the data storage instruction includes: Based on the storage method of the tensor to be stored, the shape size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determining multiple third requests for storing the tensor to be stored, where the multiple third requests are divided into X storage request groups based on the storage method of the tensor to be stored, and the size of the data stored in each storage request group among the X storage request groups is required to be aligned in multiples of Y, and X and Y are positive integers; Sequentially sending the third requests in the k-th storage request group among the X storage request groups, sequentially obtaining data from the buffer, and obtaining the third data returned by the k-th storage request group, where k = 1, 2,..., X; In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, padding the third data to obtain fourth data, where the size of the obtained fourth data is aligned in multiples of Y; In response to k being equal to 1, writing the third data returned by the first storage request group into the storage space. In response to k being less than X, write the third data returned by the (k + 1)-th storage request group into the storage space based on the fourth data corresponding to the k-th storage request group.

21. A processor, comprising a first instruction parsing unit and a first execution unit, wherein, the first instruction parsing unit is configured to receive and parse a data loading instruction, wherein the data loading instruction is used to execute loading a tensor to be processed from a storage space where an original tensor is located into a buffer area, the tensor to be processed and the original tensor are multi-dimensional tensors, a data storage format of the original tensor is the same as a data storage format of the tensor to be processed, the data storage format is used to indicate a storage order and a dimension arrangement of the tensor in the storage space, and the data loading instruction includes a loading manner of the tensor to be processed, a shape size of the tensor to be processed, the data storage format of the tensor to be processed, and a starting coordinate of the tensor to be processed in a coordinate system determined by the original tensor as input parameters; the first execution unit executes the data loading instruction after the first instruction parsing unit parses the data loading instruction, wherein, when the first execution unit executes the data loading instruction, the following operations are included: determine a plurality of first requests for loading the tensor to be processed based on the loading manner of the tensor to be processed, the shape size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor, wherein the plurality of first requests are divided into X loading request groups based on the loading manner of the tensor to be processed, and a size of data loaded by each loading request group in the X loading request groups is required to be aligned in multiples of Y, and X and Y are positive integers; sequentially send the first requests in the k-th loading request group among the X loading request groups, sequentially obtain data from the storage space, and obtain first data returned by the k-th loading request group, wherein k = 1, 2,..., X; in response to the size of the first data returned by the k-th loading request group not being aligned in multiples of Y, perform padding on the first data to obtain second data, wherein a size of the second data obtained by padding is aligned in multiples of Y; write the second data into the buffer area.

22. A processor, comprising a second instruction parsing unit and a second execution unit, wherein, The second instruction parsing unit is configured to receive and parse a data storage instruction, where the data storage instruction is used to execute writing a tensor to be stored in a buffer into the storage space where the original tensor is located. The tensor to be stored and the original tensor are multi-dimensional tensors, and the data storage format of the original tensor is the same as that of the tensor to be stored. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage space. The data storage instruction includes the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer as input parameters; After the second instruction parsing unit parses the data storage instruction, the second execution unit executes the data storage instruction. When the second execution unit executes the data storage instruction, the following operations are included: Based on the storage method of the tensor to be stored, the shape and size of the tensor to be stored, the data storage format of the tensor to be stored, and the starting coordinates of the tensor to be stored in the buffer, determine a plurality of third requests for storing the tensor to be stored. The plurality of third requests are divided into X storage request groups based on the storage method of the tensor to be stored. The size of the data stored in each storage request group among the X storage request groups is required to be aligned in multiples of Y. X and Y are positive integers; Sequentially send the third requests in the k-th storage request group among the X storage request groups, and sequentially obtain data from the buffer to obtain the third data returned by the k-th storage request group, where k = 1, 2,..., X; In response to the size of the third data returned by the k-th storage request group not being aligned in multiples of Y, pad the third data to obtain fourth data, where the size of the padded fourth data is aligned in multiples of Y; In response to k being equal to 1, write the third data returned by the first storage request group into the storage space; In response to k being less than X, based on the fourth data corresponding to the k-th storage request group, write the third data returned by the (k + 1)-th storage request group into the storage space.

23. An electronic device, comprising: One or more processors; A memory, including one or more computer program modules; Wherein, the one or more computer program modules are stored in the memory and configured to be executed by the one or more processors. The one or more computer program modules are used to implement the data loading method according to any one of claims 1-14, or the data storage method according to any one of claims 15-20.

24. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they implement the data loading method according to any one of claims 1-14, or implement the data storage method according to any one of claims 15-20.

Citation Information

Patent Citations

  • Data processing method and device, equipment and storage medium

    CN118643253A

  • Processor operating method and device, electronic device and program product

    CN119005274A