Data Loading Method, Data Storage Method, Processor, Electronic Device, and Medium
By using coordinate information tables and state machines to optimize data loading and storage requests in parallel processors, the problem of invalid data operations is solved, and data bandwidth and hardware computing efficiency is improved.
Patent Information
- Application Number
- CN202510704090.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In parallel processors, when data is loaded and stored in the prior art, some data is loaded or stored insensitive to the expert network, resulting in invalid data operations and reducing data bandwidth and hardware computing efficiency.
By obtaining the coordinate value and size of the target data, using the coordinate information table to divide the request, data loading or storing is preferred in the channel number dimension and adjacent dimensions, ensuring that the data is continuously or discontinuously separated in memory, and a state machine is used to determine the initial state and data length of the request.
It improves data memory access bandwidth and hardware computing efficiency, reduces invalid data operations, and improves the performance of hardware computing units.
Smart Images

Figure CN120235254B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] A tensor is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. For example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor can have any number of dimensions. Tensor operations are widely used in processors such as parallel processors.
[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processor devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of parallel processors require a large amount of data, and this data is generally stored in the storage components of parallel processors. For example, the storage component can be memory. Through data loading instructions, this data can be extracted from the storage component to the buffer for calculation, and through data storage instructions, the data in the buffer can be stored in memory.
[0004] For a parallel processor deploying a network such as a mixture-of-experts model, when loading or storing data, if all data is loaded into the buffer of the hardware computing unit deploying a certain expert network or stored in memory, and some of this data may be insensitive to this expert network, this will bring a lot of invalid data loading and storage, greatly reducing the data bandwidth when loading data and reducing the hardware computing efficiency. Summary of the Invention
[0005] At least one embodiment of the present disclosure provides a data loading method for loading target data into a buffer. The data loading method includes: obtaining a first coordinate value and a first size, where the target data includes a plurality of data parts, and the coordinate ranges of each data part in the channel number dimension of the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The sum of the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; obtaining a coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the plurality of dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first size, and the coordinate information table, determining a plurality of requests for loading the target data; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the target data into the buffer; where, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, requests are divided in the first dimension, the sub-data loaded by each request is either all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory. The at least one dimension includes the first dimension. When loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0006] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the first size not being equal to the shape size of the original tensor in the channel number dimension, and / or the first coordinate value not being equal to the second coordinate value of the starting coordinate of the original tensor in the channel number dimension, it is determined that continuous loading cannot be performed on the target data in the channel number dimension.
[0007] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the first coordinate value, the first size, and the coordinate information table, determining a plurality of requests for loading the target data includes: based on the first coordinate value, the first size, and the coordinate information table, determining the first request sent among the plurality of requests and the initial state of the first request entering the state machine; based on the initial state, combining the coordinate information table, and using the state machine to determine each request among the plurality of requests except the first request.
[0008] For example, in the data loading method provided by at least one embodiment of the present disclosure, each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the length of the data loaded by the request. Determining the first request sent among the multiple requests and the initial state of the first request entering the state machine based on the first coordinate value, the first dimension, and the coordinate information table includes: determining the initial state of the first request entering the state machine based on the first coordinate value; taking the first coordinate value as the coordinate value of the request initial coordinate corresponding to the first request in the dimension of the number of channels; determining the coordinate values of the request initial coordinate corresponding to the first request in the at least one dimension from the coordinate information table; determining the data read address of the first request according to the shape and size of the original tensor and the request initial coordinate corresponding to the first request; determining the starting address of writing the target data in the buffer as the data write address of the first request; and determining the length of the data loaded by the first request according to the first coordinate value and the first dimension.
[0009] For example, in the data loading method provided by at least one embodiment of the present disclosure, each dimension in the at least one dimension has a corresponding coordinate information table. Determining the coordinate values of the request initial coordinate corresponding to the first request in the at least one dimension from the coordinate information table includes: for each coordinate information table corresponding to each dimension in the at least one dimension, selecting a first starting coordinate from the coordinate information table corresponding to the dimension, and taking the first starting coordinate as the coordinate value of the request initial coordinate corresponding to the first request in the dimension.
[0010] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the initial state, in combination with the coordinate information table, using the state machine to determine each request other than the first request among the multiple requests includes: based on the initial state, in combination with the coordinate information table, using the state machine to determine the request initial coordinates corresponding to each request and the length of the data loaded by each request; determining the data read address of each request according to the shape and size of the original tensor and the request initial coordinates corresponding to each request; and determining the data write address of each request according to the length of the data loaded by each request.
[0011] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the initial state, in combination with the coordinate information table, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: based on the initial state, in combination with the relationship between the first dimension, the first coordinate value, the second coordinate value, and the third coordinate value, determining the state transition of the state machine, where the state transition includes a transition from the current state where the current request is located to the next state entered by the next request, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the channel number dimension, and the difference between the second coordinate value and the third coordinate value is the shape size of the original tensor in the channel number dimension; according to the state transition of the state machine, determining the coordinate value of the request initial coordinate corresponding to the next request in the channel number dimension and the data length loaded by the current request; and determining the coordinate values of the request initial coordinate corresponding to the next request in the at least one dimension according to the coordinate information table.
[0012] For example, in the data loading method provided by at least one embodiment of the present disclosure, each dimension in the at least one dimension has a corresponding coordinate information table, and determining the coordinate values of the request initial coordinate corresponding to the next request in the at least one dimension according to the coordinate information table includes: for any one of the at least one dimension, in response to the coordinate value of the request initial coordinate corresponding to the current request in the second dimension being the last starting coordinate in the coordinate information table corresponding to the second dimension, updating the coordinate value of the request initial coordinate corresponding to the next request in the any one dimension from the current starting coordinate to the next starting coordinate, where the next starting coordinate is the next starting coordinate adjacent to the current starting coordinate in the coordinate information table corresponding to the any one dimension and in the preset order, and the second dimension is adjacent to the any one dimension and lower than the any one dimension; where, when the coordinate values of the request initial coordinates corresponding to the respective requests in the any one dimension are updated, they are sequentially updated in the preset order according to the multiple starting coordinates included in the coordinate information table corresponding to the any one dimension.
[0013] For example, in the data loading method provided by at least one embodiment of the present disclosure, obtaining the coordinate information table includes: receiving a mode parameter, where the mode parameter is used to indicate loading the target data into the buffer area; receiving the memory read address of the coordinate information table; and reading the coordinate information table from the memory according to the memory read address of the coordinate information table.
[0014] For example, in the data loading method provided by at least one embodiment of the present disclosure, each dimension in the at least one dimension has a corresponding coordinate information table. In the storage unit indicated by the memory read address of the coordinate information table corresponding to each dimension, N starting coordinates are stored, where N is a positive integer greater than 1. Reading the coordinate information table from the memory according to the memory read address of the coordinate information table includes: for any one of the at least one dimension, in response to the shape size of the target data in the any one dimension being greater than N: reading the first storage unit indicated by the memory read address of the coordinate information table corresponding to the any one dimension to obtain the N starting coordinates stored in the first storage unit; adding a preset value to the memory read address of the coordinate information table corresponding to the any one dimension, and reading the second storage unit indicated by the addition result to obtain the N data in the second storage unit indicated by the addition result; in response to the shape size M of the target data in the any one dimension being less than 2×N, selecting the lower M - N data from the N data as the starting coordinates for generating a request; in response to M being greater than or equal to 2×N, using the N data as the N starting coordinates for subsequent request generation.
[0015] For example, in the data loading method provided by at least one embodiment of the present disclosure, the target data is used for data processing of an expert mixture model. The expert mixture model includes a routing module and a plurality of expert networks. The coordinate information table is obtained through the routing module. Each dimension in the at least one dimension has a corresponding coordinate information table. The coordinate information table corresponding to each dimension is used to indicate the data that needs to be input to the target expert network for processing by the target expert network on the dimension; each expert network is trained to process specific tasks and data features.
[0016] For example, in the data loading method provided by at least one embodiment of the present disclosure, sequentially sending the plurality of requests and sequentially writing the sub - data returned by each request into the buffer area to load the target data into the buffer area includes: for any one request, in response to all the sub - data loaded by the any one request being located in the memory, sending the any one request to the memory; in response to all the sub - data loaded by the any one request not being located in the memory, converting the any one request into writing a plurality of predetermined values into the buffer area, where the number of the plurality of predetermined values is determined by the data length to be loaded specified by the any one request.
[0017] At least one embodiment of the present disclosure provides a data loading method for loading target data into a buffer. The data loading method includes: receiving a data loading instruction indicating to execute loading of the target data into the buffer, where the data loading instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first size. The target data includes multiple data parts, and the coordinate ranges of each data part in the channel number dimension in the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The sum of the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; after parsing the data loading instruction, use an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: reading the coordinate information table from the memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first size, and the coordinate information table, determine multiple requests for loading the target data; sequentially send the multiple requests, and sequentially write the sub-data returned by each request into the buffer to load the target data into the buffer; where, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, divide the requests in the first dimension, the sub-data loaded by each request is either all located in the memory or none of them is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory, the at least one dimension includes the first dimension, and when loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0018] For example, in the data storage method provided by at least one embodiment of the present disclosure, which is used to write target data in a buffer into memory, the data storage method includes: obtaining a first coordinate value and a first size, where the target data includes multiple data parts, and the coordinate ranges of each data part in the channel number dimension of the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate, and the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; obtaining a coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first size, and the coordinate information table, determining multiple requests for storing the target data; sequentially sending at least some of the multiple requests, and sequentially writing the data to be written indicated by each request into the memory to store the target data in the memory; where, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when storing the target data, continuous storage cannot be performed in the channel number dimension, partitioning requests in the first dimension, and each sub-data to be written by each request belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the sub-data to be written by the request belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory, and the at least one dimension includes the first dimension, and when loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0019] For example, in the data storage method provided by at least one embodiment of the present disclosure, the request initial coordinate corresponding to each request is used to determine the data storage address of the sub-data to be stored by the request in the memory. Sequentially sending at least some of the multiple requests and sequentially writing the data to be written indicated by each request into the memory to store the target data includes: for any one request, in response to the first request coordinate value of the request initial coordinate corresponding to the any one request in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determining not to send the any one request; in response to the first request coordinate value being greater than or equal to the second coordinate value and the first request coordinate value being less than the third coordinate value, determining to send the any one request; where the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0020] At least one embodiment of the present disclosure provides a processor, including an instruction parsing unit and an execution unit. Wherein, the instruction parsing unit is configured to receive and parse a data loading instruction, and the data loading instruction is used to load target data into a buffer area. The data loading instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first size. The target data includes a plurality of data parts, and the coordinate ranges of each data part in the channel number dimension in the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension. After the instruction parsing unit parses the data loading instruction, the execution unit executes the data loading instruction. When the execution unit executes the data loading instruction, the following operations are included: reading the coordinate information table from the memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the plurality of dimensions included in the coordinate system except the channel number dimension; determining a plurality of requests for loading the target data based on the first coordinate value, the first size, and the coordinate information table; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the target data into the buffer area; wherein, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, requests are divided in the first dimension, the sub-data loaded by each request are all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request are from the original tensor and are continuously stored in the memory. The at least one dimension includes the first dimension. During data loading, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0021] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data loading method according to at least one embodiment of the present disclosure, or the data storage method according to at least one embodiment of the present disclosure.
[0022] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data loading method according to at least one embodiment of the present disclosure or the data storage method according to at least one embodiment of the present disclosure is implemented.
[0023] In the data loading method provided by at least one embodiment of the present disclosure, the starting coordinates of the data to be loaded in other dimensions except the number of channels are obtained through the coordinate information table. Therefore, the target data to be loaded can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data loaded into a certain expert model. For example, as described above, only a certain token or some tokens can be loaded into the buffer area of the hardware computing unit deploying the corresponding expert model through the coordinate information table, so that more flexible data loading can be realized according to the coordinate information table.
[0024] In addition, the data loading method provided by at least one embodiment of the present disclosure divides requests according to whether tensors can be continuously loaded in the number-of-channels dimension. Each request is used to load sub-data that are all located in memory or all not located in memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, and further improving the hardware utilization rate of the computing unit and the hardware performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0026] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU);
[0027] Figure 2 It is a schematic structure of a tensor;
[0028] Figure 3 It is a schematic structural diagram of an MOE layer;
[0029] Figure 4 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;
[0030] Figure 5A It is a schematic diagram of the target data provided by an embodiment of the present disclosure;
[0031] Figure 5B It is a schematic diagram of the target data provided by another embodiment of the present disclosure;
[0032] Figure 6 Schematic diagram of a state machine provided by an embodiment of the present disclosure;
[0033] Figure 7 Schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure;
[0034] Figure 8 Schematic flowchart of a data loading method provided by at least one embodiment of the present disclosure;
[0035] Figure 9 Schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure;
[0036] Figure 10 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure;
[0037] Figure 11 Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure;
[0038] Figure 12 Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure;
[0039] Figure 13 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0041] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of some known functions and known components are omitted in this disclosure.
[0042] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU).
[0043] As Figure 1 shown, the general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processor can be a Streaming Processor Cluster (SPC), for example, including Figure 1 the shown Streaming Processor Cluster 1,..., Streaming Processor Cluster M, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing is performed between multiple streaming processor clusters through a global cache or global memory.
[0044] As Figure 1 shown, taking Streaming Processor Cluster 1 as an example, 1 streaming processor cluster includes multiple computing units, for example Figure 1 the computing units 1, computing unit 2,..., computing unit N in Figure 1 , where N is a positive integer. Each computing unit (Compute Unit, CU for short) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or calculation cores), and each computing core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as
[0045] As Figure 1 shown, each streaming processor cluster also provides a buffer for caching data of N computing units in the streaming processor cluster.
[0046] In parallel computing, computing tasks are generally executed by multiple threads. Before these threads are executed in a general-purpose graphics processing unit (or called a parallel computing processor), they are divided into multiple thread blocks, and then multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called a thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.
[0047] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the mid-level cache or global cache or global memory (such as Figure 1 the high bandwidth memory, High Bandwidth Memory, abbreviated as HBM in the figure) for read and write operations, etc.
[0048] As Figure 1 shown, general computing operations, such as computing operations on tensors, usually require a large amount of data. These data are usually stored in a memory, such as can be stored in HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored in the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.
[0049] The shape dimensions of a multi-dimensional tensor can be expressed as, for example, [N, D, H, W, C]. The N dimension represents the batch size, that is, the number of data samples captured in one training. The D dimension represents the depth, the H dimension represents the height of the input data, the W dimension represents the width of the input data, and the C dimension represents the number of channels.
[0050] For example, Figure 2 is a schematic structure of a tensor. InFigure 2 In the shown tensor, a1 represents the N - dimensional size and is equal to 1, a2 represents the D - dimensional size and is equal to 1, a3 represents the H - dimensional size and is equal to 5, a4 represents the W - dimensional size and is equal to 4, and a5 represents the C - dimensional size and is equal to 64.
[0051] For example, Figure 2 the pixel elements of the tensor in are represented as 0, 1, 2, 3,... and so on.
[0052] The Transformer model is a classic neural - network - based text - processing model, which has given rise to important models such as BERT (Bidirectional Encoder Representation from Transformers, the bidirectional encoder representation in transformers), GPT (Generative Pre - trained Transformer), and LLM (Large Language Model), which have greatly promoted the development of the NLP (Natural Language Processing) field.
[0053] The MOE structure is a very important structure in the Transformer model. It makes the model performance better by allocating different tokens (also known as token feature vectors) to appropriate experts. The MOE structure, whose full name is Mixture of Experts, is an advanced neural - network architecture that improves the overall model performance by integrating the predictions of multiple network models (such as "expert networks"). Its core idea is to distribute the input data to different expert networks for processing, and then a gating network weights and combines the outputs of the expert networks to generate the final result. The MOE network is widely used in fields such as natural language processing (NLP), computer vision (CV), and recommendation systems. Specifically, for example, Deepseek which provides related services for natural language processing. For example, in a large language model (LLM), the MOE structure can replace the feed - forward network (FFN) layer in the traditional Transformer architecture to improve the efficiency and performance of the model.
[0054] Figure 3 It is a schematic diagram of the structure of an MOE layer.
[0055] As Figure 3 shown, the MOE layer includes multiple expert networks and a routing module.
[0056] The expert network is the core computing unit in the MOE layer. Each expert network is an independent neural network, usually in the form of a feed-forward network (FFN). Each expert network is designed to focus on processing specific tasks or data features. For example, in natural language processing, some expert networks may focus on syntactic analysis, while others focus on semantic understanding.
[0057] For example, the input tensor output by the attention layer includes multiple tokens. The role of the routing module is to determine which expert networks each input token should be sent to for processing. For example, the number of expert networks each token is sent to can be preset, such as 2, 4, 6, 8, etc. Then, the processing results of multiple expert networks for the same token are weighted and normalized to obtain the final processing result of the token.
[0058] As Figure 3 shown, the MOE layer does not use the entire network for each token, but learns a mapping function with low computational cost, which determines which parts of the network (i.e., which expert networks) can most effectively process the given input. At the same time, the routing module of the MOE layer is used to selectively activate specific experts required for a given task, rather than activating the entire neural network for each task.
[0059] Therefore, for a neural network that uses the MOE layer, for example, or a neural network with a similar mechanism, it is not required that each token (element) of an input tensor be written into memory for processing. The information of some tokens is unimportant or insensitive to the model, and these data can be not loaded into memory. For example, assume that for expert network 1, judged by the routing module, according to the tasks or data features processed by expert network 1, the first token can be processed by expert network 1, and expert network 1 is not sensitive to other tokens. Therefore, in fact, only the first token needs to be loaded into expert network 1, rather than loading all tokens into expert network 1.
[0060] For a certain expert network, the tokens input to this expert network for processing are not regular. For example, for each type of input tensor, it is not the tokens at fixed positions that are input to this expert network. In other words, for different input tensors, when loading the data that needs to be loaded into a certain expert network for processing, the loading data positions are variable and need to be determined by the routing module after judging the input tensor.
[0061] Therefore, in this scenario, for example, expert networks are deployed on different hardware computing units (such as Figure 1It is processed in the streaming processor cluster), so different expert networks may have their own independent buffer areas for caching data. When loading or storing data, if all data is loaded into the buffer area of the hardware computing unit where a certain expert network is deployed or stored in the memory, some of this data may be insensitive to this expert network, which will bring a lot of invalid data loading and storage, greatly reducing the data bandwidth when loading data and reducing the hardware computing efficiency.
[0062] At least one embodiment of the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium.
[0063] In at least one embodiment, the data loading method is used to load target data into a buffer area. The data loading method includes: obtaining a first coordinate value and a first size, where the target data includes multiple data parts, and the coordinate ranges of each data part in the channel number dimension in the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The sum of the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; obtaining a coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first size, and the coordinate information table, determining multiple requests for loading the target data; sequentially sending the multiple requests, and sequentially writing the sub-data returned by each request into the buffer area to load the target data into the buffer area; where, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, partitioning the requests in the first dimension, the sub-data loaded by each request is all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory, the at least one dimension includes the first dimension, when loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0064] In the data loading method provided by at least one embodiment of the present disclosure, the starting coordinates of the data to be loaded in other dimensions except the number of channels are obtained through the coordinate information table. Therefore, the target data to be loaded can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data loaded into a certain expert model. For example, as described above, only a certain token or some tokens can be loaded into the cache area of the hardware computing unit where the corresponding expert model is deployed through the coordinate information table, so that more flexible data loading can be achieved according to the coordinate information table.
[0065] In addition, the data loading method provided by at least one embodiment of the present disclosure divides requests according to whether the tensor can be continuously loaded in the number of channels dimension. Each request is used to load sub-data that is either all located in the memory or all not located in the memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.
[0066] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0067] Figure 4 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure.
[0068] As Figure 4 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S10 - S40.
[0069] For example, the data loading method provided by at least one embodiment of the present disclosure is used to load target data into the cache area, and the cache area is, for example, the cache area in the streaming processor cluster.
[0070] In step S10, a first coordinate value and a first size are obtained.
[0071] For example, the target data includes multiple data parts. The coordinate ranges of each data part in the number of channels dimension in the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The first coordinate value is the coordinate value of the starting coordinate of each data part in the number of channels dimension, and the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the shape size of the target data in the number of channels dimension, that is, the first size copy_c.
[0072] For example, the coordinate system determined by the original tensor includes multiple dimensions, which may include an N dimension (batch dimension), a D dimension (depth dimension), an H dimension (height dimension), a W dimension (width dimension), and a C dimension (number of channels dimension). According to actual needs, for example, the multiple dimensions may include some of the N dimension, D dimension, H dimension, W dimension, and C dimension. For example, in some embodiments, the multiple dimensions may include the C dimension, W dimension, and H dimension. For example, in some other embodiments, the N dimension, D dimension, H dimension, and W dimension may be unfolded into one dimension. At this time, the multiple dimensions include two dimensions, namely the unfolded one dimension and the C dimension. Hereinafter, the unfolded one dimension may be used as the first dimension in the following text and referred to for subsequent data loading.
[0073] For example, the original tensor is stored in memory. The target data may include some data in the original tensor, or the target data may include some data in the original tensor and data outside the original tensor that is not stored in memory.
[0074] Figure 5A Schematic diagram of target data provided by an embodiment of the present disclosure.
[0075] Figure 5A In, each solid small cube represents an element, and the tensor composed of multiple solid small cubes is the original tensor. Figure 5A What is shown is a batch and a tensor at a certain depth in a batch. However, the original tensor may have multiple batches, and each batch may have multiple tensors in the depth dimension. Its structure is the same as that of Figure 5A shown, and will not be described repeatedly here.
[0076] Assume Figure 5A The element pointed by the arrow is the origin of the coordinate system of the original tensor. C, W, and H respectively represent three coordinate axes. row is the coordinate value in the H dimension direction, col is the coordinate value in the W dimension direction, and c is the coordinate value in the C dimension direction. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. The channel, width, and height values in the coordinates of other elements increase along the arrow direction. Of course, the original tensor may also include D dimension and N dimension coordinates, which are not shown here.
[0077] For example, the shape and size of the original tensor are represented by b1, b2, b3, b4, b5. Assume that b1 is the N dimension size, b2 is the D dimension size, b3 is the H dimension size, b4 is the W dimension size, and b5 is the C dimension size for description. In Figure 5A the example shown, assume b1 = 1, b2 = 1, b3 = 4, b4 = 8, b5 = 8.
[0078] For example, in Figure 5AIn the example, the cubes with a white frame on a black background form the target data, which includes some elements in the original tensor.
[0079] For example, the coordinate range of the target data in the number of channels dimension is from the first coordinate value to the first channel coordinate, and the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the shape size of the target data in the number of channels dimension. For example, in Figure 5A the example, the first coordinate value is 0, the first channel coordinate is 4, and the shape size of the target data in the number of channels dimension is 5.
[0080] As Figure 5A shown, the target data is composed of multiple discrete data parts, and the coordinate range of each data part in the C dimension is the same. As Figure 5A shown, for example, the coordinate range of each data part in the C dimension is from 0 to 4, and each data part can be composed of black cubes arranged in a column.
[0081] The starting coordinates of each data part may be different in the W dimension, H dimension, D dimension, and N dimension. For example, in the present disclosure, the starting coordinate may refer to the coordinate of the element with the smallest coordinate value in the C dimension in each data part. The first coordinate value is the coordinate value of the starting coordinate of each data part in the C dimension, and the first channel coordinate is the coordinate value of the element with the largest coordinate value in the C dimension in each data part.
[0082] For example, in Figure 5A the example, assuming that the N dimension size and the D dimension size are both equal to 1, for data part 0, its starting coordinate may be: c = 0, row = 3, col = 0. For data part 1, its starting coordinate may be: c = 0, row = 3, col = 7. And so on for other data parts.
[0083] For example, in some other embodiments, the target data includes at least some elements of the original tensor and tensor elements not stored in memory. For example, in the padding mode, in addition to including at least some content of the original tensor, the target data may also include multiple elements with a predetermined value (such as 0) added to the edge of the original tensor.
[0084] Figure 5B This is a schematic diagram of the target data provided by another embodiment of the present disclosure.
[0085] Figure 5B In, the tensor composed of multiple solid-line cubes is the original tensor, Figure 5B in, the element at the upper left corner of the original tensor is the origin of the coordinate system determined based on the original tensor, and the coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0 and col = 0. C, W, and H respectively represent three coordinate axes, and the specific meanings are the same asFigure 5A The same is not elaborated here.
[0086] For example, in Figure 5B the example of, the cubes with black backgrounds and white frames form the target data, which includes some elements in the original tensor. In addition, the target data also includes small dotted cubes outside the edge of the original tensor, and the small dotted cubes are obtained by filling operations, for example, and their values are 0, for example. Suppose the N - dimension size and the D - dimension size are both equal to 1. Figure 5B In the example of, the starting coordinates of data part 2 are c = 0, row = 0, col = - 1.
[0087] For example, referring to Figure 5A or Figure 5B , the first coordinate value of each data part is 0, the first channel coordinate is 4, the first size is 5, and the coordinate ranges of each data part in the C - dimension are the same.
[0088] In step S20, obtain the coordinate information table.
[0089] For example, the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension.
[0090] For example, if the multiple dimensions include the N - dimension, D - dimension, H - dimension, W - dimension, and C - dimension, the at least one dimension includes the N - dimension, D - dimension, H - dimension, and W - dimension. For example, if the multiple dimensions include one dimension expanded from the N - dimension, D - dimension, H - dimension, and W - dimension and the C - dimension, the at least one dimension includes the expanded one dimension.
[0091] For example, in some embodiments, each dimension in the at least one dimension can have a corresponding coordinate information table. For a certain dimension, the starting coordinates of each data part recorded in its corresponding coordinate information table are the coordinate values on this dimension. The order of the coordinate values in the coordinate information tables of each dimension is the same. For example, the first coordinate in each coordinate information refers to the starting coordinate of the first data part, and the first coordinate values in the coordinate information tables of the at least one dimension can jointly form the starting coordinate of the first data part in the at least one dimension.
[0092] For example, in some other embodiments, the coordinate information table can provide all the starting coordinates on the at least one dimension, and different starting coordinates on different dimensions can be marked by indexing or other means, and the present disclosure does not make specific limitations on this.
[0093] For example, the data storage format of the original tensor in at least one embodiment of the present disclosure may be the same as that of the target data. For example, the data storage format of the target data or the original tensor is NDHWC. Taking NDHWC as an example, the C dimension is the lowest dimension during storage / loading, followed by the H dimension, then the W dimension, then the D dimension, and finally the N dimension. Of course, the data storage format of the original tensor in at least one embodiment of the present disclosure may also be different from that of the target data. For example, the data storage format of the target data is NDHWC, and the data storage format of the original tensor is N(C / x)DHW(xC), or vice versa. The present disclosure does not make specific limitations on this.
[0094] For example, a coordinate system is determined with a certain element in the original tensor as the origin of the coordinate system. For example, referring to Figures 5A - 5B the embodiment of, the upper left vertex in the original tensor is used as the origin of the coordinate system. Of course, the present disclosure is not limited thereto.
[0095] For example, in some embodiments, step S20 may include: receiving a mode parameter, where the mode parameter is used to indicate loading the target data into the buffer; in response to the mode instruction indicating loading the target data into the buffer, receiving the memory read address of the coordinate information table; and reading the coordinate information table from the memory according to the memory read address of the coordinate information table.
[0096] For example, a data loading instruction may be received, and the data loading instruction includes a mode parameter, the memory read address of the coordinate information table, a first coordinate value, and a first size.
[0097] For example, when the mode parameter is a first value, it indicates data loading according to a conventional tensor, and the coordinate information table is not used during loading. When the mode parameter is a second value, it indicates that the data to be loaded is the target data, which includes discrete multiple data parts, and the starting coordinates need to be obtained from the coordinate information table.
[0098] The coordinate information table is read from the memory according to the memory read address of the coordinate information table. For example, if each dimension in at least one dimension has a corresponding coordinate information table, in the storage unit indicated by the memory read address table_addr of the coordinate information table corresponding to each dimension, N starting coordinates are stored, where N is a positive integer greater than 1. For example, the memory includes multiple storage units, each storage unit has a corresponding memory read address. Assuming that the bit width of each starting coordinate is 32B and the storage space size of one storage unit is 512B, then N = 16.
[0099] For example, reading the coordinate information table from the memory according to the memory read address of the coordinate information table may include: for any one of at least one dimension, in response to the shape size of the target data in any one dimension being greater than N: reading the first storage unit indicated by the memory read address of the coordinate information table corresponding to any one dimension to obtain N starting coordinates stored in the first storage unit; adding a preset value to the memory read address of the coordinate information table corresponding to any one dimension, and reading the second storage unit indicated by the accumulated result to obtain N data in the second storage unit; in response to the shape size M of the target data in any one dimension being less than 2×N, selecting the M - N data with the lowest bits from the N data as the starting coordinates for generating a request; in response to M being greater than or equal to 2×N, using the N data as the N starting coordinates for subsequent request generation.
[0100] For example, in one embodiment, N = 16. Assume that for the W dimension, the memory read address table_addr_w of the coordinate information table corresponding to the W dimension is carried in the load instruction. Read the 16 coordinate values stored in the first storage unit indicated by table_addr_w from the memory, and use these 16 coordinate values to perform subsequent request generation.
[0101] For example, assume that the size of the target data in the W dimension is greater than 16, for example, it is 17. For example, after 16 coordinate values have all been used to generate requests through step S40, or after reading the 16 coordinate values in the first storage unit indicated by table_addr_w, table_addr_w can be incremented by 1, and 16 data stored in the second storage unit indicated by table_addr_w + 1 are read from the memory. This is because the memory stores data in the smallest granularity of storage units, so each read or write operation needs to be performed in units of storage units. Then, take the 1 data with the lowest bit among the 16 data as the 17th starting coordinate. For example, read the lower 32 bits of the data returned by the address table_addr_w + 1 in the memory as the 17th starting coordinate to continue generating requests using step S40.
[0102] For example, if the size of the target data in the W dimension is equal to 32, then use the 16 data stored in the storage unit indicated by the address table_addr_w + 1 as the 16 starting coordinates of the W dimension for subsequent request generation.
[0103] For example, if the size of the target data in the W dimension is greater than 32, continue to read the data in the storage unit at the address table_addr_w + 2, and obtain the starting coordinates with reference to the above process, which will not be elaborated here.
[0104] For example, the target data is used for data processing of the mixture-of-experts model, which includes a routing module and multiple expert networks. For the mixture-of-experts model, reference can be made to the foregoing introduction about MOE.
[0105] For example, the coordinate information table is obtained through the routing module. Each dimension in at least one dimension has a corresponding coordinate information table. The coordinate information table corresponding to each dimension is used to indicate the data that needs to be input into the target expert network for processing by the target expert network in that dimension. Each expert network is trained to process specific tasks and data features.
[0106] For example, taking expert network 1 as an example, the coordinate information table is used to indicate the coordinate values in other dimensions except the C dimension of the starting coordinates of each data part in the original tensor that needs to be input into expert network 1, so as to use the coordinate information table to take these data parts as target data and load them into the buffer area of the hardware computing unit where expert network 1 is deployed.
[0107] In step S30, based on the first coordinate value, the first size, and the coordinate information table, multiple requests for loading the target data are determined.
[0108] In step S40, multiple requests are sequentially sent, and the sub-data returned by each request is written into the buffer area in order to load the target data into the buffer area.
[0109] For example, in response to the size relationship between the target data and the original tensor in the channel number dimension such that the target data cannot be continuously loaded in the channel number dimension during data loading, requests are divided in the first dimension. The sub-data loaded by each request is either all in memory or none of it is in memory. And in response to the sub-data loaded by the request being all in memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in memory.
[0110] At least one dimension includes the first dimension. During data loading, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0111] For example, for the data storage format NDHWC, the first dimension is the width dimension. For example, for the scenario where the N dimension, D dimension, H dimension, and W dimension are expanded into one dimension, the first dimension is one of the expanded dimensions.
[0112] The object to be loaded by the present disclosure, that is, the target data, is split into multiple different requests according to a certain rule and sent sequentially, and the returned data is written into the buffer area in order, thereby efficiently loading the target data into the memory. A principle during splitting is that the sub-data loaded by each request is either all in the memory (the sub-data loaded by the request belongs to the data range of the original tensor) or all not in the memory (the sub-data loaded by the request does not belong to the data range of the original tensor). This is because when the sub-data of the request is in the memory, a request is sent to the memory, and when the sub-data of the request is not in the memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by the hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.
[0113] In addition, during splitting, if the size relationship between the target data and the original tensor in the channel number dimension makes it impossible to continuously load in the channel number dimension but possible to continuously load in each dimension lower than the channel number dimension when loading the target data, then further division of the requests is performed in the first dimension, so that the requests can be split more reasonably, the continuously stored data in the memory can be retained as much as possible, the number of requests can be reduced, the target data can be loaded efficiently, the bandwidth when loading or storing data from the memory can be greatly increased, the performance can be improved, and the efficiency of the hardware computing unit can be improved.
[0114] For example, if the data storage format of the target data is NDHWC, the coordinates of the sub-data loaded by each request satisfy the following conditions: the coordinates of each element of the sub-data are different in the channel number dimension, but the same in the width dimension and each dimension higher than the width dimension.
[0115] For example, in response to the shape size copy_c of the target data in the channel number dimension not being equal to the shape size tensor_c of the original tensor in the channel number dimension, and / or the first coordinate value c_coord_b of the starting coordinate of the target data in the channel number dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the channel number dimension, it is determined that the target data cannot be continuously loaded in the channel number dimension.
[0116] For example, assuming the second coordinate value is 0, it is determined that it is not continuous in the C dimension when any of the following conditions is satisfied:
[0117] (1) The starting coordinate c_coord_b of the target data in the C dimension is not equal to 0
[0118] (2) The size copy_c of the target data in the C dimension is greater than the size tensor_c of the original tensor in the C dimension
[0119] (3) The size copy_c of the target data in the C dimension is less than the size tensor_c of the original tensor in the C dimension
[0120] For example, the sub-data loaded by each request has different coordinates in the C dimension, but the same coordinates in the W dimension, H dimension, D dimension, and N dimension. At this time, it can be understood that without considering the depth direction, the original tensor is unfolded into a one-dimensional vector composed of multiple pixels (pixel) in the order of WHDN. The sub-data loaded by each request is the data within a pixel, that is, the requests are divided in the W dimension, and the next request can be to load the data in another pixel, that is, the requests are divided in a pixel-skipping manner.
[0121] The determination method of multiple requests for loading target data will be specifically described below.
[0122] For example, in some embodiments, step S30 may include: determining the first request sent among multiple requests and the initial state when the first request enters the state machine based on the first coordinate value, the first size, and the coordinate information table; based on the initial state, combining the coordinate information table, and using the state machine to determine each request other than the first request among multiple requests.
[0123] Each request includes three parameters: a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the length of the loaded data.
[0124] The data read address of each request is obtained by accumulating the sizes of all previous requests.
[0125] For example, taking the data storage format as NDHWC as an example, the calculation formulas for the data read address Addr_1 and the data write address Addr_2 of the nth request are as follows:
[0126] Addr_1 = u_addr_base +
[0127] ((((n_coord × tensor_d + d_coord) × tensor_h + h_coord) × tensor_w + (Formula 1)
[0128] w_coord) × tensor_c + c_coord)
[0129] Addr_2 = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 (Formula 2)
[0130] Among them, u_addr_base represents the storage address in memory of the element at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in 5 dimensions, b_addr_base represents the starting address for writing the target data in the buffer, and req_size_1, req_size_2, ..., req_size_n-1 represent the data lengths loaded by each of the first n-1 sent requests.
[0131] For example, taking the data storage format as N(C / x)DHW(xC) as an example, the calculation formulas for the data read address Addr_3 and the data write address Addr_4 of the nth request are as follows:
[0132] Addr_3 = u_addr_base +
[0133] n_coord × tensor_c × tensor_d × tensor_h × tensor_w +
[0134] c_coord × tensor_d × tensor_h × tensor_w +
[0135] d_coord × tensor_h × tensor_w × xC + (Formula 3)
[0136] h_coord × tensor_w × xC +
[0137] w_coord × xC
[0138] Addr_4 = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 (Formula 4)
[0139] Among them, xC represents the size occupied by x channels. The definitions of other parameters are the same as those in Formula 1 and Formula 2, and the repeated parts will not be elaborated here.
[0140] For example, in some embodiments, determining the first request sent among multiple requests and the initial state when the first request enters the state machine based on the first coordinate value, the first dimension, and the coordinate information table may include: determining the initial state when the first request enters the state machine based on the first coordinate value and the first dimension; using the first coordinate value as the coordinate value of the initial request coordinate corresponding to the first request in the dimension of the number of channels; determining the coordinate values of the initial request coordinate corresponding to the first request in at least one dimension from the coordinate information table; determining the data read address of the first request according to the shape and dimension of the original tensor and the initial request coordinate corresponding to the first request; determining the starting address where the target data is written in the buffer as the data write address of the first request; and determining the data length loaded by the first request according to the first coordinate value and the first dimension.
[0141] For the first request sent, use the first coordinate value as the coordinate value of the initial request coordinate corresponding to the first request in the dimension of the number of channels. Determine the coordinate values of the initial request coordinate corresponding to the first request in the at least one dimension from the coordinate information table.
[0142] For example, each dimension in the at least one dimension has a corresponding coordinate information table. Determining the coordinate values of the initial request coordinate corresponding to the first request in the at least one dimension from the coordinate information table may include: for each coordinate information table corresponding to each dimension in the at least one dimension, selecting the first starting coordinate from the coordinate information table corresponding to the dimension as the coordinate value of the initial request coordinate corresponding to the first request in the dimension.
[0143] For example, assume that the at least one dimension includes the W dimension and the H dimension, and the W dimension and the H dimension have their respective corresponding coordinate information tables. Read the first starting coordinate from the coordinate information table corresponding to the W dimension as the coordinate value of the initial request coordinate corresponding to the first request in the W dimension, and read the first starting coordinate from the coordinate information table corresponding to the H dimension as the coordinate value of the initial request coordinate corresponding to the first request in the H dimension. For example, the 32 bits (bits) located at the lowest bit in the storage unit indicated by the memory read address table_addr_w of the coordinate information table corresponding to the W dimension can be used as the coordinate value of the initial request coordinate corresponding to the first request in the W dimension. Here, assume that the bit width of a coordinate value is 32 bits.
[0144] Of course, for other scenarios, other methods such as indexing can also be used to mark the first starting coordinate.
[0145] After obtaining the initial request coordinate corresponding to the first request, the data read address of the first request can be calculated with reference to the above formula. Use the starting address where the target data is written in the buffer as the data write address of the first request.
[0146] The length of the data loaded by the first request is determined according to the first coordinate value and the first dimension. For example, if the sum of the first coordinate value c_coord_b and the first dimension copy_c is less than the second coordinate value (i.e., the coordinate value of the starting coordinate of the original tensor in the channel number dimension, for example, it is 0), the length of the data loaded by the first request is the first dimension copy_c. For example, if the sum of the first coordinate value c_coord_b and the first dimension copy_c is greater than or equal to the second coordinate value, the length of the data loaded by the first request is the absolute value of the first coordinate value c_coord_b.
[0147] Considering the advantages that the state machine has a clear logical structure, is easy to maintain and expand, and is particularly suitable for handling logical scenarios with multiple conditions and multiple branches, in this disclosure, the state machine is adopted to automatically update the request initial coordinates and the loaded data length corresponding to each request, avoiding complex conditional nesting, with clear logic, easy to maintain, and strong scalability.
[0148] Figure 6 It is a schematic diagram of the state machine provided by an embodiment of this disclosure.
[0149] For example, the state machine includes a first state s0, a second state s1, and a third state s2, and the initial state of entering the state machine is determined by the first coordinate value.
[0150] For example, based on the first coordinate value, determining the initial state of the first request entering the state machine may include: in response to the first coordinate value being less than the second coordinate value of the starting coordinate of the original tensor in the channel number dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the shape dimension of the original tensor in the channel number dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.
[0151] Reference Figure 6 , the arrow in the C dimension represents the data in the channel number dimension. The data with the coordinate in the C dimension less than the second coordinate value and greater than the third coordinate value is not in the memory ( Figure 6 the black part in), and the data with the coordinate in the C dimension between the second coordinate value and the third coordinate value ( Figure 6 the white part in) is stored in the memory, which is the actual size of the original tensor data in the channel number dimension.
[0152] The state will only switch to itself and adjacent states. For example Figure 6 in, the first state s0 can jump to the first state s0 or the second state s1, the second state s1 can jump to the first state s0, the second state s1, and the third state s2, and the third state s2 can jump to the first state s0, the second state s1, and the third state s2.
[0153] Figure 6 Six cases (① to ⑥) among them mark the state transition changes that the target data has to go through for different sizes in the dimension of the number of channels.
[0154] For example, Figure 6 In case ①, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the first size is less than the second coordinate value. The state machine always loops and jumps in the first state s0 on the left.
[0155] For example, Figure 6 In case ②, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the first size is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1.
[0156] For example, Figure 6 In case ③, the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the first size is less than the third coordinate value. The state machine always loops and jumps in the second state s1.
[0157] For example, Figure 6 In case ④, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the first size is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1, and the third state s2.
[0158] For example, Figure 6 In case ⑤, the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the first size is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.
[0159] For example, Figure 6 In case ⑥, the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.
[0160] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and combine with the trigger condition to determine the state that the second sent request enters. Then, based on the state that the second sent request enters, determine the data length loaded by the second sent request, and combine with the trigger condition to determine the state that the third sent request enters and its corresponding request initial coordinates, and so on.
[0161] For example, in some embodiments, based on the initial state, in combination with the coordinate information table, using a state machine to determine each request among multiple requests except the first request may include: based on the initial state, in combination with the coordinate information table, using the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request; determining the data read address for each request according to the shape and size of the original tensor and the request initial coordinates corresponding to each request; and determining the data write address for each request according to the data length loaded by each request.
[0162] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state.
[0163] After determining the request initial coordinates corresponding to the current request, it can be determined whether the sub-data loaded by the current request is located in the memory according to the request initial coordinates.
[0164] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the channel number dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the sub-data to be loaded by the current request is located in the memory. At this time, the data read address and data write address of the current request can be determined with reference to the above formulas 1-4, and then the current request is sent to the memory to load the corresponding sub-data into the buffer.
[0165] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the channel number dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the sub-data to be loaded by the current request is not located in the memory. At this time, in response to all the sub-data to be loaded by the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values to the buffer, where the number of the multiple predetermined values is determined by the data length specified by the current request. For example, the predetermined value is 0.
[0166] For example, based on the initial state, in combination with the coordinate information table, using a state machine to determine the request initial coordinates corresponding to each request and the data lengths loaded by each request may include: based on the initial state, in combination with the relationship between the first dimension, the first coordinate value, the second coordinate value, and the third coordinate value, determining the state transition of the state machine, where the state transition includes the next state that the next request enters by jumping from the current state where the current request is located. The second coordinate value is the coordinate value of the starting coordinate of the original tensor in the dimension of the number of channels, and the difference between the second coordinate value and the third coordinate value is the shape size of the original tensor in the dimension of the number of channels; according to the state transition of the state machine, determining the coordinate value of the request initial coordinate corresponding to the next request in the dimension of the number of channels and the data length loaded by the current request; determining the coordinate value of the request initial coordinate corresponding to the next request in at least one dimension according to the coordinate information table.
[0167] For example, each dimension in at least one dimension has a corresponding coordinate information table. Determining the coordinate value of the request initial coordinate corresponding to the next request in at least one dimension according to the coordinate information table may include: for any one of the at least one dimension, in response to the coordinate value of the request initial coordinate corresponding to the current request in the second dimension being the last starting coordinate in the coordinate information table corresponding to the second dimension, the coordinate value of the request initial coordinate corresponding to the next request in any one dimension is updated from the current starting coordinate to the next starting coordinate, where the next starting coordinate is the next starting coordinate adjacent to the current starting coordinate in the coordinate information table corresponding to any one dimension and arranged in a preset order. The second dimension is adjacent to and lower than any one dimension.
[0168] For example, when the coordinate value of the request initial coordinate corresponding to each request in any one dimension is updated, it is sequentially updated according to the multiple starting coordinates included in the coordinate information table corresponding to any one dimension in a preset order.
[0169] For example, here the preset order refers to the arrangement order of the multiple starting coordinates in the coordinate information table. The orders of the coordinate information tables corresponding to different dimensions are the same, so that the starting coordinates at the same position can be combined to form a coordinate value of the request initial coordinate in the at least one dimension.
[0170] For any dimension other than the dimension of the number of channels, when an update occurs, it is sequentially updated according to the multiple starting coordinates included in the coordinate information table corresponding to that dimension. The update method includes updating in the order of the coordinate information table, returning the first starting coordinate in the coordinate information table, or remaining unchanged.
[0171] For example, for Dimension 1 among the multiple dimensions included in the target data, the coordinate value of the initial request coordinate corresponding to each request in Dimension 1 changes sequentially according to the coordinate information table table_1 corresponding to Dimension 1. For example, the coordinate value of the initial request coordinate corresponding to the first request in Dimension 1 is the first starting coordinate in the coordinate information table table_1, and the coordinate value of the initial request coordinate corresponding to the second request in Dimension 1 is the second starting coordinate in the coordinate information table table_1, and so on.
[0172] When the coordinate value of the initial request coordinate corresponding to the current request in Dimension 1 is the last starting coordinate in the coordinate information table table_1, the coordinate value of the initial request coordinate corresponding to the next request in Dimension 1 returns to the first starting coordinate in the coordinate information table table_1, and the coordinate value of the initial request coordinate corresponding to the next request in Dimension 2 is updated to the next coordinate value in the coordinate information table table_2 corresponding to Dimension 2. Dimension 2 is adjacent to Dimension 1 and higher than Dimension 1. Each dimension is updated in this way.
[0173] In a specific example, taking the second dimension as the W dimension and any one dimension as the H dimension as an example, assuming that the current request is the nth request, the coordinate value of the nth initial request coordinate corresponding to the nth request in the W dimension is the last starting coordinate in the coordinate information table table_w corresponding to the W dimension, that is, all the coordinates of the target data in the W dimension have been traversed. Then, the coordinate value of the (n + 1)th initial request coordinate corresponding to the (n + 1)th request in the W dimension returns to the first starting coordinate w_coord_b in the coordinate information table table_w, and the coordinate value of the (n + 1)th initial request coordinate in the H dimension is updated according to the coordinate information table table_h corresponding to the H dimension, and is updated to the next coordinate value in the coordinate information table table_h that is located after the coordinate value of the nth initial request coordinate in the H dimension.
[0174] The process of using the state machine to determine the initial request coordinates and the loaded data lengths corresponding to each request is specifically described below.
[0175] When the first coordinate value is less than the second coordinate value, enter the first state s0.
[0176] For example, when the current request is the first request, the state corresponding to the current request is the initial state. The state machine outputs the initial request coordinates and the loaded data length corresponding to the current request, and prepares the initial request coordinates for the next state.
[0177] For example, in response to the current request being in the first state and satisfying the first condition, it is determined that the next request of the current request enters the first state. For example, the first condition includes that the sum of the first dimension and the first coordinate value is less than the second coordinate value. Taking the second coordinate value as 0 as an example, the first condition includes that the first dimension is less than the absolute value of the first coordinate value.
[0178] For example, if the current request is in the first state s0 and the first dimension is relatively small, for example, the sum of the size of the target data in the number of channels dimension and the first coordinate value is less than the second coordinate value, the state of the next request remains the first state s0. For example, as in Figure 6 case ① s0->s0 in
[0179] In this case, it is determined that the coordinate value of the initial request coordinates corresponding to the next request in the number of channels dimension is the first coordinate value, and according to the coordinate information table, the coordinate values of the initial request coordinates corresponding to the next request in other dimensions are determined. For example, if the next request is the second request, the coordinate value of the initial request coordinates corresponding to the next request in the W dimension is updated to the second starting coordinate in the coordinate information table table_w, and the coordinate values of the initial request coordinates corresponding to the next request in other dimensions (such as the H dimension, D dimension, N dimension) remain unchanged.
[0180] For example, in this case, the data length req_size loaded by the current request is the first dimension.
[0181] For example, in response to the current request being in the first state and satisfying the second condition, it is determined that the next request of the current request enters the second state s1. For example, the second condition includes that the first dimension is greater than or equal to the absolute value of the first coordinate value. For example, as in Figure 6 case ② and case ④ in
[0182] For example, if the current request is in the first state s0 and the first dimension is relatively large, for example, the sum of the first dimension and the first coordinate value is greater than or equal to the second coordinate value, the state of the next request jumps to the second state s1. For example, assuming the second coordinate value is 0, the second condition includes that the first dimension is greater than or equal to the absolute value of the first coordinate value.
[0183] In this case, it is determined that the coordinate value of the initial request coordinates corresponding to the next request in the number of channels dimension is the second coordinate value, for example 0, and it is determined that the coordinate values of the initial request coordinates corresponding to the next request in other dimensions except the number of channels dimension remain unchanged. Since it is a jump from the first state to the second state and the data in the number of channels dimension has not been fully loaded, it is determined that the coordinate values of the initial request coordinates corresponding to the next request in other dimensions except the number of channels dimension remain unchanged.
[0184] For example, in this case, the data length req_size loaded by the current request is determined based on the first coordinate value and the second coordinate value, for example, it is the difference between the first coordinate value and the second coordinate value.
[0185] For example, in response to the status of the current request being the second status s1 and satisfying the third condition, it is determined that the next request of the current request enters the first status s0. The third condition includes that the first coordinate value is less than the second coordinate value and the sum of the first coordinate value and the first dimension is less than the third coordinate value (for example, the second coordinate value is 0 and the third coordinate value is the dimension tensor_c of the original tensor in the number of channels dimension), for example, as in Figure 6 case ② in, s1 -> s0. Since the sum of the first coordinate value and the first dimension is less than the third coordinate value, it will not enter the third status s2.
[0186] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the number of channels dimension is the first coordinate value, and the coordinate values of the request initial coordinate corresponding to the next request in other dimensions are determined according to the coordinate information table. The specific process can refer to the foregoing content and will not be elaborated here. For example, in this case, the data length req_size loaded by the current request is determined based on the first coordinate value and the first dimension. For example, if the second coordinate value is 0, the data length req_size loaded by the current request is the sum of the first coordinate value and the first dimension.
[0187] For example, when the status of the current request is the second status s1 and satisfies the fourth condition, it is determined that the next request enters the second status s1. The fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the first dimension is less than the third coordinate value, and the difference between the second coordinate value and the third coordinate value is equal to the dimension of the original tensor in the number of channels dimension. For example, when the current request is in the second status, it means that the sub-data of the request are all in the memory, that is, they are all valid data. When the first dimension is relatively small, for example, the sum of the dimension of the target data in the number of channels dimension and the first coordinate value is less than the third coordinate value, the status of the next request is still the second status s1 and will not jump to the third status s2. For example, as in Figure 6 case ③ in.
[0188] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the number of channels dimension is the first coordinate value, and the coordinate values of the request initial coordinate corresponding to the next request in other dimensions are determined according to the coordinate information table. The specific process is as described above and will not be elaborated here.
[0189] For example, in this case, the data length req_size loaded by the current request is the first dimension.
[0190] For example, in response to the status of the current request being the second status s1 and satisfying the fifth condition, it is determined that the next request enters the third status s2. The fifth condition includes that the sum of the first dimension and the first coordinate value is greater than or equal to the third coordinate value. At this time, the first dimension is relatively large, and the status of the next request jumps to the third status s2. For example, such as Figure 6 s1->s2 in case ④ and case ⑤ in
[0191] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the channel number dimension is the third coordinate value, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the channel number dimension remain unchanged.
[0192] For example, in this case, the data length req_size loaded by the current request is determined based on the coordinate value c_coord of the request initial coordinate corresponding to the current request in the channel number dimension and the third coordinate value. For example, the data length req_size loaded by the current request is the difference between the third coordinate value and the coordinate value c_coord.
[0193] For example, in response to the status of the current request being the third status, the status that the next request of the current request enters is determined according to the first coordinate value.
[0194] For example, when the first coordinate value is less than the second coordinate value, it is determined that the next request enters the first status s0, such as Figure 6 s2->s0 in case ④ in Figure 6 When the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, it is determined that the next request enters the second status s1, such as Figure 6 s2->s1 in case ⑤ in
[0195] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the channel number dimension is the first coordinate value, and the coordinate values of the request initial coordinate corresponding to the next request in other dimensions are determined according to the coordinate information table. The specific process is as described above and will not be elaborated here.
[0196] For example, in this case, the data length req_size loaded by the current request is determined based on the coordinate value of the request initial coordinate corresponding to the current request in the channel number dimension and the third coordinate value. For example, the data length req_size loaded by the current request is the difference between the coordinate value of the request initial coordinate corresponding to the current request in the channel number dimension and the third coordinate value.
[0197] In some other embodiments, the state machine can also update the remaining size remain_copy_t in five dimensions, where the remaining size represents the amount of data that has not been loaded in each dimension. When determining the state transition conditions of the state machine, the relationship between the fourth coordinate value of the request initial coordinates corresponding to the current request in the channel number dimension and the remaining size is considered for judgment.
[0198] For example, taking the transition from the first state s0 to the first state s0 as an example, the first condition may include that the sum of the fourth coordinate value c_coord of the request initial coordinates corresponding to the current request in the channel number dimension and the remaining size remain_copy_c of the target data in the channel number dimension is less than the second coordinate value, and the fourth coordinate value c_coord is less than the second coordinate value. The remaining size remain_copy_c is the number of remaining tensor data that has not been requested to be loaded in the channel number dimension of the target data.
[0199] For example, taking the transition from the first state s0 to the second state s1 as an example, the second condition may include that the sum of the fourth coordinate value c_coord of the request initial coordinates corresponding to the current request in the channel number dimension and the remaining size remain_copy_c of the target data in the channel number dimension is greater than or equal to the second coordinate value.
[0200] For example, taking the transition from the second state s1 to the first state s0 as an example, the third condition may include that the sum of the fourth coordinate value c_coord of the request initial coordinates corresponding to the current request in the channel number dimension and the remaining size remain_copy_c of the target data in the channel number dimension is less than or equal to the third coordinate value, and the first coordinate value is less than the second coordinate value.
[0201] For example, taking the transition from the second state s1 to the second state s1 as an example, the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value.
[0202] For example, taking the transition from the second state s1 to the third state s2 as an example, the fifth condition may include that the sum of the request initial coordinates corresponding to the current request and the remaining size remain_copy_c of the target data in the channel number dimension is greater than the third coordinate value.
[0203] For example, for the third state s2, the state into which the next request of the current request enters is determined according to the first coordinate value, and specific reference can be made to the foregoing content.
[0204] Of course, when all the starting coordinates in the coordinate information table have been traversed, the state machine can enter the idle state.
[0205] The state transition conditions of the specific state machine and the updated coordinate content can be changed and set according to actual needs. The logic is similar to the state machine logic described above, and no further examples will be given here.
[0206] For example, step S40 may include: for any request, in response to all the sub-data loaded for any request being located in memory, sending any request to memory; in response to all the sub-data loaded for any request not being located in memory, converting any request into writing a plurality of predetermined values to the buffer area, where the number of the plurality of predetermined values is determined by the data length to be loaded specified by any request.
[0207] If the sub-data is located in memory, load the data from memory to the buffer area; if the sub-data is not located in memory, convert it to the hardware writing a plurality of predetermined values to the buffer area. For example, the predetermined value is 0, and the number of the plurality of predetermined values is determined by the data length to be loaded specified by the request.
[0208] For example, taking the data storage format NDHWC as an example below, the specific data loading process is described. In this example, the first coordinate value is the coordinate value c_coord_b of the starting coordinate of the target data in the C dimension, the second coordinate value is 0, the third coordinate value is the size tensor_c of the original tensor in the C dimension, and the predetermined value is 0.
[0209] For example, if the size copy_c of the target data in the C dimension is less than or greater than the size tensor_c of the original tensor in the C dimension, and / or the first coordinate value c_coord_b is not equal to 0, it is determined at this time that the target data cannot be continuously loaded in the C dimension.
[0210] First, obtain the first coordinate value c_coord_b and the first size copy_c in step S10.
[0211] After that, in step S20, obtain the coordinate information table. Specifically, the coordinate information table can be read from memory, which will not be elaborated here. For example, 4 coordinate information tables can be obtained, namely the coordinate information table table_w corresponding to the W dimension, the coordinate information table table_h corresponding to the H dimension, the coordinate information table table_d corresponding to the D dimension, and the coordinate information table table_n corresponding to the N dimension. Each coordinate information table records the starting coordinates of each data part in the corresponding dimension.
[0212] After that, in step S30, determine a plurality of requests for loading the target data in combination with the above information.
[0213] Referring to the process described above, a state machine is used to determine the request initial coordinates and the length of the loaded data corresponding to each request, and based on this, the data read address and the data write address of each request are determined. For example, in this example, the coordinates of the sub-data loaded by each request in the C dimension are different, but the coordinates in the H dimension, W dimension, D dimension, and N dimension are the same.
[0214] Table 1 shows the update process of the request initial coordinates and the length of the data loaded by the request provided in an embodiment of the present disclosure.
[0215] Table 1
[0216]
[0217] As shown in Table 1, if the first coordinate value c_coord_b is less than 0, it is determined that the first request enters the first state s0. The data read address and the data write address of the first request refer to the foregoing description and will not be elaborated here.
[0218] If the first condition is satisfied, it is determined that the next request enters the first state s0. For the first condition, refer to the foregoing description.
[0219] As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is the first coordinate value c_coord_b.
[0220] The coordinate value w_coord of the request initial coordinates corresponding to the next request in the W dimension is the next coordinate in the coordinate information table table_w corresponding to the W dimension. For example, for the second request, it is the second starting coordinate. If the coordinate value of the request initial coordinates corresponding to the current request in the W dimension is the last starting coordinate in the coordinate information table table_w, the coordinate value w_coord of the request initial coordinates corresponding to the next request returns the first starting coordinate w_coord_b in the coordinate information table table_w.
[0221] The coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is updated to the next coordinate in the coordinate information table table_h corresponding to the H dimension when w_coord returns w_coord_b, and remains unchanged when the W dimension has not been updated to the last starting coordinate in the coordinate information table table_w. If the coordinate value of the request initial coordinates corresponding to the current request in the H dimension is the last starting coordinate in the coordinate information table table_h, the coordinate value h_coord of the request initial coordinates corresponding to the next request returns the first starting coordinate h_coord_b in the coordinate information table table_h.
[0222] When the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is updated to the next coordinate in the coordinate information table table_d corresponding to the D dimension when h_coord returns h_coord_b, it remains unchanged when the last starting coordinate in the coordinate information table table_h in the H dimension has not been updated. If the coordinate value of the request initial coordinate corresponding to the current request in the D dimension is the last starting coordinate in the coordinate information table table_d, the coordinate value of the request initial coordinate corresponding to the next request in the D dimension returns the first starting coordinate d_coord_b in the coordinate information table table_d.
[0223] When d_coord returns d_coord_b, the coordinate value n_coord of the request initial coordinate corresponding to the next request is updated to the next starting coordinate in the coordinate information table corresponding to the N dimension, and remains unchanged in other cases.
[0224] Moreover, at this time, the data length loaded by the first request is the size copy_c of the target data in the C dimension.
[0225] If the second condition is satisfied, it is determined that the next request enters the second state s1. For the second condition, refer to the foregoing description. At this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is 0; the coordinate value w_coord of the request initial coordinate corresponding to the next request, the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension, the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension, and the coordinate value n_coord of the request initial coordinate corresponding to the next request all remain unchanged.
[0226] Moreover, at this time, the data length loaded by the first request is the absolute value of the first coordinate value.
[0227] Since the coordinate value c_coord of the request initial coordinate corresponding to the first request in the C dimension is less than 0, none of the sub-data loaded by the first request is in the memory. In step S40, the first request is converted into an operation of writing copy_c or |c_coord_b| zeros to the buffer area, where |c_coord_b| represents the absolute value of c_coord_b.
[0228] In addition, it should be noted that if it is determined that the coordinate value c_coord of the request initial coordinate corresponding to the request is less than 0, there is no need to calculate the data read address and data write address of the request, reducing the calculation amount.
[0229] As shown in Table 1, when the first coordinate value c_coord_b is greater than or equal to 0 and less than tensor_c, it is determined that the first request enters the second state s1. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0230] If the third condition is satisfied, it is determined that the next request enters the first state s0. For the third condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the first coordinate value c_coord_b. The coordinate updates of the coordinate values w_coord, h_coord, d_coord, and n_coord of the request initial coordinate corresponding to the next request in the W dimension, H dimension, D dimension, and N dimension are the same as those when the first state s0 jumps to the first state s0 and will not be elaborated here.
[0231] Moreover, at this time, the data length loaded by the first request is the sum of the size copy_c of the target data in the C dimension and the first coordinate value c_coord_b.
[0232] If the fourth condition is satisfied, it is determined that the next request enters the second state s1. For the fourth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the first coordinate value c_coord_b; the coordinate updates of the coordinate values w_coord, h_coord, d_coord, and n_coord of the request initial coordinate corresponding to the next request in the W dimension, H dimension, D dimension, and N dimension are the same as those when the first state s0 jumps to the first state s0 and will not be elaborated here.
[0233] Moreover, at this time, the data length loaded by the first request is the size copy_c of the target data in the C dimension.
[0234] If the fifth condition is satisfied, it is determined that the next request enters the third state s2. For the fifth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is the third coordinate value tensor_c; the coordinate values w_coord, h_coord, d_coord, and n_coord of the request initial coordinate corresponding to the next request remain unchanged.
[0235] Moreover, at this time, the length of the data loaded by the first request is the difference between the size of the original tensor in the C dimension, tensor_c, and the coordinate value c_coord of the initial coordinate of the next request in the C dimension.
[0236] Since the state of the first request is the second state s1 and all the sub-data it loads are in the memory, in step S40, the first request is sent to the memory to load the sub-data in the original tensor stored in the memory into the memory.
[0237] As shown in Table 1, when the first coordinate value c_coord_b is greater than or equal to tensor_c, it is determined that the first request enters the third state s2. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0238] Determine the state that the next request enters according to the first coordinate value c_coord_b. For example, when the first coordinate value c_coord_b is less than 0, it is determined that the next request enters the first state s0; when the first coordinate value c_coord_b is greater than or equal to 0 but less than tensor_c, it is determined that the next request enters the second state s1; when the first coordinate value c_coord_b is greater than or equal to tensor_c, it is determined that the next request enters the third state s2.
[0239] As shown in Table 1, regardless of which state the next request enters, the update logic of the request initial coordinate and the length of the data loaded by the first request are exactly the same. Moreover, the coordinate updates of the coordinate values w_coord, h_coord, d_coord, and n_coord of the request initial coordinate corresponding to the next request in the W dimension, H dimension, D dimension, and N dimension are the same as those when the first state s0 jumps to the first state s0 and will not be elaborated here.
[0240] The length of the data loaded by the first request is the difference between the size of the original tensor in the C dimension, tensor_c, and the coordinate value c_coord of the initial coordinate of the next request in the C dimension.
[0241] Since the coordinate value c_coord of the initial coordinate of the first request in the C dimension is greater than tensor_c, all the sub-data loaded by the first request are not in the memory. In step S40, the first request is converted into an operation of writing c_coord-tensor_c zeros to the buffer area.
[0242] After that, continue the above process to determine the request initial coordinates and the lengths of the data loaded by subsequent requests, and execute step S40 to send multiple requests. The specific process will not be elaborated here.
[0243] For example, in a specific example, assume that the N dimension, D dimension, H dimension, and W dimension are expanded into one dimension, and this dimension corresponds to a coordinate information table for indicating the starting coordinates of the data part. At this time, in step S30, when determining multiple requests, the coordinate value of the request initial coordinate corresponding to each request in the C dimension changes according to the state of the state machine, as described above, and will not be elaborated here; the coordinate value of the request initial coordinate corresponding to each request in the first dimension is sequentially updated according to the coordinate information table. For example, the coordinate value of the request initial coordinate corresponding to the first request in the first dimension is the first starting coordinate in the coordinate information table, and the coordinate value of the request initial coordinate corresponding to the second request in the first dimension is the second starting coordinate in the coordinate information table, and so on.
[0244] Regarding the loading lengths corresponding to each request, reference can be made to the relevant description in Table 1 above, and it will not be elaborated here.
[0245] Thus, multiple requests can be determined, and multiple requests are sent in sequence, and the sub-data returned by each request is sequentially written into the buffer area to load the target data into the buffer area.
[0246] In the above embodiment, the starting coordinates of the data to be loaded in other dimensions except the number of channels are obtained through the coordinate information table. Therefore, the target data to be loaded can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data loaded into a certain expert model. For example, as described above, only a certain token can be loaded into the buffer area of the hardware computing unit deploying the corresponding expert model through the coordinate information table, so that more flexible data loading can be achieved according to the coordinate information table.
[0247] In addition, according to whether the tensor can be continuously loaded in the number of channels dimension, requests are divided. Each request is used to load sub-data that are all located in the memory or all not located in the memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, and further improving the hardware utilization rate of the computing unit and the hardware performance. [[ID=I4]]
[0248] At least one embodiment of the present disclosure also provides a data storage method for writing target data in the buffer area into the memory.
[0249] For example, the target data includes multiple data parts. The coordinate ranges of each data part in the number of channels dimension in the coordinate system determined by the original tensor are the same, and are both from the first coordinate value to the first channel coordinate. The difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the number of channels dimension. Regarding the relevant content of the target data, reference can be made to the description of the foregoing data loading method, and it will not be elaborated here.
[0250] Figure 7 Schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure.
[0251] As Figure 7 shown, the data storage method provided by at least one embodiment of the present disclosure at least includes steps S50 - S80.
[0252] First, in step S50, obtain a first coordinate value and a first size.
[0253] Regarding the first coordinate value and the first size, reference can be made to the relevant descriptions in the foregoing data loading method, which will not be elaborated here.
[0254] In step S60, obtain a coordinate information table.
[0255] Regarding the relevant description of step S60, reference can be made to the relevant descriptions in the foregoing data loading method, which will not be elaborated here.
[0256] In step S70, based on the first coordinate value, the first size, and the coordinate information table, determine multiple requests for storing target data.
[0257] In step S80, sequentially send at least some of the multiple requests, and write the data to be written indicated by each request into the memory in sequence to store the target data in the memory.
[0258] For example, in some embodiments, step S70 may include: based on the first coordinate value, the first size, and the coordinate information table, determine the first request sent among the multiple requests and the initial state when the first request enters the state machine; based on the initial state, in combination with the coordinate information table, use the state machine to determine each of the multiple requests other than the first request.
[0259] For example, in response to the first size not being equal to the shape size of the original tensor in the channel number dimension, and / or the first coordinate value not being equal to the second coordinate value of the starting coordinate of the original tensor in the channel number dimension, determine that the target data cannot be continuously loaded in the channel number dimension.
[0260] For example, referring to Figure 6 , the arrow in the C dimension represents the data of the target data in the C dimension direction. The data with the coordinate in the C dimension less than the second coordinate value and greater than the third coordinate value is not in the memory, or rather, the data range of this part in the C dimension is not within the data range of the original tensor ( Figure 6 the part shown in the black background in Figure 6The white part in the middle is stored in the memory, or in other words, the data range of this part in the C dimension is within the data range of the original tensor.
[0261] For example, the sub-data for writing all belong to the data range of the original tensor, indicating the coordinate range of the sub-data in any dimension. For example, referring to Figure 6 the C dimension, it is between the corresponding second coordinate value and the third coordinate value, that is, it belongs to the original tensor. At this time, the sub-data for writing, as a part of the original tensor, needs to be stored in the memory.
[0262] For example, the sub-data for writing all do not belong to the data range of the original tensor, indicating the coordinate range of the sub-data in a certain dimension. For example, referring to Figure 6 the C dimension, it is less than the second coordinate value or greater than the third coordinate value, that is, it does not belong to the original tensor. At this time, the sub-data for writing does not belong to the original tensor and may not be stored in the memory subsequently.
[0263] The determination method for each request can refer to the relevant content of step S30 in the foregoing data loading method, and the repeated parts will not be elaborated here.
[0264] For example, in some embodiments, step S80 may include: for any request, in response to the first request coordinate value of the request initial coordinate corresponding to any request in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determining not to send any request; in response to the first request coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining to send any request; wherein, the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0265] For example, if the first request coordinate value of the request initial coordinate corresponding to the request in the first dimension is less than the second coordinate value, or the first request coordinate value is greater than or equal to the third coordinate value, this indicates that the sub-data for storage requested does not belong to the original tensor and does not need to be stored in the memory, so this request is not sent.
[0266] For example, if the first request coordinate value is greater than or equal to the second coordinate value and less than the third coordinate value, this indicates that the sub-data for storage requested belongs to the original tensor and needs to be stored in the memory, and this request is sent to store the sub-data at the corresponding position in the original tensor.
[0267] As Figure 5A shown in the example, the target data is composed of part of the data of the original tensor. At this time, all the target data in the buffer need to be stored at the corresponding positions in the original tensor to update the original tensor. Therefore, at this time, multiple requests all need to be sent to the memory to store the sub-data for storage of each request at the corresponding positions in the memory.
[0268] As Figure 5B shown in the example, the target data not only includes partial data of the original tensor, but also includes data outside the edge of the original tensor caused by padding operations, etc., which are not stored in memory. Therefore, the request needs to be distinguished before sending. If the sub-data requested for storage does not belong to the content of the original tensor, such as Figure 5B the small dashed cube in, there is no need to send a request, and these sub-data do not need to be stored in memory; if the sub-data requested for storage belongs to the original tensor, such as Figure 5B the small solid cube in the target data that belongs to the original tensor, a request is sent to the memory to store the sub-data at the corresponding position in the original tensor.
[0269] The data storage method provided by at least one embodiment of the present disclosure obtains the starting coordinates of the data to be stored in other dimensions except the number of channels through the coordinate information table. Therefore, the target data to be stored can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data that needs to be stored in memory as the processing result of a certain expert model. For example, as described above, only the processing result corresponding to a certain token can be stored from the buffer area of the hardware computing unit where the corresponding expert model is deployed to the corresponding position in memory through the coordinate information table to update the corresponding position of the original tensor, so that more flexible data storage can be achieved according to the coordinate information table.
[0270] In addition, the present disclosure divides requests according to whether the tensor can be continuously stored in the first dimension. Each request is used to store sub-data that are all in memory or all not in memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, and further improving the hardware utilization rate of the computing unit and the hardware performance.
[0271] At least one embodiment of the present disclosure also provides a data loading method. Figure 8 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure.
[0272] As Figure 8 shown, the data loading method provided by at least one embodiment of the present disclosure includes steps S201 - S202.
[0273] For example, in step S201, a data loading instruction indicating to load target data into the buffer area is received.
[0274] For example, the data loading instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first dimension. The target data includes a plurality of data parts, and each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor, and is from the first coordinate value to the first channel coordinate. The difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension.
[0275] For the related descriptions of the target data and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0276] For example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction can be a Load instruction.
[0277] In step S202, after parsing the data loading instruction, the execution unit is used to execute the data loading instruction.
[0278] For example, after receiving the data loading instruction, the processor parses the data loading instruction, for example, decodes the data loading instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameter is ready, the execution unit executes the related operations of the data loading instruction.
[0279] For example, step S202 may include: reading the coordinate information table from the memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first dimension, and the coordinate information table, determining a plurality of requests for loading the target data; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the target data into the buffer area.
[0280] For example, in response to the size relationship between the target data and the original tensor in the dimension of the number of channels, such that when loading the target data, continuous loading cannot be performed in the dimension of the number of channels, the request is divided in the first dimension, and the sub-data loaded by each request is either all located in the memory or none of it is located in the memory. And in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory. The at least one dimension includes the first dimension. When loading data, the dimension of the number of channels is loaded prior to the first dimension and the dimension of the number of channels is adjacent to the first dimension.
[0281] Regarding "determining a plurality of requests for loading the target data based on the first coordinate value, the first size, and the coordinate information table", reference may be made to the relevant content of the foregoing step S30, which will not be elaborated here.
[0282] Regarding "sequentially sending the plurality of requests and sequentially writing the sub-data returned by each request into the buffer area to load the target data into the buffer area", reference may be made to the relevant content of the foregoing step S40, which will not be elaborated here.
[0283] In the above embodiment, the starting coordinates of the data to be loaded in other dimensions except the number of channels are obtained through the coordinate information table. Therefore, the target data to be loaded can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data loaded into a certain expert model. For example, as described above, only a certain token can be loaded into the buffer area of the hardware computing unit deploying the corresponding expert model through the coordinate information table, so that more flexible data loading can be achieved according to the coordinate information table.
[0284] In the present disclosure, the data loading instruction is divided into multiple requests. According to whether the tensor can be continuously loaded in the first dimension, the data loading instruction is converted into multiple requests, and the sub-data loaded by each request is either all located in the memory or none of it is located in the memory. The division of the data requests is more reasonable and more suitable for the storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.
[0285] At least one embodiment of the present disclosure further provides a data storage method. Figure 9 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure.
[0286] As Figure 9 shown, the data storage method provided by at least one embodiment of the present disclosure includes steps S203 - S204.
[0287] For example, in step S203, a data storage instruction indicating to execute writing target data in a buffer to memory is received.
[0288] For example, the data storage instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first dimension. The target data includes a plurality of data parts, and each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor and is from the first coordinate value to the first channel coordinate. The difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension.
[0289] For the relevant descriptions of the target data and the original tensor, reference can be made to the relevant descriptions of the foregoing data storage method, and the repeated parts will not be elaborated.
[0290] For example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction can be a Store instruction.
[0291] In step S204, after parsing the data storage instruction, the execution unit is used to execute the data storage instruction.
[0292] For example, after receiving the data storage instruction, the processor parses the data storage instruction. For example, it decodes the data storage instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameters are ready, the execution unit executes the relevant operations of the data storage instruction.
[0293] For example, step S204 may include: reading the coordinate information table from memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the plurality of dimensions included in the coordinate system except the channel number dimension; based on the first coordinate value, the first dimension, and the coordinate information table, determining a plurality of requests for storing the target data; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the target data to the memory.
[0294] For example, in response to the size relationship between the target data and the original tensor in the channel number dimension, when storing the target data, the channel number dimension cannot be stored continuously, and the request is divided on the first dimension, and each sub-data requested for writing belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the sub-data requested for writing all belong to the data range of the original tensor, the sub-data requested for writing belong to the original tensor and are stored continuously in the memory, the at least one dimension includes the first dimension, and when loading data, the channel number dimension is loaded before the first dimension and the channel number dimension is adjacent to the first dimension.
[0295] Regarding “determining multiple requests for storing the target data based on the first coordinate value, the first size, and the coordinate information table,” reference may be made to the relevant content of the aforementioned step S70 , which will not be repeated here.
[0296] Regarding "sending at least part of the multiple requests in sequence, and writing the data to be written indicated by each request into the memory in sequence to store the target data into the memory", please refer to the relevant content of the aforementioned step S80, which will not be repeated here.
[0297] At least one embodiment of the present disclosure provides a data storage method that obtains the starting coordinates of the data to be stored in dimensions other than the number of channels through a coordinate information table. Therefore, the target data to be stored can be composed of multiple discrete parts, rather than being a continuous whole in at least two dimensions. For example, taking the MOE layer as an example, the starting coordinates can indicate the data that needs to be stored in the memory as the processing result of a certain expert model. For example, as described above, the coordinate information table can be used to store only the processing result corresponding to a certain word from the cache area of the hardware computing unit that deploys the corresponding expert model to the corresponding location in the memory to update the corresponding position of the original tensor, thereby achieving more flexible data storage based on the coordinate information table.
[0298] In addition, the present disclosure divides requests according to whether the tensor can be stored continuously in the first dimension. The sub-data for each request to be stored are all located in the memory or not located in the memory. The division of data requests is more reasonable and more suitable for the loading and storage of tensor data, which greatly improves the bandwidth and efficiency of data memory access, thereby improving the hardware utilization of the computing unit and improving hardware performance.
[0299] Figure 10 This is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 10 As shown, the electronic device 300 is suitable for implementing the data loading method or data storage method provided by the embodiment of the present disclosure. It should be noted that Figure 10The components of the electronic device 300 shown are merely exemplary and not restrictive. According to actual application requirements, the electronic device 300 may also have other components.
[0300] As Figure 10 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.
[0301] For example, when the computer-readable instructions are run by the processing device 301, one or more steps of the data loading method described in any of the above embodiments, or one or more steps of the data storage method described in any of the above embodiments may be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.
[0302] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) 303 and / or cache memory, etc. For example, the computer-readable instructions may be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may, for example, include read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage medium, such as style images, and various data used and / or generated by the application programs, etc.
[0303] For example, the processing device 301, the read-only memory (ROM) 302, and the random access memory (RAM) 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0304] Typically, the following devices can be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 10 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 can alternatively implement or have more or fewer devices. For example, the processing device 301 can control other components in the electronic device 300 to perform desired functions. The processing device 301 can be a device with data processing capabilities and / or program execution capabilities such as a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU. The central processing unit (CPU) can be of the X86, ARM, RISC-V architecture, etc. The GPU can be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.
[0305] Figure 11 Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 11 shown, the processor 400 includes an instruction parsing unit 401 and a first execution unit 402.
[0306] For example, the instruction parsing unit 401 is used to receive and parse data loading instructions.
[0307] For example, the data loading instruction includes a first coordinate value, a memory read address of a coordinate information table, and a first dimension as input parameters, the target data includes multiple data parts, each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor, and is from the first coordinate value to the first channel coordinate, the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension.
[0308] For the related descriptions of the target data and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0309] For example, after the instruction parsing unit parses the data loading instruction, the first execution unit 402 executes the data loading instruction.
[0310] For example, when the first execution unit 402 executes a data loading instruction, it includes performing the following operations: reading the coordinate information table from the memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; determining a plurality of requests for loading the target data based on the first coordinate value, the first size, and the coordinate information table; and sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the target data into the buffer area.
[0311] For example, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, the requests are divided in the first dimension, the sub-data loaded by each request is either all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, and the at least one dimension includes the first dimension. When loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0312] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send a data loading instruction for computing and processing to the processor (such as a CPU or a GPU) through a unified encapsulated function library, the data loading instruction can carry the memory read address as the first coordinate value, the coordinate information table, and the first size; when the processor receives the data loading instruction, the instruction parsing unit 401 parses the data loading instruction to obtain the first coordinate value, the memory read address of the coordinate information table, and the first size as input parameters, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data loading instruction, the processor can store the input parameters in the data loading instruction in a register or memory, and when the first execution unit 402 performs computing and processing, it can obtain the input parameters from the register or memory.
[0313] Regarding the specific process of using the first execution unit 402 to execute the data loading instruction, reference can be made to steps S20 - S40 in the data loading method described above, and the repeated parts will not be elaborated.
[0314] The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method, and the repeated parts will not be elaborated.
[0315] Figure 12Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 12 shown, the processor 400 includes an instruction parsing unit 401 and a second execution unit 403.
[0316] For example, the instruction parsing unit 401 is used to receive and parse data storage instructions.
[0317] For example, the data storage instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first dimension. The target data includes a plurality of data parts, and the coordinate ranges of each data part in the channel number dimension of the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension.
[0318] For the related descriptions of the target data and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0319] For example, after the instruction parsing unit parses the data storage instruction, the second execution unit 403 executes the data storage instruction.
[0320] For example, when the second execution unit 403 executes the data storage instruction, it includes performing the following operations: reading the coordinate information table from the memory according to the memory read address, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; determining a plurality of requests for storing the target data based on the first coordinate value, the first dimension, and the coordinate information table; and sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the target data into the memory.
[0321] For example, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when storing the target data, continuous storage cannot be performed in the channel number dimension, requests are divided in the first dimension. Each sub-data to be written by each request belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the sub-data to be written by the request belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory. The at least one dimension includes the first dimension. When loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
[0322] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data storage instructions for computing and processing to the processor (such as CPU or GPU) through a unified encapsulated function library, the data storage instructions can carry the first coordinate value as an input parameter, the memory read address of the coordinate information table, and the first size; when the processor receives the data storage instruction, the instruction parsing unit 401 parses the data storage instruction to obtain the first coordinate value as an input parameter, the memory read address of the coordinate information table, and the first size, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data storage instruction, the processor can store the input parameters in the data storage instruction into a register or memory, and the second execution unit 403 can obtain the input parameters from the register or memory when performing computing and processing.
[0323] Regarding the specific process of using the second execution unit 403 to execute the data storage instruction, reference can be made to steps S70 - S80 in the data loading method described above, and the repeated parts will not be elaborated.
[0324] The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data storage method, and the repeated parts will not be elaborated.
[0325] Figure 13 It is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 13 shown, the storage medium 600 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 601 can be non-temporarily stored on the storage medium 600. For example, when the computer-readable instructions 601 are executed by the processor, one or more steps in the data loading method described above can be executed. For example, when the computer-readable instructions 601 are executed by the processor, one or more steps in the data storage method described above can be executed.
[0326] For example, the storage medium 600 can be applied to the electronic device 300. For example, the storage medium 600 can include the storage device 308 in the electronic device 300.
[0327] For example, a storage device may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor may run the computer-readable instructions to implement various functions of the processor. Various application programs and various data may also be stored in the storage medium.
[0328] For example, the storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or may be other applicable storage media.
[0329] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0330] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0331] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0332] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, it should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0333] In addition, although the operations are depicted in a specific order, this should not be construed as requiring that the operations be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0334] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0335] Regarding the present disclosure, the following points also need to be noted:
[0336] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures can refer to the general design.
[0337] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0338] The above description is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A data loading method for loading target data into a buffer, the data loading method comprising: Obtaining a first coordinate value and a first size, wherein the target data includes a plurality of data parts, each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor, and is from the first coordinate value to the first channel coordinate, and the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; Obtaining a coordinate information table, wherein the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the plurality of dimensions included in the coordinate system except the channel number dimension; Based on the first coordinate value, the first size, and the coordinate information table, determining a plurality of requests for loading the target data; Sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the target data into the buffer; Wherein, in response to the size relationship between the target data and the original tensor in the channel number dimension such that the target data cannot be continuously loaded in the channel number dimension, requests are divided in the first dimension, the sub-data loaded by each request is either all located in the memory or none of them is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, The at least one dimension includes the first dimension, and during data loading, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
2. The data loading method according to claim 1, wherein In response to the first size not being equal to the shape size of the original tensor in the channel number dimension, and / or the first coordinate value not being equal to the second coordinate value of the starting coordinate of the original tensor in the channel number dimension, it is determined that the target data cannot be continuously loaded in the channel number dimension.
3. The data loading method according to claim 1, wherein, Based on the first coordinate value, the first size, and the coordinate information table, determining a plurality of requests for loading the target data includes: Based on the first coordinate value, the first size, and the coordinate information table, determining the first request sent among the plurality of requests and the initial state of the first request entering the state machine; Based on the initial state, in combination with the coordinate information table, using the state machine to determine each request among the plurality of requests other than the first request.
4. The data loading method according to claim 1, wherein Each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data into the buffer, and the data length loaded by the request, Based on the first coordinate value, the first size, and the coordinate information table, determining the first request sent among the plurality of requests and the initial state of the first request entering the state machine includes: Based on the first coordinate value, determining the initial state of the first request entering the state machine; Use the first coordinate value as the coordinate value of the request initial coordinate corresponding to the first request in the dimension of the number of channels; Determine the coordinate values of the request initial coordinate corresponding to the first request in the at least one dimension from the coordinate information table; Determine the data reading address of the first request according to the shape and size of the original tensor and the request initial coordinate corresponding to the first request; Determine the starting address for writing the target data in the buffer as the data writing address of the first request; Determine the data length loaded by the first request according to the first coordinate value and the first size; 5. The data loading method according to claim 4, wherein, Each of the at least one dimension has a corresponding coordinate information table, Determining the coordinate values of the request initial coordinate corresponding to the first request in the at least one dimension from the coordinate information table includes: For the coordinate information table corresponding to each dimension in the at least one dimension, select the first starting coordinate from the coordinate information table corresponding to the dimension, and use the first starting coordinate as the coordinate value of the request initial coordinate corresponding to the first request in the dimension; 6. The data loading method according to claim 3, wherein, Based on the initial state, in combination with the coordinate information table, use the state machine to determine each request other than the first request among the multiple requests, including: Based on the initial state, in combination with the coordinate information table, use the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request; Determine the data reading address of each request according to the shape and size of the original tensor and the request initial coordinates corresponding to each request; Determine the data writing address of each request according to the data length loaded by each request; 7. The data loading method according to claim 6, wherein, Based on the initial state, in combination with the coordinate information table, using the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request includes: Based on the initial state, in combination with the relationship between the first size, the first coordinate value, the second coordinate value, and the third coordinate value, determine the state transition of the state machine, where the state transition includes the transition from the current state where the current request is located to the next state entered by the next request, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the dimension of the number of channels, and the difference between the second coordinate value and the third coordinate value is the shape size of the original tensor in the dimension of the number of channels; According to the state transition of the state machine, determine the coordinate value of the request initial coordinate corresponding to the next request in the dimension of the number of channels and the data length loaded by the current request; Determine the coordinate values of the request initial coordinate corresponding to the next request in the at least one dimension according to the coordinate information table; 8. The data loading method according to claim 7, wherein, Each of the at least one dimension has a corresponding coordinate information table, Determining the coordinate values of the request initial coordinate corresponding to the next request in the at least one dimension according to the coordinate information table includes: For any one of the at least one dimension, in response to the coordinate value of the request initial coordinate corresponding to the current request in the second dimension being the last starting coordinate in the coordinate information table corresponding to the second dimension, the coordinate value of the request initial coordinate corresponding to the next request in any one of the dimensions is updated from the current starting coordinate to the next starting coordinate, where the next starting coordinate is the next starting coordinate adjacent to the current starting coordinate in the coordinate information table corresponding to any one of the dimensions and in the preset order, and the second dimension is adjacent to and lower than any one of the dimensions; Among them, when the coordinate values of the request initial coordinates corresponding to each request in any one of the dimensions are updated, they are sequentially updated according to the preset order based on the multiple starting coordinates included in the coordinate information table corresponding to any one of the dimensions.
9. The data loading method according to any one of claims 1-8, wherein Obtaining the coordinate information table includes: Receiving a mode parameter, where the mode parameter is used to indicate loading the target data into the buffer; Receiving the memory read address of the coordinate information table; Reading the coordinate information table from the memory according to the memory read address of the coordinate information table.
10. The data loading method according to claim 9, wherein, Each of the at least one dimension has a corresponding coordinate information table, and the memory includes multiple storage units, and each storage unit has a corresponding memory read address. In the storage unit indicated by the memory read address of the coordinate information table corresponding to each dimension, N starting coordinates are stored, and N is a positive integer greater than 1. Reading the coordinate information table from the memory according to the memory read address of the coordinate information table includes: For any one of the at least one dimension, in response to the shape size of the target data in any one of the dimensions being greater than N: Reading the first storage unit indicated by the memory read address of the coordinate information table corresponding to any one of the dimensions to obtain the N starting coordinates stored in the first storage unit; Adding a preset value to the memory read address of the coordinate information table corresponding to any one of the dimensions, and reading the second storage unit indicated by the accumulated result to obtain the N data in the second storage unit; In response to the shape size M of the target data in any one of the dimensions being less than 2×N, selecting the M - N data located at the lower positions from the N data as the starting coordinates for generating requests; In response to M being greater than or equal to 2×N, using the N data as the N starting coordinates for subsequent request generation.
11. The data loading method according to any one of claims 1-8, wherein, The target data is used for data processing of the mixture of experts model. The mixture of experts model includes a routing module and multiple expert networks. The coordinate information table is obtained through the routing module. Each of the at least one dimension has a corresponding coordinate information table, and the coordinate information table corresponding to each dimension is used to indicate the data that needs to be input to the target expert network for processing by the target expert network on that dimension. Each expert network is trained to process specific tasks and data features.
12. The data loading method according to any one of claims 1-8, wherein, Sequentially sending the multiple requests and sequentially writing the sub - data returned by each request into the buffer to load the target data into the buffer includes: For any one request, if all the sub-data loaded in response to the any one request is located in the memory, send the any one request to the memory; If all the sub-data loaded in response to the any one request is not located in the memory, convert the any one request into writing a plurality of predetermined values to the buffer area, where the number of the plurality of predetermined values is determined by the data length to be loaded specified by the any one request.
13. A data loading method, comprising: Receiving a data loading instruction indicating to load target data to a buffer area, where the data loading instruction includes a first coordinate value as an input parameter, a memory read address of a coordinate information table, and a first dimension, the target data includes a plurality of data parts, each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor, and is from the first coordinate value to the first channel coordinate, the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension; After parsing the data loading instruction, use an execution unit to execute the data loading instruction. Wherein, using the execution unit to execute the data loading instruction includes: Reading the coordinate information table from the memory according to the memory read address of the coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension in the coordinate system, and the at least one dimension is other dimensions in the plurality of dimensions included in the coordinate system except the channel number dimension; Based on the first coordinate value, the first dimension, and the coordinate information table, determining a plurality of requests for loading the target data; Sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request to the buffer area to load the target data to the buffer area; Wherein, in response to the size relationship between the target data and the original tensor in the channel number dimension such that continuous loading cannot be performed in the channel number dimension when loading the target data, requests are divided in a first dimension, each request loads sub-data that is all located in the memory or all not located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory. The at least one dimension includes the first dimension, and when loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
14. A data storage method for writing target data in a buffer area to a memory, the data storage method comprising: Obtaining a first coordinate value and a first dimension, where the target data includes a plurality of data parts, each data part has the same coordinate range in the channel number dimension in the coordinate system determined by the original tensor, and is from the first coordinate value to the first channel coordinate, the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first dimension, and the first dimension is the shape dimension of the target data in the channel number dimension; Obtain a coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; Based on the first coordinate value, the first size, and the coordinate information table, determine multiple requests for storing the target data; Sequentially send at least some of the multiple requests, and sequentially write the data to be written indicated by each request into the memory to store the target data in the memory; Among them, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when storing the target data, continuous storage cannot be performed in the channel number dimension, divide the requests in the first dimension. Each sub-data to be written by each request belongs to the data range of the original tensor or does not belong to the data range of the original tensor. And in response to the sub-data to be written by the request belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory, The at least one dimension includes the first dimension. When loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
15. The data storage method according to claim 14, wherein, The request initial coordinate corresponding to each request is used to determine the data storage address of the sub-data to be stored by the request in the memory, Sequentially send at least some of the multiple requests, and sequentially write the data to be written indicated by each request into the memory to store the target data, including: For any one request, in response to the first request coordinate value of the request initial coordinate corresponding to the any one request in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determine not to send the any one request; In response to the first request coordinate value being greater than or equal to the second coordinate value and the first request coordinate value being less than the third coordinate value, determine to send the any one request; Among them, the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
16. A processor, including an instruction parsing unit and an execution unit, where, The instruction parsing unit is used to receive and parse a data loading instruction, where the data loading instruction is used to load target data into a buffer area. The data loading instruction includes a first coordinate value, a memory read address of a coordinate information table, and a first size as input parameters. The target data includes multiple data parts, and the coordinate ranges of each data part in the channel number dimension of the coordinate system determined by the original tensor are the same, and are all from the first coordinate value to the first channel coordinate. The sum of the difference between the first coordinate value and the first channel coordinate plus 1 is equal to the first size, and the first size is the shape size of the target data in the channel number dimension; The execution unit executes the data loading instruction after the instruction parsing unit parses the data loading instruction, When the execution unit executes the data loading instruction, the following operations are included: Read the coordinate information table from the memory according to the memory read address of the coordinate information table, where the coordinate information table is used to indicate the starting coordinates of each data part in at least one dimension of the coordinate system, and the at least one dimension is other dimensions in the multiple dimensions included in the coordinate system except the channel number dimension; Determine a plurality of requests for loading the target data based on the first coordinate value, the first size, and the coordinate information table; Send the plurality of requests in sequence, and write the sub-data returned by each request into the buffer area in sequence to load the target data into the buffer area; Wherein, in response to the size relationship between the target data and the original tensor in the channel number dimension such that when loading the target data, continuous loading cannot be performed in the channel number dimension, the requests are divided in the first dimension, the sub-data loaded by each request is either all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, The at least one dimension includes the first dimension. When loading data, the channel number dimension is loaded prior to the first dimension and the channel number dimension is adjacent to the first dimension.
17. An electronic device, comprising: A memory that stores computer-executable instructions non-transiently; A processor configured to run the computer-executable instructions, Wherein, when the computer-executable instructions are run by the processor, the data loading method according to any one of claims 1-13, or the data storage method according to claim 14 or 15 is implemented.
18. A non-transitory computer-readable storage medium, wherein, The non-transient computer-readable storage medium stores computer-executable instructions, When the computer-executable instructions are executed by a processor, the data loading method according to any one of claims 1-13, or the data storage method according to claim 14 or 15 is implemented.
Citation Information
Patent Citations
Tensor non-uniform splitting in shuffled secure multi-party operations
CN116893895A
Circular buffer for input and output of tensor computations
US20240281393A1