Chip, Method, Device, and Medium for Flexibly Accessing Data
The AI processor chip addresses inefficiencies in data access by using a software-configurable address calculation module to dynamically reorder data access, enhancing efficiency and flexibility in handling tensor operations.
Patent Information
- Application Number
- JP2025501618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-07-12
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-07-12
AI Technical Summary
Existing AI chips face inefficiencies due to rigid data access mechanisms, leading to high power consumption and reduced flexibility when adapting to modified network structures, and they struggle with latency and bandwidth issues in data transfer between different memory levels.
An AI processor chip with a memory control unit that includes an address calculation module, allowing flexible data access through software-configurable address calculation in single or multiple loops, enabling dynamic reordering of data access based on tensor operations.
Enhances calculation efficiency by reducing hardware costs and delays, improving flexibility in handling various tensor operations without the need for sequential data access, thereby optimizing performance and power usage.
Smart Images

Figure 2025522088000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to an artificial intelligence processor chip that flexibly accesses data, a method for flexibly accessing data in an artificial intelligence processor chip, an electronic device, and a non-transitory storage medium.
[0002] This application claims priority to Chinese Patent Application No. 202210836577.0, filed on July 15, 2022, and hereby incorporates by reference in its entirety the content disclosed in the above Chinese patent application as part of this application.
Background Art
[0003] Artificial Intelligence (AI) chips need to achieve better performance and fully utilize computing power to increase the utilization rate of Media Access Control (MAC) bit addresses. Therefore, the design of the data pipeline is also very important. During the calculation of neural networks, the forms of data input to or output from the calculation unit and the storage forms are diverse, which also determines different chip memory calculation architectures. For example, a Graphics Processing Unit (GPU) is a parallel computing processor that adopts a multi-level high-speed cache system consisting of an L1 cache, a shared internal memory, a register group, an L2 cache, and an external memory DRAM. Dividing these memory levels is mainly to reduce the delay of data transfer and improve the bandwidth. The L1 cache is usually divided into an L1D cache and an L1I cache, which are used to store data and instructions respectively. Generally, a corresponding L1 cache is provided for each processor core. The size of the L1 cache is 16k - 64k respectively and is different. The L2 cache is always a private cache, which does not distinguish between instructions and data. Generally, a corresponding L2 cache is provided for each processor core. The size of the L2 cache is 256k - 1M and is different. For example, the L1 cache has the highest speed but the smallest space, the L2 cache has a slower cache speed and a larger space, and the external DRAM has the largest space but the slowest speed. Therefore, by storing frequently accessed data from the DRAM to the L1 cache, the delay of transporting data from the external DRAM to the internal memory for each access can be reduced, and the efficiency of data processing can be improved. However, to ensure the generality and flexibility of the processor structure, there is a certain redundancy in the data pipeline. For example, it is necessary to fetch registers or data every time a calculation is performed and finally store the data in the register, resulting in high power consumption.
[0004] Some AI chips can achieve high efficiency by customizing the pipeline, but at the cost of losing flexibility, and there is a risk of becoming unusable if the network structure is modified. Also, some AI chips solve the problems of bandwidth and latency by adding a large on-chip buffer, but the data access form of static random access memory (SRAM) is initiated by hardware, that is, the connection between calculation and storage is coupled through hardware. Due to the problem that the policy is not flexible in this way, in some scenarios, the efficiency may decrease and software may not be able to intervene.
Summary of the Invention
Means for Solving the Problems
[0005] According to one aspect of the present application, there is provided an artificial intelligence processor chip that flexibly accesses data, including a memory that stores tensor data read from outside the processor chip, the read tensor data including a plurality of elements used for tensor operations of operators included in artificial intelligence calculations, a memory control unit that controls to read elements from the memory based on the tensor operations of the operators and send them to a calculation unit, the memory control unit including an address calculation module having an interface that receives parameters set by software, the address calculation module calculating an address in the memory in a single read loop or a multiple loop (nested read loop) based on the parameters received by the interface, reading elements from the calculated address and sending them to the calculation unit, and a calculation unit that performs tensor operations of the operators on the received elements.
[0006] In another aspect, a method for flexibly accessing data in an artificial intelligence processor chip, wherein the memory in the artificial intelligence processor chip stores tensor data read from outside the processor chip, and the read tensor data includes a plurality of elements used for tensor operations of operators involved in artificial intelligence calculations; and based on the tensor operations of the operators, controlling to read elements from the memory and send them to a calculation unit, calculating addresses in the memory in a single-read loop or a multi-read loop (nested loop) based on parameters set by received software, reading elements from the calculated addresses and sending them to the artificial intelligence processor chip, and the calculation unit performing tensor operations of the operators on the received elements. A method for flexibly accessing data in an artificial intelligence processor chip is provided.
[0007] In another aspect, an electronic device includes a memory for storing instructions and a processor for reading the instructions in the memory and executing the methods of the embodiments of the present application.
[0008] In another aspect, a non-transitory storage medium stores instructions, and when the instructions are read by a processor, the methods of the embodiments of the present application are executed by the processor. When the instructions are read by a processor, the methods of the embodiments of the present application are executed by the processor.
Advantages of the Invention
[0009] In this way, by flexibly calculating addresses in the memory according to parameters set by software, elements in the memory can be flexibly read, and it is not limited to the order of these elements or the address order stored in the memory.
Brief Description of the Drawings
[0010] To more clearly explain the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly describes the drawings that need to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings based on these drawings without creative labor.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Modes for Carrying Out the Invention
[0011] With reference to the specific embodiments of the present application in detail, examples of the present application are illustrated in the drawings. Together with the specific embodiments, the present application will be described, but it is understood that the present application is not intended to be limited to the described embodiments. On the contrary, it is intended to cover modifications, variations, and equivalents within the spirit and scope of the present application as defined by the appended claims. It should be noted that the method steps described herein can be implemented by any functional block or functional configuration, and any functional block or functional configuration can be implemented as a physical entity or a logical entity, or a combination of both.
[0012] Before using the technical solutions disclosed in each embodiment of the present disclosure, it should be understood that the types, scope of use, usage scenarios, etc. of personal information related to the present disclosure should be notified to the user in an appropriate manner based on relevant laws and regulations, and the user's permission must be obtained.
[0013] For example, in response to receiving a user's voluntary request, an operation of sending presentation information to the user and requesting execution explicitly presents to the user that it is necessary to obtain and use the user's personal information. Therefore, based on the presented information, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application, a server, or a storage medium that executes the operation of the technical solution of the present disclosure.
[0014] As an optional and non-limiting embodiment, the form of sending presentation information to the user in response to receiving a user's voluntary request may be, for example, in the form of a pop-up that can present presentation information in text. In addition, the pop-up window can also be equipped with a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0015] It should be noted that the above notification and user permission acquisition process are merely schematic and do not limit the embodiments of the present disclosure. Other forms that meet the relevant laws and regulations can also be applied to the embodiments of the present disclosure.
[0016] In addition, the data related to this technical solution (including but not limited to the data itself, data acquisition or use) should comply with the requirements of corresponding laws, regulations and related provisions.
[0017] The identification process of the above application scenario can be realized by receiving various application data input by a neural network as tensors and going through the calculations of the neural network. Currently, neural networks and machine learning systems use tensors as the basic data structure. The core of the concept of a tensor is a data container, and since most of the data it contains is always numerical data, it is a digital container. The specific numerical values in a tensor may be application data, for example, including image data, natural language data, etc.
[0018] For example, a scalar is a tensor of rank 0, such as 2, 3, 5. In a specific application scenario, for example, it is image data. 2 represents, for example, the grayscale value of one pixel in the image data, 3 represents, for example, the grayscale value of one pixel in the image data, and 5 represents, for example, the grayscale value of one pixel in the image data, etc. For example, a vector is a tensor of rank 1, such as [0, 3, 20]. A matrix is a tensor of rank 2, for example, [Number] Or it is [[2,3], [1,5]]. For example, furthermore, there can also be a third-order tensor (e.g., a: (shape: (3,2,1)), [[[1],[2]],[[3],[4]],[[5],[6]]]), or even a fourth-order tensor, etc. All of these tensors can be used to represent data in specific application scenarios, such as image data, natural language data, etc. The functions of neural networks for these application data can include image recognition (e.g., when inputting image data, identifying the animals contained in the image), natural language recognition (e.g., when inputting the user's language, being able to identify the user's intention of speaking, such as whether the user is saying to open a music player), etc.
[0019] The identification process of the above application scenarios receives various application data input by the neural network as tensors, and is realized by the calculation of the neural network. As described above, the calculation of the neural network can be composed of a series of tensor operations, and these tensor operations may be complex geometric transformations of input data of tensors of several orders. These tensor operations are also called operators, and the calculation of the neural network can be converted into a computational graph. The computational graph has a plurality of operators, and between the plurality of operators, they may be connected by lines so as to represent the dependency relationship between the calculations of each operator.
[0020] An artificial intelligence (AI) chip is a chip dedicated to neural network operations, and is mainly a chip designed to accelerate the execution of neural networks. A neural network can be expressed by pure mathematical formulas. According to these mathematical formulas, the neural network can be represented by a computational graph model. The computational graph is a visual representation of these mathematical formulas. The computational graph model can divide one composite operation into a plurality of sub-operations, and each sub-operation is called an operator (abbreviated as Operator, Op).
[0021] When a large amount of intermediate data is generated by the calculation of a neural network and stored in a Dynamic Random Access Memory (DRAM), the overall performance deteriorates due to large latency and insufficient bandwidth. By adding an L2 cache, this problem can be alleviated. The advantages are that it is invisible to programming, so it does not affect programming, and it can reduce latency. However, due to the problems of the access address and access occasion of the L2 cache, the cache miss rate is high, and when the locality is poor, it is also inconvenient to hide the data access time.
[0022] Using the on-chip SRAM of an artificial intelligence chip to store the data required for neural network calculation and the generated data, and through software, spontaneously control the data flow and hide the data transfer time between SRAM and DRAM in a preset form. Since the data access pattern of neural network calculation is flexible, if the access flexibility of SRAM is insufficient, some operators will complete using multiple calculation processes.
[0023] If data calculation and data transfer are completely combined and these operations are started through hardware, the calculation form of the operator will be hardwareized and there will be no room for software adjustment.
[0024] Therefore, a form that can access the on-chip SRAM of the artificial intelligence chip more flexibly is required.
[0025] Figure 1 shows an example of a computational graph in a neural network used to process or identify image data.
[0026] For example, a tensor with image data (e.g., chromaticity values of pixels) is input into the computational graph of an example shown in FIG. 1. The computational graph shows only some of the operators for ease of viewing by the reader. During the operation of the computational graph, first, the tensor is calculated by the Transpose operator, and then, one branch is calculated by the Reshape operator, and the other branch is calculated by the Fully connected operator.
[0027] Assuming that the tensor is first input into the Transpose operator, the Transpose operator is a tensor operation that does not change the numerical values in the input tensor. The action of the Transpose operator is to change the order of the dimensions (axis) of the array. For example, for a two-dimensional array, swapping the order of the two dimensions results in matrix transposition. The Transpose operator can be applied in the case of more dimensions. The input parameter of the Transpose operator is the order of the dimensions of the output array, and the numbers are counted from 0. Taking the input tensor of the Transpose operator as a two-dimensional matrix [[1, 2, 3], [4, 5, 6], [7, 8, 9], [10, 11, 12]], or,
Number
Number
[0028] As can be seen, the Transpose operator changes the dimension ordering, that is, it changes the shape of the tensor, but does not change the numerical values in the tensor. For example, it is still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. Of course, the Transpose operator may not change the shape of the tensor. For example, a 3*3 matrix is still a 3*3 matrix after transposition, and the numerical values in the tensor have not changed, but the ordering of the numerical values in the transposed matrix is different.
[0029] And the tensor calculated by the above Transpose operator
Number
[0030] The specific operation of the Reshape operator is to change the shape attribute of the tensor, and an m*n matrix a can be arranged into a matrix b of size i*j. For example, the Reshape operator (Reshape(A, 2, 6), where A is the input tensor) changes the shape of the above tensor
Number
Number
[0031] As can be seen from the above, the Reshape operator changes the shape of the tensor, but does not change the numerical values in the tensor. For example, it is still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12.
[0032] The Fully connected operator (also called the Full Connection operator) can be regarded as a special convolutional layer or as a tensor product. It is an operation that extracts features using the entire tensor input as a feature map. That is, it performs a linear transformation from one feature space to another, and the output tensor is the weighted sum of the input tensor. For example, the Fully connected operator multiplies the input tensor (the output tensor of the Transpose operator)
Number
Number
[0033] In the prior art, when using an artificial intelligence chip to perform the calculation process from the Transpose operator to the Fully connected operator in the computational graph shown in FIG. 1, the artificial intelligence chip first needs to read the input tensor of the Transpose operator into the memory in the chip's storage unit, and use the input tensor
Number
Number
Number
[0034] Then, the operation of the Fully connected operator is executed. The input tensor (the output tensor of the Transpose operator)
Number
Number
Number
[0035] That is, for the calculation of the Transpose operator and the subsequent calculation of the Fully connected operator, each hardware unit of the chip needs to cooperate to perform the processes of reading, calculating, storing, re-reading, re-calculating, and re-storing. However, in this way, the calculation efficiency of the entire process is very low, and the flexibility is also very low.
[0036] The present disclosure proposes a form of an artificial intelligence processor chip that can flexibly access data. By utilizing the software configuration and related parameters of the artificial intelligence processor chip and the read operation of the chip that can flexibly access data, it can replace the operations of several operators and improve the calculation efficiency.
[0037] FIG. 2 shows a schematic diagram of an artificial intelligence processor chip that can flexibly access data according to an embodiment of the present application.
[0038] As shown in FIG. 2, an artificial intelligence processor chip 200 that flexibly accesses data includes a memory 201 that stores tensor data read from outside the processor chip 200. The read tensor data includes a plurality of elements used for tensor operations of operators included in artificial intelligence calculations. The memory 201, and a storage control unit 202 that controls to read elements from the memory and transmit them to a calculation unit 203 based on the tensor operation of the operator. The storage control unit 202 includes an address calculation module 2021. The address calculation module 2021 has an interface that receives parameters set by software. Based on the set parameters, the address calculation module 2021 calculates the address in the memory in a single-read loop or a multi-read loop (nested loop), reads elements from the calculated address, and transmits them to the calculation unit 203. The storage control unit 202, and a calculation unit 203 that performs a tensor operation of the operator on the received elements.
[0039] According to this embodiment, an address calculation module 2021 is set in the storage control unit 202. The address calculation module 2021 has an interface that receives parameters set by software. Based on the set parameters, the address calculation module 2021 calculates the address in the memory in a single-read loop or a multi-read loop (nested loop), reads elements from the calculated address, and can transmit them to the calculation unit 203. In this way, the address in the memory can be flexibly calculated and read by the parameters set by software. That is, the address flexibly calculated in this way may be different from the storage order. Unlike the prior art, it is not necessary to read elements in the stored order, and the user may read them in a new read order set by software according to the parameters.
[0040] In this way, by flexibly calculating the address in the memory by the parameters set by software, the elements in the memory can be flexibly read, and it is not limited to the order of these elements or the address order stored in the memory.
[0041] Combined with the example of FIG. 1, the Transpose operator changes the dimension ordering but does not change the elements in the tensor. For example, it is still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. That is, based on the set parameters, when it is possible to calculate the addresses in the memory in a single-read loop or a multi-read loop (nested loop) and read the elements from the calculated addresses and send them to the calculation unit 203, the user can use software to set a new read order according to the form of the tensor transposed by the Transpose operator from the stored 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12.
[0042] Specifically, for the input tensor [Number] If it is set as, usually, the memory in the storage unit of the chip stores in a continuous manner, that is, stores as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 (for example, the storage addresses are, for example, (hexadecimal) 00000001, 00000002, 00000003, 00000004, 00000005, 00000006, 00000007, 00000008, 00000009, 00000010, 00000011, 00000012). Then, the user can use software to set a new read order according to the parameters, and the address calculation module 2021 can calculate the addresses in the memory based on the set parameters. For example, the order of the addresses in the memory calculated based on the set parameters is 00000001, 00000004, 00000007, 00000010, 00000002, 00000005, 00000008, 00000011, 00000003, 00000006, 00000009, 00000012 respectively, that is, it can be directly replaced by the transposition operation of the Transpose operator according to the address read order.
[0043] In this way, the address in the memory can be flexibly calculated and read according to the parameters set by the software. That is to say, the address calculated in such a flexible manner may be different from the memory order, and unlike the prior art, it is not necessary to read the elements in the stored order, but the user may read them in a new reading order set by the software according to the parameters.
[0044] In one embodiment, the parameters set by the software instruct to replace the tensor operation of the operator in the form of reading elements from the address in the memory based on the tensor operation of the Transpose operator. Combining with the example of FIG. 1, that is, the parameters set by the software instruct to replace the tensor operation of the Transpose operator in the form of reading elements from the address in the memory based on the tensor operation of the Transpose operator.
[0045] In one embodiment, the parameters set by the software include a value representing the address separated between the address of the first element read in the first step of each reading loop and the initial address in the memory of the input tensor, a value representing the number of read steps in one reading loop, and a value representing the stride between each step in the single reading loop. However, here, the stride is similar to the step width / stride in the neural network concept.
[0046] In this embodiment, when there is only one reading loop, the parameters set by the software may notify the start address, how many elements are read in total, and how many addresses are left between each element when reading. That is to say, it may also be three parameters.
[0047] For example, when performing the operation of the Gather operator, the operation of the Gather operator is to select several numerical values as the output tensor from several numerical values in the input tensor. As can be seen from the above, the operation of the Gather operator is a tensor operation that does not change the numerical values in the input tensor. In this case, the input tensor is [1, 2, 3, 4, 5, 6, 7, 8], and the operation of the Gather operator is to select [1, 3, 5, 7] from [1, 2, 3, 4, 5, 6, 7, 8].
[0048] In the prior art, to complete the operation of the Gather operator, first, the chip stores the input tensor [1, 2, 3, 4, 5, 6, 7, 8] in the memory as memory addresses (hexadecimal), for example, 00000001, 00000002, 00000003, 00000004, 00000005, 00000006, 00000007, 00000008 continuously, and then the calculation unit of the chip performs the operation of the Gather operator to obtain [1, 3, 5, 7], and stores [1, 3, 5, 7] as addresses (hexadecimal numbers), for example, 00000009, 00000010, 00000011, 00000012, and then the chip needs to continue the operation of further operators on the result [1, 3, 5, 7] by the calculation unit.
[0049] However, according to this embodiment, the address calculation module directly calculates the addresses of 1, 3, 5, 7 from the input tensor [1, 2, 3, 4, 5, 6, 7, 8] according to the parameters set by software, reads [1, 3, 5, 7] from these addresses, and can execute the operation of further operators on the result [1, 3, 5, 7] by the calculation unit. Specifically, the parameters set by software may include a value 4 representing the number of times to perform the step of reading in one read loop (indicating a total of 4 reads), and a value 2 representing the stride between each step in the single read loop (adding 2 addresses each time reading to perform the next read).
[0050] Therefore, based on these parameters, the address calculation module can calculate the address order 00000001, 00000003, 00000005, 00000007 to be read (starting from the 00000001 address, adding two addresses each time it reads to perform the next read, and performing a total of 4 reads). The calculation unit reads the addresses stored at 00000001, 00000003, 00000005, 00000007, that is, 1, 3, 5, 7, based on the calculated address order.
[0051] In this way, based on the parameters set by software, the calculation unit reads the addresses stored at 00000001, 00000003, 00000005, 00000007, that is, 1, 3, 5, 7, based on the calculated address order, directly replacing the operation of the Gather operator. In the prior art, it is possible to save the time and hardware cost for calculating the operation of the Gather operator, the time and hardware cost for storing the result tensor of the operation of the Gather operator, and the time and hardware cost for reading each element from the address storing the result tensor of the operation of the Gather operator.
[0052] For a more complex and flexible reading mode, there may be multiple loops (nested reading loops). In one embodiment, the parameters set by software may further include a value representing the number of read steps in each layer of the reading loop and a value representing the stride between each step within each layer of the reading loop. Each layer of the reading loop is performed in a nested form from the outside to the inside. In one embodiment, the parameters set by software may further include a value representing the address separated from the address of the first element read in the first step in each layer of the reading loop and the initial address in the memory of the input tensor. In this way, the reading form can be made more flexible.
[0053] In this embodiment, there is a nested loop (nested read loop). The nesting of loops means that the outer loop is executed once, and after the inner loop finishes execution, it enters the second outer loop and the inner loop is executed again. The above outer loop and inner loop are examples of nested double loops. For example, when reading a two-dimensional matrix, the outer loop can control up to which column to read, and the inner loop can control which row value in a column to read. In the C language, a nested statement of multiple for loops can be adopted to execute a nested loop (nested read loop).
[0054] For example, in combination with the example of FIG. 1, an example of a nested double loop is given. For the input tensor
Number
[0055] Instead of performing the operation of this Transpose operator, that is, to read 1, 4, 7, 10, 2, 5, 8, 11, 3, 6, 9, 12 in sequence, a double read loop can be set. The first-level read loop is set as the inner loop, and the second-level read loop is set as the outer loop. That is, in the present disclosure, for the nesting of multiple loops, the larger the number of levels, the outer loop, and the smaller the number of levels, the inner loop.
[0056] The parameters set by software include a value representing the number of read steps in the read loop for each layer (the number of steps in the read loop for the second layer (outer) is 3, that is, it traverses 3 columns of the tensor, and the number of steps in the read loop for the first layer (inner) is 4, that is, it traverses all rows in one column), and a value representing the stride between each step in the read loop for each layer (the stride between each step in the read loop for the second layer is 1, that is, in the first step, it starts reading from 00000001, and in the second step, it adds 1 address to 00000001 and starts reading from 00000002, which is equivalent to; the stride between each step in the read loop for the second layer is 3, that is, in the first step, it starts reading from 00000001, and in the second step, it adds 3 addresses to 00000001 and starts reading from 00000004, which is equivalent to). The read loop for each layer is performed in a nested form from the outside to the inside.
[0057] For example, the pseudo-code for the nested double read loop is as follows.
[0058] For loop_1 from 1 to loop_1_cnt_max For loop_0 from 1 to loop_0_cnt_max Here, loop_1 represents the read loop for the second layer, and loop_0 represents the read loop for the first layer. Based on the above parameter settings, loop_1_cnt_max is 3, and loop_0_cnt_max is 4.
[0059] In this way, next, the address calculation module reads from memory addresses 00000001, 00000002, 00000003, 00000004, 00000005, 00000006, 00000007, 00000008, 00000009, 00000010, 00000011, 00000012 where 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 are stored, based on these parameters set by this software.
[0060] Specifically, assuming that based on the parameters set by this software, in the second-level reading loop, the initial address read by the compiler is calculated. In this example, for instance, the initial address where the input tensor is stored in memory is 00000001, and in the second-level reading loop, it starts reading from 00000001. In the first step of the second-level reading loop, all steps of the first-level reading loop are executed. That is, in the 4 steps of the first-level reading loop, reading is performed 4 times according to the address with a stride of 4. That is, in the first step of the second-level reading loop, since the address calculation module calculates the reading addresses 00000001, 00000004, 00000007, 00000010, the elements read in the first step of the second-level reading loop are 1, 4, 7, 10 stored at addresses 00000001, 00000004, 00000007, 00000010 respectively.
[0061] The stride between the initial address of the second step and the initial address of the first step of the second-level read loop is 1. That is, this time 1 address is added to 00000001, that is, reading starts from the 00000002 address. In the second step of the second-level read loop, all steps of the first-level read loop are executed. That is, in the 4 steps of the first-level read loop, reading is performed 4 times according to 4 addresses with a stride, that is, in the second step of the second-level read loop, since the address calculation module calculates the read addresses 00000002, 00000005, 00000008, 00000011, the elements read in the second step of the second-level read loop are 2, 5, 8, 11 stored at addresses 00000002, 00000005, 00000008, 00000011 respectively.
[0062] The stride between the initial address of the third step and the initial address of the second step of the second-level read loop is 1. That is, this time 1 address is added to 00000002, that is, reading starts from the 00000003 address. In the third step of the second-level read loop, all steps of the first-level read loop are executed. That is, in the 4 steps of the first-level read loop, reading is performed 4 times according to 4 addresses with a stride, that is, in the third step of the second-level read loop, since the address calculation module calculates the read addresses 00000003, 00000006, 00000009, 00000012, the elements read in the third step of the second-level read loop are 3, 6, 9, 12 stored at addresses 00000003, 00000006, 00000009, 00000012 respectively.
[0063] Thus, based on the parameters set by these software, the order of the addresses calculated by the address calculation module is 00000001, 00000004, 00000007, 00000010, 00000002, 00000005, 00000008, 00000011, 00000003, 00000006, 00000009, 00000012. Therefore, the elements read in order from these addresses are 1, 4, 7, 10, 2, 5, 8, 11, 3, 6, 9, 12.
[0064] Based on the parameters set by these software, since the number of steps of the second-layer reading loop is 3, the address calculation module can stop after executing the address calculation of the second-layer reading loop and the first-layer reading loop within each second-layer reading loop three times.
[0065] In this embodiment, the situation of the loop is set only once by the parameters set by the software. For example, the number of steps of the second-layer reading loop can be set to 1, that is, all steps of the first-layer reading loop can be executed only once.
[0066] Also, the way of taking the values of the above specific parameters is only an example. That is, the meaning is directly represented by numbers, but this is not a limitation. The meaning can also be represented by other numbers or non-numeric content. For example, (since the chip usually counts from 0, etc.), a stride or the number of steps being 1 can be represented by 0, or a stride or the number of steps being 1 can be represented by A. As long as the chip can infer that the value represents the corresponding meaning.
[0067] In this way, by using parameters set by software and calculating addresses in cooperation with an address calculation module, various tensor operations can be directly replaced. By setting the address calculation process for nested loops with more than one loop, the address reading calculated by the address calculation module can be made more flexible, and the order address where each numerical value of the input tensor is stored is not limited to itself. The operation of the operator can be directly replaced, saving the time and hardware cost for calculating the operation of the operator in the prior art, the time and hardware cost for storing the result tensor of the operator's operation, and the time and hardware cost for reading each element from the address storing the result tensor of the operator's operation. Thereby, the calculation delay is reduced, the hardware calculation cost is reduced, and the execution efficiency of the artificial intelligence chip is improved.
[0068] Therefore, by using parameters set by software and calculating addresses in cooperation with an address calculation module, an address sequence that can be arranged more flexibly can be calculated, and furthermore, various tensor operations can be replaced. In one embodiment, the tensor operation to be replaced may be a tensor operation that does not change the numerical values in the input tensor. The address calculation module calculates the address in the memory to be read each time it reads in a single read loop or a multiple read loop (nested loop) based on the parameters received through the interface, thereby replacing the tensor operation with the read operation.
[0069] In one embodiment, the tensor operations may include operations of transpose operator, reshape operator, broadcast operator, gather operator, reverse operator, concat operator, cast operator, etc. Of course, operations of other types of operators can be flexibly realized by the address calculation module and the parameters set by software in the embodiments of the present application.
[0070] Not only can specific tensor operations be replaced by using parameters set by software and calculating addresses in cooperation with an address calculation module, but also addresses loaded in various flexible forms can be calculated, thereby realizing a flexible loading function that exceeds the fixed limitations of loading according to the order of the hardware itself.
[0071] Hereinafter, specific applications in actual scenarios will be described in combination with examples of specific chip hardware and examples of parameters.
[0072] FIG. 3 shows an exploded schematic diagram of an artificial intelligence processor chip that flexibly accesses data according to an embodiment of the present application.
[0073] FIG. 3 shows the members inside the processing engine (PE) 300 of the artificial intelligence processor chip and the parameters used.
[0074] The processing engine 300 may include an arrangement unit 301, a calculation unit 302, and a storage unit 303. The storage control unit 304 is used to set the calculation unit 302 and the storage unit 303. The calculation unit 302 is mainly used for convolution / matrix calculation / vector calculation, etc. The storage unit 303 is used for the interaction between internal data and external data of the processing engine 300 and for accessing the data of the calculation unit 302 in the processing engine 300, and includes an on-chip SRAM memory 3031 (the size is, for example, 8MB, but this is not a limitation) and an access control module 3032.
[0075] The processing engine 300 further includes a storage control unit 304, and the storage control unit 304 realizes the following specific functions.
[0076] The sram_read function reads data from SRAM 3031, transmits it to the calculation unit 302, and calculates the read data. For example, it is used to perform tensor operations such as convolution / matrix calculation / vector calculation. The sram_write function is used to obtain calculation result data from the calculation unit 302, write it, and store it in SRAM 3031. The sram_upload function is used to transfer the data stored in SRAM 3031 to the outside of the processing engine 300 (for example, another processing engine or DRAM). The sram_download function is used to download the data outside the processing engine 300 (data from another processing engine or DRAM) to SRAM 3031.
[0077] That is, the sram_upload function and the sram_download function are used to perform data interaction with the devices outside the processing engine 300. The sram_read function and the sram_write function are used for data interaction between the calculation unit 302 and the storage unit 303 inside the processing engine 300.
[0078] SRAM 3031 is a shared memory in the processing engine 300. Its size is not limited to 8MB and is mainly used to store intermediate data (including data for calculation and calculation results) in the processing engine 300. SRAM 3031 can be divided into multiple banks to improve the overall data bandwidth.
[0079] The crossbar 3041 is a complete interconnection structure between the internal memory control access interface of the processing engine 300 and the SRAM multi - bank. The crossbar 3041 and the access control module 3032 of the storage unit 303 control the address calculated by the address calculation module 3042 to read elements from the corresponding address in the SRAM memory 3031 in the storage unit 303.
[0080] As will be understood hereinafter, the calculation pipeline (calculation unit 302) and the data pipeline (storage unit 303) of the processing engine 300 are configured separately. To complete the calculation of an operator, a plurality of modules need to cooperate. For example, in the calculation of one convolution, an sram_read is configured to input feature data to the calculation unit 302, an sram_read is configured to input weights to the calculation unit 302, and the calculation unit 302 performs matrix convolution calculation and needs to configure an sram_write to output the calculation result to the SRAM memory 3031 in the storage unit 303. In this form, the selection of the calculation form can be made more flexible.
[0081] Specifically, the memory control unit 304 is configured to control to read data from the memory 3031 based on the tensor operation of the operator and transmit it to the calculation unit. The memory control unit 304 includes an address calculation module 3042. The address calculation module 3042 has an interface for receiving a parameter data_noc set by software, and based on the set parameter, calculates the address in the memory 3031 in a single-read loop or a multi-read loop (nested loop), reads elements from the calculated address, and transmits them to the calculation unit 302.
[0082] Most of the calculations of the artificial intelligence AI have relatively regular addressing. For example, calculations such as matrix multiplication, fully connected, and convolution all read data from addresses that store tensors regularly. Therefore, it is conceivable to realize various complex address calculations depending on the form of the parameters set by software.
[0083] In one embodiment, when only one read loop is set, the parameters set by software include a value representing the number of read steps in one read loop and a value representing the stride between each step in the single-read loop.
[0084] In one embodiment, when setting the nest of multiple read loops, the parameters set by software include a value representing the number of read steps in each hierarchical read loop and a value representing the stride between each step within each hierarchical read loop, and each hierarchical read loop is performed in a nested form from the outside to the inside.
[0085] By setting the nests of the above-mentioned multiple read loops, the addresses stored in the same tensor or the same address can be read multiple times, realizing various complex address calculations, and obtaining the flexibility to read addresses more flexibly.
[0086] For example, when adopting the form of an eight-fold loop nest (loop_7~Loop_0 in order from the outer loop to the inner loop), its pseudo code is shown below.
[0087] For loop_7 from 1 to loop_7_cnt_max For loop_6 from 1 to loop_6_cnt_max For Loop_5 from 1 to loop_5_cnt_max For Loop_4 from 1 to loop_4_cnt_max For Loop_3 from 1 to loop_3_cnt_max For Loop_2 from 1 to loop_2_cnt_max For Loop_1 from 1 to loop_1_cnt_max For Loop_0 from 1 to loop_0_cnt_max In the register, the following parameters related to address specification are set for a value loop_xx_cnt_max (xx represents the hierarchical number of the read loop) representing the number of read steps in each hierarchical read loop and a value jump_xx_addr (xx represents the hierarchical number of the read loop) representing the stride between each step within each hierarchical read loop.
[0088] When the above eight reading loops are executed, they are executed as follows. First, execute a total of loop_7_cnt_max steps of the Loop7 hierarchical loop. In each step of the Loop7 hierarchical loop, execute a total of loop_6_cnt_max steps of the Loop6 hierarchical loop. In each step of the Loop6 hierarchical loop, execute a total of loop_5_cnt_max steps of the Loop5 hierarchical loop. In each step of the Loop5 hierarchical loop, execute a total of loop_4_cnt_max steps of the Loop4 hierarchical loop. In each step of the Loop4 hierarchical loop, execute a total of loop_3_cnt_max steps of the Loop3 hierarchical loop. In each step of the Loop3 hierarchical loop, execute a total of loop_2_cnt_max steps of the Loop2 hierarchical loop. In each step of the Loop2 hierarchical loop, execute a total of loop_1_cnt_max steps of the Loop1 hierarchical loop. In each step of the Loop1 hierarchical loop, execute a total of loop_0_cnt_max steps of the Loop0 hierarchical loop. As can be seen from the above, the innermost Loop0 hierarchical loop executes a total of loop_0_cnt_max * loop_1_cnt_max * loop_2_cnt_max * loop_3_cnt_max * loop_4_cnt_max * loop_5_cnt_max * loop_6_cnt_max * loop_7_cnt_max steps, and the Loop1 hierarchical loop in the upper layer executes a total of loop_1_cnt_max * loop_2_cnt_max * loop_3_cnt_max * loop_4_cnt_max * loop_5_cnt_max * loop_6_cnt_max * loop_7_cnt_max steps. Thus, the outermost Loop7 layer loop executes a total of loop_7_cnt_max steps.
[0089] As will be understood hereinafter, the eight - layer loop (nested loop) set as described above has a multiplicative relationship from the innermost to the outermost from an abstract perspective. That is, the number of times of looping inside is equal to the product of the number of times of looping outside and itself.
[0090] As shown in FIG. 4, taking the example of reading with a nested double - loop for a tensor in a two - dimensional space. FIG. 4 shows an example of reading an input tensor according to an embodiment of the present application with a nested double - loop.
[0091] The input tensor is
Number
[0092] Then, base_address is the starting address, which may be pre - calculated by the compiler. Generally, it is the initial address at the address where the input tensor is stored in the SRAM (that is, the storage position of the first element, which is address 0 in this example). The software sets a parameter loop_0_cnt_max = 4, indicating that the number of steps of the first - layer (innermost) read loop is 4 or the size of the loop is 4. The software sets another parameter jump0_addr = 2, indicating that the stride between each step in the first - layer read loop is 2. The software sets another parameter loop_1_cnt_max = 2, indicating that the number of steps of the second - layer (outermost) read loop is 2 or the size of the loop is 2. The software sets another parameter jump1_addr = 9, indicating that the stride between each step in the second - layer read loop is 9.
[0093] The suspected code is as follows.
[0094] For loop_1 from 1 to 2 For loop_0 from 1 to 4 For each reading loop, a corresponding counter loop_xx_cnt (xx represents the hierarchical number of the reading loop) is set. For example, loop_0_cnt is incremented in the order from 1 to 4, and for example, loop_1_cnt is incremented in the order from 1 to 2.
[0095] Based on the above-set parameters, together with FIG. 5 (FIG. 5 shows a schematic diagram for calculating an address based on parameters set by software according to an embodiment of the present application, and sram_addr represents addresses 0 - 15 stored in SRAM), the address calculation module calculates the address as follows.
[0096] First, execute the first step (a total of two steps) of the second-layer (outer) reading loop loop_1. Starting from the base_address initial address 0, execute the first step (for example, 0_0 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 0 from address 0. The parameter jump0_addr = 2 indicates that the stride between each step in the first-layer reading loop is 2. So, execute the second step (for example, 0_1 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 2 from address 2 (address 0 + 2). Execute the third step (for example, 0_2 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 4 from address 4 (address 2 + 2). Execute the fourth step (for example, 0_3 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 6 from address 6 (address 4 + 2). Then, the execution of the four steps of the first-layer (internal memory) reading loop loop_0 ends.
[0097] Next, execute the second step (a total of two steps) of the second-layer (outer) reading loop loop_1. The parameter jump1_addr = 9 indicates that the stride between each step within the second-layer reading loop is 9. Thus, as shown by the arrow in Figure 5, starting from the address 9 obtained at the base_address initial address 0 + 9, execute the first step (for example, 1_0 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 9 from address 9. The parameter jump0_addr = 2 indicates that the stride between each step within the first-layer reading loop is 2. Therefore, execute the second step (for example, 1_1 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, read element 11 from address 11 (address 9 + 2), execute the third step (for example, 1_2 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, read element 13 from address 13 (address 11 + 2), execute the fourth step (for example, 1_3 in Figure 5, a total of four steps) of the first-layer (internal memory) reading loop loop_0, and read element 15 from address 15 (address 13 + 2). Then, the execution of the four steps of the first-layer (internal memory) reading loop loop_0 ends.
[0098] So far, the execution of the total two steps of the second-layer (outer) reading loop loop_1 has both ended, and the address calculation and address reading of the address calculation module end. In this way, the order of reading elements is 0 - 2 - 4 - 6 - 9 - 11 - 13 - 15.
[0099] Of course, in one embodiment, the parameters set by software may further include a value representing an address separated from the address of the first element read by the single reading loop and the initial address in the memory of the input tensor. In this way, the initial address of the reading loop for each layer can be set more flexibly.
[0100] As can be seen from the above, by software to set the corresponding parameters of the double-reading loop nest, it is possible to realize the flexible reading of elements from the SRAM address.
[0101] Similarly, a mechanism of a loop nest for more than double readings can be adopted, which is not limited here.
[0102] When setting the nested eight-layer loop as described above, from an abstract perspective, it is a multiplicative relationship from the inside to the outside. That is, the number of times looped inside is equal to the product of the number of times looped outside. This is a completely aligned form, that is, each layer of the nested loop is regular, and the read address is also regular.
[0103] However, in some special scenarios, it may not be completely aligned. For example, the order of address reading may follow the first rule at one address and the second rule different from the first rule at another address. Two different loop nest forms can exist. In this case, the parameters set by software may include conditions for the parameters, and the parameters take different values when the conditions are met and when the conditions are not met.
[0104] In one embodiment, the parameter is a value representing the number of reading steps in a specific first-layer reading loop, and the condition is that several steps have been performed in another reading loop outside the specific first layer. For example, one setting can be added to the specified reading loop and bound to another reading loop to solve the unaligned situation. For example, for the Loop5 reading loop, there are settings loop_1_cnt_max0 and loop_1_cnt_max1 for two different step numbers of the Loop1 reading loop, which are linked to the loop_5 reading loop.
[0105] For loop_7 from 1 to loop_7_cnt_max For loop_6 from 1 to loop_6_cnt_max For Loop5 from 1 to loop_5_cnt_max For Loop4 from 1 to loop_4_cnt_max For Loop3 from 1 to loop_3_cnt_max For Loop2 from 1 to loop_2_cnt_max { if(loop_5_cnt==loop_5_cnt_max) For Loop1 from 1 to loop_1_cnt_max_1 else For Loop1 from 1 to loop_1_cnt_max_0 } For Loop0 from 1 to loop_0_cnt_max That is, the parameters set by the software may include a condition (loop_5_cnt == loop_5_cnt_max) for a parameter (the value of the number of steps to read in the read loop in the Loop1 layer, that is, loop_1_cnt_max). The parameter loop_1_cnt_max takes different values loop_1_cnt_max_1 and loop_1_cnt_max_0 when the condition (performing up to the last step in the read loop of the Loop5 layer outside the read loop of the Loop1 layer, that is, loop_5_cnt == loop_5_cnt_max) is satisfied and when the condition (loop_5_cnt <> loop_5_cnt_max) is not satisfied.
[0106] That is, when executing up to Loop5, in each step of Loop5, Loop1 is executed at least once. When Loop5 has not been executed up to the last step, that is, when loop_5_cnt < > loop_5_cnt_max, the number of steps of Loop1 executed in the step of that Loop5 is always loop_1_cnt_max_0. When Loop5 is executed up to the last step, that is, when loop_5_cnt == loop_5_cnt_max, the number of steps of Loop1 executed in the step of that Loop5 is always loop_1_cnt_max_1.
[0107] In this way, the form of calculating the address for reading can be made more flexible.
[0108] An example of reading using a nested triple loop when the tensor in the two-dimensional space is not completely aligned as described above is given. As shown in FIG. 6, FIG. 6 shows an example of performing a triple loop nested read that is not completely aligned for the input tensor according to the embodiment of the present application.
[0109] The input tensor is
Number
[0110] As will be understood hereinafter, when reading 0-8-1-9-2-10-3-11-4-12, it is a completely aligned reading form and there is one rule. However, the reading form when reading 16-17-18-19-20 and the reading form when reading 0-8-1-9-2-10-3-11-4-12 are not completely aligned and there is another rule. In this case, consider using the parameters set by the software to achieve an uncompletely aligned loop nest reading.
[0111] Specifically, set a triple loop nest (Loop2, Loop1, Loop0), calculate the address to be read, and realize the above reading order.
[0112] The parameters set by the software may be as follows. The number of steps loop_2_cnt_max of the outermost Loop2 is 2, the stride jump2_addr is 16, the number of steps loop_1_cnt_max of the inner layer Loop1 is 5, the stride jump1_addr is 1, the number of steps loop_0_cnt_max_0 of the innermost Loop0 is 2, the stride jump0_addr is 8, and it is set to bind the specified Loop2 and Loop0. The set condition is that when performing until the last step in Loop2 (loop_2_cnt == loop_2_cnt_max), the number of steps of Loop0 changes from loop_0_cnt_max_0, that is, 2 to loop_0_cnt_max_1 and the value is 1.
[0113] The reason for setting the maximum number of steps loop_2_cnt_max of the outermost Loop2 to 2 is that in the first step, it is executed in the order of reading 0-8-1-9-2-10-3-11-4-12, while in the second step, it is executed in the order of reading 16-17-18-19-20. Since the execution order and rules are different in the two steps, it is necessary to consider how to combine the inner reading loop Loop1 and Loop0 in the second step to achieve different reading orders.
[0114] Specifically, based on the parameters set by the above software, this nested triple loop is executed, and the obtained calculated addresses are as follows.
[0115] First, execute the first step (a total of 2 steps) of the outermost Loop2, and in this first step, execute 5 steps of Loop1.
[0116] In the first step of Loop1, execute all steps of Loop0, that is, execute 2 steps from 0. Since the stride of each step is 8, first read 0, and then add 8 addresses to read 8. In this way, read 0-8 in 2 steps.
[0117] In the second step of Loop1, the stride is 1, that is, execute all steps of Loop0 from 1, that is, execute 2 steps from 1. Since the stride of each step is 8, first read 1, and then add 8 addresses to read 9. In this way, read 1-9 in 2 steps.
[0118] In the third step of Loop1, the stride is 1, that is, execute all steps of Loop0 from 2, that is, execute 2 steps from 2. Since the stride of each step is 8, first read 2, and then add 8 addresses to read 10. In this way, read 2-10 in 2 steps.
[0119] In the fourth step of Loop1, the stride is 1, that is, all steps of Loop0 are executed starting from 3, that is, two steps are executed starting from 3. Since the stride of each step is 8, first, 3 is read, and then 8 addresses are added to read 11. In this way, 3 - 11 is read in two steps.
[0120] In the fifth step of Loop1, the stride is 1, that is, all steps of Loop0 are executed starting from 4, that is, two steps are executed starting from 4. Since the stride of each step is 8, first, 4 is read, and then 8 addresses are added to read 12. In this way, 4 - 12 is read in two steps.
[0121] Then, the second step (a total of two steps) of the outermost Loop2 is executed, adding a stride of 16 to the initial address 9, that is, reading from 25.
[0122] In this case, the condition loop_2_cnt == loop_2_cnt_max is satisfied. Therefore, the number of steps of Loop0 is loop_0_cnt_max_1, that is, it is 1 step instead of 2 steps. In the second step, five steps of Loop1 are executed, and in each step of Loop1, a loop for reading Loop0 with a step number of 1 is executed.
[0123] Specifically, in the first step of Loop1, all steps of Loop0 are executed, that is, one step is executed starting from 16, indicating that only one read is performed. Then, since the stride 8 is not used, only 16 is read.
[0124] In the second step of Loop1, the stride is 1, that is, since one step of Loop0 is executed starting from 16 + 1 = 17, that is, 17 is read.
[0125] In the third step of Loop1, the stride is 1. That is, since 17 + 1 = 18, one step of Loop0 is executed, that is, 18 is read.
[0126] In the fourth step of Loop1, the stride is 1. That is, since 18 + 1 = 19, one step of Loop0 is executed, that is, 19 is read.
[0127] In the fifth step of Loop1, the stride is 1. That is, since 19 + 1 = 20, one step of Loop0 is executed, that is, 20 is read.
[0128] In this way, by performing a triple nested read loop set by software parameters, a complex address read order of 0 - 8 - 1 - 9 - 2 - 10 - 3 - 11 - 4 - 12 - 16 - 17 - 18 - 19 - 20 is realized.
[0129] Of course, in the above, the set conditions are that in a read loop of another layer outside a specific layer, these several steps are performed, and taking as an example that different numbers of steps are respectively set in the read loop of a specific layer when the conditions are met and not met, but this application is not limited to this, and considering other conditions and changes in other parameters that meet the conditions, a more complex address read order can be flexibly realized.
[0130] Therefore, based on the parameters set by the software of this application, the address in the memory is flexibly calculated, and elements are read from the calculated address and sent to the calculation unit, thereby realizing a flexible address reading form, improving the calculation efficiency of the artificial intelligence processor chip, and reducing costs. In some cases, a specific tensor operation itself in artificial intelligence calculation can also be replaced, thereby simplifying the operation of the operator.
[0131] To calculate the specific address of the read loop nest, in one embodiment, in the first-level or each-level read loop, based on the number of times the current read step has been performed and the respective stride of the first-level or each-level read loop, the currently read address is calculated.
[0132] Specifically, for the address calculation of which address in the SRAM will ultimately be read, based on the parameters set by the above software, the address calculation unit can calculate the address Address to be read each time by determining the position of one point in the multi-dimensional (read loop) space coordinate system.
[0133]
Number
[0134]
Number
[0135] That is, when the actual address calculation module calculates an address, if only the number of read steps currently performed in the first-level or each-level read loop and the stride of each first-level or each-level read loop are known, the address to be currently read can be obtained.
[0136] FIG. 7 shows a schematic diagram of the internal structure of the SRAM according to the embodiment of the present application.
[0137] In addition, the SRAM may be divided into a plurality of banks. In order to increase the data writing and reading speeds, data can be written from the outside and directly placed in different banks, and the read data, calculation data during the process, and result data can be directly placed in different banks. The address formatting mode in the SRAM is configurable. By default, the most significant bit of the address may be used to distinguish different banks, or interleaving of other granularities may be performed in a form where an address hash can be set. In terms of hardware design, finally, a multi-bit bank selection signal bank_sel is generated to select different SRAM banks sram_bank (such as sram_bank0 - sram_bank3) for the data of each port port0, port1, port2, port3, etc.
[0138] Multi-port access can use handshake signals, and the data pipeline supports backpressure (when the incoming traffic is larger than the outgoing traffic, backpressure is required, or when the downstream stage is not ready, if the current stage performs data transfer, it needs to have more backpressure than the upstream stage. In that case, the upstream stage needs to hold the data as it is until the handshake is successful in order to update the data). The memory control unit 304 includes a crossbar (fully cross-connected interconnection path) structure that can access multiple banks of SRAM in parallel and has separate read and write functions, corresponding to a form in which two-stage crossbars are cascade-connected, which alleviates the winding problem in hardware implementation. The single-port SRAM used for the lower-layer memory saves power consumption and area. The crossbar structure can access multiple banks where the read data, the calculated data during the process, and the final result data are respectively stored in parallel or simultaneously, thereby increasing the read and write speeds and improving the execution efficiency of the artificial intelligence chip.
[0139] In this way, by using the parameters set by software and cooperating with the address calculation module to calculate the address, various tensor operations can be directly replaced, and by setting the address calculation process of more than one loop nest, the address reading calculated by the address calculation module can be made more flexible, and the order address where each numerical value of the input tensor is stored is not limited to itself. The operation of the operator can be directly replaced, saving the time and hardware cost for calculating the operation of the operator, the time and hardware cost for storing the result tensor of the operator's operation, and the time and hardware cost for reading each element from the address storing the result tensor of the operator's operation in the prior art, thereby reducing the calculation delay, reducing the hardware calculation cost, and increasing the execution efficiency of the artificial intelligence chip.
[0140] To summarize the above, by adopting a configuration in which the data pipeline and the calculation pipeline are separated, complete control of the software is realized, and flexibility is maximized. By using the addressing form of the multiple-loading loop, various and complex address access patterns can be realized. By adopting the form of the multiple asymmetric loop, a non-aligned configuration can be realized, and more complex address access patterns can be realized. By dividing the on-chip shared memory into multiple banks, the software form is used to maintain the formation pattern between the banks. By separating the data pipeline and the calculation pipeline, the flexibility of data access is ensured. In addition, the data transfer and the calculation pipeline can be realized and hidden from each other, and the parallel effect of different modules can be realized. By reasonably dividing the intermediate data of the compiler, different types of data are stored in different banks of the SRAM. For example, different types of data can be accessed simultaneously. When it is necessary to access these data simultaneously, these data can be read in parallel from different banks of the SRAM at the same time, improving the efficiency. After the calculation is started, the utilization rate of the media access control address (MAC) of the convolution calculation can reach almost 100%.
[0141] FIG. 8 shows a flowchart of a method for flexibly accessing data in an artificial intelligence processor chip according to an embodiment of the present application.
[0142] As shown in FIG. 8, a method 800 for flexibly accessing data in an artificial intelligence processor chip includes a step 801 of storing tensor data read from outside the processor chip in a memory in the artificial intelligence processor chip, where the read tensor data includes a plurality of elements used for tensor operations of operators involved in artificial intelligence calculations; and a step 802 of controlling to read elements from the memory and send them to a calculation unit based on tensor operations of the operators, where based on parameters set by received software, an address in the memory is calculated in a single-read loop or a multi-read loop (nested loop), and elements are read from the calculated address and sent to the calculation unit in the artificial intelligence processor chip; and a step 803 of performing tensor operations of the operators by the calculation unit on the received elements.
[0143] In this way, by flexibly calculating the address in the memory according to parameters set by software, elements in the memory can be flexibly read, and it is not limited to the order or address order of these elements stored in the memory.
[0144] In one embodiment, the parameters set by software include a value representing the quantity of elements to be read for tensor data read in a single-read loop and a value representing the stride between each step in the single-read loop.
[0145] In one embodiment, the parameters set by software include a value representing the number of read steps in each hierarchical read loop and a value representing the stride between each step in each hierarchical read loop, and each hierarchical read loop is performed in a nested form from the outside to the inside.
[0146] In one embodiment, the parameters set by software include a value representing the address separated from the address of the first element read in a single-read loop and the initial address in the memory of the input tensor.
[0147] In one embodiment, the parameters set by software include conditions for the parameters, and the parameters take different values when the conditions are met and when the conditions are not met.
[0148] In one embodiment, the parameter is a value representing the number of read steps in a read loop of a specific first layer, and the condition is that in a read loop of another layer outside the specific first layer, some of these steps have been performed.
[0149] In one embodiment, the method 800 further includes calculating the currently read address based on the number of read steps currently performed in the read loop of the first layer or each layer and the respective strides of the read loops of the first layer or each layer.
[0150] In this way, the form of calculating the address for reading can be made more flexible.
[0151] In one embodiment, the parameter set by software instructs to replace the tensor operation of the operator by a form of reading elements from an address in memory based on the tensor operation of the operator.
[0152] In one embodiment, the tensor operation is a tensor operation that does not change the numerical values in the input tensor, and based on the parameters received by the interface, in a single read loop or a multiple read loop (nested loop), calculates the address in memory to be read each time it is read, and replaces the tensor operation with the read.
[0153] In one embodiment, the tensor operation includes at least one of the operations of the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator.
[0154] In this way, the calculation unit, based on the calculated address order, reads the addresses stored in the calculated addresses according to the parameters set by software, and directly replaces the operations of several tensor operators. In the prior art, the time and hardware cost for calculating the operations of these tensor operators, the time and hardware cost for storing the result tensors of the operations of these tensor operators, and the time and hardware cost for reading each element from the addresses storing the result tensors of the operations of these tensor operators can be saved.
[0155] In one embodiment, the memory is divided into a plurality of banks each storing data that can be accessed in parallel, and the method further includes a crossbar with separated read and write functions for accessing the data stored in the plurality of banks of the memory in parallel.
[0156] In this way, by using the parameters set by software and calculating addresses, various tensor operations can be directly replaced, and by setting the address calculation process of more than one nested loop, the calculated address reading can be made more flexible, and it is not limited to the sequential addresses where each numerical value of the input tensor is stored. The operations of the operators can be directly replaced, and in the prior art, the time and hardware cost for calculating the operations of the operators, the time and hardware cost for storing the result tensors of the operations of the operators, and the time and hardware cost for reading each element from the addresses storing the result tensors of the operations of the operators can be saved, thereby reducing the calculation delay, reducing the hardware calculation cost, and improving the execution efficiency of the artificial intelligence chip.
[0157] FIG. 9 shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application.
[0158] The electronic device may include a processor (H1) and a storage medium (H2) coupled to the processor (H1) and storing computer-executable instructions that, when executed by the processor, perform the steps of each method of the embodiments of the present application.
[0159] The processor (H1) may include, for example, but is not limited to, one or more processors or microprocessors.
[0160] The storage medium (H2) may include, for example, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy (registered trademark) disks, solid state hard disks, removable disks, CD-ROMs, DVD-ROMs, Blu-ray disks, etc.).
[0161] In addition, the electronic device may further include a data bus (H3), an input / output (I / O) bus (H4), a display (H5), and input / output devices (H6) (such as keyboards, mice, speakers, etc.).
[0162] The processor (H1) can communicate with external devices (such as H5 and H6) via a wired or wireless network (not shown) through the I / O bus (H4).
[0163] The storage medium (H2) can further store at least one computer-executable instruction that, when executed by the processor (H1), executes each function and / or method step in the embodiments described in this technology.
[0164] In one embodiment, the at least one computer-executable instruction may be compiled or configured as a software product, and when one or more computer-executable instructions are executed by the processor, they execute each function and / or method step in the embodiments described in this technology.
[0165] FIG. 10 shows a schematic diagram of a non - transitory computer - readable storage medium according to an embodiment of the present disclosure.
[0166] As shown in FIG. 10, instructions are stored in the computer - readable storage medium 1020, and the instructions are, for example, computer - readable instructions 1010. When the computer - readable instructions 1010 are executed by a processor, the above - described methods can be executed with reference to the above - described methods. The computer - readable storage medium includes, but is not limited to, for example, volatile memory and / or non - volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or high - speed cache memory. The non - volatile memory may include, for example, read - only memory (ROM), hard disk, flash, etc. For example, the computer - readable storage medium 1020 can be connected to a computing device such as a computer, and when the computing device executes the computer - readable instructions 1010 stored in the computer - readable storage medium 1020, it can execute various methods as described above.
[0167] The present application provides the following items.
[0168] Item 1. An artificial intelligence processor chip for flexibly accessing data, A memory for storing tensor data read from outside the processor chip, the read tensor data including a plurality of elements used for tensor operations of operators included in artificial intelligence calculations, and the memory, A storage control unit that controls to read elements from the memory and transmit them to a calculation unit based on the tensor operation of the operator. The storage control unit includes an address calculation module having an interface that receives parameters set by software. The address calculation module calculates an address in the memory in a single read loop or a multiple read loop (nested loop) based on the parameters received by the interface, reads an element from the calculated address, and transmits it to the calculation unit. A calculation unit that performs a tensor operation of the operator using the received elements. A processor chip.
[0169] Item 2. The parameters set by the software include a value representing the quantity of elements to be read for the tensor data read in a single read loop and a value representing the stride between each step in the single read loop. Alternatively, the parameters set by the software include a value representing the number of read steps in each hierarchical read loop and a value representing the stride between each step in each hierarchical read loop. The read loop for each hierarchy is performed in a nested form from the outside to the inside. The processor chip according to Item 1.
[0170] Item 3. The parameters set by the software include a value representing an address separated from the address of the first element read in a single read loop and the initial address in the memory of the input tensor. The processor chip according to Item 1.
[0171] Item 4. The parameters set by the software include conditions for the parameters, and the parameters take different values when the conditions are met and when the conditions are not met. The processor chip according to Item 1.
[0172] Item 5. The parameter is a value representing the number of read steps in a read loop of a specific single layer, and the condition is that these several steps have been performed in a read loop of another layer outside the specific single layer. The processor chip according to Item 4.
[0173] Item 6. The address calculation module calculates the currently read address based on the number of read steps currently performed in the read loop of a single layer or each layer, and the respective stride of the read loop of a single layer or each layer. The processor chip according to any one of Items 2 - 5.
[0174] Item 7. The parameter set by the software instructs to replace the tensor operation of the operator by the form of reading elements from the address in the memory based on the tensor operation of the operator. The processor chip according to Item 1.
[0175] Item 8. The tensor operation is a tensor operation that does not change the numerical values in the input tensor. The address calculation module replaces the tensor operation in the read by calculating the address in the memory that is read each time in a single - read loop or a multi - read loop (nested loop) based on the parameter received by the interface. The processor chip according to Item 7.
[0176] Item 9. The tensor operation includes at least one of the operations of a transpose operator, a reshape operator, a broadcast operator, a gather operator, a reverse operator, a concat operator, and a cast operator. The memory is divided into a plurality of banks each storing data that can be accessed in parallel, and the storage control unit includes a crossbar with separated read and write functions for accessing the data stored in the plurality of banks of the memory in parallel. The processor chip according to Item 8.
[0177] Item 10. A method for flexibly accessing data in an artificial intelligence processor chip, comprising: The memory in the artificial intelligence processor chip stores tensor data read from outside the processor chip, and the read tensor data includes a plurality of elements used for tensor operations of operators included in artificial intelligence calculations; Based on the tensor operation of the operator, controlling to read elements from the memory and send them to a calculation unit, including calculating an address in the memory in a single-read loop or a multi-read loop (nested loop) based on parameters set by received software, reading elements from the calculated address, and sending them to the artificial intelligence processor chip; The calculation unit performs tensor operations of the operator using the received elements; A method comprising the above.
[0178] Item 11. The parameters set by the software include a value representing the quantity of elements read for the read tensor data in a single-read loop and a value representing the stride between each step in the single-read loop; Alternatively, the parameters set by the software include a value representing the number of read steps in each hierarchical read loop and a value representing the stride between each step in each hierarchical read loop, and each hierarchical read loop is performed in a nested form from the outside to the inside. The method according to Item 10.
[0179] Item 12. The method according to Item 10, wherein the parameters set by the software include a value representing an address separated from the address of the first element read in a single-read loop and the initial address in the memory of the input tensor.
[0180] Item 13. The method according to item 10, wherein the parameter set by the software includes a condition for the parameter, and the parameter takes different values depending on whether the condition is satisfied or not.
[0181] Item 14. The method according to item 13, wherein the parameter is a value representing the number of read steps in a read loop of a specific first layer, and the condition is that several steps have been performed in a read loop of another layer outside the specific first layer.
[0182] Item 15. The method according to any one of items 11 - 14, further comprising calculating the currently read address based on the number of read steps currently performed in the read loop of one layer or each layer and the respective strides of the read loops of one layer or each layer.
[0183] Item 16. The method according to item 10, wherein the parameter set by the software instructs to replace the tensor operation of the operator by a form of reading elements from an address in the memory based on the tensor operation of the operator.
[0184] Item 17. The tensor operation is a tensor operation that does not change the numerical values in the input tensor, and based on the parameter received by the interface, in a single read loop or a multiple read loop (nested loop), the tensor operation is replaced by the read by calculating the address in the memory to be read each time. The method according to item 16.
[0185] Item 18. The tensor operation includes at least one of an operation of a transpose operator, an operation of a reshape operator, an operation of a broadcast operator, an operation of a gather operator, an operation of a reverse operator, an operation of a concat operator, and an operation of a cast operator. The memory is divided into a plurality of banks each storing data that can be accessed in parallel. The method further includes accessing in parallel the data stored in the plurality of banks of the memory via a crossbar in which a read function and a write function are separated, the method according to Item 17.
[0186] Item 19. An electronic device, a memory storing instructions, a processor that reads instructions in the memory and executes the method according to any one of Items 10-18, the electronic device.
[0187] Item 20. A non-transitory storage medium storing instructions, wherein when the instructions are read by a processor, the method according to any one of Items 10-18 is executed by the processor, the non-transitory storage medium.
[0188] It should be noted that the above specific embodiments are examples and not limitations. A person skilled in the art can, based on the concept of the present application, integrate and combine several steps and devices from the various embodiments separately described above to achieve the effects of the present application. Such integrated and combined embodiments are also included in the present application, and the integration and combination are not described one by one here.
[0189] It should be noted that the advantages, merits, effects, etc. mentioned in the present disclosure are examples and not limitations. It is not considered that these advantages, merits, effects, etc. are required for each embodiment of the present application. Also, the specific details of the above disclosure are not limitations, but only for the purpose of illustration and for ease of understanding. The above details do not limit that the present application must be implemented using the above specific details.
[0190] The block diagrams of the devices, apparatuses, instruments, and systems according to the present disclosure are used only as exemplary examples, and are not intended to require or imply that they need to be connected, arranged, and configured as shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, instruments, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," and "having" are open words meaning "including but not limited to," and can be used interchangeably therewith. As used herein, the terms "or" and "and" mean the term "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. As used herein, the term "for example" means the phrase "for example, without limitation" and can be used interchangeably therewith.
[0191] The step flowcharts and the above-described methods in the present disclosure are described only as exemplary examples, and are not intended to require or imply that the steps of each embodiment must be performed in the given order. As those skilled in the art will recognize, the order of the steps in the above embodiments can be performed in any order. Words such as "then," "and," and "next" are not intended to limit the order of the steps. These words are used only to guide the reader in reading the description of these methods. Also, for example, any reference to a single element using the articles "a," "one," or "the" is not to be construed as limiting that element to being singular.
[0192] Also, the steps and apparatuses in each embodiment of this specification are not limited to a certain embodiment. In fact, based on the concepts of this specification, new embodiments can be conceived by combining some of the related steps and some of the apparatuses in each embodiment of this specification, and these new embodiments are also included within the scope of this specification.
[0193] Each operation of the above-described method can be performed by any suitable means capable of performing the corresponding function. Such means can include, but are not limited to, hardware circuits, application-specific integrated circuits (ASICs), or processors, and can include various hardware and / or software components and / or modules.
[0194] Various exemplary logic blocks, modules, and circuits described herein can be implemented or illustrated using a general-purpose processor, a digital signal processor (DSP), an ASIC, a field programmable gate array signal (FPGA), or other programmable logic device (PLD), discrete gates or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a microprocessor cooperating with a DSP core, or any other such configuration.
[0195] The steps of a method or algorithm described in connection with the present disclosure can be directly incorporated into hardware, software modules executed by a processor, or a combination of the two. The software modules can be present in any form of tangible storage medium. Some examples of storage media that can be used include random access memory (RAM), read only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, and the like. The storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium can be integral to the processor. The software modules can be a single instruction, or many instructions, and can be distributed over several different code segments, different programs, and across multiple storage media.
[0196] The methods disclosed herein include operations that implement the described methods. The methods and / or operations can be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of operations is specified, the order and / or use of specific operations can be modified without departing from the scope of the claims.
[0197] The above functions can be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as instructions on a suitable computer-readable medium. The storage medium may be any available appropriate medium accessible by a computer. By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser.
[0198] Accordingly, a computer program product can perform the operations provided herein. For example, such a computer program product can be a computer-readable tangible medium having tangible (and / or encoded) instructions thereon that can be executed by a processor to perform the operations described herein. The computer program product can include packaging materials.
[0199] Software or instructions may be transmitted by a transmission medium. For example, software may be transmitted by a transmission medium such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, RF, or microwave from a website, server, or other remote source.
[0200] Furthermore, modules and / or other suitable means for implementing the methods and techniques described herein may be downloaded by a user terminal and / or a base station when appropriate and / or obtained in other ways. For example, such devices may be coupled to a server to facilitate the transmission of means for implementing the methods described herein. Alternatively, the various methods described herein may be provided via a storage member (e.g., a physical storage medium such as RAM, ROM, CD, or floppy disk) so that a user terminal and / or a base station can obtain the various methods while being connected to the device or providing a storage member to the device. Furthermore, the methods and techniques described herein can be utilized in other suitable techniques of the device.
[0201] Other examples and embodiments are within the scope and spirit of the present disclosure and the appended claims. For example, due to the nature of software, the above-described functions can be realized using software executed by a processor, hardware, firmware, hardwire, or any combination thereof. Features for realizing the functions can also be physically located at various positions, including being distributed such that some parts of the functions are realized at different physical positions. Furthermore, as used in this specification and the claims, for example, the enumeration of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A, B, and C), so the "or" used in the enumeration of terms starting with "at least one" means a disjunctive enumeration. Additionally, the term "exemplary" does not mean that the examples described are preferred or superior to other examples.
[0202] Various changes, substitutions, and modifications to the technology described herein can be made without departing from the teachings of the disclosed technology as defined by the appended claims. Also, the claims of the present disclosure are not limited to the specific forms of the above processes, machines, manufactures, configurations of events, means, methods, and acts. Using the corresponding forms described herein, processes, machines, manufactures, configurations of events, means, methods, or acts that are presently existing or later developed and that are substantially the same in function or produce substantially the same result can be realized. Accordingly, the appended claims include such processes, machines, manufactures, configurations of events, means, methods, or acts within their scope.
[0203] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the present application. Various changes to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the present application. Accordingly, the present application is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0204] The above description has been provided for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed. Although several exemplary embodiments and examples have been described above, those skilled in the art will recognize some variations, modifications, alterations, additions, and subcombinations thereof.
Claims
1. An artificial intelligence processor chip that flexibly accesses data, A memory for storing tensor data read from outside the processor chip, wherein the read tensor data includes a plurality of elements used for tensor operations of operators included in artificial intelligence calculations, a memory, A storage control unit that controls to read elements from the memory and send them to a calculation unit based on the tensor operation of the operator, wherein the storage control unit includes an address calculation module having an interface for receiving parameters set by software, and the address calculation module is based on the parameters received by the interface. A storage control unit that calculates an address in the memory in a single read loop or a nested multiple read loop, reads an element from the calculated address, and sends it to the calculation unit, A calculation unit that performs a tensor operation of the operator with the received elements, and a processor chip that flexibly accesses data.
2. The parameters set by the software include a value representing the quantity of elements read for the read tensor data in a single read loop and a value representing the stride between each step in the single read loop, or, The parameters set by the software include a value representing the number of read steps in each hierarchical read loop and a value representing the stride between each step in each hierarchical read loop, and each hierarchical read loop is performed in a nested form from the outside to the inside. The processor chip according to claim 1.
3. The parameters set by the software include a value representing an address separated from the address of the first element read in a single read loop and the initial address in the memory of the input tensor. The processor chip according to claim 1.
4. The parameters set by the software include conditions for the parameters, and the parameters take different values when the conditions are satisfied and when the conditions are not satisfied. The processor chip according to claim 1.
5. The parameter is a value representing the number of read steps in a read loop of a specific one layer, and the condition is the number of read steps performed in a read loop of another layer outside the specific one layer. The processor chip according to claim 4.
6. The address calculation module calculates the currently read address based on the number of read steps currently performed in the read loop of one layer or each layer and the respective strides of the read loops of one layer or each layer. The processor chip according to any one of claims 2 to 5.
7. The parameter set by the software instructs to replace the tensor operation of the operator by a form of reading elements from the address in the memory based on the tensor operation of the operator. The processor chip according to claim 1.
8. The tensor operation is a tensor operation that does not change the numerical values in the input tensor. The address calculation module replaces the tensor operation in the read by calculating the address in the memory that is read each time in the nesting of the single read loop or the multiple read loops based on the parameter received by the interface. The processor chip according to claim 7.
9. The tensor operation includes at least one of the operations of the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator. The memory is divided into a plurality of banks that respectively store data that can be accessed in parallel. The storage control unit includes a crossbar in which the read function and the write function are separated in order to access the data stored in the plurality of banks of the memory in parallel. The processor chip according to claim 8.
10. A method for flexibly accessing data in an artificial intelligence processor chip, wherein the memory in the artificial intelligence processor chip stores tensor data read from outside the processor chip, and the read tensor data includes a plurality of elements used for the tensor operation of the operator included in the calculation by the artificial intelligence. Controlling to read elements from the memory and send them to a calculation unit based on the tensor operation of the operator, including calculating the addresses in the memory in a single-read loop or a multi-read loop (nested loops) based on parameters set by the received software, reading elements from the calculated addresses, and sending them to the calculation unit of the artificial intelligence processor chip. The method for flexibly accessing data in an artificial intelligence processor chip includes the calculation unit performing a tensor operation of the operator on the received elements. **Claim 11** The parameters set by the software include a value representing the quantity of elements to be read for the tensor data read in a single-read loop and a value representing the stride between each step in the single-read loop, or The parameters set by the software include a value representing the number of read steps in the read loop for each layer and a value representing the stride between each step in the read loop for each layer. The read loops for each layer are performed in a nested form from the outside to the inside. The method according to claim 10. **Claim 12** The parameters set by the software include a value representing the address separated from the address of the first element read in a single-read loop and the initial address in the memory of the input tensor. The method according to claim 10. **Claim 13** The parameters set by the software include conditions for the parameters, and the parameters take different values when the conditions are met and when the conditions are not met. The method according to claim 10. **Claim 14** The parameter is a value representing the number of read steps in the read loop of a specific one layer, and the condition is the number of read steps performed in the read loop of another layer outside the specific one layer. The method according to claim 13. **Claim 15** The method according to any one of claims 11 to 14 further includes calculating the currently read address based on the number of read steps currently performed in the read loop of one layer or each layer and the respective strides of the read loop of one layer or each layer. **Claim 16** The parameter set by the software is the method according to claim 10, which instructs to replace the tensor operation of the operator by a form of reading elements from an address in the memory based on the tensor operation of the operator.
17. The tensor operation is a tensor operation that does not change the numerical values in the input tensor, and based on the parameter received by the interface, calculates the address in the memory that is read each time it is read in the nested loop for single reading or multiple readings, so as to replace the tensor operation with the reading. The method according to claim 16.
18. The tensor operation includes at least one of the operations of the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator. The memory is divided into a plurality of banks that respectively store data that can be accessed in parallel. The method further includes accessing in parallel the data stored in the plurality of banks of the memory via a crossbar in which the read function and the write function are separated. The method according to claim 17.
19. A memory for storing instructions, An electronic device including a processor that reads instructions in the memory and executes the method according to any one of claims 10 to 18.
20. A non-transitory storage medium in which instructions are stored, When the instructions are read by a processor, a non-transitory storage medium in which the method according to any one of claims 10 to 18 is executed by the processor.
Citation Information
Patent Citations
Arithmetic processing unit and image processing unit
JP2020017179A
Accessing Data in Multidimensional Tensors Using Adders
JP2020521198A
Neural Network Processor
JP2022514680A
Address Generation Unit Using Nested Loops To Scan Multi-Dimensional Data Structures
US20100145992A1