Chips, methods, devices, and media for flexible access to data
Patent Information
- Application Number
- JP2025501618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-07-12
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-07-12
AI Technical Summary
【0009】 このように、ソフトウェアにより設定されるパラメータによって、メモリにおけるアドレスを柔軟に算出することにより、メモリにおける要素を柔軟に読み込み、メモリに記憶されるこれらの要素の順序又はアドレス順序に限られない。
Smart Images

Figure 0007909684000021 
Figure 0007909684000022 
Figure 0007909684000023
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to an artificial intelligence processor chip that flexibly accesses data, a method for flexibly accessing data in an artificial intelligence processor chip, an electronic device, and a non-temporary storage medium.
[0002] This application claims priority to Chinese Patent Application No. 202210836577.0, filed on July 15, 2022, and hereby incorporates by reference in its entirety the content disclosed in the above Chinese patent application as part of this application.
Background Art
[0003] For artificial intelligence (AI) chips, the design of the data pipeline is crucial to achieving better performance, fully utilizing computing power, and increasing the utilization of media access control bit addresses (MACs). During neural network computation, the data provided and stored in the input and output sections of the computation unit are diverse, which also determines different chip memory and computation architectures. For example, a Graphics Processing Unit (GPU), a parallel computing processor, employs a multi-stage high-speed cache system with a memory hierarchy consisting of L1 cache, shared memory, register groups, L2 cache, and external memory DRAM. These memory hierarchies are primarily divided to reduce data transport delays and improve bandwidth. The L1 cache is typically divided into L1D cache and L1I cache, used to store data and instructions respectively, and generally, each processor core has its own corresponding L1 cache. The size of L1 caches varies from 16k to 64k, while L2 caches are always private caches that do not distinguish between instructions and data. Generally, each processor core has its own L2 cache, and the size of L2 caches varies from 256k to 1M. For example, L1 caches are the fastest but have the smallest space, L2 caches are slower but have the largest space, and external DRAM has the largest space but the slowest speed. Therefore, by storing frequently accessed data from DRAM to the L1 cache, the delay of transporting data from external DRAM to internal memory each time it is accessed can be reduced, improving the efficiency of data processing. However, in order to ensure versatility and flexibility, the data pipeline has a certain degree of redundancy. For example, it is necessary to retrieve data from a register each time a calculation is performed and finally store the data in a register, which results in high power consumption.
[0004] Some AI chips can achieve high efficiency by customizing their pipelines, but at the cost of losing flexibility and potentially becoming unusable if the network structure is modified. Others solve bandwidth and latency issues by adding large on-chip buffers, but the data access to static random access memory (SRAM) is hardware-initiated, meaning that computation and storage are coupled via hardware. Because of this lack of policy flexibility, efficiency may decrease in some scenarios, and software intervention may be impossible. [Overview of the project] [Means for solving the problem]
[0005] In accordance with one aspect of this application, an artificial intelligence processor chip for flexible access to data is provided, comprising: a memory for storing tensor data read from outside the processor chip, wherein the read tensor data includes a plurality of elements used in tensor operations of operators to be calculated by artificial intelligence; a storage control unit that controls the reading of elements from the memory and transmission to a calculation unit based on the tensor operations of the operators; the storage control unit includes an address calculation module having an interface for receiving parameters set by software; the address calculation module includes a storage control unit that calculates an address in the memory in a single read loop or multiple loops (nested read loops) based on the parameters received by the interface, reads elements from the calculated address and transmits them to the calculation unit; and a calculation unit that performs tensor operations of the operators on the received elements.
[0006] In another embodiment, a method for flexibly accessing data in an artificial intelligence processor chip is provided, the method comprising: a memory in the artificial intelligence processor chip storing tensor data read from outside the processor chip, the read tensor data including a plurality of elements used in tensor operations of operators to be calculated by artificial intelligence; controlling the reading of elements from the memory and transmitting them to a calculation unit based on the tensor operations of the operators; calculating an address in the memory in a single read loop or a multiple read loop (nested loop) based on parameters set by the receiving software, reading elements from the calculated address and transmitting them to the artificial intelligence processor chip; and the calculation unit performing tensor operations of the operators on the received elements.
[0007] In another embodiment, the present invention provides an electronic device comprising a memory for storing instructions and a processor for reading instructions from the memory and executing the methods of each embodiment of this application.
[0008] In another embodiment, a non-temporary storage medium in which instructions are stored, When the aforementioned instruction is read by the processor, the method of each embodiment of this application is executed by the processor. [Effects of the Invention]
[0009] In this way, by flexibly calculating memory addresses using parameters set by the software, elements in memory can be read flexibly, and the order of these elements stored in memory is not limited to their order or address order. [Brief explanation of the drawing]
[0010] To more clearly illustrate the embodiments of this disclosure or the technical concepts in the prior art, the drawings that may be used in the description of the embodiments or the prior art are briefly described below. Obviously, the drawings in the following description are only a few embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these without any creative effort. [Figure 1] Figure 1 shows an example of a computational diagram in a neural network used to process or identify image data. [Figure 2] Figure 2 shows a schematic diagram of an artificial intelligence processor chip that flexibly accesses data according to an embodiment of this application. [Figure 3] Figure 3 shows a schematic exploded view of an artificial intelligence processor chip that flexibly accesses data, according to an embodiment of this application. [Figure 4] Figure 4 shows an example of reading an input tensor using a nested double loop according to the embodiment of this application. [Figure 5] Figure 5 shows a schematic diagram illustrating how an address is calculated based on parameters set by software, according to an embodiment of this application. [Figure 6] Figure 6 shows an example of reading an input tensor using a nested triple loop that is incompletely aligned, according to an embodiment of this application. [Figure 7] Figure 7 shows a schematic diagram of the internal structure of the SRAM according to the embodiment of this application. [Figure 8] Figure 8 shows a flowchart illustrating a method for flexibly accessing data in an artificial intelligence processor chip according to an embodiment of this application. [Figure 9] Figure 9 shows a block diagram of an exemplary electronic device suitable for realizing the embodiment of this application. [Figure 10] Figure 10 shows a schematic diagram of a non-temporary computer-readable storage medium according to an embodiment of the present disclosure. [Modes for carrying out the invention]
[0011] Examples of the present application are illustrated in the drawings with detailed reference to specific embodiments of the present application. While the present application is described in conjunction with specific embodiments, it should be understood that it is not intended to be limited to the embodiments described. Conversely, changes, modifications, and equivalents included within the spirit and scope of the present application as defined by the attached claims should be overridden. It should be noted that the method steps described herein can be implemented by any functional block or functional configuration, and any functional block or functional configuration can be implemented as a physical entity, a logical entity, or a combination of both.
[0012] It is understood that before using any of the technical methods disclosed in each embodiment of this disclosure, the user must be notified in an appropriate manner, in accordance with applicable laws and regulations, of the types of personal information related to this disclosure, the scope of use, and the usage scenarios, and the user's permission must be obtained.
[0013] For example, in response to a user's voluntary request, the system sends the user information and requests that the user perform an operation that explicitly informs the user that it will require the acquisition and use of the user's personal information. Therefore, the user can autonomously choose, based on the information presented, whether or not to provide personal information to electronic devices, applications, servers, or software or hardware such as storage media that perform the operation of the proposed technology of this disclosure.
[0014] As an optional and non-limiting embodiment, the form in which information is sent to the user in response to a voluntary request from the user may be, for example, a pop-up that can present the information as text. The pop-up window may also include a selection control that allows the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0015] The above notice and user permission acquisition process are merely general in nature and do not limit the embodiments of this disclosure. Other forms that comply with applicable laws and regulations may also be applied to the embodiments of this disclosure.
[0016] In addition, the data related to this technical solution (including, but not limited to, the data itself, data acquisition or use) should comply with the requirements of corresponding laws, regulations and related provisions.
[0017] The identification process of the above application scenario can be realized by receiving various application data input by a neural network as tensors and going through the calculations of the neural network. Currently, neural networks and machine learning systems use tensors as the basic data structure. The core of the concept of a tensor is a data container, and since most of the data contained is always numerical data, it is a digital container. The specific numerical values in a tensor can be application data, for example, including image data, natural language data, etc.
[0018] For example, a scalar is a 0 - order tensor, such as 2, 3, 5. In a specific application scenario, for example, in image data, 2 represents, for example, the grayscale value of one pixel in the image data, 3 represents, for example, the grayscale value of one pixel in the image data, and 5 represents, for example, the grayscale value of one pixel in the image data, etc. For example, a vector is a 1 - order tensor, such as [0, 3, 20]. A matrix is a 2 - order tensor, for example,
Number
[0019] The identification process of the above application scenarios receives various application data input by the neural network as tensors, and is realized by the calculation of the neural network. As described above, the calculation of the neural network can be composed of a series of tensor operations, and these tensor operations may be complex geometric transformations of input data of tensors of several orders. These tensor operations are also called operators, and the calculation of the neural network can be converted into a computational graph. The computational graph has a plurality of operators, and between the plurality of operators, they may be connected by lines so as to represent the dependency relationship between the calculations of each operator.
[0020] An artificial intelligence (AI) chip is a chip dedicated to neural network operations, and is mainly a chip designed to accelerate the execution of neural networks. A neural network can be expressed by pure mathematical formulas. According to these mathematical formulas, the neural network can be represented by a computational graph model. The computational graph is a visual representation of these mathematical formulas. The computational graph model can divide one composite operation into multiple sub-operations, and each sub-operation is called an operator (abbreviated as Operator, Op).
[0021] When neural network computations generate a large amount of intermediate data and store it in Dynamic Random Access Memory (DRAM), the overall performance degrades due to high latency and insufficient bandwidth. Adding an L2 cache can mitigate this problem, offering the advantages of being invisible to the programmer, thus reducing latency. However, issues with L2 cache access addresses and access occasions lead to a high cache miss rate, and hiding data access times becomes inconvenient when locality is poor.
[0022] By using on-chip SRAM on the artificial intelligence chip, the data required and generated during neural network computation can be stored, and the data flow can be spontaneously controlled through software, hiding the data transport time between SRAM and DRAM in a pre-configured manner. Because the data access patterns for neural network computation are flexible, if the access flexibility of the SRAM is insufficient, some operators may be completed using multiple computation processes.
[0023] By fully integrating data computation and data transport, and initiating these operations via hardware, the computational form of the operators becomes hardware-based, leaving no room for software adjustment.
[0024] Therefore, a form that allows for flexible access via on-chip SRAM on the artificial intelligence chip is necessary.
[0025] Figure 1 shows an example of a computational diagram in a neural network used to process or identify image data.
[0026] For example, a tensor containing image data (e.g., pixel chromaticity values) is input into the example calculation diagram shown in Figure 1. This calculation diagram shows only some of the operators for the reader's convenience. During the calculation in this diagram, the tensor is first recalculated using the Transpose operator, then one branch is recalculated using the Reshape operator, and the other branch is recalculated using the Fully connected operator.
[0027] If the tensor in question is first input to the Transpose operator, the Transpose operator is a tensor operation that does not change the numerical values in the input tensor. The action of the Transpose operator is to change the order of the dimensions (axis) of the array. For example, swapping the order of two dimensions in a two-dimensional array is matrix transposition. The Transpose operator can be applied to more dimensions. The input parameter of the Transpose operator is the order of the dimensions of the output array, and the indices are counted from 0. The input tensor of the Transpose operator can be, for example, a two-dimensional matrix [[1, 2, 3], [4, 5, 6], [7, 8, 9], [10, 11, 12]], or
number
number
[0028] As can be seen, the Transpose operator changes the order of dimensions, that is, changes the shape of the tensor, but does not change the numerical values in the tensor. For example, they are still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. Of course, the Transpose operator may not change the shape of the tensor. For example, a 3x3 matrix remains a 3x3 matrix after transposition, and the numerical values in the tensor remain unchanged, but the order of the numerical values in the transposed matrix is different.
[0029] Then, the tensor calculated using the Transpose operator described above.
number
[0030] The specific operation of the Reshape operator is to change the shape attribute of a tensor, so that an m*n matrix a can be arranged into a matrix b of size i*j. For example, the Reshape operator (Reshape(A,2, 6), where A is the input tensor) can be used to transform the above tensor
number
number
[0031] As can be seen from the above, the Reshape operator changes the shape of the tensor but does not change the numerical values in the tensor; for example, they are still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12.
[0032] The Fully Connected Operator (also called the Full Connection Operator) can be viewed as a special convolutional layer or as a tensor product, and is an operation that extracts features from the entire tensor input as a feature diagram. That is, it performs a linear transformation from one feature space to another, and the output tensor is a weighted sum of the input tensors. For example, the Fully Connected Operator takes the input tensor (the output tensor of the Transpose Operator)
number
number
[0033] In conventional technology, when performing the calculation process from the Transpose operator to the Fully connected operator in the calculation diagram shown in Figure 1 using an artificial intelligence chip, the artificial intelligence chip first needs to load the input tensor of the Transpose operator into the memory of the chip's storage unit, and then the input tensor
number
number
number
[0034] Then, the fully connected operator operation is performed. Input tensor (Output tensor of the Transpose operator)
number
number
number
[0035] In other words, the calculation of the Transpose operator and the subsequent Fully Connected operator requires the cooperation of each hardware part of the chip to perform the processes of reading, calculating, storing, rereading, recalculating, and restoring. However, the calculation efficiency of the entire process is very low, and the flexibility is also very low.
[0036] This disclosure proposes a form of artificial intelligence processor chip that provides flexible access to data. By utilizing the software configuration and related parameters of the artificial intelligence processor chip, the read operations of the chip, which provide flexible access to data, can replace the calculations of several operators and improve computational efficiency.
[0037] Figure 2 shows a schematic diagram of an artificial intelligence processor chip that flexibly accesses data according to an embodiment of this application.
[0038] As shown in Figure 2, the artificial intelligence processor chip 200, which has flexible access to data, includes a memory 201 that stores tensor data read from outside the processor chip 200, the read tensor data including multiple elements used in tensor operations of operators to be calculated by artificial intelligence, and a memory control unit 202 that controls the reading of elements from the memory and transmission to the calculation unit 203 based on the tensor operations of operators, and includes an address calculation module 2021, the address calculation module 2021 which has an interface for receiving parameters set by software, the memory control unit 202 which calculates the address in the memory in a single read loop or a multiple read loop (nested loop) based on the set parameters, reads the elements from the calculated address and transmits them to the calculation unit 203, and a calculation unit 203 which performs tensor operations of operators on the received elements.
[0039] In this embodiment, the memory control unit 202 is configured with an address calculation module 2021. The address calculation module 2021 has an interface for receiving parameters set by software, and based on the set parameters, it can calculate the address in memory in a single read loop or a multiple read loop (nested loop), read the element from the calculated address, and transmit it to the calculation unit 203. In this way, the address in memory can be flexibly calculated and read using parameters set by software. In other words, the address calculated in this way may differ from the storage order, and instead of having to read the elements in the order they are stored, as in the prior art, the elements may be read in a new read order set by the user using parameters in the software.
[0040] In this way, by flexibly calculating memory addresses using parameters set by the software, elements in memory can be read flexibly, and are not limited to the order or address order of these elements stored in memory.
[0041] As illustrated in the example in Figure 1, the Transpose operator changes the dimensional sorting order but does not change the elements in the tensor, for example, they are still 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. In other words, if the memory address can be calculated in a single read loop or a multiple read loop (nested loop) based on the set parameters, and the elements can be read from the calculated address and sent to the calculation unit 203, the user can set a new read order in software using parameters to read these elements from the stored 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 according to the form of the tensor transposed by the Transpose operator.
[0042] Specifically, the input tensor is
number
[0043] In this way, memory addresses can be flexibly calculated and read using parameters set by the software. In other words, the addresses calculated in this way do not have to be in the order in which they are stored, and elements do not have to be read in the order in which they are stored, as in conventional technology. Instead, they can be read in a new reading order set by the user using parameters in the software.
[0044] In one embodiment, the parameters set by the software instruct the Transpose operator to be replaced in a manner that reads elements from memory addresses based on the Transpose operator's tensor operation. The example in Figure 1 illustrates this: the parameters set by the software instruct the Transpose operator to be replaced in a manner that reads elements from memory addresses based on the Transpose operator's tensor operation.
[0045] In one embodiment, the parameters set by the software include a value representing the address separated between the address of the first element read in the first step of each hierarchical read loop and the initial address in memory of the input tensor; a value representing the number of read steps in one read loop; and a value representing the stride between each step in a single read loop. Here, the stride is analogous to the step width / stride in the neural network concept.
[0046] In this embodiment, if there is only one read loop, the parameters set by the software may specify the starting address, the total number of elements to read, and the number of addresses to leave between each element after reading. In other words, there may be three parameters.
[0047] For example, when performing the Gather operator operation, the Gather operator selects a number of values from a set of values in the input tensor to be used as the output tensor. As can be seen from the above, the Gather operator operation is a tensor operation that does not change the values in the input tensor. In this case, the input tensor is [1,2,3,4,5,6,7,8], and the Gather operator operation is the selection of [1,3,5,7] from [1,2,3,4,5,6,7,8].
[0048] In conventional technology, in order to complete the Gather operator operation, the chip first stores the input tensor [1,2,3,4,5,6,7,8] in memory as hexadecimal addresses, for example 00000001, 00000002, 00000003, 00000004, 00000005, 00000006, 00000007, 00000008, then the chip's calculation unit performs the Gather operator operation to obtain [1,3,5,7], and stores [1,3,5,7] as hexadecimal addresses, for example 00000009, 000000010, 00000011, 00000012, and then the chip's calculation unit must continue to perform further operator operations on the result [1,3,5,7].
[0049] However, in this embodiment, the address calculation module can directly calculate the addresses of 1, 3, 5, and 7 from the input tensor [1,2,3,4,5,6,7,8] using parameters set by software, read [1,3,5,7] from these addresses, and perform further operator operations on the result [1,3,5,7] by the calculation unit. Specifically, the parameters set by software may include a value of 4 representing the number of reading steps performed in one reading loop (representing a total of 4 readings) and a value of 2 representing the stride between each step in a single reading loop (adding 2 addresses each time a reading is performed before the next reading).
[0050] Therefore, the address calculation module can calculate the order of addresses to be read, 00000001, 00000003, 00000005, and 00000007, based on these parameters (starting with address 00000001, adding two addresses each time it is read, and performing a total of four reads). Based on the calculated address order, the calculation unit reads the addresses stored at 00000001, 00000003, 00000005, and 00000007, namely 1, 3, 5, and 7.
[0051] In this way, the calculation unit reads the addresses stored at 00000001, 00000003, 00000005, and 00000007, i.e., 1, 3, 5, and 7, based on the calculated address order according to the parameters set by the software, and directly replaces the calculation of the Gather operator. This saves the time and hardware cost of calculating the Gather operator, the time and hardware cost of storing the tensor resulting from the Gather operator, and the time and hardware cost of reading each element from the address where the tensor resulting from the Gather operator is stored, compared to the conventional technology.
[0052] For more complex and flexible reading configurations, multiple loops (nested reading loops) may exist. In one embodiment, the parameters set by the software may include a value representing the number of reading steps in the reading loop of each level, and a value representing the stride between each step in the reading loop of each level. The reading loops of each level are performed in a nested configuration from the outside to the inside. In one embodiment, the parameters set by the software may further include a value representing the address separated from the address of the first element read in the first step of the reading loop of each level and the initial address in memory of the input tensor. In this way, the reading configuration can be made more flexible.
[0053] In this embodiment, there are multiple loops (nested reading loops). Loop nesting means that the outer loop is executed once, and after the inner loop finishes executing, the program enters the outer loop a second time and re-executes the inner loop. The outer and inner loops described above are an example of nested double loops. For example, when reading a two-dimensional matrix, the outer loop can control which columns to read, and the inner loop can control which row of a given column to read. In the C language, multiple loops (nested reading loops) can be executed by employing nested for loop statements.
[0054] For example, consider the example in Figure 1 and give an example of a nested double loop. The input tensor is
number
[0055] Instead of using the Transpose operator, a double read loop can be set up to read 1, 4, 7, 10, 2, 5, 8, 11, 3, 6, 9, and 12 in order. The first-level read loop is designated as the inner loop, and the second-level read loop as the outer loop. In other words, in this disclosure, for nested loops, a larger number of levels corresponds to an outer loop, and a smaller number of levels corresponds to an inner loop.
[0056] The parameters set by the software are a value representing the number of steps to read in the read loop of each level (the number of steps in the second level (outer) read loop is 3, i.e., traversing 3 columns of the tensor, and the number of steps in the first level (inner) read loop is 4, i.e., traversing all rows in 1 column), and a value representing the stride between each step in the read loop of each level (the stride between each step in the second level read loop is 1, i.e., in the first step, This may include: starting to read from 00000001, in the second step adding one address to 00000001 and starting to read from 00000002, with a stride of 3 between each step of the second-level read loop, i.e., starting to read from 00000001 in the first step, in the second step adding three addresses to 00000001 and starting to read from 00000004), and the read loops at each level are performed in a nested form from the outside to the inside.
[0057] For example, the following is pseudocode for nested loops for double reading.
[0058] For loop_1 from 1 to loop_1_cnt_max For loop_0 from 1 to loop_0_cnt_max Here, loop_1 represents the second-level reading loop, and loop_0 represents the first-level reading loop. Based on the above parameter settings, loop_1_cnt_max is 3 and loop_0_cnt_max is 4.
[0059] Thus, the address calculation module then reads from the memory addresses 00000001, 00000002, 00000003, 00000004, 00000005, 00000006, 00000007, 00000008, 00000009, 00000010, 00000011, and 00000012, where 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12 are stored, based on these parameters set by the software.
[0060] Specifically, assuming that the initial address read by the compiler is calculated in the second-level read loop based on the parameters set by this software, in this example, for example, the initial address where the input tensor is stored in memory is 00000001, and the second-level read loop starts reading from 00000001. In the first step of the second-level read loop, all the steps of the first-level read loop are executed. In other words, in the four steps of the first-level read loop, four reads are performed according to the addresses with four strides. Therefore, in the first step of the second-level read loop, the address calculation module calculates the read addresses 00000001, 00000004, 00000007, and 00000010. Thus, the elements read in the first step of the second-level read loop are 1, 4, 7, and 10, which are stored at addresses 00000001, 00000004, 00000007, and 00000010, respectively.
[0061] The stride between the initial address of the second step of the second-level read loop and the initial address of the first step is 1. That is, one address is added to 00000001, meaning that reading begins from address 00000002. In the second step of the second-level read loop, all steps of the first-level read loop are executed. That is, in the four steps of the first-level read loop, four reads are performed according to the addresses with a stride of four. That is, in the second step of the second-level read loop, the address calculation module calculates the read addresses 00000002, 00000005, 00000008, and 00000011. Therefore, the elements read in the second step of the second-level read loop are 2, 5, 8, and 11, which are stored at addresses 00000002, 00000005, 00000008, and 00000011, respectively.
[0062] The stride between the initial address of the third step of the second-level read loop and the initial address of the second step is 1, meaning that one address is added to 00000002, and reading begins from address 00000003. In the third step of the second-level read loop, all steps of the first-level read loop are executed. That is, in the fourth step of the first-level read loop, four reads are performed according to the address with a stride of four, meaning that in the third step of the second-level read loop, the address calculation module calculates the read addresses 00000003, 00000006, 00000009, and 00000012, so the elements read in the third step of the second-level read loop are 3, 6, 9, and 12, which are stored at addresses 00000003, 00000006, 00000009, and 00000012, respectively.
[0063] Thus, based on the parameters set by this software, the order of addresses calculated by the address calculation module is 00000001, 00000004, 00000007, 00000010, 00000002, 00000005, 00000008, 00000011, 00000003, 00000006, 00000009, 00000012. Therefore, the elements read in order from these addresses are 1, 4, 7, 10, 2, 5, 8, 11, 3, 6, 9, and 12.
[0064] Based on the parameters set by this software, the number of steps in the second-level read loop is 3, so the address calculation module can stop after performing address calculations for the second-level read loop and the first-level read loop within each second-level read loop three times.
[0065] In this embodiment, the state of a loop that is executed only once can be set by parameters set by the software. For example, the number of steps in the second-level reading loop may be set to 1, meaning that all steps in the first-level reading loop are executed only once.
[0066] Furthermore, the specific parameter values described above are merely examples. In other words, while the meaning is directly represented by a number, this is not the only way. Meaning can be represented by other numbers or non-numerical content. For example, (since chips usually count from 0) a stride or step count of 1 could be represented by 0, or a stride or step count of 1 could be represented by A. The important thing is that the chip can infer that its value represents the corresponding meaning.
[0067] In this way, by using parameters set by software and working together with the address calculation module to calculate addresses, various tensor operations can be directly replaced. Furthermore, by setting up an address calculation process with more than one nested loop, the address reading calculated by the address calculation module can be made more flexible and is not limited to sequential addresses where each numerical value of the input tensor is stored. By directly replacing operator operations, the time and hardware cost of calculating operator operations, the time and hardware cost of storing the result tensor of the operator operations, and the time and hardware cost of reading each element from the address where the result tensor of the operator operations is stored are saved in conventional technology. This reduces calculation delay, lowers hardware calculation costs, and improves the execution efficiency of the artificial intelligence chip.
[0068] Therefore, by using parameters set by the software and working together with the address calculation module to calculate the address, a more flexible address sequence can be calculated, and various tensor operations can be replaced. In one embodiment, the tensor operation that is replaced may be a tensor operation that does not change the numerical value in the input tensor, and the address calculation module replaces the tensor operation during reading by calculating the address in the memory being read each time a read is performed in a single read loop or a multiple read loop (nested loop) based on the parameters received by the interface.
[0069] In one embodiment, tensor operations may include operations on the transpose operator, reshape operator, broadcast operator, gather operator, reverse operator, concat operator, cast operator, and so on. Of course, operations on other types of operators can be flexibly implemented by parameters set by the address calculation module and software of the embodiment of this application.
[0070] Of course, by using parameters set by software and working together with an address calculation module to calculate addresses, it is possible not only to replace specific tensor operations, but also to calculate addresses that are read in various flexible forms. This enables flexible reading functionality that goes beyond the specific limitations imposed by the hardware's own reading order.
[0071] The following describes specific applications in real-world scenarios, along with examples of chip hardware and parameters.
[0072] Figure 3 shows a schematic exploded view of an artificial intelligence processor chip that provides flexible data access according to an embodiment of this application.
[0073] Figure 3 shows the internal components and parameters used in the processing engine (PE) 300 of the artificial intelligence processor chip.
[0074] The processing engine 300 may include a placement unit 301, a calculation unit 302, and a storage unit 303. A storage control unit 304 is used to configure the calculation unit 302 and the storage unit 303. The calculation unit 302 is mainly used for convolution / matrix calculation / vector calculation, etc. The storage unit 303 includes an on-chip SRAM memory 3031 (for example, 8MB in size, but not limited to this) and an access control module 3032 for interaction between the processing engine 300's internal data and external data, and for accessing the data of the calculation unit 302 in the processing engine 300.
[0075] The processing engine 300 further includes a memory control unit 304, which implements the following specific functions.
[0076] The sram_read function reads data from SRAM 3031, sends it to the calculation unit 302, and performs calculations on the read data. For example, it is used to perform tensor operations such as convolution, matrix calculation, and vector calculation. The sram_write function is used to obtain calculation result data from the calculation unit 302, write it, and store it in the SRAM 3031. The sram_upload function is used to transport data stored in the SRAM 3031 to an external location (for example, another processing engine or DRAM) outside of the processing engine 300. The sram_download function is used to download data from outside the processing engine 300 (data from other processing engines or DRAM) to the SRAM 3031.
[0077] In other words, the sram_upload and sram_download functions are used for data interaction with devices outside the processing engine 300. The sram_read and sram_write functions are used for data interaction between the calculation unit 302 and the storage unit 303 inside the processing engine 300.
[0078] SRAM 3031 is a shared memory in the processing engine 300. Its size is not limited to 8MB, and it is primarily used to store intermediate data (including data and calculation results) in the processing engine 300. SRAM 3031 can be divided into multiple banks to improve the overall data bandwidth.
[0079] The crossbar 3041 is a fully interconnected structure between the internal memory control access interface of the processing engine 300 and the SRAM multibank. The crossbar 3041 and the access control module 3032 of the storage unit 303 control the address calculated by the address calculation module 3042 to read an element from that address in the SRAM memory 3031 of the storage unit 303.
[0080] As can be seen, the calculation pipeline (calculation unit 302) and data pipeline (storage unit 303) of the processing engine 300 are configured separately. Multiple modules must cooperate to complete the calculation of operators. For example, in a single convolution calculation, it is necessary to configure sram_read to input feature data to the calculation unit 302, configure sram_read to input weights to the calculation unit 302, configure sram_write for the calculation unit 302 to perform matrix convolution calculation and output the calculation result to the SRAM memory 3031 in the storage unit 303. This configuration allows for greater flexibility in selecting the calculation method.
[0081] Specifically, the memory control unit 304 is configured to control the system to read data from memory 3031 and transmit it to the calculation unit based on the tensor operation of the operator. The memory control unit 304 includes an address calculation module 3042, which has an interface for receiving a parameter data_noc set by software. Based on the set parameter, it calculates the address in memory 3031 in a single read loop or a multiple read loop (nested loop), reads the element from the calculated address, and transmits it to the calculation unit 302.
[0082] Most calculations performed by artificial intelligence (AI) involve relatively regular addressing; for example, matrix multiplication, full connectivity, and convolution all involve reading data from addresses that have been systematically stored as tensors. Therefore, it is conceivable that various complex address calculations can be achieved depending on the form of parameters set by the software.
[0083] In one embodiment, when only one read loop is configured, the parameters set by the software include a value representing the number of read steps in one read loop and a value representing the stride between each step in a single read loop.
[0084] In one embodiment, when setting up nesting of multiple read loops, the parameters set by the software include a value representing the number of read steps in the read loop of each level and a value representing the stride between each step in the read loop of each level, and the read loops of each level are performed in a nested form from the outside to the inside.
[0085] By setting up the nesting of multiple read loops described above, it becomes possible to read addresses stored in the same tensor or the same address multiple times, enabling the calculation of various complex addresses and providing greater flexibility in reading addresses.
[0086] For example, when adopting a nested octave loop configuration (loop_7 to Loop_0 in order from the outer loop to the inner loops), the pseudocode is shown below.
[0087] For loop_7 from 1 to loop_7_cnt_max For loop_6 from 1 to loop_6_cnt_max For Loop_5 from 1 to loop_5_cnt_max For Loop_4 from 1 to loop_4_cnt_max For Loop_3 from 1 to loop_3_cnt_max For Loop_2 from 1 to loop_2_cnt_max For Loop_1 from 1 to loop_1_cnt_max For Loop_0 from 1 to loop_0_cnt_max The register contains the following address-related parameters: loop_xx_cnt_max (where xx represents the level number of the read loop), which represents the number of steps read during the read loop for each level; and jump_xx_addr (where xx represents the level number of the read loop), which represents the stride between each step in the read loop for each level.
[0088] When the above eight-tiered reading loop is executed, it is executed as follows: First, the total loop_7_cnt_max steps of the Loop7 hierarchical loop are executed, and in each step of the Loop7 hierarchical loop, the total loop_6_cnt_max steps of the Loop6 hierarchical loop are executed, in each step of the Loop6 hierarchical loop, the total loop_5_cnt_max steps of the Loop5 hierarchical loop are executed, in each step of the Loop5 hierarchical loop, the total loop_4_cnt_max steps of the Loop4 hierarchical loop are executed, in each step of the Loop4 hierarchical loop, the total loop_3_cnt_max steps of the Loop3 hierarchical loop are executed, in each step of the Loop3 hierarchical loop, the total loop_2_cnt_max steps of the Loop2 hierarchical loop are executed, in each step of the Loop2 hierarchical loop, the total loop_1_cnt_max steps of the Loop1 hierarchical loop are executed, and in each step of the Loop1 hierarchical loop, the total loop_0_cnt_max steps of the Loop0 hierarchical loop are executed. As can be seen from the above, the innermost Loop0 layer loop executes a total of loop_0_cnt_max* loop_1_cnt_max* loop_2_cnt_max* loop_3_cnt_max* loop_4_cnt_max* loop_5_cnt_max* loop_6_cnt_max* loop_7_cnt_max steps, and the Loop1 layer loop above it executes a total of loop_1_cnt_max* loop_2_cnt_max* loop_3_cnt_max* loop_4_cnt_max* loop_5_cnt_max* loop_6_cnt_max* loop_7_cnt_max steps. Thus, the outermost Loop7 layer loop executes a total of loop_7_cnt_max steps.
[0089] As can be seen from this, the octave loops (nested loops) set up as described above have a multiplicative relationship from the inside to the outside when viewed from an abstract perspective. In other words, the number of times the inner loop loops is equal to the product of that number of times the outer loop loops are looped.
[0090] As shown in Figure 4, an example is reading a tensor in two-dimensional space using a nested double loop. Figure 4 shows an example of reading an input tensor using a nested double loop according to the embodiment of this application.
[0091] input tensor
number
[0092] Then, base_address is the starting address, which may be calculated in advance by the compiler, and is generally the initial address at the address where the input tensor is stored in SRAM (i.e., the storage location of the first element, which in this example is address 0). The software sets one parameter loop_0_cnt_max=4 to indicate that the number of steps in the first-level (inner) read loop is 4 or the loop size is 4, the software sets another parameter jump0_addr=2 to indicate that the stride between each step in the first-level read loop is 2, the software sets another parameter loop_1_cnt_max=2 to indicate that the number of steps in the second-level (outer) read loop is 2 or the loop size is 2, and the software sets another parameter jump1_addr=9 to indicate that the stride between each step in the second-level read loop is 9.
[0093] The pseudocode is as follows:
[0094] For loop_1 from 1 to 2 For loop_0 from 1 to 4 A corresponding counter, loop_xx_cnt (where xx represents the hierarchy number of the read loop), is set for each read loop. For example, loop_0_cnt is incremented from 1 to 4, and loop_1_cnt is incremented from 1 to 2.
[0095] Based on the parameters set above, and also referring to Figure 5 (Figure 5 shows a schematic diagram of how the address is calculated based on the parameters set by the software according to this embodiment of the application, where sram_addr represents addresses 0-15 stored in SRAM), the address calculation module calculates the address as follows.
[0096] First, the first step of the second-level (outer) read loop loop_1 (a total of two steps) is executed, starting from the initial address 0 of base_address. Then, the first step of the first-level (internal memory) read loop loop_0 (for example, 0_0 in Figure 5, a total of four steps) is executed, reading element 0 from address 0. The parameter jump0_addr=2 indicates that the stride between each step in the first-level read loop is 2. Therefore, the second step of the first-level (internal memory) read loop loop_0 (for example, 0_1 in Figure 5, a total of 4 steps) is executed, reading element 2 from address 2 (address 0+2). The third step of the first-level (internal memory) read loop loop_0 (for example, 0_2 in Figure 5, a total of 4 steps) is executed, reading element 4 from address 4 (address 2+2). The fourth step of the first-level (internal memory) read loop loop_0 (for example, 0_3 in Figure 5, a total of 4 steps) is executed, reading element 6 from address 6 (address 4+2). Then the execution of the 4 steps of the first-level (internal memory) read loop loop_0 is completed.
[0097] Next, the second step (a total of two steps) of the second-level (outer) read loop loop_1 is executed. The parameter jump1_addr=9 indicates that the stride between each step in the second-level read loop is 9. As shown by the arrow in Figure 5, starting from address 9 obtained at the initial base_address address 0+9, the first step (for example, 1_0 in Figure 5, a total of four steps) of the first-level (internal memory) read loop loop_0 is executed, and element 9 is read from address 9. The parameter jump0_addr=2 indicates that the stride between each step in the first-level read loop is 2. Therefore, the second step of the first-level (internal memory) read loop loop_0 (for example, 1_1 in Figure 5, a total of 4 steps) is executed, reading element 11 from address 11 (address 9+2). The third step of the first-level (internal memory) read loop loop_0 (for example, 1_2 in Figure 5, a total of 4 steps) is executed, reading element 13 from address 13 (address 11+2). The fourth step of the first-level (internal memory) read loop loop_0 (for example, 1_3 in Figure 5, a total of 4 steps) is executed, reading element 15 from address 15 (address 13+2). Then the execution of the 4 steps of the first-level (internal memory) read loop loop_0 is completed.
[0098] Up to this point, the execution of both steps in the second-level (outer) reading loop_1 has finished, and the address calculation and address reading processes of the address calculation module are also complete. Thus, the order in which the elements are read is 0-2-4-6-9-11-13-15.
[0099] Of course, in one embodiment, the parameters set by the software may further include a value representing an address separated from the address of the first element read by the single read loop and the initial address of the input tensor in memory. In this way, the initial addresses of the read loops at each level can be set flexibly.
[0100] As can be seen from the above, by setting the corresponding parameters of the double-read loop nest using software, it is possible to flexibly read elements from SRAM addresses.
[0101] Similarly, a mechanism with more than two nested read loops can be employed, and is not limited to this approach.
[0102] As described above, when we set up nested octaves, from an abstract perspective, there is a multiplicative relationship from the inside to the outside. That is, the number of times the inner loop is looped is equal to the product of that number of times the outer loop is looped. This is a perfectly aligned form, meaning that each level of the nested loop is regular, and the addresses read are also regular.
[0103] However, in some special cases, the order may not be perfectly aligned; for example, the order for reading addresses may follow a first rule for some addresses and a second rule different from the first for other addresses. Two different nested loop configurations can exist. In this case, the parameters set by the software may include conditions on the parameters, and the parameters may take different values depending on whether the condition is met or not.
[0104] In one embodiment, the parameter is a value representing the number of steps to read in a particular one-level read loop, and the condition is that several steps have been reached in a read loop of a different level outside of that particular one level. For example, one setting can be added to a given read loop and bound to another read loop to resolve unaligned situations. For example, for the Loop5 read loop, there are settings loop_1_cnt_max0 and loop_1_cnt_max1 for two different number of steps in the Loop1 read loop, which work in conjunction with the loop_5 read loop.
[0105] For loop_7 from 1 to loop_7_cnt_max For loop_6 from 1 to loop_6_cnt_max For Loop5 from 1 to loop_5_cnt_max For Loop4 from 1 to loop_4_cnt_max For Loop3 from 1 to loop_3_cnt_max For Loop2 from 1 to loop_2_cnt_max { if(loop_5_cnt==loop_5_cnt_max) For Loop1 from 1 to loop_1_cnt_max_1 else For Loop1 from 1 to loop_1_cnt_max_0 } For Loop0 from 1 to loop_0_cnt_max In other words, the parameters set by the software may include a condition (loop_5_cnt == loop_5_cnt_max) on the parameter (the value of the number of steps read in the Loop1 level read loop, i.e., loop_1_cnt_max), and the parameter loop_1_cnt_max takes different values, loop_1_cnt_max_1 and loop_1_cnt_max_0, depending on whether the condition (loop_5_cnt < > loop_5_cnt_max) is met or not.
[0106] In other words, when executing up to Loop 5, in each step of Loop 5, Loop 1 is executed at least once, and if Loop 5 is not executed to the last step, i.e., loop_5_cnt < > loop_5_cnt_max, then the number of Loop 1 steps executed in that Loop 5 step is loop_1_cnt_max_0. When Loop 5 is executed to the last step, i.e., loop_5_cnt == loop_5_cnt_max, then the number of Loop 1 steps executed in that Loop 5 step is loop_1_cnt_max_1.
[0107] In this way, the method for calculating the address to read data can be made more flexible.
[0108] As an example of reading data using a nested triple loop when the input tensor in two-dimensional space is not perfectly aligned as described above, Figure 6 shows an example of performing a triple-loop nested read operation on an input tensor that is not perfectly aligned according to the embodiment of this application.
[0109] input tensor
number
[0110] As can be seen, when reading 0-8-1-9-2-10-3-11-4-12, the reading format is perfectly aligned and follows one rule. However, the reading format for 16-17-18-19-20 is incompletely aligned and follows a different rule than the reading format for 0-8-1-9-2-10-3-11-4-12. In this case, we consider using software-configured parameters to achieve a loop-nested reading that is not perfectly aligned.
[0111] Specifically, a triple loop nest (Loop2, Loop1, Loop0) is set up to calculate the address to be read, and the above reading order is implemented.
[0112] The parameters set by the software may be as follows: The number of steps for the outermost Loop2, loop_2_cnt_max, is 2, and the stride jump2_addr is 16; the number of steps for the innermost Loop1, loop_1_cnt_max, is 5, and the stride jump1_addr is 1; the number of steps for the innermost Loop0, loop_0_cnt_max_0, is 2, and the stride jump0_addr is 8; and the conditions for setting the binding between the specified Loop2 and Loop0 are that when Loop2 completes to the last step (loop_2_cnt == loop_2_cnt_max), the number of steps for Loop0 changes from loop_0_cnt_max_0, i.e., from 2 to loop_0_cnt_max_1, and the value is 1.
[0113] The reason for setting the number of steps in the outermost Loop2, loop_2_cnt_max, to 2 is that in the first step, the data is read in the order 0-8-1-9-2-10-3-11-4-12, while in the second step, it is read in the order 16-17-18-19-20. Because the order and rules of execution differ in the two steps, it is necessary to consider how to combine the inner reading loops Loop1 and Loop0 in the second step to achieve a different reading order.
[0114] Specifically, based on the parameters set by the above software, this nested triple loop is executed, and the resulting address is as follows:
[0115] First, the first step of the outermost Loop2 (a total of two steps) is executed, and in that first step, the five steps of Loop1 are executed.
[0116] In the first step of Loop1, all steps of Loop0 are executed, that is, two steps starting from 0 are executed, and since the stride of each step is 8, first 0 is read, then 8 addresses are added together to read 8, and in this way 0-8 are read in two steps.
[0117] In the second step of Loop1, the stride is 1, meaning that all steps of Loop0 are executed from 1, and two steps are executed from 1. Since the stride of each step is 8, first 1 is read, then 8 addresses are added together to read 9, and in this way, 1-9 are read in two steps.
[0118] In the third step of Loop1, the stride is 1, meaning that all steps of Loop0 are executed starting from 2, and two steps are executed starting from 2. Since the stride of each step is 8, first 2 is read, and then 10 is read by adding the 8 addresses. In this way, 2-10 is read in two steps.
[0119] In the fourth step of Loop1, the stride is 1, meaning that all steps of Loop0 are executed from 3, and two steps are executed from 3. Since the stride of each step is 8, first 3 is read, and then 11 is read by adding the 8 addresses. In this way, 3-11 is read in two steps.
[0120] In the fifth step of Loop1, the stride is 1, meaning that all steps of Loop0 are executed starting from 4, and two steps are executed starting from 4. Since the stride of each step is 8, first 4 is read, and then 12 is read by adding the 8 addresses, thus reading 4-12 in two steps.
[0121] Then, the second step of the outermost Loop2 (a total of two steps) is executed, adding a stride of 16 to the initial address 9, i.e., reading from 25.
[0122] In this case, the condition loop_2_cnt == loop_2_cnt_max is satisfied. Therefore, the number of steps in Loop0 is loop_0_cnt_max_1, which means it is 1 step, not 2 steps. In this second step, the 5 steps of Loop1 are executed, and in each step of Loop1, a Loop0 reading loop with 1 step is executed.
[0123] Specifically, in the first step of Loop1, all steps of Loop0 are executed, meaning that one step is executed starting from 16, and the value is read only once. In this case, since stride 8 is not used, only 16 is read.
[0124] In the second step of Loop1, the stride is 1, meaning that 16+1=17, so one step of Loop0 is executed, which means that 17 is read.
[0125] In the third step of Loop1, the stride is 1, meaning that 17+1=18, so one step of Loop0 is executed, which means that 18 is read.
[0126] In the fourth step of Loop1, the stride is 1, meaning that 18+1=19, so one step of Loop0 is executed, which means that 19 is read.
[0127] In the fifth step of Loop1, the stride is 1, meaning that 19+1=20, so one step of Loop0 is executed, which means that 20 is read.
[0128] In this way, by performing a triple nested read loop using parameters set by the software, a complex address read sequence of 0-8-1-9-2-10-3-11-4-12-16-17-18-19-20 is achieved.
[0129] Of course, the above example illustrates that the set conditions are met by performing these steps in a read loop at a different level outside of a specific level, and that different numbers of steps are set in the read loop at that specific level depending on whether the conditions are met or not. However, this application is not limited to this, and can flexibly realize more complex address read sequences by considering other conditions and changes in other parameters that satisfy the conditions.
[0130] Therefore, by flexibly calculating addresses in memory based on parameters set by the software of this application, and reading elements from the calculated addresses and sending them to the calculation unit, a flexible address reading configuration can be realized, thereby increasing the calculation efficiency of the artificial intelligence processor chip and reducing costs. In some cases, it is also possible to replace specific tensor operations in artificial intelligence calculations, thereby simplifying the operation of operators.
[0131] In one embodiment, to calculate the specific address of a nested read loop, the currently read address is calculated based on how many read steps have been performed in one or each level of the read loop, and the stride of each level of the read loop.
[0132] Specifically, regarding the calculation of which address in the SRAM to ultimately read, the address calculation unit can calculate the address to be read each time by determining the position of a single point in a multi-dimensional (read loop) spatial coordinate system based on the parameters set by the software mentioned above.
[0133]
number
[0134]
number
[0135] In other words, when an actual address calculation module calculates an address, it can determine the address to be read by knowing only the number of read steps currently performed in the read loop of one level or each level, and the stride of each read loop of one level or each level.
[0136] Figure 7 shows a schematic diagram of the internal structure of the SRAM according to the embodiment of this application.
[0137] Furthermore, SRAM may be divided into multiple banks, and to speed up data writing and reading, data can be written externally and placed directly into different banks, and read data, calculated data during the process, and result data can be placed directly into different banks. The address creation method in SRAM is configurable, and by default, the most significant bit of the address may be used to distinguish different banks, or interleaving at other granularities may be performed in a way that allows setting an address hash. In terms of hardware design, a multi-bit bank selection signal bank_sel is generated to ultimately select the different SRAM bank sram_bank (sram_bank0-sram_bank3, etc.) for each port such as port0, port1, port2, port3, etc.
[0138] Multi-port access can use handshake signals, and the data pipeline supports back pressure (back pressure is required when inlet traffic is greater than outlet traffic, or when the downstream stage is not prepared, the current stage must have more back pressure than the preceding stage when transferring data, in which case the preceding stage must hold the data as is until the handshake is successful in order to update the data). The memory control unit 304 includes a crossbar (fully cross-connected path) structure in which the read and write functions are separated, allowing parallel access to multiple banks of SRAM, and corresponds to a configuration in which two crossbars are cascaded, mitigating the winding problem of hardware implementation. Single-port SRAM used for lower-layer memory saves power consumption and area. The crossbar structure allows parallel or simultaneous access to multiple banks in which read data, computational data in the process, and final result data are stored, thereby increasing read and write speeds and improving the execution efficiency of the artificial intelligence chip.
[0139] In this way, by using parameters set by software and working in cooperation with the address calculation module to calculate addresses, various tensor operations can be directly replaced. Furthermore, by setting up address calculation processes with more than one loop nest, the address reading calculated by the address calculation module can be made more flexible and is not limited to sequential addresses where each numerical value of the input tensor is stored. By directly replacing operator operations, the time and hardware cost of calculating operator operations, the time and hardware cost of storing the result tensor of operator operations, and the time and hardware cost of reading each element from the address where the result tensor of operator operations is stored are saved in conventional technology. This reduces calculation delay, reduces hardware calculation costs, and improves the execution efficiency of the artificial intelligence chip.
[0140] In summary, the separation of the data pipeline and the computation pipeline enables complete software control and maximizes flexibility. The addressing configuration of the multiple read loops allows for various and complex address access patterns. The multiple asymmetric loop configuration enables unaligned configurations and allows for even more complex address access patterns. Dividing the on-chip shared memory into multiple banks maintains the creation configuration between banks through software. Separation of the data pipeline and the computation pipeline ensures flexibility in data access. Furthermore, data transport and computation pipelines can be hidden from each other, enabling parallel effects for different modules. The compiler's rational division of intermediate data allows different types of data to be stored in different banks of SRAM, for example, enabling simultaneous access to different types of data. If simultaneous access to this data is required, it can be read simultaneously and in parallel from different banks of SRAM, increasing efficiency. After computation is initiated, the utilization rate of the convolutional computation medium access control address (MAC) can reach nearly 100%.
[0141] Figure 8 shows a flowchart illustrating a method for flexibly accessing data in an artificial intelligence processor chip according to an embodiment of this application.
[0142] As shown in Figure 8, the method 800 for flexibly accessing data in an artificial intelligence processor chip includes: step 801, in which the memory in the artificial intelligence processor chip stores tensor data read from outside the processor chip, wherein the read tensor data includes multiple elements used in tensor operations of operators to be computed by artificial intelligence; step 802, in which, based on the tensor operations of the operators, the system controls the system to read elements from memory and transmit them to the compute unit, calculating the address in memory in a single read loop or a multiple read loop (nested loop) based on parameters set by the receiving software, reading the elements from the calculated address and transmitting them to the compute unit in the artificial intelligence processor chip; and step 803, in which the compute unit performs tensor operations of the operators on the received elements.
[0143] In this way, by flexibly calculating memory addresses using parameters set by the software, elements in memory can be read flexibly, and the order of these elements stored in memory is not limited to their order or address order.
[0144] In one embodiment, the parameters set by the software include a value representing the number of elements read for the tensor data read in a single read loop, and a value representing the stride between each step in the single read loop.
[0145] In one embodiment, the parameters set by the software include a value representing the number of reading steps in the reading loop of each hierarchical level, and a value representing the stride between each step in the reading loop of each hierarchical level, and the reading loop of each hierarchical level is performed in a nested manner from the outside to the inside.
[0146] In one embodiment, the parameters set by the software include a value representing an address separated from the address of the first element read in a single read loop and the initial address of the input tensor in memory.
[0147] In one embodiment, the parameters set by the software include conditions on the parameters, and the parameters take different values depending on whether the conditions are met or not.
[0148] In one embodiment, the parameter is a value representing the number of read steps in a particular one-level read loop, and the condition is that several of these steps have been reached in a read loop at a different level outside of that particular one level.
[0149] In one embodiment, the method 800 further includes calculating the address currently being read based on the number of read steps currently performed in one-level or each-level read loop and the respective stride of one-level or each-level read loop.
[0150] In this way, the method for calculating the address in order to read data can be made more flexible.
[0151] In one embodiment, the parameters set by the software instruct the software to replace the tensor operation of the operator by reading elements from memory addresses based on the tensor operation of the operator.
[0152] In one embodiment, tensor operations are tensor operations that do not change the numerical values in the input tensor. Based on the parameters received by the interface, the address in the memory being read is calculated each time a read is performed in a single read loop or a multiple read loop (nested loop), and the tensor operation is changed with each read.
[0153] In one embodiment, the tensor operation includes at least one of the following: the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator.
[0154] In this way, the calculation unit reads the addresses stored at the calculated addresses based on the calculated address order according to the parameters set by the software, and directly replaces the calculation of several tensor operators. This saves the time and hardware cost of calculating these tensor operator calculations, the time and hardware cost of storing the resulting tensor, and the time and hardware cost of reading each element from the address where the resulting tensor is stored, compared to the conventional technology.
[0155] In one embodiment, the memory is divided into multiple banks, each storing data that can be accessed in parallel, and the method further includes a crossbar with separated read and write functions for parallel access to the data stored in the multiple banks of memory.
[0156] In this way, by using parameters set by software and calculating addresses, various tensor operations can be directly replaced, and by setting up address calculation processes for more than one nested loop, the reading of the calculated addresses can be made more flexible and is not limited to sequential addresses where each numerical value of the input tensor is stored. By directly replacing operator operations, the time and hardware cost of calculating operator operations, the time and hardware cost of storing the result tensor of operator operations, and the time and hardware cost of reading each element from the address where the result tensor of operator operations is stored are saved in conventional technology, thereby reducing calculation delay, reducing hardware calculation costs, and improving the execution efficiency of the artificial intelligence chip.
[0157] Figure 9 shows a block diagram of an exemplary electronic device suitable for realizing the embodiment of this application.
[0158] The electronic device may include a processor (H1) and a storage medium (H2) coupled to the processor (H1) and storing computer-executable instructions that, when executed by the processor, perform the steps of each embodiment of this application.
[0159] The processor (H1) may include, but is not limited to, one or more processors or microprocessors.
[0160] The storage medium (H2) may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, and computer storage media (e.g., hard disks, floppy disks, solid-state hard disks, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).
[0161] In addition, the electronic device may further include a data bus (H3), an input / output (I / O) bus (H4), a display (H5), and input / output devices (H6) (e.g., a keyboard, mouse, speaker, etc.).
[0162] The processor (H1) can communicate with external devices (H5, H6, etc.) via a wired or wireless network (not shown) using an I / O bus (H4).
[0163] The storage medium (H2) can further store at least one computer executable instruction that, when executed by the processor (H1), performs a step of each function and / or method in the embodiments described herein.
[0164] In one embodiment, the at least one computer executable instruction may be compiled or configured as a software product, and one or more computer executable instructions, when executed by a processor, perform the steps of each function and / or method in the embodiments described herein.
[0165] Figure 10 shows a schematic diagram of a non-temporary computer-readable storage medium according to an embodiment of the present disclosure.
[0166] As shown in Figure 10, the computer-readable storage medium 1020 stores instructions, such as computer-readable instructions 1010. When a computer-readable instruction 1010 is executed by a processor, it can be executed by referring to the methods described above. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or high-speed cache memory (cache). Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the computer-readable storage medium 1020 can be connected to a computing device such as a computer, and the computing device can execute the computer-readable instructions 1010 stored in the computer-readable storage medium 1020 using the various methods described above.
[0167] This application provides the following items.
[0168] Item 1. An artificial intelligence processor chip that provides flexible access to data, A memory that stores tensor data read from outside the processor chip, wherein the read tensor data includes multiple elements used in tensor operations of operators to be calculated by artificial intelligence, A storage control unit that controls the reading of elements from the memory and transmission to the calculation unit based on the tensor operation of the operator, the storage control unit includes an address calculation module having an interface for receiving parameters set by software, the address calculation module calculates an address in the memory in a single read loop or a multiple read loop (nested loop) based on the parameters received by the interface, reads elements from the calculated address and transmits them to the calculation unit, A calculation unit that performs a tensor operation of the operator on the received elements, Processor chip.
[0169] Item 2. The parameters set by the software include a value representing the number of elements read into the tensor data read in the single read loop, and a value representing the stride between each step in the single read loop. Alternatively, the parameters set by the software include a value representing the number of reading steps in the reading loop for each level, and a value representing the stride between each step in the reading loop for each level. The processor chip described in Item 1 has read loops for each layer, performed in a nested manner from the outside to the inside.
[0170] Item 3. The processor chip described in Item 1, wherein the parameters set by the software include a value representing an address separated from the address of the first element read in a single read loop and the initial address of the input tensor in memory.
[0171] Item 4. The processor chip described in Item 1, wherein the parameters set by the software include conditions for the parameters, and the parameters take different values when the conditions are met and when they are not.
[0172] Item 5. The processor chip described in Item 4, wherein the parameter is a value representing the number of read steps in a particular one-level read loop, and the condition is that several of these steps have been performed in a read loop of another level outside of the particular one-level.
[0173] Item 6. The processor chip described in any one of items 2-5, wherein the address calculation module calculates the currently read address based on the number of read steps currently performed in a read loop of one level or each level, and the stride of each read loop of one level or each level.
[0174] Item 7. The processor chip according to Item 1, wherein the parameters set by the software instruct the software to replace the tensor operation of the operator by reading elements from addresses in the memory, based on the tensor operation of the operator.
[0175] Item 8. The processor chip according to Item 7, wherein the tensor operation is a tensor operation that does not change the numerical value in the input tensor, and the address calculation module replaces the tensor operation with each read in a single read loop or multiple read loop (nested loop) by calculating the address in the memory that is read each time a read is made in a single read loop or multiple read loop (nested loop) based on the parameters received by the interface.
[0176] Item 9. The processor chip according to Item 8, wherein the tensor operation includes at least one of the operations of the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator, the memory is divided into a plurality of banks each storing parallel-accessible data, and the memory control unit includes a crossbar with separated read and write functions for parallel access to the data stored in the plurality of banks of the memory.
[0177] Item 10. A method for flexibly accessing data in an artificial intelligence processor chip, The memory in the artificial intelligence processor chip stores tensor data read from outside the processor chip, and the read tensor data includes multiple elements used in tensor operations of operators to be calculated by the artificial intelligence. The control to read elements from the memory and transmit them to the calculation unit based on the tensor operation of the operator includes calculating the address in the memory in a single read loop or a multiple read loop (nested loop) based on parameters set by the receiving software, reading the elements from the calculated address and transmitting them to the artificial intelligence processor chip. The calculation unit performs a tensor operation of the operator on the received elements, Methods that include...
[0178] Item 11. The parameters set by the software include a value representing the number of elements read into the tensor data read in the single read loop, and a value representing the stride between each step in the single read loop. Alternatively, the method according to item 10, wherein the parameters set by the software include a value representing the number of reading steps in the reading loop of each hierarchy and a value representing the stride between each step in the reading loop of each hierarchy, and the reading loop of each hierarchy is performed in a nested manner from outside to inside.
[0179] Item 12. The method according to Item 10, wherein the parameters set by the software include a value representing an address spaced apart between the address of the first element read in a single read loop and the initial address of the input tensor in memory.
[0180] Item 13. The method according to Item 10, wherein the parameters set by the software include conditions for the parameters, and the parameters take different values when the conditions are met and when the conditions are not met.
[0181] Item 14. The method described in Item 13, wherein the parameter is a value representing the number of read steps in a particular one-level read loop, and the condition is that several of these steps have been performed in a read loop of another level outside of the particular one-level.
[0182] The method described in any one of items 11-14, further comprising calculating the address currently being read based on the number of read steps currently performed in the read loop of the 1st tier or each tier and the stride of the read loop of the 1st tier or each tier.
[0183] Item 16. The method of Item 10, wherein the parameters set by the software instruct to replace the tensor operation of the operator in a manner that reads elements from addresses in the memory, based on the tensor operation of the operator.
[0184] Item 17. The method according to Item 16, wherein the tensor operation is a tensor operation that does not change the numerical value in the input tensor, and the tensor operation is replaced in each read by calculating the address in the memory that is read each time a read is performed in a single read loop or a multiple read loop (nested loop) based on the parameters received by the interface.
[0185] Item 18. The method of Item 17, wherein the tensor operation includes at least one of the operations of the transpose operator, the reshape operator, the broadcast operator, the gather operator, the reverse operator, the concat operator, and the cast operator, the memory is divided into a plurality of banks each storing parallel-accessible data, and the method further includes parallel access to the data stored in the plurality of banks of the memory via a crossbar with separated read and write functions.
[0186] Item 19. Electronic device, Memory for storing instructions, Electronic device including a processor that reads instructions in the memory and performs the method described in any one of items 10-18.
[0187] Item 20. A non-temporary storage medium on which instructions are stored, A non-temporary storage medium in which, when the aforementioned instruction is read by the processor, the method described in any one of items 10-18 is performed by the processor.
[0188] The specific embodiments described above are examples and not limitations. Those skilled in the art can integrate and combine several steps and apparatus from the various embodiments described separately above to achieve the effects of the present invention, based on the concept of the present invention. Such integrated and combined embodiments are also included in the present invention, and their integration and combination will not be described in detail here.
[0189] Furthermore, the advantages, merits, and effects mentioned in this disclosure are illustrative and not limiting, and it is not considered necessary for each embodiment of this application to possess these advantages, merits, and effects. In addition, the specific details of the above disclosure are for illustrative and comprehensible purposes only, and are not limiting, and the above details do not limit the implementation of this application using the above specific details.
[0190] The block diagrams of the devices, apparatus, equipment, and systems relating to this disclosure are provided for illustrative purposes only and are not intended to require or imply that connections, arrangements, or configurations must be made as shown in the block diagrams. As a person skilled in the art will recognize, these devices, apparatus, equipment, and systems can be connected, arranged, or configured in any way. The words “including,” “equipped,” and “having” are open terms meaning “including but not limited to” and can be used interchangeably with them. The terms “or” and “and” as used herein mean the terms “and / or” unless the context explicitly indicates otherwise and can be used interchangeably with them. The term “for example” as used herein means the phrase “for example, not limited to” and can be used interchangeably with them.
[0191] The step flowcharts and methods described herein are illustrative examples only and are not intended to require or imply that the steps of each embodiment must be performed in a given order. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as “then,” “and,” and “next” are not intended to limit the order of the steps. These words are used only to guide the reader to read the description of these methods. Also, any reference to a singular element using, for example, the article “one,” “1,” or “it” is not intended to limit that element to singular.
[0192] Furthermore, the steps and apparatus in each embodiment of this specification are not limited to any particular embodiment. In practice, new embodiments can be conceived based on the concepts of this specification by combining some of the relevant steps and some of the apparatus in each embodiment of this specification, and these new embodiments are also included within the scope of this specification.
[0193] Each operation of the method described above can be performed by any suitable means capable of performing the corresponding function. This means may include, but is not limited to, hardware circuits, application-specific integrated circuits (ASICs), or processors, and may include a variety of hardware and / or software components and / or modules.
[0194] Various exemplary logic blocks, modules, and circuits implemented or described by general-purpose processors, digital signal processors (DSPs), ASICs, field-programmable gate array signal (FPGA) or other programmable logic devices (PLDs), discrete gate or transistor logic, discrete hardware components or any combination thereof, designed to perform the functions described herein may be utilized. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any commercially available processor, controller, microcontroller or state machine. The processor may further be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, a microprocessor working with a DSP core, or any other such configuration.
[0195] Steps of the methods or algorithms described in connection with this disclosure can be directly incorporated into hardware, software modules executed by a processor, or a combination of the two. The software modules can reside in any form of tangible storage medium. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, and the like. The storage medium can be coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The software module may be a single instruction, a number of instructions, and may be distributed across several different code segments, different programs, and multiple storage media.
[0196] The methods disclosed herein include actions to implement the described methods. The methods and / or actions are interchangeable without departing from the claims. In other words, unless a specific order of actions is specified, the order and / or use of any particular actions can be modified without departing from the claims.
[0197] The above functions can be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions can be stored as instructions on a suitable computer-readable medium. The storage medium may be any available suitable medium accessible by a computer. By example, but not limited to, such computer-readable mediums may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other reliable medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital general-purpose discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using a laser.
[0198] Therefore, a computer program product can perform the operations given herein. For example, such a computer program product may be a computer-readable tangible medium having tangible stored (and / or encoded) instructions, which can be executed by a processor to perform the operations described herein. The computer program product may include packaging materials.
[0199] Software or instructions may be transmitted by a transmission medium. For example, software may be transmitted by a website, server, or other remote source using a transmission medium such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, RF, or microwave.
[0200] Furthermore, modules and / or other suitable means for performing the methods and techniques described herein may be downloaded and / or otherwise obtained by user terminals and / or base stations at appropriate times. For example, such equipment may be coupled to a server to facilitate the transmission of means for performing the methods described herein. Alternatively, the various methods described herein may be provided via storage materials (e.g., physical storage media such as RAM, ROM, CD, or floppy disk) so that user terminals and / or base stations can obtain the various methods while they are connected to the equipment or while they are providing storage materials to the equipment. Furthermore, the methods and techniques described herein can be utilized in other suitable techniques of the equipment.
[0201] Other examples and embodiments are within the scope and spirit of the claims and appendices of this disclosure. For example, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwire, or any combination thereof, due to the nature of the software. The features that implement the functions can also be physically located in various locations, including being distributed so that some of the functions are implemented in different physical locations. Furthermore, as used herein and in the claims, the "or" used in an enumeration of a term beginning with "at least one" means a separate enumeration, for example, the enumeration of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A, B, and C). Furthermore, the term "exemplary" does not mean that the examples described are preferred or superior to other examples.
[0202] Various changes, substitutions, and modifications of the techniques described herein can be implemented without departing from the taught techniques as defined by the appended claims. Furthermore, the claims of this disclosure are not limited to specific embodiments of the configurations, means, methods, and operations of the processes, machines, manufactures, and events described herein. By utilizing the corresponding embodiments described herein, existing or later developed configurations, means, methods, or operations of processes, machines, manufactures, and events that are substantially the same in function or produce substantially the same results can be realized. Accordingly, the appended claims include such configurations, means, methods, or operations of processes, machines, manufactures, and events within their scope.
[0203] The above description of the disclosed embodiments is provided so that a person skilled in the art can prepare or use the present application. Various modifications to these embodiments will be very obvious to a person skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the present application. Accordingly, the present application is not intended to be limited to the embodiments shown herein, but rather to adhere to the broadest scope that is consistent with the principles and novel features disclosed herein.
[0204] The above description is provided for illustrative and explanatory purposes. Furthermore, this description is not intended to limit the embodiments of the present application to the forms of the present disclosure. Although several exemplary embodiments and examples have been described above, those skilled in the art will be able to recognize several variations, modifications, changes, additions, and subcombinations thereof.
Claims
1. A processor chip that provides flexible access to data, A memory that stores tensor data read from outside the processor chip, wherein the read tensor data includes multiple elements used in tensor operations of operators to be calculated by artificial intelligence, A storage control unit that controls the reading of elements from the memory and transmission to the calculation unit based on the tensor operation of the operator, the storage control unit includes an address calculation module having an interface for receiving parameters set by software, the address calculation module calculates an address in the memory in a nest of a single read loop or multiple read loop based on the parameters received by the interface, reads elements from the calculated address and transmits them to the calculation unit, A calculation unit that performs the tensor operation of the operator on the received elements, A processor chip, including...
2. The parameters set by the software include, or, a value representing the number of elements read into the tensor data read in the single-read loop, and a value representing the stride between each step in the single-read loop. The parameters set by the aforementioned software include a value representing the number of steps to be read in the read loop for each level, and a value representing the stride between each step in the read loop for each level. The loading loops for each level are performed in a nested manner, from the outside to the inside. The processor chip according to claim 1.
3. The parameters set by the aforementioned software include a value representing an address separated from the address of the first element read in a single read loop and the initial address of the input tensor in memory. The processor chip according to claim 1.
4. The parameters set by the aforementioned software include conditions for those parameters, The aforementioned parameter takes different values depending on whether the condition is met or not. The processor chip according to claim 1.
5. The aforementioned parameter is a value that represents the number of reading steps in a particular single-level reading loop, The aforementioned condition is the number of read steps performed in a read loop for another hierarchy outside of the specific hierarchy mentioned above. The processor chip according to claim 4.
6. The address calculation module calculates the currently read address based on the number of read steps currently performed in the read loop of one level or each level, and the stride of each read loop of one level or each level. A processor chip according to any one of claims 2-5.
7. The parameters set by the software instruct the software to replace the tensor operation of the operator by reading elements from the address in the memory, based on the tensor operation of the operator. The processor chip according to claim 1.
8. The aforementioned tensor operation is a tensor operation that does not change the numerical values in the input tensor. The address calculation module calculates the address in the memory that is read each time a read is performed in a nested single-read loop or multiple-read loop, based on the parameters received by the interface, thereby replacing the tensor operation during the read. The processor chip according to claim 7.
9. The tensor operation includes at least one of the following: the operation of the transform operator, the operation of the reshape operator, the operation of the broadcast operator, the operation of the gather operator, the operation of the reverse operator, the operation of the concat operator, and the operation of the cast operator. The memory is divided into multiple banks, each storing data that can be accessed in parallel. The memory control unit includes a crossbar with separated read and write functions in order to access data stored in multiple banks of the memory in parallel. The processor chip according to claim 8.
10. A method for flexibly accessing data on a processor chip, The memory in the processor chip stores tensor data read from outside the processor chip, and the read tensor data includes multiple elements used in tensor operations of operators to be calculated by artificial intelligence. Based on the tensor operation of the aforementioned operator, the system may control the reading of elements from the memory and sending them to the calculation unit; based on the parameters set by the received software, the system may calculate the address in the memory in a single-read loop or a multiple-read loop (nested loop), read the elements from the calculated address, and send them to the calculation unit of the processor chip; The calculation unit performs a tensor operation of the operator on the received elements, Methods that include...
11. The parameters set by the software include, or, a value representing the number of elements read into the tensor data read in the single-read loop, and a value representing the stride between each step in the single-read loop. The parameters set by the aforementioned software include a value representing the number of steps to be read in the read loop for each level, and a value representing the stride between each step in the read loop for each level. The loading loops for each level are performed in a nested manner, from the outside to the inside. The method according to claim 10.
12. The parameters set by the aforementioned software include a value representing an address separated from the address of the first element read in a single read loop and the initial address of the input tensor in memory. The method according to claim 10.
13. The parameters set by the aforementioned software include conditions for those parameters, The aforementioned parameter takes different values depending on whether the condition is met or not. The method according to claim 10.
14. The aforementioned parameter is a value that represents the number of reading steps in a particular single-level reading loop, The aforementioned condition is the number of read steps performed in a read loop for another hierarchy outside of the specific hierarchy mentioned above. The method according to claim 13.
15. The above method further, This includes calculating the currently read address based on the number of read steps currently performed in a read loop at one level or each level, and the stride of each read loop at one level or each level. The method according to any one of claims 11-14.
16. The parameters set by the software instruct the software to replace the tensor operation of the operator by reading elements from the address in the memory, based on the tensor operation of the operator. The method according to claim 10.
17. The tensor operation is a tensor operation that does not change the numerical value in the input tensor, and the tensor operation is replaced with each read by calculating the address in the memory that is read each time a read is made in a nested single read loop or multiple read loop, based on the parameters received by the interface. The method according to claim 16.
18. The tensor operation includes at least one of the following: the operation of the transform operator, the operation of the reshape operator, the operation of the broadcast operator, the operation of the gather operator, the operation of the reverse operator, the operation of the concat operator, and the operation of the cast operator. The memory is divided into multiple banks, each storing data that can be accessed in parallel. The above method further, This includes parallel access to data stored in multiple banks of the memory via a crossbar that separates read and write functions, The method according to claim 17.
19. Memory for storing instructions, Electronic device comprising a processor that reads instructions in the memory and performs the method according to any one of claims 10-14.
20. A non-temporary storage medium in which instructions are stored, When the instruction is read by the processor, the method according to any one of claims 10-14 is executed by the processor. Non-temporary storage medium.
Citation Information
Patent Citations
Arithmetic processing unit and image processing unit
JP2020017179A
Accessing Data in Multidimensional Tensors Using Adders
JP2020521198A
Neural Network Processor
JP2022514680A
Address Generation Unit Using Nested Loops To Scan Multi-Dimensional Data Structures
US20100145992A1