Flexible tensor traversal unit
Patent Information
- Application Number
- JP2025574702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-18
- Filing Date
- 2024-09-18
- Publication Date
- 2026-09-30
Smart Images

Figure 2026532580000001_ABST
Abstract
Description
Technical Field
[0001] Cross-Reference to Related Applications This application claims the priority of U.S. Provisional Patent Application No. 63 / 538,040 filed on September 18, 2023. The disclosure of the prior application is considered part of the disclosure of the present application and is incorporated into the disclosure of the present application by reference. Background Art
[0002] The present specification generally relates to tensor traversal units (TTU). A tensor traversal unit (TTU) is a special hardware component used for performing tensor operations in machine learning, for example, neural network computation. A tensor includes a single numerical array or multiple numerical arrays, and a tensor operation is an arithmetic operation for manipulating tensors. Tensor traversal units are typically integrated into hardware accelerators such as graphics processing units (GPU), tensor processing units (TPU), and field programmable gate arrays (FPGA).
[0003] A neural network is a machine learning model that utilizes one or more layers of the model to generate an output corresponding to a received input. Some neural networks include one or more hidden layers in addition to outer layers. The output of each hidden layer is used as an input to the next layer in the network, that is, the next hidden layer or the output layer of the network. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters. Summary of the Invention
[0004] This specification describes techniques related to determining memory addresses of tensor elements of an N-dimensional tensor using one or more hardware counters and one or more hardware adders.
[0005] In general, one innovative aspect of the subject matter described herein can be embodied in a way that is performed by a computer, the method of which is performed by a computer, comprising taking one or more instructions to distribute the tensor elements of an N-dimensional tensor between destination memory units, the N-dimensional tensor having a plurality of tensor elements arranged across each of the N dimensions, where N is an integer of 1 or more, the method further comprises determining the address offset value of the corresponding memory bank in the destination memory unit for each tensor element of the N-dimensional tensor by using one or more counters and one or more adders, the address offset value of each tensor element being determined as a linear combination of partial address offset values counted by one or more counters, and the method further comprises writing the tensor elements of the N-dimensional tensor to the corresponding memory bank in the destination memory unit according to the determined address offset values.
[0006] These embodiments and other embodiments may each optionally include one or more of the following features. In some embodiments, the destination memory unit includes one or more cell groups, each including one or more memory cells, each including one or more memory banks, and the instruction includes a first value defining the number of consecutive tensor elements written to each memory cell, a second value defining the step value taken between memory cells when writing tensor elements to different memory cells, and a third value defining the total number of cell groups in the destination memory unit for storing tensor elements. In some embodiments, determining the address offset value for each tensor element includes determining a fourth value defining the total number of tensor elements written to each cell group by dividing the total number of tensor elements by the third value. In some embodiments, the partial address offset value includes three values counted by three consecutively arranged counters. In some embodiments, the three counters include a least significant digit counter that increments by 1 for each tensor element of an N-dimensional tensor up to a first value. In some embodiments, the three counters include a second least-decipher counter that increments by 1 with each increment signal from the least-decipher counter up to a third value. In some embodiments, the three counters include a third least-decipher counter that increments by 1 with each increment signal from the second least-decipher counter up to a fourth value divided by the first value. In some embodiments, determining the address offset value for each tensor element of an N-dimensional tensor involves using one or more adders to calculate the sum of (i) the partial address offset value counted by the least-decipher counter, (ii) the partial address offset value counted by the second least-decipher counter obtained by multiplying the second value by a constant corresponding to the number of memory banks contained in the cell, and (iii) the partial address offset value counted by the third least-decipher counter obtained by multiplying the first value.In some embodiments, obtaining an instruction to distribute the tensor elements of an N-dimensional tensor among memory units includes reading the source tensor elements of the source tensor from the source memory unit, performing a data type conversion on the source tensor elements, and storing the converted source tensor elements in a shift buffer. In some embodiments, obtaining an instruction to distribute the tensor elements of an N-dimensional tensor among memory units includes obtaining a base memory address value for accessing the N-dimensional tensor and storing the base memory address value in an address queue. In some embodiments, writing the tensor elements of an N-dimensional tensor to the corresponding memory banks of the memory units includes, for each tensor element, using a set of multiplexers to write the tensor element to the corresponding memory bank identified by the determined address offset value of the tensor element. In some embodiments, the method further includes obtaining one or more separate instructions to distribute the tensor elements of another N-dimensional tensor among destination memory units, and writing the tensor elements of the other N-dimensional tensor to the memory banks of the destination memory units, while bypassing determining the address offset values of the tensor elements of the other N-dimensional tensor using one or more counters and one or more adders. In some embodiments, the method further includes outputting data indicating the determined address offset value of a particular tensor element of the N-dimensional tensor.
[0007] This aspect and other embodiments of other aspects include corresponding systems and computer programs configured to perform actions of the method, encoded in a computer storage device. One or more computer systems can be configured in this way by software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform actions when it is running. One or more computer programs can be configured in this way by having instructions that cause the device to perform actions when executed by a data processing device.
[0008] The subject matter described herein can be implemented in particular embodiments to achieve one or more of the following advantages: (i) one or more counters that each count up to the respective maximum value specified in a memory access instruction, and (ii) one or more arithmetic units that apply arithmetic operations, such as an adder that applies addition to the values counted by the counters that calculate the address offset value, so that a special-purpose computing unit can quickly determine a memory address value for accessing a storage medium. The memory address values do not need to be contiguous on the destination storage medium. The memory address values do not need to be ordered on the destination storage medium in the same way as the source storage medium from which the data was read.
[0009] Therefore, special-purpose computing units facilitate easier and faster data distribution across storage media according to flexible data distribution patterns suited to the needs of specific computing tasks. By using counters and adders to determine memory address values, the number of computation cycles in the processor, the number of instructions the processor needs to execute, or both of those required to write data in other ways can be reduced, increasing processor bandwidth for other computing tasks.
[0010] Details of one or more embodiments of the subject matter described herein are given in the accompanying drawings and the following description. Other potential features, aspects and advantages of the subject matter will become apparent from the description, drawings and claims.
[0011] Similar reference numbers and names in various drawings refer to the same elements. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram of an exemplary computing system. [Figure 2] This diagram illustrates the transfer of a tensor from the first memory unit to the second memory unit. [Figure 3] This is an illustrative diagram showing how tensor elements of a tensor are distributed across memory units. [Figure 4] This flowchart illustrates an exemplary process for distributing tensor elements of an N-dimensional tensor between destination memory units. [Figure 5] This is an illustrative diagram showing how tensor elements of an N-dimensional tensor are distributed among destination memory units. [Modes for carrying out the invention]
[0013] Figure 1 shows a block diagram of an exemplary computing system 100. Generally, the computing system 100 processes an input 101 to produce an output 116. For example, the computing system 100 may be configured to perform linear algebra calculations. The input 101 may include any suitable data, instructions, or both that can be processed by the computing system 100. The computing system 100 includes a processing unit 102, a storage medium 104, and a tensor traversal unit 106. The computing system 100 can use the tensor traversal unit 106 to traverse a tensor while processing the input 101 to produce the output 116.
[0014] An N-dimensional tensor may be a vector, a matrix, or a multidimensional matrix. For example, a one-dimensional tensor is a vector, a two-dimensional tensor is a matrix, and a three-dimensional tensor is a three-dimensional matrix composed of multiple two-dimensional matrices. Each dimension of an N-dimensional tensor may contain one or more tensor elements, and each tensor element may store its own data value.
[0015] The processing unit 102 is configured to process instructions for execution within the computing system 100, including instructions 112 stored in the storage medium 104, or other instructions stored in other storage devices and provided to the computing system 100 via a communication bus. The processing unit 102 may include one or more processors. The storage medium 104 stores information within the computing system 100. In some embodiments, the storage medium 104 is a volatile memory unit. In some embodiments, the storage medium 104 is a non-volatile memory unit. The storage medium 104 may also be another form of computer-readable medium, such as an array of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. When the instructions are executed by the processing unit 102, they cause the processing unit 102 to perform one or more computational tasks.
[0016] The tensor traversal unit 106 may be implemented as an application-specific integrated circuit. Generally, when the processing unit 102 executes an instruction to store a particular tensor element of a tensor, the tensor traversal unit 106 determines the address offset value of the particular tensor element of the tensor based on the tensor index of the particular tensor element of the tensor. For example, the data 114 may include any intermediate data generated by the computing system 100 during processing of the input 101 to produce the output 116, or it may include the output 116 instead.
[0017] In other words, the tensor traversal unit 106 can convert a set of N-dimensional tensor indices into a sequence of address offset values in a one-dimensional address space. The tensor traversal unit can perform such a conversion by counting one or more partial address offset values based on the dimensional index of the tensor elements, and then converting the address offset values of the tensor elements into a linear combination of the counted partial address offset values.
[0018] The tensor traversal unit 106 can efficiently and programmatically generate a sequence of addresses that reference a sequence of tensor elements. The address sequence corresponds to a sequence of tensor elements that will be accessed, for example, by one or more memory access instructions specifying loop nesting in a software traversal routine. A sequence of tensor elements, even if contiguous within a tensor, may or may not be stored in physically contiguous memory banks when written to memory. The example shown in Figure 5 and described below provides an example of how a sequence of tensor elements can be written to a memory bank of a destination memory unit that is not physically contiguous within the destination memory unit.
[0019] The tensor traversal unit 106 includes a hardware counter unit 122 and a hardware adder unit 124. The hardware counter unit 122 may include one or more counters. Each counter may include digital circuitry configured to perform counting operations. Each counter may count a partial address offset value, which is stored as a count value in the counter's count register, which is incremented each time an increment signal is received at the input, up to a maximum value specified in the counter's maximum count register.
[0020] In embodiments with multiple hardware counters, the hardware counters can be arranged sequentially in a chain such that the output signal of a first counter, which may be called the least significant digit counter, is provided as an input increment signal to a second counter, which may be called the second least significant digit counter. Correspondingly, the output signal of the second counter is provided as an input to a third counter, which may be called the third least significant digit counter, and so on, up to the last counter in the chain, which may be called the Nth least significant digit counter, if any exist.
[0021] In some embodiments, each hardware counter includes a count register configured to store a value and increment the value stored therein whenever an input signal is received (for example, implemented as a group of flip-flops or latches). In some embodiments, each hardware counter also includes a maximum count register configured to receive and store the maximum value (for example, implemented as a group of flip-flops or latches). An output signal of the hardware counter is provided when the value stored in the count register reaches the maximum value, i.e., when the value stored in the count register exceeds the maximum value. Furthermore, when the value in the hardware counter's count register reaches the maximum value of the maximum count register, the count register may be configured to reset the value stored in the count register to a starting value, such as zero.
[0022] For example, as described below, the value stored in the count register of the least significant digit counter may be increased by a predetermined amount, such as being incremented by 1, each time an input signal is received by the hardware counter unit 122. The value stored in the count register of the second least significant digit counter is incremented each time the least significant digit counter provides an output signal, that is, each time the value stored in the count register reaches the value stored in the maximum count register of the least significant digit counter, and correspondingly, the value stored in the count register is reset to a starting value such as zero. Then, the second least significant digit counter provides an output signal when the value of its count register reaches the maximum value stored in the maximum count register of the second least significant digit counter. When the second least significant digit counter provides an output signal, the value stored in the count register of the third least significant digit counter is sequentially incremented up to the Nth least significant digit counter, that is, the last counter in the chain.
[0023] The hardware adder unit 124 may include one or more hardware adders. Each hardware adder may include a digital circuit configured to perform addition. For example, as described below, one or more adders may multiply a partial address offset value by a given coefficient (e.g., by shift addition) to determine a weighted partial address offset value, and then add the weighted partial address offset values to determine a total address offset value for a tensor element of a tensor.
[0024] In addition to or instead of the hardware adder unit 124, some embodiments of the tensor traversal unit 106 may include an arithmetic logic unit (ALU), a hardware multiplier unit, or other hardware units having a suitable digital circuit that can be configured to perform addition operations and multiplication operations. When included, the ALU and / or multiplier unit can be used to calculate a total address offset value of a tensor element of a tensor from a partial address offset value counted by the hardware counter unit 122.
[0025] The computing system 100 described in the present specification may optionally operate in cooperation with other similar or different computing systems to perform calculations across one or more layers of a multi-layer neural network. A calculation process performed within a neural network layer may include multiplication between an input activation tensor including input activations and a weight tensor including weights of the neural network layer. The calculation may include multiplying input activations by the weights to generate products in one or more cycles, and accumulating the products over a number of cycles.
[0026] In general, computations on a given neural network layer require several instructions 112. When these instructions 112 are received and executed by the processing unit 102, they cause the processing unit 102 to perform computational tasks associated with traversing a tensor. For example, such computational tasks include moving data between two distinct memory address locations in the same or different memory units, storing activations at a memory address location in the first memory unit, and storing weights at a memory address location in the second memory unit. Through these instructions 112, the processing unit 102 can access the data stored in the first and second memory units. Through these instructions 112, the processing unit 102 can move data as needed to achieve the execution of a particular tensor operation by the system 100, which includes moving data between two memory resources of different widths, for example, from the first memory unit to the second memory unit or vice versa.
[0027] For example, the first memory unit may be a narrow memory unit, and the second memory unit may be a wide memory unit, and either or both of the first and second memory units are included in the storage medium 104. In some cases, data read from the narrow memory unit can be broadcast to a processing unit associated with the wide memory unit, and data can be read in parallel from the wide memory unit to support parallel computing.
[0028] The terms "wide" and "narrow" generally refer to the approximate width (bits / bytes) of one or more memory units. As used herein, "narrow" may refer to one or more memory units each having a size or width of less than 16 bytes (or 16 bits), and "wide" may refer to one or more memory units each having a size or width greater than 16 bytes (or 16 bits), although in some embodiments this may be less than 128 bytes or 256 bytes (or 128 bits or 128 bits).
[0029] Figure 2 is an illustrative diagram showing the movement of a tensor from the first memory unit 202 to the second memory unit 218. The computing system 100 can perform the operations shown in Figure 2 as a result of the processing unit 102 executing instructions (or multiple instructions) to move data, namely, instructions (or multiple instructions) to read data from the first memory unit 202 (also called the source memory unit) and then write data to the second memory unit 218 (also called the destination memory unit).
[0030] In the example in Figure 2, the first memory unit 202 is a narrow memory unit, and the second memory unit 218 is a wide memory unit. Each memory unit has several memory banks. Each memory bank can store tensor elements of a tensor. For example, the first memory unit can store the input activations of an input activation tensor in its memory bank at memory address locations corresponding to each tensor element of the input activation tensor. Similarly, the first memory unit can also store the weights of a weight tensor, or the output activations of an output activation tensor, in its memory bank.
[0031] The operation of moving data occurs when the computing system executes a memory access instruction to read data from the first memory unit 202 (step 204). Correspondingly, the computing system requests to read the tensor elements of a tensor stored in the memory bank of the first memory unit 202 (e.g., the input activations of an input activation tensor) (step 206). The computing system retrieves the tensor elements (step 208). In some cases, the tensor elements read from the first memory unit also need to have the data type corresponding to the second memory unit. Therefore, the computing system optionally performs a data type conversion, for example, to convert an integer to a floating-point number (step 210). The computing system buffers the tensor elements in one or more shift buffers (step 212). For example, the buffering of tensor elements can be done in a first-in, first-out (FIFO) order. Alternatively, the buffering of tensor elements can be done according to the order defined by the instructions received by the computing system. The computing system executes a memory access instruction to write data to the second memory unit 218 (step 214). Since the tensor is stored in the second memory unit 218, a set of memory addresses in the second memory unit 218 must be generated for the tensor based on the tensor indices of the tensor elements of the tensor. As will be further described below, the computing system uses a tensor traversal unit to provide the address of each tensor element of the tensor based on the index of the dimension associated with the tensor element (step 216).
[0032] Figure 3 is an illustrative diagram showing the distribution of tensor elements of one or more tensors across memory units 318. For example, memory unit 318 may be a destination memory unit, such as the second memory unit 218 in Figure 2, and the tensor may be a tensor read from a source memory unit, such as the first memory unit 202 in Figure 2.
[0033] The computing system receives a memory access instruction 314 to write the tensor to the memory unit 318. As briefly described above, the N-dimensional tensor may be a vector, a matrix, or a multidimensional matrix. The computing system may execute an algorithm that includes nested loops to perform tensor operations by iterating through one or more nested loops to traverse the N-dimensional tensor. In one exemplary computing process, each loop in the loop nesting may be responsible for traversing a particular dimension of the N-dimensional tensor.
[0034] The memory access instruction 314 includes a set of values that define a data distribution pattern for distributing the tensor elements of a tensor across the memory unit 318. Furthermore, the memory access instruction 314 may also include a base memory address value 302 and one or more loop index variables 304. For example, the memory access instruction 314 may be provided to a computing system via a communication bus, and upon arrival, cause the computing system to perform a memory access operation.
[0035] The base memory address value 302 can be used by the computing system to locate a memory address in the memory unit 318 to be accessed. For example, a memory access instruction 314 may include the memory address of a specific tensor element of a tensor, for example, the first tensor element. The received base memory address value may be stored in the computing system's address queue 306. The address queue 306 may include several address registers for storing base memory address values received by one or more memory access instructions.
[0036] One or more loop index variables 304 are used by the computing system to traverse the tensor elements of the tensor in order to write the tensor elements of the tensor to a destination memory unit, for example, the second memory unit 218 in Figure 2. The for loop is an example of nested loops, in which three loops tracked by three loop index variables (e.g., i, j, and k) can be nested to traverse a three-dimensional tensor. In this example of a three-dimensional tensor, the outer for loop may be used to traverse the loop tracked by variable i, the middle for loop may be used to traverse the loop tracked by variable j, and the inner for loop may be used to traverse the loop tracked by variable k. Thus, the first tensor element accessed may be (i=0, j=0, k=0), the second tensor element may be (i=0, j=0, k=1), and so on. Other examples of nested loops may include more or fewer loop index variables.
[0037] As will be further explained below with reference to Figures 4 and 5, the computing system uses hardware counter unit 322 and hardware adder unit 324 to calculate an address offset value for each tensor element of the tensor, and then writes the tensor elements to memory unit 318 according to the memory address value determined from the address offset value and the base memory address value.
[0038] In some embodiments, the computing system includes a set of multiplexers, including a multiplexer device 308 that controls which memory bank of a memory unit 318 stores which tensor elements. For example, each memory bank of a memory unit can be connected to a multiplexer, and the multiplexer can be used to select the appropriate memory bank for storing a tensor element based on the memory address value of the tensor element.
[0039] When tensor elements are received from the shift buffer 312, the computing system controls the multiplexer device 308 to store each tensor element in the appropriate memory bank corresponding to the memory address value of the tensor element (calculated using the hardware counter unit 322 and the hardware adder unit 324). In this way, the computing system distributes the tensor elements of the tensor across the memory units 318 according to the data distribution pattern specified by the memory access instruction 314.
[0040] In some embodiments, the computing system includes a bypass path that bypasses the use of both the hardware counter unit 322 and the hardware adder unit 324 when calculating memory address values. In these embodiments, the system can operate in bypass mode by following the bypass path instead of using the hardware counter unit 322 and the hardware adder unit 324. In other words, tensor elements received from the shift buffer 312 are written directly to the memory unit 318, i.e., without redistribution. For example, the memory access instruction 314 may include an indicator or field that tells the computing system to either enable or disable the hardware counter unit 322 and the hardware adder unit 324.
[0041] That is, the memory access instruction 314 may enable the hardware counter unit 322 and the hardware adder unit 324, thereby causing the computing system to use these units to calculate the address offset value. Alternatively, the memory access instruction 314 may disable the hardware counter unit 322 and the hardware adder unit 324, thereby causing the computing system to enter a bypass mode in which the computing system does not use these units to calculate the address offset value. When operating in bypass mode, the computing system can write the tensor elements of the tensor to the memory unit 318 according to a fixed order, for example, the order in which the tensor elements of the tensor were popped from the shift buffer 312, or the order in which the tensor elements of the tensor were read from the first memory unit 202 in Figure 2. If redistribution of tensor elements is not required, the time and / or amount of computation required to traverse the tensor can be further reduced by bypassing the use of the hardware counter unit 322 and the hardware adder unit 324 when calculating the address offset value.
[0042] Figure 4 is a flowchart illustrating an exemplary process 400 for distributing tensor elements of an N-dimensional tensor between destination memory units. Process 400 may be performed by a system of one or more computers, for example, the computing system 100 in Figure 1. The system includes a tensor traversal unit. The tensor traversal unit includes a hardware counter unit having one or more counters. The tensor traversal unit also includes a hardware adder unit having one or more adders.
[0043] The system obtains one or more instructions to distribute the tensor elements of an N-dimensional tensor among destination memory units (step 402). An N-dimensional tensor can contain one or more elements arranged across each of the N dimensions, where N is an integer greater than or equal to 1. An N-dimensional tensor can also be part of a larger tensor having the same or different dimensions, and therefore the distributed tensor elements can be a subset of a larger group of tensor elements. For example, the system may include a processing unit, e.g., processing unit 102 in Figure 1, that executes instructions to write the tensor elements of the tensor to a destination memory unit, e.g., storage medium 104 in Figure 1.
[0044] One or more instructions may represent instructions for handling nested loops, each of which may be iterated over using a corresponding loop index variable. The program may specify the value of each of the loop index variables. Furthermore, one or more instructions may include a base memory address for accessing a destination memory unit. For example, an instruction may include a base memory address corresponding to the address of a destination memory unit to which a particular tensor element of a tensor, e.g., a first tensor element, should be written.
[0045] Specifically, one or more instructions include a set of values that define, or otherwise specify, a data distribution pattern in which these tensor elements of an N-dimensional tensor should be distributed among destination memory units.
[0046] In some cases, the destination memory unit may include multiple memory banks. Each memory bank has a different memory address. One or more memory banks may be grouped into memory cells (so each memory cell includes multiple memory banks). One or more memory cells may be grouped into cell groups (so each cell group includes multiple memory cells).
[0047] In these cases, one or more instructions may include three values that collectively define the data distribution pattern: a first value ("unit byte" value) that defines the number of consecutive tensor elements written to each memory cell; a second value ("cell stride" value) that defines the step value (or length of interval) taken between memory cells when writing tensor elements to different memory cells; and a third value ("group count" value) that defines the total number of cell groups in the destination memory unit for storing the tensor elements.
[0048] In some of these cases, one or more instructions may also include a fourth value ("access bytes" value) that defines the total number of tensor elements written to each cell group. In other cases, the system can calculate the fourth value by dividing the total number of tensor elements of the N-dimensional tensor by a third value ("group count" value).
[0049] Figure 5 is an illustrative diagram illustrating the distribution of tensor elements of an N-dimensional tensor among the destination memory units 500. As shown in the figure, the N-dimensional tensor contains a total of 18 tensor elements, and the destination memory unit 500 contains three cell groups, each cell group containing two memory cells (e.g., the first cell group containing cell 0 and cell 1), and each memory cell containing four memory banks (e.g., the first memory cell contains four memory banks, each capable of storing a tensor element of the tensor). Thus, the destination memory unit 500 contains a total of 3 * 2 * 4 = 24 memory banks to store the 18 tensor elements of the N-dimensional tensor.
[0050] In the example in Figure 5, the system obtains one or more instructions that include, or specify, a first value ("unit byte" value) of 2, a second value ("cell stride" value) of 2, a third value ("group count" value) of 3, and a fourth value ("access byte" value) of 6.
[0051] The system determines the address offset value of the corresponding memory bank in the destination memory unit for each tensor element of the N-dimensional tensor by using one or more counters and one or more adders (step 404). Generally, the system determines the address offset value for each tensor element by using one or more adders to compute a linear combination of the partial address offset values counted by one or more counters.
[0052] Continuing with the example in Figure 5, the system may include three counters, including a least significant digit counter, a second least significant digit counter, and a third least significant digit counter, which are arranged consecutively in the chain. For each counter, the system can determine the maximum value from a set of values contained in (or specified by) one or more instructions obtained by the system in step 402. The system may configure the counters so that each counter counts up to its respective maximum value by, for example, storing the maximum value in the counter's maximum count register.
[0053] The least significant digit counter is configured to increment by 1 each time an input signal is received by the least significant digit counter. For example, the least significant digit counter may receive such an input signal during each memory access cycle in which a tensor element of an N-dimensional tensor is read. The maximum value of the least significant digit counter can be equal to a first value ("unit byte" value). Thus, the least significant digit counter can start from zero (included), count up to a first value (not included), and then wrap around, i.e., reset to zero.
[0054] The second least significant digit counter is configured to increment by 1 with each increment signal from the least significant digit counter. For example, the least significant digit counter can provide such an increment signal whenever the value counted by the least significant digit counter reaches the maximum value of the least significant digit counter. The maximum value of the second least significant digit counter can be equal to a third value ("group count" value). Thus, the second least significant digit counter can start from zero (included), count up to the third value (not included), and then wrap around, i.e., reset to zero.
[0055] The third least significant digit counter is configured to increment by 1 with each increment signal from the second least significant digit counter. For example, the second least significant digit counter can provide such an increment signal whenever the value counted by the second least significant digit counter reaches its maximum value. The maximum value of the third least significant digit counter can be equal to the result of dividing the fourth value ("access byte" value) by the first value ("unit byte" value). Thus, the third least significant digit counter can start from zero (included), count up to the division result (not included), and then wrap around, i.e., reset to zero.
[0056] Therefore, as shown in the "Counters" row of the table in the lower half of Figure 5, when the first tensor element (D0) is read by the system, the three counters count three partial address offset values (0, 0, 0). The least significant digit counter increments by 1 when the second tensor element (D1) is read by the system, i.e., incrementing each tensor index of the tensor by 1. Correspondingly, the three counters now count three partial address offset values (0, 0, 1). The second least significant digit counter increments by 1 in response to the increment signal provided by the least significant digit counter when the third tensor element (D2) is read by the system, and correspondingly, the three counters now count three partial address offset values (0, 1, 0). The three counters continue counting in this manner until the last tensor element (D17) is read by the system, at which point the three partial address offset values are (2, 2, 1).
[0057] The system determines an address offset value for each tensor element of an N-dimensional tensor and can calculate the sum of (i) the partial address offset value counted by the least significant digit counter, (ii) the partial address offset value counted by a second least significant digit counter obtained by multiplying the result of the second value by a constant corresponding to the number of memory banks contained in the cell (equal to 4 in the example in Figure 5), and (iii) the partial address offset value counted by a third least significant digit counter obtained by multiplying the result of the first value, as shown in the following equation. Address offset value = Offset 0 + Offset 1 * 2nd value * 4 + Offset 2 * 1st value In the formula, offset 0 represents the partial address offset value counted by the least significant digit counter, offset 1 represents the partial address offset value counted by the second least significant digit counter, and offset 2 represents the partial address offset value counted by the third least significant digit counter. The first value is the unit byte value, and the second value is the cell stride value.
[0058] The system can calculate this sum by using one or more adders. For example, to calculate (iii), the system can provide the partial address offset value counted by the third least significant digit counter and the first value as inputs to an adder and perform a shift addition using the adder.
[0059] Therefore, as shown in the "Position" row of the table in the lower half of Figure 5, the address offset value of the first tensor element (D0) is 0 + 0*2*4 + 0*2 = 0. The address offset value of the second tensor element (D1) is 1 + 0*2*4 + 0*2 = 1. The address offset value of the third tensor element (D2) is 0 + 1*2*4 + 0*2 = 8. The address offset value of the fourth tensor element (D3) is 1 + 1*2*4 + 0*2 = 9. (The rest is omitted.)
[0060] The system writes the tensor elements of the N-dimensional tensor to the corresponding memory bank in the destination memory unit according to the total memory address determined from the address offset value (step 406). In some embodiments, the total memory address of a particular tensor element is determined as the sum of the base memory address included in the instruction and the address offset value of that particular tensor element. For example, the total memory address of the fourth tensor element (D3) in the example in Figure 5 is equal to 9, which is the sum of the base memory address and the address offset value. The fourth tensor element can then be written to the corresponding memory bank in the destination memory unit having this total memory address.
[0061] In these embodiments, the total memory address can similarly be calculated using one or more adders. For example, the inputs to the adder for a particular tensor element are the base memory address and the address offset value. The output is the total memory address of the memory banks in the destination memory unit for storing the particular tensor element.
[0062] In this way, the system distributes the tensor elements of the N-dimensional tensor in a data distribution pattern defined by the first, second, and third values. As shown in the upper half of Figure 5, for example, the four memory banks of the first cell (cell 0) store the first, second, seventh, and eighth tensor elements (D0, D1, D6, and D7) of the N-dimensional tensor, the two memory banks of the second cell (cell 1) store the thirteenth and fourteenth tensor elements (D12 and D13) of the N-dimensional tensor, and the remaining two memory banks of the second cell are empty.
[0063] The system may optionally output data to another computing system indicating, for example, the address offset value, the total memory address, or both, for a specific tensor element of the N-dimensional tensor, so that the other computing system can use the address offset value, the total memory address, or both to access that specific element of the N-dimensional tensor in the destination memory unit.
[0064] This specification uses the term “configured” in relation to a system and computer program components. When one or more computer systems are configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof is installed on the system that causes the system to perform that operation or action while it is running. When one or more computer programs are configured to perform a particular operation or action, it means that one or more programs, when executed by a data processing device, contain instructions that cause the device to perform that operation or action.
[0065] The subject matter and functional embodiments described herein can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware, and include structures disclosed herein and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer program instructions, i.e., as one or more modules of computer program instructions encoded in a tangible non-temporary storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage board, a random or serial access memory device, or one or more combinations thereof. Alternatively or additionally, the program instructions can be encoded into artificially generated propagating signals, such as mechanically generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information for transmission to a receiving device suitable for execution by a data processing device.
[0066] The term "data processing device" refers to data processing hardware and encompasses all types of devices, machines, and equipment for processing data, including, for example, programmable processors, computers, or multiple processors or computers. A device may also be, or further include, special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, a device may optionally include code that constitutes the execution environment for computer programs, such as processor firmware, protocol stacks, database management systems, operating systems, or one or more combinations thereof.
[0067] Computer programs, which may be called or described as programs, software, software applications, apps, modules, software modules, scripts, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A program may, but may not, correspond to a file in a file system. A program may be stored in a single file dedicated to a program of interest, in part with other programs or data, for example, in a file containing one or more scripts stored in a markup language document, or in multiple collaborative files, for example, in a file containing one or more modules, subprograms, or parts of code. A computer program can be deployed to run on one computer, or on multiple computers located in one location or distributed across multiple locations and interconnected by a data communication network.
[0068] In this specification, the term “database” is used broadly to refer to any collection of data. The data does not need to be structured in any particular way, or not structured at all, and can be stored in one or more storage locations. Therefore, for example, an index database may contain multiple collections of data, each of which may be organized and accessed in a different way.
[0069] Similarly, in this specification, the term “engine” is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, while in other cases, multiple engines may be installed and run on the same one or more computers.
[0070] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to act on input data and produce outputs. Alternatively, the processes and logic flows may be performed by special-purpose logic circuits, such as FPGAs or ASICs, or by a combination of special-purpose logic circuits and one or more programmed computers.
[0071] A computer suitable for running computer programs can be based on a general-purpose or dedicated microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory, random-access memory, or both. The basic elements of a computer are the central processing unit for executing and running instructions, and one or more memory devices for storing instructions and data. The central processing unit and memory can be complemented or incorporated by dedicated logic circuits. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely connected to them to receive data from them, transmit data to them, or both. However, a computer is not required to have such devices. Furthermore, to give some examples, a computer can be embedded in other devices, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices, such as Universal Serial Bus (USB) flash drives.
[0072] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0073] To provide user interaction, embodiments of the subject matter described herein can be implemented in a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, to which the user can provide input to the computer. Other types of devices can also be used to provide user interaction. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and the input from the user can be received in any form, including acoustic input, voice input, or haptic input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from a web browser. The computer can also interact with the user by sending text messages or other forms of messages to a personal device, such as a smartphone running a messaging application, and receiving response messages from the user in return.
[0074] A data processing device for implementing a machine learning model may, for example, include a dedicated hardware accelerator unit for processing the general and numerical computation portions of machine learning training or operation (i.e., inference, workload).
[0075] Machine learning models can be implemented and deployed using machine learning frameworks, such as the TensorFlow framework or the JAX framework.
[0076] Embodiments of the subject matter described herein can be implemented in a computing system that includes a backend component, for example, a data server; or in a computing system that includes a middleware component, for example, an application server; or in a computing system that includes a frontend component, for example, a graphical user interface, a web browser, or a client computer having an application that enables a user to interact with embodiments of the subject matter described herein; or in a computing system that includes one or more such backend, middleware, or frontend components in any combination thereof. The components of the system can be interconnected by any form or medium of digital data communication, for example, a communication network. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and for example, the Internet.
[0077] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with each other. In some embodiments, the server sends data, such as an HTML page, to a user device for the purpose of displaying data to a user interacting with a device acting as a client and receiving user input from that user. Data generated on the user device, such as the results of user interactions, can be received by the server from the device.
[0078] While this specification includes details of many specific embodiments, these should not be construed as limiting the scope of any invention or claimable content, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Specific features described herein in the context of individual embodiments can also be implemented in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment can also be implemented separately or in any preferred subcombination in multiple embodiments. Furthermore, features may be described above as functioning in a particular combination, and even if initially claimed as such, one or more features from the claimed combination may be removed from the combination, and the claimed combination may cover a subcombination or a variation of a subcombination.
[0079] Similarly, while operations are shown in the drawings and described in a specific order in the claims, this should not be understood as requiring that such operations be performed in a specific or sequential order shown, or that all shown operations be performed, in order to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.
[0080] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the desired results can still be obtained even if the actions described in the claims are performed in a different order. As an example, the process shown in the attached diagram does not necessarily require the actions to be performed in the specific order or sequential order shown to obtain the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method performed by a computer, wherein the method is The method includes obtaining one or more instructions to distribute tensor elements of an N-dimensional tensor between destination memory units, wherein the N-dimensional tensor has a plurality of tensor elements arranged across each of the N dimensions, and N is an integer of 1 or more, and the method further includes The method includes determining the address offset value of the corresponding memory bank in the destination memory unit for each tensor element of the N-dimensional tensor using one or more counters and one or more adders, wherein the address offset value of each tensor element is determined as a linear combination of partial address offset values counted by the one or more counters, and the method further includes A method performed by a computer, comprising writing the tensor elements of the N-dimensional tensor to the corresponding memory bank of the destination memory unit according to the determined address offset value.
2. The destination memory unit includes one or more cell groups, each including one or more memory cells, each including one or more memory banks, and the instruction is, A first value that defines the number of consecutive tensor elements written to each memory cell, A second value that defines the step value taken between memory cells when writing tensor elements to different memory cells, The method according to claim 1, further comprising: a third value defining the total number of cell groups in the destination memory unit for storing tensor elements.
3. Determining the address offset value for each tensor element means that The method according to claim 2, comprising determining a fourth value that defines the total number of tensor elements to be written to each cell group by dividing the total number of tensor elements by the third value.
4. The method according to any one of claims 1 to 3, wherein the partial address offset value includes three values counted by three consecutively arranged counters.
5. The method according to claim 4, wherein the three counters include a least significant digit counter that increments by 1 in each tensor element of the N-dimensional tensor up to the first value.
6. The method according to claim 5, wherein the three counters include a second least significant counter that increments by 1 in each increment signal from the least significant counter up to the third value.
7. The method according to claim 6, wherein the three counters include a third least significant digit counter that increments by 1 in each increment signal from the second least significant digit counter until the fourth value is divided by the first value.
8. Determining the address offset value for each tensor element of the N-dimensional tensor means that The method according to claim 7, comprising using one or more adders to calculate the sum of (i) a partial address offset value counted by the least significant digit counter, (ii) a partial address offset value counted by the second least significant digit counter obtained by multiplying the second value by a constant corresponding to the number of memory banks contained in the cell, and (iii) a partial address offset value counted by the third least significant digit counter obtained by multiplying the first value by the second value.
9. Obtaining the instruction to distribute the tensor elements of the N-dimensional tensor among the memory units is: Reading the source tensor elements of a source tensor from the source memory unit, Performing a data type conversion on the aforementioned source tensor elements, The method according to any one of claims 1 to 8, comprising storing the converted source tensor elements in a shift buffer.
10. Obtaining the instruction to distribute the tensor elements of the N-dimensional tensor among the memory units is: Obtaining the base memory address value for accessing the aforementioned N-dimensional tensor, The method according to any one of claims 1 to 9, comprising storing the base memory address value in an address queue.
11. Writing the tensor elements of the N-dimensional tensor to the corresponding memory bank of the memory unit means that for each tensor element, The method according to any one of claims 1 to 10, comprising using a set of multiplexers to write the tensor elements to the corresponding memory bank identified by the determined address offset value of the tensor elements.
12. Obtaining one or more other instructions to distribute the tensor elements of another N-dimensional tensor among the destination memory units, The method according to any one of claims 1 to 11, further comprising writing the tensor elements of the other N-dimensional tensor to the memory bank of the destination memory unit, while bypassing the determination of the address offset values of the tensor elements of the other N-dimensional tensor using the one or more counters and the one or more adders.
13. The method according to any one of claims 1 to 12, further comprising outputting data indicating the determined address offset value of a specific tensor element of the N-dimensional tensor.
14. A system comprising one or more computers and one or more storage devices for storing instructions, wherein, when the instructions are executed by the one or more computers, they are operable to cause the one or more computers to perform the operation of each of the methods described in any one of claims 1 to 13.
15. One or more non-temporary computer storage media on which computer program instructions are encoded, which, when executed by multiple computers, cause the multiple computers to perform the operation of each of the methods described in any one of claims 1 to 13.