System and method for int9 quantization
The programmable hardware architecture addresses inefficiencies in machine learning systems by converting data to a symmetric int9 format, simplifying ALU calculations and reducing processing time through format-independent operations.
Patent Information
- Application Number
- JP2025033372
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-04-29
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-04-28
AI Technical Summary
Current hardware-based machine learning systems face inefficiencies in performing arithmetic logic unit (ALU) calculations due to the need to track and manage different quantized data formats, leading to increased complexity and latency, especially when mixing format types, and integer rescaling is time-consuming.
A programmable hardware architecture that converts data to a symmetric int9 format, eliminating the need to track format types and enabling efficient ALU operations by extending int8 or uint8 data to int9, allowing for simplified processing and reduced latency.
This approach simplifies ALU calculations by eliminating the need to track format types, reduces processing time, and enhances the efficiency of machine learning operations by using a symmetric quantization method.
Smart Images

Figure 2025102765000001_ABST
Abstract
Description
Background Art
[0001] Hardware-based machine learning (ML) systems typically include multi-core / subsystems (blocks and tiles), each having its own processing unit and on-chip memory (OCM). The ML system can process quantized numerical values for various calculations. For example, quantized data stored in a memory unit, such as double data rate (DDR) memory, can be transmitted to the processing tile so that the data can be processed by the processing unit for various ML operations.
[0002] In general, floating-point numbers (data) are converted, for example, to a quantized data format for storage in DDR and for subsequent processing by, for example, an ML system. The quantized format can include, but is not limited to, signed integers, unsigned integers, etc., used in arithmetic logic unit (ALU) calculations. For various calculations, a mixture of quantized format types is often used by, for example, an ML system.
[0003] Unfortunately, at present, there is no mechanism for performing ALU calculations on a mixture of quantized format types without introducing an offset. Therefore, the format type of each operand is tracked when a mixture of quantized format types is used in ALU calculations, increasing the complexity and latency associated with the ALU calculations. Further, in an ML system, integer values may need to be rescaled before being input to the processing unit. However, mathematical division in a processor is often time-consuming and inefficient in terms of time.
[0004] The foregoing examples of the related art and the limitations related thereto are for illustrative purposes only and not exclusive. Other limitations of the related art will become apparent from the interpretation obtained from the specification and the knowledge obtained from the drawings.
Brief Description of the Drawings
[0005] When read in conjunction with the accompanying figures, the aspects of the present disclosure are best understood from the following detailed description. Note that, in accordance with the convention in the industry, various features are not drawn to scale. In fact, the dimensions of the various features may be arbitrarily increased or decreased for clarity of explanation.
[0006]
Figure 1
[0007]
Figure 2A
Figure 2B
Figure 2C
Figure 2D
[0008]
Figure 3A
Figure 3B
[0009]
Figure 4
[0010]
Figure 5
Best Mode for Carrying Out the Invention
[0011] In the following description, many different embodiments or examples are provided so as to implement different features of the present subject matter. Specific examples of components and arrangements are described below to simplify the present disclosure. These are, of course, merely examples and are not intended to be limiting. Further, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for the purpose of simplicity and clarity and does not in itself define a relationship between the various embodiments and / or configurations described.
[0012] Before various embodiments are described in more detail, it should be understood that such embodiments are not limiting since elements in such embodiments can be different. It should also be understood that the specific embodiments described and / or illustrated herein can be readily separated from a particular embodiment and, optionally, combined with any of several other embodiments or replaced with elements in any of several other embodiments described herein. It should also be understood that the terms used herein are for the purpose of describing certain concepts and are not intended to be limiting. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood in the technical field to which the embodiments belong.
[0013] A new programmable hardware architecture for machine learning (ML) is proposed, including at least a host, a memory, cores, a data streaming engine, an instruction streaming engine, and an interface engine. The memory is configured to store floating-point numbers in a quantized format including, but not limited to, int8, uint8, etc. According to some embodiments, the quantized data stored in the memory is converted to the int9 format to represent different quantized data format types, such as uniformly int8, uint8, etc., and similarly, to provide symmetric quantization of the data (i.e., quantization is symmetric about zero) while eliminating the need to perform offset calculations. Converting the data to the int9 format type simplifies the complexity by enabling the inference engine to perform ALU calculations on homogeneous int9 format type operands without the need to keep track of the format type for the quantized operands, and similarly, as a result, it is understood that it brings a faster processing time.
[0014] In some embodiments, when data is read from a memory unit, e.g., DDR, an 8-bit numerical value is converted to the int9 format type based on, as a non-limiting example, whether the numerical value stored in the memory unit is int8 or uint8. In some embodiments, 9 bits are used, the int8 format type is a sign-extended version to the int9 format type, and the uint8 format type is copied to the least significant bits of the 9-bit data, and its most significant bit (i.e., bit order 9) is unsigned, e.g., set to zero. Since 8 bits are not sufficient to store the full int9 range, it is understood that a software component can ensure that the int9 values written to a memory unit, e.g., 8-bit DDR, are within the appropriate range of int8, uint8, etc.
[0015] In some embodiments, the software component performs an operation that restricts the range of int9 values to be within the range of int8, uint8, etc., or alternatively, performs an operation on the int9 value to represent it as two int9 values, one within the int8 range and the other within the uint8 range. Representing an int9 value as two int9 values, one within the int8 range and the other within the uint8 range, enables the least significant bits of the int9 value to be copied to the 8-bit DDR while preserving the information.
[0016] In some embodiments, the core of the programmable hardware architecture interprets a plurality of ML commands / instructions for the ML operations and / or data received from the host, and configures the activities of the streaming and inference engines based on the data in the received ML commands. The inference engine may include a cryptographic operation engine and an irregular operation engine. The cryptographic operation engine is an engine optimized for efficiently processing encrypted data with normal operations, such as matrix operations, such as multiplication, matrix operations, tanh, sigmoid, etc. On the other hand, the irregular operation engine is an engine optimized for efficiently processing sporadic data with irregular operations, such as memory transpose, operations on irregular data structures (e.g., trees, graphs, and priority queues). According to some embodiments, the core may adjust some of the instructions received from the host to be processed. In some embodiments, the core may be a general-purpose processor, such as a CPU.
[0017] In some embodiments, the core is specifically configured to split a plurality of ML commands for efficient execution between the core and the inference engine. The ML commands and related data to be executed by the inference engine are transmitted from the core and memory to the instruction streaming engine and data streaming engine for efficient streaming to the inference engine. As described above, the data read from the memory unit is converted to the int9 format. The data and instruction streaming engine is configured to send one or more data streams and ML commands to the inference engine in response to programming instructions received from the core. The inference engine is configured to process the instruction / data stream received from the data / instruction streaming engine for ML operations according to the programming instructions received from the instruction / data streaming engine.
[0018] It is understood that the data input to the dense operation engine of the inference engine may need to be rescaled prior to certain operations, such as tanh, sigmoid, etc. To rescale in an efficient manner, the data input in the int32 format is multiplied by an integer scale value and later shifted. To achieve the highest possible precision and lowest possible error in the calculations, the relationship between the integer scale value and the shift value is obtained based on the size of the register storing the integer data, e.g., int32.
[0019] Referring now to FIG. 1, an example diagram of a hardware-based programmable system / architecture 101 configured to support machine learning is shown. The diagram shows the components as functionally separated, but such a representation is for illustrative purposes only. It is clear that the components depicted in this figure can be arbitrarily combined and divided into separate software, firmware, and / or hardware components. Further, it is clear that such components can be executed on the same host or multiple hosts, regardless of how they are combined or divided, and multiple hosts can be connected by one or more networks. Each of the engines within architecture 101 is a dedicated hardware block / component that includes one or more microprocessors and an on-chip memory unit that stores data and software instructions programmed by the user for various machine learning operations. As will be described in detail below, when the software instructions are executed by the microprocessor, each of the hardware components becomes a dedicated hardware component for training certain machine learning functions. In some embodiments, architecture 101 is a single chip, e.g., a system-on-chip (SOC).
[0020] In the example of FIG. 1, architecture 101 may include a host 110 coupled to a memory (e.g., DDR) 120 and a core engine 130. Memory 120 may be coupled to a direct memory access (DMA) engine (not shown) and a network interface controller (NIC) (not shown) to receive external data. Memory 120 may be internally connected to a data streaming engine 140. Core 130 is coupled to an instruction streaming engine 150 that is coupled to data streaming engine 140. Core 130 is also coupled to a general-purpose processor 165. In some embodiments, general-purpose processor 165 may be part of core 130. Instruction streaming engine 150 and data streaming engine 140 are coupled to an inference engine 160 that includes a cryptographic operation engine 161 and an irregular operation engine 163. It is understood that inference engine 160 may include an array for performing various calculations. It is understood that any description of the array for performing various calculations in inference engine 160 is for illustrative purposes and should not be construed as limiting the scope of the embodiments. For example, in some embodiments, the array for performing various calculations may exist outside of inference engine 160.
[0021] It is understood that the external data may be in a floating-point format, for example, 32-bit floating-point. Thus, if the data is stored in memory 120, for example, 8-bit DDR, the data may be converted to an integer format type, for example, int8, uint8, etc. It is understood that uint8 is in the range from 0 to 255, while int8 is in the range from -128 to 127. On the other hand, int9 is in the range from -256 to 255, and thus can represent both int8 and uint8 without any offset calculation. Using int9 as the uint8 range and the int8 range enables the data to be copied to standard 8-bit DDR. The description regarding the use of 32-bit floating-point and 8-bit DDR is for illustrative purposes and should not be construed as limiting the scope of the embodiments. The floating-point data is ultimately quantized to int9 instead of int8 or uint8. Further, since the int9 range covers both positive and negative values, it results in a zero offset, further simplifying the rescaling of int9 numerical values in the ML system. Thus, when data is being read from memory 120, for example, 8-bit DDR, the data is converted to the int9 format. When the data is converted to the int9 format, it is understood that there is no need to track the type of operand when a mixture of different format types is used in the calculation. For example, using int9 eliminates the need to track whether the operand in the executed calculation is int8, uint8, etc.
[0022] In some embodiments, it is further understood that the memory 120, e.g., DDR, may store a floating-point number, e.g., a 32-bit floating-point number, as four 8-bit values. Thus, when data is read from the memory 120, e.g., an 8-bit DDR, into the on-chip memory, quantization is performed from a 32-bit floating-point number to int9 in either the general-purpose processor 165 or the irregular arithmetic engine 163. In some embodiments, the registers in the general-purpose processor 165 and / or the irregular arithmetic engine 163 store a 32-bit width that holds a 32-bit floating-point value. Thus, for use in the ML system, the floating-point number may be converted to an int9 value. However, the 32-bit floating-point number is first scaled to convert it to the int9 format. For example, the appropriate scale may be as follows. Scale = (upper limit range of the floating-point number - lower limit range of the floating-point number) / (upper limit range of int9 - lower limit range of int9) = (End - (-End)) / (255 - (-255)) = 2End / (2(255)) = End / 255.
[0023] It is understood that the same scale can be used when it is extended to include -256 of the lower limit range of int9. FIG. 2A shows the symmetric quantization and mapping of a 32-bit floating-point number to the full range of int9, while FIG. 2B shows the mapping of FIG. 2A to include -256 of int9. As shown, the same scale may be used for both FIGS. 2A and 2B. FIG. 2C shows mapping a 32-bit floating-point number to an int8 representation within int9 and the int9 range. It is understood that the scale for the int8 representation within the int9 range is different from the scale determined above. In some embodiments, the scale for representing int8 using 9 bits may be as follows. Scale = (upper limit range of the floating-point number - lower limit range of the floating-point number) / (upper limit range of int8 - lower limit range of int8) = (End - (-End)) / (127 - (-127)) = End / 127.
[0024] Figure 2D shows mapping a 32-bit floating point number to an int9 and a uint8 representation within the int9 range. It is understood that the uint8 representation within the int9 range has the same scale as those in FIGS. 2A and 2B.
[0025] In some embodiments, when transferring data from the memory 120 to an array, such as the inference engine 160, etc., the data being transferred is either an extended sign or extended zero depending on whether the data being transferred is int8 or uint8. That is, the data is converted from one format type, such as int8, uint8, etc., to another format type, such as int9. As a non-limiting example, when converting data from int8 or uint8 to the int9 format, 8-bit data is converted to 9-bit data by extending the number of bits by 1 bit. It is determined whether the data to be converted is signed, such as int8, or unsigned, such as uint8. If the data to be converted is signed, the most significant bit of the 9-bit data for int9 is the extended sign, and if the data to be converted is unsigned, the most significant bit of the 9-bit data for int9 is set to zero. It is understood that int8 or uint8 is directly copied to the lower bits of the int9 data (i.e., lower 8-bit order). It is understood that the int9 data may be referred to as extended data compared to the int8 or uint8 format type. In this example, the extended data of the int9 format type is stored in the inference engine 160 that is the operand for the operation. In some embodiments, the extended data may be stored in the on-chip memory (OCM) of the inference engine 160 to be processed by the processing tiles of the ML computer array. In some embodiments, it is understood that a floating-point number, such as a 32-bit floating-point number, may be converted to an integer representation, such as int9. In one exemplary embodiment, the floating-point number is appropriately quantized and scaled as illustrated in FIGS. 2A to 2D so as to be converted to the int9 format type. As illustrated, one scaling value may be used to represent the floating-point number for int8 of the int9 format type, while a different scaling value may be used to represent the floating number for uint8 of the int9 format type.It is understood that the 16-bit floating-point numbers stored in the memory unit 120, e.g., DDR, remain the same when stored in the OCM of the inference engine 160 from the memory unit 120. As a non-limiting example, the lower 7 bits of the 16-bit floating-point number are the same as the lower 7 bits of its mantissa, the 8th bit is extended but not used in the operation, and the remaining mantissa bits are followed by the exponent after the 9th and 10th bits, and after the signed bit, there are additional extended bits that are not used in any operation.
[0026] It is understood that the inference engine 160 may include a plurality of processing tiles arranged in a plurality of rows and columns, e.g., an 8-row × 8-column two-dimensional array. Each processing tile may include at least one OCM, one POD unit, and one processing engine / element (PE). Here, the OCM within the processing tile is configured to receive data from the data streaming engine 140 in a streaming manner. As described above, it is understood that the received data may be in the int9 format. The OCM enables efficient local access to the data for each processing tile. The processing units, e.g., POD and PE, are configured to perform dense or sparse calculations of ML operations, respectively, on the data received in the OCM.
[0027] It is understood that the OCM of each processing tile may receive int9 format type data for various ALU operations associated with ML operations. In some embodiments, the format type of the data stored in the memory 120, e.g., whether it is signed or unsigned, is tracked such that appropriate instructions can be scheduled to be streamed for execution by appropriate processing units, e.g., each POD / PE of the processing tile. That is, various ALU operations are performed on the data received in int9 format by the processing tile. The data received in int9 format may be operands of various ALU operations. The results of various ALU operations in int9 format type may be stored in their respective OCMs.
[0028] In some embodiments, the inference engine 160 includes a dense operation engine 161 optimized to efficiently process dense data, e.g., data received from the memory 120 in int9 format, with normal operations such as matrix operations, e.g., multiplication, matrix manipulation, tanh, sigmoid, etc. On the other hand, the inference engine 160 may also include an irregular operation engine 163 optimized to efficiently process sporadic data of int9 format type with irregular operations such as memory transpose, additional operations, operations on irregular data structures (e.g., trees, graphs, and priority queues). According to some embodiments, the core 130 may condition some of the instructions received from the host 110 processed by a general-purpose processor 165, e.g., a CPU, etc.
[0029] In some embodiments, core 130 is configured to execute any software code written through a general high-level language. Core 130 is configured to process a plurality of performance non-critical operations, such as data / instruction pre-work, data collection, data mapping, etc. In some embodiments, performance non-critical operations can be processed by core 130, and performance critical operations (e.g., matrix multiplication) can be processed by inference engine 160. Core 130 may be configured to classify received ML commands into performance critical and non-critical operations / tasks. That is, core 130 is configured to divide a plurality of ML commands between core 130 and inference engine 160 for their efficient execution. In some embodiments, core 130 may be configured to allocate / divide a plurality of ML commands (also referred to as tasks or subtasks) to various components, such as inference engine 160, for processing. In some embodiments, core 130 is configured to allocate one or more locations in memory 120 for storing tasks / commands, data, results after data is processed, etc., so that they can be accessed and used by other components in core 130 or architecture 101, such as inference engine 160. Therefore, rather than relying on or requiring host 110 to execute certain ML commands or operations, core 130 and inference engine 160 are configured to execute the entire ML algorithm and the operations therewith. By supporting and executing the entire ML operations in programmable hardware architecture 101, core 130 eliminates the performance overhead of transferring data to host 110 and back to execute any unsupported ML operations, reduces the burden on host 110, and achieves higher performance.
[0030] In some embodiments, the ML commands and associated data, for example, in their int8 format, executed by the inference engine 160 are transmitted from the core 130 and the memory 120 to the instruction streaming engine 150 and the data streaming engine 140 for efficient streaming to the inference engine 160. In some embodiments, the data / instruction streaming engines 140-150 are configured to transmit one or more data streams and programming instructions to the inference engine 160 in response to the ML commands received from the core 130. In some embodiments, it is understood that the format type of the data stored in the memory 120, for example, whether it is signed or unsigned, is tracked such that appropriate instructions can be scheduled to be streamed to the cryptographic operation engine 161 and / or the irregular operation engine 163 of the inference engine 160. That is, various ALU operations are performed on the data received in the int9 format by the engines within the inference engine 160. The data received in the int9 format may be operands for various ALU operations. The results of the various ALU operations in the int9 format type may be stored within the cryptographic operation engine 161 and / or the irregular operation engine 163 of the inference engine 160. In some embodiments, the results may be stored in the appropriate OCM of the processing tiles of the ML computer array.
[0031] In some embodiments, it is understood that the result of the ALU operation within the inference engine 160 is stored in the memory component of each processing tile within the inference engine 160, e.g., within the OCM. The result stored in the inference engine 160 may be transmitted for storage to the memory unit 120, e.g., the DDR. However, before storing the result, if the value of the result exceeds the upper limit of the data format type within the memory unit 120, e.g., the maximum value, the value may be adjusted to the upper limit range for the data, and if the value of the result is less than the lower limit range of the memory unit 120, the value may be adjusted to the lower limit range for the data, e.g., the minimum. When storing the result from the OCM of each processing tile to the memory unit 120, it is understood that the most significant bit of the int9 result is dropped.
[0032] In some embodiments, it is understood that the results of the processes stored in each OCM may be transmitted for storage and returned to the memory unit 120, e.g., DDR. However, before storing the results, if the value of the results exceeds the upper limit of the data format type within the memory unit 120, e.g., the maximum value, the value may be adjusted to the upper limit range for the data, and if the value of the results is less than the lower limit range of the memory unit 120, the value may be adjusted to the lower limit range for the data, e.g., the minimum. That is, the data may be clamped to be within an appropriate range, e.g., within the int8 range, uint8 range, etc. When storing the results from the OCM of each processing tile to the memory unit 120, it is understood that the most significant bit of the int9 results is dropped. Further, when transferring data from the OCM of each inference engine, e.g., inference engine 160, to the memory unit 120, e.g., DDR, the software module may track whether the data stored in the memory unit 120 is signed or unsigned so that the int9 data format type can be accurately interpreted as int8 for data that was in the int8 format within the memory unit 120, uint8 for data that was in the uint8 format within the memory unit 120, etc.
[0033] Referring now to FIGS. 3A and 3B, as described in FIGS. 1 - 2D, an example of a process is shown that supports converting data stored in a memory, e.g., DDR, from a first format, e.g., int8, uint8, floating point, etc., to a second format type, e.g., int9. The figures show functional steps in a particular order for purposes of illustration, but the process is not limited to any particular order or arrangement of steps. One of ordinary skill in the art will understand that the various steps depicted in this figure can be omitted, rearranged, combined, and adapted in various ways.
[0034] As shown in FIGS. 3A and 3B, at stage 310, the number of bits stored in memory unit 120, e.g., in DDR, is extended by one bit to form extended data, e.g., int9. Thus, int8 or uint8, which contains 8 bits, is extended to 9 bits. In some embodiments, it is understood that the data stored in memory unit 120 is a floating-point number. At stage 320, it is determined whether the data stored in memory 120 is signed, e.g., int8, or unsigned, e.g., uint8. At stage 330, in response to determining that the data is signed, the extended data is sign-extended. On the other hand, at stage 340, in response to determining that the data is unsigned, the most significant bit of the extended data is set to zero. At stage 350, the data is copied to the lower bits, and thus to all bits except the most significant bit. At stage 360, the extended data is copied to inference engine 160, e.g., to the OCM of inference engine 160. At stage 370, it is understood that whether the data stored in memory unit 120, e.g., in DDR, is signed or unsigned is tracked, and thus at stage 380, appropriate instructions for the extended data are scheduled. At stage 382, various ALU operations may be performed on the extended data. At stage 384, the result of the ALU operation is stored in the OCM. At stage 386, the result of the ALU operation stored in the OCM is also stored / copied to memory unit 120, e.g., in DDR. At stage 388, the most significant bit of the result is dropped before storing the result from the OCM to the DDR. Optionally, at stage 390, it is understood that based on the range of the numeric format type stored in memory unit 120, the value of the result of the ALU may be adjusted before storing from the OCM to memory unit 120, e.g., in DDR. For example, if the value of the result stored in the OCM exceeds the upper limit range of the numeric type stored in memory unit 120, e.g., int8 or uint8, etc., the result is adjusted and modified for the maximum or upper limit range of the numeric, e.g., int8, uint8, etc.
[0035] Figure 4 shows a diagram of an example architecture of a POD. It is understood that the number of components, the size and number of bits of the components, the matrix size, etc. shown in Figure 4 are for illustrative purposes and are not intended to limit the scope of the embodiments. In the following description, matrix multiplication is used as a non-limiting example, but it is understood that the POD may also be configured to perform other types of dense computational tasks of ML operations. In the example of Figure 4, the POD includes a computational POD instruction control 699 configured to control the loading of data / instructions to various components, such as registers, tanh / sigmoid units 614, etc. The POD includes a matrix multiplication block 602 that is a two-dimensional array having X rows and Y columns, and it is understood that each element / cell in the array has a certain number of registers (e.g., MIPS or microprocessors without using interlocked pipeline stages). The matrix multiplication block 602 is configured to multiply two matrices, a matrix A composed of X rows and Z columns and a matrix B composed of Z rows and Y columns, to generate a matrix C composed of X rows and Y columns. Even if the data stored in the memory unit 120 is of different types of formats, such as int8, uint8, floating point, etc., it is understood that the data to be multiplied may be of the int9 format type stored in their respective OCMs.
[0036] In the example of FIG. 4, the POD further includes three types of registers that supply matrix data to the matrix multiplication block 602 for matrix multiplication, namely, the A register 604, the B register 606, and the C register 608. The A register 604 includes a bank of registers, for example, m registers, each configured to hold a row / column of the A matrix that is supplied to a column of the array of matrix multiplication blocks 602. Each A register may have a plurality of entries, for example, k elements, each having a certain number of bit widths and supporting a certain read or write operation per cycle. It is understood that the data may be in the int9 format type in each register even if the data stored in the memory unit 120 is of different format types, such as int8, uint8, floating type, etc. That is, the data is converted from a certain format of the memory unit 120 to a different format type, such as int9, so as to be stored in the respective OCMs of the processing tiles used in the ALU calculations of the PE and / or POD operations. The entries enable each A register to prefetch the next portion of the A matrix before they are needed for the calculations by the matrix multiplication block 602. The B register 606 includes a bank of registers, for example, n registers, each configured to hold a row / column of the B matrix that is supplied to a row of the array of multiplication blocks 602. Similar to the A register 604, each B register may have a plurality of entries, for example, k elements, each having a certain number of bit widths and supporting a certain read or write operation per cycle. The entries enable each B register to prefetch the next portion of the B matrix before they are needed for the calculations by the matrix multiplication block 602. The C register 608 is configured to hold the result of the matrix-multiplication - the C matrix - generated by the multiplication block 602. The C register 608 includes a plurality of banks, each configured to hold a row / column of the C matrix. The C matrix is configured to have m×n components.
[0037] During matrix multiplication processing, the matrix multiplication block 602 is configured to load the components of matrices A and B into the A and B registers only once from the OCM (instead of reading each row or column of the matrix), thus saving the memory access time to the OCM. Specifically, the multiplication operation of each matrix has a unique structure, where the rows of the first matrix are multiplied by all the columns in the second matrix, and the columns in the second matrix are multiplied by all the rows in the first matrix. Since the matrix multiplication block 602 executes the matrix multiplication operation, the rows of the A register 604 remain the same, while the columns of the B register 606 are supplied to the matrix multiplication block 602 one by one so as to be multiplied by the rows in the A register 604. At the same time, the columns of the B register 606 remain the same, while the rows of the A register 604 are supplied to the matrix multiplication block 602 one by one so as to be multiplied by the columns of the B register 606. Therefore, the matrix multiplication block 602 is configured to simultaneously multiply each row of the first matrix by all the columns of the second matrix and multiply each column of the second matrix by all the rows of the first matrix. These outputs from these multiplications are accumulated and stored in the C register until the matrix multiplication process is completed.
[0038] As shown in the example of FIG. 4, the A register 604, B register 606, and C register 608 are each associated with a corresponding OCM streamer 603, 605, or 607, respectively, and each of the OCM streamers is programmed to stream data from the OCM to the corresponding register to ensure that matrix multiplication operations can be performed by the matrix multiplication block 602 in a simplified manner. Each OCM streamer has an address range of the OCM to be read and a stride that is tracked for the next read. Registers of type A or B are configured to send a per-bank next line ready signal to their corresponding streamer, and the bit pattern of the signal indicates which bank requests the next line of data. The corresponding streamer of the A or B register responds to the read signal by sending the corresponding line of data from the OCM to the register. The streamer sends a completion signal to its corresponding register when it has sent the last line of data to be transmitted. If all of the banks of the register have lines of data, the A or B register sends a ready signal to the matrix multiplication block 602 indicating that the next set of A or B registers is ready to be read into the matrix multiplication block 602 for matrix multiplication. In some embodiments, each register bank has a valid bit that notifies the matrix multiplication block 602 which values are valid and should be operated on.
[0039] When the matrix multiplication is completed, for example, when the end of a row of matrix A and the end of a column of matrix B are reached, the matrix multiplication block 602 notifies the C register 608 that all the accumulations within the entries of the C register 608 are completed and the entries are ready to be written back to the OCM via their corresponding streamers 607. Each bank of the C register 608 then sends the data to the OCM. If the OCM is not ready to receive the data from the bank of the C register 608, the transmission is postponed until the PE is ready to receive the data from the bank and is retried in the next cycle. In some embodiments, the C register 608 is pre-loaded with data or reset to zero prior to the next set of accumulations during the next matrix multiplication operation. Such pre-loading enables adding a bias as part of the next matrix multiplication. In some embodiments, each PE is configured to receive the output C matrix from the matrix multiplication block 602 of the POD, process it, and write it to the OCM.
[0040] According to one example, the result of the process stored in each OCM may be transmitted for storage and returned to the memory unit 120, for example, DDR. However, before storing the result, if the value of the result exceeds the upper limit of the format type of the data in the memory unit 120, for example, the maximum value, the value may be adjusted to the upper limit range for the data. If the value of the result is smaller than the lower limit range of the memory unit 120, the value may be adjusted to the lower limit range for the data, for example, the minimum. That is, the data may be clamped to be within an appropriate range, for example, within the int8 range, uint8 range, etc. It is understood that when storing the result from the OCM of each processing tile to the memory unit 120, the most significant bit of the int9 result may be dropped. Further, when transferring data from the OCM of each inference engine, for example, the inference engine 160, to the memory unit 120, for example, DDR, the software module may track whether the data stored in the memory unit 120 was signed or unsigned so that the int9 data format type can be accurately interpreted as int8 for data that was in the int8 format in the memory unit 120, uint8 for data that was in the uint8 format in the memory unit 120, etc.
[0041] In some embodiments, the inference engine 160 is configured to fuse / integrate these subsequent matrix multiplication operations by each PE with the corresponding matrix multiplication operations by the PODs, so that it first transmits and stores the output to the OCM and then immediately executes these subsequent matrix multiplication operations on the output from the matrix multiplication block 602 without reading the C matrix from the OCM again for these subsequent matrix multiplication operations. By bypassing the round trip to the OCM, the fusion of the subsequent matrix multiplication operations with the matrix multiplication operations saves time and improves the efficiency of the inference engine 160. For example, in some embodiments, it is understood that additional normal operations, such as rectified linear unit (ReLU), quantization, etc., may be required for the output C matrix. Thus, a switching mechanism may be integrated within the POD architecture to determine whether additional normal operations are required, and if so, instead of writing the output C matrix to another memory location, the output can be manipulated. For example, if a rectified linear operation is required, the output C matrix is streamed to a ReLU unit 601 configured to perform a ReLU operation on the C matrix. Similarly, if quantization is required, the output C matrix or the output of the ReLU unit 601 is streamed to a quantization unit 612 configured to quantize the result from the C matrix or the ReLU operation.
[0042] In some embodiments, the scale value, shift value, and / or offset value required for quantization / resquantization operations may be statically set by core 130 and may be different from different ML operations. In some embodiments, these values may be part of the ML model downloaded to the core, and the values corresponding to the ML operations may be read from the model and written to appropriate registers before the quantization operation begins. For the input to quantization 612 and / or the tanh / sigmoid unit 614, and later, for direct storage to their respective OCM blocks, it is understood that resquantization performs rescaling of the output values stored in the C register 608. It is understood that resquantization may be performed on the output data, e.g., the C register 608 in this example, but in other examples, resquantization can be performed on other outputs from other registers. Therefore, performing resquantization on the data stored in the C register 608 is for illustrative purposes and should not be construed as limiting the scope of the embodiments. In some embodiments, a single scaling value is applied to all elements of the output. It is understood that the scaling operation, which is a division operation, may be replaced by integer multiplication and shift operations. It is further understood that the relationship between the value of the integer multiplication (also referred to as the integer scale value) and the shift value determines the accuracy and error within the system. In some embodiments, the relationship between the integer scale value and the shift value is obtained, and the largest possible value for the integer scale value and its corresponding scale value is selected based on the size of the register that stores the result of the multiplication (the multiplication of the output from the C register 608 and the integer scale value). In some embodiments, the output from the C register 608 may be denoted as V, and the quantization multiplier may be denoted as x, where x can be greater than or less than 1. It is understood that the relationship between the integer scale value and the shift value determines the quantization multiplier. The relationship between the integer scale value and the shift value is approximately given by the following equation (1). x ~ integer scale value / (2 シフト値 ) (1). Therefore, integer scale value = int(x * 2シフト値 ) It is (2). It is understood that the largest integer scale value is limited by the size of the register that holds the result of the integer multiplication, and thus the integer scale value limits the output of the C register 608, for example, the V value. For example, if V is 32 bits and the register size is 64 bits, the integer scale value must be less than the largest 32-bit integer, or else there will be an overflow. That is, the largest possible value is 2,147,483,647. The largest possible values for other sizes may be different, and it is understood that the examples provided above are for illustrative purposes only and not intended to limit the scope of the embodiments. Therefore, the condition shown in the following formula (3) will be met. Integer scale value / Largest possible value < 1 (3)
[0043] In some embodiments, to obtain the largest possible integer scale value, formulas (2) and (3) are repeatedly performed as a whole. First, the shift value is 0, and with each iteration, the shift value is incremented by a value, for example, 1, 2, 5, 6, 7, 11, etc. The shift value determines the possible integer scale value, and the iteration is performed one more time as long as the condition specified by formula (3) holds. The process is repeated until formula (3) is no longer true, at which point the previous shift value and its corresponding integer scale value are selected. It is understood that either the previous shift value and its corresponding integer scale value can be selected, even if the largest previous integer scale value and its corresponding scale value provide the highest accuracy given the register size. The above process of selecting the largest possible integer scale value and its corresponding scale value is shown in Python.
Table 1
[0044] It will be appreciated that when an integer scale value and its corresponding scale value are selected, a quantization / requantization operation may be performed. The output of the C register 608 is multiplied by the integer scale value. The result of the multiplication is shifted by a shift value as selected above to form scaled integer data. When the data is scaled, additional operations may be performed, such as a tahn operation, a sigmoid operation, a rounding operation, a clipping / clamping operation, etc. In some embodiments, a rounding operation is performed by considering the most significant bits that are dropped due to the shift operation, and the remaining result is rounded based on the dropped most significant bits. It will be appreciated that the scaled integer data may be further adjusted based on the range for the integer data. For example, if the integer data stored in the memory unit 120 is int8, and the scaled integer data exceeds the upper limit of int8, the scaled integer data is changed and adjusted to the maximum or upper limit of int8. Similarly, if the integer data stored in the memory unit 120 is uint8, and the scaled integer data exceeds the upper limit of uint8, the scaled integer data is changed and adjusted to the maximum or upper limit of uint8. On the other hand, if the scaled integer data has a value lower than the minimum or lower range of the data stored in the memory unit 120, e.g., int8 or uint8, the scaled integer data is adjusted and changed to the minimum or lower range of the integer data within the memory unit 120.
[0045] Referring now to FIG. 5, a method for rescaling integer data in a machine learning operation is shown. It is understood that the method illustrated in FIG. 5 is a flow of the method for the operation as described in FIG. 4. At stage 510, as described in equation (1), the relationship between the integer scale value and the shift value is determined. At stage 520, the condition shown in equation (3) is no longer true, and thus, the shift value is iteratively increased until the value is greater than or equal to 1, and the corresponding integer scale value is obtained with respect to equation (2). At stage 530, before equation (3) is no longer true, the shift value and its corresponding integer scale value are selected. In some non-limiting examples, it is understood that stages 510 to 530 are executed during the compilation stage and before any inference by the inference engine 160. At stage 540, an integer value, for example, in int32 format, is received from, for example, the C register 608. At stage 550, the received integer value is multiplied by the selected integer scale value. At stage 560, the result of the multiplication is shifted by the shift value corresponding to the selected integer scale value. At stage 570, further operations, such as tanh, sigmoid, rounding, clipping, clamping, etc., may be performed. At stage 580, the value of the scaled integer data may be adjusted based on the range of the integer data stored in the memory unit 120. For example, when int8 type data is stored in the memory unit 120, for example, in DDR, then, if the scaled integer data exceeds the upper limit of the int8 data type, the scaled integer data is changed to the maximum or upper limit value of the int8 type data. Similarly, when uint8 type data is stored in the memory unit 120, for example, in DDR, then, if the scaled integer data exceeds the upper limit of the uint8 data type, the scaled integer data is changed to the maximum or upper limit value of the uint8 type data. On the other hand, if the scaled integer data is less than the lower limit of the int8 data type, the scaled integer data is changed to the minimum or lower limit of the int8 data type stored in the memory unit 120, for example, in DDR.Similarly, if the scaled integer data is less than the lower limit of the uint8 data type, the scaled integer data is changed to the minimum or lower limit of the uint8 data type stored in the memory unit 120, e.g., DDR. Accordingly, higher precision and accuracy are achieved based on the size of the register size.
[0046] The foregoing description of various embodiments of the subject matter according to the claims is provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the subject matter of the claims to the exact forms disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling others of ordinary skill in the art to understand the invention for various embodiments and the various modifications suitable for the particular uses contemplated. [Other possible items] [Item 1] A method for converting data stored in a memory from a first format to a second format for machine learning (ML) operations, extending the number of bits of the data stored in a double data rate (DDR) memory by one bit to form extended data; determining whether the data stored in the DDR memory is signed or unsigned data; in response to determining that the data is signed, adding a sign value to the most significant bit of the extended data and copying the data to the lower bits of the extended data; in response to determining that the data is unsigned, copying the data to the lower bits of the extended data and setting the most significant bit to an unsigned value; storing the extended data in on-chip memory (OCM) of a processing tile of a machine learning computer array and comprising. [Item 2] The method according to item 1, wherein the data is an unsigned integer. [Item 3] The method according to item 1, wherein the data is a signed integer. [Item 4] The method according to any one of items 1 to 3, wherein the data is 8 bits and the extended data is 9 bits. [Item 5] The method according to any one of items 1 to 4, wherein the extended data is int9 data. [Item 6] Tracking whether the data stored in the DDR memory is signed or unsigned; Scheduling appropriate instructions for the extended data based on whether the data is signed or unsigned; The method according to any one of items 1 to 5, further comprising: [Item 7] The method according to item 6, further comprising executing an arithmetic logic unit (ALU) operation on the extended data as an operand. [Item 8] The method according to item 7, further comprising storing the result of the operation in the OCM of the processing tile of the machine learning computer array. [Item 9] The method according to item 8, further comprising storing the result stored in the OCM in the DDR memory. [Item 10] Before storing the result in the DDR memory, if the value of the result exceeds the maximum value of the range for the data, adjusting the value of the result to the maximum value, and if the value of the result is lower than the minimum value of the range for the data, adjusting the value of the result to the minimum value of the range for the data. The method according to item 9. [Item 11] Before storing the result from the OCM in the DDR memory, further comprising dropping the most significant bit of the result. The method according to item 9 or 10. [Item 12] The data stored in the DDR memory is the integer representation of the data regarding the floating-point data, according to the method described in any one of Items 1 to 11. [Item 13] The floating-point data is scaled and quantized to form the data in the first format, according to the method described in Item 12. [Item 14] A first scaling value is used to convert the floating-point data into the int8 format, and a second scaling value is used to convert the floating-point data into the uint8 format, according to the method described in Item 12 or 13. [Item 15] A double data rate (DDR) memory configured to store integer data in a first format, A machine learning processing unit having a plurality of processing tiles and comprising: Each processing tile has an on-chip memory (OCM) configured to receive and maintain extended data converted from the integer data in the first format from the DDR memory for various ML operations. The extended data includes one additional bit compared to the integer data in the first format. When the integer data in the first format is signed, the most significant bit of the extended data is signed. When the integer data in the first format is unsigned, the most significant bit of the extended data is set to an unsigned value. The least significant bit of the extended data is the same as the integer data in the first format. [Item 16] The integer data in the first format is either int8 or uint8, according to the system described in Item 15. [Item 17] The extended data is int9, according to the system described in Item 15 or 16. [Item 18] Whether the integer data in the first format stored in the DDR memory is signed or unsigned is tracked, and appropriate instructions are scheduled according to whether the integer data in the first format is signed or unsigned. The system according to any one of items 15 to 17. [Item 19] The extended data is an operand for an operation. The system according to any one of items 15 to 18. [Item 20] The result of the operation is stored in the OCM. The system according to item 19. [Item 21] The result of the operation stored in the OCM is further stored in the DDR memory. The system according to item 20. [Item 22] If the value of the result exceeds the maximum value, the value of the result is adjusted to the maximum value of the range for the integer data in the first format. If the value of the result is lower than the minimum value of the range for the data, before storing the result in the DDR memory, the value of the result is adjusted to the minimum value of the range for the integer data in the first format. The system according to item 21. [Item 23] Before storing the result in the DDR memory, the most significant bit of the result is dropped. The system according to item 21 or 22. [Item 24] The integer data in the first format is an integer representation of floating-point data, and the floating-point data is scaled and quantized to form the integer data in the first format. The system according to any one of items 15 to 23. [Item 25] A first scaling value is used to convert the floating-point data to the int8 format, and a second scaling value is used to convert the floating-point data to the uint8 format. The system according to item 24.
Claims
1. A method for converting data stored in a storage unit from a first format to a second format, the method comprising: extending the data stored in the storage unit by one or more bits to form extended data; determining whether the data stored in the storage unit is signed or unsigned; adding a sign value to the signed bits of the one or more extended bits of the data in response to determining that the data is signed; setting the signed bits of the one or more extended bits of the data to an unsigned value in response to determining that the data is unsigned; storing the extended data in on-chip memory (OCM) of a processing tile of a computing unit; A method comprising the above steps.
2. The method according to claim 1, wherein the data is an unsigned integer.
3. The method according to claim 1, wherein the data is a signed integer.
4. The method according to any one of claims 1 to 3, wherein the data is 8 bits and the extended data is 9 bits.
5. The method according to any one of claims 1 to 4, wherein the extended data is int9 data.
6. tracking whether the data stored in the storage unit is signed or unsigned; scheduling appropriate instructions for the extended data based on whether the data is signed or unsigned; The method according to any one of claims 1 to 5, further comprising the above steps.
7. The method according to claim 6, further comprising performing an arithmetic logic unit (ALU) operation on the extended data as an operand.
8. The method according to claim 7, further comprising storing the result of the operation in the OCM of the processing tile of the computing unit.
9. The method according to claim 8, further comprising storing the result stored in the OCM in the storage unit.
10. Before the step of storing the result in the storage unit, if the value of the result exceeds the maximum value of the range for the data, adjusting the value of the result to the maximum value, and if the value of the result is lower than the minimum value of the range for the data, adjusting the value of the result to the minimum value of the range for the data; the method according to claim 9, further comprising this step.
11. Before the step of storing the result from the OCM in the storage unit, further comprising the step of dropping the extended one or more bits of the result; the method according to claim 9.
12. The data stored in the storage unit is an integer representation of the data regarding floating-point data; the method according to any one of claims 1 to 11.
13. The floating-point data is scaled and quantized to form the data in the first format; the method according to claim 12.
14. A first scaling value is used to convert the floating-point data to the int8 format, and a second scaling value is used to convert the floating-point data to the uint8 format; the method according to claim 13.
15. A storage unit configured to store integer data in a first format, A processing unit having a plurality of processing tiles Comprising, each processing tile For various operations, having an on-chip memory (OCM) configured to receive and maintain extended data converted from the integer data in the first format from the storage unit, the extended data including at least one or more additional bits compared to the integer data in the first format, and when the integer data in the first format is signed, the signed bit of the extended data is signed, and when the integer data in the first format is unsigned, the signed bit of the extended data is set to an unsigned value; a system.
16. The integer data in the first format is either int8 or uint8; the system according to claim 15.
17. The extended data is int9; the system according to claim 15 or 16.
18. Whether the integer data in the first format stored in the storage unit is signed or unsigned is tracked, and appropriate instructions are scheduled according to whether the integer data in the first format is signed or unsigned. The system according to any one of claims 15 to 17.
19. The extended data is an operand for an operation. The system according to any one of claims 15 to 18.
20. The result of the operation is stored in the OCM. The system according to claim 19.
21. The result of the operation stored in the OCM is further stored in the storage unit. The system according to claim 20.
22. If the value of the result exceeds the maximum value, the value of the result is adjusted to the maximum value of the range for the integer data in the first format. If the value of the result is lower than the minimum value of the range for the data, before storing the result in the storage unit, the value of the result is adjusted to the minimum value of the range for the integer data in the first format. The system according to claim 21.
23. Before storing the result in the storage unit, the signed bit of the result is dropped. The system according to claim 21 or 22.
24. The integer data in the first format is an integer representation of floating-point data, and the floating-point data is scaled and quantized to form the integer data in the first format. The system according to any one of claims 15 to 23.
25. A first scaling value is used to convert the floating-point data to the int8 format, and a second scaling value is used to convert the floating-point data to the uint8 format. The system according to claim 24.
26. Extending the data stored in the storage unit by one or more bits to form extended data; Determining whether the data is signed or unsigned data; When the data is signed, adding a sign value to the signed bit of the extended one or more bits of the data; When the data is unsigned, setting the signed bits of the one or more extended bits of the data to an unsigned value; storing the extended data; A method comprising: **Claim 27** The method according to claim 26, wherein the data is an unsigned integer. **Claim 28** The method according to claim 26, wherein the data is a signed integer. **Claim 29** The method according to any one of claims 26 to 28, wherein the data is 8 bits and the extended data is 9 bits. **Claim 30** The method according to any one of claims 26 to 29, wherein the extended data is int9 data. **Claim 31** tracking whether the data stored in the storage unit is signed or unsigned; scheduling appropriate instructions for the extended data based on whether the data is signed or unsigned; The method according to any one of claims 26 to 30, further comprising: **Claim 32** The method according to claim 31, further comprising, as an operand, executing an arithmetic logic unit (ALU) operation on the extended data. **Claim 33** The method according to claim 32, further comprising storing the result of the operation in an on-chip memory (OCM) of a processing tile of a computing unit. **Claim 34** The method according to claim 33, further comprising storing the result stored in the OCM in the storage unit. **Claim 35** Before storing the result in the storage unit, if the value of the result exceeds the maximum value of the range for the data, adjusting the value of the result to the maximum value, and if the value of the result is lower than the minimum value of the range for the data, adjusting the value of the result to the minimum value of the range for the data. The method according to claim 34, further comprising: **Claim 36** Before storing the result from the OCM to the storage unit, further comprising dropping the extended one or more bits of the result. The method according to claim 34 or 35. **Claim 37** The method according to any one of claims 26 to 36, wherein the data stored in the storage unit is an integer representation of the data related to floating point data. **Claim 38** The method according to claim 37, wherein the floating-point data is scaled and quantized to form the data in a first format. **Claim 39** The method according to claim 38, wherein a first scaling value is used to convert the floating-point data into an int8 format, and a second scaling value is used to convert the floating-point data into a uint8 format.
Citation Information
Patent Citations
Information processor
JP1999328001A
Architecture for dense operations in machine learning inference engine
US20190244130A1
Information processing device, method, and program
WO2020049681A1