Unified programming interface for re-grained tile execution
An integrated programming interface dynamically adapts tensor operations to runtime conditions, optimizing execution across heterogeneous hardware, thus reducing deployment time and costs for AI applications.
Patent Information
- Application Number
- JP2025067132
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI application deployment across different ISAs and processors is time-consuming, costly, and inefficient due to the need for code customization and lack of dynamic adaptation to varying hardware resources.
An integrated programming interface that determines tensor operation input sizes at runtime, selects a split configuration based on input tensor size and runtime conditions, and uses a lookup table for optimal execution planning, allowing deployment across heterogeneous hardware resources.
Enhances performance by reducing execution time, saving development costs, and improving efficiency, enabling easy deployment across different ISAs and processors without statically tied optimizations.
Smart Images

Figure 2025108592000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments generally relate to an application programming interface (API). More specifically, embodiments relate to an integrated programming interface for regrained tile execution.
Background Art
[0002] An instruction set architecture (ISA) generally defines the types of data, registers, and hardware support that are supported for the operation of a processor, such as data processing, memory operations, arithmetic operations, control flow operations, etc. Recent developments in artificial intelligence (AI) may have led to the result that an extension of the ISA more explicitly supports the training and inference operations of neural networks. Therefore, software developers can customize the code of AI applications to utilize the new computing capabilities and accelerated execution facilitated by the extended ISA. However, code customization can be time-consuming, costly, and inefficient, especially when applications are deployed across different ISAs and processors.
Brief Description of the Drawings
[0003] Various advantages of the embodiments will become apparent to those skilled in the art by reading the following specification and the appended claims, and by referring to the following drawings.
[0004]
Figure 1
[0005]
Figure 2
[0006]
Figure 3
[0007]
Figure 4
[0008]
Figure 5
[0009]
Figure 6
[0010]
Figure 7
[0011]
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0012] Referring to FIG. 1, an arrangement scenario 10 is shown. Here, an application developer 12 generates an application 14 for placement in an execution environment 16 having a computing system 20 (e.g., a backend platform including one or more processor cores, not shown). The application 14 may include training of a neural network (e.g., convolutional neural network / CNN, deep neural network / DNN, etc.) (e.g., iterative selection of weights of network layers) and / or real-time operation of a neural network (e.g., for drawing inferences regarding image recognition, natural language processing / NLP, etc.). In an embodiment, the application 14 is a portable application designed to read and write configuration settings of the application 14 to a folder accessible to the computing system 20.
[0013] In the example shown, application 14 includes one or more common tensor operations 18 (18a-18b, for example, matrix multiplication operations, convolution operations, normalization operations, normalized linear unit / relu operations, exponential linear unit / elu operations, and / or other complex instruction set computer / CISC operations). Generally, a tensor can be a multi-dimensional data array that facilitates the automatic classification of input data by a neural network. The multi-dimensional nature of tensors generally requires the use of matrix-based mathematical operations. Here, the size of the matrix (e.g., the length of the columns and / or rows) varies. It may not be possible to specify the input tensor size when application 14 is created by application developer 12.
[0014] Accordingly, the tensor operations 18 shown have unspecified tensor input sizes when application 14 is created by application developer 12. Rather, computing system 20 can determine the tensor input size 22 (e.g., the length of the input columns and / or rows) of tensor operations 18 at runtime (e.g., during neural network training and / or inference). In an embodiment, computing system 20 also determines one or more runtime conditions 24 (e.g., expected power consumption, matrix sparsity, hardware resource availability, etc.) and selects a split configuration 26 of tensor operations 18 based on the tensor input size 22.
[0015] In one example, the partitioning configuration 26 defines a first set of matrix shapes for the first tensor operation 18a (e.g., a combination of the number of columns and width of a “tile”) and a second set of matrix shapes for the second tensor operation 18b. The partitioning configuration 26 may also define a first set of hardware resources for the first tensor operation 18a (e.g., a pool of computing cores) and a second set of hardware resources for the second tensor operation 18b. Here, the first set of hardware resources and the second set of hardware resources are different types of hardware resources. For example, the first set of hardware resources may be a relatively lightweight (e.g., “light”) pool of computing cores that includes scalar cores. On the other hand, the second set of hardware resources may be a relatively heavy pool of computing cores. The computing system 20 may use the partitioning configuration 26 to generate an output 28 (e.g., optimized code for performing training or inference based on the runtime tensor size and available computing resources) from the application 14.
[0016] Accordingly, the solution shown is less time-consuming, less expensive, and more efficient from the perspective of the application developer 12. In fact, since the optimizations and transformations are not statically tied to specific tensor sizes or tensor cores by the application developer 12, the same application 14 can be much more easily deployed across different ISAs and processors. Further, performance is improved by considering the runtime conditions 24 when the partitioning configuration 26 is generated. For example, leveraging knowledge about the sparsity of a matrix (e.g., the distribution of zero values in the matrix) may enable selecting a relatively lightweight pool of computing cores and / or a different floating-point format for the operations.
[0017] Figure 2 shows a method 30 for operating a performance-enhanced computing system. The method 30 may be implemented in one or more modules as a set of logic instructions. The logic instructions may be stored in a machine-readable or computer-readable storage medium such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc., in configurable logic such as a programmable logic array (PLA), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), etc., or in function-fixed logic hardware using circuit technologies such as application specific integrated circuits (ASICs), complementary metal oxide semiconductor (CMOS) technology or transistor-transistor logic (TTL) technology, or any combination thereof.
[0018] For example, the computer program code for performing the operations shown in method 30 may be written in any combination of one or more programming languages. The one or more programming languages may include object-oriented programming languages such as Java (registered trademark), Smalltalk (registered trademark), or C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. Further, the logic instructions may include assembly instructions, ISA instructions, machine instructions, machine-dependent instructions, microcode, state-setting data, configuration data for integrated circuits, state information for personalizing electronic circuits, and / or other structural components specific to the hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0019] The processing block 32 shown detects tensor operations in the application, where the tensor operations have unspecified input tensor sizes. In an embodiment, tensor operations are common and include matrix multiplication operations (e.g., matmul), convolution operations (e.g., conv2d, conv2d_transpose, conv3d), normalization operations (e.g., l2_normalize), normalized linear unit operations, exponential linear unit operations, etc., or any combination thereof. The tensor operations can be detected by analyzing and / or compiling the application 14 for execution. At block 34, the runtime input tensor sizes are determined. In an embodiment, the input tensor sizes are determined by analyzing input data (e.g., input images, utterances, etc.) to the neural network and analyzing data output from previous layers in the neural network.
[0020] Based at least in part on the input tensor sizes and one or more runtime conditions, a split configuration is selected for the tensor operation in block 36. In one example, the runtime conditions include the expected power consumption, matrix sparsity, and / or hardware resource availability. Further, the split configuration can define a first set of matrix shapes for a first operation (e.g., tile size), a second set of matrix shapes for a second operation, etc. In an embodiment, the split configuration further defines a first set of hardware resources for a first operation, a second set of hardware resources for a second operation, etc., where the first set of hardware resources and the second set of hardware resources are different types of resources. Such an approach to tile size and resource selection allows the matrix computation granularity to be on-the-fly and changed (e.g., re-granulated) in real time.
[0021] Block 36 may include looking up in a lookup table at least one of an input tensor size and runtime conditions. In this regard, since each tensor operation is generally well understood in terms of computational, memory, and communication patterns, an offline tuning / benchmarking process may be able to capture a mapping that is closest to optimal for different tensor sizes and sets of available resources. Thus, at runtime, based on the detected tensor size and the enumeration of available tensor cores, the table lookup is performed by the runtime engine and, along with the generation of optimized code within any optimized tensor core (e.g., optimal tile size, etc.), may obtain the most optimal / closest-to-optimal execution / splitting plan. When the tuning process is done offline, the runtime scheduling overhead can be minimized.
[0022] Accordingly, the method 30 shown provides a solution for re-granulated tile execution that improves performance by considering runtime conditions when generating a split configuration. For example, leveraging knowledge about the expected power consumption may make it possible to map tensor operations to a more power-efficient core pool. The method 30 shown is also less time-consuming, less expensive, and more efficient from the perspective of an application developer. In fact, the same application may be much easier to deploy across different ISAs and processors since optimizations tied to specific tensor input sizes are not incorporated into the application by the application developer.
[0023] Figure 3 shows a split configuration 40 in which an integrated dynamic dispatcher 42 (e.g., a "granulation device") receives a computational granular portable application 44. The dynamic dispatcher 42 can be considered "integrated" to the extent that, for example, the dispatcher 42 uses an integrated programming model such as ONEAPI to configure the application 44 to execute across a set of heterogeneous hardware resources (e.g., CPU, graphics processing unit / GPU, FPGA, dedicated accelerators, etc.). In an embodiment, the dispatcher 42 includes a pre-compiled plan lookup table 46 that contains benchmark data that facilitates the selection of the split configuration of the application 44 at runtime. Additionally, a set of pre-compiled granular optimization libraries ("libs") 48 can include, for example, a vector performance library, a 16×16 (e.g., 16 elements × 16 elements) performance library, a 32×32 (e.g., 32 elements × 32 elements) performance library, etc.
[0024] In the example shown, the split configuration defines / specifies the use of a scalar core 50 (e.g., selected from an integrated pool of lightweight computational cores), a 16-wide vector computational lane 52 (e.g., selected from an integrated pool of heavy computational vector cores), a 16×16 tensor core 54 (e.g., selected from an integrated media computational 2-dimensional / 2D tensor core pool), a 32×32 tensor core 56 (e.g., selected from an integrated pool of heavy computational GPU / 2D tensor cores), etc. In an embodiment, the dispatcher 42 generates customized modules 58 (e.g., a vector module, a 16×16 module, a 32×32 module) at runtime for optimal execution of the application 44 in heterogeneous tensor cores, potentially for collaborative execution.
[0025] Figure 4 shows a split configuration for a matrix multiplication (matmul) operation between an activation matrix 60 (e.g., a matrix X representing the activation of a neural network layer) and a weight matrix 62 (e.g., a matrix W representing the weights applied to the activation). Here, Y = X.W. In particular, during training with a larger batch size, loading the activation matrix 60 can put pressure on the core memory bandwidth. This may also be the case for the weight matrix 62. Based on the generated offline tuning plan, the runtime engine may decide to split the X.W operation, where X is a 52×32 element matrix and W is a 32×64 element matrix, into three parts.
[0026] Rows 0 to 31 of matrix X, together with matrix W, are read by the "heavy-loading" 32×32 compute element tensor core. The tile matmul operation is performed twice to generate rows 0 to 31 of the output matrix 64.
[0027] Rows 32 to 47 of matrix X and columns 0 to 31, as well as the upper half of matrix W, are read by a 16×16 tensor core to generate the partial sum of the upper half of the output matrix 64.
[0028] At the same time, rows 32 to 47 of matrix X and columns 32 to 63, as well as the lower half of matrix W, are read by another 16×16 tensor core to generate the partial sum of the upper half of the output matrix 64.
[0029] During the reduction stage between the two cores, pairs of partial sums are added to generate the final result for rows 32 to 47 of the output matrix 64.
[0030] Four 16-wide vector units read the last four rows of matrix X, and each vector unit processes one quarter of the columns of matrix W. Each vector unit generates one quarter of the columns of the output matrix 64 and outputs them to rows 48 to 51.
[0031] Referring to FIG. 5, a performance-enhanced computing system 151 is shown. System 151 may generally be part of an electronic device / platform having computing capabilities (e.g., personal digital assistant (PDA (registered trademark)), notebook computer, tablet computer, convertible tablet, server), communication capabilities (e.g., smartphone), imaging capabilities (e.g., camera, camcorder), media playback capabilities (e.g., smart TV), wearable capabilities (e.g., wristwatch, glasses, hat, shoes, jewelry), vehicle capabilities (e.g., car, truck, bike), robotic capabilities (e.g., autonomous robot), etc., or any combination thereof. In the example shown, system 151 includes a host processor 153 (e.g., a CPU having multiple cores, not shown) having an integrated memory controller (IMC) 155 coupled to a system memory 157.
[0032] The system 151 shown further includes an input / output (IO) module 159 and a graphics processor 161 implemented on a semiconductor die 163 as a system-on-chip (SoC) together with the host processor 153. The IO module 159 shown communicates with, for example, a display 165 (e.g., touch screen, liquid crystal display / LCD, light emitting diode / LED display), a network controller 167 (e.g., wired and / or wireless), and a mass storage 169 (e.g., hard disk drive / HDD, optical disk, solid state drive / SSD, flash memory).
[0033] In an embodiment, host processor 153, graphics processor 161, and / or IO module 159 execute program instructions 171 obtained from system memory 157 and / or mass storage 169 to perform one or more aspects of method 30 (Figure 2) already described. Thus, execution of the instructions 171 shown can cause computing system 151 to detect tensor operations in an application. Here, the tensor operations have an unspecified tensor input size, determine the input tensor size at runtime, and select a split configuration of the tensor operations based at least in part on the input tensor size and one or more runtime conditions. In an embodiment, the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation. Also, the split configuration can define a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, etc., where the first set of hardware resources and the second set of hardware resources are different types of resources. In one example, to select the split configuration, the instructions 171, when executed, cause computing system 151 to look up in a lookup table the input tensor size and at least one of the runtime conditions.
[0034] Accordingly, the system 151 shown is considered to be performance enhanced to the extent of providing a solution for re-granulated tile execution that takes into account runtime conditions at least when generating the split configuration. For example, using knowledge about the expected availability of hardware resources can make it possible to reduce the execution time. Also, the system 151 shown saves application development time, reduces costs, and improves efficiency. In fact, the same application can be much more easily deployed across different ISAs and processors because optimizations tied to a particular tensor input size or a particular tensor core are not incorporated into the application by the application developer.
[0035] Figure 6 shows semiconductor package device 173. The device 173 shown includes one or more substrates 175 (e.g., silicon, sapphire, gallium arsenide) and logic 177 (e.g., transistor arrays and other integrated circuit (IC) components) coupled to the one or more substrates 175. The logic 177 may be implemented at least partially within configurable logic or function-fixed logic hardware. In one example, the logic 177 implements one or more aspects of the method 30 (Figure 2) already described. Thus, the logic 177 may detect tensor operations in an application. Here, a tensor operation has an unspecified tensor input size, determines the input tensor size at runtime, and selects a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions. In one example, to select a split configuration, the logic 177 looks up at least one of the input tensor size and the runtime conditions in a look-up table.
[0036] Accordingly, the device 173 shown is considered to be performance-enhanced to the extent that it provides a means for resolving re-granulated tile execution that takes into account runtime conditions, at least when generating a split configuration. For example, utilizing knowledge about the expected availability of hardware resources may make it possible to reduce the execution time. Also, the device 173 shown saves the application development time, reduces costs, and improves efficiency. In fact, the same application may be much more easily placed across different ISAs and processors since the tensor input size is not incorporated into the application by the application developer.
[0037] In one example, logic 177 includes transistor channel regions disposed (e.g., embedded) within one or more substrates 175. Thus, the interface between logic 177 and one or more substrates 175 may not become a step junction. Logic 177 may further be considered to include an epitaxial layer that grows on the initial wafer of one or more substrates 175.
[0038] FIG. 7 shows a processor core 200 according to one embodiment. The processor core 200 may be a core of any type of processor, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, or other device that executes code. Only one processor core 200 is shown in FIG. 7, but the processing element may alternatively include more than one of the processor cores 200 shown in FIG. 7. The processor core 200 may be a single-threaded core, or, for at least one embodiment, the processor core 200 may be multi-threaded in that each core may include more than one hardware thread context (or may include a “logical processor”).
[0039] FIG. 7 also shows a memory 270 coupled to the processor core 200. The memory 270 may be any of a variety of memories known to those skilled in the art or otherwise available to those skilled in the art (including various layers of the memory hierarchy). The memory 270 may contain one or more code 213 instructions executed by the processor core 200. Here, the code 213 may implement one or more aspects of the method 30 (FIG. 2) already described. The processor core 200 follows the program sequence of instructions indicated by the code 213. Each instruction may enter the front-end portion 210 and be processed by one or more decoders 220. The decoder 220 may generate micro-operations, such as fixed-width micro-operations, in a predefined format as its output, or may generate control signals that reflect other instructions, micro-instructions, or the original code instructions. The front-end portion 210 shown also includes register naming logic 225 and scheduling logic 230. They generally allocate resources and queue the operations corresponding to the transformed instructions for execution.
[0040] The processor core 200 is shown to include execution logic 250 having a set 255-1 to 255-N of execution units. Some embodiments may include multiple execution units dedicated to a particular function or set of functions. Other embodiments may include only one execution unit, or may include one execution unit capable of performing a particular function. The execution logic 250 shown executes the operations specified by the code instructions.
[0041] After the execution of the operation specified by the code instruction, the backend logic 260 retires the instruction of the code 213. In one embodiment, the processor core 200 permits out-of-order execution but requires in-order retirement of instructions. The retirement logic 265 can take various forms known to those skilled in the art (e.g., reorder buffer, etc.). Thus, the processor core 200 is transformed during the execution of the code 213 with respect to at least the output generated by the decoder, the hardware registers and tables utilized by the register renaming logic 225, and any registers (not shown) modified by the execution logic 250.
[0042] Although not shown in FIG. 7, the processing element may include other elements on the chip having the processor core 200. For example, the processing element may include memory control logic together with the processor core 200. The processing element may include I / O control logic and / or may include I / O control logic integrated with the memory control logic. The processing element may also include one or more caches.
[0043] Referring now to FIG. 8, a block diagram of an embodiment of a computing system 1000 according to one embodiment is shown. Shown in FIG. 8 is a multiprocessor system 1000 including a first processing element 1070 and a second processing element 1080. Although two processing elements 1070 and 1080 are shown, it should be understood that one embodiment of the system 1000 may also include only one such processing element.
[0044] The system 1000 is shown as a point-to-point interconnect system. Here, the first processing element 1070 and the second processing element 1080 are coupled via a point-to-point interconnect 1050. It should be understood that any or all of the interconnects shown in FIG. 8 may be implemented as a multi-drop bus instead of a point-to-point interconnect.
[0045] As shown in FIG. 8, each of processing elements 1070 and 1080 may be a multi-core processor including a first processor core and a second processor core (i.e., processor cores 1074a and 1074b, and processor cores 1084a and 1084b). Such cores 1074a, 1074b, 1084a, 1084b may be configured to execute instruction code in a manner similar to that described above in connection with FIG. 7.
[0046] Processing elements 1070, 1080 may each include at least one shared cache 1896a, 1896b. Shared caches 1896a and 1896b may each store data (e.g., instructions) utilized by one or more components of the processor, such as cores 1074a, 1074b and 1084a, 1084b. For example, shared caches 1896a, 1896b may locally cache data stored in memories 1032, 1034 so that components of the processor can access it more quickly. In one or more embodiments, shared caches 1896a, 1896b may include one or more intermediate-level caches, such as level 2 (L2), level 3 (L3), level 4 (L4), or other-level caches, last-level cache (LLC), and / or combinations thereof.
[0047] Although only two processing elements 1070, 1080 are shown, it should be understood that the scope of the embodiments is not so limited. In other embodiments, one or more additional processing elements may be present in a given processor. Alternatively, one or more of the processing elements 1070, 1080 may be elements other than a processor, such as an accelerator or a field programmable gate array. For example, one or more additional processing elements may include one or more additional processors that are the same as the first processor 1070, one or more additional processors that are heterogeneous or asymmetric with respect to the first processor 1070, an accelerator (e.g., a graphics accelerator or a digital signal processing (DSP) unit, etc.), a field programmable gate array, or any other processing element. There may be various differences between the processing element 1070 and the processing element 1080 with respect to a wide range of value criteria, including architectural characteristics, microarchitectural characteristics, thermal characteristics, power consumption characteristics, etc. These differences may effectively manifest as asymmetry and heterogeneity between the processing elements 1070, 1080. In at least one embodiment, the various processing elements 1070, 1080 may be present in the same die package.
[0048] The first processing element 1070 may further include a memory controller logic (MC) 1072, and point-to-point (P-P) interfaces 1076 and 1078. Similarly, the second processing element 1080 may include an MC 1082 and P-P interfaces 1086 and 1088. As shown in FIG. 8, the MCs 1072 and 1082 couple the processors to their respective memories, namely, memories 1032 and 1034. Those memories may be part of the main memory locally attached to their respective processors. Although the MCs 1072 and 1082 are shown as being integrated within the processing elements 1070 and 1080, in an alternative embodiment, the MC logic may be separate logic external to the processing elements 1070, 1080 rather than being integrated within the processing elements 1070, 1080.
[0049] The first processing element 1070 and the second processing element 1080 may each be coupled to the I / O subsystem 1090 via P-P interconnects 1076, 1086. As shown in FIG. 8, the I / O subsystem 1090 has P-P interfaces 1094 and 1098. Further, the I / O subsystem 1090 has an interface 1092 that couples the I / O subsystem 1090 to the high-performance graphics engine 1038. In one embodiment, a bus 1049 may be used to couple the graphics engine 1038 to the I / O subsystem 1090. Alternatively, a point-to-point interconnect may couple these components.
[0050] Similarly, the I / O subsystem 1090 may be coupled to the first bus 1016 via an interface 1096. In one embodiment, the first bus 1016 may be a bus such as a Peripheral Component Interconnect (PCI) bus, or a PCI Express bus or another third-generation I / O interconnect bus, although the scope of this embodiment is not so limited.
[0051] As shown in FIG. 8, various I / O devices 1014 (e.g., a bioscaner, a speaker, a camera, a sensor) may be coupled to the first bus 1016 together with a bus bridge 1018. The bus bridge 1018 may couple the first bus 1016 to a second bus 1020. In one embodiment, the second bus 1020 may be a Low Pin Count (LPC) bus. In one embodiment, various devices may be coupled to the second bus 1020, including, for example, a keyboard / mouse 1012, one or more communication devices 1026, and a data storage unit 1019 such as a disk drive or other mass storage device that may include code 1030. The code 1030 shown may implement one or more aspects of the method 30 (FIG. 2) already described. Further, audio I / O 1024 may be coupled to the second bus 1020, and a battery 1010 may supply power to the computing system 1000.
[0052] Note that other embodiments are contemplated. For example, instead of the point-to-point architecture of FIG. 8, the system may implement a multi-drop bus, or another such communication topology. Also, the elements of FIG. 8 may alternatively be split using more or fewer integrated chips than shown in FIG. 8.
[0053] Additional Considerations and Examples
[0054] Example 1 includes a performance-enhanced computing system that includes a network controller, a processor coupled to the network controller, and a memory coupled to the processor, where the memory includes a set of executable program instructions that, when executed by the processor, cause the computing system to detect tensor operations in an application, where the tensor application has an unspecified input tensor size, determines the input tensor size at runtime, and selects a split configuration for the tensor operations based at least in part on the input tensor size and one or more runtime conditions.
[0055] Example 2 includes the computing system of Example 1, where the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
[0056] Example 3 includes the computing system of Example 2, where the split configuration further defines a first set of hardware resources for a first operation and a second set of hardware resources for a second operation, where the first set of hardware resources and the second set of hardware resources are different types of resources.
[0057] Example 4 includes the computing system of Example 1, where the one or more runtime conditions include one or more of expected power consumption, matrix sparsity, or hardware resource availability.
[0058] Example 5 includes the computing system of Example 1, where, to select a partitioning configuration, the instructions, when executed, cause the computing system to look up in a look-up table an input tensor size and at least one of one or more runtime conditions.
[0059] Example 6 includes the computing system of any one of Examples 1 to 5, where the tensor operations include one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation.
[0060] Example 7 includes a semiconductor device including one or more substrates, and logic coupled to the one or more substrates, where the logic is at least partially implemented by one or more configurable logics or fixed-function hardware logics, the logic is coupled to the one or more substrates to detect tensor operations in an application, where the tensor operations have unspecified input tensor sizes, determine input tensor sizes at runtime, and select a partitioning configuration of the tensor operations based at least in part on the input tensor sizes and one or more runtime conditions.
[0061] Example 8 includes the semiconductor device of Example 7, where the partitioning configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
[0062] Example 9 includes the semiconductor device of Example 8, where the partitioning configuration further defines a first set of hardware resources for a first operation and a second set of hardware resources for a second operation, where the first set of hardware resources and the second set of hardware resources are different types of resources.
[0063] Example 10 includes the semiconductor device of Example 7, where the one or more runtime conditions include one or more of an expected power consumption, a sparsity of a matrix, or an availability of hardware resources.
[0064] Example 11 includes the semiconductor device of Example 7, where, to select a split configuration, logic coupled to one or more substrates looks up in a look-up table an input tensor size and at least one of one or more runtime conditions.
[0065] Example 12 includes the semiconductor device of any one of Examples 7 to 11, where the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation.
[0066] Example 13 includes at least one computer-readable storage medium including a set of executable program instructions that, when executed by a computing system, cause the computing system to detect a tensor operation in an application, where the tensor operation has an unspecified input tensor size, determine an input tensor size at runtime, and select a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions.
[0067] Example 14 includes the at least one computer-readable storage medium of Example 13, where the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
[0068] Example 15 includes the at least one computer-readable storage medium of Example 14, where the split configuration further defines a first set of hardware resources for a first operation and a second set of hardware resources for a second operation, where the first set of hardware resources and the second set of hardware resources are different types of resources.
[0069] Example 16 includes the at least one computer-readable storage medium of Example 13, where the one or more runtime conditions include one or more of an expected power consumption, a sparsity of a matrix, or an availability of hardware resources.
[0070] Example 17 includes at least one computer-readable storage medium of Example 13, wherein, to select a split configuration, the instructions, when executed, cause a computing system to look up in a lookup table an input tensor size and at least one of one or more runtime conditions.
[0071] Example 18 includes at least one computer-readable storage medium of any one of Examples 13 to 17, wherein the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation.
[0072] Example 19 includes a method of operating a performance-enhanced computing system, the method comprising detecting a tensor operation in an application, wherein the tensor operation has an unspecified input tensor size, detecting, determining a runtime input tensor size, and selecting a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions.
[0073] Example 20 includes the method of Example 19, wherein the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
[0074] Example 21 includes the method of Example 20, wherein the split configuration further defines a first set of hardware resources for a first operation and a second set of hardware resources for a second operation, wherein the first set of hardware resources and the second set of hardware resources are different types of resources.
[0075] Example 22 includes the method of Example 19, wherein the one or more runtime conditions include one or more of an expected power consumption, matrix sparsity, or hardware resource availability.
[0076] Example 23 includes the method of Example 19, where selecting a partitioning configuration includes looking up in a lookup table the input tensor size and at least one of one or more runtime conditions.
[0077] Example 24 includes the method of any one of Examples 19 to 23, where the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation.
[0078] Example 25 includes means for performing the method of any one of Examples 19 to 24.
[0079] Accordingly, the techniques described herein provide an integrated programming interface targeting general tensor operations and a runtime procedure for re-granulating a computational architecture based on available computing resources. As a result, the application is properly mapped to available tensor cores and written only once.
[0080] Embodiments are applicable for use in all kinds of semiconductor integrated circuit ("IC") chips. Examples of these IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, system-on-chip (SoC), SSD / NAND controller ASICs, and the like. Additionally, in some of the drawings, signal lines are represented by lines. Some of these may be different to show more component signal paths, may have numerical labels to show multiple component signal paths, and / or may have arrows at one or more endpoints to show the main flow direction of information. However, this should not be construed in a limiting sense. Rather, such additional details may be used in connection with one or more exemplary embodiments to make the understanding of the circuit easier. Whether having additional information or not, any signal line shown may actually include one or more signals that may have multiple directions of movement, for example, may be implemented in any suitable type of signal scheme such as digital or analog lines implemented as differential pairs, optical fiber lines, and / or single-ended lines.
[0081] Exemplary sizes / models / values / ranges may be given, but embodiments are not limited thereto. As manufacturing technologies (e.g., photolithography) mature over time, it is expected that smaller devices can be manufactured. Additionally, well-known power / ground connections to IC chips and other components may or may not be shown in the figures in order to simplify the illustration and description and not obscure certain aspects of the embodiments. Further, to avoid obscuring the embodiments, the configurations may be shown in the form of block diagrams, and in view of the fact that the details regarding the implementation of such block diagram configurations highly depend on the computing system in which the embodiments are to be implemented, i.e., such details are within the purview of those of ordinary skill in the art. When specific details (e.g., circuits) are described to illustrate exemplary embodiments, it will be apparent to those of ordinary skill in the art that the embodiments can be practiced without these specific details or with variations of these specific details. Accordingly, the detailed description should be regarded as illustrative rather than limiting.
[0082] The term "coupled" may be used herein to refer to any kind of direct or indirect relationship between the components of interest and may apply to electrical, mechanical, fluidic, optical, electromagnetic, electromechanical, or other connections. Additionally, terms such as "first," "second," etc. may be used herein solely for ease of explanation and, unless otherwise specified, do not imply any particular temporal or chronological meaning.
[0083] The recitation of a list of items joined by the term "one or more of" in the present application and the claims may mean any combination of the recited terms. For example, the phrase "one or more of A, B, or C" may mean A; B; C; A and B; A and C; B and C; or A, B, and C.
[0084] Those skilled in the art will understand from the above description that the broad technology of the embodiments can be implemented in various forms. Therefore, although the embodiments have been described in relation to those specific examples, the true scope of the embodiments should not be so limited. This is because other modified forms will become apparent to those skilled in the art upon examining the drawings, the specification, and the following claims. [Other possible items] (Item 1) A network controller, A processor coupled to the network controller, A memory coupled to the processor Comprising a computing system, When the memory is executed by the processor, the computing system Includes a set of executable program instructions that cause the detection of tensor operations in the application, The tensor operations are Have an unspecified input tensor size, Determine the input tensor size at runtime, Select a split configuration of the tensor operations based at least in part on the input tensor size and one or more runtime conditions, Computing system. (Item 2) The split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation, the computing system according to item 1. (Item 3) The split configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, where the first set of hardware resources and the second set of hardware resources are different types of resources, the computing system according to item 2. (Item 4) The above-described one or more runtime conditions include one or more of the expected power consumption, matrix sparsity, or hardware resource availability of the computing system according to item 1. (Item 5) To select the above-described split configuration, when executed, the above-described instruction causes the computing system to search in a lookup table for the above-described input tensor size and at least one of the above-described one or more runtime conditions of the computing system according to item 1. (Item 6) The above-described tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation of the computing system according to item 1. (Item 7) One or more substrates, A semiconductor device comprising logic coupled to the above-described one or more substrates, The above-described logic is at least partially implemented by one or more configurable logics or fixed-function hardware logics, The above-described logic is coupled to the above-described one or more substrates, Detect a tensor operation in an application, the above-described tensor operation having an unspecified input tensor size, Determine the above-described input tensor size at runtime, Select a split configuration of the above-described tensor operation based at least in part on the above-described input tensor size and one or more runtime conditions. Semiconductor device. (Item 8) The above-described split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation of the semiconductor device according to item 7. (Item 9) The above-described split configuration further defines a first set of hardware resources for the above-described first operation and a second set of the above-described hardware resources for the above-described second operation, and the first set of the above-described hardware resources and the second set of hardware resources are different types of resources. The semiconductor device according to item 8. (Item 10) The semiconductor device according to item 7, wherein the one or more runtime conditions include one or more of the expected power consumption, the sparsity of the matrix, or the availability of hardware resources. (Item 11) The semiconductor device according to item 7, wherein, in order to select the split configuration, the logic coupled to the one or more substrates looks up at least one of the input tensor size and the one or more runtime conditions in a look-up table. (Item 12) The semiconductor device according to item 7, wherein the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation. (Item 13) At least one computer-readable storage medium comprising a set of executable program instructions that, when executed by a computing system, cause the computing system to detect a tensor operation in an application, wherein the tensor operation has an unspecified input tensor size, determine the input tensor size at runtime, select a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions and execute. At least one computer-readable storage medium. (Item 14) The split configuration according to item 13, defining a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation. (Item 15) The above-described partitioning configuration further defines a first set of hardware resources for the first operation and a second set of the hardware resources for the second operation, and the first set of the hardware resources and the second set of the hardware resources are different types of resources. At least one computer-readable storage medium according to item 14. (Item 16) The at least one computer-readable storage medium according to item 13, wherein the one or more runtime conditions include one or more of expected power consumption, sparsity of a matrix, or availability of hardware resources. (Item 17) The at least one computer-readable storage medium according to item 13, wherein, in order to select the above-described partitioning configuration, when the instructions are executed, the computing system is caused to search a lookup table for the input tensor size and at least one of the one or more runtime conditions. (Item 18) The at least one computer-readable storage medium according to item 13, wherein the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a rectified linear unit operation, or an exponential linear unit operation. (Item 19) Detecting a tensor operation in an application, the tensor operation having an unspecified input tensor size; Determining the input tensor size at runtime; And selecting a partitioning configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions. (Item 20) The partitioning configuration according to item 19, which defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation. (Item 21) The above-mentioned split configuration further defines a first set of hardware resources for the first operation and a second set of the hardware resources for the second operation, and the first set of the hardware resources and the second set of the hardware resources are different types of resources. The method according to item 20. (Item 22) The above-mentioned one or more runtime conditions include one or more of the expected power consumption, matrix sparsity, or hardware resource availability. The method according to item 19. (Item 23) The step of selecting the above-mentioned split configuration includes the step of searching in a lookup table for at least one of the above-mentioned input tensor size and the above-mentioned one or more runtime conditions. The method according to item 19. (Item 24) The above-mentioned tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation. The method according to item 19.
Claims
1. A computing system comprising: a network controller; a processor coupled to the network controller; and a memory coupled to the processor, wherein the memory, when executed by the processor, causes the computing system to: detect a tensor operation in an application, the tensor operation having an unspecified input tensor size; determine the input tensor size at runtime; and select a partitioning configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions. The computing system having a set of executable program instructions for causing the above to be executed.
2. The computing system of claim 1, wherein the partitioning configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
3. The computing system of claim 2, wherein the partitioning configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, and the first set of hardware resources and the second set of hardware resources are different types of resources.
4. The computing system of claim 1, wherein the one or more runtime conditions include one or more of an expected power consumption, matrix sparsity, or hardware resource availability.
5. The computing system of claim 1, wherein, to select the partitioning configuration, the executable program instructions, when executed, cause the computing system to search a lookup table for the input tensor size and at least one of the one or more runtime conditions.
6. The computing system according to any one of claims 1 to 5, wherein the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a rectified linear unit operation, or an exponential linear unit operation.
7. A semiconductor device comprising: one or more substrates; and logic coupled to the one or more substrates, wherein the logic is at least partially implemented by one or more configurable logics or fixed-function hardware logics, and the logic is coupled to the one or more substrates, Detect a tensor operation in an application, the tensor operation having an unspecified input tensor size, Determine the input tensor size at runtime, Select a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions, A semiconductor device.
8. The semiconductor device according to claim 7, wherein the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
9. The semiconductor device according to claim 8, wherein the split configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, and the first set of hardware resources and the second set of hardware resources are different types of resources.
10. The semiconductor device according to claim 7, wherein the one or more runtime conditions include one or more of expected power consumption, matrix sparsity, or hardware resource availability.
11. The semiconductor device according to claim 7, wherein, to select the split configuration, the logic coupled to the one or more substrates searches a look-up table for at least one of the input tensor size and the one or more runtime conditions.
12. The semiconductor device according to any one of claims 7 to 11, wherein the tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation.
13. In a computing system, Detecting a tensor operation in an application, the tensor operation having an unspecified input tensor size, Determining the input tensor size at runtime, Selecting a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions Causing to be executed, A program.
14. The program according to claim 13, wherein the split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation.
15. The split configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, and the first set of hardware resources and the second set of hardware resources are different types of resources. The program according to claim 14.
16. The one or more runtime conditions include one or more of expected power consumption, matrix sparsity, or hardware resource availability. The program according to claim 13.
17. To select the split configuration, the computing system is made to search a look-up table for the input tensor size and at least one of the one or more runtime conditions. The program according to claim 13.
18. The tensor operation includes one or more of matrix multiplication operation, convolution operation, normalization operation, normalized linear unit operation, or exponential linear unit operation. The program according to any one of claims 13 to 17.
19. A computer-readable storage medium storing the program according to any one of claims 13 to 18.
20. Detecting a tensor operation in an application, the tensor operation having an unspecified input tensor size; detecting; Determining the input tensor size at runtime; Selecting a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions. Method.
21. The split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation. The method according to claim 20.
22. The split configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, and the first set of hardware resources and the second set of hardware resources are different types of resources. The method according to claim 21.
23. The one or more runtime conditions include one or more of expected power consumption, matrix sparsity, or hardware resource availability. The method according to claim 20.
24. The step of selecting the split configuration includes the step of looking up in a lookup table the input tensor size and at least one of the one or more runtime conditions, the method according to claim 20.
25. The tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation, the method according to any one of claims 20 to 24.
26. A means for detecting a tensor operation in an application, the tensor operation having an unspecified input tensor size, and means for performing the detecting step, means for performing the step of determining the input tensor size at runtime, and means for performing the step of selecting a split configuration of the tensor operation based at least in part on the input tensor size and one or more runtime conditions. A semiconductor device.
27. The split configuration defines a first set of matrix shapes for a first operation and a second set of matrix shapes for a second operation, the semiconductor device according to claim 26.
28. The split configuration further defines a first set of hardware resources for the first operation and a second set of hardware resources for the second operation, and the first set of hardware resources and the second set of hardware resources are different types of resources, the semiconductor device according to claim 27.
29. The one or more runtime conditions include one or more of an expected power consumption, a sparsity of a matrix, or an availability of hardware resources, the semiconductor device according to claim 26.
30. The means for performing the step of selecting the split configuration includes means for performing a lookup in a lookup table of the input tensor size and at least one of the one or more runtime conditions, the semiconductor device according to claim 26.
31. The tensor operation includes one or more of a matrix multiplication operation, a convolution operation, a normalization operation, a normalized linear unit operation, or an exponential linear unit operation, the semiconductor device according to any one of claims 26 to 30.
Citation Information
Patent Citations
Systems and methods for implementing chained tile operations
JP2019197531A
Runtime of cublas matrix multiplication on GPU
US20170046307A1
Heterogeneous hardware accelerator architecture for processing sparse matrix data with skewed non-zero distributions
US20180189239A1
Accelerated mathematical engine
US20190026078A1
Methods and apparatus to improve utilization of a heterogeneous system executing software
US20190317741A1